<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Hung Pham]]></title><description><![CDATA[Hung Pham]]></description><link>https://hunpeolabs.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Hung Pham</title><link>https://hunpeolabs.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 30 Sep 2026 10:57:32 GMT</lastBuildDate><atom:link href="https://hunpeolabs.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[You Can Run a Company With AI Agents. But Who Is Running the Agents?]]></title><description><![CDATA[AI agents can research, code, test, and review at remarkable speed. The harder question is whether anyone still owns the outcome.
🌐 Languages: English · Tiếng Việt bên dưới
Scroll through your feed a]]></description><link>https://hunpeolabs.hashnode.dev/you-can-run-a-company-with-ai-agents-but-who-is-running-the-agents</link><guid isPermaLink="true">https://hunpeolabs.hashnode.dev/you-can-run-a-company-with-ai-agents-but-who-is-running-the-agents</guid><category><![CDATA[AI]]></category><category><![CDATA[aiagent]]></category><category><![CDATA[applied ai]]></category><category><![CDATA[startup]]></category><dc:creator><![CDATA[Hung Pham]]></dc:creator><pubDate>Sat, 22 Aug 2026 04:04:22 GMT</pubDate><content:encoded><![CDATA[<p><em>AI agents can research, code, test, and review at remarkable speed. The harder question is whether anyone still owns the outcome.</em></p>
<p>🌐 <strong>Languages:</strong> English · Tiếng Việt bên dưới</p>
<p>Scroll through your feed and you can feel that something has shifted. One founder has built a product over the weekend; another is running a small army of AI agents from a laptop. Research, design, code, testing, launch copy work that once required a team now seems to happen somewhere between Friday night and Monday morning. The one-person startup no longer sounds like a thought experiment. It sounds like something you could begin tonight with an idea, a laptop, and a few terminal windows.</p>
<p>The excitement makes sense. For years, ideas moved faster than the people available to build them. You could see the product clearly but had no designer. You found a designer but still needed an engineer. By the time the first version was ready, you were already looking for someone to test it, explain it, sell it, or support the people using it. There were always more ideas than hands. Then AI arrived. At first, it helped us think; soon it could write code, inspect systems, plan features, run tests, and review the work of other agents. After a long drought, it felt like rain. We were so relieved to have water that few of us stopped to ask where it would all go.</p>
<p>Imagine a founder building a booking app for small hair salons. The idea is straightforward: customers choose a service, pick a time, and make an appointment; the salon sees the booking and assigns a stylist. A few years ago, someone with this idea might have spent months finding the right people before writing a single line of code. Now the founder can divide the work among agents. One researches the market, another designs the interface, a third creates the database, while others handle authentication, scheduling, notifications, and testing. Within a week, the product is running.</p>
<p>Every morning, the founder opens his laptop and finds that it has grown overnight. One day there is a calendar for managing appointments; the next, reminder emails. Soon there is a revenue report, followed by suggestions for discount codes, a loyalty programme, and a dashboard filled with charts. A steady stream of “completed” statuses makes the product feel as though it is moving at extraordinary speed. Keep assigning tasks, and perhaps it will continue to grow, polish itself, and eventually find its own way into the market.</p>
<p>That feeling lasts until a real salon tries it. A customer books an appointment for three o’clock, then cancels because something comes up. The appointment disappears from the customer’s phone, but the three o’clock slot remains blocked in the salon’s system. Another customer tries to book the same time and cannot. It sounds like a small bug, so the founder asks an agent to investigate.</p>
<p>The first agent decides that the problem comes from state synchronisation. It changes the API, adjusts the frontend data, and adds a test. A review agent examines the work and concludes that the appointment state model is not robust enough, then proposes reorganising the booking flow. The database agent notices that the current schema may be difficult to extend and creates another table. Finally, the testing agent runs the suite and returns a reassuring wall of green. By the end of the afternoon, a minor cancellation bug has produced changes across more than forty files.</p>
<p>Every explanation sounds reasonable, and each agent has done its part according to the way it understood the task. Yet somewhere among the analyses, new code, schema changes, and passing tests, one simple question remains unanswered: after the first customer cancelled, could a second customer actually book the three o’clock slot?</p>
<p>This is the quiet gap between work being produced and a problem being solved. The frontend agent cares whether the interface updates. The backend agent looks at the API response. The database agent checks whether the records remain consistent. The testing agent knows only the cases that someone thought to encode as tests. Every part has someone working on it, yet the customer’s complete experience belongs to no one.</p>
<p>The same thing happens in an ordinary restaurant during the lunch rush. Customers are arriving faster than the tables can turn. Someone takes orders near the door, another calls them into the kitchen, cooks move between hot pans, and servers weave through the room carrying plates. Finished dishes cover the counter. Everyone looks busy, and the restaurant appears to be operating at full capacity. Yet table three has been waiting for almost half an hour, table seven receives the same dish twice, and a plate of fried rice sits cooling at the edge of the kitchen because nobody remembers who ordered it.</p>
<p>From inside the kitchen, judged by the number of dishes prepared, this looks like an exceptionally productive lunch service. From table three, where no food has arrived, it looks completely broken. Software built with AI can create the same illusion. We count tasks closed, lines of code generated, tests executed, and agents kept busy, while the user cares about something much simpler: did the meal they ordered reach the right table? When an agent reports that a task is complete, it may only mean the dish has left the pan. It does not mean the dish reached the customer, that it was the right dish, or that the customer could actually eat it.</p>
<p>We understand this distinction instinctively in everyday life. Nobody evaluates a restaurant by counting how many times the cooks moved their spatulas; we care whether customers received the right food, how long they waited, and how often a dish had to be sent back. We would not judge a plumber by the number of metres of pipe replaced either. We want to know whether the tap still leaks, whether anything else was damaged, and who will return if there is water on the floor again tomorrow.</p>
<p>Yet in the world of AI, volume is remarkably persuasive. Five thousand lines of code sound more impressive than five. Ten agents feel more powerful than one. A three-page report appears more trustworthy than a tiny change with almost nothing to explain. But if five lines fix the cause while five thousand leave a human reviewing code for two days and repairing side effects for another week, which result was more efficient? If ten agents produce ten interpretations of the same task, do we have a stronger team or simply ten new directions in which to get lost?</p>
<p>An honest measure of AI’s effectiveness should not begin with how much it produces. It should begin with what remains after the agent says it is done. How much cleanup is left for a human? How often does the work need to be redone? Does anyone understand why the system changed? When something breaks, do we know where to begin looking? Most importantly, did the completed work help the user do what they came to do, or did it merely make the repository busier?</p>
<p>This is the less glamorous side of the one-person startup. One founder may be able to operate many agents, but that does not make the different roles inside a company disappear. Consider the owner of a small restaurant. She may buy the ingredients, work the register, and step into the kitchen when the lunch rush begins. The restaurant has one owner, but the responsibilities remain distinct. At the register, she needs to know which tables have paid. In the kitchen, she needs to know which orders are waiting. Before a plate leaves the counter, she still checks that it is going to the right table. One person can perform many roles; that does not make those roles interchangeable.</p>
<p>A founder working with AI faces the same reality. Agents can research, code, test, and write content, but at some point the founder must step out of the role of the person trying to move faster and look at the product as a whole. AI can perform the work, but it cannot own the consequences. When the system fails in the middle of the night, the agent will not answer the customer’s call. When data is changed incorrectly, the agent will not have to explain what happened. When a technical decision costs the company weeks of recovery work, the agent will not live with the consequences. Approval may require only one click, but the responsibility behind that click still belongs to a person.</p>
<p>None of this means dragging AI into the same heavy processes, endless meetings, and paperwork that have slowed teams down for years. A restaurant does not need a twenty-page operating manual to deliver the right meal to the right table. It needs an order with a table number, the name of the dish, and someone who checks the plate before it leaves the kitchen. Calling a plumber does not require a project plan either. A clear conversation is enough: fix the leak; if you need to break through the wall or replace the entire pipe, call me first; turn the water on when you are finished; and if it leaks again tomorrow, come back and make it right.</p>
<p>These expectations feel ordinary because, in everyday life, we understand that even a skilled and fast worker must know what the job is, how far they are allowed to go, and what must be true before the work can honestly be called finished. It is only when we enter the world of AI carried away by its speed and apparent limitlessness that we forget how useful ordinary clarity can be.</p>
<p>Return to the booking bug. The founder does not need another elaborate process. He only needs to keep the task anchored to the original story: a customer cancelled, but the slot did not reopen. Reproduce that exact failure first. Change only what is necessary to release the time slot. If the fix requires altering the database or touching the payment flow, stop and ask before proceeding. When the change is ready, cancel the appointment with one account and use another account to book the same time. A second agent may review the work, but it should not use the review as an opportunity to redesign the entire system.</p>
<p>This time, the agent may change only a few files, and its report may be too short to show off. But another customer can now book the three o’clock appointment. That is what the user needed, and it is the moment when the work can truthfully be called complete. The result is no longer judged from the perspective of the person or agent who produced it. It is judged from the perspective of the person waiting to receive it.</p>
<p>This is where the image of water becomes clearer. AI resembles a mountain stream running down a steep slope: fast, forceful, and carrying more energy than we have ever had at our disposal. Every agent opens another current. Every prompt creates another tributary. Code, documentation, tests, analyses, and ideas keep flowing downhill until the repository slowly becomes a lake.</p>
<p>At first, watching the waterline rise feels like progress. Every day brings more features, more files, and more reports. Once enough water has gathered, however, it becomes difficult to tell which streams are clean and which are carrying mud, where the lake is safely deep and where the pressure is quietly weakening the bank. The problem is not that we have too much water. It is that we have not built channels to carry it to the right fields, installed gates that can close before it spills over, or placed anyone at the edge of the lake to watch the waterline and say, “This field has enough. Send the rest somewhere else.”</p>
<p>A farmer does not boast about how many cubic metres of water passed through the land that day. The farmer waits for the harvest. That may be the most important lesson in working with AI: do not mistake flow for results, activity for progress, or an agent’s declaration of “completed” for a user receiving what they actually needed.</p>
<p>The age of AI may indeed produce very small companies capable of work that once required large teams. That is worth being excited about. But the people who go furthest may not be those running the most agents, writing the longest prompts, or generating the most code. They will be the ones who know what is worth delegating, when to stop, which results can be trusted, and how to keep speed from becoming an illusion of productivity.</p>
<p>Fast-moving water has never created a harvest on its own. It matters only when it reaches the right field, at the right time, in the right amount.</p>
<h2>🇻🇳 Phiên bản tiếng Việt</h2>
<p>Dạo này, chỉ cần lướt một vòng mạng xã hội là có thể cảm nhận rất rõ một cơn sốt đang lan ra. Người này kể chuyện làm xong một sản phẩm chỉ trong cuối tuần, người kia khoe đội ngũ AI agent có thể nghiên cứu thị trường, thiết kế giao diện, viết code, chạy test và chuẩn bị nội dung ra mắt gần như không cần thêm ai. “Startup một thành viên” từ một ý tưởng nghe có phần viển vông bỗng trở thành thứ rất nhiều người tin rằng mình có thể bắt đầu ngay tối nay, chỉ với một chiếc laptop và vài cửa sổ terminal.</p>
<p>Sự hào hứng ấy hoàn toàn dễ hiểu. Đã có một thời, thứ giữ chân chúng ta không phải là thiếu ý tưởng mà là thiếu người. Có ý tưởng nhưng không biết thiết kế; có bản thiết kế lại không có người viết code; sản phẩm vừa thành hình thì tiếp tục thiếu người kiểm thử, viết nội dung, nghiên cứu thị trường hoặc nói chuyện với khách hàng. Ý tưởng lúc nào cũng nhiều hơn số đôi tay có thể biến chúng thành hiện thực. Rồi AI xuất hiện, giống như một trận mưa lớn đến sau quãng hạn kéo dài. Chúng ta vui vì cuối cùng cũng có nước, nên chẳng mấy ai nghĩ đến chuyện nếu mưa cứ tiếp tục thì nước sẽ chảy về đâu.</p>
<p>Hãy thử hình dung một người đang làm ứng dụng đặt lịch cho các tiệm tóc nhỏ. Ý tưởng ban đầu rất gọn: khách chọn dịch vụ, chọn giờ, đặt lịch; chủ tiệm nhìn thấy lịch hẹn và sắp xếp nhân viên. Trước đây, một người có ý tưởng như vậy có thể mất vài tháng chỉ để tìm đủ người bắt đầu. Bây giờ, anh ta giao việc cho từng agent. Một agent nghiên cứu thị trường, một agent dựng giao diện, một agent thiết kế database; những agent khác tiếp tục lo đăng nhập, lịch hẹn, thông báo và kiểm thử. Chưa đầy một tuần, sản phẩm đã có thể chạy.</p>
<p>Mỗi sáng mở máy là một lần anh ta thấy sản phẩm lớn thêm. Hôm nay có trang quản lý lịch, ngày mai có email nhắc hẹn, hôm sau nữa đã xuất hiện báo cáo doanh thu. AI còn rất nhiệt tình đề xuất thêm mã giảm giá, chương trình khách hàng thân thiết và một dashboard đầy biểu đồ. Những dòng trạng thái “completed” nối tiếp nhau tạo ra cảm giác mọi thứ đang tiến lên rất nhanh. Chỉ cần tiếp tục giao việc, dường như sản phẩm sẽ tự lớn, tự hoàn thiện rồi tự tìm được đường ra thị trường.</p>
<p>Cảm giác ấy kéo dài cho đến khi một tiệm tóc thật sự dùng thử. Một khách đặt lịch lúc ba giờ chiều, sau đó huỷ vì có việc bận. Trên điện thoại của khách, lịch hẹn đã biến mất; nhưng trong hệ thống của tiệm, khung giờ ba giờ vẫn bị giữ. Người khác muốn đặt đúng giờ đó thì không thể. Một lỗi rất nhỏ, ít nhất là khi nghe qua.</p>
<p>Founder giao cho agent kiểm tra. Agent đầu tiên cho rằng nguyên nhân nằm ở trạng thái đồng bộ nên sửa API, điều chỉnh dữ liệu phía giao diện và bổ sung một bài test. Agent review nhìn vào thay đổi rồi nhận xét mô hình trạng thái hiện tại chưa đủ tốt, từ đó đề xuất tổ chức lại luồng đặt lịch. Agent phụ trách database thấy cấu trúc cũ sẽ khó mở rộng nên tạo thêm một bảng mới. Agent kiểm thử chạy lại toàn bộ test và trả về một bản báo cáo xanh mướt. Chỉ trong một buổi chiều, một lỗi huỷ lịch đã tạo ra hơn bốn mươi thay đổi nằm rải rác khắp hệ thống.</p>
<p>Mọi lời giải thích đều có vẻ hợp lý. Agent nào cũng làm đúng phần việc theo cách nó hiểu. Thế nhưng giữa những bản phân tích, những đoạn code mới và những bài test vừa được bổ sung, vẫn còn một câu hỏi rất đơn giản chưa ai trả lời: sau khi khách thứ nhất huỷ lịch, khách thứ hai đã thật sự đặt được khung giờ ba giờ hay chưa?</p>
<p>Chính ở đó, sự khác biệt giữa “có nhiều việc được làm” và “có một vấn đề được giải quyết” bắt đầu lộ ra. Agent frontend quan tâm giao diện đã cập nhật chưa. Agent backend nhìn vào phản hồi của API. Agent database kiểm tra tính nhất quán của dữ liệu, còn agent kiểm thử chỉ biết những trường hợp đã được viết thành test. Mỗi phần đều có người chăm sóc, nhưng câu chuyện trọn vẹn của người khách cần đặt lịch lại không thực sự thuộc về ai.</p>
<p>Nó giống một quán cơm vào giờ trưa. Ngoài cửa, khách bắt đầu đông; người ghi món gọi liên tục vào bếp, đầu bếp không ngừng đảo chảo, nhân viên bưng bê chạy qua chạy lại. Đĩa thức ăn làm xong phủ kín mặt bàn, nhìn ai cũng bận và có vẻ quán đang hoạt động hết công suất. Thế nhưng bàn số ba đã chờ gần nửa tiếng vẫn chưa có món, bàn số bảy lại nhận hai phần giống nhau, còn một đĩa cơm rang nằm nguội ở góc bếp vì chẳng ai nhớ nó thuộc về bàn nào.</p>
<p>Nếu đứng trong bếp và đếm số đĩa đã nấu, đó là một buổi trưa vô cùng năng suất. Nếu đang ngồi ở bàn số ba, câu chuyện hoàn toàn khác. Phần mềm làm bằng AI cũng vậy: chúng ta đếm số task được đóng, số dòng code được tạo, số test đã chạy và số agent đang hoạt động, trong khi người dùng chỉ quan tâm món họ gọi có được mang đến đúng bàn hay không. Một agent báo “đã hoàn thành” nhiều khi mới chỉ có nghĩa món ăn đã rời khỏi chảo. Nó chưa chắc đã đến đúng người, càng chưa chắc người đó có thể dùng được.</p>
<p>Ngoài đời, chúng ta vốn hiểu chuyện này rất tự nhiên. Không ai đánh giá một quán ăn bằng số lần đầu bếp đảo chảo; điều đáng quan tâm là khách có nhận đúng món không, phải chờ bao lâu và món có bị trả lại không. Cũng chẳng ai đánh giá người thợ sửa nước bằng số mét ống anh ấy đã thay. Chủ nhà chỉ muốn biết chiếc vòi còn rò hay không, những thứ xung quanh có bị làm hỏng không và nếu ngày mai nước lại chảy ra sàn thì ai sẽ quay lại xử lý.</p>
<p>Vậy mà bước sang thế giới AI, chúng ta rất dễ bị khối lượng làm cho choáng ngợp. Năm nghìn dòng code nghe ấn tượng hơn năm dòng. Mười agent tạo cảm giác mạnh hơn một agent. Một bản báo cáo dài ba trang trông đáng tin hơn một thay đổi nhỏ đến mức gần như không có gì để kể. Nhưng nếu năm dòng code xử lý đúng nguyên nhân, còn năm nghìn dòng khiến con người mất hai ngày để đọc lại và thêm một tuần sửa lỗi, đâu mới là hiệu quả thật sự? Nếu mười agent tạo ra mười cách hiểu khác nhau về cùng một công việc, ta đang có một đội ngũ mạnh hơn hay chỉ có thêm mười hướng để đi lạc?</p>
<p>Nếu cần một KPI thực tế cho AI, có lẽ đừng bắt đầu từ việc nó tạo ra được bao nhiêu. Hãy nhìn vào phần còn lại sau khi nó nói “xong”. Con người có phải dọn dẹp nhiều không, kết quả có phải làm lại không, có ai hiểu vì sao hệ thống thay đổi và khi sự cố xảy ra, chúng ta có biết nên bắt đầu tìm từ đâu không? Quan trọng hơn cả, thứ vừa được hoàn thành có thật sự giúp người dùng làm được điều họ cần hay chỉ khiến repository trở nên bận rộn hơn?</p>
<p>Đây cũng là phần ít hào nhoáng nhất trong câu chuyện startup một thành viên. Một người có thể vận hành nhiều agent, nhưng điều đó không khiến các vai trò trong một công ty tự nhiên biến mất. Ở một quán ăn nhỏ, chủ quán có thể vừa nhập hàng, vừa thu ngân, vừa vào bếp khi đông khách. Quán chỉ có một người chủ, nhưng khi đứng ở quầy, cô ấy phải biết bàn nào đã thanh toán; khi bước vào bếp, cô ấy phải biết món nào đang chờ; trước khi đưa đồ ăn ra, cô ấy vẫn nhìn lại xem có đúng bàn hay không. Một người đảm nhiệm nhiều vai trò không có nghĩa những vai trò ấy trở thành một.</p>
<p>Founder dùng AI cũng vậy. Anh ta có thể nhờ agent nghiên cứu, viết code, kiểm thử và làm nội dung, nhưng vẫn cần có lúc bước ra khỏi vai trò của người đang muốn làm thật nhanh để nhìn lại toàn bộ sản phẩm. AI có thể giúp thực hiện công việc, nhưng nó không thể làm chủ hậu quả thay con người. Khi hệ thống gặp lỗi lúc nửa đêm, agent không phải người nghe điện thoại của khách hàng. Khi dữ liệu bị thay đổi sai, agent không phải người giải thích chuyện gì đã xảy ra. Và khi một quyết định kỹ thuật khiến cả sản phẩm phải mất nhiều tuần sửa lại, agent cũng không phải người sống cùng hậu quả của quyết định ấy.</p>
<p>Điều đó không có nghĩa chúng ta phải kéo AI trở lại những quy trình nặng nề, đầy họp hành và giấy tờ. Một quán cơm không cần viết tài liệu hai mươi trang để nhân viên mang đúng món đến đúng bàn; họ chỉ cần một tờ order ghi rõ số bàn, tên món và một người nhìn lại trước khi món rời khỏi bếp. Người gọi thợ sửa vòi cũng không cần lập kế hoạch dự án. Họ chỉ cần nói rõ: hãy sửa chỗ đang rò; nếu phải đục tường hay thay cả đường ống thì gọi lại trước; sửa xong mở nước kiểm tra; nếu ngày mai vẫn còn rò thì quay lại xử lý.</p>
<p>Những điều ấy nghe bình thường vì trong đời sống, chúng ta hiểu rằng một người dù giỏi và nhanh đến đâu cũng cần biết mình đang làm việc gì, được phép đi đến đâu và khi nào công việc mới thật sự kết thúc. Chỉ khi bước vào thế giới AI, bị cuốn theo tốc độ và cảm giác vô hạn, chúng ta mới quên mất sự bình thường ấy.</p>
<p>Quay lại lỗi đặt lịch, founder không cần tạo thêm một tầng quy trình phức tạp. Anh ta chỉ cần giữ công việc ở đúng với câu chuyện ban đầu: khách đã huỷ nhưng khung giờ chưa được mở lại; trước hết hãy tái hiện đúng lỗi đó; chỉ sửa phần liên quan đến việc giải phóng khung giờ; nếu buộc phải thay đổi database hoặc luồng thanh toán thì dừng lại và báo trước. Sau khi sửa, hãy dùng một tài khoản huỷ lịch rồi dùng tài khoản khác đặt lại đúng giờ đó. Một agent khác có thể kiểm tra phần thay đổi, nhưng không được nhân cơ hội viết lại toàn bộ giải pháp.</p>
<p>Lần này, có thể agent chỉ sửa vài file. Bản báo cáo cũng chẳng dài đến mức đáng đem đi khoe. Nhưng khung giờ ba giờ chiều đã được đặt lại thành công. Đó mới là điều người dùng cần, và cũng là lúc công việc có thể thật sự được gọi là hoàn thành. Sự khác biệt nằm ở chỗ kết quả không còn được đánh giá từ phía người tạo ra nó, mà từ phía người đang chờ nhận nó.</p>
<p>Đến đây, hình ảnh dòng nước mới hiện ra rõ ràng hơn. AI giống một con suối trên dốc cao, chảy nhanh, mạnh và mang theo nguồn năng lượng mà trước đây chúng ta chưa từng có. Mỗi agent mở thêm một dòng chảy, mỗi prompt tạo thêm một nhánh nước; code, tài liệu, test và ý tưởng cứ thế đổ xuống. Repository dần trở thành một cái hồ.</p>
<p>Ban đầu, nhìn mặt hồ lớn lên khiến chúng ta tin rằng mình đang tiến bộ. Mỗi ngày có thêm tính năng, thêm file, thêm báo cáo. Nhưng khi nước đã quá nhiều, ta bắt đầu không biết dòng nào sạch, dòng nào đang mang theo bùn đất, chỗ nào đủ sâu và chỗ nào đã âm thầm làm bờ hồ yếu đi. Vấn đề không phải chúng ta có quá nhiều nước. Vấn đề là chưa có con mương đưa nước đến đúng thửa ruộng, chưa có chiếc van để đóng lại khi nước bắt đầu tràn và cũng chưa có ai đứng trên bờ để nói rằng chỗ này đã đủ rồi.</p>
<p>Một người nông dân không khoe rằng hôm nay ruộng của mình nhận được bao nhiêu mét khối nước. Điều họ chờ là mùa thu hoạch. Có lẽ đó cũng là bài học quan trọng nhất khi làm việc với AI: đừng nhầm dòng chảy với kết quả, đừng nhầm sự bận rộn với tiến bộ và đừng nhầm lời thông báo “đã hoàn thành” với việc người dùng đã nhận được điều họ cần.</p>
<p>Thời đại AI hoàn toàn có thể tạo ra những công ty rất nhỏ nhưng làm được những việc từng cần cả một đội ngũ lớn. Điều đó đáng để hào hứng. Nhưng người đi xa có lẽ không phải người mở được nhiều agent nhất, viết được prompt dài nhất hay tạo ra nhiều code nhất. Đó sẽ là người biết việc nào đáng giao, lúc nào cần dừng, kết quả nào có thể tin và đủ tỉnh táo để không biến tốc độ thành một ảo giác về năng suất.</p>
<p>Bởi nước chảy nhanh chưa bao giờ tự làm nên mùa màng. Nó chỉ trở nên có ý nghĩa khi đến đúng nơi, vào đúng lúc và vừa đủ cho điều đang cần được nuôi lớn.</p>
]]></content:encoded></item><item><title><![CDATA[AI Agent Kit - A Story That Began With a Small Worry]]></title><description><![CDATA[🌐 Languages: English · Tiếng Việt bên dướiAI Agent Kit did not begin with a dream of replacing people with AI. It began with a much more ordinary question: if I hand part of a codebase to AI, how do ]]></description><link>https://hunpeolabs.hashnode.dev/ai-agent-kit-a-story-that-began-with-a-small-worry</link><guid isPermaLink="true">https://hunpeolabs.hashnode.dev/ai-agent-kit-a-story-that-began-with-a-small-worry</guid><category><![CDATA[AI]]></category><category><![CDATA[ai-agent]]></category><category><![CDATA[aiagentkit]]></category><category><![CDATA[applied ai]]></category><dc:creator><![CDATA[Hung Pham]]></dc:creator><pubDate>Fri, 21 Aug 2026 21:14:34 GMT</pubDate><content:encoded><![CDATA[<p>🌐 <strong>Languages:</strong> English · Tiếng Việt bên dưới<br />AI Agent Kit did not begin with a dream of replacing people with AI. It began with a much more ordinary question: <strong>if I hand part of a codebase to AI, how do I know it is doing the right thing?</strong></p>
<p>At first, I simply wanted a safer, less repetitive way to bring Claude Code or Codex into a repository. Every project required the same explanations: where the code lived, which commands verified a change, which files belonged to the tool, which ones belonged to the developer, and how to recover when an update went wrong. Back then, I was not thinking about coordinating multiple agents or analyzing software architecture. I just wanted AI to enter a codebase carefully—like a new teammate who pauses at the door, looks around, asks where things belong, and only then starts fixing what they were asked to fix, instead of walking in and rearranging the furniture.</p>
<p>Real repositories, however, are rarely as tidy as a freshly cleaned room. Sometimes I have several pieces of work in progress at once: one worktree contains dozens of modified files, another branch is exploring a different direction, while the agent has been asked to change only one small part of the system. In that situation, the promise “I will only change what is necessary” is not enough to reassure me. It is like asking someone to repair a pipe in a house that is still being renovated. Materials are scattered across the floor, one room is halfway through being painted, and the wiring has just been replaced somewhere else. The person doing the repair may be excellent at their job, but if they cannot tell what is already in progress from what must remain untouched, one small fix can disrupt several days of work.</p>
<p>Experiences like that led AI Agent Kit to care about ownership, dry runs, backups, and rollback. Before making a change, the tool needed to show me what it intended to touch. If I had already edited a file, it needed to preserve my work or report a conflict rather than treating the version bundled with the package as the only correct one. I pictured it as someone coming to maintain a lived-in home: before taking anything apart, they photograph its original state, mark what belongs to the owner and what they installed themselves, and prepare a way to put everything back if the new approach does not work out.</p>
<p>One small bug from that period stayed with me. The file that stored Repository Intelligence state was accidentally included in the very worktree signature used to verify the repository. Every time the tool updated its state, the signature changed, and the system concluded that its own data had become stale. It was like a security guard making a round, leaving footprints behind, then spotting those same footprints and reporting that an intruder had entered the building. It was not a dramatic bug, but it showed me how easily a verification system can produce a confident warning from a flawed measurement.</p>
<p>After that, I became more careful about how AI Agent Kit described what it knew. If CodeGraph or CocoIndex was unavailable, the tool should not speak as though it had seen the entire repository. At the same time, I did not want every task to stop simply because one supporting tool was missing. The more honest response was to say: this part is <code>DEGRADED</code>; this is the scope I can currently inspect, and these are the limits of the conclusion. It is like inspecting a house when the final room is still locked. I can examine the wiring, plumbing, and structure everywhere else, but I cannot honestly write in the report that I inspected the whole building.</p>
<p>The more I used AI in real work, the more I saw that a reminder inside a prompt was not a strong enough boundary. I could say, “Ask before changing anything sensitive,” but as the conversation grew longer, the agent could still forget. I might ask it to prepare a local change, but in its effort to finish the job properly, it could suggest committing, pushing, or opening a pull request. Its intent was not bad. It simply did not feel the boundaries of ownership in the same way a person would. To me, being invited to repair the kitchen does not mean being handed the keys to every room in the house.</p>
<p>That is why approval, policy, and capability gradually moved out of the prompt and into the runtime. An agent needed to know which keys it had been given, which doors required permission before opening, and which areas were entirely outside the scope of the task. When an important action took place, AI Agent Kit needed to leave a record of who authorized it, whether that authority was still valid, and whether the final result stayed within what had been approved. I did not do this to turn every change into a ceremony. I simply did not want an agent’s authority to depend on whether it happened to remember one sentence from the beginning of a long conversation.</p>
<p>There was one point when all the important checks had passed, yet the worktree still contained more than seventy changes. If I looked only at the test results, I could have said the work was complete. Looking at the repository, I knew it was not ready. It felt like starting a car and hearing the engine run perfectly while the hood was still open, tools were scattered across the seats, and several parts had not yet been put back. “The engine starts” was true, but it was not the whole truth about the state of the car.</p>
<p>That experience pushed me to build the Final Task Report and evidence system more carefully. A final report could not stop at listing which tests were green. It also needed to say which commit the evidence belonged to, whether the worktree had changed, what had not been checked, and what was still blocking a release. If the code changed after a review, the old review could not continue to be treated as though nothing had happened. It is similar to a building inspection certificate: if I alter a load-bearing wall after the inspection, the old certificate no longer describes the house as it stands today.</p>
<p>When I began using multiple agents and worktrees in parallel, I ran into a different kind of problem. Each agent could do its own part correctly, yet the final pieces still failed to fit together. One agent might research from an older commit while another had already changed the foundation beneath it. Both branches could pass their tests independently, only to reveal during integration that they had been built from different states of the project.</p>
<p>At one point, the v1.5 branch looked close to complete. But when I checked the release history, I discovered that it did not correctly continue from v1.4.1. Some capabilities that had already shipped—such as Architecture Pulse, routing, and evidence verification—were no longer fully present in the new branch. It was like having two teams renovate the same house from different blueprints. One team had finished strengthening the ground floor. The other had built a beautiful upper floor using a blueprint printed before that foundation work happened. Each part looked sound on its own; only when they were placed together did it become clear that they no longer shared the same history.</p>
<p>Those collisions taught me that multiple agents do not naturally become a team simply because they work in the same repository. If several people are renovating one house, each person needs to know which area belongs to them, which blueprint is current, and who decides what becomes part of the final structure. The Repository Team Control Plane in v1.5 grew out of that practical need. An agent writing code gets its own workspace. Its claim on a task has a limited lifetime. Another agent cannot quietly take over simply because the first one has gone silent for a while. Before integration, the result must be checked against the exact state from which it began, and important changes need an independent review.</p>
<p>Memory came from an equally familiar experience. Across many sessions, I found myself repeating earlier decisions: why I had avoided a particular dependency, which behavior needed to remain backward compatible, or which approach had already been tried and had not worked. If nothing was saved, every new agent had to start from the beginning. But if everything was saved, the system became like a cabinet overflowing with old notes. Some had once been important but had since expired. Some were only guesses made during an investigation. Others belonged to a different project and had somehow ended up in the wrong drawer.</p>
<p>Governed Shared Memory in v1.3 was therefore not built to help agents remember more. I wanted it to work like a carefully maintained project notebook. An agent could suggest something worth keeping, but before that information became a shared convention, it needed to be reviewed. Every note had to explain where it came from, where it applied, and when it should be revisited. If a decision was no longer correct, I needed to be able to replace or revoke it rather than letting agents repeat it forever.</p>
<p>Later, I ran into a problem that ordinary tests could not answer. The code still ran, lint was clean, and an individual pull request showed no obvious defect, yet the architecture was quietly becoming harder to change. One module started depending on another in the wrong direction. A small file gradually became the place every change had to pass through. A new dependency cycle appeared, but the total number of cycles stayed the same because an older one had just been fixed. The number looked stable; in reality, one problem had merely been replaced by another.</p>
<p>It reminded me of the electrical wiring and plumbing hidden behind the walls of a house. The lights can still work and the water can continue to flow even while the lines behind the walls become increasingly tangled. Nothing seems wrong on the first day. But the next time something needs repair, replacing one outlet requires opening an entire wall, or fixing a pipe in the kitchen affects the bathroom. A codebase can behave the same way. Tests tell me that the house works today, but they do not necessarily tell me whether it will still be easy to repair tomorrow.</p>
<p>Architecture Pulse in v1.4 was built to look behind those walls. It compares the structure of a repository before and after an agent works, looking for dependency cycles, crossed boundaries, hotspots, and the likely reach of a change. While building it, I also discovered that I had sometimes put the “lock” in the wrong place. One CI guard was too broad and blocked even the creation of an in-memory baseline, although the real danger was writing an untrusted baseline out as an artifact. It was like trying to stop someone from carrying documents out of a room, but locking the reading desk instead of the door. The answer was not to remove the lock. It was to move it to the boundary that actually needed protection.</p>
<p>Another test worked on POSIX systems but failed on Windows. That reminded me that a key which opens the door in my own house may not fit a similar-looking door somewhere else. AI Agent Kit could not work well only in the environment I used every day and call that enough. Filesystem behavior, paths, symlinks, and package boundaries needed to be tested on the platforms where people would actually run the tool.</p>
<p>By then, I thought AI Agent Kit covered many of the important parts of AI-assisted engineering: agents had clear scope, actions left evidence, multiple agents had a way to coordinate, memory was governed, and architecture could be observed. Then I realized that all of those systems began after I had already decided what product to build.</p>
<p>In practice, some of my projects begin with only a short idea. An agent often responds immediately by proposing features, choosing a technology stack, splitting the work into a backlog, and preparing implementation. At first, that momentum feels good. But it is also a little like hiring an excellent construction team and asking them to build a house before I have decided who will live there, how many rooms they need, or how they move through their day. The house may follow the blueprint perfectly, arrive on schedule, and be structurally sound, yet still fail the people who eventually have to live in it.</p>
<p>I did not want a rough idea to turn immediately into a polished-looking specification. I wanted time to understand the problem, test assumptions, and separate what I knew from what I was still guessing. If the business requirements had not been approved, the agent should not move on to design by itself. If the design changed something that had already been agreed upon, the affected parts needed to be reviewed again. Product Genesis in v1.6 grew from that need. It is the stage where I sit down with the first rough sketch and ask who the house is for, why it needs to exist, and what truly matters before choosing the materials.</p>
<p>Once Product Genesis was working, I noticed that entering the process still did not feel natural enough. To begin, people had to know the name of a skill, remember a workflow, or use the right command. But that is not how I usually begin an idea. Most of the time, I simply say, “I have been thinking about a product like this…” I did not want to learn the language of the tool before the tool could understand mine.</p>
<p>That is why v1.6.1 made the first doorway easier to enter. I can describe an idea in Vietnamese or English as part of a normal conversation. AI Agent Kit recognizes that I am beginning a product idea, finds the existing workspace if the journey is already underway, or starts a new one when needed. The entrance is easier, but that does not mean every door inside is left open. Important decisions still require human approval, evidence still has to match the real state of the work, and AI still needs to know when to stop because there is not enough information to continue responsibly.</p>
<p>Looking back, nearly every part of AI Agent Kit began with a specific collision between what seemed complete and what was actually true. A tool invalidated its own signature. A worktree passed its tests while still containing more than seventy changes. Two good branches no longer shared the same history. A test worked on one machine but failed on Windows. An agent remembered a great deal but could not tell what had become stale. A pull request was green while the architecture beneath it was becoming harder to change. And finally, an entire team could move quickly while the question “Should we build this at all?” remained unanswered.</p>
<p>That is why the original question still sits at the center of AI Agent Kit: if I hand part of a codebase to AI, how do I know it is doing the right thing? After everything I have experienced, however, the meaning of “right” has grown much wider. It no longer means only that the code runs. It means entering the right place, touching the right part, using the right authority, working from the right information, leaving the right evidence, and still moving toward something I genuinely want to build.</p>
<h2>Explore AI Agent Kit</h2>
<ul>
<li><p>📦 <a href="https://www.npmjs.com/package/@hunpeolabs/ai-agent-kit">AI Agent Kit on npm</a></p>
</li>
<li><p>💻 <a href="https://github.com/phamhungptithcm/ai-agent-kit">Source code on GitHub</a></p>
</li>
<li><p>⭐ If you find the project useful, consider starring the repository.</p>
</li>
</ul>
<hr />
<h2>🇻🇳 Phiên bản tiếng Việt</h2>
<h3>AI Agent Kit - Câu chuyện bắt đầu từ một nỗi lo rất nhỏ</h3>
<p>AI Agent Kit không ra đời từ giấc mơ thay con người bằng AI. Nó bắt đầu bằng một câu hỏi gần gũi hơn nhiều: <strong>nếu giao một phần codebase cho AI, làm sao biết nó đang làm đúng?</strong></p>
<p>Thời gian đầu, mình chỉ muốn việc đưa Claude Code hay Codex vào một repository trở nên an toàn và đỡ lặp lại hơn. Mỗi dự án lại phải giải thích từ đầu: code nằm ở đâu, lệnh nào dùng để kiểm tra, file nào do công cụ quản lý, phần nào thuộc về người phát triển và nếu một lần cập nhật xảy ra lỗi thì có thể quay lại bằng cách nào. Khi ấy, mục tiêu chưa phải điều phối nhiều agent hay đánh giá kiến trúc. Mình chỉ muốn AI bước vào codebase một cách cẩn thận, giống như một thành viên mới biết đứng ở cửa quan sát, hỏi vị trí đồ đạc rồi mới bắt đầu sửa thứ được giao, thay vì vừa bước vào nhà đã tự ý kê lại bàn ghế.</p>
<p>Nhưng repository thật hiếm khi gọn gàng như một căn phòng vừa được dọn sạch. Có lúc mình đang làm dở nhiều việc cùng lúc, một worktree chứa hàng chục file đã thay đổi, một nhánh khác đang thử hướng mới, còn agent chỉ được giao sửa đúng một phần nhỏ. Trong hoàn cảnh đó, lời hứa “mình sẽ chỉ thay đổi những file cần thiết” chưa làm mình yên tâm. Nó giống như gọi một người đến sửa đường ống trong căn nhà đang được cải tạo: trên sàn còn vật liệu, một phòng đang sơn dở, hệ thống điện vừa được thay ở khu vực khác. Người thợ có thể rất giỏi, nhưng nếu không biết đâu là phần đang thi công và đâu là thứ cần giữ nguyên, một việc sửa nhỏ cũng có thể làm xáo trộn công việc của nhiều ngày.</p>
<p>Từ những tình huống như vậy, AI Agent Kit bắt đầu quan tâm đến ownership, dry-run, backup và rollback. Trước khi thay đổi, công cụ cần cho mình xem nó định chạm vào đâu. Nếu file đã được mình sửa, nó phải biết giữ lại hoặc báo xung đột, chứ không được xem phiên bản đi kèm package là câu trả lời duy nhất. Mình hình dung nó như một người đến bảo trì ngôi nhà: trước khi tháo một món đồ, người đó chụp lại trạng thái ban đầu, đánh dấu thứ gì thuộc về chủ nhà, thứ gì do mình lắp đặt và chuẩn bị cách lắp lại nếu phương án mới không phù hợp.</p>
<p>Có một lỗi nhỏ trong quá trình phát triển khiến mình nhớ khá lâu. File lưu trạng thái của Repository Intelligence vô tình bị tính vào chính chữ ký dùng để kiểm tra repository. Mỗi lần công cụ cập nhật trạng thái, chữ ký lại thay đổi và hệ thống tự kết luận dữ liệu của mình đã cũ. Nó giống như một người bảo vệ đi tuần, để lại dấu chân rồi nhìn thấy chính dấu chân ấy và báo rằng vừa có người lạ xâm nhập. Lỗi không lớn, nhưng nó cho mình thấy một hệ thống kiểm tra vẫn có thể tạo ra cảnh báo rất chắc chắn từ một cách đo chưa đúng.</p>
<p>Sau lần đó, mình thận trọng hơn với cách AI Agent Kit diễn đạt điều nó biết. Nếu CodeGraph hoặc CocoIndex chưa sẵn sàng, công cụ không nên nói như thể đã nhìn thấy toàn bộ repository. Nhưng mình cũng không muốn chỉ vì thiếu một công cụ hỗ trợ mà mọi công việc đều phải dừng. Cách hợp lý hơn là nói rõ: phần này đang ở trạng thái <code>DEGRADED</code>, mình chỉ nhìn được trong phạm vi này và kết luận hiện tại có giới hạn như thế nào. Nó giống như kiểm tra một căn nhà khi chưa mở được căn phòng cuối cùng. Mình vẫn có thể xem phần điện, nước và kết cấu ở những nơi còn lại, nhưng không thể viết vào báo cáo rằng đã kiểm tra toàn bộ ngôi nhà.</p>
<p>Càng dùng AI trong công việc thật, mình càng thấy lời nhắc trong prompt không phải là một ranh giới đủ chắc. Mình có thể nói “hãy hỏi trước khi thay đổi phần nhạy cảm”, nhưng khi cuộc trò chuyện dài lên, agent vẫn có thể quên. Mình chỉ nhờ chuẩn bị thay đổi local, nhưng vì muốn hoàn thành công việc cho trọn vẹn, agent có thể đề nghị commit, push hoặc mở pull request. Ý định của nó không xấu, chỉ là nó không cảm nhận được ranh giới sở hữu như con người. Với mình, việc được mời vào sửa căn bếp không đồng nghĩa với việc có chìa khóa của mọi phòng trong nhà.</p>
<p>Vì vậy approval, policy và capability dần được đưa ra khỏi prompt để trở thành một phần của runtime. Agent cần biết mình đang được giữ chìa khóa nào, cánh cửa nào phải hỏi trước khi mở và khu vực nào hoàn toàn nằm ngoài phạm vi công việc. Khi một hành động quan trọng xảy ra, AI Agent Kit cần để lại dấu vết cho biết ai đã cho phép, quyền đó còn hiệu lực không và kết quả sau cùng có đúng với phần đã được duyệt hay không. Mình không làm vậy để biến mọi thay đổi thành thủ tục. Mình chỉ không muốn quyền hạn của agent phụ thuộc vào việc nó có nhớ đúng một câu đã xuất hiện từ đầu cuộc trò chuyện.</p>
<p>Có lần các kiểm tra quan trọng đều đã pass, nhưng worktree vẫn còn hơn bảy mươi thay đổi. Nếu chỉ nhìn vào test, mình có thể nói công việc đã hoàn thành. Nhưng nhìn vào repository, mình biết mình chưa thể gọi nó là sẵn sàng. Cảm giác ấy giống như chạy thử một chiếc xe và thấy động cơ hoạt động tốt, trong khi nắp ca-pô vẫn mở, dụng cụ còn nằm trên ghế và nhiều bộ phận chưa được lắp lại. “Máy đã nổ” là một thông tin đúng, nhưng chưa phải toàn bộ sự thật về trạng thái của chiếc xe.</p>
<p>Từ đó, mình bắt đầu xây Final Task Report và hệ thống evidence cẩn thận hơn. Báo cáo cuối không chỉ cần nói test nào đã xanh. Nó còn phải nói bằng chứng gắn với commit nào, worktree còn thay đổi không, điều gì chưa được kiểm tra và phần nào vẫn đang chặn release. Nếu code thay đổi sau khi review, kết quả review cũ không thể tiếp tục được dùng như thể chưa có gì xảy ra. Nó giống như giấy kiểm định của một căn nhà: nếu sau khi kiểm định mình tiếp tục sửa phần chịu lực, tờ giấy cũ không còn mô tả đúng căn nhà hiện tại.</p>
<p>Khi bắt đầu dùng nhiều agent và nhiều worktree song song, mình gặp một kiểu rắc rối khác. Mỗi agent có thể làm đúng phần của mình, nhưng kết quả cuối cùng vẫn không ghép lại được. Một agent nghiên cứu trên commit cũ, một agent khác đã thay đổi phần nền. Hai nhánh riêng lẻ đều pass test, nhưng khi đưa về cùng một nơi mới thấy chúng đã dựa trên hai trạng thái khác nhau của dự án.</p>
<p>Có lần nhánh v1.5 trông khá hoàn chỉnh nhưng khi kiểm tra lịch sử phát hành, mình phát hiện nó không đi tiếp đúng từ v1.4.1. Một số khả năng đã có trong phiên bản trước như Architecture Pulse, routing và evidence verification không còn đầy đủ trong nhánh mới. Nó giống như hai đội cùng cải tạo một ngôi nhà từ hai bản vẽ khác nhau. Một đội đã sửa xong tầng trệt, đội còn lại xây tầng trên rất đẹp nhưng lại dùng bản vẽ được in trước khi phần móng được gia cố. Từng phần nhìn riêng đều ổn; chỉ khi đặt chúng lên nhau mới thấy ngôi nhà không còn cùng một lịch sử.</p>
<p>Những va chạm ấy khiến mình hiểu rằng nhiều agent không tự nhiên trở thành một đội chỉ vì chúng cùng làm trên một repository. Nếu nhiều người cùng sửa một ngôi nhà, mỗi người cần biết khu vực của mình, bản vẽ nào đang có hiệu lực và ai chịu trách nhiệm quyết định phần nào được lắp vào công trình chung. Repository Team Control Plane trong v1.5 được phát triển từ nhu cầu rất thực tế đó. Agent viết code có không gian làm việc riêng. Quyền giữ một task có thời hạn. Một agent khác không thể âm thầm tiếp quản chỉ vì agent đầu tiên tạm thời không phản hồi. Trước khi tích hợp, kết quả phải được đối chiếu với đúng trạng thái ban đầu và những phần quan trọng cần một lượt kiểm tra độc lập.</p>
<p>Memory cũng bắt đầu từ một trải nghiệm rất quen. Qua nhiều phiên làm việc, mình phải giải thích lại những quyết định đã có: vì sao không dùng một dependency, phần nào cần giữ tương thích, cách nào từng thử nhưng không hiệu quả. Nếu không lưu, mỗi agent mới lại phải đọc từ đầu. Nhưng nếu lưu tất cả, hệ thống sẽ giống một chiếc tủ chứa đầy giấy ghi chú cũ. Có tờ từng rất quan trọng nhưng nay đã hết hạn. Có tờ chỉ là một suy đoán trong lúc tìm hiểu. Có tờ thuộc về dự án khác nhưng vô tình được đặt nhầm ngăn.</p>
<p>Governed Shared Memory trong v1.3 vì thế không được xây để agent nhớ nhiều hơn. Mình muốn nó giống một cuốn sổ làm việc được chăm sóc cẩn thận. Agent có thể đề xuất một điều đáng ghi lại, nhưng trước khi trở thành quy ước dùng chung, thông tin đó cần được xem xét. Mỗi ghi chú phải cho biết nó đến từ đâu, dùng trong phạm vi nào và khi nào cần xem lại. Nếu một quyết định không còn đúng, mình có thể thay thế hoặc thu hồi nó thay vì để agent tiếp tục lặp lại mãi về sau.</p>
<p>Sau đó mình gặp một vấn đề mà test thông thường không trả lời được. Code vẫn chạy, lint vẫn sạch và pull request nhìn riêng không có lỗi rõ ràng, nhưng kiến trúc đang âm thầm trở nên khó thay đổi. Một module bắt đầu phụ thuộc ngược vào module khác. Một file nhỏ dần trở thành điểm mà mọi thay đổi đều phải đi qua. Một dependency cycle mới xuất hiện, nhưng tổng số cycle không đổi vì một cycle cũ vừa được sửa. Nhìn vào con số, mọi thứ có vẻ đứng yên; nhìn vào bản chất, một vấn đề cũ đã được thay bằng một vấn đề mới.</p>
<p>Điều đó làm mình liên tưởng đến hệ thống điện và nước nằm sau những bức tường. Một căn nhà vẫn có thể sáng đèn và nước vẫn chảy bình thường, dù đường dây bên trong đã bắt đầu được nối chồng chéo. Trong ngày đầu, không ai thấy vấn đề. Nhưng đến lần sửa tiếp theo, muốn thay một ổ điện lại phải mở cả mảng tường; muốn sửa đường nước ở bếp lại ảnh hưởng đến phòng tắm. Codebase cũng vậy. Test cho mình biết căn nhà hôm nay vẫn hoạt động, nhưng chưa chắc cho biết ngày mai nó còn dễ sửa hay không.</p>
<p>Architecture Pulse trong v1.4 được xây để nhìn vào phần nằm sau những bức tường đó. Nó so sánh cấu trúc repository trước và sau khi agent làm việc, tìm các vòng phụ thuộc, ranh giới bị vượt qua, điểm nóng và phạm vi ảnh hưởng. Trong quá trình phát triển, mình cũng nhiều lần đặt “ổ khóa” sai chỗ. Có lúc CI guard được thiết kế quá rộng và chặn cả việc tạo baseline trong bộ nhớ, dù thứ cần ngăn chỉ là ghi một baseline không đáng tin thành artifact. Nó giống như muốn ngăn người lạ mang tài liệu ra khỏi phòng, nhưng lại khóa luôn chiếc bàn nơi người bên trong đang đọc tài liệu. Cách sửa không phải bỏ ổ khóa, mà là đặt nó lại đúng cánh cửa.</p>
<p>Một test khác chạy tốt trên POSIX nhưng thất bại trên Windows. Điều đó nhắc mình rằng một chiếc chìa khóa mở được cửa nhà mình chưa chắc mở được cùng loại cửa ở nơi khác. AI Agent Kit vì thế không thể chỉ hoạt động tốt trong môi trường mình dùng hằng ngày rồi xem đó là đủ. Các giới hạn về filesystem, đường dẫn, symlink và package cần được kiểm tra trên những nền tảng mà người dùng thực sự sẽ chạy.</p>
<p>Đến lúc này, mình từng nghĩ AI Agent Kit đã bao phủ khá nhiều phần quan trọng: agent có phạm vi, hành động có bằng chứng, nhiều agent có cách phối hợp, memory được kiểm soát và kiến trúc có thể quan sát. Nhưng rồi mình nhận ra tất cả những điều đó đều bắt đầu sau khi mình đã quyết định sẽ xây sản phẩm gì.</p>
<p>Trong thực tế, có những lúc mình bắt đầu chỉ bằng một ý tưởng rất ngắn. Agent thường phản hồi nhanh: đề xuất tính năng, chọn công nghệ, chia backlog rồi chuẩn bị implementation. Cảm giác ban đầu khá thích vì mọi thứ tiến lên ngay lập tức. Nhưng nó cũng giống như thuê một đội xây dựng rất giỏi rồi yêu cầu họ dựng nhà khi mình chưa nghĩ rõ ai sẽ sống trong đó, cần bao nhiêu phòng hay thói quen sinh hoạt hằng ngày ra sao. Ngôi nhà có thể được xây đúng bản vẽ, đúng tiến độ và rất chắc chắn, nhưng sau cùng vẫn không phù hợp với người sử dụng.</p>
<p>Mình không muốn một ý tưởng vừa được nói ra đã lập tức biến thành một bản đặc tả có vẻ hoàn chỉnh. Mình muốn có thời gian để hiểu vấn đề, kiểm tra giả định và tách điều mình biết khỏi điều mình mới đang đoán. Nếu BRD chưa được duyệt, agent không nên tự bước sang thiết kế. Nếu thiết kế làm thay đổi điều đã thống nhất trước đó, những phần liên quan cần được xem lại. Product Genesis trong v1.6 ra đời từ nhu cầu ấy. Nó giống giai đoạn mình ngồi xuống với bản phác thảo đầu tiên, hỏi căn nhà này dành cho ai, vì sao cần xây và điều gì thực sự quan trọng trước khi bắt đầu chọn vật liệu.</p>
<p>Khi Product Genesis đã hoạt động, mình lại thấy cách bước vào hành trình vẫn chưa đủ tự nhiên. Để bắt đầu, người dùng phải biết tên skill, nhớ workflow hoặc nói đúng câu lệnh. Trong khi cách mình thường bắt đầu một ý tưởng chỉ đơn giản là: “Mình đang nghĩ đến một sản phẩm như thế này…” Mình không muốn phải học ngôn ngữ của công cụ trước khi công cụ có thể hiểu ngôn ngữ của mình.</p>
<p>Vì vậy v1.6.1 làm cho cánh cửa đầu tiên trở nên nhẹ hơn. Mình có thể kể về ý tưởng bằng tiếng Việt hoặc tiếng Anh như một cuộc trò chuyện bình thường. AI Agent Kit nhận ra đó là một ý tưởng sản phẩm, tìm lại workspace nếu hành trình đang dang dở hoặc bắt đầu một hành trình mới khi cần. Cánh cửa dễ mở hơn, nhưng không vì thế mà mọi căn phòng đều được mở sẵn. Những quyết định quan trọng vẫn cần con người phê duyệt, bằng chứng vẫn phải gắn với trạng thái thật và AI vẫn phải biết dừng khi chưa đủ thông tin để đi tiếp.</p>
<p>Nhìn lại, mỗi phần của AI Agent Kit thường bắt đầu từ một va chạm rất cụ thể. Một công cụ tự làm chữ ký của mình hết hạn. Một worktree pass test nhưng vẫn còn hơn bảy mươi thay đổi. Hai nhánh đều tốt nhưng không đi cùng một lịch sử. Một test chạy trên máy này nhưng thất bại trên Windows. Một agent nhớ rất nhiều nhưng không biết điều gì đã cũ. Một pull request xanh toàn bộ nhưng kiến trúc bên trong đang khó sửa hơn. Và cuối cùng, một đội có thể xây rất nhanh trong khi câu hỏi “có nên xây điều này không?” vẫn chưa được trả lời.</p>
<p>Vì vậy câu hỏi ban đầu vẫn còn ở trung tâm AI Agent Kit: nếu giao một phần codebase cho AI, làm sao biết nó đang làm đúng? Chỉ là sau từng trải nghiệm, chữ “đúng” với mình đã rộng hơn. Nó không còn chỉ là code chạy. Nó còn là bước vào đúng chỗ, chạm vào đúng phần, dùng đúng quyền, dựa trên đúng thông tin, để lại đúng bằng chứng và vẫn đi về hướng mà mình thật sự muốn xây.</p>
<h2>Khám phá AI Agent Kit</h2>
<ul>
<li><p>📦 <a href="https://www.npmjs.com/package/@hunpeolabs/ai-agent-kit">AI Agent Kit trên npm</a></p>
</li>
<li><p>💻 <a href="https://github.com/phamhungptithcm/ai-agent-kit">Mã nguồn trên GitHub</a></p>
</li>
<li><p>⭐ Nếu thấy dự án hữu ích, bạn có thể dành tặng repository một ngôi sao.</p>
</li>
</ul>
]]></content:encoded></item></channel></rss>