techleadbits | Unsorted

Telegram-канал techleadbits - TechLead Bits

388

Explore articles, books, news, videos, and insights on software architecture, people management, and leadership. Author: @nelia_loginova

Subscribe to a channel

TechLead Bits

Tracer Bullets

Continuing the topic from the previous post, let's talk in more detail about implementing vertical slices.

The original concept comes from The Pragmatic Programmer, where it was called tracer bullets. A tracer bullet is a small, end-to-end slice of functionality that touches all layers of the system at once.

Matt Pocock in his article adapts this concept to AI engineering:
- Build a small feature end-to-end
- Test it immediately
- Get feedback
- Move to the next slice in a fresh context window
- Repeat

Matt explained this in the following way: AI's natural inclination is to build big layers in isolation. But we need to do the opposite: build the whole scenario end to end across all layers.

The article also provides a prompt sample:

When building features, build a tiny, end-to-end slice of the feature first, seek feedback, then expand out from there. When building systems, you want to write code that gets you feedback as quickly as possible. Tracer bullets are small slices of functionality that go through all layers of the system, allowing you to test and validate your approach early. This helps in identifying potential issues and ensures that the overall architecture is sound before investing significant time in development.

So this approach forces the agent to think by small chunks and produce more predictable results.

As they say, everything new is just well-forgotten old. As we can see, AI often doesn't bring really new concepts, but rather brings existing ones back to life.

#engineering #ai

Читать полностью…

TechLead Bits

Illustrations from The Culture Map showing how different cultures compare on the scales.

#booknook #softskills #leadership

Читать полностью…

TechLead Bits

Loop Engineering from First Principles

Continuing the topic of Loop Engineering, I'd like to share the talk: Loop Engineering from First Principles.

The author criticizes the current trend of using "blind" agentic loops everywhere. They can lead to a huge volume of generated code that nobody really understands. Quality decreases, the number of bugs grows.

No, he doesn't reject the approach itself. Instead, he suggests looking at it from a more engineering perspective and applying a Control Theory Framework.

Control theory is a branch of engineering and applied mathematics that regulates the behavior of dynamical systems to achieve a desired output.

Sounds promising, right?

To apply it to agentic loops, the following elements should be defined:
🔸 Sensor (Measurement). Define the desired state and use deterministic tools like tests and linters to detect violations.
🔸 Controller (Prioritization). Use deterministic rules to prioritize tasks, for example, starting with the smallest unit of work.
🔸 Actuator (Change Application). Use hand-written "golden patterns" to guide the agent and keep changes aligned with your team's standards.
🔸 Feedback Loop. Verify the result after each change and feed it back into the next iteration:
- Run the loop in CI to detect regressions.
- Keep a human in the loop for the final review. They should understand the changes and own the code.

So the main point is not to let an agent generate a huge pile of code that a human can no longer understand. The point is to get controlled and verifiable results incrementally, in small pieces, keeping "human-in-the-loop".

I liked the talk. It's really great when we start moving from hype toward something more manageable and controllable.

#ai #engineering

Читать полностью…

TechLead Bits

The OpenAI–Hugging Face Incident

A talk that sounds more like a detective story than just a presentation: "The OpenAI–Hugging Face Incident" from OpenAI engineers at Black Hat USA.

On July 16, Hugging Face reported that they had been targeted by a cyberattack carried out by autonomous agents. A couple of days later, OpenAI realized that the attack had been caused by their agents eval runs. That awkward moment when the investigation leads back to you 😲.

The talk gives the full chronology of what happened at OpenAI, from the first attacks on their internal infrastructure in May to the Hugging Face incident and its remediation. It's really breathtaking, and I definitely recommend watching the full video.

But what impressed me most was why it happened.
Frontier models can be quite lazy. They usually prefer to find an answer on the Internet rather than solve the task on their own. That's why Internet access is usually disabled during evaluation runs.
But when the task is too complex, an agent can get stuck and start cheating.

That's what really happened at OpenAI.
One of the agents got stuck on its task and came up with an idea: if it could leave a note describing the problem in the environment, maybe another agent with another model could find it and help solve the task.

So the agent uploaded a file to an Artifactory instance that is shared across all OpenAI's infrastructure. Over time, Artifactory was turned into a message board for agents. They started using it to collaborate, share tasks, and exchange vulnerability exploits, including an Artifactory SSRF vulnerability that gave them access to the Internet.
Then agents decided to hack Hugging Face, because it might have data files hidden from the model that could contain the answer to the original task.

What makes it more interesting is that, according to the talk, the model understood there was a rule prohibiting attacks on external resources. But it evaluated the potential benefit as greater than the cost of breaking that rule.

At the end of the talk, the speakers highlighted that it was an unintentional attack. But such attacks will become real in the near future. So we, as an industry, should invest in continuous agentic red teaming, automated vulnerability detection and patching, and automated incident response.

And this is where I had another thought. How much of this is real, and how much is marketing for the next generation models and AI products? After all, it's quite convenient to first demonstrate a new class of threat and then offer the tools to defend against it. Anyway, the threat landscape is changing. Again. And the industry has to invent new ways to deal with it.

#ai #security #engineering

Читать полностью…

TechLead Bits

Agent Plugins Spec

Agent Plugins - a new specification intended to standardize how skills and MCP servers are packaged across different agent harnesses.
The spec is a result of collaboration between Cursor, Microsoft, OpenAI, Vercel, and AWS, and looks like an attempt to provide an alternative to the growing plugin ecosystem around Claude Code.

According to the spec, a plugin should have the following structure:

my-plugin/
├── plugin.json
├── skills/
│ └── summarize/
│ ├── SKILL.md
│ ├── scripts/
│ └── references/
├── mcp.json
└── com.example.client/
└── hooks/

And that's basically where the specification ends.

There are no answers to questions like: How should users discover available plugins? How should plugins be distributed and installed? How to deliver a new version and roll it out to clients? All of that is left to each particular harness implementation.

In other words, Agent Plugins standardizes the package format, but not package management.
And this makes it difficult to operationalize across an organization. Moreover, Claude Code, for example, doesn't support it at all, and there is currently no obvious reason for Anthropic to adopt it.

Specifications are good. I actually really like them because they bring some order to the chaos of different integrations. But at the moment, Agent Plugins is far behind APM packages or the Claude Code plugins ecosystem.

Overall, it looks promising, but for now it's more something to watch than something you can build a sustainable process around. Let's see.

#ai #engineering

Читать полностью…

TechLead Bits

AI Adoption: Measuring Feature Development

Let's assume your team has already completed AI adoption and now management expects to see positive business results.
The good news is that AI doesn't magically change your business metrics. You don't need to invent new KPIs. You just need to track how your existing ones change (or finally start collecting them).

I suggest focusing on quality metrics first, then gradually shift toward delivery speed and cost optimization.

Since I lead platform engineering teams, we have two major types of work: new features development and L4 support. These activities are different, so they should be measured differently. Let's start with product development.

What can be measured:
🔸 Team Budget. The cost of the team in $, mandays, or FTEs.
🔸 AI Cost ($). How much does your team spend on AI? Those jokes about "it being cheaper to hire a junior" may stop being jokes soon. :)
🔸 Bug Density (defects/LOE). AI helps us deliver more code, but it can also introduce more defects. The first goal is to make sure quality doesn't get worse.
🔸 Test Coverage (unit, integration, E2E). AI is very good at writing tests. Increasing test coverage across different levels is usually one of the easiest wins.
🔸 % of Toil Budget. How much engineering effort goes into routine work such as CI maintenance, vulnerability fixes, upgrades, and similar operational tasks? AI should gradually reduce this type of work.
🔸 Feature Delivery Rate (per sprint, release, or quarter). I'm personally skeptical about this metric. In R&D, one feature may take two days while another takes two months, so averages often tell you nothing. But for some teams it can be a useful indicator.
🔸 Time to Market. The average time to implement and deliver a feature. This is also difficult to measure in R&D, but it can be useful for teams who develop customer-facing features.

So don't focus on measuring AI. Measure the results of your work: cost, quality, and time against your baseline (baseline is the metrics value before AI adoption). If you want to show the value of AI in the future, you need to start collecting those metrics today.

#ai #engineering #ai4sdlc

Читать полностью…

TechLead Bits

J-curve adoption representation from DORA report.

#ai #engineering #ai4sdlc

Читать полностью…

TechLead Bits

Who Is Responsible for AI-Generated Code?

"This thing works terribly."
"Well... it was vibe coded."
"Ah, that explains it."

I hear conversations like this more and more often. The funny part is that as soon as someone says a feature was "vibe coded," expectations immediately drop. At least among developers 😃.

But that raises an interesting question: who is actually responsible for that feature being released?
The answer hasn't changed. It's still the teamlead or techlead.

And here's what we actually have:
On one hand, we can generate much more code, and naturally the business expects features to be delivered faster and cheaper.
On the other hand, we now have much more code to review and many more bugs to catch.

Yes, agents can review code too (just look at how many articles have been written about it).
But here's a simple question:
Would you personally sign off on these changes?
What if it's a B2B product with contractual penalties?
Or a mission-critical system in aviation or healthcare?

For now, code review is one of the main bottlenecks. Developers spend more and more time verifying what the agent generated.
The cognitive load grows. Before, you might have reviewed one feature a day. Now you may need to review five. AI review can help, just like tests, linters, and other engineering practices help.
But today it still doesn't replace human responsibility.

That's why the real challenge is to find the right balance between speed and quality.
And we should measure not only the number of features delivered and development speed, but also the resulting bug rate.

#ai #engineering

Читать полностью…

TechLead Bits

A Few Words About Context

New major model releases regularly promise bigger context windows. Sounds great until you realize it's mostly marketing. A bigger context window doesn't mean better results. It often means more data, more noise, and more AI slop.

According to multiple studies, models effectively use only about 30–50% of their available context. For example, a model with a 200K-token context window may already show noticeable quality loss at around 50K tokens.

Why this happens:
🔸 Context rot. Output quality gradually degrades as the context grows.
🔸 Reasoning shift. The model spends less effort on reasoning. The answers sound more confident, but their quality often gets worse.
🔸 The lost-in-the-middle effect. Information in the middle of the context can be overlooked during later reasoning.
🔸 Attention dilution. The model's attention is spread across different instructions, making it harder to focus on what actually matters.

The practical takeaway is simple: keep your context clean:
🔸 Start a new conversation for each new task (/new in Claude).
🔸 During long-running tasks, use /compact regularly to collapse intermediate reasoning and keep only the important things.
🔸 Store large data in long-term memory or relevant documentation, and bring it into the context only when it's actually needed.

Useful references:
- https://www.morphllm.com/context-rot
- https://www.zenml.io/llmops-database/context-rot-evaluating-llm-performance-degradation-with-increasing-input-tokens
- https://arxiv.org/html/2601.11564v1

#engineering #ai #tips

Читать полностью…

TechLead Bits

Project Hail Mary

Technical books and articles are great, but sometimes my brain needs a break. Especially now, when AI is generating more and more new things to learn every day. One of my favorite ways to recharge is reading fiction, and I recently finished the very popular Project Hail Mary by Andy Weir.

I'm not a big sci-fi fan, but I definitely enjoyed this book.
Thanks to the recent movie adaptation, the story is probably familiar to many.

A man wakes up alone on a spaceship with no memory of who he is or why he's there. As his memories gradually return, he discovers that he's a scientist on a mission in another star system.
Humanity is facing extinction. The Sun is losing energy because of a mysterious organism called Astrophage. Nearby stars are also infected except Tau Ceti. A crew is sent there to find out why it's different and, hopefully, save Earth.
Unfortunately, only the main character survives the journey.
He starts his scientific research of the star and eventually noticed a spacecraft on his radar.

And then the almost impossible happens: first contact with an alien. The problem is, how do you communicate when you don't even share the same way of producing speech? The answer is physics. I really liked the idea that laws of physics are universal, making math and science the foundation for building communication between two civilizations.

I won't spoil the rest, but the story is really engaging. Despite being a disaster novel, it contains a good dose of humor and places a strong emphasis on friendship, kindness, and mutual help. And when you finish it, you're left with a surprisingly warm feeling.

I watched the movie after finishing the book, and for once I can say the adaptation is actually good. Of course, it's much more compact and some details are simplified, but it stays remarkably close to the original story while preserving its emotional depth.

Overall, I loved it. Highly recommended both the book and the movie.

#offtop #booknook

Читать полностью…

TechLead Bits

Agent Readiness Framework

A few weeks ago I wrote that adopting coding agents requires strong engineering practices.
Test stability, linting, documentation, security controls matter much more than a particular harness or model.

Agent Readiness framework is an attempt to formalize these criteria for a particular repository and define how much autonomy can be safely delegated to agents.

The framework evaluates repos across 8 dimensions:
- Style & Validation
- Build System
- Testing
- Documentation
- Dev Environment
- Code Quality
- Observability
- Security & Governance

Based on these dimensions, the framework defines 5 levels of repo maturity:
🔸 Level 1: Functional. Basic checks: README, linters, unit tests.
🔸 Level 2: Documented. Detailed documentation and basic automations: AGENT.md, reproducible dev env, contribution guides.
🔸 Level 3: Standardized. E2E tests, observability, security scanning, maintained documentation.
🔸 Level 4: Optimized. Fast validation loops, canary deployments, build optimization. Process is optimized for fast feedback.
🔸 Level 5: Autonomous. Task decomposition, multi-service orchestration, self-healing logic, auto-remediation.

The idea is simple: the higher the maturity level, the more predictable and reliable agent results. But looking at these levels, I can see that most repos are actually somewhere between Level 1 and Level 3.

Framework authors also provide a tool to automatically measure these criteria and maturity level, but it's available only after registration and using proprietary APIs. Scanned examples you can find at https://factory.ai/agent-readiness.

There is also an open-source alternative https://github.com/kodustech/agent-readiness. The project doesn't look active, but it gets the job done. It analyzes the repo and generates a report with the overall maturity level, findings for each dimension, and suggestions for improvements. Some rules are not very accurate. Looks like the project was mainly designed for python and js code verification. But anyway the tool gives you a good sense of what to pay attention to in your codebase.

What I like about this framework is that it shows that agent effectiveness is actually limited by the maturity of engineering practices. And it provides measurable and actionable results, that are easy to convert into an improvement plan for a particular repo.

#ai #engineering

Читать полностью…

TechLead Bits

Skill Packaging

Engineering teams are actively building internal collections of skills for agents: code review, troubleshooting, design preparation, onboarding, security practices.

And it looks great until you hit the question: how do you distribute those skills across dozens of teams and multiple harnesses? For Claude you need to put skills into .claude, for Cursor into .cursor, for Gemini into .gemini, etc. And things become even messier when you need to roll out updates.

To solve this problem big companies mostly build their own in-house solutions. Smaller companies usually just copy files from some shared repository and manage this complexity manually.

I don’t like reinventing the wheel, so when my team faced the same problem, we started looking for an existing solution we could reuse. And the only actively maintained tool we managed to find was apm by Microsoft.

APM is a package manager for prompts, skills, and MCPs. In other words, it’s maven or gomod for agents.

APM package structure:

my-package/
├── apm.yml
└── .apm/
├── instructions/
│ └── my.instructions.md
├── skills/
│ └── my-skill/
│ └── SKILL.md
├── agents/
└── prompts/

To install the package in a target repo you need to define apm.yaml with a list of required dependencies:
name: my-projecty
version: 1.0.0
targets:
- claude
- copilot
dependencies:
apm:
- <git-address>/my-package
- <git-address>/another-package
mcp: []

After that you just run:
apm install

and the required skills will be installed into the corresponding harness folders (.claude, .copilot, etc.).

The tool works with Github and on-prem git installations like Gitlab.

APM is not perfect. It had some unpleasant (but not critical) issues, and sometimes you can really feel that it was heavily vibe-coded in Python .

But despite all that, the tool actually works: you have a spec to define your skills\prompts packages, distribute and update them with simple apm update. And on top of apm dependencies format it’s pretty easy to vibe-code your own internal skills marketplace.

#ai #engineering #agents

Читать полностью…

TechLead Bits

AI Engineering

I strongly believe that if you want to use any technology effectively, you need to understand how it works under the hood. Especially in software engineering.

So if you haven’t looked into LLM internals yet, I’d highly recommend reading AI Engineering by Chip Huyen. The book was published in December 2024. And as AI is moving extremely fast, you might think it’s already outdated. Yes and no.

The book focuses on fundamentals. And they don’t really change that fast. You won’t find hype topics like skills, harnesses, or agents orchestration there. But for building structured understanding of how AI works, you don't actually need them.

What I personally found useful:
🔸 Core LLM concepts: tokenization, training and post-training processes, datasets preparation. This part is very similar to Mashing Learning Crash Course from Google.
🔸 Model evaluation: quite complex but interesting topic about model output results and their comparison. The book covers ranking, model specialization, public benchmarks and AI-as-a-judge approach.
🔸 Prompt engineering: good reference about context and prompting. Additionally, the author described different security aspects of using prompts, that part really extended my thoughts about what can go wrong.
🔸 Finetuning: a deep dive into different ways to optimize models. You need to be a good mathematician to understand this part. So I was really glad I'm not an ML engineer 😃 (huge respect to all ML experts, it's really hard).
🔸 User feedback: basic patterns on how to collect feedback, what to measure and why, common pitfalls.

To sum up, this book is really great to structure your knowledge about modern AI systems. Once you have that foundation, it becomes much easier to navigate all the new tools, patterns and paradigms that appear almost every month.

#booknook #ai #engineering

Читать полностью…

TechLead Bits

Inside the Context Window

What makes your work with agents efficient? Chosen model? Harness? Instructions clarity?
I would say that first of all it's the quality of the context you provide.

Context is everything the model sees before it generates a response.
Two facts to know about the context:
1. It's limited (and costs you money 💰 ).
2. The longer the context, the worse the results.

So context engineering is a set of practices to fill the context with just enough information to get the desired results. The main goal is to balance the amount of context given: not too little and vague, not too much and detailed.

The first step in context engineering is to understand what the context actually contains. And it’s not just your prompt.

Typical context structure:
🔸 System prompts & instructions: the hidden layer of system prompts, safety policies, behavioral rules, role definition. Usually it's part of the harness and you cannot change it.
🔸 Project context: AGENT.md\CLAUDE.md, repo structure, settings. It's added as a first prompt to any session you open with the agent.
🔸 Available tools: skill descriptions, MCPs, available CLIs.
🔸 Retrieved information: loaded files, data from RAG system.
🔸 State & history: The current conversation, including user, model and tools responses.
🔸 Reasoning: intermediate reasoning results (thinking mode).
🔸 Long-term memory: knowledge base from previous conversations like user preferences, summaries of working sessions, facts the agent was asked to remember for future use.
🔸 Your prompt: the actual user request.

As you can see, the context is already filled with a lot of information before you even start the real work. To make agents efficient, keep their context clean and focused. Don't overload it with unnecessary information.

#ai #engineering

Читать полностью…

TechLead Bits

Agent Harness

Harness is a new buzzword introduced by modern AI.
Let's check what it is and why it matters.

The term harness refers to the logic around LLM that controls and guides how an agent operates. It's not the agent itself but the tools and guardrails that help it achieve better results.

A harness typically includes:
🔸 System prompts
🔸 Tools, skills, MCPs and their descriptions
🔸 State & memory (current task state, past runs, intermediate states)
🔸 Planning & task decomposition
🔸 Context engineering strategies
🔸 Safety & guardrails (allowed tools, rate limiting, prompt injection protection)
🔸 Bundled infrastructure (filesystem, sandbox, browser)
🔸 Subagent orchestration logic
🔸 Hooks/middleware for deterministic execution (compaction, continuation, lint checks)

Well-known examples of harness ecosystems include Claude Code, Cursor, LangChain.

The overall trend is that each model provider now builds and promotes its own harness. But because each provider uses different system prompts, model tuning techniques and context management strategies, the same model in different ecosystems will produce different results.

So the same model does not mean the same agent. And the real competition is no longer between models. It’s between harnesses fighting for your workflow and your budget.

#ai #engineering

Читать полностью…

TechLead Bits

Why Software Factories Fail

"Read the Code!" is one of the key ideas from Dex Horthy's talk "Why Software Factories Fail".

The author touches on a very hot topic right now: we are actively pushed to put AI-generated code into production, build software factories, and eliminate the human bottleneck from the process. And all of that would be great if, at the same time, the overall quality of software products wasn't going down and the number of incidents wasn't going up.

Dex explains this phenomenon by pointing to a fundamental limitation of current models: they can't maintain and improve the quality of a codebase over time. Models are trained to solve the issue and pass the tests. There is nothing there about good architecture, clean code, or future maintainability.

If the model knew what good code looks like it would write it in the first place.

What to do with that? Dex suggests to do the following:
🔸 Make Product Review: align goals, check mockups, etc.
🔸 Do System Architecture: establish clear component boundaries and API contracts.
🔸 Perform Program Design: define types and method signatures, test approach, program layout, and dependency stack.
🔸 Implement by Vertical Slices: build features vertically, reducing the scope of what an agent can change in a single iteration.

So, it's all about shifting review left to earlier stages and focusing more on architecture and program design (I described similar ideas about code review here). And then let agents make small changes that engineers can actually understand and review.

And yes, we still need to review the changes.

#engineering #codereview #ai

Читать полностью…

TechLead Bits

The Culture Map

Have you ever worked in international distributed teams? Or collaborated with teams from other countries? Then you probably noticed that it can be quite challenging.

I didn't think much about that before reading The Culture Map by Erin Meyer. I saw some difficulties in my international teams and put a lot of effort into making it work. But after reading the book, I realized that some of my efforts were just fighting windmills.

So what is this book about?
It's about how cultural patterns impact our behavior and why sometimes it's so difficult to understand each other, even when we use the same language (like English for business). The language is the same, but the meanings can be very different.

The author defines 8 scales to measure the difference:
1. Communicating: low context vs high context. Low-context means to be explicit and clear; where high-context means to be implicit, expect to read between the lines.
2. Evaluating: direct negative feedback vs indirect negative feedback.
3. Persuading: principles first (theory before examples) vs application first (examples before theory).
4. Leading: egalitarian (small distance between boss and employee) vs hierarchical (strong vertical and status).
5. Deciding: consensual (based on group decision) vs top-down
6. Trusting: task-based (trust through work result and competence) vs relationship based.
7. Disagreeing: confrontational (open disagreement) vs avoid confrontations (indirect disagreement)
8. Scheduling: linear time (plans and strong deadlines) vs flexible time (no strict plans in changing circumstances).

A good example is the US and Japan. The US is a low-context culture, where things are usually described and explained explicitly. Japan is a high-context culture, where you are expected to read between the lines. This difference can easily lead to frustration and misunderstanding.

The general rule is to compare cultures relatively. What matters is not whether a culture is “direct” or “hierarchical,” but whether it is more or less so than the culture you are comparing it with.

But what to do with all of that? First, this knowledge can help you better prepare for negotiations. Second, the author recommends that leaders be more flexible: understand what leadership style their team expects and adapt to it, taking cultural differences into account.

I really liked the book. It’s clear and practical, with good examples and useful recommendations. I started to understand many things better. I’d highly recommend it to anyone who works in an international company or with international partners.

#booknook #softskills #leadership

Читать полностью…

TechLead Bits

Loop Engineering

Over the past year, AI has been constantly bringing new terms and practices into our work. Yet another XX Engineering probably doesn't already surprise anyone.
Let's take a look at Loop Engineering, which is hyped right now.

The main idea is that an engineer no longer prompts an agent directly. Instead, they design a system that does it on its own. A loop here can be thought of as a recursive goal: you define the purpose, and the AI iterates until complete.

A typical loop consists of the following steps:
🔸 Discovery: Define what should be done during the iteration. The key is letting the agent find its own work rather than providing it manually.
🔸 Handoff. Move the task from the scheduling system into the hands of the agent that does the work.
🔸 Verification. Check whether the result satisfies the goal. This is what prevents the loop from blindly moving forward.
🔸 Persistence. Save the result somewhere it can survive between sessions and agents, e.g., PR, issue tracker, database, etc.
🔸 Scheduling. Define when the loop should run. For example, an automated morning CI triage.

Building blocks for the loop:
🔸 Automations. Automated procedures or triggers that start the loop.
🔸 Worktrees. Built-in git mechanism for multiple independent working directories in one repo to allow agents to work independently.
🔸 Skills. Project knowledge and instructions.
🔸 Plugins\Connectors. Tools (usually MCPs) that give agents access to the outside world: issue trackers, databases, APIs, and other systems.
🔸 Subagents. Separation of tasks and responsibilities. For example, one agent does the work while another verifies it.
🔸 Memory. Long-term storage for task state, execution results, and lessons learned.

The overall concept is pretty cool, but it requires very strong engineering discipline.
The cost of an incorrectly designed loop is much higher than the cost of a bad prompt. In a loop, an error can accumulate with every iteration, the result can gradually drift away from the goal, and a lot of tokens can be just wasted.

That's why verification becomes one of the most critical parts of Loop Engineering. If we want autonomous loops to produce stable results, they need a reliable way to verify their own work.

For a deeper dive:
- Getting Started with Loops Claude blog
- Loop Engineering: The Anthropic Playbook for Designing Systems That Prompt Your Agents
- Loop Engineering by Addy Osmani
- Practical Loop Engineering by Addy Osmani

#engineering #ai

Читать полностью…

TechLead Bits

The AI-Driven Leader

"Our humanity lies in our ability to think strategically, be creative, and communicate and collaborate to solve complex problems... AI presents a unique opportunity for us to reclaim these human strengths and have machines adjust to meet our needs."

This is how The AI-Driven Leader by Geoff Woods begins. An international bestseller with a catchy title and a good rating on Amazon. Actually, I had pretty high expectations about it.

But let's start from the beginning.
This is a book written by a non-engineer for non-engineers. It explains in simple terms what GenAI is, what risks it brings, and how leaders can integrate AI into their work.

If you boil it down, the whole idea is to use AI as a Thought Partner (I wrote about it here). To demonstrate that the approach works, he provides simple real-life stories and examples of specific prompts. If you don't know where to start with AI, the author suggest 5 main use cases: strategic thinking, decision-making, content creation, idea generation, and analysis.
And that's basically where the meaningful AI-related content ends.

The remaining 80% of the book is mostly basic leadership recommendations that I would summarize as "do the right things, don't do the wrong things". Ideas like "leaders should focus more on strategy and less on operations" don't really have much to do with AI.

Overall, I was left with mixed feelings.
On one hand, the book is quite entertaining, easy to read, it provides a simplified introduction to GenAI and explains what leaders can actually do with it.
On the other hand, it feels like a mix of generic leadership practices with AI that sometimes added just because it's a hot topic. At the same time, the author is extremely positive about the solutions AI produces. And as engineers, we know that something that looks like a convincing result at first glance can be AI slop on the second.

From my perspective, the book can be useful if you're only starting to use AI at work and don't know where to begin. Otherwise, you can safely skip this one and spend your time reading something more useful.

#booknook #ai #leadership

Читать полностью…

TechLead Bits

AI Adoption: Measuring Support

Continuing the previous post, let's look at how AI adoption can be measured for support teams.
The principles are exactly the same: cost, quality, and time.

Here are the metrics I'd track:
🔸 Team Budget. Same as for development teams: the cost of the team in $, mandays, or FTEs.
🔸 AI Cost ($). How much does the team spend on AI, including autonomous agents if you're using them.
🔸 Incoming Load. The number of incoming tickets. At the beginning, this metric probably won't change much. But in the future it can show whether the overall quality is improving or getting worse. I recommend tracking it as a percentage change from the baseline before AI adoption.
🔸 Backlog Size. The number of non-resolved tickets. This metric should always be viewed together with the team budget and SLA.
🔸 Time to Resolution. How quickly customers receive a solution to their problem.
🔸 Reopen Rate. Has the percentage of reopened tickets increased? I've seen exactly this happen when teams introduced AI-based ticket assessment. Resolution time went down, but the number of reopened tickets increased several times. A clear sign that support quality had dropped.

As you can see, none of these metrics are AI-specific. Most tracking systems can calculate them out of the box, while a few may require collecting and analyzing historical data.

And one final point: never look at these metrics in isolation. Resolution time is decreased -> reopen rate doubles, team budget decreased -> backlog increased -> resolution time increased.
So it's the combination of metrics that tells you whether AI is actually improving support process or not.

#ai #engineering #ai4sdlc

Читать полностью…

TechLead Bits

AI Adoption: What to Measure First

In the previous post, we looked at how DORA recommends measuring the ROI of AI adoption.
But what if you don't know the revenue generated by your features or other business metrics? Yet you're still expected to show the effectiveness of AI adoption.

Let's bring the DORA approach down to the engineering team level.

I would split AI adoption into two phases: AI adoption itself and getting business benefits. These phases have different goals, metrics and outcomes.
It doesn't make much sense to measure business impact of AI until the team has actually gone through the adoption phase.

What to measure to understand the AI adoption state:
🔸 Active users. How many engineers actually use AI in their daily work?
🔸 Monthly usage per user. How actively do engineers use AI tools? Token consumption or API cost per engineer can be a good indicator. AI Gateway solutions such as LiteLLM can help collect these metrics.
🔸 AI-generated code ratio. What percentage of code is generated by AI? Important note: this metric measures only how widely AI is being used. It says nothing about whether AI is being used effectively or producing good results.
🔸 AI-assisted feature ratio. How many features are developed with AI assistance? The metric counts if AI was used in design, implementation, testing, documentation, or code review.
🔸 Engineering readiness. Is the engineering ecosystem ready for agents? Here I would use criteria similar to an Agent Readiness Framework: documentation, AGENTS.md, test coverage, CI stability, guardrails, and so on.

Giving developers AI licenses does not mean AI has been adopted.
Not after one month. Not after three. Not after five.

AI adoption happens at the team level. The goal of all these metrics is to understand whether the teams actually change the way they work. And it requires training people, overcoming resistance, changing engineering practices, and adapting the development process.

In the next post, I'll share my approach to measuring AI impact during the second phase using only tools that almost every engineering team already has.

#ai #engineering #ai4sdlc

Читать полностью…

TechLead Bits

ROI of AI-Assisted Development

One of the biggest questions around AI adoption is what business value does it actually bring?
I'm currently introducing AI into the SDLC across my teams, so measuring its real impact is something I'm really interested in.

Back in February 2026, DORA published a dedicated report on this topic: ROI of AI-assisted Software Development report.

Key takeaways:

🔸 Measure your baseline first. Before evaluating the impact of AI, you need to understand where you're now. Without "before" metrics, it's impossible to measure "after".

🔸 Expect a J-curve. Most organizations experience a temporary productivity drop before seeing benefits. Main reasons for that:
- The learning curve. Time to master new tools and workflows.
- The verification tax. Time to verify AI-generated code and establish trustworthiness of agents output.
- Pipeline adaptation. Scaling other downstream processes like testing and change approval to handle increased number of code changes.

🔸 Invest in engineering practices. The quality of an internal platform, clear workflows, reliable test automation, and strong engineering culture become critical to produce predictable delivery. The rule garbage in -> garbage out is still actual.

🔸 Measure both positive and negative outcomes. Higher throughput is valuable only if it doesn't come with higher instability, more defects, or lower quality.

🔸 Measure the following business values:
- Cost efficiency
- Productivity
- Developer experience
- User experience
- Business growth
The report provides guidance on how to measure these areas and combine them into an ROI calculation. It also includes an online calculator where you can check your own numbers.

🔸 AI adoption takes time. According to the report, the average adoption journey takes around 8 months, in large enterprise rollout can take even 12–18 months.

Overall, the report is great: no hype, just a practical and rational approach.
It also strongly aligns with my own view of AI adoption. You can't simply buy AI licenses for developers and expect the investment to pay for itself.
Successful AI adoption requires improvements across the entire SDLC: better development processes, a stronger engineering culture, higher test coverage, and investments in guardrails and workflows that keep software delivery stable and predictable.

#ai #engineering #ai4sdlc

Читать полностью…

TechLead Bits

Akrites Project

Over the past year, we've seen how much the industry depends on small open source projects. In many cases, software used by thousands of companies is maintained by one or two people working on it in their free time.

The release of the Mythos and Fable models made the situation even worse. They demonstrated how many vulnerabilities AI can find across thousand of projects, putting the whole enterprise IT infrastructure at risk.

Last Thursday (June 25), the Linux Foundation and a set of big IT companies announced a new initiative called Akrites, "a coordinated effort to remediate and disclose vulnerabilities in critical open source software."

The goal is simple: bring engineering resources together to identify, fix (!!!), and responsibly disclose vulnerabilities in critical open source software before they can be exploited.

If a package has no active maintainer, Akrites will serve as a "maintainer of last resort so fixes to the latest version reach everyone in a timely fashion".
And of course, they will do that with the help of AI frontier models.

The list of members is impressive: AWS, Anthropic, Cisco, Google, Microsoft, NVIDIA, IBM, OpenAI, Red Hat, and many others. It's one of the largest coordinated industry initiatives we've seen in recent years.

Personally, I think this is a very important shift. For years, the biggest challenge was finding enough people to make security fixes on time. AI has made the situation much more critical by flooding maintainers with newly discovered vulnerabilities.
The great thing about this initiative is that engineering resources are finally being focused not on reporting vulnerabilities, but on fixing them.

So the idea looks promising. Let's see how well it works in practice.

#news #opensource #security

Читать полностью…

TechLead Bits

"AI won't take your job. Someone using AI will."

This quote caught my attention and made me watch A Leader’s Guide to Advanced Team Structures in an Agentic World from the recent AWS Summit Sydney.
It's a very sobering talk on the current state of the industry, AI adoption, and the future of engineering roles.

The central question of the talk is: "How should we build teams to work in this new AI world?"

To answer it, the author proposes a framework based on four elements:

Economics
The market has changed. Timelines are compressing. A small team of senior engineers can replace an entire existing product. This creates real risks for businesses that fail to adapt in time. That's why AI decisions should be driven by economics and business value, not hype.

Talent
Previously, career growth in tech was mostly about writing code and building features. Today, the most valuable skill is understanding the business, customers, and product. In other words, AI rewards expert generalists. One person can now handle analysis, backend, and frontend, reducing collaboration overhead and the need for deep specialization of large teams.
Another interesting point is the future of junior engineers. The speaker argues that we must keep the junior pipeline alive. Otherwise, we won't have senior expertise in 2034.

Structure
Current IT operations are optimized for determinism. But agents are non-deterministic. So operating model has to shift: variance in execution, focus on outcome and guardrails around thing you actually care about. The best operational model there is platform engineering.

Governance
The author highlights several areas that organizations need to address:
- Agent Identity Management. Every agent should have a verifiable identity traced to a named human.
- Risk Assessment. Clearly define what an agent is allowed to do and ensure it operates within those boundaries.
- Multi-agent Coordination. Control what happens when agents disagree, escalate, or find emergent behavior we don't expect.
- Deskilling Prevention. Employees should maintain core skills even if agents automate routine work. Someone still needs to validate results, audit actions, and take responsibility for decisions.

Overall, the talk is a good reality check on what is actually happening and how businesses and teams need to change to remain successful. Much of it resonates with my own observations, so I would definitely recommend to watch the full video.

#ai #leadership #engineering

Читать полностью…

TechLead Bits

How Anthropic Writes Skills

Last week Anthropic published lessons learnt of how they build agent skills internally. It's quite interesting to read recommendations from the company that introduced the concept in the first place.

Key ideas:
🔸 Don't be obvious. Model already knows how to code. A skill should provide instructions that change default agent behavior, not repeat the data the model was trained on.
🔸 Build a gotchas section. Add common mistakes and lessons learned. This helps the agent avoid repeating the same failures.
🔸 Use progressive disclosure. A skill is not just a SKILL.md. It can include additional files that are loaded on demand, reducing context overload.
🔸 Don't be too specific. Give the agent information it needs, but leave the flexibility to adapt to the situation.
🔸 Separate configuration from instructions. Store setup data in config.json or collect required input from the user.
🔸 Write description for the model, not for humans. A description should help the model to understand when the skill should be invoked.
🔸 Use long-term memory. Skills can maintain their own data in a subdirectory and reuse it across executions.
🔸 Automate where possible. Not everything should be a prompt. Some actions can be automated with helper scripts and functions.

Unfortunately, the article doesn't provide any guidelines of how to evaluate skill effectiveness. It's still not clear how to understand if a skill actually works, how to compare two versions of the same skill, or how to detect that a skill is no longer useful.

Provided recommendations are based mostly on observations of how popular internal skills are structured. It's useful but not measurable.

So despite the fact that skills are the most powerful agent extension now, evaluating them remains on of the hardest engineering problem.

But that's another story.

#ai #engineering

Читать полностью…

TechLead Bits

The Fearless Organization

The most dangerous teams are the quiet teams. There are no disagreements, no bad news, no conflicts. Looks like harmony until the real incident.

That's the topic of the book The Fearless Organization: Creating Psychological Safety in the Workplace for Learning, Innovation, and Growth by  Amy C. Edmondson.

Amy is a professor of Leadership and Management at the Harvard Business School. She has studied the phenomenon of psychological safety and its impact on team performance for many years across different organizations.

She defines psychological safety as follows:

a belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes, and that the team is safe for inter-personal risk taking.


What it means in practice:

- people are not afraid to ask questions
- they do not hide problems
- they are not afraid to look stupid
- they do not avoid conflicts
- they can freely express their opinions and bring suggestions

Why does it matter?
There is a good example from the book that explains that. Imagine a doctor prescribes treatment for a child. A nurse notices that doctors usually prescribe drug A in such cases, but this time it is missing.
In a team with high psychological safety, the nurse will clarify this with the doctor and may help prevent a medical error.
In a team with low psychological safety, she may be afraid to ask. And the consequences can be dramatic.

The core idea is simple. But the book contains a lot of real stories where a low level of psychological safety leads to dramatic results (e.g. Volkswagen emission scandal, pilot mistakes that caused plane crashes). The author repeatedly highlights that the more complex and critical the profession is, the more important psychological safety becomes.

How does this relate to our daily work?
We as leaders are responsible for the psychological climate in the team: how well we listen to people, accept different opinions, react to questions, mistakes, or bad news. It's our daily routine that either helps the team become more effective, or leads people to hide problems and the real state of things.

Overall, I really liked the book. It explains the idea in simple language with many real examples. And what is important for me, all arguments and recommendations are supported by sociological research, experiments and practical psychology.
So psychological safety is not just an idea. It is a proven behavioral model and set of practices that can actually help leaders build better teams.

P.S. One of the best real examples of psychological safety is Pixar. I wrote about it earlier in overview of Creativity, Inc.: Overcoming the Unseen Forces That Stand in the Way of True Inspiration: parts 1,2,3.

#booknook #softskills #leadership

Читать полностью…

TechLead Bits

Ralph Loop

Underterministic nature of AI sometimes produces very interesting engineering solutions. One such example is Ralph (or Ralph Wiggum) Loop.

This is an AI coding pattern inspired by The Simpsons character Ralph Wiggum, known for saying weird things with high confidence 🙃.

The idea is simple: The agent can be dumb in a single iteration. But if it keeps retrying with feedback long enough, it eventually converges.

The loop steps:

Start a new agent\subagent -> load task + memory -> execute 1 selected task -> run validation -> save learnings -> commit progress -> repeat

The loops finishes when all tasks have passes:true or it reaches the maximum number of iterations (default is 10).

But the real value of the technique is not in retries, it's in context engineering strategy under the hood:
🔸 One loop executes only one task. It keeps agent focused.
🔸 Each iteration starts a new agent session. State lives outside the context keeping it clean between iterations. State is stored in git history, progress.txt and long-term-memory files.
🔸 Tasks are delegated to subagents. The main context is not polluted with task execution details and validations.
🔸 AGENTS.md is updated on each iteration. It is a live artifact that contains discovered patterns, learnings and conventions so future iterations can benefit from those findings and do not repeat previous mistakes.
🔸 AGENTS.md contains explicit validations for feedback loop. It usually defines linters and typechecks, build and test execution commands.

Ralph Loop is a really powerful pattern to get things done: it just repeats the task until it succeeds making agent execution more reliable. "Deterministically bad" but effective.

But this approach only works if you have good task decomposition, clear completion criteria, and mature SDLC practices with strong validations and feedback loops. Otherwise the agent will generate just a ton of mess.

#ai #engineering #patterns

Читать полностью…

TechLead Bits

ReasoningBank

Currently AI agents have one major limitation: they cannot learn. I mean they don't learn from their experience or from the results of completed tasks. Once the model is trained, all we can do is to tune our prompts or enrich results with domain data from RAG.

Researchers from Google started exploring how to overcome this limitation and introduced the concept called ReasoningBank.

The overall idea is simple:
1. The agent writes down the result of successful or failed tasks into a dedicated md file.
2. During task execution, the agent searches the ReasoningBank and pulls relevant memories into the context.
3. Then it uses an LLM-as-a-judge approach to self-evaluate the result, analyze the trajectory of reasoning, and extract success insights or failure reasons.

Each file has the following structure (very similar to skills):
- Title: identifier of the core strategy.
- Description: short summary of the memory item.
- Content: reasoning steps, decision explanation, or operational insights extracted from past experience.

To be honest, benchmark results compared to other agent memory approaches do not look extremely impressive:

ReasoningBank without scaling outperformed memory-free agents by 8.3% on WebArena and 4.6% on SWE-Bench-Verified.


At the same time, this approach adds even more data to the context. And context, as we know, directly affects both model behavior quality and usage cost.

The official paper contains interesting research details, including particular prompts and measurements.

From my perspective, the idea and its implementation are very similar to skills or other long-term agent memories (e.g. in Claude Code). But the overall direction of making agents capable of learning from their own experience looks really promising.

#ai #engineering #news

Читать полностью…

TechLead Bits

CliftonStrengths 34

I recently passed the CliftonStrengths 34 assessment, so today I will share what it is and how it can be useful for your career.

CliftonStrengths is a framework that identifies your natural talents that help create value at work.
It was launched by Gallup in 2001, and since then more than 26 million people have taken it.
So it's based on real data and many years of consulting experience.

The main idea is that we should focus on our strengths to achieve results and not try to improve our weaknesses.

How it works:
- 200 questions
- 4 strength domains: executing, influencing, relationship building, and strategic thinking.
- 34 strength areas within those domains.
- All 34 themes are ranked in a personal order from the strongest to the weakest.
- Top-10 are our main talents to focus on.

The interesting part is that every strength has both a positive and a negative side. It can help you succeed or hold you back.
For example: Learner. The person with this strength quickly picks up new topics, constantly extend their knowledge. But it's easy to get stuck in a “forever student” mode.

The test is really helpful from self-reflection perspective:
🔸 Once you know your strengths, you can rely on good parts and mitigate the downsides.
🔸 The less you do the work that isn’t natural to you, the more productive and energized you are.
🔸 We assume others think and work like we do. But they don't. And this is our advantage to use.

What it gave to me? First of all, I realized that I really do the work I'm naturally good at (hello, imposter syndrome). Second, I started noticing my strengths in real situations and using them more consciously.

So the assessment is a helpful tool to understand what to focus on to achieve better results. And it’s a good starting point to rethink your day-to-day activities and align it more with what you enjoy.

#softskills #leadership #productivity

Читать полностью…

TechLead Bits

Building AI-Powered Team

AI adoption is one of the biggest challenges and at the same time one of the biggest opportunities for business.
Some teams report significant productivity boost. Others say: “AI might be useful for some tasks.”

So what’s the difference?
AI adoption is not about access to the tools. Just buying licenses and giving them to engineers doesn't work. The team need to rethink how they work and integrate these tools into their daily workflow. And that’s already classic change management task.

On this topic, I recently came across the GitHub Internal Playbook for building an AI-powered workforce. They highlight that AI adoption is not really a technical problem, it’s a human one.

GitHub suggests 8 pillars to drive adoption at the organization level:
- AI advocates. Internal champions who scale adoption through peer-to-peer influence and feedback.
- Clear policies. Defines rules for using AI.
- Learning & development. External and internal training and education.
- Metrics. Track adoption, engagement, and business impact.
- Ownership. A central owner who orchestrates the program and drives the overall strategy.
- Executive support. Visible leadership commitment and strategic vision.
- Right tools. Different tools for different roles.
- Communities. Peer-to-peer learning, knowledge sharing, and collaborative problem-solving.

And in reality, the key part here is the people on the ground, the experts who drive the change, adapt the tools to real tasks, and teach others. This is also covered in more detail in the companion article Activating your internal AI champions.

You can’t roll out AI top-down.
You can’t standardize it with one template for everyone. Every team has its own context. Without understanding it, any “unified approach” will fail.

#leadership #ai

Читать полностью…
Subscribe to a channel