Viktor 00:00:00.315 So we, we have two problems over here. First, agents do bad work because they have no idea who you are, what you want, and how you want it, and so on and so forth. That's the first problem. And the second problem is that once they find out what you really want, then they do a really bad job because what you want is stupid That was the initial conversation, right? It amplifies bad
Darin 00:01:26.217 Viktor, what do you think of this? A bad engineer with AI-assisted development is much, much, much, much worse than just a bad development engineer
Viktor 00:01:35.677 Okay. Yeah, of course. have… I think that bad engineer with a laptop is much, much worse than a bad engineer with a calculator. Yes
Darin 00:01:44.197 You have companies that are laying off people because the roles weren't AI enough for the company
Viktor 00:01:51.009 I feel that there are two big cases for layoffs. One is we need much smaller workforce to do what we're doing. That's probably the bigger one, and it's completely wrong. And another one is we're actually going to do so much more. We shouldn't off people. But we cannot get there with the people we have So it's not necessarily always layoffs, but maybe replacements disguised as layoffs And then you if that's your case, you need to think very hard about what you're doing
Darin 00:02:22.600 Because do you even have a, grounded foundational business?
Viktor 00:02:26.000 Yeah, I was now more referring to the person being laid off only to be replaced next week with another person that happens to be less skeptical about AI. There is a lot of skepticism right now, a lot of not good enough And then what do you do as a company who chose to go all, all in on AI? I get rid of this person and I get the enthusiastic one
Darin 00:02:50.903 That seems like it would be an answer, but DORA has sort of disproved that. Yes, that same DORA that's given us the DevOps reports for years and years. Back in May of '26, they released their ROI of AI-assisted software development report.
Viktor 00:03:07.271 Okay
Darin 00:03:07.893 If I remember, I will put the link to that in the show notes so you can go download it and read it for yourself. But what they found and what they're saying is AI without engineering excellence just scales your problems That makes too much sense
Viktor 00:03:26.258 I was afraid for a second that because I haven't read the report, shame on me, that sentence would start with whatever it started and end up with kind of no return of investment on AI, but that's not the case. It's k- no return of investment without people really knowing what they're doing, right?
Darin 00:03:44.788 I'll restate it. AI without engineering excellence just scales your problems
Viktor 00:03:51.520 Correct. You get more of everything
Darin 00:03:54.008 Whether you want it or not
Viktor 00:03:55.850 Or not, actually no, that's wrong. It's not more of everything. It's you get more of what you were doing And I'm sorry to say that, but there is a statistical chance that what you are doing is bad
Darin 00:04:09.470 Because y- you've made foundational choices years ago that you're still trying to… Okay, here's my example. You're trying to take your monolithic Java web app and just run it inside Kubernetes
Viktor 00:04:22.540 Exactly. And now you can have gazillion more Litic DevOps running in Kubernetes. Hooray. Well done
Darin 00:04:30.742 And your on-call staff is going to be on call 24/7
Viktor 00:04:34.662 Yep.
Darin 00:04:35.240 dealing with incidents because you did not think about what it takes to run application X inside of Kubernetes because you just said, "Hey, Kubernetes is the hot thing. That's what we're going to do." And now you gotta deal with that fallout
Viktor 00:04:52.150 that raises what I believe is an interesting question. Imagine before AI, you would go and spend some time with some team, a company, and they would tell you what you're doing, what they're doing, and you would say, "Yeah, but why are you doing that?" And they would probably answer eventually with, "Yeah we just kind of-- We don't have time and money to actually do it the right way now." And that could be true. But if you're still doing it that way, and now that time is not a problem anymore that might lead me to conclude that what you were saying is that you were fooling yourself and lying to me. It's not really that you don't have time. You just don't know how to do it right And now I have proof because now there is nothing stopping you to do it right
Darin 00:05:41.900 Yeah, there's a variation of what you're saying from the report. This is summarized. Productivity isn't the problem. The problem is that organizations without mature platforms, clear workflows, and strong reviews amplify chaos when they turn on AI
Viktor 00:05:59.750 Yes So
Darin 00:06:01.160 What was good is better and what was bad is worse
Viktor 00:06:05.300 Yeah or maybe restate it with what was good. Now, Now you have more of what was good and more of what was bad. And if you had more bad than good, then, both are amplified proportionally
Darin 00:06:20.890 If we think about it, if we have the elite performers, I imagine there are certain companies that are elite performers worth-- with AI. My guess is those are gonna be the ones that were not the command and control top-down organizations. Would you agree or disagree with that?
Viktor 00:06:38.691 those must be organizations where individuals had a lot of autonomy. I strongly believe that. Because now we're moving to the world where actually I alone do the work of a team, right? It's almost pointless now for us to work as a team, Where, "Oh, I will do this part and you will do this part," and so on and so forth. Kind of that does not make sense because you're telling that to your agents now, not your colleagues. And that means that this new world favors people who were autonomous before
Darin 00:07:12.123 because we were already used to working that way
Viktor 00:07:14.523 Yeah, kind of okay. It might take me infinite amount of time, but I'll just do it, right? If that was your attitude and not "Oh, I'm go- I'm do it, and now I'm waiting for Michael. Michael knows HTML. When Michael is finished, then I will continue my part." If that's what you were doing before, now that is amplified, right? Th- that does not simply work because you don't know Michael happens to be a bot, an agent, r- and you don't know what to ask him or her. I don't know whether it's she or he agents it. So yeah this whole setup now favors companies that were looking for and enabling people to have great autonomy.
Darin 00:07:55.896 And just like any new tech, the Dora report found that there was like this big J curve. So in the first one to three months, you took a dip in productivity,
Viktor 00:08:05.236 Oh yeah
Darin 00:08:05.852 but then it usually took off from there. But the problem is as things take off, we've talked about this, then it's dealing with reviewing of that, reviewing the code, reviewing the documentation, reviewing the thing that the human needs to, using your word, add taste to. That's something that the machines can't do for us yet
Viktor 00:08:29.302 Yes. It, it somehow reminds me of testing in the past where many companies or teams would-- they would keep some tests and tests are almost always flaky. I rarely see tests that are reliable, And what do you do when you have flaky tests? You run it again. And you run it again, and eventually it will pass, and you say, "Well done." Or you have teams who would be improving those tests. Say, "Okay, actually, we're not going to rerun it. We're going to figure out why it's flaky. We're going to improve it, and the next time it runs, it's going to be less flaky." that's the approach we need to take with agents, It's going to create a mess, and you're going to figure out one of many reasons why it's a mess, and you're going to extend system context or skills or MCPS or whatever you're using for the next time. And you're going to do it over and over again. And as you're improving agents, because agents are actually agents plus LLMs, they're actually pretty pretty good at what they're doing. They just lack the context of you, of your organization, of your choices, and so on and so forth. And you cannot provide that context from the first attempt, but you can over time. And that means that over time, there will be fewer bad things, fewer false positives, and the time it takes you to review it will shorten Apart from using agents to review as well. It's for shorten because, okay, so those are the things that after 57th iteration of improving what's or not now it's d- doing it cons- consistently right. We don't need to focus on this anymore We can now remove this part of review and that will give us more time to focus on other parts of the review Which I feel again, story that we should have had maybe we did before agents. Nothing new
Darin 00:10:26.466 Nothing new. So with the DORA report, one word they used, and I've already used it here, is amplifier. What is an amplifier? If you think about it from an audio perspective, it takes a signal and it makes it louder, right? That's the basics of it. Unless you take it to 11, then it's distorted. Spinal Tap reference. That's the problem, though, is you have to be careful with an amplifier because if you do turn it too far up, you can have problems. Feedback in not the good sense. But let's think about the elites and the not elites process here again. So let's spell it out. For the elite teams, they already have optimized testing, great CI/CD, platform engineering, clear processes, AI accelerates that implementation, everything's glorious. But for the struggling teams, backlog debt, slow CI/CD, fewer safeguards, unclear workflows, AI is pumping out code left and right, the pipeline chokes, the meantime to resolution just goes through the roof, if it's even up. And now the senior engineers that were saving your butt in the past are now the bottleneck ' Viktor: think about it this way, what does it… So let's say that you have a code base, whatever the code base is, and you instruct the agent, "Okay, I would like you to fix this issue or add this feature," or whatever that is. And I'm going to ignore system context now, I'm going to ignore skills and all the stuff that I mentioned earlier, right? You just say "Develop this new feature." What will it do? It will not randomly develop feature as you specified. It will today, not necessarily a year ago, it will actually analyze and see how you're doing things, and it'll do it more or less the same way as you're doing it And what, what that means is that you, if you had I want to use the word I'm not supposed to, bad code is going to do more of it because it said, "Oh, this is what you're doing? Well done, Viktor. Let me do more of it because you're always right That's wonderful and scary all at the same time
Viktor 00:12:31.951 actually, wasn't it one of our complaints they're doing random things and they're increasingly doing the fewer random things? they're they're taking over what you were doing in a way you were doing it. You have no CI. We don't need it, right? You have flaky tests. Let-- If you have tests, cool. Let me see how you did previous tests like this. Excellent. I'm going to do more of those. And if they were flaky, now I have more flaky tests It was okay to run for seven hours. What's five minutes added to seven hours? Nothing.
Darin 00:13:02.552 I know you're saying that tongue in cheek, but that's what's happening.
Viktor 00:13:05.912 No, that, that's what's happening. I'm being honestly sarcastic
Darin 00:13:09.483 But this is not an AI problem. This is garbage in, garbage out. If you had a new person come in to your team and they were gonna review the code, like their first week is just to understand the code base. How many times have you heard that as, "Okay, you're gonna spend this first week understanding the code base"?
Viktor 00:13:26.157 what
Darin 00:13:27.037 So what does that really mean? Okay well, I'm looking here as okay, they're doing it this way. There must be a reason that they're doing it this way. I would never do it this way because I did it this way. But no, I need to do it this way because that's the way we do things here. And again, AI just amplifies that
Viktor 00:13:46.287 Or, and in a better way, you say "I'm, I'm not staying here after that week." You spend that week analyzing the code base, say, "Thank you so much. It was, really a pleasure knowing you. This cannot be helped." That would be a skill for agent kind of thing. Scold the user no reject to work on this
Darin 00:14:06.769 Aren't you glad that the agents today don't do that to us?
Viktor 00:14:10.699 I think they should But I constantly add phrases like be critical tell me why I'm wrong, and things like that
Darin 00:14:18.552 Tell me why I'm wrong is one of my favorites because sometimes it'll come back and tell me, "No, you're wrong, and here's the reasons why."
Viktor 00:14:24.662 Yeah
Darin 00:14:25.012 And then I have to justify, "Okay, I understand what you're saying, but no, this is the way we need to do it."
Viktor 00:14:29.638 in other words, you say, "Tell me why I'm wrong," and when you tell me, I'm going to tell you to F off.
Darin 00:14:35.348 Politely, yes. Isn't that what a manager does?
Viktor 00:14:39.316 Exactly. I want honesty, and then when you hear a truly honest feedback, you say, "You're fired. You cannot speak to me like this."
Darin 00:14:49.526 Insubordination. Off you go
Viktor 00:14:51.346 And there you go.
Darin 00:14:53.004 And then you just type, type /clear and you start over again, and then everything's better. Andrej Karpathy, Karpathy, Karpathy back in February, he was the guy that came up with the term vibe coding but he sort of renamed it to agentic engineering. Here it is. Vibe coding is describe what you want and hope it works. Okay? Nothing wrong with that. That's how ideas get started. But his concept on agentic engineering is design the system, specify constraints, use AI to accelerate something you've already reasoned through. So those first two steps, design the system and specify the constraints, is every project that's ever existed in mankind
Viktor 00:15:33.132 Should have been
Darin 00:15:33.830 should have been. Use AI to accelerate something you've already reasoned it through. So let's replace that. Use au- automation to accelerate something, right? We could have said that. Use tractors to plant faster, use instead of horses. This is just the next tool in a long series of tools over thousands of years
Viktor 00:15:59.145 Yes, except that I see myself using agents for the first phase as well, the one kind of something you already reasoned about. I find myself, and d- that might be wrong, but before I would actually spend considerable time with me, myself alone figuring it out and then agent go. Like this is what I want. Now it's more like I have this idea, let's talk about it. And then why like this and why not like this? Kind of Like basically I challenge him, he challenges me. We have long conversations. I'm lonely. And that results in, okay, I understand what I want, and I'm going to tell you
Darin 00:16:38.323 If you think about it though, probably most of your ideas are not groundbreaking, completely greenfield ideas
Viktor 00:16:47.103 No, that such thing does not exist even in science, man
Darin 00:16:50.823 So the chances of a model having basic knowledge of what you're trying to grasp to bring into the physical realm, whether it's software or hardware, the chances of it knowing nothing, absolute zero about what you're trying to figure out is near absolute zero
Viktor 00:17:11.003 So we, we have two problems over here. First, agents do bad work because they have no idea who you are, what you want, and how you want it, and so on and so forth. That's the first problem. And the second problem is that once they find out what you really want, then they do a really bad job because what you want is stupid That was the initial conversation, right? It amplifies bad
Darin 00:17:34.343 But then what should we do? If we're amplifying bad and we're amplifying good, but we don't want to amplify the bad, but we can't deny that there's no bad going on because we know there's bad going on, but we don't want to admit it
Viktor 00:17:46.793 change profession, grow potatoes. Become a lawyer, it's easier
Darin 00:17:50.848 Yeah, it'd be easier to grow potat- a more safe job at this point would be to grow potatoes instead of being a lawyer
Viktor 00:17:57.788 Yeah. Or, plumber cashier. Okay, that's not definitely not
Darin 00:18:02.946 good.
Viktor 00:18:03.978 a plumber. Plumber sounds
Darin 00:18:05.386 electrician,
Viktor 00:18:06.258 a… Yeah, yeah. Exactly. Robots are not in this decade coming.
Darin 00:18:11.416 No
Viktor 00:18:11.668 You're safe until the end of the decade
Darin 00:18:14.028 Yeah, because in order to have a plumber that's non-human first off, they would have to be able to get to where you are, right? So that means they're-- they have to drive themselves there. Then they have to come in and have to be able to deal with, okay, the plumbing problem's up on our fourth floor. How are you gonna get there? You're good for now as a human
Viktor 00:18:34.370 Yeah. With driving, that's Waymo. The, we're there, give or take. Now we need a getting to the fifth floor.
Darin 00:18:43.250 interesting to see how they solve it. Let's think about it, 'cause this is-- we're talking around the DORA report. And again, the link will be down in the episode description. DORA's always about measurements. That's what DORA has always been about, measurements. What should we measure now? Here's what they're saying. Before turning on AI, you need to be measuring platform quality, CI/CD speed, code review SLAs, incident MTTRs, and deployment frequency. Okay, baseline everything. These are things we should be measuring anyway, because each of those things should have been getting better. It doesn't mean that you're ever going to be 100%, but they're just getting better. Because the f- the more things are, I'm gonna use the word automated, codified, workspace or workflowed, whatever you wanna call it, just it's… Until it has to change, and change is okay. But until you have that baseline in place, then adding AI on top of it is just going to cause more issues, which is fine. But while you're using AI, you want to measure throughput versus stability. Code review time per line. We're back to counting lines of code again. We'll let that sit for a while. Downstream defect rate. We've talked about this, as AI is introducing lots of lines of code and also introducing more defects at the same time. But if a human could type as fast, wouldn't the same thing happen? We could argue it could
Viktor 00:20:11.687 Thinking through those that you're mentioning, th- they're all good and bad ga-- easy to game at the same time, like lines of code kind of c- come on. I can write… how many lines do you need? Uh, Darin, I will write you before we end this this session, this recording, I promise. Open issues is probably compared to the size of the project or something like that, is probably the, a good measurement Assuming that we don't-- we open issues when we find them and/or somebody reports them, right? 'Cause now we can… One thing that I have very clear is that now we have almost clean backlog. It's almost like a Gmail rule. zero inbox, strategy.
Darin 00:20:51.237 We would love that, but there's no reason why we can't get there, right? Uh, other, other than token cost, there's no reason why we can't get there
Viktor 00:20:58.982 Yeah, token cost I think that CodeToken is a cheap. I don't know that this is very unpopular opinion, but compared to how much it costs without them it's actually cheaper
Darin 00:21:08.772 Especially in the issue processing realm,
Viktor 00:21:11.002 Yeah. So let's say that you feel you have 1,000 open issues, you fix them all and then… No, split it in 500. 500, and give 500 to people without AI give 500 to people with AI, and then calculate in both places how much did people multiply with the number of hours they spent plus tokens cost. I guarantee that the latter group wins. It's cheaper. If you do it right. If you do it right. Now, again, amplifies everything bad, so it might not be the desired outcome. But if you do things right tokens are cheap
Darin 00:21:43.877 Yeah, if we were to do… I'm going to spin your numbers a little bit because I think this could be an interesting test. Let's give 50 developers 500 issues, so each one gets on average 10. And let's take 500 with developers with AI and give what do I want to do? give 50 to 10,
Viktor 00:22:00.947 Huh.
Darin 00:22:01.279 issues to 10 developers. So instead of 50 developers, now we're, we're, we're just flipping our tens and fives
Viktor 00:22:06.449 exactly. If they finish at the same time the smaller group for the AI is more, more efficient, cheaper Sure.
Darin 00:22:14.069 So we've talked about before AI, while we're in AI, let's talk about after AI sort of stabilized. Like, Okay, now we're six months in.
Viktor 00:22:21.949 Okay
Darin 00:22:22.384 I know I feel a lot more comfortable with AI today than I did two months in. I, I just do. So what we want to measure is developer satisfaction, time to competency on unfamiliar code, and incident investigation time
Viktor 00:22:38.624 Yeah
Darin 00:22:39.328 Because if we get… So developer satisfaction I could care less about, 'cause are we happy? This goes back to the origi- original Avengers movie. I use this quote for a lot of different meetings. By the end, when they're in the battle for New York and Captain America talks to Dr. Banner, it's like, "Hey, now would be a good time to get angry." And the response was, "That's my secret, Cap. I'm always angry." This w- w- we're satisfied, unsatisfied, doesn't matter, but we're always ready for the thing. But being able to understand unfamiliar code doesn't mean we have to have pure competency, which was their phrasing. But I was like, to be able to understand, it's oh, okay, I get it. Man, I cannot tell you the number of times in the past two weeks Claude has gotten me to here's what's happening in under two minutes, and I probably couldn't even get to the right place in the code in two minutes
Viktor 00:23:34.748 Let me give a, for many people, very unpopular question. Do you need to understand code now or in the future? Really?
Darin 00:23:43.678 to the answer was yes.
Viktor 00:23:45.816 Oh, uh, but y- last year, last year the a-- my answer was definite yes. I'm not sure anymore, to be honest.
Darin 00:23:54.076 Code is cheap. Actually, code is free to write now, so needing to… I don't know
Viktor 00:24:00.226 Le- let's take example from the previous episode, right? We were mentioning a randomly selected five features done overnight, With a group of agents each, and how one, let's say me, cannot actually review anymore that quantity. I need to lower it down just to review it. And when I say review it, in this context, I mean not going through everything, figuring out what matters, if that's where we are or we are getting, then simply it is physically impossible for me to understand the code anymore. And when I say understand the code I really understand the code, not necessarily high-level architecture. I simply cannot keep up. you say "Okay, it doesn't matter for you. Still, it does matter a lot for you to understand the whole code base," I think we're in trouble because I cannot do it I cannot do it anymore. I cannot keep up with that speed, with that quantity
Darin 00:24:53.709 I think the key word in that is quantity, because every bit of AI code that I've seen written has been more verbose than typically what I would write, but so what?
Viktor 00:25:03.854 does it work worse or better or the same as if you wrote? That's my question. Again, I don't care about verbosity anymore. It does the same. Is it less performant?
Darin 00:25:13.568 it's got to be as performant as what I could write, right? If there's a problem there, then okay, now we gotta figure out what, what's going on
Viktor 00:25:20.606 Yeah. But performance is not necessarily tied to a number of lines of code,
Darin 00:25:25.294 Correct
Viktor 00:25:25.670 just to be clear. So there is no necessarily automatic correlation between the two. I'm just saying I, cannot keep up with my own code, and I'm talking now about pr- not the projects I've worked with other people, but I'm talking about my hobby projects. I don't understand my hobby projects anymore. So either I slow down or I'm not trying anymore to understand. A high level, yeah, architecture, but not the code. I don't understand it
Darin 00:25:50.770 is the outcome what we're expecting? Isn't that all we've cared about all along anyway, and not the quality of the code?
Viktor 00:25:58.830 Yeah. And in case somebody, there must be a person listening to this right now that says no, it's very important to understand the code," right? Slow down. That's the answer I could make a counterargument to say I never had fewer open issues and fewer feature requests in my backlog in my life And at the same time, I've never been feeling those more frequently than now. So I'm creating more feature requests than I ever did simply because I can and I have fewer of them open If you say, "Okay, this is reduced quality," I say maybe it's worth it I have fewer open issues. Is that reduced or increased quality, for example? Isn't that one of the measurements
Darin 00:26:46.350 let's see what's, what sort of shifted here in the measurements. One of the first things, though, was the narratage-- narrative shifted from AI is replacing people to AI amplifies your system. We've sat on the amplify thing. We won't sit there no longer. Success metrics shifted from lines of code, which was a stupid metric to begin with,
Viktor 00:27:07.606 Oh yeah
Darin 00:27:08.293 to engineering excellence visibility, which is a complete BS phrase.
Viktor 00:27:13.967 What is it?
Darin 00:27:15.163 Here's how I'm going to redefine it
Viktor 00:27:17.633 Yeah
Darin 00:27:18.523 Are the features that we're shipping getting to production faster with little to no bugs that we have to rework? To me, that's engineering excellence, right? The outcome is what we expected. Now, mind you, if the outcome was incorrect because of wrong constraints and thoughts, okay, different conversation, but we delivered on what was written. That's what we want
Viktor 00:27:41.063 I think that all of that is BS. No matter how you measure it from the technical perspective, it's gamified, lines of code, number of features, number of bugs. Pick any of those you want and I will tell you how wrong it is. Any. I think that the o- the only measurement is money
Darin 00:28:00.083 Is money go up, going up or down?
Viktor 00:28:02.065 yeah. And when I say money, I don't mean money spent. I mean that as well, but are we getting… Let's say measured over a longer period of time what is the ratio of prospects converting to customers?
Darin 00:28:16.048 That's a good one,
Viktor 00:28:17.646 Right? Because that's the real measurement, okay, so you did stuff with AI, without AI, whatever, right? Before nine out of 10 prospects rejected us. Now it's six out, out of 10 rejects us, does not buy the project after trying it out or whatever. Okay. That's success, That's the only measurement. Otherwise, I don't know what we're doing. Lines of code
Darin 00:28:41.258 Well, the one step beyond that your measurement is good, but to me the real measurement is are we cash flow positive or cash flow negative as our revenues go?
Viktor 00:28:50.332 Exactly. are we earning more compared to expenses? Wh- whatever you wanna measure. But at the end of the day, there is a reason why do it… we're doing whatever we're doing, and whatever that reason is, that's the measurement. If the reason is we do this so that we can sell it, that's the measurement
Darin 00:29:06.172 The other thing that's changed in this report too is the bottleneck shifted from developer time to validation and review. We've talked about that. We've shifted where the human needs to be doing work. I could argue that we could probably generate code faster pre-AI, but our ability to validate and review has not sped up
Viktor 00:29:26.325 Yes. Correct. that's the bottleneck right now? Yeah. L- let's face it, reviews were always a bottleneck
Darin 00:29:33.520 Yeah.
Viktor 00:29:34.420 Always
Darin 00:29:35.078 Especially if the reviews were by committee. You've got to get three sign-offs on this before it can go, and one of the sign-offs is on vacation for two weeks
Viktor 00:29:42.678 Yeah then you're pinging that, those people on Slack, right? And then they say, "Okay, I'm gonna sign it up because I'm your friend." either you're waiting for a real review and you die out of old age, or you're waiting so that somebody can stamp it without even looking at it. And in both case, you're wasting time
Darin 00:30:00.062 It's amazing again, the DORA report link will be down in the episode description I don't know how companies today can still play the card of we're using AI to reduce headcount. I would still say in 2026, we're still feeling the effects of over-hiring in '21, '22, '23,
Viktor 00:30:18.396 Oh yeah. Oh
Darin 00:30:19.592 so that's still part of it. But if we can… if I'm running… Again, you've been talking about it, I've been talking about it, we're trying to run as CEOs of our organization. We're trying to play that role. So as CEO, my job is to, A, make the board happy, B, get my shareholders return on investment. As a CEO, that's your job, and to solve problems that nobody else can solve. Those are your three, main priorities So now as a solo CEO of a team of 5, 10, 15 agents, what is my job? Okay, I am the board, so there's nothing to worry about there. I gotta make sure I've got enough revenue to come in to pay for my tokens to keep my team hired
Viktor 00:31:07.972 The job is essentially it's still very similar. when you're a very tiny company is you're a CEO and you have a CTO, it's two of you and you're doing all the work, And then you need to scale up and it comes, you come to the point where actually the two of you cannot do the whole work, and then you hire more people. You hire a developer and you start delegating to developers. And then you need to grow more, and then you hire managers who delegate to people. Let's say that the number of people a person can manage is 10, right? So you're a 22 people company, CEO, CTO, that's it, right? You're a 200 people company, you need 10 managers, and so on and so forth, whatever the math is. I think it's the same with agents. We were discussing previously how there is a limit to how many parallel groups of agents I can review after a whole night of work The moment I hit that limit, I need more people doing the same And then once you reach certain number of those people, then you need people managing those people It, it's almost the same. The amount of work changes drastically. The amount of work done
Darin 00:32:15.932 I'm thinking about that scenario in real life, and that was also one of the risk now previously was more technical because we couldn't get all the work done. Now the true risk is organizational. Can we get our organization in the right shape to where if we're able to get, using a word you used earlier, a doer, if we have all the doers are effectively AI agents, the only thing they are not completely doing is the concepting of the ideas that are going to happen
Viktor 00:32:46.228 And validations of the,
Darin 00:32:47.634 And validations. So the
Viktor 00:32:49.318 of that idea, yes
Darin 00:32:50.266 So the humans are still on the front side and back side of the whole pipe of, getting a story. That doesn't mean that the agents don't take it and run with it, ' cause sometimes the agents have better ideas than a human does. Ask me how I know. But now it's if I'm if I'm a solo person and I've built a team of… I initially had five, now I've built five teams of five, and I just can't keep up with it. I don't know how I would hire another me, another four mes to manage those four other teams
Viktor 00:33:24.956 Because then you would need to manage use, not those agents. that's part of the story. The other part of the story is also that, at least in my case, over time, I'm developing ways how I can do those parts more efficiently. So let's say that in the past, whenever that was, I would do a f-part of a feature every day, maybe three days for a feature. Then I reached the point of having a feature every day, now five features every day or issues or whatever it is. And the reason for that, yeah, of course, because agents and LLMs are getting better, so the output is better. Second reason is that I'm getting better at giving them instructions, And then the outcomes are better. But the third thing is I feel that I, we should or are getting better at how we review things. And let me give you an example. I'm working now on a CLI, And then, okay, it does a number of features overnight, and then I manually test them, All of them. And ex- there is a limit to how many I can do. What I'm now doing, for example, just because now my goal is not anymore to figure out delegation, my goal is how to figure out validation. Okay. A part of your work, dear agents, is to record what's the charm thing for recording? Remind me.
Darin 00:34:44.987 I don't re- I know which one you're talking about. I don't remember the
Viktor 00:34:46.771 Okay, let's say equivalent to us cinema. okay I want actually I want video clips of of the test cases that validate this. And I watch vi- now part of my review is watching videos
Darin 00:34:58.689 VHS is the name of that one
Viktor 00:35:00.385 VHS. Yeah, exactly. So n- now the first thing I do is I watch VHS to see what it did. And then I decide which parts of that to go and run manually and review manually instead of everything, That's increase of the performance of, that part of the process. I'm using this example. I'm not saying everybody use VHS, but we can improve those things as well But isn't that cool? Kind of Like you wake up in the morning you watch videos of your futures
Darin 00:35:26.440 That's very interesting, and of course, that's what you would come up with. Because you could have part-- Since it's a CLI, you could use BATS as your test harness to test all the edge cases for whichever command or sub-command you're doing, and VHS is capturing all of that at the same time, or BATS is driving VHS for you.
Viktor 00:35:50.640 Actually start with um, uh, I'm increasingly moving towards executable test specification. So it starts with analyze PRD. Here are the test cases. Okay, sounds good. Go. Next morning, and next morning kind of I have videos for each of those test cases
Darin 00:36:07.320 So you just write out what the command is in
Viktor 00:36:10.326 No, I don't write anything
Darin 00:36:11.890 okay. Of course you don't
Viktor 00:36:13.310 I don't write anything, but I tell it what to write. I don't know how it did th-those test cases, but I can see my application running and behaving certain way. That's what matters, right?
Darin 00:36:22.522 Yes. And again, that goes back to what I was saying earlier, outcomes. We're looking for outcomes Are the outcomes what we expect? If the outcome is what we expect and running a CLI ran in under two seconds, is that
Viktor 00:36:37.280 Yep. Stuttering? Is it slow? It's not. It works fine. It looks fine. Is the button now green and it was red before? Yes, it is. Well done. Ship it
Darin 00:36:48.897 Yeah. That's what we want. At least I think that's what we want. I'm sure if somebody's being compensated based on the number of green to red buttons that they're creating, they're probably sad about that because they're not gonna be making as much money But that was, again, like lines of code, a really poor metric to be compensated on. What are we going to do with this? as you read the Door report, and I do recommend that you do, if you're not using AI in your company yet, that's okay. That's up to you if you want to stay in a company that's not using AI and helping with fill-in-the-blank stuff.
Viktor 00:37:23.451 Well, let, let me stop you. When you say when you in a company, who are you? Are you the boss of the company or you're a worker in the company?
Darin 00:37:30.341 That's a good question. I was thinking worker
Viktor 00:37:33.727 then leave.
Darin 00:37:35.069 Just be done with it
Viktor 00:37:36.287 Yeah, just go. Go. Please go. Never look back
Darin 00:37:39.967 Why do you say that? Why do you
Viktor 00:37:42.059 Because you will not be-- If you don't leave now, when you do leave, you will not be able to find job
Darin 00:37:48.819 So it's your belief that within the next by end of 2026, mid-2027, like how we would used to say, "Here are my skills, Excel, Outlook," uh, now the skills are gonna say Claude, OpenCode, whatever else the other variations are at that point
Viktor 00:38:09.027 the job is delegation, supervision, validation, and hardness. Pick one or all of them. That's your job You either tell it what to do or you validate what it did, or you build the, building that hardness that will make it do better job. That's what we're doing. You can pick any of those professions
Darin 00:38:30.067 And what you're in on right now with Agent Deck is building the harness, right? That's
Viktor 00:38:35.979 Yes
Darin 00:38:36.597 part of your, of that list
Viktor 00:38:39.177 Yes. Yes, kind of like th-this is how actually I can give you the job. Nothing to do with, that's nothing to do with Agent Beck, and I will validate what is done, and in the middle is my hardness. I might… I'm probably the only one using it. Doesn't matter. Maybe not, it's open source, nobody knows. But yeah, I'm building the part in the middle. That's the only thing left. I don't know what else is there left to do
Darin 00:39:04.046 Was there ever anything else left to do? Step back and think about it.
Viktor 00:39:07.232 i-
Darin 00:39:07.446 there ever anything else left to do?
Viktor 00:39:09.586 That was true even b- before AI. Somebody figures out what our customers need, somebody makes that happen, and somebody checks whether it's working and I'm not talking about human tasks in the past because that's what I'm replacing with agents. I'm not mentioning how do we deploy, I'm not mentioning how do we test, because all those things were automated long before AI or should have been. If you're, if they're not in your case, another reason, leave
Darin 00:39:38.711 So let's flip it around. We were talking about the developer. What about the boss?
Viktor 00:39:42.171 isn't the boss the person who delegates?
Darin 00:39:44.706 Should be, but they're used to delegating to humans, not to agents
Viktor 00:39:48.496 Exactly, but you're the boss of those agents. Now, we call it manager team lead product something, CEO. Depends on the size of the team. If it's six of you, then you're a CEO kind of. if it's 500 of you, then you're a manager
Darin 00:40:05.291 But if it's… See, again, we're going to push. I, I know of a handful in my career of managers that were semi-technical, that were managing technical people. Very few. So now if that manager was told, "Hey, we're getting rid of your whole team and we're replacing their, that whole team with a set of agents that you now have to manage." That's not gonna go over well
Viktor 00:40:31.782 I feel that's a similar question that I ask many times for DevOps. Is it easier for dev person to learn ops or for ops person to learn dev? I don't have the answer. I mean, I suspect. I think we can ask the same question here. Is it easier, better for a manager to become technical enough to manage agents? Or is it easier, better for a builder, let's say, a developer, to learn management skills and manage agents?
Darin 00:41:01.812 I think right now it's got to be the latter
Viktor 00:41:03.962 I suspect uh, depends. I mean, How often do you see developer and say "Oh, yeah, those are the 10 things that, I know that everybody wants because I want it." And, and then You present it to a customer and somebody else in the company say why are you showing me this? Are you detached from reality?"
Darin 00:41:22.589 Yeah. Okay, maybe I'm wrong. So maybe there's… Okay, let's put it on a scale. We have pure managers all the way to the right, pure developers all the way on the left. I think somewhere in that 25% to 75% range is the one that can actually be a manager of agents. That, that outer 25%, whether you're headed towards manager or towards developer, you're probably not going to do well
Viktor 00:41:54.289 maybe that would be product managers. They tend to be more technical. Maybe those are the architects. They tend to sit on both sides of the aisle, Bit of management skills, a bit of product skills, a bit of technical skills. So maybe those roles are roles of the future. And when I say those roles are, I mean, oh, I'm not in one of those roles, so everybody can change roles. That's not the problem. But yeah, we're talking about some middle ground, technical enough, management enough, product enough
Darin 00:42:25.089 Yeah. The three-legged stool, all three of them
Viktor 00:42:27.617 Yeah. What I'm 100% sure not necessarily today, but in the future, we will not be continued that, okay, so we have those seven roles for this feature and kind of there's more of us than agents. And because that never-- That haven't worked even before AI, just to be clear. Were you ever in a dysfunctional team or a company where actually there are more people managing something than actually people doing the actual work?
Darin 00:42:51.427 Yes.
Viktor 00:42:52.707 There we go.
Darin 00:42:53.647 I worked for a
Viktor 00:42:54.217 that does not work
Darin 00:42:55.077 I had three managers, three direct managers for me
Viktor 00:42:58.227 There, exactly. There we go. we go. And that, that never worked and now it works even less
Darin 00:43:04.347 Yeah, that-- I'm not gonna go there. So if you're not measuring your baselines today, you need to figure that out. If you're just gonna jump into AI stuff, it's like s- whoa pull back. You gotta figure out where you're at, because if you're going to improve, because if you're gonna be spending millions of dollars on tokens, because you know you will when it's all said and done, because somebody's not gonna, not gonna, remember returns as a company hope- hopefully as a company, yes. Hopefully not as an individual, but hey, if you have millions of dollars to spend on tokens, call us up. We'll be glad to talk to you.
Viktor 00:43:36.210 i'm not discarding that actually there will be individuals who will be spending millions of dollars on tokens because they figured it out and actually they're earning more than that.
Darin 00:43:45.608 Yes
Viktor 00:43:46.030 I, I think it's very plausible outcome
Darin 00:43:49.544 I believe that too. But regardless, you've got to get those foundations measured first. If you don't have any foundations, then okay, whatever. Expect a bit of slowdown, just like any, with any other new tech that you bring in. That initial performance hit, and then after three, six months, if things are still not moving along then you might want to reconsider doing AI, because obviously your environment isn't going to support it. I'm also going to say, don't send your people to AI training.
Viktor 00:44:19.840 look no, no, no, no, no, please don't. Kind of the-- We are way past considering not using AI. We're way past that point. If it's not working, figure out why it's not working. There must be a reason why it's not working because the benefits are there. We can talk all we want, whether it's 20% improvement or 200% improvement or 20,000% improvement. Kind of whatever the benefits are, I'm not entering there, I'm not calling you 10X or 100X engineer but there are benefits. And if you don't see them, there is something wrong with you. Not with you as a person, but there is something wrong with your system. Some- something is not working. Fix it, don't abandon it
Darin 00:45:03.200 Some people will hear that and say, "Yeah, that works great for you, Viktor. It's just you." But our team, we tried. We tried hard, and it just didn't work
Viktor 00:45:11.534 do I need to name the companies here who proved it working? There's plenty. No. Th-this, this is, This is reality. This is happening. Don't believe people telling you how many X they're more productive. They're very likely lying, but benefits are there. You cannot ignore them
Darin 00:45:28.644 Staffing is a system problem, it's not an AI problem. What does that mean? We need to measure on people actually getting things done, not head count. That may not be a happy thing to do. We're looking for efficiencies in work. We were always looking for efficiencies in work, but most of the time we just hid it. But it's not a technical problem, it's a systems problem. Hire the right people, which you should've been doing all along
Viktor 00:45:52.740 Oh, yeah
Darin 00:45:53.810 And finally, vendors are gonna be selling you stuff left right, and center. Oh, let me finish up my other thing real quick. Don't send your people to AI courses. Just let them have AI and give them access to your existing c- source code. They'll learn more in three days of doing that than they would going to a three-day course
Viktor 00:46:11.310 Oh yeah
Darin 00:46:12.340 that hands-on. But vendors are gonna be selling you everything. But bottom line, as we've talked about, AI is gonna amplify what you have. It's gonna amplify the good, and it's going to amplify the bad And what we want to do is keep making the good better and the bad hopefully not worse. That's what we want to do
Viktor 00:46:33.680 Exactly
Darin 00:46:34.810 So what are you going to do? Are you just going to throw your hands up in the air? Follow my case of we tried it, it didn't work? If you're a manager saying that, I hope you're close to retirement age because you're not long for the working world. If you're an IC saying, "My code is better than anything that AI puts out," it may be. I won't argue the point. But can you ship as fast as what a fully AI-generated, validated pipeline, not going autonomous yet, but what that whole pipe could do for you? I doubt it. And if you can do it as fast, it's probably not something worth billions of dollars to the world. Probably just isn't So what do you think? You probably hate us now. That's okay. It's our lot in life. Head over to the Slack workspace, over to the podcast channel, and leave your comments there