Viktor
00:00:00.315
So we, we have two problems over here. First, agents do bad work because they have no idea who you are, what you want, and how you want it, and so on and so forth. That's the first problem. And the second problem is that once they find out what you really want, then they do a really bad job because what you want is stupid That was the initial conversation, right? It amplifies bad
Darin
00:01:26.217
Viktor, what do you think of this? A bad engineer with AI-assisted development is much, much, much, much worse than just a bad development engineer
Viktor
00:01:35.677
Okay. Yeah, of course. have… I think that bad engineer with a laptop is much, much worse than a bad engineer with a calculator. Yes
Darin
00:01:44.197
You have companies that are laying off people because the roles weren't AI enough for the company
Viktor
00:01:51.009
I feel that there are two big cases for layoffs. One is we need much smaller workforce to do what we're doing. That's probably the bigger one, and it's completely wrong. And another one is we're actually going to do so much more. We shouldn't off people. But we cannot get there with the people we have So it's not necessarily always layoffs, but maybe replacements disguised as layoffs And then you if that's your case, you need to think very hard about what you're doing
Viktor
00:02:26.000
Yeah, I was now more referring to the person being laid off only to be replaced next week with another person that happens to be less skeptical about AI. There is a lot of skepticism right now, a lot of not good enough And then what do you do as a company who chose to go all, all in on AI? I get rid of this person and I get the enthusiastic one
Darin
00:02:50.903
That seems like it would be an answer, but DORA has sort of disproved that. Yes, that same DORA that's given us the DevOps reports for years and years. Back in May of '26, they released their ROI of AI-assisted software development report.
Darin
00:03:07.893
If I remember, I will put the link to that in the show notes so you can go download it and read it for yourself. But what they found and what they're saying is AI without engineering excellence just scales your problems That makes too much sense
Viktor
00:03:26.258
I was afraid for a second that because I haven't read the report, shame on me, that sentence would start with whatever it started and end up with kind of no return of investment on AI, but that's not the case. It's k- no return of investment without people really knowing what they're doing, right?
Viktor
00:03:55.850
Or not, actually no, that's wrong. It's not more of everything. It's you get more of what you were doing And I'm sorry to say that, but there is a statistical chance that what you are doing is bad
Darin
00:04:09.470
Because y- you've made foundational choices years ago that you're still trying to… Okay, here's my example. You're trying to take your monolithic Java web app and just run it inside Kubernetes
Viktor
00:04:22.540
Exactly. And now you can have gazillion more Litic DevOps running in Kubernetes. Hooray. Well done
Darin
00:04:35.240
dealing with incidents because you did not think about what it takes to run application X inside of Kubernetes because you just said, "Hey, Kubernetes is the hot thing. That's what we're going to do." And now you gotta deal with that fallout
Viktor
00:04:52.150
that raises what I believe is an interesting question. Imagine before AI, you would go and spend some time with some team, a company, and they would tell you what you're doing, what they're doing, and you would say, "Yeah, but why are you doing that?" And they would probably answer eventually with, "Yeah we just kind of-- We don't have time and money to actually do it the right way now." And that could be true. But if you're still doing it that way, and now that time is not a problem anymore that might lead me to conclude that what you were saying is that you were fooling yourself and lying to me. It's not really that you don't have time. You just don't know how to do it right And now I have proof because now there is nothing stopping you to do it right
Darin
00:05:41.900
Yeah, there's a variation of what you're saying from the report. This is summarized. Productivity isn't the problem. The problem is that organizations without mature platforms, clear workflows, and strong reviews amplify chaos when they turn on AI
Viktor
00:06:05.300
Yeah or maybe restate it with what was good. Now, Now you have more of what was good and more of what was bad. And if you had more bad than good, then, both are amplified proportionally
Darin
00:06:20.890
If we think about it, if we have the elite performers, I imagine there are certain companies that are elite performers worth-- with AI. My guess is those are gonna be the ones that were not the command and control top-down organizations. Would you agree or disagree with that?
Viktor
00:06:38.691
those must be organizations where individuals had a lot of autonomy. I strongly believe that. Because now we're moving to the world where actually I alone do the work of a team, right? It's almost pointless now for us to work as a team, Where, "Oh, I will do this part and you will do this part," and so on and so forth. Kind of that does not make sense because you're telling that to your agents now, not your colleagues. And that means that this new world favors people who were autonomous before
Viktor
00:07:14.523
Yeah, kind of okay. It might take me infinite amount of time, but I'll just do it, right? If that was your attitude and not "Oh, I'm go- I'm do it, and now I'm waiting for Michael. Michael knows HTML. When Michael is finished, then I will continue my part." If that's what you were doing before, now that is amplified, right? Th- that does not simply work because you don't know Michael happens to be a bot, an agent, r- and you don't know what to ask him or her. I don't know whether it's she or he agents it. So yeah this whole setup now favors companies that were looking for and enabling people to have great autonomy.
Darin
00:07:55.896
And just like any new tech, the Dora report found that there was like this big J curve. So in the first one to three months, you took a dip in productivity,
Darin
00:08:05.852
but then it usually took off from there. But the problem is as things take off, we've talked about this, then it's dealing with reviewing of that, reviewing the code, reviewing the documentation, reviewing the thing that the human needs to, using your word, add taste to. That's something that the machines can't do for us yet
Viktor
00:08:29.302
Yes. It, it somehow reminds me of testing in the past where many companies or teams would-- they would keep some tests and tests are almost always flaky. I rarely see tests that are reliable, And what do you do when you have flaky tests? You run it again. And you run it again, and eventually it will pass, and you say, "Well done." Or you have teams who would be improving those tests. Say, "Okay, actually, we're not going to rerun it. We're going to figure out why it's flaky. We're going to improve it, and the next time it runs, it's going to be less flaky." that's the approach we need to take with agents, It's going to create a mess, and you're going to figure out one of many reasons why it's a mess, and you're going to extend system context or skills or MCPS or whatever you're using for the next time. And you're going to do it over and over again. And as you're improving agents, because agents are actually agents plus LLMs, they're actually pretty pretty good at what they're doing. They just lack the context of you, of your organization, of your choices, and so on and so forth. And you cannot provide that context from the first attempt, but you can over time. And that means that over time, there will be fewer bad things, fewer false positives, and the time it takes you to review it will shorten Apart from using agents to review as well. It's for shorten because, okay, so those are the things that after 57th iteration of improving what's or not now it's d- doing it cons- consistently right. We don't need to focus on this anymore We can now remove this part of review and that will give us more time to focus on other parts of the review Which I feel again, story that we should have had maybe we did before agents. Nothing new
Darin
00:10:26.466
Nothing new. So with the DORA report, one word they used, and I've already used it here, is amplifier. What is an amplifier? If you think about it from an audio perspective, it takes a signal and it makes it louder, right? That's the basics of it. Unless you take it to 11, then it's distorted. Spinal Tap reference. That's the problem, though, is you have to be careful with an amplifier because if you do turn it too far up, you can have problems. Feedback in not the good sense. But let's think about the elites and the not elites process here again. So let's spell it out. For the elite teams, they already have optimized testing, great CI/CD, platform engineering, clear processes, AI accelerates that implementation, everything's glorious. But for the struggling teams, backlog debt, slow CI/CD, fewer safeguards, unclear workflows, AI is pumping out code left and right, the pipeline chokes, the meantime to resolution just goes through the roof, if it's even up. And now the senior engineers that were saving your butt in the past are now the bottleneck ' Viktor: think about it this way, what does it… So let's say that you have a code base, whatever the code base is, and you instruct the agent, "Okay, I would like you to fix this issue or add this feature," or whatever that is. And I'm going to ignore system context now, I'm going to ignore skills and all the stuff that I mentioned earlier, right? You just say "Develop this new feature." What will it do? It will not randomly develop feature as you specified. It will today, not necessarily a year ago, it will actually analyze and see how you're doing things, and it'll do it more or less the same way as you're doing it And what, what that means is that you, if you had I want to use the word I'm not supposed to, bad code is going to do more of it because it said, "Oh, this is what you're doing? Well done, Viktor. Let me do more of it because you're always right That's wonderful and scary all at the same time
Viktor
00:12:31.951
actually, wasn't it one of our complaints they're doing random things and they're increasingly doing the fewer random things? they're they're taking over what you were doing in a way you were doing it. You have no CI. We don't need it, right? You have flaky tests. Let-- If you have tests, cool. Let me see how you did previous tests like this. Excellent. I'm going to do more of those. And if they were flaky, now I have more flaky tests It was okay to run for seven hours. What's five minutes added to seven hours? Nothing.
Darin
00:13:09.483
But this is not an AI problem. This is garbage in, garbage out. If you had a new person come in to your team and they were gonna review the code, like their first week is just to understand the code base. How many times have you heard that as, "Okay, you're gonna spend this first week understanding the code base"?
Darin
00:13:27.037
So what does that really mean? Okay well, I'm looking here as okay, they're doing it this way. There must be a reason that they're doing it this way. I would never do it this way because I did it this way. But no, I need to do it this way because that's the way we do things here. And again, AI just amplifies that
Viktor
00:13:46.287
Or, and in a better way, you say "I'm, I'm not staying here after that week." You spend that week analyzing the code base, say, "Thank you so much. It was, really a pleasure knowing you. This cannot be helped." That would be a skill for agent kind of thing. Scold the user no reject to work on this
Viktor
00:14:10.699
I think they should But I constantly add phrases like be critical tell me why I'm wrong, and things like that
Darin
00:14:18.552
Tell me why I'm wrong is one of my favorites because sometimes it'll come back and tell me, "No, you're wrong, and here's the reasons why."
Darin
00:14:25.012
And then I have to justify, "Okay, I understand what you're saying, but no, this is the way we need to do it."
Viktor
00:14:29.638
in other words, you say, "Tell me why I'm wrong," and when you tell me, I'm going to tell you to F off.
Viktor
00:14:39.316
Exactly. I want honesty, and then when you hear a truly honest feedback, you say, "You're fired. You cannot speak to me like this."
Darin
00:14:53.004
And then you just type, type /clear and you start over again, and then everything's better. Andrej Karpathy, Karpathy, Karpathy back in February, he was the guy that came up with the term vibe coding but he sort of renamed it to agentic engineering. Here it is. Vibe coding is describe what you want and hope it works. Okay? Nothing wrong with that. That's how ideas get started. But his concept on agentic engineering is design the system, specify constraints, use AI to accelerate something you've already reasoned through. So those first two steps, design the system and specify the constraints, is every project that's ever existed in mankind
Darin
00:15:33.830
should have been. Use AI to accelerate something you've already reasoned it through. So let's replace that. Use au- automation to accelerate something, right? We could have said that. Use tractors to plant faster, use instead of horses. This is just the next tool in a long series of tools over thousands of years
Viktor
00:15:59.145
Yes, except that I see myself using agents for the first phase as well, the one kind of something you already reasoned about. I find myself, and d- that might be wrong, but before I would actually spend considerable time with me, myself alone figuring it out and then agent go. Like this is what I want. Now it's more like I have this idea, let's talk about it. And then why like this and why not like this? Kind of Like basically I challenge him, he challenges me. We have long conversations. I'm lonely. And that results in, okay, I understand what I want, and I'm going to tell you
Darin
00:16:38.323
If you think about it though, probably most of your ideas are not groundbreaking, completely greenfield ideas
Darin
00:16:50.823
So the chances of a model having basic knowledge of what you're trying to grasp to bring into the physical realm, whether it's software or hardware, the chances of it knowing nothing, absolute zero about what you're trying to figure out is near absolute zero
Viktor
00:17:11.003
So we, we have two problems over here. First, agents do bad work because they have no idea who you are, what you want, and how you want it, and so on and so forth. That's the first problem. And the second problem is that once they find out what you really want, then they do a really bad job because what you want is stupid That was the initial conversation, right? It amplifies bad
Darin
00:17:34.343
But then what should we do? If we're amplifying bad and we're amplifying good, but we don't want to amplify the bad, but we can't deny that there's no bad going on because we know there's bad going on, but we don't want to admit it
Darin
00:17:50.848
Yeah, it'd be easier to grow potat- a more safe job at this point would be to grow potatoes instead of being a lawyer
Darin
00:18:14.028
Yeah, because in order to have a plumber that's non-human first off, they would have to be able to get to where you are, right? So that means they're-- they have to drive themselves there. Then they have to come in and have to be able to deal with, okay, the plumbing problem's up on our fourth floor. How are you gonna get there? You're good for now as a human
Viktor
00:18:34.370
Yeah. With driving, that's Waymo. The, we're there, give or take. Now we need a getting to the fifth floor.
Darin
00:18:43.250
interesting to see how they solve it. Let's think about it, 'cause this is-- we're talking around the DORA report. And again, the link will be down in the episode description. DORA's always about measurements. That's what DORA has always been about, measurements. What should we measure now? Here's what they're saying. Before turning on AI, you need to be measuring platform quality, CI/CD speed, code review SLAs, incident MTTRs, and deployment frequency. Okay, baseline everything. These are things we should be measuring anyway, because each of those things should have been getting better. It doesn't mean that you're ever going to be 100%, but they're just getting better. Because the f- the more things are, I'm gonna use the word automated, codified, workspace or workflowed, whatever you wanna call it, just it's… Until it has to change, and change is okay. But until you have that baseline in place, then adding AI on top of it is just going to cause more issues, which is fine. But while you're using AI, you want to measure throughput versus stability. Code review time per line. We're back to counting lines of code again. We'll let that sit for a while. Downstream defect rate. We've talked about this, as AI is introducing lots of lines of code and also introducing more defects at the same time. But if a human could type as fast, wouldn't the same thing happen? We could argue it could
Viktor
00:20:11.687
Thinking through those that you're mentioning, th- they're all good and bad ga-- easy to game at the same time, like lines of code kind of c- come on. I can write… how many lines do you need? Uh, Darin, I will write you before we end this this session, this recording, I promise. Open issues is probably compared to the size of the project or something like that, is probably the, a good measurement Assuming that we don't-- we open issues when we find them and/or somebody reports them, right? 'Cause now we can… One thing that I have very clear is that now we have almost clean backlog. It's almost like a Gmail rule. zero inbox, strategy.
Darin
00:20:51.237
We would love that, but there's no reason why we can't get there, right? Uh, other, other than token cost, there's no reason why we can't get there
Viktor
00:20:58.982
Yeah, token cost I think that CodeToken is a cheap. I don't know that this is very unpopular opinion, but compared to how much it costs without them it's actually cheaper
Viktor
00:21:11.002
Yeah. So let's say that you feel you have 1,000 open issues, you fix them all and then… No, split it in 500. 500, and give 500 to people without AI give 500 to people with AI, and then calculate in both places how much did people multiply with the number of hours they spent plus tokens cost. I guarantee that the latter group wins. It's cheaper. If you do it right. If you do it right. Now, again, amplifies everything bad, so it might not be the desired outcome. But if you do things right tokens are cheap
Darin
00:21:43.877
Yeah, if we were to do… I'm going to spin your numbers a little bit because I think this could be an interesting test. Let's give 50 developers 500 issues, so each one gets on average 10. And let's take 500 with developers with AI and give what do I want to do? give 50 to 10,
Darin
00:22:01.279
issues to 10 developers. So instead of 50 developers, now we're, we're, we're just flipping our tens and fives
Viktor
00:22:06.449
exactly. If they finish at the same time the smaller group for the AI is more, more efficient, cheaper Sure.
Darin
00:22:14.069
So we've talked about before AI, while we're in AI, let's talk about after AI sort of stabilized. Like, Okay, now we're six months in.
Darin
00:22:22.384
I know I feel a lot more comfortable with AI today than I did two months in. I, I just do. So what we want to measure is developer satisfaction, time to competency on unfamiliar code, and incident investigation time
Darin
00:22:39.328
Because if we get… So developer satisfaction I could care less about, 'cause are we happy? This goes back to the origi- original Avengers movie. I use this quote for a lot of different meetings. By the end, when they're in the battle for New York and Captain America talks to Dr. Banner, it's like, "Hey, now would be a good time to get angry." And the response was, "That's my secret, Cap. I'm always angry." This w- w- we're satisfied, unsatisfied, doesn't matter, but we're always ready for the thing. But being able to understand unfamiliar code doesn't mean we have to have pure competency, which was their phrasing. But I was like, to be able to understand, it's oh, okay, I get it. Man, I cannot tell you the number of times in the past two weeks Claude has gotten me to here's what's happening in under two minutes, and I probably couldn't even get to the right place in the code in two minutes
Viktor
00:23:34.748
Let me give a, for many people, very unpopular question. Do you need to understand code now or in the future? Really?
Viktor
00:23:45.816
Oh, uh, but y- last year, last year the a-- my answer was definite yes. I'm not sure anymore, to be honest.
Viktor
00:24:00.226
Le- let's take example from the previous episode, right? We were mentioning a randomly selected five features done overnight, With a group of agents each, and how one, let's say me, cannot actually review anymore that quantity. I need to lower it down just to review it. And when I say review it, in this context, I mean not going through everything, figuring out what matters, if that's where we are or we are getting, then simply it is physically impossible for me to understand the code anymore. And when I say understand the code I really understand the code, not necessarily high-level architecture. I simply cannot keep up. you say "Okay, it doesn't matter for you. Still, it does matter a lot for you to understand the whole code base," I think we're in trouble because I cannot do it I cannot do it anymore. I cannot keep up with that speed, with that quantity
Darin
00:24:53.709
I think the key word in that is quantity, because every bit of AI code that I've seen written has been more verbose than typically what I would write, but so what?
Viktor
00:25:03.854
does it work worse or better or the same as if you wrote? That's my question. Again, I don't care about verbosity anymore. It does the same. Is it less performant?
Darin
00:25:13.568
it's got to be as performant as what I could write, right? If there's a problem there, then okay, now we gotta figure out what, what's going on
Viktor
00:25:25.670
just to be clear. So there is no necessarily automatic correlation between the two. I'm just saying I, cannot keep up with my own code, and I'm talking now about pr- not the projects I've worked with other people, but I'm talking about my hobby projects. I don't understand my hobby projects anymore. So either I slow down or I'm not trying anymore to understand. A high level, yeah, architecture, but not the code. I don't understand it
Darin
00:25:50.770
is the outcome what we're expecting? Isn't that all we've cared about all along anyway, and not the quality of the code?
Viktor
00:25:58.830
Yeah. And in case somebody, there must be a person listening to this right now that says no, it's very important to understand the code," right? Slow down. That's the answer I could make a counterargument to say I never had fewer open issues and fewer feature requests in my backlog in my life And at the same time, I've never been feeling those more frequently than now. So I'm creating more feature requests than I ever did simply because I can and I have fewer of them open If you say, "Okay, this is reduced quality," I say maybe it's worth it I have fewer open issues. Is that reduced or increased quality, for example? Isn't that one of the measurements
Darin
00:26:46.350
let's see what's, what sort of shifted here in the measurements. One of the first things, though, was the narratage-- narrative shifted from AI is replacing people to AI amplifies your system. We've sat on the amplify thing. We won't sit there no longer. Success metrics shifted from lines of code, which was a stupid metric to begin with,
Darin
00:27:18.523
Are the features that we're shipping getting to production faster with little to no bugs that we have to rework? To me, that's engineering excellence, right? The outcome is what we expected. Now, mind you, if the outcome was incorrect because of wrong constraints and thoughts, okay, different conversation, but we delivered on what was written. That's what we want
Viktor
00:27:41.063
I think that all of that is BS. No matter how you measure it from the technical perspective, it's gamified, lines of code, number of features, number of bugs. Pick any of those you want and I will tell you how wrong it is. Any. I think that the o- the only measurement is money
Viktor
00:28:02.065
yeah. And when I say money, I don't mean money spent. I mean that as well, but are we getting… Let's say measured over a longer period of time what is the ratio of prospects converting to customers?
Viktor
00:28:17.646
Right? Because that's the real measurement, okay, so you did stuff with AI, without AI, whatever, right? Before nine out of 10 prospects rejected us. Now it's six out, out of 10 rejects us, does not buy the project after trying it out or whatever. Okay. That's success, That's the only measurement. Otherwise, I don't know what we're doing. Lines of code
Darin
00:28:41.258
Well, the one step beyond that your measurement is good, but to me the real measurement is are we cash flow positive or cash flow negative as our revenues go?
Viktor
00:28:50.332
Exactly. are we earning more compared to expenses? Wh- whatever you wanna measure. But at the end of the day, there is a reason why do it… we're doing whatever we're doing, and whatever that reason is, that's the measurement. If the reason is we do this so that we can sell it, that's the measurement
Darin
00:29:06.172
The other thing that's changed in this report too is the bottleneck shifted from developer time to validation and review. We've talked about that. We've shifted where the human needs to be doing work. I could argue that we could probably generate code faster pre-AI, but our ability to validate and review has not sped up
Viktor
00:29:26.325
Yes. Correct. that's the bottleneck right now? Yeah. L- let's face it, reviews were always a bottleneck
Darin
00:29:35.078
Especially if the reviews were by committee. You've got to get three sign-offs on this before it can go, and one of the sign-offs is on vacation for two weeks
Viktor
00:29:42.678
Yeah then you're pinging that, those people on Slack, right? And then they say, "Okay, I'm gonna sign it up because I'm your friend." either you're waiting for a real review and you die out of old age, or you're waiting so that somebody can stamp it without even looking at it. And in both case, you're wasting time
Darin
00:30:00.062
It's amazing again, the DORA report link will be down in the episode description I don't know how companies today can still play the card of we're using AI to reduce headcount. I would still say in 2026, we're still feeling the effects of over-hiring in '21, '22, '23,
Darin
00:30:19.592
so that's still part of it. But if we can… if I'm running… Again, you've been talking about it, I've been talking about it, we're trying to run as CEOs of our organization. We're trying to play that role. So as CEO, my job is to, A, make the board happy, B, get my shareholders return on investment. As a CEO, that's your job, and to solve problems that nobody else can solve. Those are your three, main priorities So now as a solo CEO of a team of 5, 10, 15 agents, what is my job? Okay, I am the board, so there's nothing to worry about there. I gotta make sure I've got enough revenue to come in to pay for my tokens to keep my team hired
Viktor
00:31:07.972
The job is essentially it's still very similar. when you're a very tiny company is you're a CEO and you have a CTO, it's two of you and you're doing all the work, And then you need to scale up and it comes, you come to the point where actually the two of you cannot do the whole work, and then you hire more people. You hire a developer and you start delegating to developers. And then you need to grow more, and then you hire managers who delegate to people. Let's say that the number of people a person can manage is 10, right? So you're a 22 people company, CEO, CTO, that's it, right? You're a 200 people company, you need 10 managers, and so on and so forth, whatever the math is. I think it's the same with agents. We were discussing previously how there is a limit to how many parallel groups of agents I can review after a whole night of work The moment I hit that limit, I need more people doing the same And then once you reach certain number of those people, then you need people managing those people It, it's almost the same. The amount of work changes drastically. The amount of work done
Darin
00:32:15.932
I'm thinking about that scenario in real life, and that was also one of the risk now previously was more technical because we couldn't get all the work done. Now the true risk is organizational. Can we get our organization in the right shape to where if we're able to get, using a word you used earlier, a doer, if we have all the doers are effectively AI agents, the only thing they are not completely doing is the concepting of the ideas that are going to happen
Darin
00:32:50.266
So the humans are still on the front side and back side of the whole pipe of, getting a story. That doesn't mean that the agents don't take it and run with it, ' cause sometimes the agents have better ideas than a human does. Ask me how I know. But now it's if I'm if I'm a solo person and I've built a team of… I initially had five, now I've built five teams of five, and I just can't keep up with it. I don't know how I would hire another me, another four mes to manage those four other teams
Viktor
00:33:24.956
Because then you would need to manage use, not those agents. that's part of the story. The other part of the story is also that, at least in my case, over time, I'm developing ways how I can do those parts more efficiently. So let's say that in the past, whenever that was, I would do a f-part of a feature every day, maybe three days for a feature. Then I reached the point of having a feature every day, now five features every day or issues or whatever it is. And the reason for that, yeah, of course, because agents and LLMs are getting better, so the output is better. Second reason is that I'm getting better at giving them instructions, And then the outcomes are better. But the third thing is I feel that I, we should or are getting better at how we review things. And let me give you an example. I'm working now on a CLI, And then, okay, it does a number of features overnight, and then I manually test them, All of them. And ex- there is a limit to how many I can do. What I'm now doing, for example, just because now my goal is not anymore to figure out delegation, my goal is how to figure out validation. Okay. A part of your work, dear agents, is to record what's the charm thing for recording? Remind me.
Viktor
00:34:46.771
Okay, let's say equivalent to us cinema. okay I want actually I want video clips of of the test cases that validate this. And I watch vi- now part of my review is watching videos
Viktor
00:35:00.385
VHS. Yeah, exactly. So n- now the first thing I do is I watch VHS to see what it did. And then I decide which parts of that to go and run manually and review manually instead of everything, That's increase of the performance of, that part of the process. I'm using this example. I'm not saying everybody use VHS, but we can improve those things as well But isn't that cool? Kind of Like you wake up in the morning you watch videos of your futures
Darin
00:35:26.440
That's very interesting, and of course, that's what you would come up with. Because you could have part-- Since it's a CLI, you could use BATS as your test harness to test all the edge cases for whichever command or sub-command you're doing, and VHS is capturing all of that at the same time, or BATS is driving VHS for you.
Viktor
00:35:50.640
Actually start with um, uh, I'm increasingly moving towards executable test specification. So it starts with analyze PRD. Here are the test cases. Okay, sounds good. Go. Next morning, and next morning kind of I have videos for each of those test cases
Viktor
00:36:13.310
I don't write anything, but I tell it what to write. I don't know how it did th-those test cases, but I can see my application running and behaving certain way. That's what matters, right?
Darin
00:36:22.522
Yes. And again, that goes back to what I was saying earlier, outcomes. We're looking for outcomes Are the outcomes what we expect? If the outcome is what we expect and running a CLI ran in under two seconds, is that
Viktor
00:36:37.280
Yep. Stuttering? Is it slow? It's not. It works fine. It looks fine. Is the button now green and it was red before? Yes, it is. Well done. Ship it
Darin
00:36:48.897
Yeah. That's what we want. At least I think that's what we want. I'm sure if somebody's being compensated based on the number of green to red buttons that they're creating, they're probably sad about that because they're not gonna be making as much money But that was, again, like lines of code, a really poor metric to be compensated on. What are we going to do with this? as you read the Door report, and I do recommend that you do, if you're not using AI in your company yet, that's okay. That's up to you if you want to stay in a company that's not using AI and helping with fill-in-the-blank stuff.
Viktor
00:37:23.451
Well, let, let me stop you. When you say when you in a company, who are you? Are you the boss of the company or you're a worker in the company?
Viktor
00:37:42.059
Because you will not be-- If you don't leave now, when you do leave, you will not be able to find job
Darin
00:37:48.819
So it's your belief that within the next by end of 2026, mid-2027, like how we would used to say, "Here are my skills, Excel, Outlook," uh, now the skills are gonna say Claude, OpenCode, whatever else the other variations are at that point
Viktor
00:38:09.027
the job is delegation, supervision, validation, and hardness. Pick one or all of them. That's your job You either tell it what to do or you validate what it did, or you build the, building that hardness that will make it do better job. That's what we're doing. You can pick any of those professions
Darin
00:38:30.067
And what you're in on right now with Agent Deck is building the harness, right? That's
Viktor
00:38:39.177
Yes. Yes, kind of like th-this is how actually I can give you the job. Nothing to do with, that's nothing to do with Agent Beck, and I will validate what is done, and in the middle is my hardness. I might… I'm probably the only one using it. Doesn't matter. Maybe not, it's open source, nobody knows. But yeah, I'm building the part in the middle. That's the only thing left. I don't know what else is there left to do
Viktor
00:39:09.586
That was true even b- before AI. Somebody figures out what our customers need, somebody makes that happen, and somebody checks whether it's working and I'm not talking about human tasks in the past because that's what I'm replacing with agents. I'm not mentioning how do we deploy, I'm not mentioning how do we test, because all those things were automated long before AI or should have been. If you're, if they're not in your case, another reason, leave
Darin
00:39:38.711
So let's flip it around. We were talking about the developer. What about the boss?
Viktor
00:39:48.496
Exactly, but you're the boss of those agents. Now, we call it manager team lead product something, CEO. Depends on the size of the team. If it's six of you, then you're a CEO kind of. if it's 500 of you, then you're a manager
Darin
00:40:05.291
But if it's… See, again, we're going to push. I, I know of a handful in my career of managers that were semi-technical, that were managing technical people. Very few. So now if that manager was told, "Hey, we're getting rid of your whole team and we're replacing their, that whole team with a set of agents that you now have to manage." That's not gonna go over well
Viktor
00:40:31.782
I feel that's a similar question that I ask many times for DevOps. Is it easier for dev person to learn ops or for ops person to learn dev? I don't have the answer. I mean, I suspect. I think we can ask the same question here. Is it easier, better for a manager to become technical enough to manage agents? Or is it easier, better for a builder, let's say, a developer, to learn management skills and manage agents?
Viktor
00:41:03.962
I suspect uh, depends. I mean, How often do you see developer and say "Oh, yeah, those are the 10 things that, I know that everybody wants because I want it." And, and then You present it to a customer and somebody else in the company say why are you showing me this? Are you detached from reality?"
Darin
00:41:22.589
Yeah. Okay, maybe I'm wrong. So maybe there's… Okay, let's put it on a scale. We have pure managers all the way to the right, pure developers all the way on the left. I think somewhere in that 25% to 75% range is the one that can actually be a manager of agents. That, that outer 25%, whether you're headed towards manager or towards developer, you're probably not going to do well
Viktor
00:41:54.289
maybe that would be product managers. They tend to be more technical. Maybe those are the architects. They tend to sit on both sides of the aisle, Bit of management skills, a bit of product skills, a bit of technical skills. So maybe those roles are roles of the future. And when I say those roles are, I mean, oh, I'm not in one of those roles, so everybody can change roles. That's not the problem. But yeah, we're talking about some middle ground, technical enough, management enough, product enough
Viktor
00:42:27.617
Yeah. What I'm 100% sure not necessarily today, but in the future, we will not be continued that, okay, so we have those seven roles for this feature and kind of there's more of us than agents. And because that never-- That haven't worked even before AI, just to be clear. Were you ever in a dysfunctional team or a company where actually there are more people managing something than actually people doing the actual work?
Viktor
00:42:58.227
There, exactly. There we go. we go. And that, that never worked and now it works even less
Darin
00:43:04.347
Yeah, that-- I'm not gonna go there. So if you're not measuring your baselines today, you need to figure that out. If you're just gonna jump into AI stuff, it's like s- whoa pull back. You gotta figure out where you're at, because if you're going to improve, because if you're gonna be spending millions of dollars on tokens, because you know you will when it's all said and done, because somebody's not gonna, not gonna, remember returns as a company hope- hopefully as a company, yes. Hopefully not as an individual, but hey, if you have millions of dollars to spend on tokens, call us up. We'll be glad to talk to you.
Viktor
00:43:36.210
i'm not discarding that actually there will be individuals who will be spending millions of dollars on tokens because they figured it out and actually they're earning more than that.
Darin
00:43:49.544
I believe that too. But regardless, you've got to get those foundations measured first. If you don't have any foundations, then okay, whatever. Expect a bit of slowdown, just like any, with any other new tech that you bring in. That initial performance hit, and then after three, six months, if things are still not moving along then you might want to reconsider doing AI, because obviously your environment isn't going to support it. I'm also going to say, don't send your people to AI training.
Viktor
00:44:19.840
look no, no, no, no, no, please don't. Kind of the-- We are way past considering not using AI. We're way past that point. If it's not working, figure out why it's not working. There must be a reason why it's not working because the benefits are there. We can talk all we want, whether it's 20% improvement or 200% improvement or 20,000% improvement. Kind of whatever the benefits are, I'm not entering there, I'm not calling you 10X or 100X engineer but there are benefits. And if you don't see them, there is something wrong with you. Not with you as a person, but there is something wrong with your system. Some- something is not working. Fix it, don't abandon it
Darin
00:45:03.200
Some people will hear that and say, "Yeah, that works great for you, Viktor. It's just you." But our team, we tried. We tried hard, and it just didn't work
Viktor
00:45:11.534
do I need to name the companies here who proved it working? There's plenty. No. Th-this, this is, This is reality. This is happening. Don't believe people telling you how many X they're more productive. They're very likely lying, but benefits are there. You cannot ignore them
Darin
00:45:28.644
Staffing is a system problem, it's not an AI problem. What does that mean? We need to measure on people actually getting things done, not head count. That may not be a happy thing to do. We're looking for efficiencies in work. We were always looking for efficiencies in work, but most of the time we just hid it. But it's not a technical problem, it's a systems problem. Hire the right people, which you should've been doing all along
Darin
00:45:53.810
And finally, vendors are gonna be selling you stuff left right, and center. Oh, let me finish up my other thing real quick. Don't send your people to AI courses. Just let them have AI and give them access to your existing c- source code. They'll learn more in three days of doing that than they would going to a three-day course
Darin
00:46:12.340
that hands-on. But vendors are gonna be selling you everything. But bottom line, as we've talked about, AI is gonna amplify what you have. It's gonna amplify the good, and it's going to amplify the bad And what we want to do is keep making the good better and the bad hopefully not worse. That's what we want to do
Darin
00:46:34.810
So what are you going to do? Are you just going to throw your hands up in the air? Follow my case of we tried it, it didn't work? If you're a manager saying that, I hope you're close to retirement age because you're not long for the working world. If you're an IC saying, "My code is better than anything that AI puts out," it may be. I won't argue the point. But can you ship as fast as what a fully AI-generated, validated pipeline, not going autonomous yet, but what that whole pipe could do for you? I doubt it. And if you can do it as fast, it's probably not something worth billions of dollars to the world. Probably just isn't So what do you think? You probably hate us now. That's okay. It's our lot in life. Head over to the Slack workspace, over to the podcast channel, and leave your comments there