00:00:00.122 This episode is sponsored by Sauce Labs
Shubha 00:00:03.395 the discussion was about how three years back in QA, everyone was saying, "If you are not doing test automation, if you're not an automation engineer, your role is going to go away." or, you know, you were saying everyone who was live manual testing engineer, that if you're not moving to QA test automation and not automating your pipelines, your role will go away. Now, you can say the same thing about a test automation engineer now, because hypothesis everyone has is that AI will drive all the automation, your test automation pipeline, and there's no need for these SDETs in your organization. This is DevOps Paradox, episode number 370. Will AI replace QA testers?
Darin 00:01:50.898 Viktor, do you think the role of QA software tester is going to exist in three years as we know it today?
Viktor 00:01:58.484 All roles will exist in the nearby future, and none of them will be the same
Darin 00:02:04.304 All of them
Viktor 00:02:05.544 All of them, and none of them will be the same
Darin 00:02:08.464 On today's show we have with us Shubha Govil from Sauce Labs. Shubha, how are you doing?
Shubha 00:02:14.814 I'm doing well. Great to be on your show, Viktor and Darin. This is amazing. I must start by saying I'm probably the least DevOps native person you have on your show , so let's see how it goes.
Viktor 00:02:27.642 Okay, now we caught you. Now we know that you haven't listened to all the episodes. Now we know it for certain, because if you did, you would know that you're not the least. It's physically impossible.
Darin 00:02:40.286 And by the way, Viktor, you haven't listened to any of them, so she's listened to more than you have.
Viktor 00:02:45.586 That's true. I, I listened to it live while it's happening
Darin 00:02:49.624 back to my statement thinking about software testers, QA, Viktor says everything's still gonna be there, but they're not going… My phrasing now, they're not gonna be doing what they're doing today. Does that seem reasonable?
Shubha 00:03:02.948 So the world we are living in and with the disruption and change that we are experiencing this industry, I actually truly believe that is the statement that we can predict, we can say that all these roles will go away or this role will change XYZ. However, the fact is every minute when our teams are building great products and are thinking about how to solve the problems our customers have, there are many role each of them are playing. So I'm going to put a slightly different hypothesis here. And my hypothesis is that, yes, all of the roles will exist and change, but also each of us will play many different roles. Little bit coming from the product management side of the role, that's where I have been in my career all along. Product management is not same as it was three years back or as it was six years back or even 20 years back when I started, when there were no product managers in the industry. That role was mostly a program manager or project management. Now, QA teams at that time had a very prominent role being a separate organization, defining how the quality of product looks like, and their role has since evolved every few years as well. So just this morning, I was having discussion with one of the industry analysts, and the discussion was about how three years back in QA, everyone was saying, "If you are not doing test automation, if you're not an automation engineer, your role is going to go away." or, you know, you were saying everyone who was live manual testing engineer, that if you're not moving to QA test automation and not automating your pipelines, your role will go away. Now, you can say the same thing about a test automation engineer now, because hypothesis everyone has is that AI will drive all the automation, your test automation pipeline, and there's no need for these SDETs in your organization. What I believe is truly there is the expertise needed to understand what quality really looks like, how should you define the quality for your product, what that strategy should be, as well as value in bringing someone from business, which many times a manual tester understands your business much better, They can walk through each step of your application just the way it's intended to be, it's intended for the user, and provide that perspective in terms of what that quality criteria should look like, what that coverage should look like, because no engineering organization can test everything. So someone needs to define what are the most critical pieces that you should focus on, where to drive coverage. And so going back to your point, Viktor, I believe these roles are going to change. Many of them roles will now… a product manager might be doing some of that quality work, and at the same time, even a QA engineer could be the one bringing that perspective for business and saying, "By the way, this is the requirement that we must address. Otherwise, the code that you have created is not going to work as intended to solve the business need." So that's my point. That's where I will argue.
Viktor 00:06:48.726 I feel that what we're going to see is in certain way extension of the existing platform engineering story. And by that-- So if you step back forget about AI. Where we were moving is that there were, on a very high level, two types of profiles, let's say, right? One would be a person that contributes to the platform its own expertise "Hey, I'm a database expert," or, "I'm a testing expert," or whatever. "I'm going to plug in the service that will actually enable you, who is not an expert in that, to do something," whatever that something is. And then that enabled engineers to say, "Okay, so nobody's an expert in everything, but now I'm a, let's say, a developer. I can write code. That I do that, but I can also deploy that code somewhere without spending seven years learning Kubernetes. I can do this, I can do that," right? And I feel that trend will continue, except that now we have two roles in terms that one is teaching certain aspects of the life cycle to agents, call it prompts, skills, MCPs, whatever. It can be thousand different ways how to do that, but Plugging its knowledge into agents' capabilities. And another role is I do everything from start to beginning of a feature, I'm not going to I write code, I hand it to you, you do this, and then you do that, and then we after we pass seven hands, we have a feature complete. No. I'm going to do feature from beginning to the end. o- other people jobs is to provide all the means how agents can do that well, in a way. That's my near term prediction, let's say next couple of years. I feel that that's where we are going
Shubha 00:08:37.637 I will bring it slightly different nuance to that perspective. this idea that a full stack engineer or developer or even a product manager, because now they have the tools to build everything from all the way from requirements to testing and get something quickly, a feature out quickly, will happen in some places, Having said that, what is not going away in my mind is each role brings an expertise And agents can do stuff for you, but they are not ready to do that thinking as a human who brings that expertise based on their years of learning, seeing the examples of the use cases, how customer are using their application, or how they need to solve one particular problem, or the architecture differences over the years they learned and what works, what does not work, or the data they have based on understanding of their product. And because of that, this uniqueness each role brings will also exist. And in many places, there might be some early capabilities a end-to-end person can build, but to really have true enterprise production-ready apps that are rolled in production with confidence, you will have to have real teams play each of these roles and bring their expertise on the table
Viktor 00:10:12.781 Definitely. Where we might agree or disagree, I'm not yet sure, is whether that expertise will be acted upon, as a chain. Kind of, "I do this, I'm an expert in that. I do that and because I'm an expert in this." or we will be spending more and more time providing that expertise in advance to agents, to models. Okay, so I'm very good at testing. Cool. Let me create all the skills That you might need so that when you write tests, I don't need to be the one babysitting you. You're a junior everything,
Shubha 00:10:50.537 Yeah.
Viktor 00:10:51.280 right? I'm going to teach you how to not to be junior in a way
Shubha 00:10:55.452 that's happening across the board. Now, the challenge that comes with teaching the context and skill, or at least the current AI tools that I'm seeing the behavior of those tools, you can bring lot of expertise and skills. I will tell you last night I rolled out a new skill for my product team on prototyping, Again, from scratch to prototyping using our UI design our, components and things like that. Now, agents are getting good in picking up the nuances of each skill and being able to deliver some outputs that are expected from those skills or what the end goal for that project or initiative or thing might be. What they are still struggling with, I'm going to put that there, is how to apply it for the new problem set, because the uniqueness of customer problem set is not well understood. So it takes a lot more time and cycles to get them to a place, and they will get there, but today I'm not seeing them getting there
Viktor 00:12:10.302 I will actually go further and say that it reminds me of, you know how when, again, without agents, how when people develop something, doesn't matter whether it's tests or code or whatever it is, and can, "Oh, I'm done." No, you're not done. This is never going to be done. This is done when you sunset it, And I think that the same thing applies to, agent capabilities or whatever we wanna call it. No, you're not done with that skill. You never will be. a constant work that we learn more over time, we observe how it behaves, we improve it, we add new ones, we remove things. It's just a constant work. I feel that we are moving towards all of us being mentors to juniors. That might be a good description kind of. okay, so I have seven agents ri- working on this task, and I and, and, and they're dummies. They do terrible work. But I have seven of them, right? It's, uh, How can I leverage the fact that I have infinite number of them doing stuff for me,
Shubha 00:13:17.530 Yeah. And Viktor, I will say this, what you are telling me is what I tell my teams as well. Now, here's the other contrary side which comes into play is, do I truly see that today each of them by becoming the mentor are spending the time in most valuable area to prioritize for the work? Or today their time and energy is better spent some areas leveraging the skills that agent can learn and make them do work for you and mentor them. But are they really prioritizing their time for most valuable use cases where these agents are really good at? And that is the struggle to some extent that industry is looking at is, are we getting there to bring the right use cases, right value that AI agents can bring in the SDLC or in the day-to-day working of our teams?
Viktor 00:14:19.343 Of course we're not. Look, I, I don't know how it's with younger people, but my generation, we became software engineers. So whatever type, doesn't matter, right? Because we did not want to do those things. I don't want to speak with other people. I'm asocial. That's kind of like typical engineer. I don't wanna mentor anybody. I don't wanna manage anybody, right? Kinda We now have to learn the skills that we were trying to avoid as engineers. The basically you're putting me now in a position to do the very set of things that I wanted not to do. Hence, that's why I'm here
Shubha 00:14:58.017 yes. But, and it's also like you are making me sit in front of a tool and watch it do the things that I want to do, because I take pride in doing those things. I take pride in building, and I want to do it because that's exciting for me, versus the things that I actually do want to offload and are not exciting for me. So that's the balance I see is many times teams have to make where decide what is the right value that tool. And the way I look at it is like there is the AI capability itself or the value that AI part of the tool adds. There is the workflow that you need to address and the experience of that full workflow, which might be created by part your overall system, which might or might not be AI, your system and your process that is established, and part what AI brings. in the end, the overall experience of your workflow, combined with your overall system process, along with the AI tools, needs to give you the right value and make sure you can address your use case. Otherwise, I feel that's where the… There's little bit of overhype that every use case can be solved
Viktor 00:16:31.885 Wait, wait, wait.
Shubha 00:16:32.627 the… Okay.
Viktor 00:16:35.393 What do you mean by little bit overhyped? There was never more hype than, than, than we are experiencing right now. So if this is little hype, then there, there was no hype before this
Shubha 00:16:48.333 Uh, you, You are right. Uh, I, think the right way of putting it is, yes, this is the time of overhype we live in. if I go back to your earlier point about a developer not wanting to become a manager, and that's why sitting and deciding to do the job of coding, this overhype is also very taxing and stressful in many teams. And the burnout that at times that we are hearing on many places that people are writing and talking about
Darin 00:17:19.379 Let's look at it a different way. 18 months ago, generative AI was just getting off the ground. Let's call that beginning of 2025. Now we're able to write what appears to be production-ready code, production-ready tests. We think we can do testing, but bottom line is if we're generating that much more, that ratio of testing to deployment is becoming much wider because now instead of 10 features this time, I've got 100 features or 200 features, so nobody can actually get through that. But the dirty little secret that's not really a secret, and nothing's going to ever change this, everybody's still testing in production. I don't care what anybody says. Is that what y'all are seeing at Sauce Labs? That, sure, we're trying to do as best as we can a little bit of testing, but bottom line, no real testing occurs until it's in production
Shubha 00:18:12.850 this is a unfortunate truth in many places. However one of the things that I would like to go to this analogy of absolutely where the amount of code that is being written has 10Xed in many places. The data itself is showing the quality itself is degrading. There are a lot more production incidents reported. going back to the aspect of the bottleneck and whether that bottleneck is addressed early on or is code shipped to production to find the bugs and issues. What we are seeing is many customers are truly aligned with us on this equation that quality is an important aspect of their software development life cycle For them, it means shift left early they… you detect the capability, especially when you think about enterprise customers in regulated industry, where the cost of every bug shipping in the production can be in billions these enterprise customers are looking at this code quality early in their life cycle where validation… So if you think about a conveyor belt, where now the speed of conveyor belt you have increased, but you still have one human standing and only two eyes looking at the thing that goes past that conveyor belt, you have no quality. this is where the validation tax, as some folks are coining that term, is coming into play. We are seeing other side from many of our customer who truly believe in quality and are working with us to enable, ensure they have the right tools. And this is where another place where AI is actually very useful. So I'm not advocating that AI tools are not the right tools for the industry or SDLC. What I'm advocating is AI tools for the right use cases. And this is where testing is one of those use cases, where if you think about a simple retail application or where there's a checkout workflow. I add certain items to my cart, I want to check it out, and if that transaction fails in actual production where the user is not able to check out those items, it's a business loss. It's a business revenue loss, and so that is a very critical workflow. If I have a tool that can help me scan that workflow very quickly, tell me not only the steps to ensure I am writing, covering all the test cases for that workflow, run the execution of this test codes based on the simple intent that I want to check out two items and add to my cart and check it out. And then execute those tests, not only find what's missing in those test cases in simple functional testing, but also being able to do it across the browsers, devices. Because my user might not be just coming from one type of device or browser or OS. And based on that data, next time if the locator for that checkout button changes, is able to go and update the… Give me a better test script to now also test and ensure that every time I have new release or new updates, I'm providing a great functional code that actually works in production. And that's the least confidence that we are seeing that our customers truly believe and why they love to have the right tools for enabling them for these scenarios
Darin 00:22:23.002 What about this? What we don't see In all of this is testing debt, right? That's not factored into any sprint velocity charts or anything like that. Usually those kinds of things show up in production incidents, rollbacks, complaints, whatever the case is. If we are testing in production, and again, we kn- look, we're always testing in production. Nothing's ever perfect. Are you seeing that as being a conscious decision that they are, okay, they're trying to test and they're just doing canaries to test that way? Or we just shipped it and we found out that it was busted
Shubha 00:23:00.203 the fact, as you're calling out is absolutely true. You can never have the 100% coverage for your test cases. So there are things that will ship to production which is not tested, not validated. this is essentially is like you are deferring that bug and depending on when you find it and when it really hits your business outcomes It actually gets fixed. But what you are also doing by just testing in production, you are adding to tech debt. You are not just adding to the verification debt at that point, because you are building more and more code assuming things are working. And the later you wait to fix it, the more tech debt you have already added. You have already spent your engineering's valuable time creating that code which in first place does not work. So yes, in some industries that can definitely… If you are playing in a B2C environment where the consumer apps, which is what we all see like on our phones, Lot of the times our day-to-day consumer apps don't work as we want, and many times what you are calling out, it's just running the production in Canary a- and then shipping the code into production is the phenomena happening. But many industries cannot have these type of production challenges. Yes, they still have to balance how much overall test coverage that they drive, and this is where also the production observability and monitoring comes in, how you ensure the end-to-end visibility and ownership of quality across the organization versus playing the siloed game where only QA is responsible for quality or only production is responsi-- uh, the engineering is responsible for SRE team or only product is responsibility. In the end, it's the ownership of the whole team together to drive accountability and ensure whatever you are shipping is not done at the time of engineering delivering the code. Shipping is done when you have user adoption and user reducing the support tickets you are getting and being able to see actual usage on your product. And again, I come from a product background, so my point of view is, in the end, product owns ensuring the right business outcome, which means also working together with every part of these siloed teams to ensure there is the right quality to drive, in the end, customer adoption, the value we are offering to the customer and the user, the reason why they pay us, the reason why the business exists.
Darin 00:26:03.026 I'm thinking everything you said that product is the glue that talks to all the other sideload teams. There's a problem in that statement in that those teams aren't talking to each other, and now product has become the bottleneck to getting anything done
Shubha 00:26:18.035 that's a interesting way of putting and thinking about it. Living in the product world what I truly believe is if anyone is seeing the silos where teams themselves are not talking to each other, each one and many teams work in the triads of engineering, design, and product. Any of these roles, and QA teams being part of that, they have to bring it to drive the end outcome together. They have to break those silos. So breaking the silos is not just product's responsibility. There's something else not working in that environment if the teams are not talking to each other and can't have a honest discussion about what to release and how to ship a product with quality in the end. So th- that's maybe different thinking Darin, than what you are saying
Viktor 00:27:09.479 I will go out on a limb here and say that actually I have nothing against silos. I think that silos are fantastic. It's just that silos are most of the time not done as they should be. So here's an example. For me, AWS is a silo. I cannot talk to them. I'm not in a multi-billion dollar company, right? I cannot talk to them. I cannot contact them. If I open an issue, God knows when it will be answered. But, and that's definitely a silo for me. But I don't mind because that silo works, right? They have their own services. They ensure that services work. They do whatever they do. I consume them. I'm fine. I don't necessarily have to ever talk to somebody from AWS. If I do talk to them, that's because something is terribly wrong, The problem is when silos, instead of giving us services that we can consume, oh, okay, so here's a service for a database. You don't need to ever contact me again. It just works. Just run the database yourself. I'm fine with that silo. The problem is that silos are not designed like that. They're designed like you need to open a Jira ticket. You need to be religious because only prayer will help you that I ever answer to that Jira ticket, and let's see what happens, right? That's a wrong type of silo. if I could have your silo as a service, and I consume it, whatever that something is I don't mind. I'm happy
Shubha 00:28:47.054 A- and I think you're calling out very interesting nuance to point, which is the problem is not the silo, it's the lack of clear communication to make some process work between the silos
Viktor 00:29:00.256 Correct
Shubha 00:29:00.634 or consistency of information shared that's happening to get the next silo to do something or act or… I truly believe that is a process and organizational problem. if we are creating teams that are operating in the processes where communication is not flowing, and that is absolutely the wrong silo to have if communication is not flowing. So yeah, I'm agreeing with you in it by this nuanced aspect of the silo as you are calling out.
Viktor 00:29:38.924 Going back to previous conversation about tech debt and et cetera. Until-- In the past, whenever I would speak with a company and I would say, "Hey, but why don't you do this? Why don't you do that? Why don't you move to Kubernetes? Why don't you go to cloud? why don't you rewrite your mainframe into something that makes more sense?"
The Answer Is Always The Same 00:29:56.971 "We cannot. we are so overwhelmed with what we are doing with, other things that we simply we will never have time to do this." We, we were mentioning quality, right? Big part of quality is that you have infinite number of issues that are unresolved sitting there somewhere in Jira, I understand, y-y-you cannot do everything. You don't have time. You need to focus on something. But now you can, right? Now you can. Now you can-- Okay let's see what happens if you give hundred issues to agents to solve. I'm not saying deploy to production without uh, doing anything, but mentioned earlier 10X, right? Let's say that it's 20% help. That's already enough for me to have 20% of my time not focused on putting down fires and thinking about the s- what's really important, what matters, what can I really do instead of writing yet another label selector,
Shubha 00:30:56.138 I will first say something and I will counter-argue myself, and here's why. I do see, and day-to-day in our, our own work here at Sauce Labs, there are examples in our customers' examples where there is a lot of tech debt, where something was sitting on an old code, old system that was not easy to work with, and it was just sitting as a tech debt. And AI tools definitely helped getting some of that tech debt either by converting from one language to other, like a old framework-based tool that was there and someone built and took effort to move it. Having said that, this is the second part of it. The amount of things that you need to prioritize or amount of areas that you have to prioritize to drive that growth does not mean you have to prioritize every single tech debt. There are a lot of the things you'll have can say, "Does not matter anymore. You-- We don't need to fix." Now, even the ones that you prioritized, the amount of time, not everything can be resolved. So an example I will tell you where AI generated a very long document and ways of solving a age-old problem And the-- once the spec itself is very long to s- figure out how to solve, sometimes even the decision-making and the time spent on it becomes really valuable and waste of your time y- and energy if you are-- have to now go through this new document that's being created, the new spec that's being created to solve an age-old problem, which y- your engineering team might have taken, say, like a few months, but now they have to add this time to first understand how the AI is asking you to solve for this problem. Is it the right mechanism? Is it the right process to solve it? Does it-- Did it consider all the guardrails that you would have considered? Did it understand all the dependencies you had on the other systems? And your own team now take hours and hours of time to decipher some of that information before they can act on it. And that is a additional burden that's being added because there is a belief that if I throw in all the 10,000 tech debt projects that I had, and I'm just making up that number, I don't believe anyone would have that many. But you are also adding more work for you to understand which ones are the right ones to solve and how it's solving and can it truly help you solve. And that is sink cost
Viktor 00:33:59.005 What I'm trying to say is that you're not necessarily adding more work to you. You are switching your attention from whatever you were doing before to something else. Like in my example right, The more code, let's say not 100%, whatever the percentage is, the more code itself can be written by AI, the more time I can spend on the things you just mentioned, right? The less getters and setters in Java I write, the more time I have to think about, okay, so what is this really about? Should we really solve it or should we not solve it, I find myself, let's say that I was 80/20% before 80% just typing, you know, on a keyboard typing, typing, typing, and now I'm 80% just not necessarily even touching the keyboard. It's kind of, okay, so what am I really doing, right? And that's that shift. It's not that the work became longer, that I defied the laws of physics and now there is 48 hours in a day for me, right? It's more that the type of the work we do changes more into design and thinking and figuring out why do we do this and what are we doing rather than typing
Shubha 00:35:13.869 A-and that is very true, and that is helpful use of time if you truly get that time back. the challenge at many places is people have not yet seen the proof of being able to get that additional thinking time. some of the data points, I think CodeRabbit published that they are spending 1.7x time more for defects now than the human written code to validate and check the defects. to some extent what I believe, and I could be totally wrong on this one I do see lot of clear use cases where tech debt can be resolved. Having said that, also being mindful of how much true time you are getting back in your day to drive strategy thinking and actually doing the right work versus creating AI-generated code or specs or design or areas where it takes you more time to now validate that, ensure that it's going to deliver you the right outcome before you actually go and use it and implement it. And that's the part, it's a continuous drift also that the more you add, the more AI slop to some extent gets created. And the more you-- AI slop gets created, the more validation gap, the more tech debt c-get created. And to some extent, I feel a very important role that comes into play here is some that you alluded earlier, is being able to see and monitor agents and mentor them to do exactly what you want them to do versus letting them loose and create these long form documents or huge lines of code, which itself might be difficult to manage and maintain and support for future.
Viktor 00:37:13.152 That reminds me of, I was in quite a few situations where we would get new people in a company juniors, they're not battle tested, they're not really good yet, and then half of the company would complain and say, "Okay, now I need to spend my valuable time to babysit them," This is a waste of my time." my answer to that is that this is-- what you're now saying is very short-term thinking, right? You're investing your time in those newbies so that after that investment, those newbies will be able to do the work for you, You need to spend now so that you benefit from those people later
Shubha 00:38:01.104 I agree with that statement very much, right? And that is true whether it's human or agents. I agree with that as well It's what I'm questioning right now, and honestly based on current capabilities, Not predicting future here. Based on current capabilities, how much time to mentor new agents versus time to also ensure those are the right things these agents should be trained on.
Viktor 00:38:35.026 It's a horrible what we need to do to make any decent work out of agents. It's just horrible. However, that being said, that amount of time that I need to spend babysitting agents is huge, but infinitely smaller than a year ago The capabilities a year ago were nowhere near where they're now. And if I extrapolate and say, okay, let's say it will slow down, but it will not slow down to a crawl A year from now, I, I cannot even imagine where are we going to be a year from now. Ju- just extrapolating where we were, Darin said earlier at the very beginning, early 2025, that's a year and a half ago. Year and a half ago, I was, " I'm not touching this. This is horrible. Why anybody would want this?" That was a year and a half ago
Shubha 00:39:26.741 Yeah, and I agree with that hypothesis that a year from now, the capabilities will be much more improved than where are today. And that does mean we have to train some agent for some skills, But it comes down to how we iterate on the new versions of… So as new tools come, new models get refreshed, quad code every day, you probably are seeing the amount of updates we get on the tool uh,
Viktor 00:39:56.099 I stopped even watching release notes. I cannot follow those guys. I cannot
Shubha 00:40:00.951 Yeah. So with the-- it's about iteration versus saying today we have to spend time on mentoring all the right skills, all the agents, getting everyone because the team size is small. We need agent for everything, and each of us should be having agents to deliver everything. Where I'm spending most of my time and I'm advocating for, is finding the right areas where these agents are really good at and mentoring and gi-giving them the right skills they need, ensuring that's enabled the right context, the right details for them to enable you and empower your teams, even the smaller size of the teams, to do better and continue to mentor only for those right skills. As you see more new things coming, iterate on it. This is so fast-changing. Experimentation and failing and learning what's not working is the best way we can train ourselves and say, "Yes, we are ready for that one year out agent world where these tools will be so good that I will have my new set of junior team all running on these agents." So that's my point of view.
Viktor 00:41:23.411 Except that whoever says we are ready I, I, I'm not ready. I don't know what's going on even anymore. I'm not ready for anything. I'm lost
Shubha 00:41:32.571 I think that's a fair statement for a big part of the tech world right now. On one hand, the urgency to move fast is there, right? But move fast and build the right thing is there. On the other hand, it still comes down to making sure it's building the right thing, it's doing the right outcome.
Darin 00:41:58.011 And those right outcomes, going back into testing, we're used to doing standardized testing. We want to optimize for less manual test authoring, right? We… There's no reason t- uh, Viktor And I were talking about this earlier today. Even years ago, we had code generation. It wasn't that great, but it was code generation. Y- I could argue, and we didn't talk about this earlier today, Viktor, I could argue that we could do test generation of, or code generation of tests pretty well with those tools. Not great, but better than just the code itself. We could really scaffold something out.
Shubha 00:42:34.127 Absolutely
Darin 00:42:35.521 But now there's this phrase I've been hearing, I don't know if it came from y'all or not at Sauce Labs, intent-driven testing. Did
Shubha 00:42:43.017 Yep.
Darin 00:42:43.291 of Sauce Labs?
Shubha 00:42:44.637 That did come out of Sauce Labs and, uh
Darin 00:42:48.101 w- what is that about? Because I'm trying to understand, we don't need faster tests. I mean, Obviously we want tests to be faster. We always want tests to be faster, but we want the test to be correct, which is something we've also always wanted. So what is intent-driven testing?
Shubha 00:43:03.515 Great question there. So one of the things going back to the example that I was calling out earlier is cart checkout, An application, a retail application that has a simple scenario of checking out a cart. Now, the intent of that application code itself is allow someone to select one item or two item or XYZ number of items, add to the cart and checkout. It's a simple intent for that application. Now, in order for… to fulfill that intent, that application has certain capabilities, some buttons to click, some things to select, and in the end, based on that, a test case needs to be generated, which a lot of the time either was done manually or in automation, someone writing the test case that is followed to run across the automation line through frameworks like Selenium and Playwright and Appium for mobile application Now that itself can be… AI is very good at solving that. So some of the tools from Sauce Labs that our customers are using are for intent-based testing test case writing. It's a authoring tool that as the amount of VIP coded applications is increasing, this tool helps our customers to write test cases which are intent-driven. You just need to, in plain English, specify, "I need to check out a backpack, add it to the cart, add my address, add my credentials, write me a test case, run it, execute it against multiple devices, browser, OSes," which is where Sauce Labs provides the execution aspect for these intent-driven test cases. And then also automate the insights and analytics. Tell me what worked, what failed, how flaky some of the test cases were. Now, one thing we are finding that AI is really good in writing these test cases, which are framework agnostic, release agnostic. So some change happens in your release, it detects what was the change. Now button moved, color changed. Being able to update the test scripts. The second part which is-- AI is really good at is also enabling those insights and being able to ask these questions in plain simple language, and developers getting those real-time insights in their workflow to go back and adjust the test cases if they want to manually look at it. So there is human-in-the-loop aspect. What we truly believe in intent-based test case environment is that it is not only writing the test-- authoring the right test cases based on the intent of application, they execute well, and automated insights are further provided so that a human in the loop that is working and accepting this process can work with those insights and data to say, "Do I really want to take this fix it's recommended? Do I want to do something different?" And even more, we also have tools in the production environment where crash test and error and crash reporting data that helps further augment. So even the bugs that might be escaping to production, as we were talking earlier, we have data to tell you where to further shift left and fix these issues. And that's really we-- what we mean by driven test automation.
Darin 00:46:56.864 It's interesting that you bring up shift left again. You said it earlier. We can't leave the people actually creating the specs out of the deal. was in the SDLC in real life before, but now w- it's more important than ever. And if we don't get that, plus everything you were just explaining too, it just brought me back to doing BDD, behavior-driven development, 20 years ago. This is just, to me, that's just the next evolution of BDD
Shubha 00:47:26.561 That is so true. I know the tools like the Cucumber and other, the BDD world, that was the promise of those tools, right? What--
Darin 00:47:36.111 which did not deliver
Shubha 00:47:37.335 Which did not deliver. And we hear this from many customers, and this is why every time we share with them our product and what they are using in their environment, they really see the difference between what is the promise of these tools was versus where AI is now enabling this intent-driven test case authoring, but a full loo-loop to ensure there's the right quality that is from pre-production to production based on the data and insights that are added on top to bring, again, I will say shift left. But I also, I believe there's almost this aspect of shifting up, where you have to drive observability across the SDLC, truly integrated-- integrating quality from all the way requirements and spec all the way to the production.
Darin 00:48:31.915 but now you're saying with that, that w- we have to break down silos. That's the implicit statement under the hood there.
Shubha 00:48:38.271 Yes, I still believe in breaking down the silos, but as long as those silos can communicate well and have the right level of governance to talk to each other, I think I'm okay with silos too.
Darin 00:48:51.821 I don't think either of those things will ever be achieved in my lifetime
Shubha 00:48:55.921 You mean breaking down the silos
Darin 00:48:57.693 Breaking down the silos or communicating well between the two, between the silos. I don't think either will work well. I just don't think it'll ever happen
Shubha 00:49:05.513 I have to accept, unfortunately, the way we humans are, both of these will exist. So
Darin 00:49:11.621 Yeah. That's what I'm saying.
Shubha 00:49:12.953 view
Darin 00:49:13.031 will never change
Viktor 00:49:14.601 so the solution to what you just said is just if we get rid of humans, we'll be fine
Darin 00:49:20.041 Which is where we're headed, right? That's what we want anyway
Viktor 00:49:22.925 Good
Shubha 00:49:23.567 Okay, let's bring in agents, mentor them, get them ready for this next world
Viktor 00:49:29.087 Yeah the exactly. Bring in agents, mentor them, ensure that they're doing the right thing, and dedicate your life to growing tomatoes or something like that
Shubha 00:49:37.611 Oh, I would love that.
Darin 00:49:39.301 Oh, dear. It's beginning to feel like to me that companies that can figure out how to actually distribute testing, right? This is one thing. You're shifting up, shifting left, all the shifts. If we can actually distribute this testing responsibility that has been the bane of existence for most places, the C-levels will say look, we don't have time to run it through full testing," or not necessarily that. The reality is all the other timelines slip, and QA had two weeks to test, but now they only have two hours, That's the most reality. But now if we're able to distribute testing across the organizations, key point, without losing any kind of quality, More than likely we're gonna ship faster and more reliably
Shubha 00:50:19.987 one thing to add to that point of view is the ownership of quality, as you are saying, is not just on QA anymore and because the amount of the time pressure, because velocity of the code and how quickly now something needs to be validated. It needs to be ingrained at every cycle from requirements to code writing to how code is reviewed, validated, to the functional testing, and every stage of that ensuring that there is ownership for quality. And that is where truly this shift left can happen once there is ownership and accountability to ensure whatever you are shipping is not just lines of code, but something that can add value and show what the needed requirement for that particular capability is and can work.
Darin 00:51:15.912 And I also think that the companies that still try to keep QA in their own little box as a separate function or whatever you want to call it uh, they're gonna find themselves completely bottlenecked and unable to compete.
Shubha 00:51:26.844 A-agree.
Darin 00:51:27.932 bottom line. That's… I don't know how it's gonna shift. You shift there again. if we can get to a point to where… Okay, if we can get rid of all the… See, this is the problem. If we can go ahead and eliminate all the existing, sorry, executives, and sorry to you, Shubham, as a chief product officer uh, if we can get rid of all the chiefs and then promote all the people at the bottom moving up, and they become the chiefs, eventually, this is the true-- We've always argued about top-down, bottom-up. Done correctly, and bringing in agents at the bottom, we can push all the humans out the top, and we can go grow tomatoes or potatoes or whatever you want
Shubha 00:52:05.889 On one hand, again, this is the dichotomy of, or to some extent, the name of your podcast, the paradox here is, right? As a industry, we all have reached this consensus to some extent that AI adoption needs to happen, and it needs to happen fast. And whether it's each role getting eliminated and humans moving up and getting to the tomato farm and the agents being the new worker bees. What we all have not reached yet is consensus on the value of doing that. And that's one point that we all agree, we need AI and it can do a lot of great things in… especially in SDLC. There are lot of use cases, quality validation being a big use case there that AI can help drive and expedite But many of us are still not able to explain why, whether execs move up or the next layer move up, what value it will add by just having agents in the workflow. So my true belief is that yes, there might be smaller teams, there might not be as much hierarchy in the organization structure. There will be lot more agents in the workflow. I'm not saying agents won't be in the workflow, but I don't see us getting to a tomato farm, yes, at this point or even in a year from now
Darin 00:53:42.109 So you talked about a few of the products that are available through Sauce Labs. What are some other ones that if people are like, "Look, we know we've needed to do QA automation all along and we've just never done it. We're still manually wiring up Selenium or Playwright now." I am thankful for Playwright because Selenium is not my favorite. I'll just say it politely.
Shubha 00:54:01.789 Darin, you do know S- uh, Sauce Labs name come from Selenium?
Darin 00:54:06.024 I did know that, and I'm-- that's why I'm being as polite as I possibly can right
Shubha 00:54:09.664 Okay. I just had to call that out. No,
Darin 00:54:12.144 No, I I know. And I'm, I'm thankful for it, right? I mean, Back when we had nothing else, Selenium was phenomenal. It was the best thing at the time, and it's still very good
Shubha 00:54:22.606 Yes. To call out some of the other products that you asked. One of the things that we recently released earlier this year is an API to access our real device infrastructure. So many of those who don't know Sauce Labs have been in the business for 18 plus years enabling customers for largest of the automation. So top Fortune 1000 customers running their automation testing on us, especially in regulated financial, retail software industry large automations are run on Sauce Labs infrastructure. We provide both virtual emulator simulators and real devices, real physical devices, the phone devices in the cloud for customers to be able to leverage. We are the first ones to open up our real device infrastructure through an API, making it programmable. So now the customers who wanted to run long soak tests or get the deeper device level analytics, whether it's thinking about the heat or memory or some of the other parameters based on how their app is reacting and working, we can expose that data, and especially AI intelligence coming on these phones. So that's another area that I would like to call out. Sauce Labs also plays in beta testing, visual testing, accessibility, as well as you think about the a crash and an- error reporting and analytics to enable largest of the game industry players especially. That's where big part of our work is to enable the developers to be able to get error and crash test data and applying AI on top of it as well to give them fast analytics to fix… Going back to shift left. It's not just production, but production data, take it back and fix it early in the chain.
Darin 00:56:22.753 Okay. For your physical devices, round number, how many physical devices are you managing today?
Shubha 00:56:30.039 We have thousands of physical devices in our data centers across US EMEA, and we have a new data center in India as well. So that's the thousands of multiple OS vendors, industry vendors, devices that we are managing in our Real Device Cloud, along with virtual and simulators as well
Viktor 00:56:53.034 So next time I need a phone or a laptop or something, I'll just visit you guys
Shubha 00:56:58.032 That's the best way, Viktor
Darin 00:57:00.185 That was going to be my question. It's like, how many… So I said devices, physical devices. I really meant to ask how many physical phones do you have? Because to me, that would be the absolute worst because back when I was doing mobile development in the early 2010s, just working on devices I know it's gotten better, just like everything else has gotten better over the years. But I cannot imagine trying to keep phones and tablets. That's probably just as… I think tablets are probably worse than phones now
Shubha 00:57:31.735 That is true, and we do manage tablets as well. the best part for the developers is think about the complexity of these thousands of devices and the battery popping heating up, and being able to go fix that in a cloud. So lot of the times our customers depend on us to manage these hundreds and thousands of devices to and browser combinations for them so that they don't have to worry about the pain of managing their own grids.
Darin 00:58:04.442 If for no other reason, if you've got a mobile app and you need to test on multiple devices. I mean, It's one thing with Apple and iOS, right? 'Cause there's only so many.
Shubha 00:58:14.050 Yes.
Darin 00:58:14.552 But Android…
Shubha 00:58:16.070 The combinations are crazy. Yes. Yes. And especially outside US, the variety of vendors and devices
Darin 00:58:26.815 what is… Okay, I know we're running a little bit long here, but I have one question. So out of the Android devices what vendor outside the US, w- because, Android's big here too, but not like Apple. Is Samsung normally the number one or is it another
Shubha 00:58:42.709 Samsung is the number one, what we are seeing in many of the markets. Now the OPPO is another one that shows up quite often. But those are the most common devices that are tested against, right? So here's the other aspect of it, like depending on how you're driving your test coverage, you might choose certain versions of devices certain vendors and because they might address all the use cases that you need to add- address for most of your users. So OPPO and Samsung are the two largest that we hear on the Android side. And of course, Pixel phones are popular as well
Darin 00:59:23.237 And I'm assuming you're glad that the BlackBerry devices no longer exist
Shubha 00:59:27.133 Yes. Uh, Interesting tidbit fact one of the earliest product that I worked in the industry was to create applications for BlackBerry. So I have dealt with those.
Darin 00:59:40.563 Do not miss it. Do not miss it.
Shubha 00:59:42.485 miss Yes
Darin 00:59:44.173 So Sauce Labs can be found at saucelabs.com. In case you don't know how to spell it, that's S-A-U-C-E-L-A-B-S.com, and all of Shuba's information will be down in the episode description. Shuba, anything else that you'd like to say before we wrap it up today?
Shubha 01:00:03.391 I just want to thank all the listeners and great time chatting with you both Darin and Viktor. These are very important topics of discussion, especially as the SDLC is shifting and the transformation AI is bringing, but also the impact that we talked about on the roles and quality and how we all should be thinking about reducing the validation and verification tax.