Jake Aaron Villarreal: I'm Jake Aaron Villarreal, born and raised in Silicon Valley. Here to take you behind the scenes to share what it's like to be a startup founder, the journey they're on, the problems they face, the products they build as they look to transform industries. I'm excited to have with us today a very important person, Gennady Pekhimenko, co-founder and CEO of CentML. Gennady, welcome to the show.
Gennady Pekhimenko: Thank you, Jake. Thank you.
Jake Aaron Villarreal: Well, a little bit more about Gennady. He's not just a founder and runs a company, startup. He's also a professor at the University of Toronto. He currently leads CentML, which optimizes machine learning workloads for faster and more cost-effective performance. CentML's key products include CServe, an LLM inference framework, and a platform for orchestrating and scheduling ML workloads. They've raised over $30 million in funding from investors like Gradient Ventures and Nvidia and others that we can talk about as well. CentML focuses on improving AI compute efficiency. Gennady is also a co-creator of the MLPerf benchmark suite and holds several patents related to machine learning and systems optimization. We're excited to have you on here today. And I guess before we do that, where are you joining us from today?
Gennady Pekhimenko: I'm joining from Toronto, Canada, right from downtown.
Jake Aaron Villarreal: Hundreds of AI startups are launching every month, battling to build their founding teams. As a leader, your job is to get results. When it comes to hiring, that's where it gets tough. So, you go out and you try a recruitment firm, but they don't understand your story. They're off target. And when they send you candidates, it's a waste of time. We believe you should never have your time wasted. That's why we launched Match Relevant, because your story is more than just an open role. It's your founder's journey, the problem you're solving, the product you're building, and why it matters. When we work with companies, we make sure we understand your whole story. So, we go out and do a search, we're on target, it's worth their time, they're interested, and more importantly, it's worth yours. And when it comes to hiring engineers, we work to make sure we get it right by deploying a team of seasoned CTOs that have built some of Silicon Valley's best companies. They can collaborate with you in the technical interviewing process. They can be a sounding board or they can run it for you. When it comes to building teams, there's no time to waste. Let's make it count. If you have a role that needs to be filled, book a time with a hiring guide at matchrelevant.com and learn how we do it.
Okay, great. Such a great place. I I'm not sure what's in the water up in Toronto, but it seems like a lot of inventors and creators and innovators and AI seem to be coming out of that area and Canada in general. So, wherever you're getting that insight and intelligence, uh, love to have some of that. It's incredible what's happening in the world today and what's happening around AI. We're going to dive into...
[Gennady's Journey: From Russia to AI Innovator]
...that. Um, but before we do, tell us a little bit about yourself. Uh, Gennady, what was your sort of origin story? How did you, how did you get into technology? And what was kind of your background that led you to ultimately really becoming a professor?
Gennady Pekhimenko: Yeah. So, back from the time I was a high school student, I was also always very excited about math in general and programming. So I was doing a lot of math and programming competitions back home in Russia, participating in a lot of things. So not surprisingly, when I finished high school, I went into essentially Computer Science department in Moscow State University and studied there. I always liked uh building um systems that work. So I don't um, even though I had a lot of math background, I was like always like, "Oh, I really want to see things I use and deployed." It was so exciting. So the journey during my undergrad years was all about like you know learning how to program well, how to actually build larger system, learning things like compilers and operating system.
And then eventually as I was very excited to continue my education and learn things and do research, I moved uh to actually Canada to University of Toronto, where I uh did my master's in computer science. And not surprisingly, I already was familiar with some of the basic things in AI uh like back propagation, but then I moved you know to Toronto and this is where I've seen all these famous people working in systems and especially in machine learning. And that was for reference 2006, 2007. So it's been before the AI revolution happened. It was one of the periods of this long AI winters where you know majority of the people didn't care about AI and they didn't think it has a bright future. Um but I was very impressed on where they were already. I actually used machine learning back in 2007, 2008 for my master thesis, right? Um um and I try to apply to system level problem like compiler optimizations. It worked reasonably well, but not to the level where I was like feeling to change my research direction.
And you know after my master's I continued working in an industry. I work at IBM and IBM Research, and then uh but always have that desire to go for research. So ended up going and I really like academia at the time. The lesson I learned uh you know if you want to be really strong in academia like and become a professor you really need to go to a top top tier schools to get your PhD. So I have to go to US and the closest of the top schools was actually Carnegie Mellon University and it's a very famous place for...
[The Evolution of AI and Gennady's Research]
...computer science, actually where Jeff Hinton started some of his career as well, right. So I went there, work on a lot of interesting problem in systems, computer architecture. And um you know got my PhD there. And while doing it, I was involved a lot with working with industry companies like Qualcomm, Nvidia, Intel, and, and Microsoft obviously. So that gives me a lot of good perspective on where the industry is moving. And in parallel with me working on my PhD and publishing a lot of papers, I noticed that AI actually starts to pick up, right. First it was AlexNet that was like a huge breakthrough in 2012, 2013, then there was like a follow-up you know works on uh you know different very advanced convolutional neural networks to address that. It was clear we're coming to the level where that actually become practical. So after I finished my PhD, I was very convinced that for my future academic journey, I'm actually going to work uh on AI related problems and system level optimizations for AI.
And when I start my research group and that was one of the topic and something I did myself at Microsoft Research. It was obvious to me after a few years there's like huge gaps in industry that exist, right? There is like no good benchmark. This is why we work on MLPerf with a lot of other people. There are no good tools to measure uh you know performance of these workloads. It was really wild west back then in 2016, 2017. You read a blog post, everyone was claiming they're 1,000X better than everyone else. It was like ridiculous, like it was completely very unhealthy, very not... It's like "Okay, I don't like that. I want an order into this thing." So we built it academically, and as we built it, I realized by going into industry, giving talks like Google, Amazon, uh Facebook, Meta, that we actually built something interesting that the big industry doesn't have. But unfortunately they won't use it directly if I keep it at the you know student uh quality of the code. So it needs to become an enterprise software. And the only way you can do it uh you can do it by you know building you know uh your own IP.
Luckily I was always hiring graduate students that like to build artifacts. So there was like, when I was saying there's a time to start building something, there was no shortage of interest among my own students, current and former students. So some of them uh left university joined the company and some of them uh left industry like companies like Intel, AMD, Amazon and join us as well. So with them we started building, and essentially I always like as a person, I always like to make things more efficient. It was always somewhere inside of me. I don't like when things are broken and not very efficient. And the AI really needs it. It's really so expensive that it really needs a lot of that.
Jake Aaron Villarreal: Yeah. Well, there's a lot here to talk about. Um, and we're going to talk about ML. You're talking about your company...
[Balancing Academia and Entrepreneurship]
...and what, what you're solving and who it's helping. But before we kind of do that, I want to thank you for, you know, sharing a little bit more about your background. You know, it's pretty uncommon from what we see as someone to be a, a professor as well as leading a startup. So, just from a bandwidth perspective, just from a sanity perspective, how are you able to do that? And what's the, maybe some of the, I don't know, some of the, the hacks or tricks or optimization that you use personally to be able to manage you know, really two different areas of your life when it comes to school professor and leader of a company?
Gennady Pekhimenko: Yeah. Just to be honest, like you can't really do both full-time. It's just invisible. Right. So what happens in my case, I'm, I'm, I was on sabbatical uh until July. And now I'm on a leave at the university, right. So I got my tenure there and I got on sabbatical, I'm on leave. What does it mean in practice? I had a you know very reduced load. So I don't need to teach, I have not participated in any committees. The only responsibility that stays with you is your students. Right? So, the... but it's still a lot. Like, like before I started a company, I had 15 graduate students. So it's not a small, a small group. Like it's a full-time job just manage them. So what helped me a lot is that a lot of them actually joined our company in different roles. Okay, a lot of them are just my colleagues here, some co-founders, some are founding engineers, some my interns. So, a lot of them are helping and their research one way or another is aligned uh with like something we're doing here. So, it's not like, like I need to learn about completely different problems.
That said, there is about, you know, 40 to 45% of uh my other students that do research on things not necessarily directly related to what the company is doing. And I still do meet with them every week pretty much to make sure I know about, you know, what their research is like. It's definitely challenging. I was also, I'm still a faculty member at Vector Institute. So, you had to learn how to wear uh multiple hats and manage things. Then unfortunately only 24 hours in a day. So you still like you know need to balance your, your workload. But as being a professor and researcher before you learn how to manage your time well. So it does help a lot but there are periods of time you're like super loaded. That's for sure.
Jake Aaron Villarreal: Yeah. Well congratulations for being able to make that leap into what I think is one of the most transformative times of technology. You talked about it earlier in AI, and you know, I went through the dot-com boom and bust and it's been 24 years since the real big transformations come. You know we had cloud, we had mobile, but this is really impacting almost every sector. And there's so many opportunities not just from companies that want to figure out what they can do with their data and build models and you know optimize their company, generate revenue, reduce headcount, whatever it happens to be...
[Understanding AI's Three Pillars: Algorithms, Data, and Compute]
...but also just in general a lot of startups that are innovating and building and growing. And a big part of that is the cost you know, and you talked about this earlier in our pre-con conversation which was, you know, three areas of AI... and I'll talk a little bit about it, you can fill in the details. But you know, one is the algorithms around the AI space, two was the data, and then three is the inference and the compute. And costs start to escalate when it comes to compute. So I want you to talk a little bit about what problem you're helping solve in the AI space and where you fit in to that three stool, three-legged stool.
Gennady Pekhimenko: Yeah. So definitely as you said there are, there's three major pieces here. One is algorithms and models, something that belongs to the pure ML, pure AI invention and Gen AI is an example of it. Then the second one is data. This is something data scientists work for many uh for many years. You need a lot of data to make this models to function. And the third one is compute that involves both hardware and software. So our company CentML is essentially uh um contributing to the third piece. So as um you might have heard from me before, is even that compute exists, you need a lot of it, right? And to make things worse, you um if you don't do the right work on optimizing your model and mapping it well on that hardware, you can waste a lot of it. So what CentML fundamentally helps you to do is using the minimal amount compute to get the maximum possible performance you can get, and that dramatically reduce cost.
We also try to simplify the ease of use because uh important aspect of AI is uh that the people that it will be available to a wider uh number of people. Like one um of the special moments in that you know AI journey was the appearance of ChatGPT, right? Like everyone knows about it, everyone remembers when this happened. What people forget is that the GPT models were actually available for a long time. And some, some of them were open sourced by OpenAI. And the reason I believe people don't understand how transformative it is is because the interface to those models was using Python programming language, right? Like PyTorch and other models, which yes we as programmers use and like we already have seen the benefits. But the minute the API changed to a human language, like, like typing in English and you get your answer, a lot of things change and it's not trivial. There's a lot of engineering work goes into that. Right? So essentially uh we are trying to help...
[Cost Implications in AI: Training vs. Inference]
...with that aspect as well. We try to reduce the complexity of building future AI and generative AI applications for people that involves reducing cost, improving performance, and uh making it easier for them to develop new things using that very powerful technology. So at the bottom of it, like we essentially have two product. One of them is called CServe, which is an LLM inferencing engine. So this is essentially a framework where you can bring your favorite model, open source model or your own personal model like LLaMA, Falcon, Mixtral, and tell us what the workload you want to build. You say, "Is it a chatbot? Is it question and answering? Is it a retrieval augmented generation, RAG? Is it something else, like an agent?" We would provide you all the right settings to set it up. And pretty much with no lines of code you need to write, you can get everything running um you know on the hardware of your choice. As a part of that process uh we help you to find the best hardware for your model. This is what's missing for a lot of people. A lot of people don't understand what exactly the complexity there, how to do it. They just trust something they read. But uh in reality, that, that what's the best for you, it's changing every day. So we are helping with our planner tool within CServe to find what's, what's good for you.
Another product from us is called CentML platform or CCluster. So essentially what it does, it helps companies like enterprises to manage multiple AI workloads they have. So training, fine-tuning, inference, we would help to orchestrate, schedule, co-run. This, again, to improve the efficiency even further because you can lose efficiency on running individual model, right, inference or training, but there's also efficiency loss when you co-running multiple things on multiple different chips at the same time, right. We're also a big contributor to the open source. So we have a product called Hidet, which is a machine learning compiler that we open source. We have a tool for performance analysis of AI workloads called DeepView. That's also open source and you know engineers are playing with it already. We also contributed to the open source products like uh vLLM um out of University of California, Berkeley. So that's a good uh research prototype that a lot of people are actively using right now in developing new models, testing them uh for inference. Right. So we are big contributors to that.
Jake Aaron Villarreal: Talk a little about the cost. When you, when you're thinking about inference. I mean, if you look at training data, you know, you start off with gathering the data, and then the model has to be trained, and then inference is providing what the system has learned and giving that back to you. So if you're putting in questions on ChatGPT, you know, it's quickly giving you the answers at a very fast pace, which is great, but there's a ton of expense behind making that machine work. Um, can you give us some kind of rough numbers of what the listeners might kind of understand when it comes about co- when you talk about cost?
Gennady Pekhimenko: Yeah. So I'll give you a perspective between training, fine-tuning, and inference so you get a feeling for the cost. So everyone heard about training cost because it's one, one-time big, massive, and everyone knows about it. So the common example people quote right now is GPT-4 rumored to cost almost $100 million to train, including all the processes associated with it. Right? Uh recently uh you might have heard the news how Elon Musk, xAI was building a supercomputer with 100,000 of H100s. So this is typically done for training. The thing about training, it's usually a, a period of time, say a few weeks or a month, where you're using many of the chips all together to train a new model, and this is where huge capital costs come in from. There are other aspects of it, it's like post-training labeling and all this data, but like ultimately it's a, a lot of compute cost.
Inference, it felt for many people like, "Ah, that's like a lightweight, just running one input, one output." So that gives you a feeling for what the inference cost would be. So one of the largest models that are available in the open source today is the one uh the model called LLaMA. And the LLaMA 3.1 has this largest open source model I've ever seen, was a 405 billion parameters model. For reference, rumors are this is actually smaller than the models that OpenAI is using. So, but still a good representative uh model because it's uh you know, people are talking about trillion parameter models. So that model alone, to run it, you actually need two huge Nvidia DGX machines. So those are like huge servers with eight GPUs each, right, running like kilowatts of, multiple kilowatt, 10 kilowatts of power probably. Each of these machines cost, you know uh MSRP price is going to be on the order of 300,000 uh you know US dollars. And you need two machines to serve one of your single requests to do it. It's not going to do it uh, called the time, because you just send it from time to time. But anyway, you need that much just to serve like one request. And this is like a huge amount of compute. This is like orders and orders more than what you normally do if you do just a Google search for example, right? That would do it using CPUs. So the next, so first of all you need a lot of compute to even do single request.
The ne- the next problem with that is the scaling law. The people who use the model, they, if it's a good model, it's exponentially growing. That's what ChatGPT observed very quickly, right? Like when you had millions of users that are generating traffic all the time, you actually need many of these machines all the time. So this is why we observed very quickly that you can spend many, many millions of dollars of inference because you constantly run it on many inputs. Because it's humans generated inputs. It's agents uh and other requests doing it. Like imagine a bank doing uh fraud detection. So we'll take any PDF...
[Identifying Ideal Customers for CentML]
...statement from any customer every like month or every week and run it through this tool. So the number of requests going through this model is enormous, right? There's billions and billions of requests. So a huge baseline cost multiplied by a huge number of requests sent. And don't be surprised that inference very quickly outpaced in cost training. Nvidia recently report that of all the GPU cycles that they've seen in the space, 45%, 40 to 45% is already in inference, not training, right. It's already spent there and it's keep growing the relative portion of it. So the number of computers getting more. Why this happens, it's very natural. We train the models to be used. So after we build something good, it's going to be usually used more, and the cost of using it would be out-, outpacing. Also it's constant, so training you train for two months and you might stop. With inference, you get this like you know steady increase in traffic because a good model will be used by more and more and more and more people. There's more and more requests goes into that, right. So that's why you need a lot of uh compute to do it and that's why it's so expensive. So it's the sheer size of the model multiplied by the scale of how many people are using those models.
Jake Aaron Villarreal: Yeah. Wow. That's incredible. So, you talked a little bit about your product and your different products that are helping the industry and companies. What's an ideal customer for you when you go to market and you're going to talk to them? Who are you talking to in that customer and what are you showing them that they say, "Oh, wow. We haven't thought about trying that," or "That's something new that we should explore"?
Gennady Pekhimenko: Yeah, great question. Um, I think there is a dream customer obviously, but there is a reality. So essentially there's several customers that are ideal for us. And the major characteristic of the customer that we really care about, if it's in one word, it's maturity. So we want the customer that actually has a use case, right, for AI. Because we're not the company that build use cases for them, right. And uh in general uh and that's very important, being at a big enterprise like Deloitte or a small uh you know Gen AI startup, maturity of them to do it is the key requirement for us, right? So they need to know what they're doing with AI, then we can help them. So if they have that scale or they plan to have that scale, this is where we're helpful.
What does it mean in reality? So for me as for the customer, like one of the strategic uh you know focuses of us uh for us is enterprise customers. Why? So my personal belief for AI and Gen AI to be successful is that it has to be adopted by this Fortune 500, Fortune 1000 companies, right? That has to happen. Why is that? Because the fact that OpenAI can burn investor money or Google or Meta can burn their own research money doesn't prove actually that the technology is useful in practice, right. They can experiment with anything. You've seen how there are people like using for example, spending billions of dollars on AR, VR technologies, right, you know, and metaverse, and not necessarily there's an outcome for a regular user immediately. So for me the big success of Gen AI if it's used for companies like Deloitte, Bloomberg, Walmart, Visa, right? So if all these companies generate revenue by that, for me they are like the second wave of adopters that had real data, real customers. And if they see value from that, that's a real, real big success, or the financial institution. So for us, one type of ideal customer is one of those enterprises. They reached the stage where they know we want to do it, they want to do it and they had a model in mind and now it's the time to scale it. And that's where we need it. Because when you experiment, cost is like minor. Well like, I'm experimenting with one customer, one model, it's like the cost is like hundreds of dollars. Who cares about optimization? The thing they see is the minute they said, "Oh, now time to apply it for like at scale, for many customers, for many workloads." That's where the price tag goes up and companies like us are very useful. So that's one ideal customer.
Another uh really good customer for us is uh um uh small and medium-sized cloud providers. As you know, there's the top three cloud providers, Google, you know, GCP, AWS, Azure, right? But there are a lot of others. There are guys like Oracle Cloud, Lambda Labs, Nebius, San Francisco Compute, Crusoe, Denver Data Works, uh, there's plenty of them. I talked to more than a dozen of them, right? All these guys had hardware, usually Nvidia GPU, sometimes AMD GPUs, right? Maybe some other chips, but mostly Nvidia GPUs these days. They want to sell it to their customers, right, and they want to have a software stack built on top of that. But majority of the junior companies, they are not...
[The Importance of Solving Real Problems in AI]
...like AWS or Google. They don't have a mature software stack like SageMaker or Vertex, right? So they want something like that and they're ideal customers for us as well. So I had many conversations with companies like that because they want to offer a mature software on top of their hardware and they don't have resources of their own. So we are a good fit for them. So we partner with many of them, right, and had a joint go-to- uh market motion. And obviously we're not going to drop our friends, other Gen AI startups. They all come and they are usually easy to work, right, in the sense that they know what they want, right? They know what they're doing. They're just small teams and they need us as a part of their pipeline, right. So we never say no if they had a reasonable use cases. We always onboard them as customers as well. So there's several types, so there's enterprises, medium size cloud providers, Gen AI startups, each of them can benefit from us. In addition to that, we obviously partner with large guys as well. So, we partner with companies like uh Google Cloud, GCP, with AWS, we partner with Snowflake, and of course Nvidia and AMD, right? So, how you cannot partner with them?
Jake Aaron Villarreal: Yeah. No, that's great. Thanks for sharing that. Um, why is this space important to you?
Gennady Pekhimenko: So this space um I always like to pick the problems like forget about me being um you know in the AI space. Like just me as a researcher in general, like why I picked that before it was booming and there was money in the space. Right. I just see that this problem exists and it's real. So I finished my PhD and before starting as a professor, I went to Microsoft Research like for one year. And that, my role there, I was in systems research group. I had a privilege to work with amazing people there and get a complete freedom to pick the topics I would like to work on as a researcher. I went around, and that's what I like to do with industry all my life, like I look at where the real problem are. And I pick the one that exists, but they don't even understand that is a problem yet. But I'm like carefully looking because that's what's going to drive my research and in this case even my startup for years to come.
So what I noticed in 2016 was like plenty of people waiting for the time on the Nvidia cluster with the old GPUs that existed back then. And they were like complaining. All these ML researchers were complaining about, "Oh my God, this is like, like how long it takes, it's like there's a long line of people, everyone wants this GPU." And I was like, "Guys, how is... like, it seems to be very precious resource, like you really need this compute. How well it's running?" And they're like, "What do you mean?" It's like for them it doesn't even make sense, my question. Like, "It's running on a GPU. GPU is busy, like your question doesn't make sense." And then I realized that there's a huge semantic gap between the ML researchers and systems people, computer architects like me, who knows that you might be running and running like 10% utilization. So what I found very quickly there, to get access to the...
[Building a Strong Team from Academia]
...servers, that utilization is typically between 10 to 15%. So all of them waited in line, but they were only using 10% of the, or 15% of the cycles available of those machines for many reason: of the inefficient software stack, how they write the models, many reasons. So it's like "Oh my God, it seems like everyone needs this resource." There is, was already a trend for like four years uh back then, every year the model is growing by 10x. Like, like a clear exponential trajectory of this models growing. People need that compute. Compute improves maybe at 2x per year at best. So you're going to get like a 3 to 5x gap growing every year, and after that it's continued that exponential low for another six years. So the models we used to train like back then, the cost was hundreds of dollars, maybe a thousand at max. Right now we're talking about 100 million. That's what this law cost.
So for me it was really serious problem that exists in industry, not a made-up problem by academics. And um it was really interesting. It was not like simple solution, see like a huge disconnect between ML researchers and like systems people. Neither were aware of what others are doing. So it's like, "Oh my God, there's such an great opportunity there." So I got motivated because it was like a real big problem and it was not easy to solve, clearly not easy to solve. So I was like, "Okay, that this is like why I'm doing research." So, and then while doing that, I learned it's like "Well, actually doing quite well in this space, so maybe I should you know commercialize this."
Jake Aaron Villarreal: Yeah, it's a fascinating story. Um when you started the company, you ended up bringing on your own students. What a great way to hire. What a great way to really build expertise into your company really from the beginning. Um and now...
[Hiring Strategies for Startups]
...you're roughly at 40, 45 employees. How big are you today?
Gennady Pekhimenko: It's about 45 people. Yeah.
Jake Aaron Villarreal: Okay.
Gennady Pekhimenko: Across this.
Jake Aaron Villarreal: Yeah, you know, technology is one aspect of a startup or a company. It's the people behind it that drive the innovation and continue to make it uh successful based on how they execute. What's been something you can share in terms of your learning when it comes to making the right choices on people? Maybe it's your interview process, maybe it's uh understanding the domain space and you can really calibrate if they bring value to the table. But what's really worked for you in making the decisions for those key hires that other founders might be able to learn from?
Gennady Pekhimenko: Yeah. Well, first of all, I definitely had to be honest. I had an advantage, right? Like, like a lot of people as founders build and there's like two or three founders and that's how they start the company and they need to fund raise in this conditions. I first of all was an academic. I already had some IP built, right? So I had an evidence to investors that can build a valuable IP, right? And in addition, there was like a team that was willing to join me. Like they were like already like day one when we started company I had nine people like in the company, right. So it is an advantage, and but it was not like a random advantage. I know, I build my team so that there will be a lot of good people like that. So why is it important, like a lot of um um, it's really like there are certain general qualities that are needed for you to hire. And then a specialized one in the AI space, being able to go with the research and follow the state of the art, the latest and greatest, requires certain skill set, and I happen to hire a lot of you know students to have the skill set already. So that's helpful. Um, um that's not a general advice, but more like, in general, I really like to hire the people that um had substantial potential for growth.
So I'm not hiring like a, a toolman that can do one feature, nothing else. I like the people that actually had ability to grow and learn. A lot of people are afraid of that. There is a risk they can, you invest into them growing, right, and then they leave you. Like, I don't, [I'm not] afraid of that. I normally like you know graduate students when they reach their peak potential, and I'm okay with that, but uh I think it's very important to create the atmosphere where people can grow, right, so it's very important. So obviously uh uh their ability to build in this very sophisticated technological space is very important for me to hire. So we look for those technical abilities. What we try to avoid in our interviews is focusing on what they memorize. So like a lot of people ask is like, "Oh how great man is how we solve..." Like people, person can be trained similarly to a model on solving fixed standard task quite well and not being able to do anything practical afterwards. So we always try to see how uh how their thinking process is as important as them getting the correct answer in the questions. So how they think out loud it's very important and sometimes if they get the questions too quickly, you know they just know it or they memorize it. So you just change the problem immediately, right?
So um that's an important one. Another thing is very important here. People don't work in isolation, at least not in startups, right? And definitely not in our company. So they need to be a good team players. So it's very critical. We don't hire like those you know uh single-man type of teams that like they do everything and don't talk to anyone. Uh it can be very, it can kill the atmosphere here. And remember, startup has pros and cons versus Big Tech. It's like our advantage is we can move faster. Like, like Google, Amazon, they like big corporations. So steering them up is difficult. We are very adjacent and we can quickly move around. We are very agile. But uh you have no tolerance for mistakes. So Google can make like multiple like mistakes here and there and they're still going to be successful. I can't afford that. So every single hire is super critical. I can't afford to hire someone relatively...
[Competitive Advantages in AI]
...senior to lead the team and say, "Oh, like a year later, I'll change them and things move like as normal." No, it could be a disaster. So, the selection process has to be actually even higher bar to what you would expect in this like you know famous companies, right? So, you need to find someone who's good technically, who's a good team player and the people that actually like to have, they're willing to take some risk. Like, let me be honest, startups is about taking risks. Like it's not like it used to be like "we are paying really well." Like we in Canada can match like, like, like all the top companies. And so similar in the Bay Area, pretty much everyone we can match, right. But like, you still taking a risk on the shares, right. So you don't get in standard RSUs that you know it's convertible tomorrow. But that's where you take your risk, right. So the people should be reasonable risk takers. There are some people that just like slow pace, they would prefer large companies. But if you want to move fast, grow fast, and take a little bit of a risk associated with that, that's the people you're looking for in a startup. And people need to be decent generalist. Like you're not going to do in one thing for like three years. That, things would adjust and things move quickly.
Jake Aaron Villarreal: Yeah. Well, you know, you've talked about your product and you talked about your people. What's the competitive advantage you have today? Because there's a lot of competition out there building multiple different solutions and tools to help support you know this big transformation. What's your competitive advant- what's the moat that you're building that aside... maybe it's your IP, maybe it's you're ahead of the market. Like kind of walk us through that a little bit and why people would be interested in join your, joining your company.
Gennady Pekhimenko: Yeah. Well, several things is like there was our competitive things in terms of competitors and on the market and like in terms of bringing people. So, in terms of uh our IP and technical abilities, like one of the strongest pieces that we have, I would say uh against many, many different competitors is our ability to optimize things across full system stack. So like we can take the model when it's written in something like PyTorch or JAX and then we had our own you know optimizer, graph compiler, all the layers afterwards. So we had our inference serving engine, we had our own platform, we had our own compiler. Very few company, if any in the world, actually can claim the same. Most of them would be just in one of these little layer across the stack and everything else is standard.
As you can uh hypothesize even without going technically, if I had much more uh bigger piece of system stacks I can do way more creative with my optimizations, right? I can do way more than others can't. So I can optimize the models knowing it's running with something else because I had my platform so I know I can generate the code that would run with something else. I can co-run them in a friendly fashion. Someone who does just the compiler, just the platform can't do that because the platform has to deal with whatever is given and the compiler doesn't know what's running after it.
[Challenges in the AI Industry]
Right? So, we have all these different advantages like across our competition including even like big tech companies. So that's one of the advantages. And in terms of uh uh you know advantages for people to join the company again as I said it's, we're paying competitively. Like it is exciting environment to be everyone. Like we had actually um a very good numbers in terms of attrition. Like, like it's very rarely someone wants to leave. Like extremely rarely, because it's very interesting job in a way, right? Like it's you're not going to be bored, you're going to be working around with many people have PhDs, right, people or others that are really good builders and designers. It's very diverse environment. We had a lot of different nations in the company. It's a very typical, you know, Canadian startup that is very diverse, right, with a lot of people from all the different, from all over the world. It's just a very nice conversations during lunchtime because of like how many different nations we have here.
Jake Aaron Villarreal: Yeah, that's great. What do you think are the main challenges today in the industry in AI and, and moving forward?
Gennady Pekhimenko: Yeah, there are several major challenges that still exist, right? And like one of the challenges in the industry right now is again, I mentioned, alluded to that before, is a maturity, right? So for, for the startups to exist, they need to get to revenue relatively quickly, right? And it's very important there would be people that benefit from that. So uh the, the most tricky situation here, it's not that there is any lack of interest. Like almost any Fortune 500 companies had Gen AI strategy right now in some form. And they put money, millions of dollars in it. But they might still three months, six months away from making their first product. So we need to find the one that ready and to find it fast. And like the people in our space, that is a challenge. And I think for the industry as a whole, it's an important one to show that there's an evidence there. Because like again, yes, there's companies like OpenAI, but and Anthropic, but and Cohere. But they're not a...
[Navigating Company Growth and Market Changes]
...lot of companies that actually generate meaningful revenue, right, without cheating with the numbers. That's very important for the AI as a whole to show ROI because all these numbers, like all this money go into the area and people would be very frustrated there's no outcome. Right? So again, I, there's definitely hype around AI and eventually that hype is going to blow up. But I, I also strongly believe that there's a substance behind that, so we need to make sure when that burst happens, there's a real substance there, strong, and it would stay solid. Right. And then after that we move like business as normal, there would be less hype, but it would actually keep delivering and delivering. So my hope is that we're going to make AI similar to Industrial Revolution essentially, where it's going to transform the society for years to come. We're going to, don't like, "Oh my God," dream about it, we're going to keep building the future where AI is a part of everyday's life.
Jake Aaron Villarreal: Yeah.
Gennady Pekhimenko: And that's one of the challenge we need to solve as a community overall.
Jake Aaron Villarreal: Yeah, that's great. Well said. That's the challenge for the industry. What's your challenge as a company? What's the biggest challenge you face today?
Gennady Pekhimenko: Yeah, so for, for us at this stage is organic growth. So it's like, you can raise too much money and then you grow too fast and then like you need to keep up your growth with, there are like so many moving targets. So like, like if there is the company growth, there's individual people growth, there's a market changes. Like the market is big right now and no one owns any significant portion of the market, but there is still, as I said, lack of maturity, lack of clarity in certain things. Another obvious challenge for everyone who works in the space [is] in how frequently things change. For example, the product I mentioned, CServe, it didn't exist a year ago. We just start to build it a year ago after everything we learned by working with a few customers. What's interesting, none of the products we compete against, either in the big tech or other startups, don't exist a year ago either. Right? The everything is super fresh just because the things that LLM bring to the world change things so much and optimizations, all of that requires rethinking everything. And we already went through several layers of that rethinking, and there are more that like every six months things change. Like you've seen, then there are substantial changes where what people do and what is the most exciting things changes every, you know, 3 to six months.
A year ago people were very excited about fine-tuning. Why? Because LLaMA models proved to have an, other open source models like the one from Mosaic MPT and the one from uh the Falcon model, uh Mistral model, proved the open source can be not necessarily matching OpenAI, but can be competitive. Right? Reasonable. And with fine tuning you can make it even better than GPT-4. So it's like "Oh my God" moment, "I don't need to train my model anymore. I take the open source, fine-tune it a little bit, and then it becomes you know uh you know better than you know these models from OpenAI." That was mind-blowing for a lot of people. And why fine-tuning is exciting, because training was like 10 millions to $100 million process, fine-tuning you can do in hundred thousands, maybe half a million dollar budget. Right? So that's was important moment.
Then uh this spring uh a lot of people realized that even fine-tuning is difficult for them. Not only the cost, but also engineering capabilities. So a new thing applied called retrieval augmented generation where you can take pre-trained model and just add your data as a database. So essentially why it's cool, you have no knowledge on how to do fine-tuning. You just take your favorite model and take your data. Like you can... like we use like our website, you can take all the content of your website with your papers. All of a sudden that model combine this tool using retrieval augmented generation can answer about your website in a meaningful, very high quality English. So it takes the linguistic knowledge from the foundational model and then it takes the specifics of the data, the context from your data. So it can answer the question it would never be able to answer with ChatGPT because it's not even a part of the training data there, but it would answer it. Like you can build this in a minute, right? Like, and it would cost like, like that's what offering. That's right now.
There's another shiny thing called agents or agentic system or compound system. So the...
[Leadership Evolution in a Startup]
...way to think about it, people won't remove human from the loop. So right now we talk about single model used for a single use case, but like in all the cases the models generate output that goes to human and human might involve another model for another task. Well, imagine human is the slowest part of that equation. They won't remove it. You want these agents start to talk to each other. So that creates a world where those models talk to each other without human, and human just sees the final result. That's very exciting, right? That people building system around it. And I'm pretty sure six months from now on that will be another exciting, you know, things.
Jake Aaron Villarreal: Wow.
Gennady Pekhimenko: And you need to try to stay on top of that because things are moving and they're moving every six months.
Jake Aaron Villarreal: Well, that's a pace that uh is accelerating faster than I've seen since I've been in, in this uh space for many, many years. Um, as a leader, you've had to evolve, I'm sure, from starting from a few students to now 45. What's been a breakthrough for you as a leader in terms of leadership for your startup?
Gennady Pekhimenko: Yeah. So um I managed a decent size research groups, but it was still very different relationship. It's still more a classical like mentor, mentee relationship when you advise students. And it's more like one-on-one and there's a group. Company is very different. First of all, it's, it's way more diverse. So I hired for engineering role for many years, right, and was in part of different interviews. But all of a sudden I need to hire VP of Sales, VP of Go to Market, I need to hire a marketing person. Like this is like enormously difficult for someone with a technical background like me. Why? Because this guy's amazing at talking but it doesn't mean there's a substance there, right? So like, good luck finding a good like, like you need to learn how to read what they say, how to read their profiles. It's like, like I need to learn all of that.
Then that, that's one thing. I need to understand how to build the team, like a proper company build uh like, like building... like it consists of all these different components. Then my role as a CEO making sure that all of that is connected well, right, so that the salespeople work well with product team and with engineering team. This is all different people, so that all nice and smooth. Another challenge is obviously like, I don't, I won't believe anyone could say that this is like an easy for them. Fundraising is never easy. Right. It is like you're going through an exam...
[The Importance of Strategic Partnerships]
...right? Like, and sometimes you're going through an exam like five times a day, or even eight times a day, right, when you're fundraising. So I raise money for my research group, but it's very, very different, right? From actually raising many millions of dollars for, for the product that you're building, where like a lot of things are still up in the air and not clear yet. Uh, it was quite an experience. I had to learn the whole different part of the world I was not involved [in]. And I like it, I actually enjoy like going through the you know all the learnings, but it's definitely stressful and like it was like a, a breakthrough and important experience for me.
Jake Aaron Villarreal: Yeah. Well, you know, attracting people, whether it's new employees or customers or even advisers or investors, it takes a certain skill. You got to be able to present, build trust, have people believe in your vision. Talk a little bit about the all-star team you've put together when you talk about your investors, the advisors. Name drop some of the, the names that are part of your, your, your, your team there.
Gennady Pekhimenko: Yeah. So, like um uh an important part of when I picked the investors, I actually didn't pick the one that gives me the most money or the best price for the company. In fact, I turned down the offers in each of the round of finding people with the bigger, higher bids, right? The reason why is very important to build the proper relationship with them. The people that joins you, they join your board, right? And you're building and growing the company with them. So, they are part of the important decisions made in the board meetings. They are in regularly, sometimes weekly, bi-weekly calls with you. And like how to attract customers, how to hire people. So they are literally the early stage investors, pre-seed, seed, series A, they help you grow the company. So that's extremely important. I can't overemphasize how many people don't realize how critical is to be careful about selecting people. So it's not about just like, "Oh like I need to get this many uh much money, after that I get the money I will do it myself. I come and bring them results two years later." It doesn't work like that. You actually building company together with them. If they're good investor, they will help you with growth. They don't just give you money, right? They give you connections, advice, a lot of different things, right? And that's very important.
And obviously like, you trying to hire the best engineers in the world to run in your company. Like we recently, for example, hire like one of the like probably top five people in the world in machine learning compilers, and very excited about her uh joining us and like being a team lead of our compiler team. So all of that is a good experience for me, trying to get people of that caliber excited to join us with all their like...
[Future Roadmap for CentML]
...20, 30 years of experience, right? And they had no shortage of options, especially in the Bay Area, right? They can go anywhere. So they really decide what they believe in. And you can [show] them, like they understand the tech, they understand what you're building, they look at the quality of your team, they go and assess you as well. So every interview with them, like you interview them, they interview you as well in parallel. Right? It's all bidirectional all the time. And it's important to realize where you bring this talent. So it's again, it's like a personal challenge that you need to embrace and face and try to be good at it.
Jake Aaron Villarreal: Yeah. Got it. Well, we won't talk about the names, but it sounds like you've built a great team here and really excited to see where you go. Uh as we head into 2025, what are you excited about? What's on the roadmap for CentML?
Gennady Pekhimenko: Yeah. So like we build the product that we see there significant potential they will, can improve. Like another import- uh the very important things for us moving in 2025 is keep scaling it. So trying to get you know the feedback from the first dozen of customers we have, and you know, adjust and improve. It's like the iterations are critical. So like we hope to get more and more customers through the team, keep improving, iterating the product. So for us 2025 is going to be the year of growth. Of not, we're not building the product right now off the ground. We are actually uh, the product is used by the customers. We're getting feedback and we acquire more and more. So it's the growth of the customer base is the growth of revenue is u- ideally future growth of the company. Hopefully you know uh new fundraises, uh new offices open. So we have two offices right now. We hopefully will get more.
Jake Aaron Villarreal: Yeah. Well, amazing story. I love what you're creating. I think there's a lot of companies will hear this and uh for sure we'll be reaching out to you. And if they reach out to us, we'll send them to you. Uh I want to thank you. Yeah, I want to thank you so much for your time and sharing your story. Um and for all the listeners that are listening, appreciate your time spending it with us today. I'm your host Jake Aaron Villarreal signing off for now. We can't wait to catch up with you all on the next episode. Until then, take care. If you like what we're doing, don't forget to subscribe, leave a review on Apple Podcasts or wherever you listen, and follow us on YouTube, where we go behind the scenes to learn what it takes to be a startup founder.