Jake Aaron Villarreal: Welcome to our podcast, Born in Silicon Valley, where we interview startup founders exploring their journeys, their success, challenges, and lessons learned. We hope you'd be inspired in discovering what it takes to build a thriving startup. I'm your host, Jake Aaron Villarreal, and here with us today we have Vikas Nair, the founder and CEO of Openlayer. Vikas, welcome to the show.
Vikas Nair: Thanks so much for having me. I'm excited to be here.
Jake Aaron Villarreal: Great. Well, a little bit about Vikas. If you're building AI models, first and foremost, you don't want to miss this episode because it's how they're helping companies reduce the bugs that are in those AI models and how they go about doing it. So, a little bit about Vikas. He's a product engineer and founder working in the AI space. He began his career working at Apple on the Vision Pro as a machine learning engineer building multi-modal AI systems. From there, he and two other engineers spun out of Apple to build an advanced system for testing AI models. They joined the startup accelerator Y Combinator shortly after to create a company out of the endeavor called Openlayer, and have since shipped their product to many enterprises and startups to help them build more powerful and performant AIs. Vikas, you're in the right spot at the right time, and we're going to dive into your company. Before we do that, where are you joining us from today?
Vikas Nair: Well, right now I'm calling from Cleveland, Ohio. Uh, which is sort of the armpit of America. Uh, I don't want to offend my fellow Cleveland brethren by saying that. Uh, I'm here home for the holidays, but normally in San Francisco, uh, right in the, in the heart of San Francisco.
Jake Aaron Villarreal: Cool. So, you're, you're originally from Ohio?
Vikas Nair: Yep. Born and raised Midwestern guy.
Jake Aaron Villarreal: There you go. All right. And um, you've, now you're in San Francisco. For a lot of companies, they move out to Silicon Valley because it's about being where technology innovation's happening or at least what they call the mecca. Others, it's the company lives there and that's why you have to go there. What was the reasoning for you to, to end up in Silicon Valley?
Vikas Nair: So, I'd say that the motivation for ending up in Silicon Valley came before we founded this company. Uh it came from when, you know, I, and also when my co-founders had all collectively joined Apple, as you mentioned, to, to work on, to work there. Um it continued because a lot of the, you know, efforts that are happening in the startup community obviously exist in San Francisco, and as well sort of a lot of the development on AI goes on in San Francisco. And so now, you know, we get the opportunity to work closely and personally with many other peer companies who are building in the AI space simply by being, you know, next door neighbors to them. Um, yeah.
Jake Aaron Villarreal: Yeah, that makes a lot of sense. Um, we've had guests on that were next door to OpenAI and they were able to walk next door and grab lunches and see Sam Altman and really just being around... it's easier to, to take calls, to meet up, to learn, to grow together and lots of uh reasons for being close by proximity to others that are innovating.
Vikas Nair: We've been invited to, to be able to attend some Q&A with OpenAI, too. And um that wouldn't have happened if, you know, we weren't located here in San Francisco. They have a fantastic office in the Dogpatch, I believe. Uh and Sam Altman gave a Q&A there. So, we got a lot of insider intel on what's going on there.
Jake Aaron Villarreal: That's cool. You know, let's talk a little bit about Apple. You were there at Apple. Uh you joined Apple, incredible company, innovating a lot.
Vikas Nair: Absolutely.
Jake Aaron Villarreal: Um talk a little bit about the Vision Pro. I mean, we see Oculus and we see now Apple coming out with their headset. Uh, this multi-mixed reality visionwear that looks incredible, maybe a little pricey, but still looks really cool. What were you helping build within that product?
Vikas Nair: So, Apple's angle in the VR/AR space is unique. Um, they were the trend setters for what is termed XR or MR, which is short for mixed reality. So there were a lot of these other, you know previously Quest, for example, which was sort of the flagship product in that space, was building specifically VR experiences. So you would have this, you know, headset which I would say was bulky and would cover your vision, right, and replace the world around you entirely. Apple, in the meantime, wanted to create a more ambient device by blending both what's called VR and what's called AR.
So, for those who aren't super familiar, VR, of course, is a fully virtual world. And sorry, XR or AR, rather, is an augmentation around your real world. So, you're still seeing the world around you. You're still interacting with the same world that we all interact with day-to-day, but it's augmented by adding visual overlays on top of that experience. So for example, you might be walking around, you know, the sort of canonical example of this is the Google Glass project, which many would say is a failed project, or the Snapchat goggles, where you would have notifications that could come into view um as you're walking by on the street or whatever it is. Apple wanted to take a blend of those two things.
So they wanted to have a, a lightweight headset that you could put on, and in its primary... in, in one mode, you would be in a fully AR experience. So, you'd be sitting in your apartment and you'd be seeing your apartment, but you might have notifications that come in view on your, on the top right corner, like your messages. Um, and then they have a dial that as you turn it, your world is slowly replaced and you enter a virtual world instead and you... in something that's completely different than your apartment. So, you could be living in a, you know, some other kind of place entirely. You could put objects that would not normally exist in the real world. Um, so that was the key differentiator.
And the other difference is that they didn't want to have handsets, hand controllers like the Oculus had. They wanted to, and this is prior to the Quest 3, which came out after Apple shipped or announced their ship, you know, their ship. They, they wanted you to be able to just interact very seamlessly with it by using your hand gestures. And so with all of this, what we were building as part of, you know, uh, for the Vision Pro was a completely new type of interaction modality, an interaction experience where you would use your hands, you would use your eyes, because they're beaming lasers into your eyes so they can track your, your, you know, where you're looking, and you can use your voice to interact with the world around you in a more seamless and less clunky way.
So I was part of the Siri team, and what we were doing was because we don't have hand controllers, because, you know, there are no, you know hardware devices to, to... for example, type or something like that or point, you need to be able to completely just use your, your, your, your senses to compose, for example, a text message. So, what that means is that you need to be able to, if you get a text message and that comes in the AR overlay above your vision, you need to be able to just look at the text message. And as you look at it, if you gaze at it for more than a second, that bubble will activate because it knows you're looking at it. And then you could just naturally just start speaking and dictating.
But normal dictation, as we know, is kind of finicky, right? It often makes word, you know, errors. You know, uh mis-, you know, if you say my name for example, my name is kind of complicated. It doesn't sound like a normal American name. So it'll often say "because" instead of... so normally you had to go back and you had to stop dictating. You have to go retype what you really, what you meant, right? With, with the you know, uh your, your gaze as your, as another input modality, what we can actually do is you can just look at the word that was misspoken or that was mis-transcribed rather, and you could just naturally re-speak it. So I could just look at my name and, or "because" rather, and I could say "Vikas" again, and it should re-correct it. Or if I want to make a new line, I could just look down and it'll start, and keep talking, and it'll naturally create a new line.
So that was one of the big experiences that I was working on there. And as well there were other, you know, there were other kinds of experiences that involve like object placement. So you might be looking around your room for example, and you might see, you know, a fridge. And you might say, "Hey, I want to keep... in this virtual world that exists in my Vision Pro, I want to have a calendar, for example, that's persisted on, on the fridge or something like this." The exact details I probably can't go deep into and this is just a toy example. But I could just look at, for example, this object and I could say, "Place my calendar on this fridge." And maybe I have two fridges in the room. And, you know, we, we have the entire scene mapped, so we'll know that there's two fridges. So how do we know which fridge you want to put your calendar on? Well, you can use your eyesight to dereference which object you're talking about. That's called deictic dereferencing. It uses modalities like your voice. It uses the indirect and direct objects of the sentence you just spoke. And as well, it uses your, your eyesight, where you're looking, in order to determine how to resolve that query. So yeah, that, that's sort of um the, the, the, the kind of work I was working on there.
Jake Aaron Villarreal: Wow, that's fascinating. And that product, uh, I'm actually excited to see how it works and how it functions and if there's a use for it within, you know, our company potentially. Uh but personally, it looks really, really exciting. What was the inspiration behind working there and identifying an opportunity to leave Apple and to start Openlayer?
Vikas Nair: It's a great question. So Apple, first of all, was my dream company. You know, I've grown up loving Apple products. I grew up watching all of the Steve Jobs and then the Tim, Tim Cook keynotes um in high school uh and before even. I love Jony Ive. I'm a huge design guy. That translates still in my work today. And it translated a little bit into my work there, even though there was some, you know, separation between engineering and design.
The actual impetus for us to leave came from a funny story which is that we were... for this team that I mentioned, you know, for, for these teams around the Vision Pro where we, we were building these multimodal experiences, of course things go wrong, right? The model is not ever 100% correct. Um and so the job of an ML, machine learning engineer, AI engineer is to... is much, is much less around designing the architectures of these models and more around collecting the correct data so that the model is able to be robust to all sorts of edge cases, right? Because one person's apartment looks very different than another person's apartment, and there's many hundreds of millions of people who use Apple products around the world. So you would expect that, you know, these kind of companies when you, when you're building these models, you need to really collect tons of eclectic data.
And even if you do, do that right, the, there are very often errors will occur. And so much of the actual job of being an ML engineer, as I was saying, is diving into those errors and understanding why those errors are happening and figuring out what data needs to be collected to help mitigate those errors or what needs to be changed, changed in the model. So we were building uh tools for each of the teams that we worked... We actually worked on a few different teams during our time there. Between the three of us, we worked at probably 15 different teams. Many of them were in the Vision Pro, but many of them were also outside um in just general Siri land. So working on, for example, the voice translation model for Siri, like translating your spoken words into text.
And for each of these teams... for the Vision Pro teams, which were very small teams that had not yet shipped a product that people were using, as, and, and only had maybe five engineers on, you know, a given team... and as well for the larger teams that were working on Core Siri, which is shipping to hundreds of millions of users, anybody who has an iPhone. In both of those cases, we kept having to rebuild much of the same tool set for evaluating these models. Which surprised us because you would think that there's sort of, you know, a sort of cohesive... this is an incredibly critical part of the process. You would think that there's a standardized tool for, for doing this and a robust tool. But in fact really all that, you know, we were building over and over again was a relatively simple way of running tests through the model and analyzing, categorizing the kinds of errors that the model was making.
And so no... you know, so, so from that we, we sort of had the inspiration to say like, "Hey, this is, if this is a problem that we're having at Apple, almost certainly it's a problem that exists in industry." And in fact, it is. It's, it's sort of an unsolved problem. Uh because the best you can really do is to sort of, you know, try to guess at, try, try to collect categories of errors and, and then you're kind of left up to your own senses to understand, "Hey, like, it seems like a lot of these errors have this thing in common. They all seem to happen uh with people who have a specific type of accent maybe, um from like some random town in England and we just didn't collect enough data um from people who speak in that dialect, and so we need to recollect more data." And this process of doing what's called error analysis is really finicky, it's an unsolved problem, requires manual labor. And so we set out to try to figure out a way to mitigate that uh labor that goes into understanding and detecting these errors.
And we took this idea, we, we left Apple. We started to iterate on this idea and build an MVP of what we believed to be the best-in-class way of doing this, which I can go, you know, deeper into later. Um, and we applied to this, uh, startup accelerator called Y Combinator, which many of your audience might be familiar with. Um, and from there we started to build a company around it called Openlayer.
Jake Aaron Villarreal: Talk about the team as we get into it.
Vikas Nair: Mh.
Jake Aaron Villarreal: Um because there's a collective brain power around your company that I think is, is worth mentioning. Where did they go to school? What, what's the makeup of your current team today?
Vikas Nair: So two of uh my other co-founders were, as I might have mentioned, also at Apple. Uh we're also all working on the, on the Vision Pro and other teams. Um those two guys came from, one of them came from, from Yale, um where he studied computer science. Uh and another went to both Carnegie Mellon and Cornell where he studied computer science and also AI. Uh and then we recruited, once we, once we started to build this company, uh once we raised a seed round. We took some time to just, you know, one of the advices that we had gotten early on is, is, is not to, not to hire fast. Um not to hire before you feel like you have an idea of what you need to build and there's a real need for it. So, we took some time to really make sure that we were ready and that we really needed the, the talent.
When we did, we hired some people, some people that we had known from our network and some people that we, we had recruited uh cold. We didn't know them. We brought on first uh one of my co-founders' close friends who was from Brazil. My co-founder is from Brazil as well. He was a, a sort of PhD dropout. He went to um uh ETH Zurich for his master's and then Duke for uh his PhD and he dropped out after a month. It was the classic Silicon Valley story. But uh he, and he also had you know, spent some time working on AI at some startups uh around the world. So we recruited him to be sort of like a you know jack of all trades, somebody who is very special, specialized in, in, in machine learning development and what goes into that, uh but who could also play a role in helping us build out our backend.
Um and so we started by hiring him. And then we also hired another, we recruited somebody else who my other co-founder uh knew from college, from his time at Yale. Uh this guy is brilliant. Um his name is Parth. He uh joined AWS, joined Amazon to work on AWS as a backend engineer there for a few years and he was super excited to, to you know, change gears and try something, try, try actually working for you know, a startup that was something exciting, something, something early stage. So he joined on as our sort of primary backend engineer. Uh shortly after we also hired um somebody who was a, acted as a designer and as a product lead, uh her name is Erica. We did not know her. Um but she's been also an invaluable asset to the team. She came from Tufts first and then she went to Harvard to specialize in their Graduate School of, of Design. Um with an intersection of both design and how it applies to AI development. Um so it's sort of a perfect um match there. Uh she joined right after and that completed our team.
Jake Aaron Villarreal: Yeah, that's amazing. You know, you have an idea, you bring it to market, but you're told to hold and wait until you have a product that's viable and then you get funding or you hire and then get funding, however you want to break it down. What was the timeline from bringing the idea through Y Combinator to where you are today?
Vikas Nair: So, we, when we applied to Y Combinator, we actually really did not have much. We had an idea, um and we had built an MVP. And I say that, you know, with what, everything that an MVP means, it was very bare bones. It was, it was just a simple demonstration of a unique technique that you could use to understand errors. It's called explainability. Um where you can sort of see, for any prediction that a model makes, what were the reasons that the model made that prediction. You know, maybe it used a specific feature of the data. Maybe it really heavily weighted on, for example, if it's a model that classifies whether or not a user is going to churn on your platform. We can understand with some, with some science, was it the user's age that caused you to think it was, they were going to churn? Was it the user's geography? So on. So we built this demo that showed this capability. We showed that to Y Combinator. They were excited by the idea, and they were excited by, you know, our experiences and our, you know, from our time at Apple, from, from what we had learned building, you know, models at, at scale.
And then we spent the three months of Y Combinator, during the summer batch of 2019. We spent um those three months really focusing on building out um a more viable product that could actually be implemented you know for a customer, um and that could be you know used to bootstrap somebody who's building a startup at the early stage or even working on you know, working at an enterprise, could help them actually start testing their models in a serious way. So we spent three months, we spent two months of that, of that, of that summer batch building out this more fleshed out MVP.
Then what happened is that we um decided to show this MVP... we, we were curious about working with enterprise design partners, people who, you know, were building models at scale like we were. We wanted them to get their hands on the product and understand you know um uh, or offer feedback rather, to help us build you know, a better product. And so one of the partners um at YC who was a visiting partner, um his name is Yuri Sagalov. He, we, he was uh somebody who gave a talk about how to run enterprise pilots and how to interact with people, folks at enterprises. And so we you know, reached out to him, had some follow-up questions that we wanted to ask, uh because we were starting that process, trying to reach out through our network uh of people at, to people at enterprises. And he gave us some really great advice. And he was really keen also on investing. And that was something that, you know, we were a little bit blindsided with to, to say to be candid. Uh you know, we didn't see that coming. We, we started to think, "Okay, we got to start thinking about you know this because this is coming around the corner." Typically a lot of YC you know companies they raise at the end of YC, which was only a month later, they raise after Demo Day.
And so we started to collect interest in that domain, trying to, tried to you know, collect just from him and from some of his other uh buddies who are, who are also uh we felt would be really helpful in trying to get us you know, access to, to more enterprises. And through that, we, we're sort of in a parallel process where we're reaching out to, to these investors, we were having these conversations, and they would introduce us to a number of different you know, ML engineers, data scientists, and so on at big enterprises. And that's actually how we met our first paying client uh at eBay. Um so so they introduced us to somebody at the eBay search team. And this is again still in around August uh August, around August of 2019. Uh they introduced us to somebody at the eBay search team and they were really you know, in love with what we were building. And we started to work with them on the MVP and, and started to build features that they thought would be interesting for their use case there. And also we met some great people at lemonade companies like Lemonade, Snapchat, Twitter, and so on. And with all of this, we, you know, had a lot of interest. We had a lot of traction. Uh, we had completed an MVP and we raised a seed round. And from then on, we focused on converting those design partners to paying clients. Um, and uh, and yeah, that's how the story sort of started uh, in 2019.
Jake Aaron Villarreal: I like how you talked about them as design partners because they're really helping you define your product and they're investing their time and energy and insights to help you make a product that's viable for the market today. What's the target audience for your product? There's a lot of AI companies that are building and popping up everywhere. Yeah, walk us through your target and and specifically now what areas are you helping solve for?
Vikas Nair: So our target, our ICP, you know, the persona, the actual person who we can help at various companies is an ML engineer, a machine learning engineer, or a data scientist or an analyst or related uh field or domain title. Um, these people's jobs, like I mentioned, a lot of their jobs are tedious. A lot of their jobs are digging through the data that their model is getting fed in and all of the predictions that the model is making and trying to understand what the errors are and why those errors are happening. So like I mentioned for example, you might have, like you know uh... and this problem, let me take a step back. Problem is relevant for big enterprises just as it's relevant for small startups. You know the, the use case, the actual domain of the problem might vary. For example, a big corporation uh they might have a model which, which... like Amazon or Netflix. They might have models for recommending content on their platform to users. And that model drives a significant amount of their revenue. For example, Amazon's search recommendation or Amazon's like product recommendation, you know, model. It decides based on a user, a bunch of user attributes, the user's, you know, basic demographics like their age, geography, and so on, as well as their previous interests, what categories of products they've purchased from before. Um, they use all of these features to try to make a prediction. You know, "is this person going to like this type of project, product? Yes or no."
And so it's hard to know, you know. First of all, you had to, like when you collect this data, you have to, you have to test that, you know, you have to send all that data through the model. You have to see, "Okay, it looks like these models are, are not getting the right prediction." Because when you're training the model, you have basically a bunch of data that's labeled where you know what the answer is. You, you, you know what the input is and you know what the answer is. You train the model on, pardon, pardon me. You train the model on, on those um input and output pairs. And then you give it some new data where you know what the answer is but you hide that from the model. And you instead uh you quiz it, basically. You see, did the model get the answer right? Yes or no. And then at the end of that, you have basically an accuracy metric. You know, like how accurate, you know, what the, the grade the model got.
And the job of the ML engineer, the, the reason it's tedious is that you have to dive through all of the incorrect answers and try to bucket them into, like, you know, "it seems like it's not good in this domain, it seems like it's not good in that domain." And then moreover, you have to understand why the model is not good in those domains. Maybe it's because it hasn't seen enough examples that look like that. Or maybe something about the model architecture is wrong um, and we need to tweak some parameters here and there to make it more attuned to those kinds of examples.
And so this really just looks like at the end of the day, what this looks like is like pouring through like CSV, you know, data, Excel sheets basically. And like kind of like, you know, as a, as a, as a... like trying to bucket, like mentally pattern matching, you know, uh, "Oh, it seems like..." And, and making hypotheses that are not necessarily grounded in much about how the model's predicting, and then you had to sort of test your hypothesis and that involves like retesting the model, retraining it. Um, and that, that process This can take like 6 months. Um, and so the whole problem here is that you know you might go six months testing, training and testing a model only to at the very end find out when you've shipped it... you know, because often a lot of these errors aren't even caught. So at the end of the process or even after you ship sometimes, is only at that point is when you realize, "Hey, this model like does not work at all for people who are you know in this demographic or, or for this kind of use case." And then you have to go all the way back. You have to go collect more data, retrain the model, so on and so forth. Your users are being affected too because they're experiencing these failures.
And so what we started this, this whole you know product started as, "Let us, let's make a really, really powerful tool that automates a lot of this process away around finding those errors and understanding them. Makes that really easy, makes that more powerful, and try to catch even more errors before they ship." And the product has evolved as well um in a lot of ways along with the advent of generative AI um and so on, but this is sort of what we started with and what our core, we're offering was, and still is to, to this day, is helping engineers identify those errors um and fix them.
Jake Aaron Villarreal: So if you're a machine learning engineer and you're building models and getting them into market, you want to make sure that they work, they're functional, they're accurate, they're not buggy, and your tool helps them expedite the speed of ensuring that they're less buggy. I like that term, "less buggy." There's kind of what they were doing and ultimately getting their product to work so their end user customer is getting the results that they want.
Vikas Nair: Yes. That's a beautiful synthesis.
Jake Aaron Villarreal: Got it. Really cool. Um, you know, we talk with currently roughly 500 AI companies. They all have their own kind of methodologies about what they're building and how they're bringing it to market. What's your view on AI and sort of the, where, where it was at, you know, in 2019 and, and, and where it's at today, and kind of what you're seeing on the inside as an engineer in the industry.
Vikas Nair: It's from... it's like fascinating to see what's changed in the past few years. It's kind of unbelievable. Um, so it's largely centered around all of the developments in generative... what's called generative AI. So I could, if you're interested and if, if your users are interested, I could give a very brief sort of crash course in how that's changed.
Jake Aaron Villarreal: So...
Vikas Nair: Basically what we had before, and I, before I, I'll say like before even this year, what we had before was a focus on what's called discriminative AI, where your model, all the model really does at the end of the day is pick an answer that's, you know, out of a set of options, almost like a multiple choice question. So for example, you know, you might have a model that takes in any kind of input. Let's say Siri. Let's go to the Siri as an example. You, you have a, or, or let's, let's start even simpler, like I, I said with the churn thing. "Are your users going to churn or not? Is this user going to churn or not? Or should I give this user a loan?" if it's like a banking model. Um or "is this user fraudulent?" if it's a you know something for like eBay's use case or like a... So you have some bit of data and the model picks an answer.
You can do actually quite a lot of powerful things with that. If you chain a bunch of these types of models together, you can create as simple of an experience [as] deciding whether or not somebody gets a loan, or as advanced of an experience as Siri or self-driving cars. If you think about Siri, all Siri is is a sequence of these types of models. There's one model which classifies whether or not, you know, for, for a given audio sample, which words, per, per frame of the audio sample, what word was spoken. And then that output of that model is text that's been transcribed. And that text goes into another classification model which says, "for this given text, what is the user asking about? Is it asking about weather, a weather request? Is it asking about a sports request? Is it asking you know about some general knowledge thing?" You know. So that model would try to guess at the intent. And then that is fed into a bunch of other upstream models that later do other kinds of things like execute, in order to execute that request.
Now what's changed is, so there was a, a paper by Google that's called "Attention Is All You Need," which is sort of a landmark paper in AI. And it described in it um this new type of architecture called a transformer, which is based off of what's called a sequence-to-sequence model, which is sort of in itself also a classification model. All this does is very simple, is given a bit of text, for example, given some initial string of data... let's say text or audio, whatever it is. Let's, let's say text though. Given like the first few words of a sentence, try to predict what you think the next word is going to be. Which you can you know sort of imagine is sort of like a classification, sort of like an, you know, a basic task. For the sentence, "Benjamin Franklin was born in the..." what do you think the next word is? And it would likely pick "year." And then you just keep chaining that. You keep doing it over and over again and you complete this sequence.
This paper basically described a finding that these models work really well uh when they are trained on, when, when you have a, a single model that is pre-trained on a very large corpus of language. So if you give a mo-, a sequence-to-sequence model uh a ton of data that's just generic data about the human language. Like for example, Wikipedia. Wikipedia is pretty generic. It's not like tuned to a specific task. Then that model is actually much better at performing the task it needs to perform, if you like later want to create a model that's very specifically trying to say certain things more or less. Like um if you want the model to speak like Siri for example, or uh output requests that are Siri-like, it's better to just have a model that has generally been trained on all of the human language um, and then fine-tune it on the very specific data that you need it to, to, to speak like.
Um and so anyways we, at, after this paper came out, more and more, this sort of caused more and more companies, and OpenAI was one of the biggest in the space, to start building uh uh these large language models, LLMs. These large pre-trained models that have been fed all of Wikipedia, that have been fed all of Reddit, all of Yahoo, all of the internet basically. And the output that we were starting to see in the mid-2010s was pretty phenomenal. I remember in college even following the you know the research on what they termed GPT models, um generative pre-trained models. And you would be able to say, like you know, you could start off uh an essay, like a college essay. You could just write that first sentence of that essay and this model would be able to complete it very realistically, creating you know, writing entire essays or articles or, or whatever it was.
And then what would h-, what had started to happen in the last one to two years was that OpenAI took this large language model that was really good, again, at just filling out the the sentences, predicting the next words, and it fine-tuned it on question-answer like tasks. So for example, if I get, you know, we're all nowadays used to what ChatGPT is, and we all interact with ChatGPT. If I tried to ask... if I tried to use a generative uh large language foundational model, just the raw GPT model in the same way I use ChatGPT, it would not go as expected. So if I asked it for example, uh "What year was Benjamin Franklin born?" It would not then just respond with an answer to that question or an attempt to answer that question. It would probably say things like, "What year was Benjamin Franklin born? What year was Thomas Jefferson born?" These are questions that historians who, you know, work in the field of US history typically ask in a, you know, introductory course to whatever. So, it would, it would go, you know, yada yada yada. It would, it would not "know," so to speak, that it's supposed to be answering a question because it's not trained on that. It's trained on Wikipedia. It would basically try to write a Wikipedia article.
So, what they did is they gave it like Yahoo Answers style questions and things like, or content rather. So it, the model was started to learn that in response to a question it gives an answer, and that was what became ChatGPT basically. So all of a sudden, now we're living in a, you know, when they shipped that, it blew up the internet. It was the biggest launch, product launch in like history. You know, basically millions and millions and millions of users day one, and it's changing the world on a day-to-day basis. They're improving the base models that, you know, ChatGPT was fine-tuned on. So ChatGPT was fine-tuned initially on GPT-3, and then they released basically GPT-4, which is just a more powerful version of GPT-3 trained on more data and more complicated architecture. Um and now, you know, ChatGPT as a result feels more intelligent.
Um they also built in other amazing functionality like GPT-V recently and like, and I'm now talking about like a month ago, uh which is that you can feed in an image, for example, and you can ask questions about that image and it can actually understand and answer questions about that image. [It] can use the web. Because you know really, at the end of the day, you, it's very easy to then just "No, instead of answering a question, let me just answer, let me just try to fill out a Google search, take, understand then all of, you know, what I'm searching for, and then use that content to give a real answer." Um it can generate code. It writes 50% of the code I write. It's not just auto-complete. It could write entire, you know, code uh uh, not just functions but entire, you know, code projects. I can even... I've seen demos where people, you know, can, you can literally say in their, you know, microphone, "Build me a website, a personal website, have the title be my first name, and uh have like these links below." And it would literally create a website using like Next.js for example to spin up a quick thing. And then you can even say like "Change the background blue, to blue instead of green" and it would do that live. So you can do amazing things nowadays.
Now the way that that's changed our work is that no longer the, the way that you evaluate these models is now, it, it first of all is sort of unsolved still, but second of all it's way more complicated of a problem. Because now there is not one definitive answer. The answer is not just easily understandable like "Oh yes, the user has churned," or "No, they have not churned." How do you grade? How do you, how do you say that, for example, an essay that you ask ChatGPT to write for you for college is, is the correct essay, is the correct answer? It's not a binary choice. It's a, it's a scale, you know, of, of zero to 100. It's a grade. So, how do you grade something like that? That is like sort of what we've been working on over the past year to adapt to people who are building in this space. Of which now everybody is building in this space because it's seeming... and we still don't know, but it's seeming that these kinds of models are actually more powerful at doing not only, you know, these new kinds of tasks that were, that are now unlocked, but also sometimes better than the old architectures were at doing things like classification.
Because you could even say like you know, you could just ask ChatGPT "this is a..." like to go back to the churn example, you could say "Here is a user I have who's from this geography, from this, you know, uh is this years, this many years old, um likes these kinds of products, or whatever it is, this has this credit score. What do you think should, should I... are they going to churn or not?" And they might actually give you a better answer than the old models did. So how do you evaluate these? It's a little bit unknown. We, there's a been a lot of research, there's, there's, there's a group from Stanford that created something called HELM a few months ago where they described a number of different metrics that you can use, um and then uh to evaluate these models. But then more interestingly over the past year a lot of the research is actually showing that you can use a GPT model or an LLM, large language model, to evaluate another large language model. Which sounds a little scary, but it's actually showing that the accuracy of the, of the responses when looked at by a human is like something in the order of 80% or 90, 90% of the time, right? Or sometimes more.
So you can for example... what this looks, what, what this means more concretely is you can take the response of an, of the LLM, large language model, that you have in your product, um and you can ask another AI, "Hey, is this answer that this AI gave relevant? Is it, is it spoken concisely? Uh is it, is it you know harmful? Is the tone correct? Uh is it leaking PII even? Like is this model, you know, uh outputting sensitive user information?" And the other model might tell you yes or no. Um, and so you can then aggregate that over large amounts of data and you can create a score uh that, that uses basically GPT to evaluate another GPT. So it's a little bit meta. It's a little bit Blade Runner, but it's cool.
Jake Aaron Villarreal: That's insane. That's I mean in some ways exciting and also scary.
Vikas Nair: Yeah.
Jake Aaron Villarreal: Um you talked about... I want to go back to the development. So engineers, you know, oftentimes we're talking to they're, you know, nervous about their future. You know, is AI going to replace what I do or maybe a portion of what I do or maybe some element that, you know, I'm going to be less valuable to a company. Um, and it's not just engineering. It sounds like there's a lot of value to having, you know, a co-pilot or a technology that's helping, you know, do maybe some mundane tasks in engineering or maybe more complex stuff in engineering. But also, you know, testing has been automated for a long time where technology can look at and, and catch things that are issues and fix them. I mean, you know, I was at Oracle for 7 years and I was just astounded to hear that, you know, in their support team now there's, your database breaks or goes down and, you know, technology goes in, fixes it before you even know it's down and, you know, spins up the database again. And that's technology and automation.
So, how do you see AI and your specific product with testing um, I don't know, integrated or supporting it, or kind of walk through that a little bit so that it's really clear about the tool, the platform you're creating for engineers and making sure that their, their, their models are efficient and less buggy. I, I guess kind of just clarify that a little bit more, because I'm curious to know, does it also take away testers? Like does it take, will AI take testing out of the equation as we go forward and evolve with technology?
Vikas Nair: That's fascinating. I think with respect to testing, yeah. I think with a lot of these jobs, not just with testing, there is definitely an existential fear that is very real um that this could replace a lot of the work. What I'll say about that part before I dive deeper into the whole testing thing, because there's some interesting stuff there as well, is that, you know, I think with any big technology that comes out, there's going to be a short run, short-term period where a lot of jobs are lost. And that's tragic. Um, and I don't know that like I'm not, I don't have any sort of ethical argument to make about whether or not we should do this or should not do this. But it does recall, you know, for example, when the internet came out or when computers came out, right? Um, one of the big uh immediate uh groups of people who are going to be affected by this are writers. Uh like the SAG, the SAG-AFTRA and the, and the writing strike were, were a big component of that was AI. You know, is AI is going to be used?
You cannot stop it. With any technology, like with the atom bomb. Um for the better, for the worse, you can't stop the development of, of, of technological progress in a, in a global world where there are multiple actors who are, you know, have their own interests and and so on and so forth. And so I think in the short run, in the short term, it does affect a lot of jobs. Um in the long term, it does probably automate away a lot of things. And hopefully we are able to find um uh a balance in abstracting the roles of everybody. So all of the people who are going to be replaced, hopefully we find some sort of balance in which they are the managers of these AI or whatever, or have some other sort of job that is um one layer up in the abstraction tree.
Um now what I'll say about testing in particular is, is there was this interesting uh talk that was given by Andrej Karpathy who was the, previously the acting, you know, Director of AI at Tesla, the Head of AI at Tesla. Um he described a, a future um MLOps stack or a, a future sort of development stack for the AI engineers at Tesla where a lot of the processes of building these AI and testing them and evaluating them and shipping them as well could be automated away, both by AI but also by just other you know software uh uh routines.
So he imagined a way in which for example you collect data autonomously. So for example, to, to go specifically deeper into Tesla. Tesla has these sensors on their cars right? These sensors are always picking up data about the driver, about the road rather, around them, so on and so forth. That data gets sent back automatically to Tesla HQ and Tesla then uses that data to retrain their models. And then automatically the errors in that retraining process and in the testing process um get surfaced and you can have some sort of automatic error detection using some tool kind of similar to Openlayer. You can automatically understand, you know, uh what are the types of errors that are going on for any bit of, any bit of data um that's sent through. And then for those errors you could theoretically automatically understand why those errors are happening also through, you know, the tools similar to ours that we're building and, and probably tools that they have there.
And from that insight about what types of errors are happening... so for example, "Hey, it seems like when it's really, really rainy and foggy." This is a really basic example, but let's say "We seem like, it seems like a lot of these errors happen when there's fog." So your automatic process could pick that up and say, "Hey, we need to, I need to collect more data that includes fog." Or you could have an AI or some other thing uh some, some other, some, some other bit of software like artificially inject fog into existing data that you already have. So you could take a normal sunny day and you could apply some sort of filter uh you know to make that day foggy or to make it rainy or whatever it is, and then you can retrain the model on that artificial data. All of this still is automatic. And then the model theoretically would be better at, at, at you know handling foggy inputs.
And then you repeat, rinse and repeat, rinse and repeat, rinse and repeat. Keep collecting more data, finding things that the model is bad at, and creating artificial synthetic examples um of those, of those bad data points, retraining the model so they're better at, at it. And until... and then finally you might have like a 100% accurate or close to that model. And then you ship that and then you see more errors that happen in the, in the wild. And this is an automatically always repeating process. He called it "Operation Vacation." He called it that because he said, "My team should be able to one day take a vacation, all of us, 5 days, 7 days, and come back when they're done and see an improved model that's been improved automatically." And I think that that's a very realistic future.
Jake Aaron Villarreal: I like "Operation Vacation." It sounds like something everyone should have, you know, just the title of it. Really cool. Um, wow. You can go very deep obviously on the tech side. Let's bring it up a little bit. Uh, in terms of, you know, the company today, you head into 2024. What are you excited about Openlayer for 2024?
Vikas Nair: What we're really excited about is we've just shipped... we spent all of this year, you know, building a new version of the product, like I mentioned, that works for generative AI, that's able to, you know, evaluate generative AI. And we've also in parallel built a version of the product that can be deployed, that can, that can evaluate your production data, which is super, super important to anybody who's shipping AI. Especially nowadays when you have this sort of unhinged agent that can speak like a human. You got to really rein it in and make sure that it's not making, you know, errors like we saw many New York Times articles talking about hallucinations, talking about aggression, talking about leaked PII that these models were, were, were exhibiting.
And so when combined with, you know, so, so our, our new product that we've shipped, that we've actually just shipped in the past month when we're recording this, is able to take any AI that's deployed to production and run an extensive suite of tests on that AI as it's being used with your real users in production. So you will basically, you know the... anybody who uses our product, when, you know, is able to get Slack and email notifications that tell them "Hey, you know, all of a sudden in the past hour your model was like hallucinating terribly" or "your model you know was responding in a way that was not relevant at all to the you know questions that it was being asked" um and so on and so forth.
So we're really, really excited because there's so much interest that we got. We launched um this product on HackerNews just recently and on Product Hunt. We got a lot of really great interest in it and we have even more clients uh streaming in. We, we had some significant growth over the past year in the number of, of not only the number of clients that we're now serving, but also the variety. Uh so you know, I said earlier that we started with eBay, now we serve both enterprises as well as small startups who are equally interested in, in using Openlayer to test their models. Um, and they also do all sorts of different things. We, we, we're working with people who are in the banking sector. Um, we're working with people like eBay who are in, you know, e-commerce. Um, we're working with people who are in core tech, who are using, you know, who are building like complex AI agents that are kind of like Siri, for example. Um, and so all of these various use cases.
So exciting to like get to see what people are working on. How amazing this is, truly like the dot-com era, hopefully with no bubble. Um where you're seeing all sorts of amazing advancements and amazing new technologies and to be able to be a part of that is really exciting. So I'm, I'm very you know optimistic about the next year like seeing even more you know interesting use cases, serving more, more and more clients of, of all nature.
Jake Aaron Villarreal: That's great. Well, it sounds like a lot of growth. Um, Vikas, if somebody wanted to find you or find your company, and more importantly, let's talk about the... or do you have open roles that people might want to find you and apply to? I mean, what's, what's growth look like from that perspective?
Vikas Nair: Yeah. So, another uh, in, in the, a new... in the new year, we're also looking to grow our company. We're looking to bring on more people who are actually focused not just on engineering and design, but also people who are focused on sales and growth. That's going to be like really critical for us moving forward. Um we need people who are able to take a complex product like this, um bring it to the table of, of you know, all sorts of different companies um with different personas, with different you know sizes and, and help us grow um in that way. So we're, we're definitely like heavily focused on getting and on recruiting in the sales front and the growth and marketing fronts um in the new year.
Uh people can find us uh on our website at Openlayer, which is openlayer... uh pretty simple... .com. Uh we have a, a full demo there. We have a lot of information about our product, uh about the clients that we serve. We have an amazing new docs website for the engineers out there who just want to get you know, started playing with it. It's fully self-serve as well. So actually anybody who's building AI, even if it's something simple, you can actually just try it out you know, uh today, if you're a student even, or if you're you know working at a company um you know, and you, you're building models, you can just sign up for free, you can add Openlayer to your, to your stack with one line of code or purely through the UI, um and you can start running, like you know, a bunch of different tests that we offer on your models to really battle test them, to see how accurate are your models really. Because it might not be as accurate as, as you think.
Jake Aaron Villarreal: Wow, that's incredible. Well, you heard it straight from Vikas. Uh, Vikas, I want to thank you for spending your time with us today and really walking us through the, the where, where AI is and where it's going. And also what companies might want to be thinking about if they want to make sure that their AIs are actually delivering on the requests that they're getting from their customers. And for machine learning engineers and engineers in general, a new tool, a new platform, you know, from, you know, an Ivy League foundation team, if you really think about it, that's been at Apple and has been delivering um real-world products to the market. So excited to uh check in with you in 6 and 12 months, see how things are going. But uh yeah, thanks for, thanks for joining us today and thanks for all, all of our listeners for listening. It means the world to me that you spent your time with us here. I'm Jake Aaron Villarreal, the host of the podcast, and can't wait to catch up with you all in the next episode. Enjoy the holidays and until then, take care.
Before we wrap up, I want to give a big shout out to all the entrepreneurs that have joined to make this podcast possible. And for all the listeners for listening, it means the world to me that you chose to spend your time with us today. I'm your host Jake Aaron Villarreal signing off for now. We can't wait to connect with you all soon on the next episode. Take care.
This show is sponsored by Match Relevant, a company that helps venture-backed startups find the best people in the market. And they do it in three simple steps. First, they sit down with founders to understand their story. Second, they tell their story into multiple candidate channels. And third, they schedule interviews within 48 hours. Find us at matchrelevant.com to learn more about how we do it.