evoilutioncast 45: How to build a solid foundation for AI on VCF?
In this episode of the IT podcast - evoilutioncast, Maciej Lelusz speaks with Frank Denneman - a very AI person in VMware by Broadcom. Frank plays a key role in the VCF division, where he shapes the roadmap for the Private AI Foundation in NVIDIA and heavily influences the division’s overall AI strategy.✔️Fancy to know whether the VCF is an infrastructure for AI? ✔️Does that put an AI construct in the DC?✔️The truth is that with AI, nothing is easy, but you can make it easier. ✔️As well as there are things to be approved by humans and things to be made by AI.✔️Listen to the conversation to find out new trends in AI infrastructure as RAG or #vector database and many more. 🤝 Episode's Partner: VMware by Broadcom𝓛𝓲𝓼𝓽 𝓸𝓯 𝓬𝓸𝓷𝓽𝓮𝓷𝓽:00:02:00 AI on VCF (VMware Cloud Foundation) platform: what's all about?00:06:50 vcf9 as a local hypervisor?00:09:10 SaaS solution as a starter on the cloudfoundation platform00:12:27 Platform, both for engineers and developers - what is this VCF platform? 00:20:58 cost spending tracking on private cloud platform00:24:41 Retrieval Augmented Generation (RAG) is the most common use case00:32:15 The most trending solutions on the market: summarizing based on augmented AI in healthcare00:36:30 digestion pipeline of data by building vector database 00:40:03 Similarity search and embedding model: how does it work?00:42:43 What gives the VCF platform to the organization as an open infrastructure model?00:49:24 Next steps on VCF? Wider integration and an easier way of consuming a platform - giving the best way of consuming the resources that you have🔔 Subskrybuj: https://bit.ly/sub_evoilutioncast
Przede wszystkim, Frank, miło Cię tutaj poznać. Kiedyś w Polsce? Dwa, trzy lata temu? Myślę, że tak. Było to nawet przed COVID-19m. Było to przed COVID-19m. Tak, bo po tym, że Demkan przyjeżdżał do VMAX... Wszystko jest blurowane. Tak, zdecydowanie. Dzień dobry wszystkim. To jest evolutioncast. Nazywam się Maciej Lelusz. Dzisiaj z nami jest Frank Denneman. Siedzę dzisiaj w jego brzuchu. Jak długo siedziałeś w VMware i w Broadcom? W totalności ponad 12 lat, ale miałem krótki stint z start-upem w ekosystemie VMware. Było to 4-3 lata na start-upie.
Back to VMware and now I'm in my eighth year, actually hitting nine in October or something like that. Cool. Twelve years. So you see that, you see the all of ACES, the products, the going out products, the Broadcom acquisition, everything there. I think I saw every CEO except Diane. Ale z powodu tego faktu, że pracowałem w społeczeństwie od 2005 roku, w zasadzie instalując ESXi i idąc od tamtej pory, więc przyjechałem do... mój pierwszy WM World był w 2006 roku, tak, i przyjechałem do TSX w Europie, więc widziałem wszystkich CEO'ów do tej pory, tak. To dobrze, to dobrze, dobry background.
Wiesz co, o czym chcę rozmawiać z tobą jest oczywiście AI, bo nikt nie mówi o czymkolwiek innym, ale... Zacznijmy od tego, wiesz, mniej lub mniej szaleństwa czasami, wiesz, historii o tym, jak to zmieni nasz świat, wiesz, bla, bla, bla. Nie. The thing is, you know, because we know each other so many years, you know, we are more, let's say, we are not people who talk about theory, right? We are more practical guys. And I think that that's the way how we should talk about technology right now, especially AI, because if you see this magical stuff that it can do, basically, you stop to thinking about real use cases, because it's so fancy with the super easy access, you know, all these pictures with Gibi, you know.
That's the concept. However, when we started to develop this platform, so it's built on top of the VCF platform, we started to think, okay, if we look at the AI ecosystem, the broad landscape, it's very wide, it's also very deep, and it's changing Names mentioned, Names mentioned, Names mentioned,
So we build, before we basically go into the use cases, we build a platform that looks at what do you need. Names mentioned in English. Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA. Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA.
Names mentioned, Names mentioned, Names mentioned, You act like a little bit like a local hypervisor, right? Because we have these two sides of the hypervisor responsibility matrix, right? The hypervisor, let's say, it's responsible to the infra and from the infra, till the, let's say, application layer and so on, their customer. Wydaje mi się, że w porządku VCF-9, a generalnie w conceptu VCF, bardzo dobrze pasuje to, że budujesz swoją prywatną eksperię na górze. Jako sprzedawca infrastruktury dajesz pracownikom, którzy konsumują tą infrastrukturę, rigidną platformę.
Więc zamiast mówić, o, tutaj jest platforma, Now you basically have to find out everything yourself. We said no. We're going to help you along the way. We're not going to build everything because we have a very strong ecosystem with a lot of partners, but we will give you enough functionality to make your first steps. In essence, what we do with Private AI Foundation is we focus on Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom
That's a Jupyter notebook or a VS Code or whatever, right? But what we do with... At the end of the build your own spectrum, the DLVM and the AI Kubernetes clusters is we give you enough resources that's aligned with the technology stack. So to go a little bit more into detail, if you build a virtual machine or a Kubernetes environment with a container runtime, Musisz być świadomy o kierowcy, kierowcy GPU, ponieważ jest wiele bibliotek, które są polegające na pewnej wersji tego kierowcy. I dodatkowe biblioteki i dodatkowe elementy, takie jak Python lub inna biblioteka, muszą być polegane na pewnej wersji CUDA, wersji setu NVIDIA.
Wspomnieliśmy o DLVM, w którym zauważyliśmy wersję gpu w hypervisorze, połączyliśmy ją z gpu w VM, a następnie dajemy wersję gpu w virtual machine. It's a very brittle stack, but we already sold that. It's like a safe place to start your journey and then go for a more sophisticated use if you need them. When you look in the market, the standards Wcześniej nie było standardów. Teraz widzimy, że niektóre biblioteki, niektóre środowiska się rozprzestrzeniają.
Ludzie bardziej często wybierają, powiedzmy, TensorFlow czy coś takiego. Widzimy, że to coraz bardziej popularne. Wybrałeś to, żeby zbudować platformę dla inżynierów, dla deweloperów, prawda? Tak. Co ta platforma robi? Myślę, że to jest dobre. Przechodzimy do tego zakresu użycia, ponieważ powiedziałeś, że jest to bezpieczne miejsce. Teraz weźmy zakres użycia tego DLVM, czy AI Cage Cluster. Zazwyczaj, co się dzieje, jeśli spojrzymy na Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA.
Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA. You can download LAMA, the one from Meta, the foundation model from Meta. That means that you have a large language model that's already pre-trained with a lot of knowledge, with a lot of data. Or you can use Mistrals or Mixtrals, one of the European models. Or you can go to China and you can take a look at DeepSeq or whatever. The interesting thing is...
Zazwyczaj lecimy na pracę kogoś innego w społeczeństwie nauki danych, a potem w zasadzie rozwiązywamy, jak mogę to naprawić dla mojego własnego usługodawstwa. Jeśli zainstalujesz model, teraz ostatnie, co chcesz zrobić, jest po prostu od razu włożyć go w produkcję. Musisz się oznaczać, że ten model w rzeczywistości zapewnia to, co mu oznacza. You need to figure out, okay, what are the heuristics, what is the behavior of this model and are there any security holes, whatever. So with that DLVM, what we do is we provide you with the ability just to spin that up using automation, because we already have these templates available.
So there's a self-service portal for the data scientist, spin it up, they can download that model into that DLVM and it becomes like a clean room, right? Names mentioned, Names mentioned Names mentioned. Names mentioned.
Tak, więc stworzyłeś to w tym projekcie, w zasadzie pozwalałeś tylko jednemu zestawowi naukowców czy rozwojowników do dostępu do tego zapoznania, a potem ubezpieczyłeś, że, hej, możemy to utrzymać, cały czas wciąż zostając lokalizowanym. Teraz jedna z kluczowych rzeczy, jest mnóstwo dobrych rzeczy o tym, więc, po pierwsze, modeli są bardzo wielkimi. The real easy way to think about it, typically a model, you basically specify it with a parameter count. So when we say Salama 3, then we say 8 billion or 33 billion or 405 billion. That's the parameter count. Every parameter typically takes up two bytes. W zasadzie, bilion byte to giga byte.
When you want to run such a model, you specify the local repository. The spin-up time for a model, getting it from storage into GPU memory, is much shorter. That can help with scaling and we can go into that as well. Jeśli zaczniemy od tego, jeśli wtedy powiedzmy, że mamy model, sprawiamy, że jest bezpieczny, bezpieczny, jest oświetlony. Teraz następną rzeczą, którą chcemy zrobić, jest to, że chcemy go uruchomić. Możesz zrobić dwie rzeczy. Możesz wywołać ten kluster AI-Kage i w zasadzie instalować wszystkie frameworky służące i wszystko to, co mamy w Fundacji Privat AI. Mamy model runtime. Więc przy użyciu CLI lub UI, naukowca albo deweloper mówi, OK, daj mi moduł do końca modelu.
Więc jedyne, co muszą zrobić, to specyfikować model. W zasadzie, gdzie jest on w galerii modelu, jaka jest moja infrastruktura, więc w zasadzie dostajemy z Veeam klasę, jak to jest GPU, to jest ilość replik, to jest CPU i to jest memoria. I to wszystko. Basically, you say, please deploy this model. We deploy a Kubernetes worker node with all the libraries. Yes, on that cluster, on that GPU, on that host in the most efficient way. So we basically automate everything as efficiently as possible. And the only thing that comes back to the developer in the CLI or in the UI is the model API endpoint, which they can use and then basically say, OK, yeah, so they can say, I'm going to use that in my application.
But another thing what you can do, if you think about it, Niemieckie rozwiązanie, ponieważ jest zbyt mało GPU-ów, głównie w przedsiębiorstwach, jest to, że mówisz, że zbieram jakieś DLVM z Jupyter Notebooka i zamiast ładowania modelu na ten DLVM, będę używał modelu runtime endpoint. Więc jeśli masz kolaborację, na przykład pięć naukowców, które próbują zrozumieć, co ten model robi, You spin it up in the model runtime and every Jupyter notebook environment just runs on a CPU but hits that model at that molecule who consumes a GPU. That's a far more efficient way of dealing with it while you still have everything local, right?
You don't do that for production. Tell me one thing here, because I see this pipeline, you know, that is going in, you know, there, there, another, this person can use it, this, how we can collaborate on that stuff, you know, that's, that's pretty amazing. But, you know, usually those infrastructures are extremely expensive, because of the equipment, GPUs, generally speaking, it's not the, it's not the cheapest hobby, let's say. And... Nope. Do we have any option in VCF plus private AI with NVIDIA? Have some kind of tracking of spending, you know, billing, something that it's allow organizations to basically tell to their users, hey, maybe, you know, five models in the same time when you don't use them, cost you too much, you know, just...
Names mentioned. Names mentioned. Now, we do have an API gateway in the model runtime, allowing us to understand what are the requests coming in, which model is it hitting. And so one of the things that we're thinking about, I'm not saying it's going to be the next version, so this is not a promise.
I need to make sure that that's the case, but we are thinking about, okay, can we somehow... Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA. Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA. Names mentioned, Maciej Lelusz, Frank Denneman, VMware, Broadcom, NVIDIA
Basically you use one and that's all, right? Exactly. So you have to figure a different way of exposing the GPU or the model, the consuming factor, and then see what other services can I allow to consume that. So the key thing what most... Wszystko, co klienci teraz rozumieją, jest to, że nie ma żadnego 1-2-1 związku. To znaczy, że kiedy wywołasz aplikację, zaczniesz ją włączyć do modelu API Endpoint, do modelu działającego na A lub mnóstwo GPU. Ale to nie jest jedyna aplikacja, która dotarła do końca tego modelu.
To, co oni robią, to oprogramowują wszystkie dokumenty i wyrażają je w pewien sposób, i powiem o tym później, do wielkiego modelu językowego. I więc klienci, przepraszam, nie klienci, ale w szczególności legalne osoby, They can use that application to directly talk to the documents and ask specific questions. Like, okay, what's going on with this? Can you find any relationship with that or whatever? And that's the nature of having a large language model, the intelligence of a large language model, but also having directly access to your data, to your personalized data that you don't want to...
Share anywhere outside of your organizational walls or maybe even on the internet, right? So another way what you can think of is coding assistant. A great example is our own environment. We downloaded a coding assistant from Hugging Face and we fine-tuned it. So fine-tuning is basically a shorter training period with fewer amount of data Names mentioned, Maciej Lelusz
What we do is when we go into retrieval augmented generation, we have that foundation model. So typically a LAMA or whatever. And then next to that is you build a what's called a vector database. And that vector database can be deployed by our DSM. It's a data service manager. It's a self-service portal to deploy a vector database. And we connect that vector database to the large language model. Now every... Names mentioned in English Names mentioned in English
We have a lot of telco providers and a lot of financial institutions that offer new types of services on a regular basis. Now, when you call up, like a call center, you will get, after basically going through all these automated steps, you will get a real person and you talk to a real person, because that's what most people like. If you ask a question nowadays, They have to go and find their information in typically one of the six, nine screens that they have in front of them, right? So they have to figure out, okay, which screen do I need to access? Which application do I need to do? And while they're talking to a customer. Now, what some of these organizations have done is they said, let's grab all of that data.
Pokazano pierwszych 25 dopasowań — doprecyzuj frazę, aby zawęzić wyniki. Transkrypcja generowana automatycznie i niesprawdzana ręcznie — może zawierać błędy.
Kliknij, aby znaleźć fragmenty, w których pada.
In this episode of the IT podcast - evoilutioncast, Maciej Lelusz speaks with Frank Denneman - a very AI person in VMware by Broadcom. Frank plays a key role in the VCF division, where he shapes the roadmap for the Private AI Foundation in NVIDIA and heavily influences the division’s overall AI strategy.
✔️Fancy to know whether the VCF is an infrastructure for AI?
✔️Does that put an AI construct in the DC?
✔️The truth is that with AI, nothing is easy, but you can make it easier.
✔️As well as there are things to be approved by humans and things to be made by AI.
✔️Listen to the conversation to find out new trends in AI infrastructure as RAG or #vector database and many more.
🤝 Episode's Partner: VMware by Broadcom
𝓛𝓲𝓼𝓽 𝓸𝓯 𝓬𝓸𝓷𝓽𝓮𝓷𝓽:
AI on VCF (VMware Cloud Foundation) platform: what's all about?
vcf9 as a local hypervisor?
SaaS solution as a starter on the cloudfoundation platform
Platform, both for engineers and developers - what is this VCF platform?
cost spending tracking on private cloud platform
Retrieval Augmented Generation (RAG) is the most common use case
The most trending solutions on the market: summarizing based on augmented AI in healthcare
digestion pipeline of data by building vector database
Similarity search and embedding model: how does it work?
What gives the VCF platform to the organization as an open infrastructure model?
Next steps on VCF? Wider integration and an easier way of consuming a platform - giving the best way of consuming the resources that you have
🔔 Subskrybuj: https://bit.ly/sub_evoilutioncast