evoilutioncast 45: How to build a solid foundation for AI on VCF?
In this episode of the IT podcast - evoilutioncast, Maciej Lelusz speaks with Frank Denneman - a very AI person in VMware by Broadcom. Frank plays a key role in the VCF division, where he shapes the roadmap for the Private AI Foundation in NVIDIA and heavily influences the division’s overall AI strategy.✔️Fancy to know whether the VCF is an infrastructure for AI? ✔️Does that put an AI construct in the DC?✔️The truth is that with AI, nothing is easy, but you can make it easier. ✔️As well as there are things to be approved by humans and things to be made by AI.✔️Listen to the conversation to find out new trends in AI infrastructure as RAG or #vector database and many more. 🤝 Episode's Partner: VMware by Broadcom𝓛𝓲𝓼𝓽 𝓸𝓯 𝓬𝓸𝓷𝓽𝓮𝓷𝓽:00:02:00 AI on VCF (VMware Cloud Foundation) platform: what's all about?00:06:50 vcf9 as a local hypervisor?00:09:10 SaaS solution as a starter on the cloudfoundation platform00:12:27 Platform, both for engineers and developers - what is this VCF platform? 00:20:58 cost spending tracking on private cloud platform00:24:41 Retrieval Augmented Generation (RAG) is the most common use case00:32:15 The most trending solutions on the market: summarizing based on augmented AI in healthcare00:36:30 digestion pipeline of data by building vector database 00:40:03 Similarity search and embedding model: how does it work?00:42:43 What gives the VCF platform to the organization as an open infrastructure model?00:49:24 Next steps on VCF? Wider integration and an easier way of consuming a platform - giving the best way of consuming the resources that you have🔔 Subskrybuj: https://bit.ly/sub_evoilutioncast
That's a Jupyter notebook or a VS Code or whatever, right? But what we do with... At the end of the build your own spectrum, the DLVM and the AI Kubernetes clusters is we give you enough resources that's aligned with the technology stack. So to go a little bit more into detail, if you build a virtual machine or a Kubernetes environment with a container runtime, Musisz być świadomy o kierowcy, kierowcy GPU, ponieważ jest wiele bibliotek, które są polegające na pewnej wersji tego kierowcy. I dodatkowe biblioteki i dodatkowe elementy, takie jak Python lub inna biblioteka, muszą być polegane na pewnej wersji CUDA, wersji setu NVIDIA.
Wspomnieliśmy o DLVM, w którym zauważyliśmy wersję gpu w hypervisorze, połączyliśmy ją z gpu w VM, a następnie dajemy wersję gpu w virtual machine. It's a very brittle stack, but we already sold that. It's like a safe place to start your journey and then go for a more sophisticated use if you need them. When you look in the market, the standards Wcześniej nie było standardów. Teraz widzimy, że niektóre biblioteki, niektóre środowiska się rozprzestrzeniają.
When you want to run such a model, you specify the local repository. The spin-up time for a model, getting it from storage into GPU memory, is much shorter. That can help with scaling and we can go into that as well. Jeśli zaczniemy od tego, jeśli wtedy powiedzmy, że mamy model, sprawiamy, że jest bezpieczny, bezpieczny, jest oświetlony. Teraz następną rzeczą, którą chcemy zrobić, jest to, że chcemy go uruchomić. Możesz zrobić dwie rzeczy. Możesz wywołać ten kluster AI-Kage i w zasadzie instalować wszystkie frameworky służące i wszystko to, co mamy w Fundacji Privat AI. Mamy model runtime. Więc przy użyciu CLI lub UI, naukowca albo deweloper mówi, OK, daj mi moduł do końca modelu.
Więc jedyne, co muszą zrobić, to specyfikować model. W zasadzie, gdzie jest on w galerii modelu, jaka jest moja infrastruktura, więc w zasadzie dostajemy z Veeam klasę, jak to jest GPU, to jest ilość replik, to jest CPU i to jest memoria. I to wszystko. Basically, you say, please deploy this model. We deploy a Kubernetes worker node with all the libraries. Yes, on that cluster, on that GPU, on that host in the most efficient way. So we basically automate everything as efficiently as possible. And the only thing that comes back to the developer in the CLI or in the UI is the model API endpoint, which they can use and then basically say, OK, yeah, so they can say, I'm going to use that in my application.
But another thing what you can do, if you think about it, Niemieckie rozwiązanie, ponieważ jest zbyt mało GPU-ów, głównie w przedsiębiorstwach, jest to, że mówisz, że zbieram jakieś DLVM z Jupyter Notebooka i zamiast ładowania modelu na ten DLVM, będę używał modelu runtime endpoint. Więc jeśli masz kolaborację, na przykład pięć naukowców, które próbują zrozumieć, co ten model robi, You spin it up in the model runtime and every Jupyter notebook environment just runs on a CPU but hits that model at that molecule who consumes a GPU. That's a far more efficient way of dealing with it while you still have everything local, right?
Basically you use one and that's all, right? Exactly. So you have to figure a different way of exposing the GPU or the model, the consuming factor, and then see what other services can I allow to consume that. So the key thing what most... Wszystko, co klienci teraz rozumieją, jest to, że nie ma żadnego 1-2-1 związku. To znaczy, że kiedy wywołasz aplikację, zaczniesz ją włączyć do modelu API Endpoint, do modelu działającego na A lub mnóstwo GPU. Ale to nie jest jedyna aplikacja, która dotarła do końca tego modelu.
Like, oh, you have a 16,000 GPU cluster or a 24,000 GPU cluster to build large models. No, no, no, use VCF, no. Nie używamy bare metalu na to. To nie jest nasz zakres używania, ale tym, na czym skupiamy się, jest strata zainteresowania. To jest mniejszy model, czy większy model, ale dodajesz tylko kilka linii do niego, albo dodajesz pewne dane do niego. Możesz to połączyć z jakimkolwiek przykładem. Nie otrzymujesz urodzenia do osoby. Zazwyczaj bierzesz tą osobę Names mentioned, Names mentioned,
Pokazano wszystkie 7 dopasowań. Transkrypcja generowana automatycznie i niesprawdzana ręcznie — może zawierać błędy.
Kliknij, aby znaleźć fragmenty, w których pada.
In this episode of the IT podcast - evoilutioncast, Maciej Lelusz speaks with Frank Denneman - a very AI person in VMware by Broadcom. Frank plays a key role in the VCF division, where he shapes the roadmap for the Private AI Foundation in NVIDIA and heavily influences the division’s overall AI strategy.
✔️Fancy to know whether the VCF is an infrastructure for AI?
✔️Does that put an AI construct in the DC?
✔️The truth is that with AI, nothing is easy, but you can make it easier.
✔️As well as there are things to be approved by humans and things to be made by AI.
✔️Listen to the conversation to find out new trends in AI infrastructure as RAG or #vector database and many more.
🤝 Episode's Partner: VMware by Broadcom
𝓛𝓲𝓼𝓽 𝓸𝓯 𝓬𝓸𝓷𝓽𝓮𝓷𝓽:
AI on VCF (VMware Cloud Foundation) platform: what's all about?
vcf9 as a local hypervisor?
SaaS solution as a starter on the cloudfoundation platform
Platform, both for engineers and developers - what is this VCF platform?
cost spending tracking on private cloud platform
Retrieval Augmented Generation (RAG) is the most common use case
The most trending solutions on the market: summarizing based on augmented AI in healthcare
digestion pipeline of data by building vector database
Similarity search and embedding model: how does it work?
What gives the VCF platform to the organization as an open infrastructure model?
Next steps on VCF? Wider integration and an easier way of consuming a platform - giving the best way of consuming the resources that you have
🔔 Subskrybuj: https://bit.ly/sub_evoilutioncast