Mentionsy Mentionsy
evoilutioncast
evoilutioncast

evoilutioncast 45: How to build a solid foundation for AI on VCF?

06.08.2025 ·52 min 59 s

In this episode of the IT podcast - evoilutioncast, Maciej Lelusz speaks with Frank Denneman - a very AI person in VMware by  Broadcom. Frank plays a key role in the VCF division, where he shapes the roadmap for the Private AI Foundation in NVIDIA and heavily influences the division’s overall AI strategy.✔️Fancy to know whether the VCF is an infrastructure for AI? ✔️Does that put an AI construct in the DC?✔️The truth is that with AI, nothing is easy, but you can make it easier. ✔️As well as there are things to be approved by humans and things to be made by AI.✔️Listen to the conversation to find out new trends in AI infrastructure as RAG or #vector database and many more. 🤝 Episode's Partner: VMware by Broadcom𝓛𝓲𝓼𝓽 𝓸𝓯 𝓬𝓸𝓷𝓽𝓮𝓷𝓽:00:02:00 AI on VCF (VMware Cloud Foundation) platform: what's all about?00:06:50 vcf9 as a local hypervisor?00:09:10 SaaS solution as a starter on the cloudfoundation platform00:12:27 Platform, both for engineers and developers - what is this VCF platform? 00:20:58 cost spending tracking on private cloud platform00:24:41 Retrieval Augmented Generation (RAG) is the most common use case00:32:15 The most trending solutions on the market: summarizing based on augmented AI in healthcare00:36:30 digestion pipeline of data by building vector database 00:40:03 Similarity search and embedding model: how does it work?00:42:43 What gives the VCF platform to the organization as an open infrastructure model?00:49:24 Next steps on VCF? Wider integration and an easier way of consuming a platform - giving the best way of consuming the resources that you have🔔 Subskrybuj: https://bit.ly/sub_evoilutioncast

Więc jedyne, co muszą zrobić, to specyfikować model. W zasadzie, gdzie jest on w galerii modelu, jaka jest moja infrastruktura, więc w zasadzie dostajemy z Veeam klasę, jak to jest GPU, to jest ilość replik, to jest CPU i to jest memoria. I to wszystko. Basically, you say, please deploy this model. We deploy a Kubernetes worker node with all the libraries. Yes, on that cluster, on that GPU, on that host in the most efficient way. So we basically automate everything as efficiently as possible. And the only thing that comes back to the developer in the CLI or in the UI is the model API endpoint, which they can use and then basically say, OK, yeah, so they can say, I'm going to use that in my application.

But another thing what you can do, if you think about it, Niemieckie rozwiązanie, ponieważ jest zbyt mało GPU-ów, głównie w przedsiębiorstwach, jest to, że mówisz, że zbieram jakieś DLVM z Jupyter Notebooka i zamiast ładowania modelu na ten DLVM, będę używał modelu runtime endpoint. Więc jeśli masz kolaborację, na przykład pięć naukowców, które próbują zrozumieć, co ten model robi, You spin it up in the model runtime and every Jupyter notebook environment just runs on a CPU but hits that model at that molecule who consumes a GPU. That's a far more efficient way of dealing with it while you still have everything local, right?

Pokazano wszystkie 2 dopasowania. Transkrypcja generowana automatycznie i niesprawdzana ręcznie — może zawierać błędy.