In an era where technology evolves at an unprecedented pace, MemCon 2024 conference stands as a testament to innovation. They invited me to talk about memory bottlenecks in GenAI compute—a challenge that resonates with every stakeholder in the memory ecosystem.
It was great to speak at MemCon 2024 with Angela Yeung from Cerebras. The event took place at the main auditorium of the Computer History Museum in MountainView. A lot of history there, with a wall of legends.
Me and Angela talked about the memory bottlenecks for GenAI compute. This a highly relevant topic for companies in the memory ecosystem.
Memory bottlenecks for GenAI compute
We all experience the latency issues in GenAI. Often, the cause is memory not getting to the compute chips (GPUs and others) fast enough. The chips have capacity for compute, they simply often are underused as memory doesn’t get to the processor on time (latency in the pipeline) and then we end up with idle time for compute.
There’s different emerging architectures so it is good engaging with the community. I’m glad I participated with Cerebras Systems on MemCon 2024. Thanks to Andy Hock and Angela Yeung, and to Griffin Marge who made it all happen and to Jeff Strater for coordinating.
At the Computer History Musem, I was glad to see that one of this year’s inductees to CHM was Barbara Jane Liskov. She was one of my profs at Massachusetts Institute of Technology when I was learning structured programming. She developed the programming language CLU. This language introduced basic concepts similar to classes and abstractions. This is now a standard in all programming languages in terms of packages and namespace declaration of procedures / methods, the basis for OOP.
MemCon 2024 has been a pivotal event for addressing the critical challenges in GenAI compute. The insights shared by experts have shed light on the importance of overcoming memory bottlenecks to enhance processing capabilities.




