AWS AI In Practice #6

Wednesday, 26 August 2026 at 17:00

AutogenAI, 123 Pentonville Rd, N1 9LG

FreeAITechDesignBusinessFinanceCulture

We’re delighted to welcome Anton Nazaruk, CTO, Cloud Combinator, Ivaylo Iliev, Partner Solutions Architect, AWS and Silvia Lehnis, Chief AI Officer, UBDS Digital Here’s what they’re bringing: Seven ways to buy GPU compute on AWS. One wrong choice and the bill climbs while your training run stalls. Anton builds production-ready data and AI platforms at Cloud Combinator, and he’s bringing the receipts - a pre-recorded walkthrough of a real distributed training run, including a node failure, replacement, and full job recovery. Tonight he’s joined by Ivaylo, giving us the complete blueprint: Capacity Blocks, SageMaker HyperPod, EFA networking, FSx storage, and Slurm or EKS orchestration, with cost control designed in from day one. Silvia and her team at UBDS Digital took an intelligent document processing solution on Amazon Bedrock and tested everything - prompting, model selection, document handling, and validation - to find out what actually moves extraction accuracy on messy, real-world documents. Tonight she’s walking us through the results: which methods earn their keep, which don’t, and the trade-offs between value and implementation effort. A big thank you to our sponsors Cloudscaler, Rayo & The Scale Factory for making this event possible. Programme: 18:00: Arrival, registration 18:15: Talks start 20:00: Networking with food and a drink provided by the generosity of our sponsors. Session 1: *From Zero to HyperPod: Choosing, Operating, and Cost-Controlling Distributed Model Training Infrastructure on AWS* with Anton Nazaruk & Ivaylo Iliev Training large models on AWS is no longer just an ML problem - it is a capacity, infrastructure, reliability, and cost-control problem. This talk gives engineers a practical framework for choosing between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, SageMaker Training Plans, and SageMaker HyperPod. We’ll look at when each option makes sense, what tradeoffs they introduce, and how to avoid common mistakes around GPU availability, quota planning, interruptions, and runaway cost. We’ll then walk through a repeatable distributed training blueprint: compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, checkpointing, and failure recovery. The session includes a demo-style walkthrough of launching a distributed training job and showing how node failure and recovery should be handled in a production-ready setup. The goal is for attendees to leave with a clear mental model of how to run distributed model training on AWS reliably, how to choose the right service or capacity model, and how to make cost and failure recovery part of the architecture from day one. Learning Takeaways • Choose between EC2 On-Demand, Spot, Savings Plans, Capacity Blocks, SageMaker Training Jobs, Training Plans, and HyperPod with a clear framework for when each makes sense. • Build a repeatable distributed training blueprint - compute fleet, EFA networking, FSx/S3 storage, Slurm or EKS orchestration, observability, and checkpointing. • Design cost control and node failure recovery into your training architecture from day one. Anton Anton Nazaruk is CTO at Cloud Combinator, where he works on cloud architecture, AI infrastructure, and distributed systems. He helps teams design production-ready platforms for data and AI workloads on AWS, with a focus on reliability, cost control, and repeatable infrastructure patterns. Ivaylo is a Partner Solutions Architect at AWS, where he helps organisations optimise their cloud solutions and accelerate their digital transformation journeys. With a focus on artificial intelligence and machine learning, he works closely with AWS partners to develop innovative solutions that address real-world business challenges. His expertise spans cloud architecture, AI implementation, and building scalable solutions for customers across diverse industries. Session 2: *Beyond Prompt Optimisation: Improving Accuracy in LLM-Based Document Extraction* with Silvia Lehnis Variable document formats make structured data extraction difficult to solve with prompting alone. Especially when the documents include everything from images, handwriting, complex financial tables repeated multiple times and without a particular form or naming conventions. This talk explores how we optimised an intelligent document processing solution on Amazon Bedrock by testing 17 different variants across prompting, model selection, document handling and validation. We will share the methods to optimise an LLM based solution that apply to many use-cases beyond documents, as well as the key trade-offs, failure modes and design decisions that had the greatest impact on extraction accuracy and reliability. Learning Takeaways • How to test and compare AI solution options, including model selection, experiment design and evaluation methods. • How the findings shaped key design decisions and the final Amazon Bedrock solution released into production. • Which methods can improve the performance of an LLM-based solution, and the trade-offs between value and implementation effort. Silvia helps high-impact organisations use data and AI to reach their goals faster and safer. She has led global and national data and AI transformations in sensitive environments across the public sector, finance, energy and academia, taking strategy through to implementation across people, process and technology - with solutions reaching up to 110,000 users and delivering £27m in savings over three years. She’s also a board member of the charity Care in Action. Do you have a story to share? If you are interested in speaking at one of our events, please check out our call for papers. We are advocates for greater inclusion & diversity in UK Tech and are especially keen to receive talk submissions from people in underrepresented groups. If you are interested in speaking at a future meetup but would like to discuss what to expect or need assistance, please contact our Inclusion & Diversity Lead Natalie Gray - graynataliej@gmail.com or DM her @natjgray Check out our website for more information about our community and our code of conduct. Remember to follow us @AWSUserGroupUK and on LinkedIn for the latest updates, and you can find videos of our past meetups here.

View event details →

Listed from meetup. Details can change — always check with the organiser.