Small Language Models for Production Applications

Deploy small language models in production for edge, privacy, and cost-optimized AI workloads. SLM fundamentals cover the size taxonomy, comparison with large models across parameters, context, cost,...

Read More
8781  Reviews star_rate star_rate star_rate star_rate star_half
$2,000USD
Duration 2 days
Course Code GAI-2105
Available Formats Classroom

Overview

Course Description

Deploy small language models in production for edge, privacy, and cost-optimized AI workloads. SLM fundamentals cover the size taxonomy, comparison with large models across parameters, context, cost, and latency, the use cases that favour SLMs, and capability tradeoffs at each size point. Deployment, quantization, and task design cover deployment options, quantization techniques, hardware sizing, local-execution tools, task-suitability frameworks, context-length constraints, chunking and summarization, prompt compression, and SLM-specific prompt optimization. Fine-tuning with LoRA and QLoRA covers when to fine-tune versus prompt-engineer, LoRA mechanics, parameter-efficient training, QLoRA for quantized training, and dataset sizing to avoid overfitting. Retrieval augmentation, evaluation, and production cover lightweight RAG architectures suited to SLM context budgets, retrieval-quality versus context-budget tradeoffs, hybrid architectures with model routing, SLM-specific failure modes, and gradual-rollout strategies. Hands-on labs produce a quantized SLM running locally, a LoRA-fine-tuned variant, a lightweight RAG system, and a production-readiness benchmark report. The course is designed for ML engineers, data scientists, software developers, and technical architects with prior LLM experience.

Skills Gained

By the end of this course, participants will be able to:

  • Recognize SLM capabilities and optimal use cases relative to large models
  • Configure quantized models for local deployment with optimized memory footprint
  • Apply prompt engineering tailored to SLM constraints
  • Evaluate when to fine-tune versus prompt-engineer for specific tasks
  • Build lightweight RAG systems with retrieval strategies suited to SLMs
  • Benchmark SLM performance for production deployment

Who Can Benefit

This course is designed for:

  • ML Engineers
  • Data Scientists
  • Software Developers
  • Technical Architects

Prerequisites

Participants should enter this course with:

  • GAI-1201 or equivalent knowledge of LLMs
  • Basic prompt engineering skills
  • Simple app development exposure with LLMs

Organizational Objectives

This course assists organizations to:

  • Reduce inference cost by routing to SLMs where capability allows it
  • Lower data-egress and privacy risk by keeping inference at the edge
  • Build a working knowledge of SLM deployment and fine-tuning across the team
  • Establish gradual-rollout patterns that catch SLM-specific failure modes before they ship

Software

All attendees must have a modern web browser and an Internet connection.

Course Details

Course Details

Module 1 - What Are SLMs and Why They Matter

By the end of this module, you will be able to compare small and large language models across parameters, context, cost, and latency; identify SLM-friendly use cases for edge, privacy, and cost; and reason about capability tradeoffs.

  • Size Taxonomy and Model Categories
  • Comparing Small and Large Models
  • Parameters, Context Length, Cost, and Latency
  • Use Cases for Edge Deployment, Privacy, and Cost Optimization
  • Trade-offs in Capability and Performance
  • Hands-on Lab: API Comparison of Large and Small Models

Module 2 - Deployment Modalities and Quantization

By the end of this module, you will be able to choose between deployment options for SLMs, apply quantization techniques, size hardware to model and quantization choices, and operate the tools commonly used to deploy SLMs.

  • Deployment Options for SLMs
  • Quantization Techniques for Inference
  • Memory Requirements by Model Size and Quantization
  • Hardware Requirements and Considerations
  • Tools for SLM Deployment
  • Hands-on Lab: Running a Quantized SLM Locally

Module 3 - Task Design and Prompt Engineering for SLMs

By the end of this module, you will be able to apply a task-suitability framework for SLMs, manage shorter context windows through chunking and summarization, optimize prompts for SLMs, and evaluate prompt effectiveness.

  • Task Suitability Framework for SLMs
  • Context Length Constraints and Management
  • Chunking, Summarization, and Prompt Compression
  • Prompt Optimization Techniques
  • Evaluating Prompt Effectiveness
  • Hands-on Lab: Prompt Engineering and Context Management

Module 4 - Fine-Tuning with LoRA/QLoRA

By the end of this module, you will be able to choose between fine-tuning and prompt engineering for SLMs, apply LoRA and QLoRA for parameter-efficient training, size datasets to avoid overfitting, and evaluate fine-tuned SLMs against the base.

  • When to Fine-Tune vs Prompt Engineer
  • LoRA Mechanics and Benefits
  • Training Efficiency with Parameter-Efficient Methods
  • QLoRA for Quantized Training
  • Dataset Requirements and Overfitting Prevention
  • Hands-on Lab: Evaluating Fine-Tuned vs Base SLMs

Module 5 - Retrieval Augmentation and Tool Use

By the end of this module, you will be able to build lightweight RAG suited to SLM context budgets, balance retrieval quality against context, choose chunking and reranking strategies, and design hybrid architectures with model routing.

  • Lightweight RAG Architecture for SLMs
  • Retrieval Quality vs Context Budget Trade-offs
  • Optimal Chunking Strategies
  • Retrieval Parameters and Reranking
  • Tool Augmentation for Enhanced Capabilities
  • Hybrid Architectures with Model Routing
  • Hands-on Lab: Building a Lightweight RAG System

Module 6 - Evaluation, Benchmarking and Production Considerations

By the end of this module, you will be able to choose evaluation metrics specific to SLMs, recognize SLM-specific failure modes, design A/B tests and gradual rollout strategies, and track emerging trends in small language models.

  • Evaluation Metrics for SLMs
  • SLM-Specific Failure Modes
  • A/B Testing Strategy and Gradual Rollout
  • Emerging Trends in Small Language Models
  • Hands-on Lab: Comprehensive Benchmarking

Schedule

FAQ

Does the course schedule include a Lunchbreak?

Classes typically include a 1-hour lunch break around midday. However, the exact break times and duration can vary depending on the specific class. Your instructor will provide detailed information at the start of the course.

What languages are used to deliver training?

Most courses are conducted in English, unless otherwise specified. Some courses will have the word "FRENCH" marked in red beside the scheduled date(s) indicating the language of instruction.

What does GTR stand for?

GTR stands for Guaranteed to Run; if you see a course with this status, it means this event is confirmed to run. View our GTR page to see our full list of Guaranteed to Run courses.

Does Ascendient Learning deliver group training?

Yes, we provide training for groups, individuals and private on sites. View our group training page for more information.

What does vendor-authorized training mean?

As a vendor-authorized training partner, we offer a curriculum that our partners have vetted. We use the same course materials and facilitate the same labs as our vendor-delivered training. These courses are considered the gold standard and, as such, are priced accordingly.

Is the training too basic, or will you go deep into technology?

It depends on your requirements, your role in your company, and your depth of knowledge. The good news about many of our learning paths, you can start from the fundamentals to highly specialized training.

How up-to-date are your courses and support materials?

We continuously work with our vendors to evaluate and refresh course material to reflect the latest training courses and best practices.

Are your instructors seasoned trainers who have deep knowledge of the training topic?

Ascendient Learning instructors have an average of 27 years of practical IT experience and have also served as consultants for an average of 15 years. To stay current, instructors spend at least 25 percent of their time learning new, emerging technologies and courses.

Do you provide hands-on training and exercises in an actual lab environment?

Lab access is dependent on the vendor and the type of training you sign up for. However, many of our top vendors will provide lab access to students to test and practice. The course description will specify lab access.

Will you customize the training for our company’s specific needs and goals?

We will work with you to identify training needs and areas of growth.  We offer a variety of training methods, such as private group training, on-site of your choice, and virtually. We provide courses and certifications that are aligned with your business goals.

How do I get started with certification?

Getting started on a certification pathway depends on your goals and the vendor you choose to get certified in. Many vendors offer entry-level IT certification to advanced IT certification that can boost your career. To get access to certification vouchers and discounts, please contact info@ascendientlearning.com.

Will I get access to content after I complete a course?

You will get access to the PDF of course books and guides, but access to the recording and slides will depend on the vendor and type of training you receive.

How do I request a W9 for Ascendient Learning?

View our filing status and how to request a W9.

Reviews

Sean is the very good instructor. I would like to take his class again in the future.

Course was great and the instructor had an answer for anything that was asked during the course.

Class was very informative, although one lab didnt but will try again later

I think the platform is very good and look forward to taking my next course in early October.

I didn't have any problem navigating Exitcertified website or lab material at all.