Let’s make it work for you 

Overview

Visual data is becoming a critical asset for modern organisations. This course explores how to build intelligent applications that can interpret, analyse, and reason over images and documents using Azure AI services and multimodal models.

We believe organisations that can combine data, AI, and cloud technologies will unlock faster and more accurate decision-making. This course focuses on applying multimodal and agent-based AI patterns to extract structured insights from visual inputs and integrate them into real-world workflows. Learners will gain practical experience in designing solutions that combine visual understanding with language models, enabling applications that move from perception to action within Azure environments.

Read more +

Prerequisites

Participants should have:

  • Basic programming experience in a language such as Python
  • A general understanding of cloud computing or artificial intelligence concepts
  • Familiarity with data handling and application development workflows is beneficial
  • No prior experience in computer vision is required

Target audience

This course is designed for:

  • Developers building intelligent and data-driven applications
  • AI engineers working with multimodal and generative AI solutions
  • Technical professionals implementing Azure-based AI services
  • Teams looking to integrate visual data into decision-making workflows
Read more +

Delegates will learn how to

By the end of this course, learners will be able to:

  • Build applications that analyse and reason over visual data using Azure AI services
  • Combine image and document inputs with language models for enhanced understanding
  • Implement multimodal and agent-based workflows for AI applications
  • Extract structured information from images and documents
  • Design pipelines that ground AI responses in visual and document context
  • Apply design patterns for orchestrating tools and services in Azure AI solutions
Read more +

Outline

Introduction to visual and multimodal AI

  • Overview of visual AI and its role in modern applications
  • Understanding multimodal models and their capabilities
  • Key use cases for image and document intelligence
  • Challenges in processing and interpreting visual data

Working with Azure AI services for visual data

  • Introduction to Azure AI vision and document intelligence services
  • Processing images and extracting metadata
  • Analysing documents for structured and unstructured information
  • Integrating visual services into application workflows

Multimodal models for images and documents

  • Combining visual and language inputs in AI systems
  • Understanding embeddings for multimodal data
  • Enhancing context through cross-modal reasoning
  • Use cases for multimodal generative AI applications

Structured information extraction from visual inputs

  • Techniques for extracting entities and key-value pairs from documents
  • Processing forms, receipts, and structured layouts
  • Handling unstructured visual data sources
  • Improving accuracy through preprocessing and validation

Grounding AI responses in visual data

  • Principles of grounding in multimodal AI systems
  • Linking model outputs to visual context
  • Reducing hallucinations through grounded responses
  • Designing systems that maintain traceability to source data

Agent-based orchestration for visual workflows

  • Introduction to agent-based AI patterns
  • Orchestrating tools and services for visual data processing
  • Designing workflows that combine reasoning and action
  • Integrating APIs and external tools into AI pipelines

Designing decision-making workflows using visual data

  • Building end-to-end pipelines for insight extraction
  • Connecting visual analysis to business decision processes
  • Automating actions based on extracted insights
  • Monitoring and improving workflow performance over time

Exams and assessments

There are no formal exams included in this course. Learners will complete interactive knowledge checks and practical exercises to reinforce their understanding of multimodal AI concepts and visual data processing techniques.

Hands-on learning

This course includes:

  • Guided labs using Azure AI services for image and document analysis
  • Practical exercises in multimodal model integration
  • Scenario-based tasks focused on real-world visual data challenges
  • Instructor-led discussions on designing production-ready AI workflows
Read more +

Why choose QA

Dates & Locations

Yellow
Need to know

Frequently asked questions

How can I create an account on myQA.com?

There are a number of ways to create an account. If you are a self-funder, simply select the "Create account" option on the login page.

If you have been booked onto a course by your company, you will receive a confirmation email. From this email, select "Sign into myQA" and you will be taken to the "Create account" page. Complete all of the details and select "Create account".

If you have the booking number you can also go here and select the "I have a booking number" option. Enter the booking reference and your surname. If the details match, you will be taken to the "Create account" page from where you can enter your details and confirm your account.

Find more answers to frequently asked questions in our FAQs: Bookings & Cancellations page.

How do QA’s virtual classroom courses work?

Our virtual classroom courses allow you to access award-winning classroom training, without leaving your home or office. Our learning professionals are specially trained on how to interact with remote attendees and our remote labs ensure all participants can take part in hands-on exercises wherever they are.

We use the WebEx video conferencing platform by Cisco. Before you book, check that you meet the WebEx system requirements and run a test meeting to ensure the software is compatible with your firewall settings. If it doesn’t work, try adjusting your settings or contact your IT department about permitting the website.

How do QA’s online courses work?

QA online courses, also commonly known as distance learning courses or elearning courses, take the form of interactive software designed for individual learning, but you will also have access to full support from our subject-matter experts for the duration of your course. When you book a QA online learning course you will receive immediate access to it through our e-learning platform and you can start to learn straight away, from any compatible device. Access to the online learning platform is valid for one year from the booking date.

All courses are built around case studies and presented in an engaging format, which includes storytelling elements, video, audio and humour. Every case study is supported by sample documents and a collection of Knowledge Nuggets that provide more in-depth detail on the wider processes.

When will I receive my joining instructions?

Joining instructions for QA courses are sent two weeks prior to the course start date, or immediately if the booking is confirmed within this timeframe. For course bookings made via QA but delivered by a third-party supplier, joining instructions are sent to attendees prior to the training course, but timescales vary depending on each supplier’s terms. Read more FAQs.

When will I receive my certificate?

Certificates of Achievement are issued at the end the course, either as a hard copy or via email. Read more here.

Let's talk

A member of the team will contact you within 4 working hours after submitting the form.

By submitting this form, you agree to QA processing your data in accordance with our Privacy Policy.