MAIAMedia AI Engine · Media AI Agent

The AI engine suite built for broadcast

MAIA (Media AI Agent) is a unified engine that analyzes video, speech, faces, on-screen text, and prompting with AI — turning vast media libraries into searchable assets.

Contact Us

On-premises deployment — your content never leaves the facility

MAIA AI Technology screenshot
Key Features

Everything You Need for AI Technology

MAIA Video — Scene Understanding

AI

Segments video into frames, shots, and scenes, then converts each scene's content into text stored as searchable metadata.

MAIA Speech — Speech Recognition (STT)

AI

Multiple STT engines — Google, Amazon, Naver Clova, Whisper — plus speaker diarization turn every utterance into searchable text.

MAIA Face — Face Recognition

AI

Automatically extracts and clusters people in video, then finds every appearance of a person across the entire archive from a single photo.

MAIA Character — Text Recognition (OCR)

AI

Detects on-screen text — captions, lower thirds, CG — and indexes it in searchable form.

MAIA Object — Object Recognition

AI

Detects 133 object classes and generates structured metadata for video at the scene level.

MAIA Prompter — AI Prompter

AI

AI matches the presenter's voice to the script in real time and scrolls automatically. No dedicated operator required.

Natural-Language Unified Search

AI

Unifies face, object, STT, and scene data so you can search the archive in everyday language.

Feature Details

Five Engines, One MAIA

MAIA Video Scene Recognition screenshot

Scene Recognition

Scene Change Detection

AI automatically segments video into frames, shots, and scenes, organizing each segment into structured metadata. It runs on-premises, so your content never leaves the facility.

  • 01Object detection
    classifies and auto-tags more than 116 classes — people, vehicles, animals, backgrounds — using panoptic segmentation.
  • 02Video summaries & scene descriptions
    generative AI writes a human-readable description for every scene.
  • 03Natural-language search
    unifies face, object, STT, and scene data so you can search the archive in everyday language.
MAIA Speech Speech Recognition (STT) screenshot

Speech Recognition (STT)

Speech-to-Text

Every word spoken on air becomes searchable text — automatically, accurately, and in real time.

  • 01Multi-engine STT hub
    choose from Google, Amazon Transcribe, Naver Clova, OpenAI Whisper, and Daglo.
  • 02Speaker diarization
    identifies who said what, and when, with timecode.
  • 03Keyword search within captions jumps straight to edit points, with AI caption editing, automatic summaries, and caption file downloads.
  • 04Deploy as cloud SaaS (usage-based billing) or fully on-premises with Whisper.
MAIA Face Face Recognition screenshot

Face Recognition

Face Recognition

Find every appearance of an on-air talent across the entire archive in seconds, not hours.

  • 01Automatic face extraction
    detects faces via landmark analysis and clusters the same person automatically.
  • 02Image-based search
    upload a single photo and instantly find every matching appearance.
  • 03Person timelines
    shows each person's appearances on a shot-level visual timeline.
MAIA Character Text Recognition (OCR) screenshot

Text Recognition (OCR)

OCR Detection

Automatically detects, indexes, and searches lower thirds, channel chyrons, and on-screen graphic text.

  • 01Full-frame analysis mode
    detects every text region on screen automatically, with no setup.
  • 02Region-select mode
    drag to define regions of interest and extract text from exactly the areas you want.
  • 03Supports multiple languages
    Korean, English, Chinese, Japanese, and more — across a wide range of fonts.
MAIA Prompter AI Prompter screenshot

AI Prompter

AI-Powered Live Prompting

A prompter that listens to the presenter, matches their speech to the script in real time, and scrolls automatically.

  • 01Real-time voice matching
    analyzes the presenter's speech and tracks the script position automatically.
  • 02No operator required
    scrolls fully automatically, with no dedicated staff.
  • 03Deploys on-premises or in the cloud (including Whisper Live).
How It Works

One Seamless Workflow

01Ingest
02AI Analysis (Video · Speech · Face · Text)
03Metadata & Indexing
04Natural-Language Search
05Use & Reuse
Performance · Specs

Proven Technology, Measured Results

116+ classes
Automatic object detection
On-premises
Content stays in your facility
MODULESVideo · Speech · Face · Character · Prompter
STTGoogle · Amazon · Naver Clova · Whisper · Daglo
VISIONScene segmentation · 116+ object classes · OCR · Face recognition
SEARCHUnified natural-language search across faces, objects, STT, and scenes
DEPLOYOn-premises · Cloud SaaS · Hybrid

Evaluating MAIA?

Tell us about your broadcast workflow and our team will design how AI Technology fits into it, together with you.