---
title: "AI video and image glossary"
description: "Plain definitions of the terms that appear across Higgsfield and every other AI video and image tool — generation modes, control, quality and commercial words."
url: "https://higgsfield.wiki/glossary/"
verified: "2026-08-26"
publisher: "Higgsfield Wiki — independent reference, not affiliated with Higgsfield AI"
---

# AI video and image glossary

Plain definitions of the terms that appear across Higgsfield and every other AI video and image tool — generation modes, control, quality and commercial words.

Terms used throughout this wiki and across the category, in plain language.

## Generation

**Text-to-image (T2I)** — producing an image from a written prompt alone.

**Text-to-video (T2V)** — producing a clip from a written prompt alone. The model invents every frame.

**Image-to-video (I2V)** — animating a supplied still. The starting frame is fixed, so you control the look and only ask for motion.

**Image-to-image (I2I)** — producing a new image anchored to one you supply, rather than from description alone.

**First frame / last frame** — supplying the opening and closing images of a clip so the model interpolates between them.

## Control and consistency

**Identity preservation** — keeping a specific person recognisably themselves through a generation or an edit. Not achievable by describing them; it requires a reference image and a mode built for it.

**Reference image** — an image supplied to anchor subject, style or composition.

**Drift** — the gradual change of a subject's appearance over a clip, or across a series. The main reason long generations are unreliable.

**Temporal coherence** — whether a video stays consistent frame to frame. Its absence shows as flicker or shimmer.

**Multi-shot** — generating several distinct camera setups that hold together as one scene.

**Seed** — the number initialising a generation. The same seed with the same prompt and settings reproduces the same output.

## Quality and processing

**Upscaling** — increasing output resolution after generation, adding plausible detail rather than recovering real detail.

**Artefact** — a visible defect introduced by generation: warped hands, garbled text, smeared edges.

**Lip sync** — matching mouth movement to speech. Done natively by models that generate audio; bolted on afterwards otherwise, which is the most common tell in AI video.

**Inpainting** — regenerating a selected region while leaving the rest untouched.

**Outpainting** — extending an image beyond its original borders.

## Commercial

**Credit** — the unit consumed by a generation. Cost per run varies by model, media type, duration and resolution.

**Watermark** — an identifying mark applied to output, common on free tiers.

**UGC** — user-generated content, or in advertising, content styled to look like it.

**Aspect ratio** — frame proportions. 9:16 vertical for short-form social, 16:9 horizontal for landscape, 1:1 square.

## Common questions

### What is the difference between text-to-video and image-to-video?

Text-to-video invents every frame from a prompt. Image-to-video starts from a still you supply, so the look is already settled and you are only asking for motion — far more controllable.

### What does identity preservation mean?

Keeping a specific person recognisably themselves through a generation. Describing someone in a prompt does not do it; it needs a reference image and a mode designed for it.

