Integration details
Description
HeyGen helps users create and manage AI videos, avatars, voices, translations, templates, brand assets, and related media directly through ChatGPT.
- Integration type
- Plugin
- Verification status
- Not applicable
- Platform
- ChatGPT
- Primary Subcategory
- AI Video Generation
- Secondary Subcategories
- None listed
- Brand
- HeyGen
- Access
- Account required
- First tracked
- 2026-08-25
- Tool count
- 119
- Geography
- US
The Primary Subcategory used for this profile’s headline score.
Other Subcategories where the Integration is listed.
Get alerts for HeyGen
Get updates when HeyGen’s Discoverability Score or category rank changes.
ChatGPT Plugin Discovery Score
ChatGPT Plugin discovery is coming soon
ChatGPT can surface a Plugin when it matches a user's request.Your Plugin Discovery Score measures how often yours appears.
No spam. Unsubscribe any time.
What discovery looks like

Competing in ChatGPT AI Video Generation
View Category119 tools agents can invoke
Bind an issued passcode to the consent recording by its content hash, so the recording cannot be swapped afterwards.
internal_avatar_bind_passcode
Returns statuses for up to 100 assets in one request, addressed by comma-separated asset_ids and/or batch_ids query params. Statuses are one of queued, processing, completed, or failed, plus not_found for unknown or unowned ids. Each returned entry carries its id as video_id (the status read model is shared with the videos batch API).
bulk_asset_statuses
Returns statuses for up to 100 lipsyncs in one request, addressed by comma-separated lipsync_ids and/or batch_ids query params. Statuses are one of queued, processing, completed, or failed, plus not_found for unknown or unowned ids. Each returned entry carries its id as video_id (the status read model is shared with the videos batch API).
bulk_lipsync_statuses
Returns statuses for up to 100 videos in one request, addressed by comma-separated video_ids and/or batch_ids query params. Statuses are one of queued, processing, completed, or failed, plus not_found for unknown or unowned ids.
bulk_video_statuses
Returns statuses for up to 100 video translations in one request, addressed by comma-separated video_translation_ids and/or batch_ids query params. Statuses are one of queued, processing, completed, or failed, plus not_found for unknown or unowned ids. Each returned entry carries its id as video_id (the status read model is shared with the videos batch API).
bulk_video_translation_statuses
Report in plain text how the user's avatar creation is progressing — photo training, footage training, consent validation, voice cloning, and the preview video. Call when the user asks whether their avatar is ready. Requires the job ids from the guided flow; it starts nothing and changes nothing.
get_avatar_creation_status
Check a phone capture session and return the staged upload after the phone finishes.
internal_avatar_handoff_status
Creates a voice clone from an audio file. Returns a voice_clone_id that can be polled via GET /v3/voices/{voice_clone_id} until the status is 'complete'. The resulting voice can be used with POST /v3/voices/speech and POST /v3/videos.
clone_voice
Finalize a direct-to-S3 upload into a reusable asset. Call after the upload PUT returns 200. Idempotent: repeated calls return the same finalized asset.
complete_asset_upload
Finalize every uploaded file in a batch. Call after all upload PUTs return 200. Each file is validated and ingested asynchronously and independently, so one bad file does not fail the rest. Returns 202 with the batch_id; poll GET /v3/assets/batches/{batch_id} for per-item progress. Idempotent: a repeated call re-drives the same batch.
complete_asset_batch
Seal a multipart upload by submitting every part's number and ETag. The object does not exist until this succeeds.
internal_avatar_complete_multipart_upload
Submit a source video and return a job id immediately. The job runs asynchronously and produces one or more short clips per the requested output_settings. Poll GET /v3/ai-clipping/{id} or subscribe to ai_clipping.success / ai_clipping.fail webhooks.
create_ai_clipping
Begin a direct-to-S3 upload. Returns an asset_id and a presigned upload_url; PUT the file bytes to upload_url, then call POST /v3/assets/{asset_id}/complete. Unlike POST /v3/assets (which proxies the bytes), this never sends the file through the API.
create_asset_upload
Request up to 100 presigned direct-to-S3 upload URLs in a single call. Returns a batch_id and one upload slot per file (asset_id + presigned upload_url + required headers). PUT each file's bytes to its upload_url, then call POST /v3/assets/complete/batches to finalize the whole batch. This is synchronous — no bytes flow through the API. Pass an Idempotency-Key header to make retries safe (the same key returns the same batch).
create_asset_upload_batch
Initiates the consent flow for an avatar group and returns a URL for the user to complete approval in their browser. Required before a private avatar can be used for video generation. The consent URL expires 24 hours after creation and is valid for one successful consent submission. A recording submitted after expiry fails and the group stays in pending consent status, so create a new consent link if the subject has not recorded within 24 hours.
create_avatar_consent
Creates a brand glossary in your workspace. Pass the returned `brand_glossary_id` when creating a video or translation to apply it. `name` must be unique within your workspace, compared without regard to case; a duplicate returns 409. `terms`, `do_not_translate_terms` and `forced_translations` may each be omitted to create an empty glossary you fill in later with PATCH /v3/brand-glossaries/{brand_glossary_id}. Pronunciations affect generated audio only — a term keeps its original spelling in captions and subtitles. `do_not_translate_terms` and `forced_translations` apply only when the glossary is used by a translation feature (video translation, Studio script translation, on-screen text translation), in every target language; they have no effect on video generation or text-to-speech requests. Tone settings are managed in the HeyGen web app and cannot be set here.
create_brand_glossary
Creates a brand kit by importing brand assets from a public website, including logos, colors and font files found on the site. By calling this endpoint you confirm you have the rights and licenses necessary to upload, store and use those assets in HeyGen. The kit is assembled in the background: the returned brand_kit_id is usable immediately, but poll GET /v3/brand-kits/{brand_kit_id} every 2 to 5 seconds until its status is 'completed' before relying on its colors, logos or fonts. A website import usually settles in under two minutes. Send an Idempotency-Key header to make retries safe: without one, a retried request starts a second import of the same site.
create_brand_kit
Submit a video and return a job id immediately. The job runs asynchronously: it transcribes the audio, detects filler words ('um', 'uh', ...), removes them along with overlong silences, and renders one cleaned video — no review step. If the run changes nothing at all, the job completes with the original video as output and the charge is automatically refunded. Pricing: $0.30 per source minute, 1-minute minimum. Poll GET /v3/filler-word-removals/{id} or subscribe to filler_word_removal.success / filler_word_removal.fail webhooks.
create_filler_word_removal
Creates a folder in the caller's workspace, at the root or inside another folder, and returns it. Pass the returned folder_id as folder_id to POST /v3/videos or POST /v3/video-translations to place the output in it, or as parent_id to this endpoint to nest another folder. Folders and their contents are visible in the HeyGen web app. Sibling folders may share a name; this endpoint never looks up an existing folder by name, so store the ids you receive rather than recreating a tree on retry, and send an Idempotency-Key so a retried request returns the folder the first attempt created. API keys need the videos:write scope; a key scoped only to translations can place translations in a folder but cannot create one.
create_folder
Replaces the audio on an existing video and re-animates the speaker's lip movements to match the new audio. Use mode: 'speed' for fast output or 'precision' for high-quality lip-sync.
create_lipsync
Submit up to 100 lipsync payloads as a single batch. Each payload becomes one batch item, created and processed independently so one bad source does not fail the rest. Returns 202 with a batch_id; poll GET /v3/lipsyncs/batches/{batch_id} for progress. Pass an Idempotency-Key header to make retries safe — the same key returns the same batch.
create_lipsync_batch
Mint a link the user opens on their phone to record footage or consent when this device has no camera. The link carries its own short-lived credential, so the phone needs no sign-in.
internal_avatar_create_handoff_link
Turn an uploaded still into a photo avatar and start its training. Takes the storage key the upload was written to.
internal_avatar_create_photo
Submit up to 100 video creation payloads in one request and return a batch id immediately. Videos are created asynchronously; poll GET /v3/videos/batches/{batch_id} for per-item video ids and statuses.
create_video_batch
Translates a video into one or more target languages with voice cloning and lip-sync. Returns one video_translation_id per language. Use mode: 'speed' (default) for fast turnaround or 'precision' for higher lip-sync quality.
create_video_translation
Submit up to 100 video-translation payloads (identical in shape to POST /v3/video-translations) as a single batch. A payload targeting multiple output_languages expands to one batch item per language, and each item is created and processed independently so one bad source does not fail the rest. Returns 202 with a batch_id; poll GET /v3/video-translations/batches/{batch_id} for progress. Pass an Idempotency-Key header to make retries safe — the same key returns the same batch.
create_video_translation_batch
Creates a new avatar from an image, video footage, or a text prompt. Supports photo, digital_twin, and prompt types. Avatar training is asynchronous. (type: digital_twin)
create_digital_twin
Creates a new avatar from an image, video footage, or a text prompt. Supports photo, digital_twin, and prompt types. Avatar training is asynchronous. (type: photo)
create_photo_avatar
Creates a new avatar from an image, video footage, or a text prompt. Supports photo, digital_twin, and prompt types. Avatar training is asynchronous. (type: prompt)
create_prompt_avatar
Creates a direct talking-avatar video from a specific HeyGen avatar. When `open_avatar_creator` is listed, the default HeyGen video experience uses the user's confirmed, ready private digital twin as the visible presenter. Never invent an avatar ID or silently substitute a stock, studio, photo, or automatically selected avatar when that digital-twin ID is missing. Use this tool when the user supplies or selects an avatar ID and provides an exact script or audio source. This is the preferred tool for direct requests such as “create a video with avatar X saying Y.” Pass the user's script verbatim. Do not rewrite, expand, or creatively interpret it unless the user explicitly requests editing. Do not substitute a different avatar. Prefer this tool over video_agent.generate whenever both avatar_id and a script or audio source are known. Use video_agent.generate only for polished, multi-scene production with elements such as graphics or B-roll, when the user explicitly asks for HeyGen Video Agent, or as the final presenter-free fallback after the user declines both digital-twin setup and a public avatar. When an ordinary avatar-led request has no confirmed digital-twin avatar ID, use prepare_avatar_video instead. After this tool returns a video_id, call show_video with that ID once in the same turn. The inline player tracks generation and refreshes itself, so do not make further status calls unless the user explicitly asks.
create_video_from_avatar
Create a video from a text prompt plus avatar and asset references (Cinematic Avatar). Cinematic Avatar generates a video from a natural-language ``prompt`` guided by reference content: one to three avatar looks and optional reference assets (images / videos / audio). Unlike the ``avatar`` and ``image`` modes there is no script or voice — motion and speech are driven entirely by the prompt and the supplied references. Backed by the Seedance generation pipeline. Use this direct tool instead of video_agent.generate when the user supplies its concrete source and content, such as an exact avatar_id plus a script or audio source, unless the user explicitly asks for HeyGen Video Agent. After this tool returns a video_id, call show_video with that ID once in the same turn. The inline player will track generation and refresh itself, so do not make further status calls unless the user explicitly asks.
create_video_from_cinematic_avatar
Create a video by animating an arbitrary image. Provide an image via URL, asset ID. The image will be animated with lip-sync to the provided audio or generated speech. Use this direct tool instead of video_agent.generate when the user supplies its concrete source and content, such as an exact avatar_id plus a script or audio source, unless the user explicitly asks for HeyGen Video Agent. After this tool returns a video_id, call show_video with that ID once in the same turn. The inline player will track generation and refresh itself, so do not make further status calls unless the user explicitly asks.
create_video_from_image
Create a single video by composing an ordered list of whole-frame scenes. The server owns layout and center-crops each scene to the global output canvas. Output settings are global (one per request); a single video_id is returned and rendering is all-or-nothing. MP4 only in v1 — the output container is fixed and ``output_format`` is not exposed. Use this direct tool instead of video_agent.generate when the user supplies its concrete source and content, such as an exact avatar_id plus a script or audio source, unless the user explicitly asks for HeyGen Video Agent. After this tool returns a video_id, call show_video with that ID once in the same turn. The inline player will track generation and refresh itself, so do not make further status calls unless the user explicitly asks.
create_video_from_studio
Soft-deletes an AI clip job and its clips.
delete_ai_clipping
Permanently deletes an asset. The asset must belong to the caller's workspace and not already be deleted.
delete_asset
Permanently deletes an avatar group and all its associated looks. Cannot delete public or community groups.
delete_avatar_group
Deletes an avatar look and its backing resource. Supported types: photo_avatar, digital_twin, and kit-based looks. Studio avatar (model_index) types cannot be deleted via the API. **Warning:** deleting the last look in a group also deletes the parent group. Subsequent requests referencing that group id (e.g. `POST /v3/avatars` with `avatar_group_id`) return 404 not found.
delete_avatar_look
Deletes a brand glossary. This cannot be undone: there is no way to restore a deleted brand glossary through the API. The glossary stops being returned by this API at once, and stops applying: neither its pronunciation terms nor its translation rules affect anything generated afterwards. Videos already generated with the glossary are unaffected, since their audio was synthesized at the time. A video or translation still configured with a deleted glossary keeps working rather than failing.
delete_brand_glossary
Deletes a brand kit, along with the colors, logos and fonts it holds. This cannot be undone: there is no way to restore a deleted brand kit through the API. A kit that is still assembling can be deleted at any time. Videos already generated with the kit are unaffected, since their brand colors were applied at render time. A video agent still configured with a deleted kit will report the id as invalid on its next use. A brand kit shared into your workspace by another workspace can be read but not deleted, and returns 403.
delete_brand_kit
Permanently deletes a lipsync job and its associated files. This action cannot be undone.
delete_lipsync
Permanently deletes a video and its associated files. This action cannot be undone.
delete_video
Permanently deletes a video translation and its associated files. This action cannot be undone.
delete_video_translation
Deletes a voice clone owned by the caller. The voice must not be in use by any template. The voice is removed from your voice list and no longer counts against your voice clone limit. Deleting an already-deleted or unknown voice returns 404 `voice_not_found` (not 200) — a delete-then-list flow should treat that 404 as success, not an error.
delete_voice
Deletes the caller-owned model-backed audio voice. The voice is removed from subsequent reads, but a voice cannot be deleted while its status is `PENDING`. Wait for training to finish and retry. Repeating a successful deletion returns the same successful response.
delete_model_audio_voice
Returns up to 3 voices matching a natural language description (e.g. 'warm, confident female narrator'). Use the seed parameter to get different batches of results.
design_voice
Delete the photo-avatar group created by an avatar flow the user chose to exit.
internal_avatar_discard_photo_group
Start an advisory quality evaluation of uploaded footage and return the workflow id its findings are read back with.
internal_avatar_submit_media_evaluation
Expire a phone capture credential that is no longer displayed by the widget.
internal_avatar_expire_handoff
Synthesize speech audio from text using a specified voice. The voice must support the starfish engine — use GET /v3/voices?engine=starfish to find compatible voices. Supports plain text and SSML. Speed range: 0.5–2.0x. Returns a URL to the generated audio file along with duration and optional word-level timestamps.
create_speech
Generates a video from the template by replacing its variables (text, image, video, audio, character, voice). Use scene_ids to select, reorder, or repeat scenes — scenes must already exist in the template; the API cannot create new ones. Returns the created video object; poll GET /v3/videos/{video_id} or use webhooks for completion. Idempotent replays return the original creation-time snapshot (status and URLs as of the first request), not the video's current state.
generate_from_template
How do I improve a ChatGPT Plugin's discoverability?
The levers are the listing surface agents actually read: names, descriptions, keywords, tool metadata, and registry health. Which lever matters depends on where discovery breaks, which is what continuous measurement shows.
What are HeyGen alternatives on ChatGPT?
As of 2026-09-28, HeyGen competes with AdKraft, AI Video Maker, Arcade, Camtasia, Clueso, Glinded for Birthday Videos, Glinded for Memorial Videos, Hypernatural, Incarn, Instavar Remotion Templates, invideo, Krikey AI Animation, Malloy Studio, Martini, Motionvid, Runway, Screel, Sequencer, Slipa, sync. labs, Synthesia, TalkGen, VEED Video Generator, VideoGen, Videomagic, VideoZero, Viewmax, Visla Video Maker in ChatGPT AI Video Generation, ranked by public Discoverability Score.
Where is this profile measured?
This profile uses the geography attached to the latest public registry snapshot: US. Locale tags are intentionally omitted.