MCP.so
Sign In

Picsart Genai MCP

@PicsArt

About Picsart Genai MCP

**Picsart MCP — 150+ AI Models for Images, Video & Audio, One Connection**

Config

Add this server to your MCP-compatible client using the configuration below.

{
  "mcpServers": {
    "picsart-gen-ai": {
      "type": "http",
      "url": "https://api.picsart.com/gen-ai/mcp"
    }
  }
}

Tools

34

Single entry point for the authenticated user's Picsart Drive. Pass `action`: - `list`: browse a folder. `folderUid` omitted = Drive root; set it to descend. Folders are returned first; files are paginated (`page`, `pageSize`<=128, optional `sort`, optional `type` filter). Set `flat: true` to list every file across all folders (folders are omitted in flat mode). - `create_folder`: requires `name`; `folderUid` = parent (omit for root). Optional `description`. - `upload`: save a file. Provide EITHER `file` (chat attachment) OR `url`+`name` (an HTTPS URL or an inline `data:` URI — data URIs are pushed to the Picsart CDN first). `folderUid` = destination, `type` = resource kind. `result.url` is the CDN-hosted URL of the saved file, ready to pass to `picsart_generate` reference params like `imageUrls`. - `move`: requires `itemUids`; `targetFolderUid` = destination (omit = root). - `delete`: requires `itemUids`; soft-deletes to trash unless `permanent` is true. - `update`: requires `itemUid` + `attributes`; sets custom key/value attributes on a file (e.g. `{ coverUrl }`). Every action except `upload` returns the current folder listing (folders, files, page math) so the widget can render; `upload` skips the re-list (the destination was already explicit) and returns only `result`, so the browser panel does not reopen as a side effect of a plain save. Requires an authenticated Picsart session (per-user Drive content): the caller's identity is forwarded to the Drive service as a `user-id` header — no bearer token is sent onward.

Removes the background from an image, returning a transparent cutout of the foreground subject. Auto-picks the newest enabled Picsart remove-bg model unless overridden via the `model` param — no need to call `picsart_list_models` first. Use this when the user asks to "remove the background", "cut out the subject", or "make the background transparent". Do NOT use this to replace the background with a new scene (use `picsart_change_bg`), upscale or sharpen the result (use `picsart_enhance`), convert raster to SVG (use `picsart_vectorize`), or generate a new image from scratch (use `picsart_generate`). Required input: `image` — a publicly-accessible URL. Local files are not supported; if you only have a local file, first make it available as a public or app-authorized URL. Optional: `model` to pin a specific remove-bg model, `outputFormat` (e.g. "png"). Example: `{ image: "https://example.com/portrait.jpg" }`. Returns `{ assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? }` as a single JSON text block plus matching structuredContent (no `resource_link` blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). `id` is the SDK's generation handle; `metadata` may include model-specific tags. Spends credits. Requires Authorization: Bearer <picsart_token>.

Replaces the background of an image with a new scene described by a prompt, keeping the foreground subject intact. Auto-picks the newest enabled Picsart change-bg model unless overridden via the `model` param — no need to call `picsart_list_models` first. Use this when the user wants to "change the background to X", "put this on a beach", "swap the background for a marble counter", or any compositing where the subject is kept and the backdrop changes. Do NOT use this to strip the background to transparency (use `picsart_remove_bg`), upscale or sharpen (use `picsart_enhance`), convert raster to SVG (use `picsart_vectorize`), or generate a brand-new image from scratch (use `picsart_generate`). Required inputs: `image` — a publicly-accessible URL, not a local file path — and `prompt` describing the new background. Optional: `model` to pin a specific change-bg model; preflight the explicit model id you plan to use (the default path may select `recraftv3-replace-bg` rather than the legacy `picsart-change-bg`). Example: `{ image: "https://example.com/product.jpg", prompt: "polished marble countertop with soft window light" }`. Returns `{ assets, id, model, created_at, prompt, summary, why_relevant, url, results: [{ url, metadata? }], drive? }` as a single JSON text block plus matching structuredContent (no `resource_link` blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). `id` is the SDK's generation handle; `metadata` may include model-specific tags (e.g. `exploreImageId` for Recraft Explore models). Spends credits. Requires Authorization: Bearer <picsart_token>.

Upscales and enhances an image — sharpens edges, denoises, and raises resolution by an optional scale factor. Auto-picks the newest enabled Picsart upscale / enhance model unless overridden via the `model` param. Use this when the user asks to "upscale", "enhance", "make it higher resolution", "sharpen", "clean up this photo", or "make this 4k". Do NOT use this to remove the background (use `picsart_remove_bg`), replace the background (use `picsart_change_bg`), convert raster to SVG (use `picsart_vectorize`), or generate a new image (use `picsart_generate`). Required input: `image` — a publicly-accessible URL, not a local file path. Optional: `model` to pin a specific enhance model, `scaleFactor` (e.g. 2 or 4) for upscale ratio. Example: `{ image: "https://example.com/photo.jpg", scaleFactor: 4 }`. Returns `{ assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? }` as a single JSON text block plus matching structuredContent (no `resource_link` blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). `id` is the SDK's generation handle; `metadata` may include model-specific tags. Spends credits. Requires Authorization: Bearer <picsart_token>.

Converts a raster image (PNG, JPG) into an SVG vector. Auto-picks the newest enabled Picsart vectorize model unless overridden via the `model` param. Use this when the user asks to "vectorize", "convert to SVG", "make this a vector", or wants a scalable version of a logo or icon. Best results on logos, icons, and simple graphics — photographic images vectorize poorly and the user should be warned. Do NOT use this to remove the background (use `picsart_remove_bg`), replace the background (use `picsart_change_bg`), upscale a raster image (use `picsart_enhance`), or generate a new image (use `picsart_generate`). Required input: `image` — a publicly-accessible URL to a PNG or JPG (not a local file path). Optional: `model` to pin a specific vectorize model. Example: `{ image: "https://example.com/logo.png" }`. Returns `{ assets, id, model, created_at, summary, why_relevant, url, results: [{ url, metadata? }], drive? }` as a single JSON text block plus matching structuredContent (no `resource_link` block for the SVG URL — the widget is the single source of visual truth, so it is not duplicated as a separate content block). `id` is the SDK's generation handle. Clients fetch the SVG from that URL. Spends credits. Requires Authorization: Bearer <picsart_token>.

Runs any Picsart AI model end-to-end to produce an image, video, audio, or text result. Spends credits. If you already have a model id/name in hand, skip straight to `picsart_generate` — no need to call `picsart_list_models` first. Optionally validate first via `picsart_model_params` (learn its inputs) and/or `picsart_preflight` (validate the payload and quote cost before spending credits); `picsart_generate` itself also rejects unsupported param values before charging. Only reach for `picsart_list_models` when you need to pick a model — e.g. no model was named, or the user wants to browse/compare visually via the model-picker widget. To browse or check model capabilities (e.g. supported aspect ratios) programmatically WITHOUT popping that widget, use `picsart_model_catalog` instead. Do NOT use this for editing operations that have dedicated tools — background removal (`picsart_remove_bg`), background replacement (`picsart_change_bg`), upscale / enhancement (`picsart_enhance`), or raster-to-SVG conversion (`picsart_vectorize`). Also do NOT use it to validate params, quote cost, or browse the catalog — those are separate tools above. Required inputs: `model` (id) and `prompt`. Model-dependent optional inputs: `duration` (video seconds), `aspectRatio` (e.g. "16:9", "9:16", "1:1"), `resolution` (e.g. "1080p", "4k"), `count` (1–10 outputs), `quality`, `style`, `negativePrompt`, `imageUrls` (for image-to-X models), `videoUrl` (for video-to-X), `enhancePrompt`, `generateAudio`, and `extra` — a free-form record for model-specific params (discover them via `picsart_model_params`). Example (image): `{ model: "flux-2-pro", prompt: "a cat in a hat", aspectRatio: "1:1", count: 1 }`. Example (video): `{ model: "kling-v3-pro", prompt: "a cat skiing down a mountain", duration: 5, aspectRatio: "16:9" }`. Returns `{ assets, id, model, created_at, prompt, summary, why_relevant, url, results: [{ url, metadata? }], drive? }` as a single JSON text block plus matching structuredContent (no `resource_link` blocks — the widget is the single source of visual truth, so result URLs are not duplicated as separate content blocks). `id` is the SDK's generation handle; `metadata` may include model-specific tags (e.g. `exploreImageId` for Recraft Explore models). Text/LLM models (mode "text" in the catalog — e.g. gemini-3-pro, gpt-5.5, claude-*) run synchronously (`async` is ignored) and return the generated text as the text content block plus `text` in structured content. VIDEO models default to async: the call returns `{ job, status: "ACCEPTED" }` immediately — poll `picsart_job_status` with the job handle until it completes (it then returns this same media payload). Never re-submit a video generation because a call seemed to hang or the host reported a timeout: the render is still running and already charged — poll instead. Pass `async: false` only for a video call you know finishes inside the host's window. ChatGPT renders images and videos with the Picsart media gallery UI; clients fetch the assets from URLs, never base64. Spends credits and writes to the user's Picsart Drive when the Drive option is enabled. Requires Authorization: Bearer <picsart_token>.

Checks a generation job started by `picsart_generate` with `async: true`. Widget-facing: widgets poll this every few seconds with the returned job handle; assistants normally call `picsart_generate` synchronously and never need this tool. While running it returns `{ status: "ACCEPTED"|"IN_PROGRESS", progress?: { percent, estimatedSecondsLeft } }`. Once finished it returns the same media payload `picsart_generate` would have returned (`{ status: "COMPLETED", assets, results, url, ... }`), or an error for FAILED/CANCELED jobs. Requires Authorization: Bearer <picsart_token>.

Render a cheap multi-frame OVERVIEW of a scene as low-res jpeg thumbnails YOU CAN ACTUALLY SEE: each sampled frame comes back as an inline image content block (labeled with its scene time), alongside the machine-readable `{ frames: [{time, url, inline}] }`. Two modes: pass `times` (PREFERRED — you usually know the interesting moments: clip seams, animation midpoints, entrance ends) to get an EXACT thumbnail per requested time, rendered concurrently; or pass `frames` (default 8, max 24) for evenly-spaced sampling across the whole timeline (one image-sequence render at 1fps — integer-second granularity only). Sits between a single full-res frame check and picsart_media_export (full encode): use it to eyeball pacing, seams, and content presence across the WHOLE timeline before exporting, instead of checking single frames repeatedly. Auth is handled by the platform automatically — no token setup needed on your end. Validate the scene first. Inline images are best-effort under a total size/count budget: a frame that fails to download, is oversized, or falls outside the budget still comes back with its `url` and `inline: false` in the JSON — a partially-inlined sheet is a SUCCESS, not an error, because the render has already happened and been charged. Do not retry this call just to get the missing pictures (that re-charges the render) — open those urls instead. Scope: this samples a scene you have already authored, to look at it. It is NOT a metadata probe — use picsart_media_probe_media first for a source clip's duration/dimensions/frame rate. Sparse-sampling frames here to discover duration is only the fallback when picsart_media_probe_media genuinely cannot determine it (every frame is a charged GPU render). Use picsart_media_query_layout to check layer geometry, overlap and stacking order without rendering. At most 8 frames come back inline per call regardless of how many you request.

Render a fully-authored MP Scene to a final file via the server-side `media-platform/v3/export` workflow, returning the output URL(s). The output container is chosen with `mediaType`: omit it (default) to render an MP4, or pass `mediaType:"png"` for a single still (`mov`/`webm`/`jpeg`/`webp`/`heif`/`gif`/`image-sequence` are also accepted). Unlike a pure scene→scene deck compiler, this performs the ACTUAL render: it translates the scene to a V3/Replay project and dispatches it to the export service over an authenticated HTTP call — no local renderer. Auth is handled by the platform automatically — no token setup needed on your end. Validate the scene (`picsart_media_validate_scene`) and check a frame (`picsart_media_contact_sheet` with `times:[t]` — sample one frame near the point you want to verify; it returns that frame as an inline image you can look at) BEFORE exporting. If you do not yet have a scene, get the whole sequence from picsart_media_quickstart rather than reverse-engineering the document shape from picsart_media_get_scene_schema.

Apply a named effect (gaussian_blur, drop_shadow, stroke, ...) to a visual layer in an MP Scene, with catalog defaults filled for any omitted params. Returns the updated scene with the effect merged into the target layer's effects[] -- re-applying the same effect id REPLACES it (idempotent), so params can be tuned without stacking duplicate effects. Each param is validated by kind and numeric range; an unknown effect id, unknown param, or out-of-range value is rejected with a clear error. Discoverable effect ids and their parameter shapes (ranges, defaults) are listed under `supports.effects` in `picsart_media_get_capabilities`. Pure: takes the full scene by value, returns a new scene; no server-side state.

Apply a composite 'look' (e.g. vintage_bw, light_leak, shimmer) to a media OR scene_ref layer. A look wraps the layer's content into an isolated nested composition so the treatment hits the composited result as one image. TWO MODES: mode:'reference' (DEFAULT) just appends a thin entry to the layer's `looks[]` array ({look, params}) and KEEPS the layer's content as the look's subject — nothing is baked; translate/preview/query expand it just-in-time (resolveLooks), or call picsart_media_resolve_looks to bake a self-contained scene. This keeps the stored scene small and the look non-destructively editable (re-tune params / drop the entry / stack multiple looks). mode:'expand' eagerly bakes the full nested sub-scene now (the original behavior). Params have catalog defaults; omit `params` for the default look. Discoverable looks + params are under `supports.looks` in picsart_media_get_capabilities. NOTE: a layer carrying its own effects/animations/mask/blendMode can't take a look (the look's scene_ref wrapper can't preserve them). Pure: takes the full scene by value, returns a new scene.

Apply a named motion preset (glow_pulse, ken_burns, scale_pop, ...) to a visual layer in an MP Scene. Returns the updated scene with the preset's animations + effects merged into the target layer. Each preset declares an `appliesTo` set of layer content kinds (text/media/color) -- mismatches are rejected with a clear `preset_kind_mismatch` error. Discoverable presets and their parameter shapes are listed under `supports.motionPresets` in `picsart_media_get_capabilities`. Pure: takes the full scene by value, returns a new scene; no server-side state.

Instantiate a scene template with the supplied parameter bindings. Two modes: `reference` (DEFAULT, canonical) packages an `MpSceneRefContent` snippet (kind:'scene_ref' with bindings + optional timeScale/trim/fit) to drop into a parent scene's `layers[]`; at translate/preview time the ref becomes its OWN nested composition (keeps its resolution/duration, fit/transform honoured — the template stays a reusable parameterized unit). For a STANDARD library template (`mpscene://montage`, …) the parent needs NO `scenes` entry; the translator resolves it from the registry. `bootstrap` returns a complete resolved standalone `MpScene` (substitutes every `$param`, synthesizes asset entries for asset-typed parameters, strips the `parameters` declaration) — use it to bake an editable starting scene. `inline` MERGES the template's resolved concrete layers + assets INTO a target `scene` you pass (no scene_ref) and returns it — the 'bring a preset's editable layers into my composition' op (ids prefixed so nothing collides, assetId refs rewired, brought-in layer starts offset by `at`). Use `picsart_media_describe_scene_template` first to learn what parameters the template accepts. Pure: returns either `sceneRef` or `scene` plus the applied parameter map. For the common video jobs — concatenating/merging N clips with per-clip trim and fit/align into one output — call picsart_media_quickstart with recipe:'concat_videos' FIRST. It returns a complete, validated parameters payload for mpscene://montage; you should not need picsart_media_describe_scene_template or picsart_media_get_scene_schema round-trips for that case.

Apply a named text-animation preset (typewriter, fade_in_chars, slide_up_lines, ...) to a text layer in an MP Scene. Returns the updated scene with the animation appended to the layer's `content.animations[]`. Discoverable presets and their parameter shapes are listed under `supports.textAnimationPresets` in `picsart_media_get_capabilities`. Agents wanting custom shapes can write `MpTextAnimation` entries directly into a scene without going through this tool. Pure: takes the full scene by value, returns a new scene; no server-side state.

Describe a single MP Scene template by URI -- returns its declared parameters (with types, defaults, ranges, required flags, descriptions) and composition dimensions. Use this before `picsart_media_apply_scene_template` so you know what to pass. Accepts `mpscene://<id>` URIs in v1; other schemes are deferred. Returns `{ id?, uri, composition, parameters }` where `parameters` is the full declarations map.

DETACH a scene_ref into editable layers: replace a resolvable scene_ref layer (or every resolvable one, if no `layerId`) with the referenced template's concrete resolved layers, inlined IN PLACE (prefixed by the ref layer's id, started at its `start`, assetId refs rewired). The result has no scene_ref for the expanded layers, so you can edit the brought-in layers directly. Inverse of the by-reference model: 'I referenced a template, now bake this instance to customize it beyond its parameters.' One level in v1 (a brought-in layer that is itself a scene_ref stays a reference — run again to go deeper). This is a SYNC op that does NOT fetch: it resolves only in-memory `mpscene://` refs (local `scenes{}` / inherited / registry). A REMOTE `https://`/`http://` ref (e.g. a CDN-hosted scene) is left as-is here — those ARE resolved automatically at translate / render-preview / query-layout time (and by the deck compiler picsart_media_overview when called over MCP, though its bare runner does not self-fetch), so detach after a translate/preview/query has inlined it, or inline the piece locally. Unresolvable refs (unfetched remote URLs / unknown ids) are left as-is in bulk mode; in single-layer mode they error. Pure.

Returns the MP SDK tool layer's capabilities: supported layer content kinds, animatable properties, effect ids, transition ids, and operational limits (max duration, max layers, max resolution). Call this first when planning a composition so subsequent calls stay within supported bounds. Omitting `sections` returns the section INDEX (names only); the full document is ~135KB, so pass `sections` to fetch only the slices you need, or ["all"] for the whole thing — PREFER the granular supports sub-keys over the whole `supports` block (e.g. ["effects","limits"] or ["generativeTemplates","limits"] for a montage plan). `looks` alone is ~62KB (full prose + full param schemas for all 17 looks) and can overflow a result cap even as the ONE section requested — ask for `looksCompact` (id/summary/params only, ~8KB) first to see what exists, then `looks` for the id you actually need. Top-level: supports, limits, engine, featureMatrix; supports sub-keys: effects, transitions, easing, blendModes, animatableProperties, textAnimationPresets, motionPresets, looks, looksCompact, layerContentKinds, sceneRefs, mask, generativeTemplates, presentation, export, templateModes, expressionAnimation. Idempotent and dependency-free; safe to cache.

Returns the JSON Schema for an MP Scene document. Use this to construct valid scenes from scratch or to remind yourself of the exact shape of layers, animations, effects, and transitions before calling picsart_media_validate_scene or checking a frame with picsart_media_contact_sheet. The returned schema is authoritative; any document that validates against it is accepted by the renderer. Pass `name` (e.g. "MpMediaContent") to get back just that one definition instead of the full ~66 KB schema — its internal $refs point into the full schema's #/definitions.

Returns the curated font catalog. Use this before authoring any text layer so `font.family` resolves to a real font: every entry has a stable `key` (passable directly in MpFont.family) and a resolved `.otf`/`.ttf` URL. The renderer has no system-font fallback — passing CSS family names like `Inter`, `Arial`, or `Helvetica` produces empty text and a `unknown_font_family` validation error. Agents may also pass a direct font URL matching `accepted_url_pattern`. The tool is dependency-free and idempotent; safe to cache.

Enumerate the curated MP Scene template catalog. Each entry summarises a reusable, parameterized scene (title cards, lower thirds, product cards, ...). Use this to discover templates by id/category/aspect before describing or applying one. Returns `{ templates: [{ id, displayName, description, category?, width, height, parameterCount }, ...] }`. Pure: no inputs, no state.

Compile a presentation DECK (an MP Scene with one top-level `track` layer of `scene_ref` slide clips) into an OVERVIEW / 'badges' board: a single MP Scene that lays the SAME slide `scene_ref`s into a grid of shrunken thumbnails — same content, a spatial projection. Each cell is the page positioned at its grid-cell centre with `fit:"none"` and a baked contain-scale (the compiler resolves each page to read its real composition size, because `scene_ref.fit` fits against the whole composition, not a cell). Defaults: a near-square grid (`ceil(sqrt(n))` columns), the deck format for board size, and a duration long enough for every slide's entrance to settle. Carries the deck's `scenes`/`assets` so the refs still resolve. Override `cols`/`gap`/`width`/`height`/`duration`. Pure scene→scene; validate / preview / translate the output. NOT a capability index or a getting-started tool despite the name — it compiles thumbnails for an existing deck scene you already authored. If you are looking for 'what can this server do' or 'how do I do X', call picsart_media_quickstart instead.

Apply a batch of incremental, ID-ANCHORED edit ops to an MP Scene and return the patched result — the token-cheap alternative to re-emitting the entire scene on every edit. Three verbs: `set` (create-or-replace a value; last-write-wins, idempotent; `replace` is accepted as a first-class alias), `remove` (delete a field/array element, or an entire anchored entity when no `path` is given), and `add` (insert a new id-keyed layer/asset/audio/marker; requires `kind` and a `value` carrying the new entity's `id`). Address the target with exactly one ANCHOR + id — `layer`, `asset`, `audio`, or `marker` — never a root index path: `{op:"set",path:"layers/3/..."}` is rejected with `use_id_anchor` and a corrected anchored form embedded in the message; omit the anchor only for true document-root fields like `composition/*`, `version`, or `scenes/<id>`. `path` is a slash-separated RFC-6901-style pointer relative to the anchored entity (e.g. `content/text`, `effects/0/params/amount`) — `~0`/`~1` escapes are honored, and dot-separated paths fail strict with a slash-form hint. All ops in one call are ATOMIC: they apply sequentially against a single clone and the first failing op aborts the WHOLE batch with no partial result (a typed error carrying `code` and `opIndex`); the input scene is never mutated, and later ops see earlier ops' results. Validation runs INSIDE this tool: the patched result goes through the exact picsart_media_validate_scene pipeline (JSON-Schema structural pass + semantic rules) and any severity:error fails the whole batch loudly with the diagnostics — a success has already passed that pipeline, so do NOT call picsart_media_validate_scene again right after a successful patch; only validate once you've made further hand-edits outside this tool. (Validation proves conformance to the scene contract, not that every referenced remote asset or scene_ref is fetchable at render time.) Id-less collections — track clips, `transitions`, `animations`, `keyframes`, text `content/animations` — carry no per-element id, so edit them with a whole-array `set` on the owning anchor's array path (e.g. `{op:"set",layer:"main_track",path:"content/layers",value:[...]}`) rather than addressing individual elements by index across calls.

Returns a remote media URL's metadata WITHOUT downloading the file: `kind` (image/video/audio), as-displayed `width`/`height` (EXIF orientation applied, plus the raw `orientation`), `durationSeconds` (mp4/mov/m4a), `contentType`, and total `bytes` — from at most two small ranged fetches (≤256KiB). Use this BEFORE authoring a scene to size compositions, pick `fit`, set clip windows to the real source duration, and reject wrong assets early — instead of shelling out to curl/ffprobe or guessing. Covers PNG/JPEG/GIF/WebP/HEIC/AVIF/SVG/TIFF dimensions and ISO-BMFF (mp4/mov/m4a) dimensions+duration, including moov-at-end files. Unrecognized containers still return `kind`/`contentType`/`bytes` plus a `notes[]` entry — dimensions are never invented. URLs are SSRF-guarded (public http(s) only; every redirect hop re-checked). Marked stochastic only because a URL's content can change between calls — probing an immutable asset is stable. Skip this call if you already know the asset's exact width/height/duration from a prior tool result (e.g. a generation tool's own response) — this is a cost-saving lookup, not a required gate. A missing `durationSeconds` means the duration is unknown — it is never emitted as 0 for a real zero-length clip. Fragmented MP4 / streaming uploads carry no duration in their container header, but this tool now recovers it for most of them by walking the file's fragment timeline, so a duration usually comes back anyway; when the field is still absent it genuinely could not be determined. That pass costs a few more small ranged reads than the two mentioned above — still bounded, still free, still no download. Do not assume 0; resolve it another way before computing trims. The response may also carry `fps`, `videoCodec`, `audioCodec` and `hasAudio` when they are derivable from the container — these are best-effort extras and are omitted rather than guessed.

Resolves where every layer ACTUALLY LANDS in the root composition at a given time — without rendering pixels. Runs the engine's Scene Query System and returns each layer's final geometry in root-composition space: `box {x,y,width,height}` (top-left pixels), `center`, `rotationDeg`, paint-order `zIndex`, `visible`/`active`, an `onCanvas` coverage flag (`full`/`partial`/`off`), `timing`, and composition/parent/nesting. This is the cheap, structured answer to 'what does my composition look like, spatially?' — use it instead of (or before) checking a rendered frame (`picsart_media_contact_sheet`) to verify positions, sizes, overlap, off-canvas layers, and stacking order, since reasoning over numbers beats eyeballing a PNG. Coordinates are ALWAYS top-left pixels in root space for BOTH engines — the query layer unifies Jet's pixel coords and V3's internal normalized coords, so a layer WITH RESOLVED GEOMETRY reports the same box on `jet` or `v3` (the `engine` arg only changes the translation path). Text is the exception (see limitations). Nested children (scene_ref/look comps) are returned too, namespaced (e.g. `a__bg`) with `nestingLevel` and `composition`; inactive layers at the queried time come back marked `active:false` with their box omitted. Set `verbose:true` to also get raw 4x4 transform matrices and the per-component transform chain (for transform-editing tools); omit it for layout reasoning. Determinism is deterministic — identical (scene,time,engine) returns an identical report. Limitations: (1) text layers — the query system does not resolve laid-out glyph bounds, so a text layer's POSITION/center is accurate but its box size is a placeholder (Jet 0x0, V3 ~2000x1); a `notes[]` entry flags any layer with a degenerate dimension. (2) multi-group V3 scenes may need a `sceneContext` (e.g. { canvasGroupId }); single-composition scenes do not. (3) GEOMETRY ONLY — this is NOT a render/export gate. The query is GL-free and never resolves effects, assets, codecs, or engine-specific support, so a clean layout here does NOT guarantee the scene renders or exports. A clean report carries this caveat in its `scope` field — confirm renderability with picsart_media_contact_sheet / picsart_media_export.

Call FIRST when the user asks how to accomplish a media task (merge/concat videos, make a contact sheet, or export/render a scene). Returns ready-to-run picsart_media_* tool-call sequences. Omit `recipe` for the index.

Expand every by-reference look (each layer's `looks[]` annotation) in an MP Scene into its concrete nested composition, returning a SELF-CONTAINED scene — the preset logic baked in, no look-catalog dependency at render time. The explicit, on-demand counterpart of the just-in-time resolution that picsart_media_translate_scene / picsart_media_contact_sheet / picsart_media_query_layout already do internally. Use it to 'flatten' a thin look-annotated scene into a portable one (e.g. to hand off, archive, or edit the expanded layers directly). Idempotent: a scene with no `looks[]` returns unchanged. Pure: takes the full scene by value, returns a new scene.

Translate an MP Scene document into an engine's project format, WRITE it to a content-addressed file, and return `{ path, cached, summary }`. The engine project itself is NOT inlined — a large scene's project can be ~100 KB of JSON and would overflow this tool's result token cap, so it goes to disk (under MP_AI_OUTPUT_DIR/projects/<hash>.<engine>.json) and you get a path plus a compact summary. The target engine is chosen by `engine` (default `v3` → a full, openable Replay file `{ meta, context:{layers,…}, actions, settings }`; `jet` → a Jet `{ compositions, activeCompositionID, ColorSpace }` project). Read the file when you need the full project. Validate the scene first; the translator does not re-validate. Deterministic: identical (scene, engine) hash to the same path (`cached:true` on a repeat).

Opens a drag-and-drop upload widget so the user can get a local image/video/audio file into this conversation as a URL — no filesystem access on your side, and no external CLI needed. Call this whenever a user wants to use a local file with any picsart_media_* tool, or after a local-path argument was refused with a `local_source_not_supported` error. Call with no arguments to just open it. Optional `purpose` labels the dropzone; `accept` restricts file kind; `detected_files` is a BEST-EFFORT, MODEL-SUPPLIED label hint ONLY — e.g. filenames you can see attached in this conversation but have no other way to reach — the widget shows it purely as a suggestion ("Drop clip.mp4 here") and NEVER filters or rejects what the user actually drops, since this hint can be wrong or hallucinated; `known_urls` seeds pickable chips in a "Use existing" tab from URLs already produced earlier in this conversation. The upload itself happens directly in the user's browser to Picsart's CDN — no bytes pass through this tool. Once the user finishes (or picks existing URLs), the widget reports the resulting URL(s) back into the conversation. This is a two-turn handshake: this call only OPENS the widget and returns no URL(s) itself — they arrive on the user's NEXT message, not in this call's result.

Validates an MP Scene document and returns a list of structured diagnostics. Use this whenever you've assembled or modified a scene, especially before rendering. Each diagnostic carries a JSON path, a stable machine-readable `code`, a human-readable `message`, and optional `hints`. A scene is considered valid when no diagnostic has severity `error`. The tool is pure: it takes the full scene in and returns diagnostics; if you keep editing, call again with the updated scene.

Lists Picsart AI models across ALL modes (image / video / audio / text) and renders the Picsart Studio model-picker widget so the USER can browse, compare, and pick a model visually. Each item carries `id`, `name`, `mode`, `inputType`, `supportedAspectRatios`/`supportedResolutions` (when the model declares an enum for that param) (and `provider`, `badges`, `description` when `verbose` is true). Use this when the user wants to SEE the available models or pick one themselves — especially when they have not committed to an output mode yet, or for cross-mode searches ("all flux models", "every model with image input"). To narrow to one output mode without a separate tool, pass the `mode` filter (image/video/audio/text) on this same tool. Ratio/resolution constraints ride along in `supportedAspectRatios`/`supportedResolutions`, so you rarely need `picsart_model_params` just to check whether a model supports a given aspect ratio or resolution. Do NOT use it to fetch a single model's FULL parameter schema (use `picsart_model_params`) or estimate per-call cost (use `picsart_preflight`). If you only need catalog knowledge for your own reasoning (no UI shown to the user), use `picsart_model_catalog` instead. Inputs (all optional): `mode` (filter to image/video/audio/text — text = LLM models that return generated text), `provider` (case-insensitive substring like "flux", "kling", "google"), `acceptsImage` (true → only models that take an image input — i2i, i2v, i2t), `acceptsVideo` (true → only models that take a video input — v2v, v2a, v2t), `acceptsAudio` (true → only models that take an audio input — a2v, sts), `inputType` (exact-match escape hatch; one of t2v/i2v/v2v/a2v/t2i/i2i/t2a/v2a/tts/sts/sfx/music/t2t/i2t/v2t), `limit` (1–100, default 20), `verbose` (default false; when true each item adds provider/badges/description). inputType codes — first letter is input modality, second is output: t2i (text→image), i2i (image→image), t2v (text→video), i2v (image→video), v2v (video→video), a2v (audio→video), t2a (text→audio), v2a (video→audio), tts (text-to-speech), sts (speech-to-speech), sfx (sound effects), music (music gen), t2t/i2t/v2t (LLM text output from text/image/video input). Example: `{ mode: "video", acceptsImage: true, limit: 10 }` returns image-to-video models. Returns `{ items, total, truncated }` — `truncated` is true when more matched than were returned; refine filters or raise `limit` (max 100) to see more. Read-only; spends no credits and works without authentication.

Returns the Picsart AI model catalog as plain data — renders NO widget or UI. Use this when YOU (the assistant) need catalog knowledge for your own reasoning: picking a model before `picsart_generate`, answering "which models support X", or comparing options — without pushing a model-picker widget into the conversation. When the user wants to SEE or browse models visually, use `picsart_list_models` instead (it renders the Picsart Studio picker). Same filters and result shape as `picsart_list_models`, but every item is rich by default: `id`, `name`, `mode`, `inputType`, `provider`, `badges`, `description`, plus `supportedAspectRatios`/`supportedResolutions` when the model declares an enum for that param — enough to answer "which models support 16:9" without `picsart_model_params`. Do NOT use it to fetch a single model's FULL parameter schema (use `picsart_model_params`) or estimate per-call cost (use `picsart_preflight`). Inputs (all optional): `mode` (filter to image/video/audio/text — text = LLM models that return generated text), `provider` (case-insensitive substring like "flux", "kling", "google"), `acceptsImage` (true → only models that take an image input — i2i, i2v, i2t), `acceptsVideo` (true → only models that take a video input — v2v, v2a, v2t), `acceptsAudio` (true → only models that take an audio input — a2v, sts), `inputType` (exact-match escape hatch; one of t2v/i2v/v2v/a2v/t2i/i2i/t2a/v2a/tts/sts/sfx/music/t2t/i2t/v2t), `limit` (1–100, default 20), `concise` (default false; when true items carry only id/name/mode/inputType plus the ratio/resolution fields, to save tokens). inputType codes — first letter is input modality, second is output: t2i (text→image), i2i (image→image), t2v (text→video), i2v (image→video), v2v (video→video), a2v (audio→video), t2a (text→audio), v2a (video→audio), tts (text-to-speech), sts (speech-to-speech), sfx (sound effects), music (music gen), t2t/i2t/v2t (LLM text output from text/image/video input). Example: `{ mode: "audio", inputType: "music" }` returns music-generation models. Returns `{ items, total, truncated }` — `truncated` is true when more matched than were returned; refine filters or raise `limit` (max 100) to see more. Read-only; spends no credits and works without authentication.

Returns the parameter schema for a specific Picsart model — a map of param name to descriptor ({ type, required, default, enum, min, max, step, label, accept }). Use this once you have a model id and need to construct the params payload for `picsart_generate` or feed `picsart_preflight` with a candidate object. Do NOT use it to discover which models exist (use `picsart_list_models`) or estimate cost (use `picsart_preflight`). Required input: `model` id. Example: `{ model: "flux-2-pro" }`. Returns `{ model, schema: { <paramName>: { type: "string"|"number"|"boolean"|"file", required?, default?, enum?, min?, max?, step?, label?, accept? } } }`. Read-only; spends no credits and works without authentication.

Free pre-flight check before `picsart_generate`: in ONE call it (1) validates a candidate params object against the model's parameter schema + inter-parameter constraints, and (2) quotes the credit cost — without running the model or charging the user. Use this after assembling params (user input, derived defaults, model swaps) and before generating, to surface bad arguments and show cost. Do NOT use it to look up which params a model accepts (use `picsart_model_params`) or to actually generate (use `picsart_generate`). Required inputs: `model` id and a `params` object (put the `prompt` inside `params`). Example: `{ model: "flux-2-pro", params: { prompt: "a cat in a hat", aspectRatio: "1:1", count: 1 } }`. Returns `{ model, valid, errors?, credits }`: `valid`/`errors` are from local validation (always present, no auth needed; `errors` only when invalid); `credits` is the dry-run cost (a number), or `null` when pricing is unavailable or the request is unauthenticated.

Returns the current Picsart credit balance for the authenticated user — `balance` (active credits available now) plus the breakdown into `resettable` (recurring monthly/period quota) and `accumulative` (top-ups and add-ons), `total` (active credits across both pools), and `overdraftUsage` (credits spent past the balance, if any). When the resettable pool has a scheduled reset, `nextResetDate` is the ISO timestamp of the next refill. Use this before expensive operations to warn the user when the balance is low, or after a 402 from `picsart_generate` to confirm the issue is credits and not something else. Do NOT use it to estimate the cost of a specific generation (use `picsart_preflight`); this tool only reports the balance, not per-call cost. Takes no input. Returns `{ balance, total, resettable, accumulative, overdraftUsage, nextResetDate? }` where each number is non-negative. Requires Authorization: Bearer <picsart_token> (per-user account data).

Overview

What is Picsart Genai MCP?

Picsart Genai MCP is the backend server invoked by the picsart-api skill from the gen-ai-skills repository. It exposes Picsart’s image, video, GenAI, and variable‑data REST APIs for use by any agent that reads SKILL.md files, including Claude Code, Cursor, and Codex. The skill complements Picsart’s gen-ai CLI; the MCP server provides programmatic access for agent‑driven workflows.

How to use Picsart Genai MCP?

Install the picsart-api skill from the gen-ai-skills repository with npx skills add PicsArt/gen-ai-skills. Agents that support the skill‑loading protocol can then call Picsart’s REST APIs through the Picsart Genai

Comments

More Media & Design MCP servers