UniMate: one model for text-driven animation across diverse skeletons
A text-to-animation model for researchers that handles many skeletons without separate training for each one.
GitHub Friedrich-M/UniMate Updated 2026-10-02 Branch main Stars 1.1K Forks 101
Python Text-to-animation UniML3D Conda

🧭 Decision Guide

Try it if you

  • You need text-driven motion for diverse skeletons using Python 3.10 and UniML3D.
    Environment Setup requires python=3.10; the Dataset section describes UniML3D as covering diverse skeletal topologies.
  • You want to run inference directly with the UniMate preview checkpoints on HuggingFace.
    The News section says HuggingFace preview checkpoints were released on 2026-09-27, and the Inference section provides a sampling command.
  • You need interpolation of existing motion or regeneration while preserving selected joints.
    The Applications section provides Motion in-betweening and Text-guided motion editing with run_sample_motion_inbetween.sh and run_sample_motion_edit.sh.

Skip it if you

  • You require reliable support for every new out-of-distribution rig.
    The README says many motions and skeletons still fail, and the official preprocessing pipeline for new OOD rigs is still TODO.
  • Your data pipeline depends on the Truebones animal motion pack without purchasing it.
    The Dataset & Data Processing section says the Truebones ZOO animal motions are commercial assets that cannot be redistributed and must be purchased directly.
  • You cannot provide the dataset/features directory used during training when running a checkpoint.
    The Inference section explicitly requires the training run's dataset/features// directory and its target-skeleton conditioning data.

Requirements

  • The Conda environment is named unimate and uses Python 3.10.
  • Run pip install "setuptools<81" and pip install -r requirements.txt --no-build-isolation.
  • Inference input comes from a training-output directory containing config.json, dataset_stats.npy, and checkpoints/.
  • The dataset/features// directory for the target skeleton must exist.
  • The mesh animation pipeline uses stage 5 and can output GLB and FBX.

First step (verbatim from README)

conda create -n unimate python=3.10 -y

Watch out

  • Without test_cases_json, sampling enumerates dataset samples from eval or train.
    The Inference section says that without --test_cases_json, it enumerates the eval split, falling back to unique (object_type, caption) pairs from train.
  • Unconditional sampling with test_cases_txt requires --cfg_scale 1.0.
    The Inference section explicitly states that unconditional sampling requires --cfg_scale 1.0.
  • Use --only_save_motion to save only NPY features instead of MP4 renders.
    The flags table says --only_save_motion skips MP4 rendering and writes only .npy features.

Not stated in the README

  • The README does not specify a minimum GPU model, VRAM requirement, or whether complete inference supports CPU.
  • The README does not provide quantitative comparisons of UniMate against other models for quality, speed, or failure rate.
  • The README does not state when the fully processed 13,006-sequence UniML3D dataset will be openly released.
  • The README does not specify the model size or training configuration of the HuggingFace preview checkpoints.
  • The README does not explain the result or performance differences between the Interactive Demo and local checkpoints.

💡 Deep Analysis

6
No, not before a separate license review, because the code license and data-asset licenses are explicitly different. I can support Python 3.10 and GPU deployment, and I am comfortable with the MIT-licensed code. However, the product will be commercial and may redistribute generated results. Can I include UniMate and its data assets as a whole?
For: An engineering lead planning to use UniMate in a commercial animation product, relying on its MIT-licensed code, Truebones motions, and Hugging Face data while needing clear redistribution and commercial-use boundaries

No, not as a complete package without a separate commercial license review. The MIT status applies to the project code and does not automatically cover training data, motion packs, or third-party assets.

  • The project metadata lists an MIT License, but the README explicitly says that Truebones ZOO motions are a commercial asset pack whose license does not permit redistribution.
  • Truebones motions must be purchased directly from Truebones and supplied in the original Truebone_Z-OO folder layout.
  • UniML3D draws from Mixamo, Objaverse-XL Rigged Animated, and Truebones, and the README does not declare one common commercial license for every source asset.
  • Code use, commercial use of training data, and redistribution of generated results are therefore separate licensing questions; the README does not provide a blanket commercial grant.
  • Project data: license — MIT License
  • Dataset & Data Processing: Truebones ZOO animal motions themselves are a commercial asset pack whose license does not permit redistribution
  • Dataset & Data Processing: please purchase the pack directly from Truebones
  • README: raw source assets are available under Mixamo-Animations-Characters, Objaverse-XL-Rigged-Animated, and Truebones-ZOO-Annotations
Not stated in the README:The README does not specify the commercial-use and redistribution boundaries for Mixamo and Objaverse-XL Rigged Animated.;It does not state whether generated NPY, GLB, FBX, and MP4 outputs inherit restrictions from source assets.
Yes, the README provides the required controls for batching, reproducibility, memory usage, and checkpoint selection. I need to generate multiple text-conditioned motions on a GPU, keep three random samples per test case, and control the seed, batch size, and checkpoint explicitly. Is UniMate’s inference interface sufficient?
For: A Python engineer maintaining GPU inference scripts who needs to generate many text-conditioned motion samples while controlling reproducibility, memory, and checkpoints with seed, batch_size, and checkpoint options

Yes, because the command-line interface directly covers the batching and experiment-control requirements you listed.

  • --num_repetitions controls how many samples are generated per test case, so it can be set to 3.
  • --seed fixes the sampling noise, while --model_path selects a specific checkpoint instead of the latest step.
  • --batch_size controls chunked inference and caps GPU memory; the README explicitly says it does not remove the total computation cost across all cases.
  • The run stores one .npy per repetition, optional MP4 renders, the conditioning T-pose, and captions.json, allowing prompts and outputs to be matched.
  • Inference flags: --num_repetitions, --batch_size, --seed, and --model_path
  • Inference: batch_size — Per-chunk inference batch; caps GPU memory regardless of how many cases there are
  • Each run writes motions/*.npy, animations/*.mp4, _tpos.png, and captions.json
python -m unimate.inference.sample \
    --exp_dir outputs/uniml3d_60frames_graph_adaln \
    --test_cases_json test_cases.json \
    --num_repetitions 3
Not stated in the README:The README does not provide a memory table by GPU model, sequence length, and batch size.;It does not state whether the same seed is fully deterministic across different CUDA, PyTorch, or checkpoint versions.
No, not for direct use, because the official preprocessing pipeline for new OOD rigs is unfinished and the project explicitly says many skeletons still fail. My target asset is a new robot or fantasy-creature skeleton outside the datasets. I only have a rigged mesh and T-pose and want to generate motion directly from natural language. Is the current version suitable?
For: A technical artist who wants to use UniMate with a new OOD robot or fantasy-creature rig outside the existing datasets, in a Python research environment rather than a standalone graphical tool

No, not for direct use. The main obstacle is not the text prompt; the new skeleton must first be converted into the project’s normalized feature representation.

  • The README TODO says the official preprocessing pipeline for new out-of-distribution rigs is still to be released.
  • Inference requires the training run’s dataset/features// directory because the T-pose and topology conditioning are loaded from it.
  • object_type must exist in the dataset; adding a new robot or fantasy-creature name to the JSON does not by itself provide support.
  • The project also states that many motions and skeletons still fail, so the current version is better suited to researchers familiar with the data pipeline than to a plug-and-play asset tool.
  • Top-level TODO: We will release an official preprocessing pipeline for new (out-of-distribution) rigs
  • Inference: The target skeleton — T-pose and topology conditioning — is taken from the dataset
  • Inference: the dataset/features// directory the model was trained on must be present
  • Test cases: object_type must exist in the dataset
  • Note: many motions and skeletons still fail
conda create -n unimate python=3.10 -y
Not stated in the README:The README does not state when the OOD-rig preprocessing pipeline will be released, what its input format will be, or whether arbitrary joint naming is supported.;It does not provide support details for robot joint limits, physical constraints, foot contacts, or collisions.
Yes, but it is better suited to research validation and prototyping than to a production animation system. I already have Mixamo, Objaverse-XL Rigged Animated, and Truebones assets, and I can use Python 3.10, Conda, and GPU inference. Is UniMate suitable for validating a text-to-motion approach without retraining for every skeleton?
For: A researcher validating cross-skeleton motion generation with Mixamo, Objaverse-XL Rigged Animated, and Truebones assets in a Python 3.10, Conda, and GPU environment, without training a separate model for every skeleton

Yes, because the project explicitly targets diverse skeleton topologies and aims to avoid per-skeleton retraining or test-time optimization.

  • UniML3D contains 13,006 text-paired motion sequences covering bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects.
  • Inference conditions on text, the target T-pose, and topology; the target skeleton is loaded from the training feature directory.
  • The project releases training and inference code together with preview checkpoints on Hugging Face, which supports reproducible experiments.
  • However, the README calls UniMate an “early step” and states that many motions and skeletons still fail, so cross-skeleton support should not be treated as a stable universal guarantee.
  • Dataset & Data Processing: 13,006 text-paired motion sequences covering bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects
  • Inference: no per-skeleton retraining and no test-time optimization
  • News: The training and inference code is released; Preview checkpoints are released
  • Note: UniMate is an early step toward text-to-animation for any skeleton, and many motions and skeletons still fail
conda create -n unimate python=3.10 -y
Not stated in the README:The README does not provide quantitative success rates or motion-quality benchmarks for each skeleton category.;It does not specify supported limits for joint count, topology complexity, or inference latency on new skeletons.
It depends: the project can export GLB and FBX, but it does not guarantee that arbitrary rigged assets will work without preprocessing. I use rigged assets from Objaverse-XL Rigged Animated, and the deliverable must be GLB and FBX rather than only NPY features. Can UniMate directly satisfy my mesh-driving workflow?
For: A technical artist integrating generated motion into existing GLB/FBX assets from Objaverse-XL Rigged Animated, with a team expecting animation files that can drive the original meshes

It depends, because UniMate supports a GLB/FBX export path, but the sampling stage first produces motion features rather than a mesh animation already bound to the asset.

  • Inference writes motions/*.npy files containing motion features shaped (T, J, 12).
  • The README says those files must be passed to stage 5 of the data pipeline to export an animated GLB and FBX.
  • The default rig is read from dataset/features//cond.npy, so the asset’s T-pose, topology, joint count, and joint ordering must match the conditioning skeleton.
  • The official preprocessing pipeline for new out-of-distribution rigs is still marked TODO, and the README does not promise direct compatibility with arbitrary third-party rigs.
  • Inference: Each run writes motions/-rep_-.npy; Generated motion features (T, J, 12)
  • Driving a rigged mesh: hand the .npy files to stage 5 of the data pipeline, which exports an animated GLB + FBX
  • Driving a rigged mesh: reads the rig from dataset/features//cond.npy by default
  • Top-level TODO: release an official preprocessing pipeline for new (out-of-distribution) rigs
bash scripts/run_animate_motion.sh objaverse \
    outputs/uniml3d_60frames_graph_adaln/samples/motions/Dog-walk-rep_0-0.npy \
    outputs/animated
Not stated in the README:The README does not define binding conventions, joint-name conversion rules, or automatic retargeting support for FBX/GLB assets outside Objaverse.;It does not document recovery or asset-repair procedures when export fails.
Yes, the README explicitly presents in-betweening and text-guided local editing as applications of the same model without extra training. I can already run UniMate’s released checkpoint and want to fix endpoint frames or selected joints while editing the rest with new text, without additional training. Does the project support this?
For: A generative-model developer researching motion in-betweening and local editing with the released checkpoint, without wanting additional training for either task

Yes, because the project implements these tasks as replacement-style sampling rather than training a separate editing model.

  • The README states that the same trained model supports motion in-betweening, text-guided motion editing, and motion expansion.
  • During sampling, part of a known motion signal is pinned while the flow ODE denoises only the remaining part, so the constraint is maintained at every step.
  • In-betweening can fix endpoints or selected keyframes; local editing fixes the original motion of selected joints and generates the rest according to new text.
  • This still depends on correct skeleton features and joint identifiers; the supplied README excerpt does not include the complete option names or a complex multi-keyframe example.
  • Applications: The same trained model does three more tasks with no extra training
  • Applications: replacement-style sampling; the constraint holds exactly rather than being encouraged by a loss
  • Project insight: supports motion in-betweening, text-guided motion editing, and motion expansion
Not stated in the README:The supplied README excerpt does not include the complete commands, argument formats, or input-file layouts for in-betweening and editing.;It does not explain how quality or editing freedom changes when many keyframes or joints are fixed simultaneously.

✨ Highlights

  • UniML3D contains 13,006 text-paired motion sequences
  • UniMate requires neither per-skeleton retraining nor test-time optimization
  • Supports bipedal, quadrupedal, avian, and serpentine skeletons
  • Provides HuggingFace preview checkpoints and inference code

🔧 Engineering

  • Use python -m unimate.inference.sample to generate skeletal motion from text
  • Use run_animate_motion.sh to export animated GLB and FBX files
  • Supports motion in-betweening and text-guided joint editing

⚠️ Risks

  • The README explicitly states that many motions and skeletons still fail
  • The Truebones animal motion pack cannot be redistributed and must be purchased separately
  • The preprocessing pipeline for new OOD rigs is still marked TODO
  • Inference requires the dataset/features directory used by training to be present

👥 For who?

  • Animation researchers and developers using Python 3.10 and Conda
  • Teams working with Mixamo, Objaverse, or Truebones assets
  • Researchers experimenting with text animation across diverse skeleton topologies