System requirements
What your machine needs to run MiniMax Converter — the minimum that gets you going, and what we recommend for the smoothest experience.
Minimum
runs the app- OS Windows 10 (1809+), macOS 11 Big Sur, or any current Linux distribution (Ubuntu 22.04+, Debian 12+, Fedora, Arch, etc.)
- CPU 64-bit dual-core processor (x86_64 or Apple Silicon ARM64) with AVX support — most CPUs from 2012 onwards qualify
- RAM 4 GB
- Disk 2 GB free for the app and the light AI models. The big voice-cloning tiers need 3–12 GB each — downloaded only if you choose them (see the table below).
- GPU Not required for most features — audio, documents, images and transcription all run on the CPU. Only the top voice-cloning tier (High) strictly needs an NVIDIA GPU or Apple Silicon.
- Display 1280 × 720
- Network Only for the first download and to fetch optional AI models the first time you use them. After that, conversions run fully offline.
Recommended
smooth & fast- OS Windows 11, macOS 14 Sonoma or later, or a current Linux distribution (Ubuntu 24.04, Fedora 40+, etc.)
- CPU Quad-core or better. 8+ cores really shines for parallel audio batch conversion (one file per core).
- RAM 8 GB. Bump to 16 GB if you transcribe long videos, upscale 4K photos, or run several conversions in parallel.
- Disk 10 GB free for comfortable everyday use — or plan around 35 GB if you want every AI voice tier installed at once.
- GPU NVIDIA GPU with 8 GB+ VRAM (CUDA), or Apple Silicon — accelerates AI voice cloning, transcription, upscaling, background removal and video encoding. Any Vulkan-capable GPU already speeds up upscaling.
- Display 1440 × 900 or larger
- Network Same as minimum — only used for the initial download and first-time AI model fetches.
AI features — download sizes & hardware
Every AI feature is an optional one-time download, fetched the first time you use it — install only what you need. Sizes below are what each feature adds on top of the base app.
| Feature | Languages | Download | On disk | Hardware |
|---|---|---|---|---|
| Speech & audio | ||||
| Transcription, subtitles, Live Transcribe & dictation (Whisper) | ~99 | 39 MB – 1.5 GB | 39 MB – 1.5 GB | Any CPU; NVIDIA GPU or Apple Silicon speeds up the bigger models a lot |
| Voice Cloner · Fast (OpenVoice) | 15 | ~150 MB + 60–115 MB/voice | ~0.3 GB | Any CPU — seconds per sentence |
| AI voice runtime (one-time, shared by the three tiers below) | — | Win 4.0 / Linux 4.9 / Mac 0.5 GB | Win 9 / Linux 12 / Mac 2 GB | Downloaded automatically with the first Good/Medium/High tier |
| Voice Cloner · Good (Chatterbox) | 23 | 2.8 GB | 3.0 GB | Works on any CPU (~30 s per sentence, 8 GB RAM); fast on NVIDIA or Apple Silicon |
| Voice Cloner · Medium (Qwen3-TTS) | 10 | 3.5 GB | 4.3 GB | CPU possible (slow, 16 GB RAM); best on NVIDIA or Apple Silicon |
| Voice Cloner · High (FireRedTTS3) | 24 | 10.8 GB | 12 GB | Needs an NVIDIA GPU (8 GB+ VRAM) or Apple Silicon (16 GB RAM) — no CPU mode |
| Vocal isolation (voice / instrumental split) | — | 28 MB | 28 MB | Any CPU; GPU-accelerated where available |
| Audio restore · voice mode (DeepFilterNet) | — | ~40 MB | ~40 MB | Any CPU |
| Images & documents | ||||
| Image upscaling 2×/4× (Real-ESRGAN) | — | ~65 MB | ~65 MB | Any CPU; a Vulkan-capable GPU makes 4K upscales many times faster |
| Background removal | — | 5 MB fast / 170 MB precise | up to 175 MB | Any CPU; accelerated via CoreML, DirectML or CUDA when present |
| OCR — scans, searchable PDFs, Screen Snip & Read (Tesseract) | 100+ | a few MB per language | < 100 MB | Any CPU |
Everything installed at once — every voice tier, the biggest Whisper model and all image AI — tops out around 34 GB on Linux, 31 GB on Windows and 24 GB on macOS. A typical setup (transcription + one or two AI features) stays under 5 GB.
What scales with hardware
Most file conversions run anywhere. The features below are the ones where better hardware makes a visible difference.
- Parallel audio batch conversion — scales with CPU cores. The app reserves 2 cores for the operating system and runs one ffmpeg process per remaining core. A 16-core machine processes 14 files in parallel; quad-cores process 2; older dual-cores fall back to serial automatically.
- Video conversion — uses your GPU encoder when one is present (NVENC / AMF / QSV / VAAPI / VideoToolbox), or your CPU otherwise. A 4K H.264 encode takes 1–2 minutes on a modern GPU and 10–20 minutes on CPU.
- Speech-to-text (subtitles & lyrics) — on Apple Silicon Macs, runs on the Neural Engine roughly 3× faster than CPU. On NVIDIA GPUs, runs on CUDA. On AMD, Intel, and older NVIDIA, runs on Vulkan. Without any of those, runs on CPU and still works fine on shorter clips.
- AI image upscaling (2× / 3× / 4×) — Vulkan GPU when present (sub-second per image on modern hardware) or CPU (several seconds to a minute per image, depending on resolution).
- Background removal — about 1–3 seconds per image on CPU; no GPU acceleration needed.
- PDF, document, image, archive and ebook conversion — minimal — runs comfortably on any supported machine.
Not supported
- 32-bit operating systems
- Windows 7, 8, and 8.1
- macOS 10.15 Catalina and earlier
- ARM Linux (no aarch64 build at the moment)
- CPUs without AVX support — required by the bundled video and transcription engines
About AI model downloads: first-time fetches are roughly 140 MB for transcription, 175 MB for background removal, and 60 MB for image upscaling. Models are cached in your home folder so they survive app updates and re-installs.
Not sure if your machine qualifies? Install on Linux for free, or grab the trial download — every feature is functional out of the box and you'll know within a few minutes whether your hardware keeps up.