Supported Formats and Limits

XMOX accepts the following file types for direct upload:

  • Audio: MP3, WAV, M4A, FLAC, OGG, OPUS
  • Video containers (audio track extracted automatically): MP4, MOV, WEBM, MKV
Maximum file size is 500 MB. For longer recordings, consider splitting the file or using a cloud integration (Zoom, Google Meet, Teams) which stream directly from the source.

Step-by-Step: Uploading a File

1
Open your XMOX workspace and click the + New Job button (or the upload icon in the toolbar).
2
In the source picker, select "Audio / Video File".
3
Click "Choose file" or drag and drop your file into the upload area. The upload progress bar will appear immediately.
4
Once uploaded, XMOX begins transcription automatically. The job card in your list will show a processing status badge while the AI works.
5
When the status changes to Ready, click the job to open the transcript. You can now read, search, and ask questions about the content.

Understanding Processing Time

Processing speed depends on file length and current server load. As a rough guide:

  • A 10-minute recording typically finishes in under 1 minute.
  • A 1-hour recording is usually ready within 3–5 minutes.
  • Very long recordings (3 hours+) may take 10–15 minutes.

You do not need to keep the page open — XMOX processes in the background. Refresh the job list when you return.

Asking Questions After Transcription

Once a job is ready, the AI chat panel on the right lets you query the transcript in plain language. Here are some effective prompts to try:

  • "Summarise the key points of this meeting."
  • "List every action item that was mentioned."
  • "What topics were discussed in the first 15 minutes?"
  • "Find every time the word 'deadline' was used."
  • "Who spoke the most and what did they focus on?"
You can also click any line in the transcript to jump to that exact moment in the audio player — useful for verifying context around a specific answer.

Getting the Best Transcription Quality

XMOX uses Gemini AI for transcription, which is highly accurate even with accented speech and technical vocabulary. For optimal results:

  • Recordings made in a quiet room with a close microphone produce the fewest errors.
  • Reduce background noise before uploading if possible (tools like Audacity's noise reduction work well).
  • Stereo recordings with multiple speakers on separate channels are handled automatically — the transcript will distinguish speakers.
  • If a word is consistently mis-transcribed, you can correct it inline and the corrected version is preserved in all exports.
Very low-quality recordings (heavy compression artefacts, severe background noise, or very soft speech) may produce reduced-accuracy transcripts. The AI will still make a best effort, but some words may be marked as uncertain.

Ready to upload your first audio file? Open your XMOX workspace and start in seconds.

Open Workspace →