News & Updates

How to Build an English UTAU Voicebank: A Complete Guide

By Dominic Hawke 15 min read 4424 views

How to Build an English UTAU Voicebank: A Complete Guide

Why Create an English Voicebank?

Most UTAU users gravitate toward Japanese phonetics because the software started there, but English singers are gaining traction. An English voicebank lets you produce lyrics that sound natural to Western ears, opening up collaborations, fan‑covers, and even game mods. Plus, the process itself teaches you a lot about recording, phoneme mapping, and sound design.

Getting Your Tools in Order

Before you hit “record,” assemble the basics.

  • UTAU – the latest stable build for your OS.
  • Resampler – preferably one that handles pitch‑shifting cleanly (e.g., Resamp or UTAU.NET).
  • Audio Interface – a decent USB mic or a small studio setup; avoid built‑in laptop mics.
  • DAW – Audacity works, but many prefer Reaper for its batch processing.
  • Folder structure – a dedicated directory for raw takes, processed wavs, and the oto.ini file.

Planning the Phoneme Set

English isn’t as tidy as Japanese’s 46‑syllable chart. You’ll need to decide whether to go with a “full” set (including diphthongs, r‑colored vowels, and some consonant clusters) or a “minimal” set that covers most pop songs.

Typical Minimal Set

  • Vowels: a, e, i, o, u, æ, ʌ, ɔ, ɪ, ɛ, ʊ, ɑ
  • Consonants: p, b, t, d, k, g, m, n, ŋ, s, z, ʃ, ʒ, f, v, θ, ð, l, r, w, j
  • Diphthongs: aɪ, eɪ, oʊ, aʊ, ɔɪ

If you’re ambitious, add “schwa” (ə) and a few tricky clusters like str or spl. The rule of thumb: record every sound you can imagine a lyric might need, then prune later.

Recording the Samples

Consistency is king. Record each phoneme at a steady pitch—most creators pick middle C (C4) or A4 (440 Hz). Keep the volume level uniform; a peak around –3 dB usually gives headroom without clipping.

Step‑by‑step workflow

  1. Warm‑up your voice. Light humming or tongue trills reduce strain.
  2. Set your DAW to record at 44.1 kHz, 16‑bit PCM. Higher rates are fine but increase file size.
  3. Record each phoneme in isolation, then immediately record a “neutral” breath to capture ambient noise.
  4. Label takes clearly: aa_01.wav, aa_02.wav, etc.
  5. After a batch, listen back and discard any take with clicks, pops, or noticeable breath.

A quick tip: use a pop filter and sit about 6‑8 inches from the mic. It evens out plosives without sacrificing presence.

Editing and Normalizing

Once you have raw files, the real work begins. Trim silence, remove breaths, then apply a gentle fade‑in/out (5‑10 ms) to avoid clicks.

  • Normalization: boost each file to the same RMS level, typically –18 LUFS for UTAU.
  • Noise reduction: a light spectral delete on the background hiss can make a huge difference.
  • Pitch correction: Only if a take is slightly off. Use a transparent algorithm; over‑processing makes the voice sound synthetic.

Save the cleaned files as 16‑bit WAVs, preserving the original naming convention.

Crafting the oto.ini File

This tiny text file tells UTAU where each sample starts, its length, and how it should be stretched. A typical line looks like:

aa.wav=0,150,60,100

Where the numbers represent start point, length, pre‑utterance, and overlap, respectively. Getting these values right is the most fiddly part, but you’ll feel the payoff when the voice sings without noticeable glitches.

Tips for accurate mapping

  • Use UTAU’s built‑in waveform viewer to line up the waveform’s peak with the start value.
  • Set pre‑utterance to roughly half the phoneme’s duration; adjust by ear.
  • Overlap should be longer for fricatives (s, f, ʃ) to smooth transitions.

Testing and Tweaking

Load your new voicebank in UTAU, select a simple melody, and type “Hello world.” If the l sounds clipped or the o feels too long, go back to oto.ini and nudge the parameters. Expect a few cycles of trial‑and‑error; it’s normal.

Packaging for Distribution

When you’re satisfied, bundle the character folder, the oto.ini, and a readme.txt that explains the phoneme set and any known quirks. Compress everything into a zip file, and consider uploading to platforms like VocaDB or the UTAU Wiki. Credit any third‑party tools you used, and you’ll likely get feedback that helps you improve future releases.

Common Pitfalls to Watch Out For

Even seasoned makers stumble on a few recurring issues.

  • Inconsistent breath noise: If some samples contain a faint hiss while others are dead quiet, the voicebank will feel uneven.
  • Missing diphthongs: English lyrics love “ow” and “ay.” Skipping them forces UTAU to splice odd combinations.
  • Over‑compression: Squashing dynamics makes the singer sound flat; keep a natural range.

When you notice any of these, isolate the offending samples and treat them as a mini‑project: re‑record, re‑edit, and replace.

Next Steps: Voicebank Expansion

Once the core set is solid, you can add expressive layers—whispers, screams, or a “soft” voice mode. Many creators also record alternative pitches (e.g., A3, C5) to reduce the reliance on resampling, which yields a richer, more authentic tone.

Building an English UTAU voicebank is a blend of technical precision and vocal artistry. It may feel daunting at first, but each step builds on the last, and the community is surprisingly supportive. With patience, a decent mic, and a clear phoneme plan, you’ll have a functional English singer ready to tackle anything from pop ballads to indie game soundtracks.

How to Make an UTAU Voicebank | A Beginner's Guide - YouTube
How To Make An Utau Voicebank
How do I make a voicebank in openutau? | Fandom
How to Make A UTAU Voicebank By using the vowel a & Numbers - YouTube

Written by Dominic Hawke

Dominic Hawke is a Chief Correspondent with over a decade of experience covering breaking trends, in-depth analysis, and exclusive insights.