Paste a Spotify track URL to download it and use as a cover source. Saved tracks (and their transcribed melody, once transcribed) live here so you can reuse them without re-downloading or re-transcribing โ pick "Cover" on one to jump to Create with it already loaded.
Tap steps to build a beat โ drag across a row to paint several at once. Tap a bass step again to cycle its note. When it sounds good, send it to Cover to turn it into a full song.
One free-text description, not a tag list โ YuE2 reads genre, instruments, vocal character, language, and tempo all together. Example: "English, warm piano pop, expressive female voice, 88 BPM." The genre chips above the field fill in a starting point you can edit. The model isn't limited to those examples โ it can also do electronic styles like techno, house, drum & bass, and happy hardcore; just describe the tempo (BPM), energy, and instrumentation you want (e.g. "aggressive happy hardcore, distorted four-on-the-floor kicks, euphoric synth stabs, pitched-up vocal chops, 170 BPM").
Use section tags โ [Verse], [Chorus], [Bridge],
[Intro], [Outro] โ each on its own line, followed by one sung line per
line. Only include words meant to be sung; stage directions or notes-to-self will get sung too.
See the project's own
full formatting guide โ for more detail.
Controls how much of the song's editable score YuE2 writes out before generating audio.
A number that determines the randomness used during generation. The same style, lyrics, and seed will always produce the same song. Change the seed to get a different take on the same lyrics/style; keep it the same while you tweak other settings to compare results fairly.
Classifier-free guidance scale โ how strictly generation follows your style/tags description versus how much freedom the model takes. Higher values stick closer to what you described (but can sound less natural if pushed too far); lower values give the model more creative freedom. Leave it blank to use YuE2's own default, which works well for most styles.
Upload, record from your microphone, or capture audio playing in another app/tab. SheetSage2 transcribes its melody into an editable ABC score, and Whisper transcribes its lyrics, so neither has to be typed by hand. The original recording is always saved alongside the song it produces, so you can go back and reuse it later. Use plan mode "melody" so the accompaniment can adapt to a new style around the transcribed melody. Recording from the browser needs this page served over HTTPS (or localhost). Models download on first use, which can take a while.
If plan mode wrote a score, it's rendered here as sheet music. Drag a note up or down to change its pitch, or click a note to jump to its position in the raw text below โ edits to either stay in sync. Paste in your own ABC score to generate from it directly.
Mixes a synthesized binaural beat underneath the generated song โ two pure sine tones a few Hz apart, one per ear, which the brain perceives as a "beat" at the difference frequency. This is plain signal synthesis, not something YuE2 itself generates (it has no way to hit an exact per-channel frequency). The 174 Hz genre chip fills in a commonly-used base frequency and switches this on; the base/beat preset dropdowns fill in the number fields below them, which you can then fine-tune directly. Needs headphones โ the stereo separation it relies on is lost on speakers.
Two different traditions are mixed into the presets, worth telling apart: the named "solfeggio" frequencies (174โ963 Hz) are a folk/wellness convention with no clinical evidence behind the specific numbers โ included because they're the frequencies people specifically ask for, not as a medical claim. 40 Hz gamma is different: there's real, if early and still unsettled, neuroscience research on rhythmic 40 Hz sensory stimulation and its effects on brain activity โ worth knowing that's a genuinely distinct (and more scientifically discussed) idea from solfeggio, not more proven for pain relief specifically.
Record or upload a few minutes of yourself talking or singing to build a reusable voice clone. Once it's ready, apply it to any generated song from that song's card โ the original song is never changed, so you always keep it alongside the version sung as you. Training runs on the same GPU as song generation, so it happens one at a time and can take several minutes.
Edited โ Save to keep these changes on the song itself, or Close to discard them.