Research and Development
Machine learning's open frontier.
Sound is one of the last hard signals left in machine learning: state of the art models, fundamentals still unsolved, and progress that would matter far beyond audio. Hard is the fun part.
“So you're not just using AI, you're actually doing machine learning.” · what we keep hearing from industry and academia leaders
01 · The Scale
One song is a million-step problem.
A paragraph of text
~300
One second of audio
44,100
A three minute song
~7,938,000
Log scale · every step right is 10x
Long horizon by construction. Coherence must survive millions of steps, and quadratic attention does not.
02 · The Structure
Coherent at every timescale, all at once.
Milliseconds
Timbre & texture
the sound itself
Seconds
Notes & beats
pitch, rhythm, groove
Tens of seconds
Phrases & motifs
themes that return
Minutes
Song form
tension, release, payoff
Long-range structure is the open problem of music generation. Progress here is progress for sequence modeling everywhere.
03 · The Signal
The same four notes, four ways.
Notation
how musicians write it
Piano roll · MIDI
how sequencers store it
Spectrogram
how models often see it
Waveform
what it really is: 44,100 numbers per second
Four views, one piece of music, and each view is its own machine learning problem. Music is multimodal by nature.
- Text
- Audio
- MIDI
- Video
- Image
- Brain
Yes, brain: recent models predict how a brain responds to sound.
If you come from computer vision, this will feel familiar. Audio is a dense temporal signal like video, and both fields fight the same battle: staying coherent over time.
- VAEs
- Tokenization
- Contrastive embeddings
- Diffusion transformers
- Flow matching
- RLHF
- State space models
- Mamba hybrids
- TRMs
- Clustering
- Deduplication
- Neural watermarking
- Transcription
- Source separation
- Representation learning
Every branch of machine learning shows up here. Same methods, new domain, open problems.
04 · The Breadth
Wider than it looks.
Circuits & gear
We build hardware too, guitar pedals included.
Chip design
Audio silicon, in contact with big companies in the space.
Classical DSP
Filters, synthesis, effects.
Music information retrieval
Transcription, tagging, search.
Deep learning
Generative models at the state of the art.
Neuroscience
How brains hear music.
From soldering irons to diffusion transformers. If it makes sound or understands it, it fits.
05 · What We Work On
Open directions, not a finished list.
New architectures for sound
Mathematical function families suited to music the way convolutions suit images, wavelets among them. We are building some from scratch.
Long horizon, less compute
Keeping generative models consistent over minutes with subquadratic methods: state space models, Mamba-style networks, TRMs, and hybrids such as diffusion Mamba.
Multimodal music intelligence
Learning across audio, text, MIDI, video, images, and brain response data.
Beyond generation
Transcription, clustering, deduplication, neural watermarking, retrieval. The full breadth of machine learning tasks, on sound.
DSP & hardware
Classical signal processing and gear built with our own hands, from audio effects and guitar pedals down to chip design.
Some projects start with our members, others with the companies we work with. Outcomes range from research papers to artistic projects, and everything in between.
From the Lab
One signal, many scales.
A glimpse of one project currently on our bench. Convolutions fit images because images are local. Musical structure nests: textures inside notes, notes inside phrases, phrases inside a song. Wavelets see all of those scales at once, and we are building architectures around them.
The signal
one line, three scales at once
Coarse
song form
Mid
phrases & notes
Fine
texture & timbre
Fixed resolution · Fourier / STFT
Multi-resolution · wavelets
Built With Industry
Many of our projects are built together with companies, and our work is supported by our partners.
Our partnersThe Longer Version
Why music is such a good place to learn machine learning.
Music and audio are still underexplored in machine learning, especially next to language and vision. There is far more research happening than most people expect, yet the field is nowhere near saturated. For a student, that gap is the opportunity: the fundamental questions are still open, and newcomers can reach the frontier fast.
Working on music means re-learning the fundamentals by adapting them to a new domain. Contrastive embeddings become audio-text models in the spirit of CLAP. Tokenization becomes neural audio codecs built on residual vector quantization. Autoencoders, diffusion models, and preference tuning all reappear, and each one has to be rethought for sound. At the same time, the state of the art stack is fully present: latent diffusion transformers, flow matching, RLHF, and beyond.
As a signal, music is closer to video than to text: continuous, high rate, and perceptual. One second of audio is 44,100 samples, and a three minute song is roughly eight million. Structure lives at every timescale simultaneously, from the milliseconds of timbre to the minutes of song form, and a model must stay coherent at all of them at once. Long-range structure is a recognized open problem in music generation, and solving it would advance sequence modeling far beyond audio. In practice, audio ML has more in common with video ML than with NLP: dense temporal perceptual signals, temporal coherence as the shared battle, and techniques that travel between the two fields constantly.
That is why efficiency is one of our main themes: subquadratic architectures such as state space models, Mamba-style networks, TRMs, and hybrids like diffusion Mamba. It is also why we design new architectures from scratch. One ongoing project explores wavelets and related multi-resolution function families: music is hierarchical by nature, and a transform that zooms may fit it the way convolutions fit images.
Music is a natural multimodal playground: audio, text, MIDI, video, images, and even brain responses, with recent models predicting how a brain reacts to sound. And generation is only part of the picture. Transcription, clustering, deduplication, neural watermarking, and retrieval are all live problems.
At Munich Music Labs, R&D is broader than machine learning alone. Teams work on classical DSP and build their own hardware, guitar pedals included. Chip design is on the table too, and we are already talking with major companies in that space. Our interests stretch from circuits and signal processing through neuroscience to state of the art deep learning. Projects run in small teams: some are started by members, others come from collaborations with companies and university chairs, and our work is supported by our partners. Outcomes range from research papers to artistic projects, and anyone can propose a direction.
Build this with us.
Join a community of 60+ members across research, creation, and organization. Pick the Research and Development department when you apply.





