Skip to content

Research and Development

Machine learning's open frontier.

Sound is one of the last hard signals left in machine learning: state of the art models, fundamentals still unsolved, and progress that would matter far beyond audio. Hard is the fun part.

“So you're not just using AI, you're actually doing machine learning.” · what we keep hearing from industry and academia leaders

01 · The Scale

One song is a million-step problem.

A paragraph of text

~300

One second of audio

44,100

A three minute song

~7,938,000

Log scale · every step right is 10x

Long horizon by construction. Coherence must survive millions of steps, and quadratic attention does not.

02 · The Structure

Coherent at every timescale, all at once.

  1. Milliseconds

    Timbre & texture

    the sound itself

  2. Seconds

    Notes & beats

    pitch, rhythm, groove

  3. Tens of seconds

    Phrases & motifs

    themes that return

  4. Minutes

    Song form

    tension, release, payoff

Long-range structure is the open problem of music generation. Progress here is progress for sequence modeling everywhere.

03 · The Signal

The same four notes, four ways.

Notation

how musicians write it

Piano roll · MIDI

how sequencers store it

Spectrogram

how models often see it

Waveform

what it really is: 44,100 numbers per second

Four views, one piece of music, and each view is its own machine learning problem. Music is multimodal by nature.

  • Text
  • Audio
  • MIDI
  • Video
  • Image
  • Brain

Yes, brain: recent models predict how a brain responds to sound.

If you come from computer vision, this will feel familiar. Audio is a dense temporal signal like video, and both fields fight the same battle: staying coherent over time.

  • VAEs
  • Tokenization
  • Contrastive embeddings
  • Diffusion transformers
  • Flow matching
  • RLHF
  • State space models
  • Mamba hybrids
  • TRMs
  • Clustering
  • Deduplication
  • Neural watermarking
  • Transcription
  • Source separation
  • Representation learning

Every branch of machine learning shows up here. Same methods, new domain, open problems.

04 · The Breadth

Wider than it looks.

  1. Circuits & gear

    We build hardware too, guitar pedals included.

  2. Chip design

    Audio silicon, in contact with big companies in the space.

  3. Classical DSP

    Filters, synthesis, effects.

  4. Music information retrieval

    Transcription, tagging, search.

  5. Deep learning

    Generative models at the state of the art.

  6. Neuroscience

    How brains hear music.

From soldering irons to diffusion transformers. If it makes sound or understands it, it fits.

05 · What We Work On

Open directions, not a finished list.

New architectures for sound

Mathematical function families suited to music the way convolutions suit images, wavelets among them. We are building some from scratch.

Long horizon, less compute

Keeping generative models consistent over minutes with subquadratic methods: state space models, Mamba-style networks, TRMs, and hybrids such as diffusion Mamba.

Multimodal music intelligence

Learning across audio, text, MIDI, video, images, and brain response data.

Beyond generation

Transcription, clustering, deduplication, neural watermarking, retrieval. The full breadth of machine learning tasks, on sound.

DSP & hardware

Classical signal processing and gear built with our own hands, from audio effects and guitar pedals down to chip design.

Some projects start with our members, others with the companies we work with. Outcomes range from research papers to artistic projects, and everything in between.

From the Lab

One signal, many scales.

A glimpse of one project currently on our bench. Convolutions fit images because images are local. Musical structure nests: textures inside notes, notes inside phrases, phrases inside a song. Wavelets see all of those scales at once, and we are building architectures around them.

The signal

one line, three scales at once

Coarse

song form

Mid

phrases & notes

Fine

texture & timbre

Fixed resolution · Fourier / STFT

Multi-resolution · wavelets

Built With Industry

Many of our projects are built together with companies, and our work is supported by our partners.

Our partners

The Longer Version

Why music is such a good place to learn machine learning.

Music and audio are still underexplored in machine learning, especially next to language and vision. There is far more research happening than most people expect, yet the field is nowhere near saturated. For a student, that gap is the opportunity: the fundamental questions are still open, and newcomers can reach the frontier fast.

Working on music means re-learning the fundamentals by adapting them to a new domain. Contrastive embeddings become audio-text models in the spirit of CLAP. Tokenization becomes neural audio codecs built on residual vector quantization. Autoencoders, diffusion models, and preference tuning all reappear, and each one has to be rethought for sound. At the same time, the state of the art stack is fully present: latent diffusion transformers, flow matching, RLHF, and beyond.

As a signal, music is closer to video than to text: continuous, high rate, and perceptual. One second of audio is 44,100 samples, and a three minute song is roughly eight million. Structure lives at every timescale simultaneously, from the milliseconds of timbre to the minutes of song form, and a model must stay coherent at all of them at once. Long-range structure is a recognized open problem in music generation, and solving it would advance sequence modeling far beyond audio. In practice, audio ML has more in common with video ML than with NLP: dense temporal perceptual signals, temporal coherence as the shared battle, and techniques that travel between the two fields constantly.

That is why efficiency is one of our main themes: subquadratic architectures such as state space models, Mamba-style networks, TRMs, and hybrids like diffusion Mamba. It is also why we design new architectures from scratch. One ongoing project explores wavelets and related multi-resolution function families: music is hierarchical by nature, and a transform that zooms may fit it the way convolutions fit images.

Music is a natural multimodal playground: audio, text, MIDI, video, images, and even brain responses, with recent models predicting how a brain reacts to sound. And generation is only part of the picture. Transcription, clustering, deduplication, neural watermarking, and retrieval are all live problems.

At Munich Music Labs, R&D is broader than machine learning alone. Teams work on classical DSP and build their own hardware, guitar pedals included. Chip design is on the table too, and we are already talking with major companies in that space. Our interests stretch from circuits and signal processing through neuroscience to state of the art deep learning. Projects run in small teams: some are started by members, others come from collaborations with companies and university chairs, and our work is supported by our partners. Outcomes range from research papers to artistic projects, and anyone can propose a direction.

Build this with us.

Join a community of 60+ members across research, creation, and organization. Pick the Research and Development department when you apply.