Audio embedding
A local CLAP-derived model represents how a file sounds and what language commonly describes it. This powers text-to-audio and audio-to-audio retrieval.
Sample converts audio into a local 512-dimension semantic representation, combines it with BPM, key, duration, instrument, timbral metadata, and your text intent, then ranks the files already on your drives.
Sample uses different methods for different jobs. That makes behavior easier to explain—and easier to trust.
A local CLAP-derived model represents how a file sounds and what language commonly describes it. This powers text-to-audio and audio-to-audio retrieval.
Separate local analysis estimates BPM, key, duration, instrument family, and timbral properties. Confidence gates prevent weak metadata from pretending to be exact.
Explicit phrases such as “120 BPM” or “C# minor” become structured constraints; descriptive language remains semantic. The two result channels are merged rather than conflated.
A local approximate-nearest-neighbor index retrieves likely matches quickly. Filters and confidence-aware eligibility rules are applied before results reach the interface.
A reference sample anchors the search. “More” and “less” concepts change the vector direction. Optional BPM and key locks narrow the result set only when the reference and candidate metadata are reliable enough.

The practical win is vocabulary independence: the pack vendor’s filename no longer has to match the phrase in your head.
Conceptual capability map, not a recall benchmark. Exact coverage depends on the query, library, metadata confidence, and model behavior.
A useful buying page should say what the system is not designed to do.
Finds audio a person might describe similarly: material, mood, instrument, texture, density, brightness, impact, or production character.
Key and BPM are estimates with confidence rules. Similarity is not beat alignment, groove matching, copyright identification, or proof that two sounds are interchangeable.
The catalog and analysis can change without renaming or moving audio. Normal library removal does not silently delete source files or preserved analysis.
Use observed retrieval time from one real session for the most defensible estimate.
recovered per year—about 1.2 hours per week, worth an estimated $2,042 of studio time.
Planning model, not a product-performance guarantee. It assumes 50 active weeks and uses only the values you enter. Time a normal session to replace the defaults with your own baseline.
Sample searches and analyzes locally on Apple Silicon. Your audio stays on your Mac. Use it offline after activation, drag results into any DAW, and request a refund within 14 days if it does not improve your workflow.