Event-based audio coding system (VPC system)
The audio coding system converts input signals into a structured event representation, addressing inefficiencies in existing systems by achieving high compression and enabling direct semantic and rhythmic analysis of speech and vocal signals.
Patent Information
- Application Number
- DE202025003859
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-19
- Estimated Expiration
- 2035-12-31
AI Technical Summary
Existing audio coding systems fail to efficiently represent speech and vocal signals as a sequence of discrete events, lacking explicit phoneme and rhythmic structures, leading to inefficient compression and limited semantic and rhythmic analysis capabilities.
An audio coding system that converts input signals into a structured event representation, comprising phonemic and rhythmic segmentation, compact event coding, and a unified file format (.vpc) to store phoneme and rhythmic events, enabling direct semantic and rhythmic analysis.
Achieves high-level compression with bit rates of 1-6 kbps, allowing intelligible speech and vocal signal reconstruction, and enabling direct semantic and rhythmic analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field
[0001] The invention relates to an audio coding system for high-level compression of speech and vocal signals. The system is based on an event-oriented representation in which not the continuous audio signal, but a sequence of discrete speech and musical events is stored. State of the art
[0002] Well-known audio coding systems such as MP3, AAC, Opus, or modern neural codecs are predominantly signal-based. They quantize temporal-frequency representations of the input signal and reduce redundant or psychoacoustically less relevant information. Lossless systems such as FLAC or ALAC reduce statistical redundancy without explicitly representing the linguistic or musical structure.
[0003] Low-bit-rate speech encoders and newer neural speech coding systems achieve bit rates in the range of approximately 1–6 kbps, but still generate signal-oriented representations. They encode acoustic feature vectors or latent representations of the input signal, but not explicit phoneme or clock events. The resulting bitstream therefore remains predominantly signal-oriented and not content-oriented.
[0004] While symbolic formats like MIDI efficiently represent musical events, they do not take into account speech phonemes or the coupling of speech and rhythmic structure characteristic of singing. In particular, they do not provide a common, reconstructible data stream that integrates both linguistic and musical-metrical events. Differentiation from the state of the art
[0005] The current state of the art in audio coding encompasses both classic, signal-based codecs (e.g., LPC, CELP, MELP, and MPEG-based systems) and newer, neural, low-bitrate speech codecs, in which continuous feature vectors or discrete latent codebook tokens of the audio signal are transmitted at a significantly reduced bit rate. These established systems predominantly operate on a temporally dense, acoustically oriented representation of the signal and do not provide an explicit sequence of phoneme or clock events as a primary, standardized data stream suitable for both speech and vocal signals.
[0006] In contrast, the audio coding system according to the invention is based on the approach of no longer storing the audio as a continuous sequence of acoustic feature vectors, but rather as an explicit event sequence consisting of phoneme and optional rhythmic events with associated linguistic and musical-metric parameters. The resulting event stream is stored in a uniform file format ".vpc" (Voice-Phoneme-Codec). This creates a content-oriented bitstream that enables both high-level compression and direct semantic and rhythmic analysis (e.g., phoneme pattern search, recognition of rhythmic motifs, or prosody statistics) on the compressed representation, which is not provided in this form by the aforementioned signal-based or latent vector-based codecs. Object of the invention
[0007] The invention is based on the objective of providing an audio coding system and an associated file format with which speech and vocal signals can be stored and transmitted at a significantly reduced data rate compared to the prior art. The system is not intended to represent the continuous acoustic signal, but rather to capture its linguistic and musical event structure in a compact form.
[0008] The aim is to maintain sufficient reconstruction quality with regard to intelligibility, rhythm, and basic character of the voice. At the same time, the generated bitstream should allow for direct semantic and rhythmic analysis. Solution to the task
[0009] To solve the task, an audio coding system is provided, designed to convert an input signal into a sequence of discrete events and store them compactly. The system includes, in particular, the following functional modules: System structure and functional modules
[0010] The audio coding system according to the invention comprises several functional modules that are jointly configured to convert a speech or singing signal into a structured event representation. 1. Module for phonemic segmentation
[0011] The system is designed to break down the input signal of a speech or singing signal into a sequence of speech sounds from a predefined phoneme inventory (for example, 28-30 phoneme classes for a target language or a cross-linguistic inventory).
[0012] For each phoneme event, at least the following parameters are recorded and stored: • Phoneme ID from inventory • Duration of the phoneme • medium or contoured energy • Fundamental frequency or pitch or pitch contour 2. Module on tactical and rhythmic segmentation (for musical content)
[0013] For content with a clear musical structure, especially vocals, the system is set up to determine additional rhythmic and metric parameters, including: • global or locally variable tempo • Time signature (e.g. 4 / 4, 3 / 4) • Emphasis patterns • Note lengths relative to the time signature
[0014] This embeds the phoneme events into an explicit metric time structure. 3. Module for compact event coding
[0015] The audio coding system is configured to convert all recorded events into a linear or block-structured event list with a time reference. Each event contains its type identifier (phoneme, pause, metric event) and its associated parameter values.
[0016] To reduce the average bit rate, the system is further configured to apply suitable entropy coding methods to the discrete parameters. These include, in particular, Huffman and arithmetic coding, which are applied to parameters such as phoneme ID, quantized durations, and quantized pitch classes. 4. Module for recording optional additional parameters
[0017] To improve reconstruction quality, the audio coding system can capture and store additional parameters for individual events or groups of events. These optional parameters include, in particular: • Formants or spectral envelope features to characterize voice timbre • Vibrato parameters • detailed melodic contours • Speaker- or singer-related identity characteristics
[0018] These additional parameters allow for a more precise description of the vocal and musical characteristics and contribute to a higher quality reconstruction of the audio signal. 5. Module for the file format (.vpc container)
[0019] The resulting compressed event stream is stored in a defined container format called ".vpc" (Voice-Phoneme-Codec). This container format includes metadata such as the phoneme inventory used, the sampling rate of the original signal, and profile information. It also contains the actual event bitstream in a compact form. reconstruction
[0020] The audio coding system includes a reconstruction module that generates an audible audio signal from the event sequence.
[0021] The module can be implemented as a parametric vocoder or as a neural synthesis system.
[0022] Phoneme type, duration, energy, pitch and optional timbre parameters control a source-filter model or a neurally conditioned model that generates a continuous waveform.
[0023] The metric parameters (tempo, time signature, accents and note lengths) control the temporal placement and weighting of the phoneme events, so that the original rhythm and the syllable or word accent structure are largely preserved in the reconstructed speech or singing sequence. Advantages of the system • Extremely low bitrate: Since only a few discrete events with relatively small parameter vectors need to be encoded per time interval, speech content can be stored or transmitted at bit rates in the range of about 1-2 kbps and vocal content at about 3-6 kbps, which represents an additional reduction compared to PCM and modern lossy codecs. • Structural and semantic information: The bitstream contains explicit linguistic and musical events, so that semantic and rhythmic analyses such as phoneme sequence or word pattern searches, recognition of rhythmic motifs, or statistical evaluations of prosody can be performed directly on the compressed data stream. • Reconstructive capability: Despite the significant reduction in data volume, event representation, in conjunction with a suitable vocoder or synthesis module, enables the generation of intelligible and rhythmically consistent speech and vocal signals suitable for telecommunications, archiving, or playback applications. • Universal, content-oriented format: The vpc format describes speech and music in a common, compact event space and is thus in a similar relationship to conventional audio signals as MIDI is to instrument performances, but extended to include the representation of phonemes and vocals. Application areas
[0024] The system is particularly suitable for: • Highly compressed speech and vocal communication over channels with very low bandwidth, e.g., specialized radio or satellite connections • Long-term archiving of large speech and song corpora where storage space and semantic searchability are of particular importance • Systems for semantic audio analysis, automatic subtitling, or music and speech forensics, where working directly on a structured event stream offers advantages. • To the best of the applicants' knowledge, no audio coding systems have been disclosed in the German and European intellectual property rights area in which speech and singing signals are primarily stored as an event sequence of phoneme and beat events with associated linguistic and musical-metric parameters in a uniform file format. • Known speech and audio codecs either operate on a signal-based basis with time-dense feature vectors or latent codebook tokens and do not provide a content-oriented, semantically evaluable event bitstream as defined by the present invention.
Claims
[1] Audio coding system for compressing audio information, characterized by , that it comprises a segmentation unit that decomposes an input signal of a speech and / or singing signal into an event sequence of discrete phoneme events and optional clock events, a parameterization unit that determines for each phoneme event at least one phoneme identifier from a predefined phoneme inventory, a duration assigned to that event, an energy quantity and a pitch or pitch contour, and a coding and storage unit that converts the entirety of the events into a compact bitstream using an entropy coding procedure and stores this in a file with the file format “.vpc” (Voice-Phoneme-Codec). [2] Audio coding system according to the preceding claim, characterized by, that the parameter unit for beat events additionally determines rhythmic parameters, in particular tempo, time signature and accent information, and assigns them to the event stream to be stored. [3] Audio coding system according to any one of the preceding claims, characterized by , that the parameterization unit determines formant and / or spectral envelope parameters for at least some of the phoneme events to reconstruct the voice color and encodes them as additional parameters. [4] Audio coding system according to any one of the preceding claims, characterized by , that the parameterization unit for vocal segments determines and provides melodic contours including vibrato and pitch contours as further parameters of the events. [5] Audio coding system according to any one of the preceding claims, characterized by that the coding and storage unit uses Huffman coding and / or arithmetic coding as its entropy coding method. [6] Audio coding system according to any one of the preceding claims, characterized by that it is coupled with a decoding unit that generates speech and / or vocal signals reconstructed from the compressed bitstream using a vocoder or another parametric or neural synthesis module. [7] Digital audio file, characterized by , that it comprises a bitstream stored in the file format “.vpc” (Voice-Phoneme-Codec) which was generated by an audio coding system according to one of the preceding claims and comprises an event sequence of discrete phoneme events and optional clock events with the respective associated parameters.