Audio processing method and device based on discrete mark prediction, equipment and medium

CN122531404APending Publication Date: 2026-08-07PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-05-19
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供一种基于离散标记预测的音频处理方法、装置、设备及存储介质,旨在解决现有生成式语音增强技术在离散标记预测过程中依赖顺序生成机制,导致推理过程难以并行执行,从而存在处理延迟较高且系统结构复杂的技术问题

Benefits of technology

[0010]Beneficial Effects: This invention relates to the field of speech processing technology and discloses an audio processing method, apparatus, device, and medium based on discrete marker prediction, comprising: acquiring an audio signal to be processed and performing time-frequency transformation processing to obtain a time-frequency representation; inputting the time-frequency representation into a feature encoding network to obtain a continuous feature representation, and dividing it into multiple sub-feature representations; inputting the multiple sub-feature representations into multiple vector quantizers respectively to obtain a set of degenerate discrete markers containing multiple sets of degenerate discrete markers; extracting spectral features and aligning the time dimension of the audio signal to be processed to obtain a conditional feature representation; embedding the multiple sets of degenerate discrete markers to obtain multiple sets of discrete marker embeddings; inputting the multiple sets of discrete marker embeddings and conditional feature representations into multiple prediction branches for parallel prediction to obtain multiple sets of target discrete markers; mapping the multiple sets of target discrete markers into multiple sets of codeword vectors and concatenating them into a quantized feature representation, and inputting it into a decoding network to obtain an enhanced audio signal. This invention can be applied to business scenarios such as fintech and healthcare. By dividing continuous features and forming multiple sets of degenerate discrete labels, and combining conditional feature representations to perform parallel prediction in multiple prediction branches, it replaces the stepwise generation method. At the same time, it maps and reconstructs multiple sets of target discrete labels into enhanced audio signals. Therefore, while ensuring speech enhancement processing capabilities, it reduces inference latency and system complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531404A_ABST
    Figure CN122531404A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech processing, and discloses an audio processing method, device and equipment based on discrete mark prediction and a medium, comprising: converting a to-be-processed audio signal into a time-frequency representation, extracting a continuous feature representation and dividing the continuous feature representation into a plurality of sub-feature representations, and obtaining a plurality of sets of degenerate discrete marks through vector quantization; performing spectral feature extraction and time alignment on the audio signal to obtain a conditional feature representation; performing embedding processing on the degenerate discrete marks and inputting the degenerate discrete marks and the conditional feature representation into a plurality of prediction branches to perform parallel prediction to obtain target discrete marks; and mapping and inputting the target discrete marks into a decoding network to obtain an enhanced audio signal. The present application can be applied to business scenarios such as financial technology and medical health, and through the construction of a plurality of sets of discrete marks and the combination of conditional features in a plurality of prediction branches for parallel prediction, the sequential generation mode is replaced, thereby reducing inference delay and system complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to an audio processing method, apparatus, device, and medium based on discrete label prediction. Background Technology

[0002] As speech enhancement technology evolves from traditional signal processing methods to generative methods, speech modeling based on discrete representations has gradually become a research hotspot. However, existing generative speech enhancement techniques generally rely on autoregressive prediction structures, which are difficult to parallelize when generating discrete tokens step by step, resulting in high inference latency and failing to meet real-time processing requirements. Furthermore, to improve generation quality, some schemes introduce complex contextual modeling structures, increasing system computational complexity, creating overall structural redundancy, and raising engineering deployment costs, thus limiting their large-scale application in real-world scenarios. Therefore, how to reduce inference latency and improve the simplicity of the system structure while ensuring speech enhancement quality has become a critical issue that urgently needs to be addressed.

[0003] In the fintech sector, voice interaction is widely used in scenarios such as intelligent customer service, telephone banking, voice-based risk control and auditing, and voice quality inspection. In these scenarios, voice data typically originates from complex communication environments, making it susceptible to background noise, channel compression distortion, and interference from multiple speakers, leading to decreased voice quality and consequently affecting speech recognition accuracy and business processing efficiency. Existing generative speech enhancement methods, when processing such voice data, suffer from significant delays in the inference process due to the reliance on sequential generation mechanisms for discrete label prediction, making it difficult to meet the real-time response and high-concurrency processing requirements of financial services. Furthermore, the complex model structure increases computational resource consumption, resulting in high deployment costs in large-scale customer service systems and limiting system scalability and stable operation.

[0004] In the healthcare field, voice technology is widely used in scenarios such as remote consultations, electronic medical record recording, medical voice-assisted input, and rehabilitation assessment. In these applications, voice signals often contain environmental noise, differences in device acquisition, and individual patient pronunciation variations, leading to unstable voice quality and affecting the accuracy of subsequent voice understanding and information extraction. Existing generative speech enhancement technologies, when applied in this field, also face processing latency issues caused by the progressive prediction of discrete labels, making it difficult to support real-time interaction requirements. Furthermore, the high complexity of the model structure hinders deployment in resource-constrained medical devices or edge environments, limiting the practical application effectiveness of these technologies. Summary of the Invention

[0005] The main objective of this invention is to provide an audio processing method, apparatus, device, and storage medium based on discrete tag prediction, aiming to solve the technical problem that existing generative speech enhancement technologies rely on sequential generation mechanisms in the discrete tag prediction process, which makes it difficult to execute the inference process in parallel, resulting in high processing latency and complex system structure.

[0006] To achieve the above objectives, the present invention provides an audio processing method based on discrete label prediction, comprising: Acquire the audio signal to be processed, and perform time-frequency transformation processing on the audio signal to be processed to obtain a time-frequency representation; The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension; The multiple sub-feature representations are respectively input into multiple vector quantizers to obtain a degenerate discrete label set, which includes multiple sets of degenerate discrete labels; The audio signal to be processed is subjected to spectral feature extraction to obtain spectral features, and the spectral features are aligned with the time dimension that matches the time scale of the degenerate discrete label set to obtain a conditional feature representation; The multiple sets of degenerate discrete tags are respectively embedded to obtain multiple sets of discrete tag embeddings; Multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. Multiple sets of target discrete markers are mapped to multiple sets of codeword vectors, and the multiple sets of codeword vectors are concatenated to form a quantized feature representation. The quantized feature representation is then input into a decoding network to obtain an enhanced audio signal.

[0007] Furthermore, to achieve the above objectives, the present invention provides an audio processing apparatus based on discrete label prediction, comprising: The audio time-frequency conversion module is used to acquire the audio signal to be processed, perform time-frequency conversion processing on the audio signal to be processed, and obtain a time-frequency representation; The feature encoding module is used to input the time-frequency representation into the feature encoding network to obtain a continuous feature representation, and to divide the continuous feature representation into multiple sub-feature representations along the feature dimension; The vector quantization module is used to input multiple sub-feature representations into multiple vector quantizers respectively to obtain a degenerate discrete label set, wherein the degenerate discrete label set includes multiple sets of degenerate discrete labels; The conditional feature construction module is used to extract spectral features from the audio signal to be processed, obtain spectral features, and align the spectral features with the time scale of the degenerate discrete label set to obtain a conditional feature representation. The embedding mapping module is used to perform embedding processing on multiple sets of the degenerate discrete tags respectively to obtain multiple sets of discrete tag embeddings; The parallel prediction module is used to embed multiple sets of discrete labels into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and input the conditional feature representation into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. The signal reconstruction module is used to map multiple sets of target discrete markers into multiple sets of codeword vectors, and to concatenate the multiple sets of codeword vectors into a quantized feature representation. The quantized feature representation is then input into the decoding network to obtain an enhanced audio signal.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an audio processing program based on discrete label prediction stored in the memory and executable on the processor, wherein the audio processing program based on discrete label prediction, when executed by the processor, implements the steps of the audio processing method based on discrete label prediction as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an audio processing program based on discrete label prediction, wherein the audio processing program based on discrete label prediction, when executed by a processor, implements the steps of the audio processing method based on discrete label prediction as described above.

[0010] Beneficial Effects: This invention relates to the field of speech processing technology and discloses an audio processing method, apparatus, device, and medium based on discrete marker prediction, comprising: acquiring an audio signal to be processed and performing time-frequency transformation processing to obtain a time-frequency representation; inputting the time-frequency representation into a feature encoding network to obtain a continuous feature representation, and dividing it into multiple sub-feature representations; inputting the multiple sub-feature representations into multiple vector quantizers respectively to obtain a set of degenerate discrete markers containing multiple sets of degenerate discrete markers; extracting spectral features and aligning the time dimension of the audio signal to be processed to obtain a conditional feature representation; embedding the multiple sets of degenerate discrete markers to obtain multiple sets of discrete marker embeddings; inputting the multiple sets of discrete marker embeddings and conditional feature representations into multiple prediction branches for parallel prediction to obtain multiple sets of target discrete markers; mapping the multiple sets of target discrete markers into multiple sets of codeword vectors and concatenating them into a quantized feature representation, and inputting it into a decoding network to obtain an enhanced audio signal. This invention can be applied to business scenarios such as fintech and healthcare. By dividing continuous features and forming multiple sets of degenerate discrete labels, and combining conditional feature representations to perform parallel prediction in multiple prediction branches, it replaces the stepwise generation method. At the same time, it maps and reconstructs multiple sets of target discrete labels into enhanced audio signals. Therefore, while ensuring speech enhancement processing capabilities, it reduces inference latency and system complexity. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an audio processing method based on discrete label prediction according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the audio processing method based on discrete label prediction according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the audio processing device based on discrete label prediction of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The audio processing method based on discrete label prediction provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain the audio signal to be processed from the client and perform time-frequency transformation to obtain a time-frequency representation. The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, which is then divided into multiple sub-feature representations. These sub-feature representations are input into multiple vector quantizers to obtain a set of degenerate discrete tags containing multiple sets of degenerate discrete tags. The server performs spectral feature extraction and time dimension alignment on the audio signal to be processed to obtain a conditional feature representation. The multiple sets of degenerate discrete tags are embedded to obtain multiple sets of discrete tag embeddings. The multiple sets of discrete tag embeddings and conditional feature representations are input into multiple prediction branches for parallel prediction to obtain multiple sets of target discrete tags. The multiple sets of target discrete tags are mapped to multiple sets of codeword vectors and concatenated into a quantized feature representation, which is then input into a decoding network to obtain an enhanced audio signal. This invention can be applied to business scenarios such as fintech and healthcare. By dividing continuous features into multiple sets of degenerate discrete labels, and combining conditional feature representations, it performs parallel predictions in multiple prediction branches, replacing the stepwise generation method. Simultaneously, it maps and reconstructs multiple sets of target discrete labels into enhanced audio signals, thus reducing inference latency and system complexity while maintaining speech enhancement processing capabilities. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the audio processing method based on discrete label prediction provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the audio processing method based on discrete label prediction proposed in this invention includes the following steps: S10, acquire the audio signal to be processed, perform time-frequency transformation processing on the audio signal to be processed, and obtain a time-frequency representation; In this embodiment, acquiring the audio signal to be processed involves collecting the original voice data source and standardizing the data format. The audio signal to be processed can come from the audio stream data in the voice interaction terminal, the acquisition interface of the communication device, or the voice data storage system. The data format is usually a continuous time domain waveform sequence, which is composed of discrete sampling points. Each sampling point corresponds to the amplitude change of the sound pressure signal. The sampling process can be completed by the analog-to-digital conversion unit. The sampling rate and quantization accuracy are set according to the business scenario. In the financial technology business scenario, it can correspond to call center voice, risk control review voice, or transaction voice command. In the medical and health business scenario, it can correspond to remote voice inquiry, health data voice input, or rehabilitation voice feedback signal. During the acquisition process, the audio signal needs to be processed in a unified channel and the amplitude range needs to be standardized so that the audio data from different sources have a consistent data scale and format expression in the subsequent processing stage.

[0016] Time-frequency transformation (TF-FDT) of audio signals involves mapping a one-dimensional time-domain signal into a two-dimensional representation containing both time and frequency dimensions. This process relies on segmenting the audio signal and performing frequency-domain mapping on each segment. Segmentation involves dividing the audio signal into multiple local time segments using a sliding window on the time axis. Each time segment has a stable spectral structure within a short time range. Discrete frequency-domain mapping is then performed on each time segment, converting the time-domain waveform into a complex spectral representation. The complex spectrum consists of amplitude and phase components. The amplitude component reflects the energy distribution of different frequency components, while the phase component reflects the relative position information of each frequency component. By combining the spectral results of multiple time segments in chronological order, a time-frequency matrix is ​​formed. The rows or columns of this matrix correspond to the frequency and time dimensions, respectively. In financial scenarios, this process can separate noise interference from the main speech in a call into different frequency structures. In healthcare scenarios, it can distinguish pronunciation details from environmental noise in the frequency domain, thus providing structured input data for subsequent processing.

[0017] The obtained time-frequency representation is manifested as a unified format output of the above frequency domain mapping results. This time-frequency representation can be expressed in the form of an amplitude spectrum matrix, a logarithmic amplitude spectrum matrix, or a complex spectrum matrix. It can also be expressed by logarithmically compressing the amplitude components to reduce the dynamic range, or by performing filter bank transformation on the frequency dimension to enhance information in specific frequency bands. Each element in the time-frequency representation corresponds to the energy or phase information at a certain time position and a certain frequency position. This representation can be directly input into the feature encoding structure for further modeling in subsequent processing. In fintech business, this representation can be used as the input representation of a speech risk recognition model. In healthcare business, it can be used as the input representation of speech feature analysis and state recognition tasks, thereby realizing the structured representation and subsequent computational processing of speech information.

[0018] In practical implementation, different time-frequency transformation paths can be adopted to adapt to different computing resources and business needs. Audio signals can be segmented using fixed-length windows, with an overlap ratio set between adjacent windows to improve temporal continuity. Alternatively, variable-length windows can be used to dynamically adjust the segmentation granularity based on speech energy distribution to adapt to different speech rates or speech structures. In the frequency domain mapping stage, a frequency domain representation based on Fourier transform can be chosen, or a filter bank transform can be used to obtain a perceptually relevant frequency distribution representation. In the spectral representation output stage, the frequency dimension can be compressed or resampled to reduce data size, and the time dimension can be truncated to meet real-time processing requirements. In fintech applications, the frequency resolution can be adjusted based on the high-frequency noise distribution characteristics in call speech. In healthcare applications, low-frequency information in speech can be enhanced to preserve pronunciation structure features, thereby achieving adaptation and optimization for different business scenarios.

[0019] This embodiment performs time-frequency transformation on the audio signal to be processed and forms a structured time-frequency representation, so that the frequency information and time information in the original time domain signal are separated and expressed, thereby enabling the distinction between the speech subject and noise interference in subsequent processing and improving the modelability of data representation.

[0020] S20, the time-frequency representation is input into the feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension; In this embodiment, inputting the time-frequency representation into the feature encoding network corresponds to mapping the frequency distribution data of the two-dimensional structure to a high-dimensional continuous feature space. The time-frequency representation is in the form of a matrix or tensor with time and frequency dimensions. Each position contains the frequency response information of the corresponding time segment. The input process sends the matrix into the input layer of the feature encoding network through the data tensor interface. The feature encoding network is composed of multiple layers of parameterized mapping units. The layers are connected by weights to form a layer-by-layer feature transformation structure. The input data undergoes local receptive region extraction, nonlinear mapping, and cross-position feature integration in the network, thereby realizing the mapping from the original frequency distribution to the abstract representation.

[0021] The obtained continuous feature representation corresponds to the high-dimensional vector or tensor structure output after performing multi-layer mapping on the time-frequency representation. This continuous feature representation is numerically continuous and can express the combination relationship between different time positions and different frequency regions. The dimension of the continuous feature representation is set by the network structure, including the channel dimension and the time dimension. The channel dimension is used to carry different feature responses, and the time dimension is used to maintain the structural information of speech changes over time. In this process, the input data is linearly transformed by weight parameters and superimposed with nonlinear functions to model the correlation between different frequency regions.

[0022] Dividing a continuous feature representation into multiple sub-feature representations along the feature dimension corresponds to segmenting the continuous feature representation along the channel dimension. The feature dimension refers to the channel dimension or embedding dimension in the continuous feature representation, which carries different feature responses. After dividing this dimension into multiple sub-intervals, each sub-interval forms an independent sub-feature representation. Each sub-feature representation maintains a complete temporal dimension structure while containing only partial channel information. The segmentation process is achieved by performing a splitting operation on the tensor along the specified dimension. The multiple sub-feature representations after the splitting are spatially independent but aligned in the temporal dimension, thus forming the basis for multiple parallel feature branch inputs.

[0023] Feature encoding networks can employ multi-layer convolutional structures to extract local frequency patterns, or networks with attention structures to enhance feature correlation across time locations. They can also combine convolutional and gating structures to jointly model frequency and temporal information. In different implementations, the number of network layers, channels, and nonlinear mapping can be adjusted based on computational resources and application requirements. The dimensionality of continuous feature representations can be configured according to subsequent processing needs; higher dimensions are used to retain more frequency detail information, while lower dimensions are used to reduce computational complexity. During feature dimension partitioning, channels can be evenly distributed, or a non-uniform partitioning method can be used to allocate different numbers of channels to different frequency regions to accommodate the varying importance of different frequency bands. In fintech, this can enhance the expression of frequency band features related to voice identity, while in healthcare, it can strengthen the expression of frequency bands related to voice physiological characteristics.

[0024] This embodiment maps the time-frequency representation to a continuous feature representation and divides it along the feature dimension, so that the original frequency distribution information is reconstructed in a high-dimensional space, and multiple structurally consistent and mutually independent sub-feature representations are formed, thereby improving the fine-grainedness of feature expression.

[0025] S30, input the multiple sub-feature representations into multiple vector quantizers respectively to obtain a degenerate discrete label set, the degenerate discrete label set including multiple sets of degenerate discrete labels; In this embodiment, inputting multiple sub-feature representations into multiple vector quantizers corresponds to performing parallel discretization mapping on the divided multiple sub-feature representations. Each sub-feature representation maintains consistency in the time dimension and is independent of each other in the feature dimension. Inputting into multiple vector quantizers means that each sub-feature representation is fed into an independently configured quantization unit. Each vector quantizer contains a codebook structure, which consists of multiple codeword vectors. Each codeword vector is used to represent a discrete position in the continuous feature space. The input sub-feature representation is compared with the codeword vectors in the codebook through a vector matching mechanism inside the vector quantizer. The comparison process is based on the distance metric between vectors. The distance metric can be Euclidean distance or other metric forms, thereby determining the codeword vector that is closest to the input sub-feature representation.

[0026] The degenerate discrete tag set is obtained by indexing the matching results output by each vector quantizer. Each matched codeword vector corresponds to a unique index in the codebook. This index serves as a discrete identifier to represent the quantization results of continuous features. Multiple vector quantizers output their respective discrete identifiers, which are organized according to the grouping relationship of sub-feature representations to form the degenerate discrete tag set.

[0027] The degenerate discrete label set includes multiple sets of degenerate discrete labels corresponding to the grouping of discrete labels output by each vector quantizer according to atomic feature partitioning relationships. Each set of degenerate discrete labels corresponds to a sub-feature representation of a discrete sequence in the time dimension. The discrete labels within each set maintain a consistent time order, and different sets do not overlap in the feature dimension, thus forming multiple parallel discrete representation sequences. This structure can map continuous features into a representation in discrete symbol space while preserving temporal information.

[0028] Vector quantizers can be implemented using an independent codebook structure. The number of codewords and the dimension of each codebook can be configured according to the feature complexity. In the implementation process, static codebooks or trainable codebooks can be used. Trainable codebooks are optimized through gradient update mechanisms to adapt to the input data distribution. Multiple vector quantizers can share some parameter structures or be completely independent to enhance expressive power. Distance metric calculation can be implemented within the vector quantizer through matrix operations to improve parallel computing efficiency.

[0029] This embodiment inputs multiple sub-feature representations into multiple vector quantizers and converts them into a degenerate discrete label set, so that continuous features can be structurally represented in discrete space, while maintaining the independence between different sub-features, thereby improving the discriminative ability of feature representation.

[0030] S40, extract spectral features from the audio signal to be processed to obtain spectral features, and align the spectral features with the time scale of the degenerate discrete label set to obtain a conditional feature representation; In this embodiment, extracting spectral features from the audio signal to be processed corresponds to converting a continuous time-domain signal into a structured representation containing time and frequency dimensions. The audio signal to be processed is represented as a continuous waveform on the time axis. By segmenting and analyzing the waveform and performing frequency decomposition on each time segment, the distribution of different frequency components at each time position is expressed, thereby forming spectral features. In terms of data structure, the spectral features are represented as a two-dimensional matrix or a three-dimensional tensor, where one dimension represents the time position and the other dimension represents the frequency component. The values ​​reflect the amplitude or correlation characteristics of the corresponding frequency component. This process is achieved by dividing the audio signal into time windows and performing frequency mapping within each time window, so that the periodic and non-periodic structures in speech can be separated and expressed in the frequency domain.

[0031] Time-dimensional alignment of spectral features to match the time scale of the degenerate discrete label set corresponds to adjusting the sampling density of the spectral features in the time dimension to be consistent with the degenerate discrete label set. The degenerate discrete label set consists of multiple discrete label sequences, each with a fixed step size and length in the time dimension. However, the time window length and sliding interval used in the extraction of spectral features may not be consistent with this step size. Therefore, it is necessary to perform time-axis resampling or time segment reorganization of spectral features through time-dimensional alignment so that the sampling position of the spectral features in the time dimension corresponds one-to-one with the degenerate discrete label set. This alignment process is achieved by downsampling, interpolating, or segmenting and aggregating the spectral features, thereby ensuring the consistency of different representations in time position.

[0032] The obtained conditional feature representation corresponds to the spectral feature structure formed after the time dimension alignment is completed. This conditional feature representation is synchronized with the degenerate discrete tag set in the time dimension and retains the original spectral information in the frequency dimension. This representation can provide auxiliary information related to speech content for subsequent processing, so that speech features at different time locations can be used together at a unified time scale.

[0033] Spectral feature extraction can be achieved by dividing the audio signal to be processed into time windows and performing frequency transformation within each time window. The length of the time window and the sliding interval can be adjusted according to the speed of speech change to achieve a balance between time resolution and frequency resolution. Spectral features can be in the form of amplitude correlation features or composite features combined with phase change information to enhance the ability to express speech details. Temporal alignment can be achieved by compressing or expanding the time axis of the spectral features. When the temporal resolution of the spectral features is higher than that of the degenerate discrete label set, the number of time steps can be reduced by downsampling or segmented aggregation. When the temporal resolution of the spectral features is lower than that of the degenerate discrete label set, the number of time steps can be increased by interpolation. In different implementations, a fixed-ratio time scaling or a non-uniform alignment method based on the time position mapping relationship can be selected according to the processing requirements to adapt to the differences in the temporal structure of different speech data.

[0034] In the healthcare sector, voice data contains vocal rhythms and frequency distribution characteristics that vary over time. Frequency structure information is obtained through spectral feature extraction, and this information is aligned with a set of degenerate discrete labels in time through temporal alignment, thereby maintaining a consistent expression of the voice structure in subsequent processing. In the fintech sector, voice interaction data contains voice content and environmental interference information. Frequency distribution characteristics are obtained through spectral feature extraction, and the spectral features are aligned with a set of degenerate discrete labels in time through temporal alignment, thereby enabling the synchronous use of different voice representations in subsequent processing.

[0035] This embodiment extracts spectral features from the audio signal to be processed and performs time-dimensional alignment, so that the spectral features are consistent with the degenerate discrete label set in the time dimension, thereby achieving synchronous expression between different representations and improving the matching degree of subsequent processing.

[0036] S50, embedding processing is performed on the multiple sets of degenerate discrete markers respectively to obtain multiple sets of discrete marker embeddings; In this embodiment, embedding multiple sets of degenerate discrete tags corresponds to mapping discrete index data into continuous vector representations. Each set of degenerate discrete tags consists of multiple discrete tag sequences, with each discrete tag representing an index position in the codebook. This index only indicates a category identifier and does not contain continuous numerical relationships. Embedding converts the discrete tags into continuous features, allowing the relationships between different tags to be represented in the vector space. The embedding process is implemented using an embedding mapping table, which is a parameterized matrix structure. Each row of the matrix corresponds to a discrete tag, and each column corresponds to an embedding dimension. The parameters in the embedding mapping table are used to map the discrete tags into continuous vectors of fixed dimensions.

[0037] Obtaining multiple sets of discrete label embeddings corresponds to performing an index mapping operation on each set of degenerate discrete labels, replacing each discrete label with the corresponding vector in the embedding map table, thus forming a continuous vector sequence. Each set of discrete label embeddings maintains the same order as the corresponding degenerate discrete labels in the time dimension and is composed of the dimensions of the embedding vectors in the feature dimension. Different sets can use different embedding maps, allowing each set to express different types of information in different feature subspaces. The embedding process is implemented through index access; the discrete label serves as the index to directly locate the corresponding vector in the embedding map table, forming a sequence structure in the time dimension, thereby completing the transformation from discrete representation to continuous representation.

[0038] This process involves learning the embedding mapping parameters. The parameters in the embedding mapping table are updated using training data so that the mapped continuous vectors can reflect the similarity between speech features. During training, degenerate discrete labels are used as input indices, and the parameters in the embedding mapping table are adjusted through a backpropagation mechanism so that the embedding vectors corresponding to similar speech features remain close in the vector space, thereby improving the expressive power of subsequent processing.

[0039] Embedding maps can be configured with independent parameter structures for each set of degenerate discrete labels, or with shared parameter structures to reduce model size. During implementation, the embedding dimension can be configured based on computational resources and expression requirements; higher dimensions enhance feature representation capabilities, while lower dimensions reduce computational overhead. Initialization of the embedding map can be random or based on pre-trained data. During training, gradient updates optimize the embedding map, gradually adapting the embedding vector distribution to the speech data features. In fintech, speech interaction data serves as input; embedding processes yield multiple sets of discrete label embeddings that convert discrete speech representations into continuous feature representations, maintaining the separability of speech information in subsequent processing. In healthcare, speech data contains vocal stability and rhythmic variations; embedding processes form continuous vector representations, enabling speech features at different time points to be distinguishable in the vector space, thus supporting stable expression during speech enhancement.

[0040] This embodiment embeds multiple sets of degenerate discrete labels, mapping the discrete identifiers into continuous vector representations, thereby making the discrete representations distinguishable in continuous space and improving the feature representation capability.

[0041] S60, the multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels respectively, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel through the multiple prediction branches to obtain multiple sets of target discrete labels. In this embodiment, inputting multiple sets of discrete label embeddings into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels corresponds to constructing a multi-path parallel processing structure. Each set of discrete label embeddings maintains a sequential form in the time dimension and is a continuous vector representation in the feature dimension. Each prediction branch receives a corresponding set of discrete label embeddings, thereby enabling different sets of features to be modeled in independent paths. The multiple prediction branches are structurally independent of each other, and can adopt the same or different structures in parameter configuration to achieve differentiated modeling of different sub-features.

[0042] Inputting the conditional feature representation into multiple prediction branches corresponds to introducing time-aligned spectral information into each prediction branch. The conditional feature representation is consistent with the degenerate discrete label in the time dimension. By inputting this representation into multiple prediction branches, each prediction branch can perform joint modeling by combining spectral information while processing discrete label embedding. The conditional feature representation is shared and used in multiple prediction branches, thereby providing consistent contextual information in different branches.

[0043] Parallel execution of prediction processing through multiple prediction branches corresponds to performing nonlinear transformations and feature mappings on the input features simultaneously in multiple prediction branches. Each prediction branch contains multiple layers of mapping units, and the layers are connected by parameters to form a feature transformation structure. The input features are modeled in the time dimension, and the input information is combined in the feature dimension to output a predicted feature representation corresponding to the time dimension. The simultaneous operation of multiple prediction branches forms a parallel computing structure.

[0044] Obtaining multiple sets of discrete target labels corresponds to converting the feature representation of each prediction branch output into a discrete label sequence. The output at each time position is mapped to obtain the corresponding discrete label, thus forming multiple sets of discrete target labels. Each set of discrete target labels is consistent with the input features in the time dimension and corresponds to the output results of different prediction branches in structure.

[0045] Prediction branches can be implemented using a network structure containing multiple layers of nonlinear mapping units. Different prediction branches can use the same number of layers and parameter scale, or different structures can be set according to the importance of sub-features. Conditional feature representation and discrete label embedding can be fused within the prediction branches. The fusion method can be achieved through vector concatenation or weighted combination, so that the two types of features are modeled in a unified representation space. Parallel prediction processing is achieved by executing multiple prediction branches simultaneously on a computing device. In the implementation process, parallel computing units can be used to improve processing efficiency. In fintech business, voice interaction data is used as input. Different voice features are processed in parallel through a multi-prediction branch structure, so that the voice representation remains stable in complex interactive environments. In healthcare business, different frequency features in voice data are modeled through different prediction branches, so that the voice structure information is preserved during parallel processing, thereby maintaining consistent feature representation in subsequent voice processing.

[0046] This embodiment improves processing efficiency and maintains consistency of feature representation by embedding multiple sets of discrete labels and conditional feature representations into multiple prediction branches and performing prediction processing in parallel. This allows different features to be modeled simultaneously in independent paths.

[0047] S70, the multiple sets of target discrete markers are mapped to multiple sets of codeword vectors respectively, and the multiple sets of codeword vectors are concatenated into a quantization feature representation. The quantization feature representation is input into the decoding network to obtain the enhanced audio signal.

[0048] In this embodiment, mapping multiple sets of discrete target labels to multiple sets of codeword vectors corresponds to restoring the prediction result in discrete index form to a continuous feature representation. The multiple sets of discrete target labels consist of multiple discrete identifier sequences, each of which has a unique position in the corresponding codebook. Through the mapping operation, the discrete identifiers are used as indexes to access the codeword vectors in the codebook, thereby obtaining a continuous vector corresponding to each discrete identifier. The multiple sets of discrete target labels correspond to different codebook structures, and the temporal order is kept consistent within each set. After mapping, multiple sets of codeword vector sequences are obtained, and each codeword vector carries continuous representation information in the feature dimension.

[0049] Concatenating multiple sets of codeword vectors into a quantized feature representation corresponds to combining different sets of codeword vector sequences in the feature dimension. Multiple sets of codeword vectors maintain alignment in the time dimension. By concatenating them in the feature dimension, vector information from different sets is fused into a unified high-dimensional representation, thus forming a quantized feature representation. The quantized feature representation corresponds to the temporal structure of speech in the time dimension and integrates multiple subspace information in the feature dimension.

[0050] Inputting the quantized feature representation into the decoding network corresponds to feeding the fused continuous feature representation as input data into the decoding structure. The decoding network consists of multiple layers of parameterized mapping units. The input quantized feature representation is transformed layer by layer in the network to realize the mapping from the feature space to the time domain signal space. The network restores the time resolution through multiple layers of nonlinear transformation and upsampling operations, so that the continuous features are gradually restored to the speech signal structure.

[0051] The enhanced audio signal corresponds to the time-domain waveform data output by the decoding network. This output forms a continuous signal in the time dimension and reflects the changes in speech intensity in the amplitude. By decoding the quantization feature representation, the time structure and frequency characteristics of the speech are recovered, thus obtaining the enhanced audio signal.

[0052] The codebook can adopt the same parameter structure as the pre-vector quantization stage. Each codebook contains a fixed number of codeword vectors, and the mapping process is implemented through index access to improve computational efficiency. The concatenation of multiple sets of codeword vectors can adopt a connection method along the feature dimension, so that the information of different groups can be fused at the same time position. Alternatively, the concatenation order of different groups can be set according to the importance of features. The decoding network can be implemented using a multi-layer mapping structure, including a layer-by-layer feature transformation and temporal resolution recovery module. By expanding the dimension and performing nonlinear transformation on the quantized feature representation of the input, the output gradually approximates the target speech waveform.

[0053] This embodiment maps multiple sets of discrete target labels to multiple sets of codeword vectors and concatenates them, enabling the discrete representation to be restored to a continuous feature representation. Then, it generates an enhanced audio signal through a decoding network, thereby realizing the reconstruction of speech information.

[0054] In one embodiment, step S20 above includes: S201, The time-frequency representation is input into the feature encoding network for multi-layer feature extraction to obtain the intermediate feature representation; S202, the intermediate feature representation is aggregated using the feature encoding network to obtain the aggregated feature representation; S203, the aggregated feature representation is mapped through the feature encoding network to obtain a continuous feature representation; S204, determine the feature dimension partitioning position corresponding to the continuous feature representation; S205, the continuous feature representation is divided along the feature dimension according to the feature dimension division position to obtain multiple sub-feature representations.

[0055] In this embodiment, the time-frequency representation is input to the feature encoding network, corresponding to feeding the data tensor, which is a coupled representation of the frequency and time dimensions, into a trainable feature extraction structure. The time-frequency representation can be organized in the form of a two-dimensional matrix or a three-dimensional tensor. The time dimension is used to express the position of speech changes over continuous time, the frequency dimension is used to express the energy distribution of different frequency bands, and the channel dimension is used to accommodate amplitude information, phase-related information, or preprocessed combined information. The feature encoding network includes an input mapping unit, a multi-layer feature extraction unit, a temporal feature aggregation unit, a feature mapping unit, and a dimension partitioning unit. These units are connected sequentially according to the data flow direction. The input mapping unit is responsible for converting the original time-frequency representation to a unified internal representation dimension. The multi-layer feature extraction unit is responsible for extracting the combination patterns of local time regions and local frequency regions. The temporal feature aggregation unit is responsible for integrating cross-frame dependencies along the time axis. The feature mapping unit is responsible for compressing or projecting the aggregated representation onto a preset feature dimension. The dimension partitioning unit is responsible for dividing the output result into multiple independent parts along the feature dimension. The corresponding implementation of this structure can be a combination of convolutional layers, normalization layers, nonlinear activation layers, residual connection layers, temporal convolutional layers, recurrent unit layers, or attention layers. However, regardless of the implementation, the input-output relationship remains consistent: the input is a time-frequency representation, and the output is a continuous feature representation and multiple sub-feature representations.

[0056] Multi-layer feature extraction corresponds to a layer-by-layer mapping of the time-frequency representation within the feature encoding network, progressively transforming lower-level frequency responses into higher-level abstract representations. This processing stage focuses on short-time frequency patterns, harmonic structures, transient changes, and noise distribution characteristics within local regions. If a convolutional structure is used, co-occurrence information in local time-frequency blocks can be extracted by sliding a two-dimensional convolutional kernel along both the time and frequency dimensions. If a residual structure is used, cross-layer identity connections can preserve original details and enhance the stability of deep representations. A combination of normalization and nonlinear activation can control feature amplitude distribution and improve nonlinear expressive power. Multi-layer feature extraction yields intermediate feature representations, which maintain the temporal order of the input time-frequency representation and are transformed into higher-level channel representations after frequency compression. Each time position in the intermediate feature representation no longer corresponds to a single frequency band energy but to a multi-dimensional response after local spatial combination, thus more fully expressing the differences between speech structure and interference components.

[0057] Temporal feature aggregation corresponds to the cross-locational information integration of intermediate feature representations along the time dimension. Speech signals exhibit continuity and dependency between adjacent time positions; relying solely on local time-frequency blocks is insufficient to fully describe the speech's articulation trends, formant variations, speech rate changes, and sustained tone structures. Temporal feature aggregation introduces trainable contextual modeling units in the temporal direction to jointly analyze the feature responses of preceding and following time positions. Using a temporal convolutional structure, dilated convolutions or multi-scale convolutional kernels can expand the temporal perception range; using a bidirectional recurrent structure, information can be accumulated simultaneously in both forward and reverse time directions, allowing each time position to receive both forward and backward context; using an attention structure, long-distance dependency modeling can be achieved through correlation weights between time positions. This processing yields an aggregated feature representation where each time position simultaneously contains local time-frequency patterns and cross-temporal contextual information, while maintaining the temporal order of the data structure, providing a stable input for subsequent mapping to continuous feature representations.

[0058] Feature mapping corresponds to converting aggregated feature representations into continuous feature representations. The channel dimension of aggregated feature representations typically contains rich intermediate responses. The feature mapping unit projects these intermediate responses onto a predefined feature space through linear transformations and nonlinear constraints, giving the output a unified dimension and decomposability. This process can be achieved through fully connected mapping at each time position, or through point convolutional structures to reorganize channels. The output continuous feature representation retains its time dimension, while the feature dimension is set by the mapping target. The dimensions in the continuous feature representation no longer directly correspond to a single frequency band, but rather to a comprehensive expression resulting from the combined effects of multiple frequency bands and multiple temporal contexts. The continuous feature representation uses continuous values, facilitating subsequent similarity matching and vector quantization during discretization.

[0059] Determining the feature dimension partitioning positions corresponding to continuous feature representations involves defining multiple segmentation boundaries along the channel dimension of the continuous feature representation. Feature dimensions carry different types of high-level speech responses, and partitioning positions are used to determine the start and end positions of each feature subspace. Partitioning positions can be set at uniform intervals, ensuring each sub-feature representation has the same dimension; or they can be set according to a preset dimension ratio, allowing different sub-feature representations to correspond to different channel ranges. The setting of partitioning positions is related to the total dimension of the continuous feature representation, the number of target sub-features, and the requirements of subsequent discretization processing. When the continuous feature representation is a matrix of time steps multiplied by the feature dimension, the partitioning positions actually correspond to several segmentation indices along the feature dimension direction; when the continuous feature representation has additional batch or channel dimensions, the partitioning positions still only affect the target feature axis and do not change the arrangement of the time and batch dimensions.

[0060] The continuous feature representation is partitioned along its feature dimension based on the partitioning position. This corresponds to decomposing the continuous feature representation into multiple sub-matrices or sub-tensors according to the aforementioned partitioning index. Each sub-feature representation retains a complete temporal dimension, ensuring that the local representation at any given time point can appear one-to-one in multiple sub-feature representations. The sub-feature representations do not overlap in channel range and numerically cover all feature dimensions of the continuous feature representation. This partitioning process disperses the high-dimensional information originally concentrated in a single continuous feature representation into multiple independent feature subspaces, providing an input foundation for subsequent independent quantization and parallel modeling from a data organization perspective. Multiple sub-feature representations maintain temporal structure consistency while preserving the independence of local feature subspaces, thus simultaneously addressing the requirements of temporal synchronization and feature decoupling.

[0061] In fintech scenarios, input data can be represented as time-frequency representations of telephone banking voice, intelligent customer service voice, or transaction voice commands, while the output data consists of continuous feature representations and multiple sub-feature representations. Since voice input in financial scenarios is often accompanied by channel compression, environmental noise, and vocalization differences, the feature encoding network needs to strengthen the separation and representation of the main voice pattern and noise pattern during multi-layer feature extraction, and maintain the continuity of command words and identity voiceprints during temporal feature aggregation, so that the output can stably represent the effective content of financial voice. In healthcare scenarios, input data can be represented as time-frequency representations of remote health inquiry voice, rehabilitation training voice, or home monitoring voice, and the output data also consists of continuous feature representations and multiple sub-feature representations. In this scenario, the feature encoding network needs to more fully preserve information related to vocal rhythm, formant changes, and continuous vocalization stability, so that the temporal variation patterns can still be distinguished after aggregation, thereby ensuring that subsequent processing stages can be performed based on a stable feature subspace.

[0062] For example, the formula for representing continuous features is:

[0063] in, This represents continuous feature representation. Let D represent the D-dimensional real vector space in which the continuous feature representation resides. D represents the feature dimension of the continuous feature representation.

[0064] The formula for continuous feature representation partitioning is:

[0065] in, This represents continuous feature representation. This represents multiple sub-feature representations obtained by partitioning along the feature dimension. N represents the number of sub-feature representation groups or the number of partitions.

[0066] This embodiment inputs the time-frequency representation into a feature encoding network to perform multi-layer feature extraction, temporal feature aggregation, and feature mapping, thereby obtaining a continuous feature representation that preserves the temporal structure and has high-dimensional continuous expressive power. Then, based on the feature dimension, the continuous feature representation is decomposed into multiple sub-feature representations, so that the frequency response, temporal context information, and high-level abstract information are recombined in a unified representation space, forming multiple independent feature subspaces, thereby improving the decomposability of speech feature representation and the adaptability to subsequent processing.

[0067] In one embodiment, step S30 above includes: S301, the multiple sub-feature representations are input into multiple vector quantizers according to their corresponding relationships; S302, for each vector quantizer, retrieve the codebook corresponding to each vector quantizer, and determine multiple distances between the sub-feature representation input to each vector quantizer and each codeword vector in the corresponding codebook; S303, determine the target codeword vector from the respective corresponding codebooks based on the multiple distances; S304, obtain the indices of the multiple target codeword vectors in their respective codebooks, and determine the multiple indices as multiple sets of degenerate discrete markers; S305, combine the multiple sets of degenerate discrete tags into a set of degenerate discrete tags.

[0068] In this embodiment, multiple sub-feature representations are input into multiple vector quantizers according to their correspondence. This corresponds to allocating an independent discretization processing unit to each sub-feature representation, assuming that the feature dimension has already been divided. The multiple sub-feature representations are consistent in the time dimension and independent in the feature dimension, with each sub-feature representation carrying local feature information with continuously taking values. Multiple vector quantizers receiving different sub-feature representations means that different feature subspaces are mapped using independent discrete representation spaces. This input relationship ensures that different sub-feature representations remain grouped independently during the discretization stage, avoiding mixing of different feature subspaces during quantization. Structurally, the vector quantizer includes an input matching unit, a codebook access unit, a distance calculation unit, and an index output unit. The input matching unit receives the corresponding sub-feature representation, the codebook access unit retrieves the codebook corresponding to the current vector quantizer, the distance calculation unit calculates multiple distances between the input sub-feature representation and each codeword vector in the codebook, and the index output unit provides a discrete identifier based on the optimal matching result. Multiple vector quantizers can use the same internal structure, or they can be configured with different codebook sizes and codeword dimensions according to the dimensions of different sub-feature representations. However, each vector quantizer only processes the corresponding set of inputs and does not change the boundary relationship between sub-feature representations.

[0069] For each vector quantizer, a codebook corresponding to that quantizer is retrieved, corresponding to the creation of an independent discrete candidate set for each input sub-feature representation. The codebook consists of multiple codeword vectors, each located at a predetermined position in the continuous feature space. These multiple codeword vectors collectively cover the main distribution area where a certain sub-feature representation might appear. The codebook is typically organized as a two-dimensional parameter matrix, where the number of rows corresponds to the number of codewords, and the number of columns corresponds to the codeword dimension. The codeword dimension is consistent with the feature dimension of the corresponding sub-feature representation, ensuring that the input vector and codeword vectors are compared in the same vector space. Multiple distances between the sub-feature representation of each input vector quantizer and each codeword vector in the corresponding codebook are obtained one by one through a distance calculation unit. The distances can be calculated using squared Euclidean distance, weighted Euclidean distance, or other continuous vector similarity measures. Each value in the distance set represents the proximity between the current sub-feature representation and a candidate codeword vector. Multiple distances together constitute the matching result of the current input in that codebook. This process transforms the continuous feature representation into a set of explicit matching evaluation results, providing a comparable basis for subsequent discretization selection.

[0070] Target codeword vectors are determined from their respective codebooks based on multiple distances, corresponding to the optimal matching selection performed within each vector quantizer. The value among the multiple distances that best reflects the closeness between the input sub-feature representation and a candidate codeword vector is selected as the decision criterion, and the codeword vector corresponding to this decision criterion is determined as the target codeword vector. The target codeword vector and the input sub-feature representation have the closest representational relationship in the feature space defined by the current codebook; therefore, the target codeword vector can be regarded as a discrete approximation of the current continuous input. The selection of target codeword vectors not only completes the assignment judgment from continuous features to discrete candidates but also simultaneously completes feature compression, because the originally continuous, high-degree-of-freedom representation is restricted to one of a finite number of discrete positions in the codebook. After multiple vector quantizers execute this selection process in parallel, multiple target codeword vectors can be obtained, each of which corresponds one-to-one with its corresponding sub-feature representation, thus maintaining the original feature grouping structure unchanged.

[0071] Multiple target codeword vectors are indexed in their respective codebooks, and these indices are then assigned as multiple sets of degenerate discrete markers. This corresponds to transing the matching results in the continuous space into a symbolic representation. Each target codeword vector has a unique storage location in the codebook, corresponding to a unique index number. This index number indicates the assignment of the current input sub-feature representation in the discrete space. The use of indexes allows subsequent processing to no longer rely on the continuous vectors themselves, but only on discrete identifiers for storage, transmission, or prediction. Multiple target codeword vectors correspond to multiple indices, each index corresponding one-to-one with a corresponding time position and a corresponding sub-feature representation. Therefore, the original arrangement relationship is maintained in the time dimension, and the original division relationship is maintained in the group dimension. After multiple indices are assigned as multiple sets of degenerate discrete markers, each set of degenerate discrete markers constitutes a discrete sequence of the corresponding sub-feature representation over the entire time range. This discrete sequence reflects the discrete expression of the input speech in the corresponding feature subspace.

[0072] Multiple sets of degenerate discrete tags are combined into a degenerate discrete tag set, corresponding to the unified organization of the discrete outputs of multiple vector quantizers. The degenerate discrete tag set contains multiple groups, each containing a sequence of discrete identifiers representing the corresponding sub-features in the time dimension, with different groups maintaining a parallel relationship. The combination process does not change the temporal order of the indices within each group, nor does it alter the independence between groups. Instead, it unifies the multiple groups in the data structure, allowing subsequent processing to access the degenerate discrete tags of each group simultaneously. This set form preserves both the sequence structure in the time dimension and the inter-group boundaries after feature dimension partitioning, thus ensuring that the discretized output possesses both temporal synchronization and group separability. After this processing, continuous high-dimensional features are transformed into a structured discrete tag set, and the speech representation is transformed from a continuous vector space to a discrete symbol space, providing a discrete input foundation for subsequent embedding, parallel prediction, and reconstruction processes.

[0073] Multiple vector quantizers and their corresponding codebooks require parameter optimization using training data. The input to the training data consists of multiple sub-feature representations, and the output is a set of degenerate discrete labels. During the training phase, the codebook parameters of each vector quantizer are updated, with the goal of making the codeword vectors closer to the distribution center of the input sub-feature representations. Batch data input can be used during training, with each sample in the batch containing multiple sub-feature representations at multiple time points. The distance calculation unit compares all inputs within the batch with the codeword vectors in the codebook in parallel, selects the target codeword vector, and then adjusts the codebook parameters using a quantization error term and a codebook update constraint term. The quantization error term constrains the deviation between the target codeword vector and the input sub-feature representations, while the codebook update constraint term controls the stable movement of the codeword vector during training. Multiple vector quantizers can update their respective codebooks independently or jointly with the preceding feature encoding structure in a joint training framework to ensure that the matching relationship between multiple sub-feature representations and their corresponding codebooks gradually converges. The codebook size, codeword dimension, batch size, and parameter update step size can all be configured based on the complexity of the input speech and available computing resources. Codebook size affects discrete representation accuracy and discrete space capacity; codeword dimension affects quantization representation capability; batch size affects training stability; and parameter update step size affects codebook convergence speed.

[0074] When input data comes from telephone customer service voice, remote identity verification voice, or transaction instruction voice in fintech businesses, multiple sub-feature representations reflect the spectral patterns, temporal variations, and local perturbations in the speech. After multiple vector quantizers discretize these sub-feature representations, a well-structured set of degenerate discrete labels can be formed, enabling high-concurrency speech input to be expressed in a unified discrete space. In this scenario, the input data is set as multiple sub-feature representations after time-frequency encoding, and the output data is set as multiple sets of degenerate discrete labels.

[0075] When the input data comes from remote inquiry voice, rehabilitation training voice, or home voice monitoring data in healthcare operations, multiple sub-feature representations reflect the vocal rhythm, frequency distribution, and local anomalous perturbations. After multiple vector quantizers discretize different subspaces, multiple sets of degenerate discrete labels can be formed, compressing complex speech variations into a stable discrete structure. In this scenario, the input data is set as multiple sub-feature representations, and the output data is set as a set of degenerate discrete labels for further processing in subsequent stages.

[0076] For example, the formula for generating degenerate discrete labels is:

[0077] in, This represents the nth group of degenerate discrete labels. This represents the sub-feature representation of the nth group. This represents the k-th codeword vector in the n-th codebook. k represents the index number of the codeword vector in the n-th codebook. This represents the L2 norm distance. This indicates selecting from all candidate codeword vectors... The index of the codeword vector with the smallest distance.

[0078] The formula for the degenerate discrete label set is:

[0079] in, This represents a set of degenerate discrete labels. This represents each group of degenerate discrete tags. N represents the number of groups of degenerate discrete tags.

[0080] This embodiment inputs multiple sub-feature representations into multiple vector quantizers and retrieves the corresponding codebooks. It calculates multiple distances between the input sub-feature representations and each codeword vector, determines the target codeword vector based on the distances, and then obtains the index of the target codeword vector in the corresponding codebook and combines them to form a degenerate discrete tag set. This allows continuous feature representations to be converted into structured discrete representations while maintaining the temporal order and grouping boundaries, thereby improving the compression degree, grouping independence, and adaptability of speech representation to subsequent processing.

[0081] In one embodiment, step S40 above includes: S401, Perform a short-time Fourier transform operation on the audio signal to be processed to extract the amplitude spectrum features and phase correlation features of the audio signal to be processed; S402, the amplitude spectrum features and the phase correlation features are fused to obtain the spectrum features; S403, determine the time scale of the degenerate discrete marker set; S404, Based on the time scale of the degenerate discrete marker set, the spectral features are downsampled and aligned in the time dimension to obtain the target aligned feature sequence; S405, the target alignment feature sequence is input into a bidirectional recurrent neural network to perform long temporal dependency modeling to obtain a temporal dependency feature sequence; S406, the time-dependent feature sequence is input into a network structure containing a self-attention mechanism for global context fusion processing to obtain conditional feature representations.

[0082] In this embodiment, a short-time Fourier transform (SFT) operation is performed on the audio signal to be processed, which corresponds to decomposing a continuous waveform in the time domain into local frequency representations expanded according to time positions. The audio signal to be processed is a sequence of sampling points on the time axis. The SFT operation segments the sequence through a sliding time window and performs frequency decomposition within each time window, thereby obtaining a set of frequency components at each time position. The amplitude spectrum feature is used to represent the energy strength of each frequency component, and the phase correlation feature is used to represent the phase change information of each frequency component. The amplitude spectrum feature reflects the distribution difference between the main speech components and interference components on the frequency axis, and the phase correlation feature reflects the time positioning, transient changes, and phase relationship between frequency components. Both the amplitude spectrum feature and the phase correlation feature are arranged in time dimension order to form a frequency domain representation sequence corresponding to the original audio time sequence. In the short-time Fourier transform operation, the time window length, window shift length, and number of frequency decomposition points jointly determine the time resolution and frequency resolution. A shorter time window length can more finely characterize rapidly changing speech segments, while a longer time window length can more stably express low-frequency structural information. The window shift length determines the degree of overlap between adjacent time positions, and the number of frequency decomposition points determines the degree of refinement of the frequency scale.

[0083] The fusion of amplitude spectrum features and phase correlation features corresponds to integrating two types of frequency domain information into a unified spectral feature representation. Amplitude spectrum features and phase correlation features naturally correspond in time position, and the fusion process can combine the two types of information into a unified feature space at the same time position. This combination process can be achieved through feature dimension concatenation or through linear transformation and weighted mapping, with the output uniformly denoted as spectral features. The spectral features maintain the correspondence between the time and frequency dimensions in the data structure, while incorporating amplitude and phase-related information into the same representation, enabling subsequent processing units to utilize both energy distribution and phase change information simultaneously. Spectral features are no longer limited to a single energy expression but form a composite representation containing multi-dimensional frequency domain information.

[0084] Determining the time scale of the degenerate discrete marker set corresponds to extracting sampling granularity information in the time dimension from the discrete representation output. The degenerate discrete marker set consists of multiple sets of degenerate discrete markers, each with a uniform time step and a uniform time length in the time direction. This time step reflects the sampling interval of the discrete representation on the time axis. Spectral features are obtained from the short-time Fourier transform, and their time step is determined by the window shift length; the two are usually inconsistent. Determining the time scale of the degenerate discrete marker set means extracting the time step, time length, or number of time positions on the discrete representation side as a reference for subsequent alignment processing. This time scale can be represented as the target number of time frames, the discrete position interval, or the time axis scaling ratio.

[0085] Based on the time scale of the degenerate discrete label set, spectral features are downsampled and aligned in the time dimension, corresponding to compressing or rearranging the temporal resolution of the spectral features to match the degenerate discrete label set. The time dimension of spectral features is typically denser than that of the degenerate discrete label set, thus requiring sampling point compression along the time axis. Downsampling reduces the number of time positions to the target scale by aggregating, extracting, or compressing the spectral features at adjacent time positions with convolutional strides. Time alignment corrects the position of the compressed spectral features on the time axis, ensuring that each time position corresponds to its corresponding position in the degenerate discrete label set. Downsampling addresses the mismatch in the number of time positions, while time alignment addresses the inconsistency in time position mapping; both output a target-aligned feature sequence. This target-aligned feature sequence corresponds one-to-one with the degenerate discrete label set in the time dimension, while retaining spectral information in the feature dimension.

[0086] The target-aligned feature sequence is input into a bidirectional recurrent neural network (NRNN) for long-term temporal dependency modeling, which corresponds to introducing forward and backward contextual information in the time dimension. The spectral information at a single time position in the target-aligned feature sequence can only reflect the acoustic state at a local moment. Speech enhancement requires utilizing continuous changes over adjacent time positions and even longer time ranges; therefore, a bidirectional recurrent neural network is used to model the target-aligned feature sequence. The bidirectional recurrent neural network includes forward recurrent units and backward recurrent units. The forward recurrent units process the target-aligned feature sequence in chronological order, while the backward recurrent units process it in reverse chronological order. The hidden states in both directions are concatenated or mapped at the same time position to form a temporal dependency feature sequence. Each time position in the temporal dependency feature sequence contains not only the local spectral features of that position but also contextual information from the preceding and following time positions, thereby enhancing the ability to describe continuous tones, transitional tones, transient interference, and long-term noise changes. Structurally, the bidirectional recurrent neural network includes an input layer, a bidirectional recurrent layer, and an output mapping layer. The input layer receives the target-aligned feature sequence, the bidirectional recurrent layer performs temporal dependency modeling, and the output mapping layer combines the bidirectional states into a unified dimensional representation. The number of recurrent layers, the number of hidden units, and the state dimension together determine the temporal modeling capability. A deeper layer can express more complex temporal correlations, and a larger number of hidden units can retain richer contextual information.

[0087] The temporally dependent feature sequence is input into a network structure containing a self-attention mechanism for global context fusion processing, which corresponds to further establishing the correlation between non-local locations across the entire time range. Bidirectional recurrent neural networks can model sequential temporal dependencies, but their representation of relationships between distant temporal locations is affected by the length of the state propagation path. The network structure containing a self-attention mechanism calculates the correlation weight between any two time locations in the temporally dependent feature sequence, and forms a new context representation by weighted aggregation of features from all time locations. This structure includes a query mapping unit, a key mapping unit, a value mapping unit, and an attention output unit. After the temporally dependent feature sequence undergoes three types of mapping, it forms a query vector, a key vector, and a value vector. The correlation between the query vector and the key vector is used to calculate the attention weight, which is applied to the value vector to output the global context fusion result. This processing yields a conditional feature representation. The conditional feature representation maintains consistency with the target-aligned feature sequence in the temporal dimension and integrates local spectral information, bidirectional temporal information, and global contextual information in the feature dimension, thus serving as a unified conditional input for subsequent processing stages.

[0088] Bidirectional recurrent neural networks (NRNNs) and network structures with self-attention mechanisms can update parameters through training. The input data consists of spectral features obtained from the audio signal to be processed through short-time Fourier transform and fusion processing, which are then aligned in the time dimension to form a target-aligned feature sequence. The output data is the conditional feature representation. During training, the recurrent weight matrix, state transition parameters, and output mapping parameters in the NRNN participate in the update, as do the query mapping parameters, key mapping parameters, value mapping parameters, and output mapping parameters in the self-attention mechanism. The training data uses a training set consisting of corresponding samples of degraded and target speech. The input is the audio signal to be processed, and the conditional feature representation's parameters are adjusted during overall training through backpropagation of errors between subsequent discrete label prediction results and audio reconstruction results. The number of layers in the NRNN can be set to single or multiple layers, and the hidden state dimension is set according to the spectral feature dimension and computing resources. The network structure with self-attention mechanism can be set to single-head or multi-head attention, and the number of heads, attention dimension, and output mapping dimension in multi-head attention are set according to the target feature dimension. The batch size, parameter update step size, and number of training rounds are configured according to the data scale and stability requirements.

[0089] When input data comes from intelligent customer service voice, telephone transaction voice, or remote verification voice in fintech business, the input data is set as the spectral features of the audio signal to be processed, and the output data is set as a conditional feature representation consistent with the discrete representation time scale. Voice data in fintech scenarios is often accompanied by bandwidth constraints, call compression, and background interference. The spectral features contain both effective speech components and noise and channel distortion information. By fusion through time dimension alignment, bidirectional recurrent neural network modeling, and a self-attention mechanism, the conditional feature representation can express the speech context relationship across time locations at a unified time scale, thereby ensuring stable representation of effective speech components in subsequent processing stages within complex business environments.

[0090] When the input data comes from remote inquiry voice, home monitoring voice, or rehabilitation training voice in healthcare applications, the input data is also set as the spectral features of the audio signal to be processed, and the output data is set as a conditional feature representation consistent with the time scale of the discrete representation. Voice data in healthcare scenarios often contains variations in vocal rhythm, breathing interference, and environmental reverberation. The target-aligned feature sequence reflects continuous vocal changes in the time dimension, and the self-attention mechanism further aggregates contextual information over a long period, enabling the conditional feature representation to preserve the continuity of the speech structure and providing a unified conditional input for subsequent processing.

[0091] For example, the formula for generating conditional feature representations is:

[0092] in, This represents the conditional feature representation. () represents the conditional feature extraction mapping function for degraded speech, corresponding to the entire process of spectral feature extraction, time alignment, long-term temporal dependency modeling, and global context fusion. y represents the degraded speech signal or the audio signal to be processed.

[0093] This embodiment extracts amplitude spectrum features and phase correlation features by performing a short-time Fourier transform operation on the audio signal to be processed, and then fuses the two types of features to obtain spectral features. Next, it performs downsampling and time dimension alignment on the spectral features according to the time scale of the degenerate discrete label set, so that the spectral information is consistent with the discrete representation in the time dimension. At the same time, it performs long-term time-series dependency modeling on the target aligned feature sequence through a bidirectional recurrent neural network, and then performs global context fusion through a network structure containing a self-attention mechanism, to obtain a conditional feature representation that simultaneously contains local spectral information, bidirectional time context information, and global time-related information, thereby improving the time matching degree and contextual expression ability of the conditional input.

[0094] In one embodiment, step S50 above includes: S501, for each group of degenerate discrete tags in the multiple groups of degenerate discrete tags, retrieve the corresponding embedding mapping table; S502, based on each degenerate discrete tag in each group of degenerate discrete tags, read the embedding vector from the corresponding embedding mapping table; S503, Arrange the read embedding vectors sequentially according to the tag arrangement order within each group of degenerate discrete tags; S504, perform intra-group assembly processing on the sequentially arranged embedding vectors to obtain each group of discrete label embeddings; S505, combine the discrete label embeddings of each group to obtain multiple groups of discrete label embeddings.

[0095] In this embodiment, embedding processing is performed on multiple sets of degenerate discrete labels, which corresponds to converting the discrete index sequence into a vector sequence that can participate in continuous operations. The multiple sets of degenerate discrete labels consist of multiple grouped discrete identifiers. Each discrete identifier corresponds to an index position in the quantization output of the previous stage. The index position only expresses the category affiliation and does not directly possess proximity relationships in continuous space, nor does it directly carry amplitude differences or directional relationships. The task of the embedding processing is to map each discrete identifier into a continuous vector of fixed dimensions, enabling the discrete representation to participate in subsequent feature fusion and prediction calculations in vector form. This processing revolves around an embedding mapping table, which is stored in matrix form. The row index of the matrix corresponds to the discrete identifier number, and the column dimension corresponds to the embedding dimension. Each row vector represents the expression result of a discrete identifier in continuous feature space. Each set of degenerate discrete labels can correspond to an independent embedding mapping table, or a shared mapping table can be used provided that the expression requirements are met. However, in the processing semantics of the current step, retrieving the corresponding embedding mapping table means that each set of degenerate discrete labels maintains a fixed correspondence with the corresponding parameterized mapping structure, thereby maintaining the independent expressive ability of each set of discrete representations in continuous space.

[0096] Based on each degenerate discrete label in each group, embedding vectors are read from the corresponding embedding map table, corresponding to an index access operation. Each degenerate discrete label serves as an index value to directly locate a row vector in the embedding map table, and the result is an embedding vector uniquely corresponding to that discrete label. The reading process unfolds position by position in the time dimension, so a group of degenerate discrete labels will generate a set of embedding vectors aligned with the time position. The dimension of the embedding vectors is determined by the number of columns in the embedding map table. Embedding vectors read from different time positions maintain a consistent dimension, allowing subsequent intra-group assembly processing to be performed under a unified tensor structure. This process realizes a single-point mapping from symbol space to continuous space, with discrete numbers as input and fixed-length continuous vectors as output. Since multiple groups of degenerate discrete labels are read separately, each set of embedding vectors is independent in origin and corresponds position by position to the original discrete label sequence in the time dimension, thus preserving the temporal order of the discrete representation while completing the construction of continuous features.

[0097] The read embedding vectors are sequentially arranged according to the label arrangement order within each group of degenerate discrete labels, corresponding to restoring the arrangement structure consistent with the degenerate discrete labels on the time axis. The degenerate discrete labels were already organized into a discrete sequence according to their time positions during the generation stage. After reading the embedding vectors, it is necessary to ensure that each embedding vector is still arranged in the same time order during output. The purpose of sequential arrangement is not to change the value of individual embedding vectors, but to maintain the positional constraints of the vector sequence in the time dimension, preventing vectors at different time positions from being swapped, misaligned, or unordered. After sequential arrangement, each group of read results is no longer an unordered set of vectors, but an ordered set of embedding vectors with time indexing significance. The resulting data can be represented as a matrix of time position number multiplied by the embedding dimension, or as a three-dimensional tensor with batch dimension, where the time dimension corresponds to the original arrangement order of the degenerate discrete labels, and the feature dimension corresponds to the embedding vector length.

[0098] The sequentially arranged embedding vectors are subjected to intra-group assembly processing to obtain each group of discrete label embeddings. This corresponds to organizing embedding vectors belonging to different time positions within the same group into a unified continuous representation unit. Intra-group assembly can be represented as stacking vectors along the time axis to form a matrix structure, or as concatenating embedding vectors from multiple time positions into a tensor structure. The processing results all point to the same goal: transforming the scattered embedding vectors within the same group into a set of discrete label embeddings that can be input as a whole into subsequent computational units without disrupting the temporal order. Each set of discrete label embeddings fully preserves the sequential structure of the degenerate discrete labels in the time dimension and fully preserves the embedding mapping results in the feature dimension. Therefore, it possesses both sequence attributes and vector space representation attributes. Intra-group assembly does not introduce new discrete labels or change the embedding vectors corresponding to each position; it only unifies the data organization, making embedding vectors from different time positions a unified input object.

[0099] The process of aggregating the discrete label embeddings of each group results in multiple sets of discrete label embeddings, corresponding to a unified cross-group organization of multiple grouping results. Each set of discrete label embeddings originates from different degenerate discrete label groups. These groups typically maintain the same length in the time dimension and can maintain the same embedding dimension in the feature dimension. Alternatively, they can be formed with different dimensions if the mapping dimension configuration allows, and then aligned through a unified mapping. The resulting multiple sets of discrete label embeddings preserve inter-group boundaries while enabling multiple grouping results to be accessed in parallel by subsequent units as a whole. The aggregated data structure can be represented as a three-dimensional tensor equal to the number of groups multiplied by the number of time positions multiplied by the embedding dimension, or as a set of tensors within multiple groups. The specific organization method is determined by the requirements of the subsequent input interface. The output result, multiple sets of discrete label embeddings, corresponds semantically to the grouped representation of multiple sets of degenerate discrete labels in a continuous vector space, completing the grouped transformation from discrete index sequences to continuous embedding sequences.

[0100] The parameters in the embedding map can be optimized through sample data rather than being manually fixed. During training, the input data consists of multiple sets of degenerate discrete labels, and the output data consists of multiple sets of discrete label embeddings. The embedding map participates in the overall network training as a parameter matrix, and parameter updates are achieved through backpropagation and gradient descent. During training, each discrete label triggers an access to the corresponding row vector in the embedding map. Errors generated by subsequent units are propagated back along the computation graph to the corresponding row vector, allowing discrete labels that frequently co-occur in the training samples and play similar roles in enhanced speech reconstruction to form a more reasonable relative positional relationship in continuous space. The parameter size of the embedding map is determined by the number of discrete labels and the embedding dimension. The number of discrete labels corresponds to the number of rows in the map, and the embedding dimension corresponds to the number of columns in the map. Increasing the embedding dimension enhances the vector expressive power, but also increases computational and storage overhead; decreasing the embedding dimension increases compression, but may decrease discriminability. If each set of degenerate discrete labels uses an independent embedding map, the parameters of each set are updated independently during training; if a shared embedding map is used, discrete labels from different sets participate in the same parameter matrix update.

[0101] This embodiment retrieves the corresponding embedding mapping table for each group of degenerate discrete markers, reads the embedding vector according to each degenerate discrete marker in each group, and completes intra-group assembly and cross-group aggregation processing while maintaining the marker arrangement order. This transforms the discrete index-based grouped representation into a multi-group discrete marker embedding with temporal order and continuous feature dimensions, thereby improving the expressive power of discrete representation in continuous space, maintaining the consistency of group boundaries and temporal positions, and enhancing the adaptability of subsequent processing to discrete speech representation.

[0102] In one embodiment, step S60 above includes: S601, the multiple sets of discrete label embeddings are input into multiple prediction branches according to their correspondence with the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches; S602, In the multiple prediction branches, the discrete label embedding and conditional feature representation of the input are subjected to feature fusion processing to obtain the branch input feature sequence; S603, Sequence modeling is performed on the input feature sequences of each of the multiple prediction branches to obtain the branch prediction feature sequences; S604, perform label mapping on the predicted feature sequences of each branch to obtain the labeled output sequence corresponding to each predicted branch; S605, determine multiple sets of target discrete labels based on the label output sequence corresponding to each prediction branch.

[0103] In this embodiment, multiple sets of discrete label embeddings are input into multiple prediction branches according to their correspondence with multiple sets of degenerate discrete labels. This corresponds to assigning an independent prediction path to each set of discrete label embeddings, given that the grouped discrete representations have already been established. The multiple sets of discrete label embeddings maintain the same arrangement order as their respective degenerate discrete labels in the time dimension and are continuous vector representations in the feature dimension. Different groups are independent of each other in terms of their source. The one-to-one correspondence between multiple prediction branches and multiple sets of degenerate discrete labels means that each prediction branch only processes one set of discrete label embeddings and does not undertake the task of predicting the discrete representations of other groups. This input relationship ensures that each prediction branch only focuses on the discrete representations in its corresponding subspace during parameter learning and feature transformation, avoiding the mixing of information from different groups in the same path. Inputting conditional feature representations into multiple prediction branches corresponds to simultaneously providing the shared auxiliary representations obtained from spectral information and time alignment processing to each prediction branch. The conditional feature representations are synchronized with the discrete label embeddings in the time dimension, and therefore can be used as contextual constraint information for parallel prediction at each time position. Discrete label embedding is a group-specific input, while conditional feature representation is a branch-shared input. These two types of inputs differ in function: the former expresses the local content of each group's discrete representation, while the latter expresses consistent time-frequency context information across groups. This input method ensures that different prediction branches can accept unified conditional constraints while maintaining group independence.

[0104] In multiple prediction branches, the discrete label embeddings and conditional feature representations of the input are fused to obtain the branch input feature sequence. This corresponds to combining group-specific information and shared context information into a unified representation within each prediction branch. Feature fusion can be performed point-by-point at each time location, so that the discrete label embedding and the corresponding conditional feature representation at each time location form a joint vector. The fused branch input feature sequence retains the input order in the time dimension and contains both discrete label embedding information and conditional feature representation information in the feature dimension. Feature fusion can be implemented using feature concatenation, weighted summation, gated modulation, or joint mapping after linear projection. When using feature concatenation, the two types of inputs are directly connected in the feature dimension, allowing subsequent units to learn the information combination relationship on their own. When using weighted summation, the two types of inputs are first mapped to a unified dimension and then fused according to weights. When using gated modulation, the channel weights or time position weights of the discrete label embeddings are controlled by the conditional feature representation. When using joint mapping after linear projection, the two types of inputs are first linearly transformed separately and then a unified representation is formed through joint projection. Regardless of the fusion method used, the output is recorded as the branch input feature sequence, and this sequence serves as the direct input to the subsequent modeling unit in each prediction branch.

[0105] Sequence modeling is performed on the input feature sequences of each prediction branch to obtain the branch prediction feature sequences, corresponding to modeling the temporal dependencies in each independent path. The branch input feature sequences contain the fusion results of the same set of discrete label embeddings and shared conditional feature representations; however, before modeling, only local correspondences are shown between time positions, and the continuous change patterns across time are not explicitly expressed. Sequence modeling performs context propagation and state updates along the temporal dimension on the branch input feature sequences within each prediction branch, ensuring that the output at each time position includes not only information from the current input position but also relevant information from preceding and following time positions. This processing can be implemented using temporal convolutional units, recurrent units, bidirectional recurrent units, attention units, or combinations thereof. When using temporal convolutional units, the branch input feature sequences undergo convolution operations within a local time window to extract short-term dependencies. When using recurrent units, the state vector is continuously updated over time, ensuring the current output includes accumulated historical information. When using bidirectional recurrent units, the state is propagated simultaneously in both the forward and reverse time directions, providing bidirectional context for each time position. When using attention units, the correlations at different positions across the entire time range are weighted, enhancing the ability to model long-distance dependencies. After sequence modeling, the branch prediction feature sequence is obtained. This sequence retains its original time length but forms a high-level representation for target-oriented discrete label prediction at each time position.

[0106] Label mapping is performed on the predicted feature sequences of each branch to obtain the label output sequence corresponding to each prediction branch. This corresponds to converting continuous high-dimensional predicted features into discriminative representations in a discrete label space. Label mapping is usually performed independently at each time point, mapping each time point vector in the branch predicted feature sequence to the category space of the corresponding discrete label set. The mapping process can be implemented through linear mapping layers, nonlinear mapping layers, or multilayer perceptron structures. The output result can be represented as the response value or probability value of each candidate discrete label at each time point. The label output sequence maintains the same time length as the branch predicted feature sequence in structure, but transforms from a continuous feature sequence into a category prediction sequence in content, with each time point corresponding to a discrete category discrimination result. The label output sequences of different prediction branches correspond to different groups of candidate discrete label spaces. Therefore, the output mapping parameters of each branch can be set independently, or a shared mapping form can be used when the number of categories is the same.

[0107] Multiple sets of discrete target labels are determined based on the label output sequences corresponding to each prediction branch, corresponding to discrete decision-making based on the output results at the current time position within each prediction branch. If the label output sequence represents class response values, the target discrete label at the current time position is determined by comparing the response magnitudes of different classes; if the label output sequence represents class probability values, the target discrete label is determined by selecting the class with the highest probability value. The determination operation is performed position-by-position in the time dimension and in parallel in the branch dimension, so each prediction branch can simultaneously form the corresponding target discrete label sequence. After multiple prediction branches complete the discrete decision-making for all time positions, the output results are aggregated into multiple sets of discrete target labels according to the branch dimension. The multiple sets of discrete target labels maintain a correspondence with the input discrete label embedding and conditional feature representation in the time dimension, and remain consistent with the original grouping of the degenerate discrete labels in the group dimension, thereby achieving parallel prediction transformation from multiple sets of degenerate discrete labels to multiple sets of discrete target labels.

[0108] Multiple prediction branches, feature fusion processing units, sequence modeling units, and label mapping units can all be updated with parameters using training data. The model structure can include branch input interfaces, fusion units, sequence modeling units, and output mapping units. Multiple prediction branches are connected in parallel. Conditional feature representations are fed into each prediction branch through a shared input interface, and multiple sets of discrete label embeddings are input into the corresponding prediction branches through their dedicated interfaces. During training, the input data consists of multiple sets of discrete label embeddings and conditional feature representations, and the output supervision data consists of reference sequences of multiple sets of target discrete labels. Each prediction branch outputs the prediction result in the corresponding class space at each time position. The training error is calculated from the difference between the prediction result and the reference target discrete label, and is commonly expressed as the cumulative classification error for each time position. The training errors of multiple prediction branches can be calculated separately and then summed as the basis for overall parameter updates. Model parameter updates are completed through backpropagation and gradient descent, with parameters of the fusion unit, sequence modeling unit, and output mapping unit all participating in the update. The number of branches is consistent with the number of multiple sets of degenerate discrete labels. The number of layers, hidden dimensions, convolutional kernel length, recurrent state dimensions, or number of attention heads in the sequence modeling unit are set according to the speech complexity and computing resources. If multiple prediction branches use the same structure, the isomorphic network can be replicated in terms of structure and the parameters can be maintained separately. If the processing difficulty of multiple prediction branches is different, different layers and different dimensions can be set so that different branches can achieve differentiated modeling for their respective input features.

[0109] For example, the formula for the branch prediction probability distribution is:

[0110] in, Let represent the target discrete label probability distribution of the output of the nth prediction branch. () represents the normalization function, used to map the prediction results to a probability distribution. () represents the neural network mapping of the nth prediction branch. Let represent the input discrete label representation corresponding to the nth group of degraded speech, which can be embedded into the nth group of discrete labels. s represents the conditional feature representation.

[0111] The formula for determining the discrete target marker is:

[0112] in, This represents the discrete label of the nth target group. This represents the output probability of the nth prediction branch on the kth candidate discrete label. This indicates the index of the candidate discrete label with the highest probability. k represents the candidate discrete label number.

[0113] After multiple prediction branches complete feature fusion, sequence modeling, and label mapping, each prediction branch outputs a probability distribution in the corresponding discrete label space at each time position. To enable multiple prediction branches to learn the mapping relationship between degenerate discrete labels and target discrete labels, a joint supervision approach is used for training. During training, the real discrete labels corresponding to the clean audio signal are used as the supervision target. The cross-entropy loss is calculated for the probability distribution output by each prediction branch, and the losses of all prediction branches are accumulated to form a multi-branch training loss function. The multi-branch training loss function can be expressed as:

[0114] in, This represents the multi-branch training loss function. N represents the number of prediction branches, and also the number of groups of discrete target labels. This represents the cross-entropy loss function, used to measure the difference between the probability distribution of the predicted branch output and the true discrete label. This represents the target discrete label probability distribution output by the nth prediction branch. This probability distribution is obtained by the nth prediction branch after label mapping, and each component in the probability distribution represents the predicted probability of a certain candidate discrete label corresponding to the current position. This represents the nth true discrete marker corresponding to the clean audio signal. This true discrete marker can be obtained from the clean audio signal through the same encoding and vector quantization process as the degraded audio signal. n represents the prediction branch number or discrete marker group number, ranging from 1 to N. The probability distribution of the nth prediction branch output is as follows: The nth set of real discrete labels corresponding to the clean audio signal is The cross-entropy loss function is denoted as By summing the cross-entropy losses of all prediction branches, the outputs of different prediction branches can be constrained simultaneously, allowing each prediction branch to learn the discrete label prediction relationship for its corresponding group while maintaining a parallel structure. During training, the parameters of multiple prediction branches, discrete label embeddings, network parameters of the conditional feature extraction part, and parameters of the feature fusion part can be jointly updated, so that the label probability distribution output by each prediction branch gradually approximates the true discrete label distribution.

[0115] This embodiment embeds multiple sets of discrete labels and inputs them into multiple prediction branches according to their correspondence with multiple sets of degenerate discrete labels. It also inputs conditional feature representations into multiple prediction branches, enabling the group-specific inputs and shared context inputs to complete joint expression within each prediction branch. Then, through feature fusion processing, sequence modeling, and label mapping, the corresponding label output sequences are simultaneously output in multiple prediction branches, and multiple sets of target discrete labels are determined respectively. This allows different sets of discrete representations to complete predictions in parallel in independent paths, thereby improving the parallel processing capability of the multi-set discrete label prediction process, maintaining the consistency of group correspondence and time dimension, and enhancing the context constraint capability in the target discrete label generation process.

[0116] In one embodiment, step S70 above includes: S701, for each group of target discrete tags in the plurality of target discrete tags, determine the corresponding codebook; S702, based on each target discrete label in each group of target discrete labels, read the corresponding codeword vector from the corresponding codebook to obtain multiple groups of codeword vectors; S703, the multiple sets of codeword vectors are sequentially arranged according to the inter-group arrangement order and the intra-group arrangement order of the multiple sets of target discrete tags; S704, concatenates multiple sets of codeword vectors arranged in order to obtain quantized feature representation; S705, the quantized feature representation is input into the decoding network, and the quantized feature representation is subjected to time-series recovery processing through the decoding network to obtain the decoded features; S706, The decoding features are reconstructed using the decoding network to obtain an enhanced audio signal.

[0117] In this embodiment, a corresponding codebook is determined for each group of discrete target labels, corresponding to establishing a continuous vector recovery basis for each group of discrete prediction results. The multiple groups of discrete target labels originate from the output results of multiple prediction branches. Each group of discrete target labels forms a discrete identifier sequence in the time dimension, and each discrete identifier corresponds to a category number. The codebook stores the mapping relationship between discrete identifiers and continuous representations. Each codeword vector in the codebook is consistent with the continuous feature subspace of the corresponding group in terms of dimension, and corresponds one-to-one with the category number of the discrete target label in terms of index. Determining the corresponding codebook does not involve regenerating codebook parameters, but rather determining the codebook entity corresponding to each group of discrete target labels from the existing parameter set, enabling subsequent reading operations to be performed under a fixed mapping relationship. Different groups of discrete target labels correspond to different codebooks, which can maintain the independence of each group of feature subspaces and prevent aliasing of different groups of discrete representations when recovering continuous features. The codebook can be stored in matrix form, with the matrix row index corresponding to the discrete category number and the matrix column dimension corresponding to the codeword vector dimension. Each row vector in the matrix represents the recovered representation of that discrete category in the continuous feature space.

[0118] Based on each target discrete marker in each group, the corresponding codeword vector is read from the corresponding codebook, resulting in multiple sets of codeword vectors. This corresponds to performing a position-wise mapping from discrete symbols to continuous vectors. Each target discrete marker serves as an index to access a storage location in the corresponding codebook, and the access result is a codeword vector with continuously taking values. This read operation is expanded position-wise along the time dimension, so each group of target discrete markers will generate a sequence of codeword vectors corresponding to the time position. The dimension of the codeword vectors is pre-defined by the codebook and is consistent with the dimension of the continuous subspace of that group before discretization. Therefore, the multiple sets of codeword vectors read maintain the same dimension within the group and preserve concatenation between groups. This process restores the discrete prediction result to a vector representation that can participate in continuous operations, while maintaining the correspondence between the time position and the discrete marker, ensuring that subsequent sequential arrangement and concatenation processing are based on positional alignment.

[0119] The codeword vectors are sequentially arranged according to the inter-group and intra-group arrangement order of the discrete target labels, corresponding to restoring the original data organization structure of the read continuous vectors. The inter-group arrangement order originates from the grouping order formed when dividing the continuous feature representation into feature dimensions during the feature encoding stage, while the intra-group label arrangement order originates from the discrete position order of each group of discrete target labels on the time axis. The sequential arrangement process is performed simultaneously in two dimensions: on the one hand, it maintains the sequential relationship of each time position within the same group, and on the other hand, it maintains the sequential relationship between different groups consistent with the original sub-feature division order. After the sequential arrangement is completed, the multiple codeword vectors are no longer a collection of scattered vectors, but form an ordered tensor structure with clear indexing significance in both the time and group dimensions. This structure directly determines the correctness of the subsequent concatenation process; if the order relationship changes, the concatenated quantized feature representation will not be consistent with the input distribution learned during the decoding network training stage.

[0120] Multiple sets of codeword vectors arranged in sequence are concatenated to obtain a quantized feature representation, which corresponds to merging the recovery results of continuous subspaces from different groups into a complete continuous feature representation. The concatenation process is performed along the feature dimension, connecting codeword vectors from different groups at the same time position sequentially along the feature axis, thus recovering a high-dimensional representation while maintaining the time dimension. Each time position vector in the quantized feature representation is composed of codeword vectors from multiple groups, and its dimension is equal to the sum of the dimensions of the codeword vectors from each group. The concatenated quantized feature representation maintains the same length as the target discrete label in the time dimension and covers each sub-feature subspace in the feature dimension. Therefore, the quantized feature representation is essentially a holistic recovery result of the discrete prediction result in continuous space. This representation simultaneously preserves the decomposition characteristics brought about by independent recovery between groups and the holistic nature brought about by feature dimension recombination, enabling the subsequent decoding network to receive data structures consistent with those in the training phase.

[0121] The quantized feature representation is input into the decoding network, which performs temporal recovery processing on it to obtain the decoded features. This corresponds to converting the discretely recovered high-dimensional continuous representation into an intermediate representation suitable for waveform reconstruction. The decoding network includes an input mapping unit, a temporal recovery unit, and an output transformation unit. The input mapping unit receives the quantized feature representation and projects it onto the hidden dimension used within the decoding network. The temporal recovery unit expands and reconstructs the feature resolution along the time dimension, gradually restoring the compressed temporal information in the quantized feature representation into a temporal structure suitable for audio output. The output transformation unit converts the restored high-dimensional temporal representation into intermediate features that are closer to the waveform space or spectral reconstruction space. Temporal recovery processing can be implemented through deconvolution units, interpolation upsampling units, subpixel rearrangement units, temporal convolution units, or cyclic state propagation units. The input quantized feature representation undergoes time length expansion, channel mapping, and nonlinear transformation in this unit to obtain the decoded features. The decoded features are more refined in the time dimension than the quantized feature representation and retain the reconstruction information for waveform generation in the feature dimension. Therefore, they belong to the intermediate representation within the decoding network used to connect to the final output.

[0122] The decoded features are reconstructed using a decoding network to obtain an enhanced audio signal, which corresponds to converting the intermediate reconstructed representation into the final time-domain audio output. The waveform reconstruction process can be implemented in different ways depending on the output format of the decoding network. If the decoding network output is a time-domain frame feature, a continuous waveform is generated through linear mapping and overlapping addition; if the decoding network output is a spectral reconstruction feature, the frequency domain representation is restored to a time-domain waveform through inverse transformation; if the decoding network has directly modeled time-domain samples, the enhanced audio signal is generated through sample-by-sample mapping. The enhanced audio signal is a continuous time series, reflecting the enhanced speech waveform in the time dimension and the reconstructed acoustic intensity changes in the amplitude dimension. Since the decoding input is a quantized feature representation recovered from multiple sets of discrete target markers, the enhanced audio signal has a traceable correspondence with the aforementioned discrete prediction results, thus completing the entire recovery process from discrete marker prediction results to continuous speech output.

[0123] For example, the formula for enhancing audio signal reconstruction is:

[0124] in, This represents the enhanced audio signal at time t. () indicates the decoding network. This represents the quantized feature representation obtained by mapping and concatenating multiple sets of discrete target labels. t represents the time variable.

[0125] This embodiment determines the corresponding codebook for each set of target discrete markers, and reads the corresponding codeword vector from the codebook according to each target discrete marker, so that the discrete prediction result is restored into multiple sets of continuous vector representations. Then, the multiple sets of codeword vectors are arranged sequentially according to the inter-group arrangement order and the intra-group marker arrangement order, and a quantized feature representation is formed by splicing, so that the grouped restoration results are recombined in a unified feature space. The quantized feature representation is then input into the decoding network, and the decoding features are obtained through time-series restoration processing. The enhanced audio signal is obtained through waveform reconstruction processing, so that the discrete prediction result can be stably converted into continuous speech output, thereby improving the integrity and reconstruction capability of the audio signal restored from the discrete representation.

[0126] In one embodiment, an audio processing apparatus based on discrete label prediction is provided, which corresponds one-to-one with the audio processing method based on discrete label prediction described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the audio processing device based on discrete label prediction of the present invention. The modules include an audio time-frequency transformation module 10, a feature encoding module 20, a vector quantization module 30, a conditional feature construction module 40, an embedding mapping module 50, a parallel prediction module 60, and a signal reconstruction module 70. Detailed descriptions of each functional module are as follows: The audio time-frequency conversion module 10 is used to acquire the audio signal to be processed, perform time-frequency conversion processing on the audio signal to be processed, and obtain a time-frequency representation. Feature encoding module 20 is used to input the time-frequency representation into the feature encoding network to obtain a continuous feature representation, and to divide the continuous feature representation into multiple sub-feature representations along the feature dimension; The vector quantization module 30 is used to input multiple sub-feature representations into multiple vector quantizers respectively to obtain a degenerate discrete label set, wherein the degenerate discrete label set includes multiple sets of degenerate discrete labels; The conditional feature construction module 40 is used to extract spectral features from the audio signal to be processed, obtain spectral features, and align the spectral features with the time scale of the degenerate discrete label set to obtain a conditional feature representation. The embedding mapping module 50 is used to perform embedding processing on the multiple sets of the degenerate discrete tags respectively to obtain multiple sets of discrete tag embeddings; Parallel prediction module 60 is used to embed multiple sets of discrete labels into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and input the conditional feature representation into the multiple prediction branches, and perform prediction processing in parallel through the multiple prediction branches to obtain multiple sets of target discrete labels. The signal reconstruction module 70 is used to map multiple sets of target discrete markers into multiple sets of codeword vectors, concatenate the multiple sets of codeword vectors into a quantization feature representation, and input the quantization feature representation into a decoding network to obtain an enhanced audio signal.

[0127] Specific limitations regarding the audio processing device based on discrete label prediction can be found in the foregoing limitations of the audio processing method based on discrete label prediction, and will not be repeated here. Each module in the aforementioned audio processing device based on discrete label prediction can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0128] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides deterministic and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements server-side functions or steps of an audio processing method based on discrete label prediction.

[0129] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of an audio processing method based on discrete label prediction.

[0130] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire the audio signal to be processed, and perform time-frequency transformation processing on the audio signal to be processed to obtain a time-frequency representation; The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension; The multiple sub-feature representations are respectively input into multiple vector quantizers to obtain a degenerate discrete label set, which includes multiple sets of degenerate discrete labels; The audio signal to be processed is subjected to spectral feature extraction to obtain spectral features, and the spectral features are aligned with the time dimension that matches the time scale of the degenerate discrete label set to obtain a conditional feature representation; The multiple sets of degenerate discrete tags are respectively embedded to obtain multiple sets of discrete tag embeddings; Multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. Multiple sets of target discrete markers are mapped to multiple sets of codeword vectors, and the multiple sets of codeword vectors are concatenated to form a quantized feature representation. The quantized feature representation is then input into a decoding network to obtain an enhanced audio signal.

[0131] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Acquire the audio signal to be processed, and perform time-frequency transformation processing on the audio signal to be processed to obtain a time-frequency representation; The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension; The multiple sub-feature representations are respectively input into multiple vector quantizers to obtain a degenerate discrete label set, which includes multiple sets of degenerate discrete labels; The audio signal to be processed is subjected to spectral feature extraction to obtain spectral features, and the spectral features are aligned with the time dimension that matches the time scale of the degenerate discrete label set to obtain a conditional feature representation; The multiple sets of degenerate discrete tags are respectively embedded to obtain multiple sets of discrete tag embeddings; Multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. Multiple sets of target discrete markers are mapped to multiple sets of codeword vectors, and the multiple sets of codeword vectors are concatenated to form a quantized feature representation. The quantized feature representation is then input into a decoding network to obtain an enhanced audio signal.

[0132] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0133] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0135] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0136] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An audio processing method based on discrete label prediction, characterized in that, Includes the following steps: Acquire the audio signal to be processed, and perform time-frequency transformation processing on the audio signal to be processed to obtain a time-frequency representation; The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension; The multiple sub-feature representations are respectively input into multiple vector quantizers to obtain a degenerate discrete label set, which includes multiple sets of degenerate discrete labels; The audio signal to be processed is subjected to spectral feature extraction to obtain spectral features, and the spectral features are aligned with the time dimension that matches the time scale of the degenerate discrete label set to obtain a conditional feature representation; The multiple sets of degenerate discrete tags are respectively embedded to obtain multiple sets of discrete tag embeddings; Multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. Multiple sets of target discrete markers are mapped to multiple sets of codeword vectors, and the multiple sets of codeword vectors are concatenated to form a quantized feature representation. The quantized feature representation is then input into a decoding network to obtain an enhanced audio signal.

2. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, The time-frequency representation is input into a feature encoding network to obtain a continuous feature representation, and the continuous feature representation is divided into multiple sub-feature representations along the feature dimension, including: The time-frequency representation is input into a feature encoding network for multi-layer feature extraction to obtain an intermediate feature representation; The intermediate feature representation is aggregated using the feature encoding network to obtain the aggregated feature representation. The aggregated feature representation is mapped using the feature encoding network to obtain a continuous feature representation; Determine the feature dimension partitioning position corresponding to the continuous feature representation; The continuous feature representation is divided along the feature dimension according to the feature dimension division position to obtain multiple sub-feature representations.

3. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, The multiple sub-feature representations are respectively input into multiple vector quantizers to obtain a degenerate discrete label set, which includes multiple sets of degenerate discrete labels, including: The multiple sub-feature representations are input into multiple vector quantizers according to their corresponding relationships; For each vector quantizer, retrieve the codebook corresponding to each vector quantizer and determine multiple distances between the sub-feature representation input to each vector quantizer and each codeword vector in the corresponding codebook; The target codeword vector is determined from its respective codebook based on the multiple distances described. Obtain the indices of the multiple target codeword vectors in their respective codebooks, and determine the multiple indices as multiple sets of degenerate discrete markers; Multiple sets of the aforementioned degenerate discrete tags are combined into a degenerate discrete tag set.

4. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, The audio signal to be processed is subjected to spectral feature extraction to obtain spectral features. These spectral features are then aligned with the time scale of the degenerate discrete label set to obtain a conditional feature representation, including: Perform a short-time Fourier transform operation on the audio signal to be processed to extract the amplitude spectrum features and phase correlation features of the audio signal to be processed; The amplitude spectrum features and the phase correlation features are fused to obtain the spectral features; Determine the time scale of the degenerate discrete tag set; Based on the time scale of the degenerate discrete marker set, the spectral features are downsampled and aligned in the time dimension to obtain the target aligned feature sequence. The target aligned feature sequence is input into a bidirectional recurrent neural network to perform long temporal dependency modeling, thereby obtaining a temporal dependency feature sequence. The temporal-dependent feature sequence is input into a network structure containing a self-attention mechanism for global context fusion processing to obtain conditional feature representations.

5. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, Embedding processes are performed on multiple sets of the degenerate discrete tags respectively to obtain multiple sets of discrete tag embeddings, including: For each group of degenerate discrete tags in the multiple groups of degenerate discrete tags, the corresponding embedding mapping table is retrieved; Based on each degenerate discrete tag in each group of degenerate discrete tags, read the embedding vector from the corresponding embedding map table; The read embedding vectors are arranged sequentially according to the tag arrangement order within each group of degenerate discrete tags; The sequentially arranged embedding vectors are subjected to intra-group assembly processing to obtain each group of discrete label embeddings; By aggregating the discrete label embeddings of each group, multiple sets of discrete label embeddings are obtained.

6. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, Multiple sets of discrete labels are embedded and input into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels, including: The multiple sets of discrete label embeddings are input into multiple prediction branches according to their correspondence with the multiple sets of degenerate discrete labels, and the conditional feature representation is input into the multiple prediction branches; In the multiple prediction branches, the discrete label embeddings and conditional feature representations of the input are fused to obtain the branch input feature sequence; Sequence modeling is performed on the input feature sequences of each of the multiple prediction branches to obtain the branch prediction feature sequences; The predicted feature sequences of each branch are labeled and mapped to obtain the labeled output sequences corresponding to each predicted branch. Multiple sets of discrete target labels are determined based on the label output sequences corresponding to each prediction branch.

7. The audio processing method based on discrete label prediction as described in claim 1, characterized in that, Multiple sets of target discrete markers are mapped to multiple sets of codeword vectors, and the multiple sets of codeword vectors are concatenated to form a quantized feature representation. The quantized feature representation is input into a decoding network to obtain an enhanced audio signal, including: For each set of target discrete tags in the multiple sets of target discrete tags, determine the corresponding codebook; Based on each target discrete tag in each group of target discrete tags, the corresponding codeword vector is read from the corresponding codebook to obtain multiple groups of codeword vectors; The codeword vectors are sequentially arranged according to the inter-group arrangement order and the intra-group arrangement order of the target discrete tags. The multiple sets of codeword vectors arranged in sequence are concatenated to obtain the quantized feature representation; The quantized feature representation is input into the decoding network, and the decoding network performs time-series recovery processing on the quantized feature representation to obtain the decoded features; The decoding features are reconstructed using the decoding network to obtain an enhanced audio signal.

8. An audio processing apparatus based on discrete label prediction, characterized in that, The audio processing device based on discrete label prediction includes: The audio time-frequency conversion module is used to acquire the audio signal to be processed, perform time-frequency conversion processing on the audio signal to be processed, and obtain a time-frequency representation; The feature encoding module is used to input the time-frequency representation into the feature encoding network to obtain a continuous feature representation, and to divide the continuous feature representation into multiple sub-feature representations along the feature dimension; The vector quantization module is used to input multiple sub-feature representations into multiple vector quantizers respectively to obtain a degenerate discrete label set, wherein the degenerate discrete label set includes multiple sets of degenerate discrete labels; The conditional feature construction module is used to extract spectral features from the audio signal to be processed, obtain spectral features, and align the spectral features with the time scale of the degenerate discrete label set to obtain a conditional feature representation. The embedding mapping module is used to perform embedding processing on multiple sets of the degenerate discrete tags respectively to obtain multiple sets of discrete tag embeddings; The parallel prediction module is used to embed multiple sets of discrete labels into multiple prediction branches corresponding to the multiple sets of degenerate discrete labels, and to input the conditional feature representation into the multiple prediction branches. The prediction processing is performed in parallel by the multiple prediction branches to obtain multiple sets of target discrete labels. The signal reconstruction module is used to map multiple sets of target discrete markers into multiple sets of codeword vectors, and to concatenate the multiple sets of codeword vectors into a quantized feature representation. The quantized feature representation is then input into the decoding network to obtain an enhanced audio signal.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and an audio processing program based on discrete label prediction stored in the memory and executable on the processor, wherein the audio processing program based on discrete label prediction, when executed by the processor, implements the steps of the audio processing method based on discrete label prediction as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an audio processing program based on discrete label prediction, which, when executed by a processor, implements the steps of the audio processing method based on discrete label prediction as described in any one of claims 1-7.