Voice processing method and device based on dual voice representation, equipment and medium
By employing a bidirectional bridging structure of a sparse bridging encoder and a dense bridging decoder, the problem of balancing compression rate and fidelity in existing speech processing is solved, achieving efficient compression and high-fidelity reconstruction, which is suitable for speech data processing in medical and financial scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-07
AI Technical Summary
In existing speech processing technologies, the single representation form makes it difficult to balance compression rate and fidelity. The unidirectional coding structure lacks high-fidelity bidirectional mapping, resulting in acoustic distortion during reconstruction. The compression strategy is also coarse and cannot adapt to different dynamic speech segments.
By employing a bidirectional bridging structure of a sparse bridging encoder and a dense bridging decoder, and through multi-scale convolution, hierarchical RVQ processing, and joint training with three types of constraints, a bidirectional reversible mapping with dual representations is established, achieving efficient compression and high-fidelity reconstruction.
It achieves efficient compression and high-fidelity voice reconstruction, reduces storage and transmission costs, ensures the integrity and reversibility of voice information, adapts to different network environments, and improves data processing efficiency and security in medical and financial scenarios.
Smart Images

Figure CN121811894A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, device, and storage medium based on dual speech representation. Background Technology
[0002] In the field of speech generation and speech representation learning, discrete speech tokenization and continuous feature representation are two core technical directions. However, both have long faced multiple key problems and shortcomings, which have limited the practical application of these technologies. On the one hand, the singularity of representation forms makes it difficult to balance compression rate and fidelity: while discrete speech tokens possess good structured characteristics, facilitating subsequent processing and model training, the quantization process inevitably causes a significant loss of information, resulting in obvious deficiencies in timbre restoration and prosodic detail representation in reconstructed speech. On the other hand, while continuous feature representation can accurately capture subtle acoustic information in speech signals and ensure fidelity, its high-dimensionality leads to high storage and transmission costs, failing to meet the efficiency requirements of rapid scene generation or large-scale model training. On the other hand, existing technologies generally lack high-fidelity bidirectional mapping mechanisms. Feature compression models represented by VQ-VAE and RVQ mostly adopt unidirectional encoding structures, lacking precise reverse decoding alignment designs, resulting in significant asymmetry between compression and decoding paths. Ultimately, this leads to hierarchical shifts between reconstructed and original features, causing acoustic distortion problems. Furthermore, the lack of hierarchical feature alignment and interpretable structures further exacerbates the technical bottleneck. Current methods, when performing multi-layer feature mapping, often only impose constraints on the final output, neglecting the alignment requirements of intermediate layer features. This leads to information degradation layer by layer during encoding and decoding, making it difficult to guarantee hierarchical consistency and reconstruction stability. Furthermore, the coarseness of compression strategies cannot be ignored. Most speech compression schemes use fixed rates or fixed codebooks, failing to fully consider the non-uniformity of speech signals in the temporal and semantic dimensions. This results in insufficient compression of dynamic speech segments (such as emotional intonation and long vowels) while over-compressing static segments, severely impacting the consistency of the overall reconstruction results. In summary, existing technologies have significant shortcomings in key dimensions such as compression efficiency, feature reconstruction quality, hierarchical consistency, and mapping reversibility, making it difficult to simultaneously achieve the goals of high compression ratio and high fidelity speech reconstruction.
[0003] In the healthcare sector, voice interaction is widely used in medical settings for remote consultation recording and archiving, electronic medical record voice input, and monitoring of voice function in rehabilitation patients. This places extremely high demands on the fidelity and integrity of the voice data. Discrete voice tokenization can lead to timbre distortion and rhythm loss, potentially causing doctors to misjudge a patient's breathing status, emotional responses (such as voice tremors caused by pain), or the recovery of their vocal function. Meanwhile, the high storage cost of continuous feature representations exacerbates the storage pressure on medical data centers and does not meet the requirements for efficient transmission and immediate retrieval of medical data. Furthermore, the lack of a bidirectional mapping mechanism and insufficient hierarchical alignment can result in the loss of crucial information (such as dosage figures in medical orders and detailed symptom descriptions) after compression and reconstruction of medical voice data. A coarse, fixed compression strategy may not adequately compress highly dynamic segments of pathological speech (such as wheezing in asthma patients or slurred speech in stroke patients) or over-compress static speech from routine consultations, further damaging the medical value of the data. Insufficient mapping reversibility can prevent accurate reconstruction of the original medical speech, affecting the validity of evidence in medical disputes and the traceability of the treatment process.
[0004] In the fintech sector, scenarios such as voice payment verification, remote account opening identity verification, and customer service call recording for evidence preservation in financial settings require both efficient storage and transmission of voice data and accurate identification of identity features and information confidentiality. Information loss in discrete voice tokenization can lead to distortion of user voiceprint features, reducing the accuracy of identity verification and increasing security risks such as fraudulent transactions and misuse. Conversely, the high-dimensionality of continuous feature representation can slow down the response speed of real-time transaction verification, affecting the user experience in payment and account opening scenarios, while also increasing the server maintenance costs for financial institutions. Furthermore, acoustic distortion caused by asymmetric bidirectional mapping can lead to misreading or loss of key information such as transaction instructions and risk warnings during customer service calls, potentially causing financial disputes. Lack of hierarchical feature alignment can exacerbate semantic deviations after speech-to-text conversion, affecting the accuracy of intelligent customer service in understanding user needs. Fixed compression strategies cannot adapt to the differences between highly dynamic segments (such as instructions given when users are emotionally agitated or parameter descriptions of complex financial products) and static segments (such as routine consultation scripts) in financial voice, potentially leading to compression distortion of key transaction information. Insufficient mapping reversibility can affect the legal validity and traceability accuracy of financial voice evidence preservation. Summary of the Invention
[0005] The main objective of this invention is to provide a speech processing method, apparatus, device, and storage medium based on dual speech representations, aiming to solve the problems in existing speech processing where the single representation leads to the difficulty in balancing compression rate and fidelity, the lack of high-fidelity bidirectional mapping in unidirectional coding structures, acoustic distortion in reconstruction, coarse compression strategies, and the inability to adapt to different dynamic speech segments.
[0006] To achieve the above objectives, the present invention provides a speech processing method based on dual speech representation, comprising: A sparse bridging encoder and a dense bridging decoder are pre-built, and the sparse bridging encoder and the dense bridging decoder are jointly trained. Obtain raw speech data, analyze the raw speech data, and generate continuous dense features; The trained sparse bridging encoder converts the continuous dense features into sparse discrete tokens; The sparse discrete tokens are reconstructed using the trained dense bridging decoder to generate reconstructed continuous dense features.
[0007] Furthermore, to achieve the above objectives, the present invention provides a speech processing apparatus based on dual speech representation, comprising: The model building and training module is used to pre-build a sparse bridging encoder and a dense bridging decoder, and to jointly train the sparse bridging encoder and the dense bridging decoder. The data acquisition and processing module is used to acquire raw speech data, analyze the raw speech data, and generate continuous dense features; A sparse discrete token module is used to convert the continuous dense features into sparse discrete tokens by a trained sparse bridging encoder. The continuous dense feature reconstruction module is used to reconstruct the features of the sparse discrete tokens by the trained dense bridging decoder to generate reconstructed continuous dense features.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a speech processing program based on dual speech representation stored in the memory and executable on the processor, wherein when the speech processing program based on dual speech representation is executed by the processor, it implements the steps of the speech processing method based on dual speech representation as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a speech processing program based on dual speech representation, wherein the speech processing program based on dual speech representation, when executed by a processor, implements the steps of the speech processing method based on dual speech representation as described above.
[0010] Beneficial Effects: This invention relates to the field of speech processing technology and can be applied to business system platforms such as healthcare and fintech. It discloses a speech processing method based on dual speech representation, comprising: pre-constructing a sparse bridging encoder and a dense bridging decoder, and jointly training the sparse bridging encoder and the dense bridging decoder; acquiring raw speech data, analyzing the raw speech data, and generating continuous dense features; converting the continuous dense features into sparse discrete tokens by the trained sparse bridging encoder; and reconstructing the features of the sparse discrete tokens by the trained dense bridging decoder to generate reconstructed continuous dense features. This invention proposes a bidirectional bridging structure containing a sparse bridging encoder and a dense bridging decoder. Through multi-scale convolution, hierarchical RVQ processing, and joint training with three types of constraints, a bidirectional reversible mapping of dual representations is established, achieving efficient compression and high-fidelity reconstruction. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a speech processing method based on dual speech representation in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the speech processing method based on dual speech representation of the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the speech processing device based on dual speech representation of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The speech processing method based on dual speech representation provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can obtain raw speech data from the user terminal, analyze the raw speech data, and generate continuous dense features; pre-construct a sparse bridging encoder and a dense bridging decoder, and jointly train the sparse bridging encoder and the dense bridging decoder; the trained sparse bridging encoder converts the continuous dense features into sparse discrete tokens; the trained dense bridging decoder reconstructs the features of the sparse discrete tokens to generate reconstructed continuous dense features. This invention proposes a bidirectional bridging structure containing a sparse bridging encoder and a dense bridging decoder. Through multi-scale convolution, hierarchical RVQ processing, and joint training with three types of constraints, a bidirectional reversible mapping with dual representations is established, achieving efficient compression and high-fidelity reconstruction. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the speech processing method based on dual speech representation provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the speech processing method based on dual speech representation proposed in this invention includes the following steps: S100. Pre-build a sparse bridging encoder and a dense bridging decoder, and jointly train the sparse bridging encoder and the dense bridging decoder. S200: Acquire raw speech data, analyze the raw speech data, and generate continuous dense features; S300, The continuous dense features are converted into sparse discrete tokens by the trained sparse bridging encoder; S400: The sparse discrete tokens are reconstructed by the trained dense bridging decoder to generate reconstructed continuous dense features.
[0016] In this embodiment, the sparse bridge encoder and dense bridge decoder are first constructed and jointly trained. The core function of the sparse bridge encoder is to transform continuous dense features into sparse discrete tokens. It includes a multi-scale convolutional feature extraction module, a temporal downsampling module, a hierarchical residual vector quantization (RVQ) module, and a code selection module. It captures the short-term and long-term dependencies of the speech signal through multi-scale convolution, reduces temporal redundancy through downsampling, and encodes high-dimensional features in blocks through hierarchical residual quantization. Finally, it selects the core encoding index with the most information content to form sparse tokens. The dense bridge decoder is responsible for reverse reconstruction. It completes the missing intermediate encoding in the quantization process with the help of the code prediction module, restores the original time scale through hierarchical residual quantization decoding and temporal upsampling, and then performs detail compensation and smoothing through a multi-scale convolutional inverse network to achieve accurate reconstruction of sparse tokens into continuous dense features. During the joint training phase, code prediction loss (based on cross-entropy constraints to ensure consistency between predicted and actual codes), feature reconstruction loss (based on mean square error constraints to reconstruct the similarity between features and original features), and intermediate layer alignment constraints (layer-by-layer supervision of features at each level of the two-layer network) are used to ensure the consistency and stability of bidirectional mapping, laying the foundation for subsequent efficient compression and high-fidelity reconstruction.
[0017] After acquiring the raw speech data, professional signal processing techniques are used to preprocess and extract features, generating continuous dense features that fully represent the acoustic details (such as timbre and prosody) and semantic information of the speech. These features serve as the core data carrier for subsequent processing. The continuous dense features are then input into a trained sparse bridging encoder. Through collaborative processing by multiple internal modules, the high-dimensional continuous features are transformed into low-dimensional sparse discrete tokens, achieving efficient compression of speech features and significantly reducing data storage and transmission costs. Finally, the sparse discrete tokens are input into a trained dense bridging decoder. Through a series of operations including encoding completion, decoding, upsampling, and detail restoration, feature reconstruction is completed, generating reconstructed continuous dense features highly similar to the original continuous dense features, ensuring the integrity and fidelity of the speech information before and after compression.
[0018] In the field of healthcare, this invention effectively solves the problems of efficient processing and secure application of voice-based medical data. In remote consultation scenarios, the voice communication between patients and doctors needs to be transmitted in real time. Especially in remote areas with low-bandwidth network environments, the high compression characteristics of sparse discrete tokens can significantly reduce the latency of voice data transmission, while ensuring the complete restoration of key symptom descriptions and medical orders in the voice, avoiding diagnostic errors caused by signal distortion. In medical voice record management, massive amounts of medical record voice data, surgical procedure voice replays, and other data can be compressed by this invention to reduce storage resource consumption, and the reconstructed voice can accurately retain medical details, meeting the needs of subsequent case tracing, medical teaching and research, etc. In addition, in assistive diagnostic and treatment equipment, such as voice-controlled medical instruments and voice-interactive diagnostic tools for disabled patients, this invention can enable the device to quickly receive and accurately recognize voice commands, ensuring the accuracy of command understanding through high-fidelity reconstruction, and improving the interaction efficiency and safety of medical equipment.
[0019] In the fintech field, this invention enables the efficient and secure implementation of voice-interactive financial services. In intelligent customer service scenarios, a large volume of user inquiries (such as account inquiries, business processing, and complaint feedback) require real-time processing and storage. Feature compression reduces the storage pressure and data transmission costs of customer service systems, while high-fidelity reconstruction ensures that customer service personnel or intelligent systems accurately extract user needs, improving response efficiency. In remote financial transactions, such as voice identity verification and agreement signing, this invention ensures efficient transmission and accurate reconstruction of user voice features, meeting the stringent requirements of financial services for accurate identity verification while adapting to different network environments, thus enhancing the convenience of user transactions. Furthermore, in the field of financial voice data security, the compressed sparse discrete token reduces the risk of data theft during transmission, and the bidirectional reversibility and high consistency of the reconstruction process ensure the complete traceability of voice information during subsequent data audits and compliance checks, aligning with the stringent standards of the financial industry for data security and compliance.
[0020] In one embodiment, S100 includes: S101. A sparse bridging encoder is constructed using a multi-scale convolutional feature extractor, a temporal downsampling module, a hierarchical RVQ encoder, and a code selector. S102. A dense bridging decoder is constructed using a code prediction module, a hierarchical RVQ decoder, a temporal upsampling module, and a multi-scale convolutional inverse network. S103. Using code prediction loss, feature reconstruction loss and intermediate layer alignment constraints, the sparse bridging encoder and dense bridging decoder are jointly trained, and a bidirectional mapping relationship between the sparse bridging encoder and dense bridging decoder is constructed.
[0021] In this embodiment, the construction of the sparse bridging encoder revolves around the goal of efficient feature compression, and its core is composed of four functional modules working together. The multi-scale convolutional feature extractor uses convolutional kernels of different sizes (1, 3, 5, etc.) to simultaneously capture short-term local dependencies and long-term global correlations in speech signals, transforming the original speech features into enhanced F1 features rich in contextual information, laying the foundation for subsequent processing. The temporal downsampler reduces the temporal resolution of features and the number of data frames through downsampling operations at a fixed ratio (e.g., 5 times), resulting in compacted F2 features, effectively reducing the computational complexity of subsequent quantization processing. The hierarchical RVQ encoder (Hierarchical Residual Vector Quantization Encoder) processes the high-dimensional F2 features in blocks, with each block being quantized layer by layer through a multi-level codebook, achieving structured compression of high-dimensional data. The code selector filters the quantized multi-level codebook indexes, retaining only the first-level or most information-dense core indexes, ultimately outputting a sparse discrete token sequence, maximizing the compression ratio while preserving key speech information.
[0022] The dense bridging decoder is built around accurate feature reconstruction, forming a reverse restoration link corresponding to the sparse bridging encoder. The Code Predictor module, based on the core index in the sparse discrete tokens, uses model inference to complete the missing mid-level codebook information during the hierarchical RVQ quantization process, providing data support for complete decoding. The Hierarchical RVQ Decoder, based on the completed codebook index, reconstructs the high-dimensional feature vector F3, restoring the feature structure before quantization. The Temporal Upsampler module restores the temporal scale of F3 to the level of the original speech features, obtaining F4 that matches the temporal dimension of the original features. The ContextualFeature Restorer performs deconvolution operations to smooth and compensate for details in F4, repairing potential information loss during decoding, and finally generating a reconstructed feature F5 that is highly consistent with the original continuous dense features.
[0023] During the joint training phase, three types of constraint mechanisms ensure the accuracy and stability of the bidirectional mapping, constructing a closed-loop association between the sparse bridging encoder and the dense bridging decoder. The code prediction loss employs the cross-entropy loss function, constraining the completed code output by the code prediction module to maintain consistency with the true-level RVQ code, ensuring the integrity of the quantized information. The feature reconstruction loss calculates the mean square error between the original continuous features (F0) and the reconstructed features (F5), forcing the reconstructed features to closely approximate the original data in acoustic details, reducing distortion. The intermediate layer alignment constraint pairs and supervises the features of corresponding intermediate layers in both networks, preventing information degradation layer by layer during encoding and decoding. This ensures the consistency and reversibility of features at each level during the bidirectional conversion from continuous features to sparse tokens and back to continuous features, ultimately achieving a balance between high compression ratio and high fidelity.
[0024] In the healthcare field, this invention's architecture is deeply adaptable to the core needs of voice-based medical scenarios. In telemedicine consultations, the voice interaction data between doctors and patients needs to be transmitted across different network environments. The high compression characteristics of the sparse bridging encoder can significantly reduce data transmission bandwidth requirements, enabling rapid transmission of consultation voice even in remote areas with low network speeds. Meanwhile, the high-fidelity reconstruction capability of the dense bridging decoder accurately restores key medical information such as symptom descriptions and diagnostic suggestions from the voice, avoiding diagnostic errors caused by signal distortion. In medical voice record management, massive amounts of surgical anesthesia voice recordings, case follow-up voices, and medical teaching voices can significantly save storage resources after compression by the sparse bridging encoder. Furthermore, the voice reconstructed by the dense bridging decoder can completely preserve medical details, meeting subsequent needs such as case tracing, medical research, and compliance auditing. In assistive medical devices, such as voice-controlled rehabilitation equipment and voice translation devices for hearing-impaired patients, this architecture enables devices to quickly receive and accurately interpret voice commands. The high stability of bidirectional mapping ensures the accuracy of command understanding, improving the interactive security and efficiency of medical devices.
[0025] In the fintech field, this invention's architecture effectively addresses the efficiency and security pain points of voice-based financial services. In intelligent financial customer service scenarios, the massive amounts of user inquiries generated daily (such as account inquiries, loan applications, and complaint feedback) can be compressed using a sparse bridging encoder, reducing storage pressure and data transmission costs for the customer service system. Meanwhile, a dense bridging decoder accurately reconstructs the user's voice request information, helping intelligent or human customer service quickly locate problems and respond efficiently. In remote financial transactions, such as voice identity verification and voice signing of electronic agreements, the sparse bridging encoder quickly compresses user voice features and transmits them securely, while the dense bridging decoder accurately restores voice biometrics, meeting the high-precision requirements of financial transactions for identity verification. It also adapts to the complex environment of mobile networks, improving the convenience of user transactions. In the field of financial voice data security and compliance, the compressed sparse discrete token reduces the risk of data leakage during transmission and storage. The high consistency and hierarchical alignment of the bidirectional mapping ensure the complete traceability of voice information during subsequent data audits and risk checks, meeting the stringent standards of the financial industry for data security and compliance, and providing technical support for the large-scale deployment of voice-based financial services.
[0026] In one embodiment, S200 includes: S201. Obtain raw voice data; S202. Preprocess the raw speech data to obtain standardized speech data; S203. The standardized speech data is subjected to feature extraction using acoustic feature extraction technology to generate continuous dense features.
[0027] In this embodiment, acquiring raw speech data is the starting point of the entire speech feature processing process. This raw data may come from various speech acquisition devices (such as microphones, recorders, smart terminals, etc.), covering speech information from different scenarios and different speakers, and its form is unprocessed raw audio signal.
[0028] The core objective of preprocessing raw speech data is to eliminate noise interference and standardize the data format to obtain standardized speech data, laying the foundation for subsequent feature extraction. The preprocessing process typically includes several key steps: First, noise suppression is performed using techniques such as spectral subtraction and adaptive filtering to remove environmental noise (e.g., background noise, equipment interference) and non-stationary noise in the speech, preserving a clean speech signal. Second, speech segmentation and endpoint detection are performed, using indicators such as energy threshold and zero-crossing rate to accurately identify the start and end positions of the speech, eliminating invalid silence segments and improving data processing efficiency. Third, sampling rate and quantization bit normalization is performed, converting speech data collected from different devices and in different formats into a unified sample rate (e.g., 16kHz) and quantization bit depth (e.g., 16-bit) to ensure data consistency. Finally, pre-emphasis processing may be performed, using a high-pass filter to enhance the high-frequency components in the speech signal, compensating for the attenuation of high-frequency information during speech transmission and enhancing the recognizability of speech features.
[0029] After preprocessing, standardized speech data needs to be processed using acoustic feature extraction techniques to generate continuous dense features. These features comprehensively capture the acoustic properties and semantic information of speech, serving as the core carrier for subsequent compression and reconstruction. Acoustic feature extraction focuses on key acoustic dimensions of speech, with common extraction targets including Mel frequency cepstral coefficients (MFCC), Mel spectrograms, and linear prediction coefficients (LPC). These features effectively characterize key information such as timbre, pitch, and prosody. During extraction, the standardized speech data is segmented into frames using a sliding window, transforming continuous speech signals into discrete frame sequences. Each frame then undergoes a series of operations, including Fourier transform, Mel filtering, and cepstral analysis, ultimately generating a high-dimensional, continuous, dense feature vector. This vector retains detailed speech information while possessing good processability, providing high-quality input for subsequent compression operations of the sparse bridging encoder.
[0030] In the field of healthcare, the process of this invention can be precisely adapted to the stringent requirements of voice-based medical scenarios. In remote consultations and chronic disease management, the voice data recorded by patients through smart terminals, including symptom descriptions and medication feedback, is often mixed with impurities such as home environment noise and equipment interference. After noise suppression and standardization in the preprocessing stage, clean and uniform voice data can be obtained. Then, continuous dense features generated by acoustic feature extraction can completely preserve key information such as tone changes and symptom details in the patient's voice, providing doctors with accurate basis for judging the condition through voice. In rehabilitation medicine, for rehabilitation training of patients with language disorders, this process can standardize the patient's pronunciation practice voice. The extracted continuous dense features can accurately capture pronunciation defects (such as inaccurate pitch and abnormal rhythm), helping rehabilitation therapists to develop personalized training plans and providing data support for the quantitative evaluation of training effects. In the construction of medical voice archives, massive amounts of surgical voice records, case follow-up voice data, etc., can be uniformly managed after preprocessing and standardization. The extracted continuous dense features provide the foundation for subsequent voice retrieval and intelligent analysis (such as extracting key medical terms and identifying nodes in the diagnosis and treatment process), improving the utilization efficiency and value of medical archives.
[0031] In the fintech field, this invention's process effectively supports the secure and efficient operation of voice-based financial services. In intelligent financial customer service scenarios, voice data submitted by users via telephone, apps, and other channels for inquiries and complaints may contain irrelevant information such as environmental noise and network interference. The preprocessing stage can quickly purify the voice signal and standardize the data format. The continuous dense features generated by acoustic feature extraction can accurately capture the user's needs and emotional tendencies, helping the intelligent customer service system quickly identify the type of user's question and improve response speed and accuracy. In voice identity verification scenarios, this process can standardize user-recorded identity verification voice (such as preset passwords), eliminating interference from different recording environments and devices. The extracted continuous dense features can accurately characterize the user's unique voice biometrics (such as voiceprints and intonation habits), providing highly recognizable feature basis for identity verification and ensuring the security of financial transactions. In financial voice compliance auditing, for voice records of business such as wealth management sales and loan approval, the preprocessed standardized data facilitates large-scale storage and management. The extracted continuous dense features can support subsequent intelligent compliance checks (such as identifying whether there are illegal sales pitches or whether risk warnings are fully disclosed), helping financial institutions efficiently meet regulatory compliance requirements and reduce business risks.
[0032] In one embodiment, S300 includes: S3011. The multi-scale convolutional feature extractor of the trained sparse bridging encoder extracts the short-term local dependencies and long-term global dependencies of the continuous dense features using multi-scale convolutional kernels. S3012. The short-term local dependency and the long-term global dependency are fused to generate context-enhanced features; S3013. The context enhancement features are downsampled in time according to a preset fixed ratio to reduce the temporal resolution of the context enhancement features and obtain downsampled features. S3014. Obtain the high-dimensional vector of the downsampled feature and divide the high-dimensional vector into several sub-high-dimensional vectors; S3015. Encode each sub-high-dimensional vector using a multi-level codebook, and perform residual quantization on the encoded high-dimensional vector to generate a complete RVQ code sequence. S3016. The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to the preset encoding filtering rules to obtain the target RVQ index. S3017. Combine all target RVQ indices to obtain sparse discrete tokens.
[0033] In this embodiment, during speech feature compression, the trained SparseBridge encoder first processes continuous dense features using a multi-scale convolutional feature extractor. This extractor employs convolutional kernels of different sizes (e.g., 1×1, 3×3, etc.). Small kernels (e.g., 1×1, 3×3) accurately capture short-term local dependencies in the speech signal, i.e., acoustic associations between adjacent speech frames (e.g., instantaneous pitch changes, short syllable transitions); large kernels (e.g., 5×5) can cover longer speech segments, capturing long-term global dependencies, such as prosodic coherence and semantic logical connections within sentences. This multi-scale capture mechanism comprehensively mines the deep information of speech features. Subsequently, the system fuses the extracted short-term local dependencies with long-term global dependencies, integrating the two types of information through weighted summation and concatenation to generate context-enhanced features (F1) that combine detail and global relevance, providing a rich and valuable feature foundation for subsequent compression processing.
[0034] Next, the context enhancement features are temporally downsampled at a preset fixed ratio (e.g., 5 times). This step reduces the number of time frames for the features, lowers the temporal resolution, and removes redundant data in the temporal dimension while retaining core information, resulting in more compact downsampled features (F2), effectively improving the efficiency of subsequent quantization processing. For the high-dimensional vector of the downsampled features, the system divides it into several independent sub-high-dimensional vectors, reducing the dimensionality complexity of individual vectors through block processing, creating conditions for accurate quantization. Subsequently, Hierarchical Residual Quantization (RVQ) is used to match a multi-level codebook for layer-by-layer encoding of each sub-high-dimensional vector. During the encoding process, residual learning continuously optimizes the quantization accuracy, ultimately generating a complete RVQ code sequence containing complete quantization information.
[0035] Finally, the code selector of the sparse bridging encoder selects the most representative core codes from the complete RVQ code sequence according to preset encoding filtering rules (such as retaining the primary encoding index with the highest information content), thus obtaining the target RVQ index. All target RVQ indices are combined sequentially to form the final sparse tokens. These tokens retain only the key quantization information of the speech features, achieving high-ratio compression. Simultaneously, through multi-scale feature extraction and precise quantization, information loss is minimized.
[0036] In the field of healthcare, the process of this invention can efficiently support the compression and processing of voice-based medical data and is adaptable to a variety of core scenarios. In telemedicine consultations, the continuous, dense features generated from the patient's description of symptoms and the doctor's treatment recommendations can be extracted using multi-scale convolution to fully preserve key medical information (such as the detailed tone of symptom descriptions and the rhythm and emphasis of medical orders). Efficient compression is achieved through temporal downsampling and hierarchical residual quantization. The generated sparse discrete tokens can be transmitted quickly in low-bandwidth network environments, avoiding the impact of network latency on consultation efficiency. Furthermore, no core medical information is lost during compression, ensuring that doctors can accurately obtain patient information. In medical voice archive storage, massive amounts of data, such as case follow-up voice recordings and surgical procedure voice recordings, can be significantly reduced in storage resources after compression using this process. The structured nature of the sparse discrete tokens facilitates the classification and retrieval of archives. When the original voice information needs to be retrieved later, it can be efficiently reconstructed based on the tokens, meeting the needs of case tracing, medical teaching and research, etc. In language disorder rehabilitation training, after processing, the sparse discrete tokens can accurately preserve the feature information related to pronunciation defects (such as local dependence on abnormal pitch and global correlation of disordered pronunciation rhythm), providing data support for rehabilitation therapists to analyze training effects and adjust training programs, thus assisting patients in efficient rehabilitation.
[0037] In the fintech field, this invention's process effectively solves the problem of compressing and efficiently utilizing voice-based financial data. In intelligent financial customer service scenarios, the continuous, dense features generated by user inquiries and complaints, after multi-scale convolution extraction, can capture key information about user needs (such as detailed descriptions of account issues and changes in tone of voice during complaints). After compression via temporal downsampling and hierarchical residual quantization, the generated sparse discrete tokens reduce the storage pressure and data transmission costs of the customer service system, while ensuring accurate reconstruction of user voice intent during subsequent reconstruction, thus helping intelligent customer service quickly locate problems and respond efficiently. In voice authentication scenarios, the sparse discrete tokens accurately retain the user's authentication voice (such as preset passwords and natural speech commands) after processing by this process. The unique voiceprint features of users (such as local dependencies on pronunciation habits and global correlations in intonation and rhythm) not only enable efficient storage and transmission of feature data, but also allow for rapid comparison of voiceprint information through token reconstruction during verification, improving the efficiency and accuracy of identity verification and ensuring the security of financial transactions. In financial voice compliance audits, sparse discrete tokens generated by compressing voice recordings from business transactions such as wealth management sales and credit approval can significantly reduce the storage and processing costs of audit data. Furthermore, the key voice information retained in the tokens can meet compliance inspection requirements (such as verifying whether risks have been fully disclosed and whether there are any illegal statements), helping financial institutions to efficiently complete compliance audits.
[0038] In one embodiment, S400 includes: S401. Based on the target RVQ index of sparse discrete tokens, the missing middle-layer RVQ code is predicted by the code prediction module of the trained dense bridging decoder. S402. Combine the predicted missing middle-level RVQ code with the target RVQ index to generate a complete hierarchical RVQ code sequence; S403. The complete hierarchical RVQ code sequence is inversely quantized using the hierarchical RVQ decoder of the trained dense bridging decoder to obtain the reconstructed feature vector; S404. The reconstructed feature vector is upsampled using the temporal upsampling module of the predicted dense bridging decoder to obtain temporally aligned features. S405. The time-aligned features are smoothed and acoustic details are compensated by a multi-scale convolutional inverse network to generate reconstructed continuous dense features.
[0039] In this embodiment, during the speech feature reconstruction stage, the target RVQ index (i.e., the core quantization index retained during encoding) in the sparse discrete tokens is first used as input. The code prediction module of the trained DenseBridge decoder then performs the prediction of the missing mid-level RVQ code. The hierarchical encoding characteristics of RVQ (Residual Vector Quantization) determine that complete quantization information contains multi-level codebook indices, while the target RVQ index only retains the first or high-information part. Based on the quantization rules and feature associations learned during training, the code prediction module optimizes the prediction accuracy through cross-entropy constraints, accurately completes the missing mid-level quantization codes, and ensures the integrity of the quantization information.
[0040] Subsequently, the predicted intermediate RVQ codes are combined with the original target RVQ index in hierarchical order to generate a complete hierarchical RVQ code sequence. This sequence contains all the key information for speech feature quantization, providing a complete data foundation for subsequent inverse quantization. Next, the complete hierarchical RVQ code sequence is inverse-quantized using a dense bridging decoder's hierarchical RVQ decoder. Following the block inverse operation logic corresponding to the encoding stage, the discrete quantized code sequence is restored to a continuous high-dimensional vector, i.e., the reconstructed feature vector (F3), completing the initial conversion from discrete tokens to continuous vectors.
[0041] To restore the temporal dimension of the original speech features, a temporal upsampling module (TemporalUpsampler) is used to upsample the reconstructed feature vector. The temporal resolution is increased by a factor corresponding to the temporal downsampling in the encoding stage (e.g., 5 times), supplementing redundant information in the temporal dimension and obtaining a time-aligned feature (F4) consistent with the time scale of the original continuous dense features. Finally, a multi-scale convolutional inverse network (Contextual Feature Restorer) is used to process the time-aligned features. This network captures multi-scale acoustic details of the speech signal through deconvolution kernels of different sizes. On the one hand, it smooths the features to eliminate distortion noise that may be generated during quantization and upsampling; on the other hand, it compensates and enhances key acoustic details such as timbre and prosody. Ultimately, a reconstructed continuous dense feature (F5) highly similar to the original continuous dense features (DenseFeatures) is generated, achieving high-fidelity reconstruction of the speech features.
[0042] In the healthcare field, this reconstruction process can accurately support the efficient restoration and application of voice-based medical data. In remote chronic disease follow-up scenarios, when the symptom description voice transmitted by patients through low-bandwidth devices is compressed and generated into sparse discrete tokens, and then reconstructed by a dense bridging decoder, the code prediction module can complete the missing quantization information, hierarchical RVQ decoding and upsampling operations can completely restore the temporal rhythm and details of the voice, and the multi-scale convolutional inverse network can accurately retain key information such as the tone and breathing rhythm of the patient when describing symptoms, helping doctors to accurately judge changes in the patient's condition through the reconstructed voice. In medical voice order archiving scenarios, when the compressed and stored order voice tokens are retrieved, this process can completely restore the voice details and logical coherence of the doctor's instructions, avoiding deviations in the execution of orders by medical staff due to missing information. In language rehabilitation assessment, after the sparse tokens of the patient's training voice are reconstructed, they can accurately reproduce the details of pronunciation defects (such as abnormal syllable pauses and inaccurate pitch), providing precise data support for rehabilitation therapists to compare training voices at different stages and adjust rehabilitation plans.
[0043] In the fintech field, this reconstruction technology can effectively ensure the security and efficiency of voice-based financial services. In intelligent financial customer service voice recall scenarios, compressed user consultation voice tokens, when requiring manual review or issue tracing, can be reconstructed with high fidelity through a series of operations such as code prediction and completion, inverse quantization, and detail compensation. This fully restores key information such as consultation content and emotional tone, helping customer service teams quickly locate the root cause of problems. In voice identity verification scenarios, the user-transmitted voiceprint feature token, after reconstruction, can accurately restore the unique acoustic characteristics of the voiceprint (such as pronunciation frequency and intonation), ensuring that the identity verification system can accurately compare user voiceprint information, preventing identity theft risks and ensuring the security of financial transactions. In financial voice compliance audit scenarios, compressed business communication voices (such as wealth management sales promotions and loan approval communications), after reconstruction, can completely retain the voice details of the entire communication process, ensuring that auditors can clearly identify compliance points such as whether there are illegal statements and whether risk warnings are complete, meeting the compliance and regulatory requirements of the financial industry.
[0044] In one embodiment, S103 includes: S1031. Using the cross-entropy loss function, calculate the difference between the missing intermediate RVQ code predicted by the code prediction module and the real intermediate RVQ code generated by the sparse bridging encoder. S1032. The difference is converted into a loss value using the cross-entropy loss function, and backpropagation is performed based on the loss value to adjust the model parameters. S1033. Using the mean squared error loss function, calculate the mean squared error between continuous dense features and reconstructed continuous dense features, and optimize the model parameters based on the mean squared error using the mean squared error loss function. S1035. Perform layer-by-layer pairing supervision on the downsampling features and context enhancement features of the sparse bridging encoder and the reconstructed feature vector and time alignment features of the dense bridging decoder to obtain the trained sparse bridging encoder and dense bridging decoder.
[0045] In this embodiment, during the model training phase, three types of loss functions and a hierarchical supervision mechanism are used to achieve accurate alignment and parameter optimization between the sparse bridging encoder and the dense bridging decoder. First, the code prediction loss is calculated and optimized using the cross-entropy loss function. The core objective is to minimize the difference between the missing intermediate RVQ code (Residual Vector Quantization) output by the code prediction module and the true intermediate RVQ code generated by the sparse bridging encoder. The cross-entropy loss function measures the similarity between the predicted and true distributions, transforming their difference into an optimizable loss value. The model adjusts its parameters through backpropagation based on this loss value, continuously improving the prediction accuracy of the code prediction module for the intermediate quantized code, ensuring the complete transmission of quantization information between discrete tokens and continuous features.
[0046] Secondly, the feature reconstruction loss is optimized by employing the Mean Squared Error Loss Function (MSE) to calculate the mean squared error between the original continuous dense features and the reconstructed continuous dense features. The MSE quantifies the numerical deviation between the two sets of features using a sum of squares and an average. The MSE loss function uses this deviation as an optimization objective to drive the model to adjust the network parameters of the encoder and decoder, reducing acoustic distortion during feature reconstruction and ensuring that the reconstructed features are highly consistent with the original features in key dimensions such as timbre and rhythm, thus achieving high-fidelity reconstruction.
[0047] Finally, there is layer-by-layer pairing supervision for intermediate layer alignment, which focuses on the consistency of features across the model's intermediate layers. Specifically, the context-enhanced features (F1) and downsampled features (F2) generated by the sparse bridging encoder are paired one by one with the reconstructed feature vectors (F3) and temporally aligned features (F4) generated by the dense bridging decoder, according to network layers. Similarity constraints and error supervision are applied to the features at each layer. This layer-by-layer supervision avoids the problem of information degradation in intermediate layers caused by focusing only on the final output. It ensures that the structure and information of features at each layer remain consistent during the bidirectional transformation from continuous features to sparse tokens and then from sparse tokens back to continuous features. Finally, through the synergistic effect of these three mechanisms, model training is completed, resulting in a stable sparse bridging encoder and a dense bridging decoder with accurate bidirectional mapping.
[0048] In the healthcare field, models optimized using this training mechanism can fully meet the high-precision requirements of medical voice data. In telemedicine consultation scenarios, the model ensures accurate prediction of mid-layer quantization codes through a cross-entropy loss function, and achieves high-fidelity reconstruction of voice features by combining a mean squared error loss function. Even under low-bandwidth transmission, it can completely restore key voice details (such as tone and logical expression) in patient symptom descriptions and doctor's treatment suggestions, avoiding the impact of information distortion on diagnostic accuracy. At the same time, the intermediate layer alignment supervision ensures the integrity of voice features at each level, enabling the reconstructed voice to accurately retain the pronunciation features of medical terms and the rhythmic emphasis of treatment instructions, providing reliable voice data support for remote consultations. In medical voice record management, the trained model can store voice data with a high compression ratio while ensuring the traceability of record voice through accurate feature reconstruction. When it is necessary to retrieve historical case voices, the reconstructed voice can clearly restore the dialogue details during the treatment process, meeting the needs of medical teaching and research, case review, and compliance auditing. In language rehabilitation training, the model's high-fidelity reconstruction capability can accurately reproduce the patient's pronunciation defects (such as abnormal pitch and syllable pause deviation), helping therapists to objectively evaluate the training effect. Meanwhile, the intermediate layer alignment mechanism ensures the stability of pronunciation feature extraction, providing accurate data for the development of personalized rehabilitation plans.
[0049] In the fintech field, the optimized model based on this training mechanism effectively balances the compression efficiency of voice data with the requirements for security and accuracy. In intelligent financial customer service scenarios, the model achieves efficient compression and high-fidelity reconstruction of user inquiry voice through dual optimization of cross-entropy loss and mean squared error loss. The compressed voice data reduces the storage and transmission costs of the customer service system, while the reconstructed voice accurately reproduces the user's intent and emotional state, helping intelligent customer service quickly locate problems and improve response efficiency. Intermediate layer alignment supervision ensures the consistency of voice features during processing, avoiding misjudgments of needs due to feature degradation. In voice authentication scenarios, the trained model accurately extracts and retains the unique features of the user's voiceprint (such as pronunciation frequency and intonation), achieving accurate comparison of voiceprint information through high-fidelity reconstruction. Simultaneously, the high compression ratio reduces the storage and transmission risks of voiceprint data, and the precise constraint of the cross-entropy loss function on the quantization code enhances the anti-interference capability of voiceprint recognition, ensuring the security of financial transactions. In financial voice compliance audits, the processed voice data achieves efficient storage and can accurately reconstruct and fully restore the details of business communication (such as risk warnings in wealth management sales and key questions and answers in credit approval). The intermediate layer alignment mechanism ensures the integrity and traceability of voice information, helping financial institutions to efficiently complete compliance checks and meet industry regulatory requirements.
[0050] In one embodiment, S300 further includes: S3021. Perform dynamic analysis of the downsampling features in both temporal and semantic dimensions to identify dynamic and static speech segments; S3022. Call the extended codebook to encode the sub-high-dimensional vectors corresponding to the dynamic speech segments to obtain the extended RVQ code; S3023. Encode the sub-high-dimensional vectors corresponding to static speech segments using a basic codebook to obtain the basic RVQ code; S3024. Integrate the extended RVQ code with the basic RVQ code to generate a complete RVQ code sequence; S3025. The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to the preset encoding filtering rules to obtain the target RVQ index. S3026. Combine all target RVQ indices to obtain sparse discrete tokens.
[0051] In this embodiment, during the sparse discrete token generation process, dynamic analysis of the downsampled features (F2) is first performed in both the temporal and semantic dimensions. Temporal dynamic analysis focuses on the temporal variation characteristics of the speech signal. By detecting indicators such as energy fluctuations and frequency change rates in speech frames, it identifies dynamic speech segments with dramatic changes in emotional intonation and long vowel extensions, as well as static speech segments with stable pronunciation and no significant changes. Semantic dynamic analysis, combined with the semantic logic of the speech, determines the core information-carrying segments (such as key expressions and semantic transitions) and auxiliary information segments in the sentence, further optimizing the segmentation accuracy of dynamic and static segments to ensure that the segmentation results both conform to signal characteristics and semantic expression rules.
[0052] Differentiated codebook encoding strategies are employed for the different segments after segmentation. For sub-high-dimensional vectors corresponding to dynamic speech segments, an extended codebook is used for encoding. The extended codebook has richer quantization dimensions and finer coding granularity, which can fully capture the complex acoustic details and semantic information in high-dynamic segments and avoid information loss due to insufficient encoding, ultimately generating extended RVQ codes (RVQ stands for Residual Vector Quantization). For sub-high-dimensional vectors corresponding to static speech segments, a basic codebook is used for encoding. The basic codebook focuses on simplicity and efficiency, reducing redundancy by simplifying the coding dimensions while ensuring that core information is not lost, thus generating basic RVQ codes.
[0053] Subsequently, following the temporal order and semantic logic of the speech segments, the extended RVQ codes and the basic RVQ codes are integrated to form a complete RVQ code sequence containing all quantized speech features. Next, the code selector of the trained sparse bridging encoder selects the most representative target RVQ index from the complete RVQ code sequence based on preset coding selection rules (such as retaining the highest-information first-level coding index or the code corresponding to the core semantics), eliminating redundant coding information. Finally, all target RVQ indices are combined in their corresponding order to obtain sparse discrete tokens that combine high compression ratio with information integrity, achieving adaptive and efficient compression of speech segments with different dynamic characteristics.
[0054] In the healthcare field, this adaptive coding technology can accurately adapt to the complex characteristics of medical voice data, empowering various core scenarios. In remote consultation scenarios, when patients describe key information such as pain levels and symptom changes, their speech often exhibits highly dynamic characteristics (such as rapid tone and intonation fluctuations). Extended codebook encoding can fully preserve the details of this core medical information, while static speech segments such as stating basic personal information are efficiently compressed using a basic codebook. The generated sparse discrete tokens can be transmitted quickly in low-bandwidth networks, while ensuring that doctors can accurately capture key details of the patient's condition through voice reconstruction, avoiding diagnostic errors. In surgical procedure voice recordings, dynamic speech segments such as the transmission of instructions for key surgical operations and emergency communication in unexpected situations are encoded using extended codebooks. After encoding with the extended codebook, key information such as rhythm and tone in speech can be fully preserved, while smooth expressions of routine operations are compressed and stored using the basic codebook. This saves storage resources and allows for accurate reconstruction of key speech details during surgical procedures in subsequent case reviews and medical teaching and research. In language rehabilitation training, high-dynamic speech segments with defects in patients' pronunciation practice (such as abnormal pitch fluctuations and abrupt syllable transitions) can be accurately captured using the extended codebook, providing precise data for therapists to analyze training effects. Meanwhile, static segments with smooth pronunciation are efficiently compressed, reducing the storage and processing costs of training data.
[0055] In the fintech field, this invention specifically addresses the compression and utilization of financial voice data with varying dynamic characteristics, improving service efficiency and security. In intelligent financial customer service scenarios, user complaints and urgent business inquiries often contain highly dynamic features (such as dramatic changes in tone due to emotional excitement and emphasis on key requests). Extended codebook encoding can fully preserve the core details and emotional inclinations of user needs, while static voice segments such as routine inquiries and information confirmations are efficiently compressed using a basic codebook. This reduces the storage and transmission pressure on the customer service system and allows intelligent or human customer service representatives to quickly and accurately pinpoint user needs by reconstructing the voice, improving response efficiency. In voice authentication scenarios, highly dynamic voice features such as the emphasis in user-preset passwords and the personalized intonation in natural speech commands... After being encoded using an extended codebook, the unique voiceprint information of users can be accurately preserved, improving the accuracy of identity verification. Meanwhile, the steady parts of the voice are compressed through the basic codebook, ensuring the efficient storage and transmission of voiceprint data and safeguarding the security of financial transactions. In the voice compliance audit of financial business, dynamic voice segments such as the emphasis on risk warnings in wealth management sales and the Q&A of key issues in credit approval can be fully preserved through extended codebook encoding, which can retain the key information required for compliance checks. Meanwhile, static segments of regular communication are efficiently compressed, reducing the storage and processing costs of audit data, while ensuring that the core details of business communication can be accurately reproduced during the audit process, meeting compliance and regulatory requirements.
[0056] In one embodiment, a speech processing apparatus based on dual speech representation is provided, which corresponds one-to-one with the speech processing method based on dual speech representation in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the speech processing device based on dual speech representation of the present invention. The modules include a model construction and training module 10, a data acquisition and processing module 20, a sparse discrete token module 30, and a continuous dense feature reconstruction module 40. Detailed descriptions of each functional module are as follows: The model building and training module 10 is used to pre-build a sparse bridging encoder and a dense bridging decoder, and to jointly train the sparse bridging encoder and the dense bridging decoder. The data acquisition and processing module 20 is used to acquire raw speech data, analyze the raw speech data, and generate continuous dense features; The sparse discrete token module 30 is used to convert the continuous dense features into sparse discrete tokens by a trained sparse bridging encoder. The continuous dense feature reconstruction module 40 is used to reconstruct the features of the sparse discrete tokens by the trained dense bridging decoder to generate reconstructed continuous dense features.
[0057] In one embodiment, the model building and training module 10 includes: A sparse bridging encoder is constructed using a multi-scale convolutional feature extractor, a temporal downsampling module, a hierarchical RVQ encoder, and a code selector. A dense bridging decoder is constructed using a code prediction module, a hierarchical RVQ decoder, a temporal upsampling module, and a multi-scale convolutional inverse network. The sparse bridging encoder and dense bridging decoder are jointly trained using code prediction loss, feature reconstruction loss, and intermediate layer alignment constraints, and a bidirectional mapping relationship between the sparse bridging encoder and dense bridging decoder is constructed.
[0058] In one embodiment, the data acquisition and processing module 20 includes: Acquire raw voice data; The raw speech data is preprocessed to obtain standardized speech data; The standardized speech data is subjected to feature extraction using acoustic feature extraction technology to generate continuous dense features.
[0059] In one embodiment, the sparse discrete token module 30 includes: The multi-scale convolutional feature extractor trained by the sparse bridging encoder extracts the short-term local dependencies and long-term global dependencies of the continuous dense features using multi-scale convolutional kernels. The short-term local dependencies and the long-term global dependencies are fused to generate context-enhanced features; The context enhancement features are downsampled in time according to a preset fixed ratio to reduce the temporal resolution of the context enhancement features and obtain downsampled features. Obtain the high-dimensional vector of the downsampled feature, and divide the high-dimensional vector into several sub-high-dimensional vectors; A multi-level codebook is used to encode each sub-high-dimensional vector, and residual quantization is performed on the encoded high-dimensional vector to generate a complete RVQ code sequence. The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to a preset encoding filtering rule to obtain the target RVQ index. Combine all target RVQ indices to obtain sparse discrete tokens.
[0060] In one embodiment, the reconstructed continuous dense feature module 40 includes: Based on the target RVQ index of sparse discrete tokens, the missing mid-layer RVQ code is predicted by the code prediction module of the trained dense bridging decoder. The predicted missing middle-level RVQ code is combined with the target RVQ index to generate a complete hierarchical RVQ code sequence; The complete hierarchical RVQ code sequence is inversely quantized using the trained dense bridging decoder to obtain the reconstructed feature vector. The reconstructed feature vector is upsampled using the temporal upsampling module of the predicted dense bridging decoder to obtain temporally aligned features. The temporally aligned features are smoothed and acoustic details are compensated by a multi-scale convolutional inverse network to generate reconstructed continuous dense features.
[0061] In one embodiment, the joint training of the sparse bridging encoder and dense bridging decoder using code prediction loss, feature reconstruction loss, and intermediate layer alignment constraints, and the construction of a bidirectional mapping relationship between the sparse bridging encoder and dense bridging decoder, includes: The cross-entropy loss function is used to calculate the difference between the missing intermediate RVQ code predicted by the code prediction module and the actual intermediate RVQ code generated by the sparse bridging encoder. This difference is then converted into a loss value using the cross-entropy loss function, and backpropagation is performed based on this loss value to adjust the model parameters. The mean squared error loss function is used to calculate the mean squared error between continuous dense features and reconstructed continuous dense features, and the model parameters are optimized based on the mean squared error using the mean squared error loss function. The downsampled features and context enhancement features of the sparse bridging encoder are paired and supervised layer by layer with the reconstructed feature vector and temporal alignment features of the dense bridging decoder to obtain the trained sparse bridging encoder and dense bridging decoder.
[0062] In one embodiment, the sparse discrete token module 30 further includes: Dynamic analysis of the downsampling features in both temporal and semantic dimensions is performed to identify dynamic and static speech segments. The extended codebook is called to encode the sub-high-dimensional vectors corresponding to the dynamic speech segments, resulting in the extended RVQ code; The basic RVQ code is obtained by encoding the sub-high-dimensional vectors corresponding to static speech segments using a basic codebook. The extended RVQ code is integrated with the basic RVQ code to generate a complete RVQ code sequence; The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to a preset encoding filtering rule to obtain the target RVQ index. Combine all target RVQ indices to obtain sparse discrete tokens.
[0063] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side speech processing method based on dual speech representation.
[0064] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements user-side functions or steps of a speech processing method based on dual speech representation. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: A sparse bridging encoder and a dense bridging decoder are pre-built, and the sparse bridging encoder and the dense bridging decoder are jointly trained. Obtain raw speech data, analyze the raw speech data, and generate continuous dense features; The trained sparse bridging encoder converts the continuous dense features into sparse discrete tokens; The sparse discrete tokens are reconstructed using the trained dense bridging decoder to generate reconstructed continuous dense features.
[0065] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: A sparse bridging encoder and a dense bridging decoder are pre-built, and the sparse bridging encoder and the dense bridging decoder are jointly trained. Obtain raw speech data, analyze the raw speech data, and generate continuous dense features; The trained sparse bridging encoder converts the continuous dense features into sparse discrete tokens; The sparse discrete tokens are reconstructed using the trained dense bridging decoder to generate reconstructed continuous dense features.
[0066] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0067] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0068] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0069] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech processing method based on dual speech representation, characterized in that, Includes the following steps: A sparse bridging encoder and a dense bridging decoder are pre-built, and the sparse bridging encoder and the dense bridging decoder are jointly trained. Obtain raw speech data, analyze the raw speech data, and generate continuous dense features; The trained sparse bridging encoder converts the continuous dense features into sparse discrete tokens; The sparse discrete tokens are reconstructed using the trained dense bridging decoder to generate reconstructed continuous dense features.
2. The speech processing method based on dual speech representation as described in claim 1, characterized in that, The pre-construction of a sparse bridging encoder and a dense bridging decoder, and the joint training of the sparse bridging encoder and the dense bridging decoder, includes: A sparse bridging encoder is constructed using a multi-scale convolutional feature extractor, a temporal downsampling module, a hierarchical RVQ encoder, and a code selector. A dense bridging decoder is constructed using a code prediction module, a hierarchical RVQ decoder, a temporal upsampling module, and a multi-scale convolutional inverse network. The sparse bridging encoder and dense bridging decoder are jointly trained using code prediction loss, feature reconstruction loss, and intermediate layer alignment constraints, and a bidirectional mapping relationship between the sparse bridging encoder and dense bridging decoder is constructed.
3. The speech processing method based on dual speech representation as described in claim 1, characterized in that, The process of acquiring raw speech data and analyzing the raw speech data to generate continuous dense features includes: Acquire raw speech data; The raw speech data is preprocessed to obtain standardized speech data; The standardized speech data is subjected to feature extraction using acoustic feature extraction technology to generate continuous dense features.
4. The speech processing method based on dual speech representation as described in claim 1, characterized in that, The process of converting the continuous dense features into sparse discrete tokens by the trained sparse bridging encoder includes: The multi-scale convolutional feature extractor trained by the sparse bridging encoder extracts the short-term local dependencies and long-term global dependencies of the continuous dense features using multi-scale convolutional kernels. The short-term local dependencies and the long-term global dependencies are fused to generate context-enhanced features; The context enhancement features are downsampled in time according to a preset fixed ratio to reduce the temporal resolution of the context enhancement features and obtain downsampled features. Obtain the high-dimensional vector of the downsampled feature, and divide the high-dimensional vector into several sub-high-dimensional vectors; A multi-level codebook is used to encode each sub-high-dimensional vector, and residual quantization is performed on the encoded high-dimensional vector to generate a complete RVQ code sequence. The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to a preset encoding filtering rule to obtain the target RVQ index. Combine all target RVQ indices to obtain sparse discrete tokens.
5. The speech processing method based on dual speech representation as described in claim 1, characterized in that, The step of reconstructing features from the sparse discrete tokens using a trained dense bridging decoder to generate reconstructed continuous dense features includes: Based on the target RVQ index of sparse discrete tokens, the missing mid-layer RVQ code is predicted by the code prediction module of the trained dense bridging decoder. The predicted missing middle-level RVQ code is combined with the target RVQ index to generate a complete hierarchical RVQ code sequence; The complete hierarchical RVQ code sequence is inversely quantized using the trained dense bridging decoder to obtain the reconstructed feature vector. The reconstructed feature vector is upsampled using the temporal upsampling module of the predicted dense bridging decoder to obtain temporally aligned features. The temporally aligned features are smoothed and acoustic details are compensated by a multi-scale convolutional inverse network to generate reconstructed continuous dense features.
6. The speech processing method based on dual speech representation as described in claim 1, characterized in that, The method of jointly training the sparse bridging encoder and dense bridging decoder using code prediction loss, feature reconstruction loss, and intermediate layer alignment constraints, and constructing a bidirectional mapping relationship between the sparse bridging encoder and dense bridging decoder, includes: The cross-entropy loss function is used to calculate the difference between the missing intermediate RVQ code predicted by the code prediction module and the real intermediate RVQ code generated by the sparse bridging encoder. The difference is converted into a loss value using the cross-entropy loss function, and the model parameters are adjusted by backpropagation based on the loss value. The mean squared error loss function is used to calculate the mean squared error between continuous dense features and reconstructed continuous dense features, and the model parameters are optimized based on the mean squared error using the mean squared error loss function. The downsampled features and context enhancement features of the sparse bridging encoder are paired and supervised layer by layer with the reconstructed feature vector and temporal alignment features of the dense bridging decoder to obtain the trained sparse bridging encoder and dense bridging decoder.
7. The speech processing method based on dual speech representation as described in claim 4, characterized in that, The process of converting the continuous dense features into sparse discrete tokens by the trained sparse bridging encoder further includes: Dynamic analysis of the downsampling features in both temporal and semantic dimensions is performed to identify dynamic and static speech segments. The extended codebook is called to encode the sub-high-dimensional vectors corresponding to the dynamic speech segments, resulting in the extended RVQ code; The basic RVQ code is obtained by encoding the sub-high-dimensional vectors corresponding to static speech segments using a basic codebook. The extended RVQ code is integrated with the basic RVQ code to generate a complete RVQ code sequence; The code selector of the trained sparse bridging encoder filters the complete RVQ code sequence according to a preset encoding filtering rule to obtain the target RVQ index. Combine all target RVQ indices to obtain sparse discrete tokens.
8. A speech processing device based on dual speech representation, characterized in that, The speech processing device based on dual speech representation includes: The model building and training module is used to pre-build a sparse bridging encoder and a dense bridging decoder, and to jointly train the sparse bridging encoder and the dense bridging decoder. The data acquisition and processing module is used to acquire raw speech data, analyze the raw speech data, and generate continuous dense features; A sparse discrete token module is used to convert the continuous dense features into sparse discrete tokens by a trained sparse bridging encoder. The continuous dense feature reconstruction module is used to reconstruct the features of the sparse discrete tokens by the trained dense bridging decoder to generate reconstructed continuous dense features.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a speech processing program based on dual speech representation stored in the memory and executable on the processor, wherein the speech processing program based on dual speech representation, when executed by the processor, implements the steps of the speech processing method based on dual speech representation as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a speech processing program based on dual speech representation, which, when executed by a processor, implements the steps of the speech processing method based on dual speech representation as described in any one of claims 1-7.