A dialect speech enhancement method and device based on cross-modal and adversarial verification

By fusing audio, lip-movement video, and text features through a cross-modal and adversarial verification framework, and combining them with a dialect knowledge graph, this approach addresses the issue of low-quality dialect speech generation in existing technologies. It achieves high-fidelity, authentic dialect speech generation, making it suitable for dialect protection and inheritance.

CN121528197BActive Publication Date: 2026-04-14UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality, high-fidelity dialect speech. Traditional speech recognition and synthesis systems perform poorly when processing dialects, and existing speech enhancement technologies cannot ensure that the enhanced speech conforms to the rhythm, tone, and cultural context of the dialect.

Method used

By constructing a cross-modal and adversarial verification framework, integrating audio, lip movement video, and text features, and introducing a multi-dimensional discriminator and dialect knowledge graph, multi-level constraints and optimizations for speech generation are achieved, ensuring the authenticity and nativeness of the generated speech.

Benefits of technology

The generated dialect speech not only improves acoustic clarity, but also conforms to the characteristics of specific dialects in terms of rhythm, tone and cultural context, making it suitable for the protection and inheritance of endangered or niche dialects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528197B_ABST
    Figure CN121528197B_ABST
Patent Text Reader

Abstract

A dialect speech enhancement method and device based on cross-modal and adversarial verification belong to the technical field of speech recognition and enhancement. The present application significantly improves the accuracy and naturalness of the dialect generation task by jointly modeling the dialect speech and lip movement video. A closed-loop enhancement framework of generation-adversarial-feedback is constructed, and a multi-dimensional adversarial verification mechanism is introduced to let the model evolve itself and ensure that the enhanced data is not only diverse but also conforms to the real usage habits of the dialect. The dialect knowledge graph and the speech generation model are innovatively integrated to achieve dual enhancement at the semantic and cultural levels, making the generated dialect speech content more consistent with its real language environment and cultural background, especially suitable for the protection and inheritance of endangered or niche dialects. The present application can be directly used in subsequent dialect recognition, synthesis or protection and other applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech recognition and enhancement technology, and particularly relates to a dialect speech enhancement method and apparatus based on cross-modal and adversarial verification. Background Technology

[0002] Dialects, as an important carrier of language and culture, bear rich information about regional history and culture. However, against the backdrop of the rapid development of artificial intelligence and speech technology, many dialects face the severe challenge of the digital divide due to factors such as scarce data resources and unique pronunciation and grammatical systems. Traditional speech recognition and synthesis systems perform poorly when processing dialects, while existing speech enhancement technologies often focus only on acoustic optimization, making it difficult to generate clear and authentic dialect speech, and easily losing the unique rhythm, vocabulary, and cultural context of dialects.

[0003] Patent CN120600003A primarily focuses on downstream applications of dialect speech, specifically how to transcribe existing speech data into text or perform language conversion. The core of these methods lies in the recognition and conversion model itself, and their performance is highly dependent on the scale and quality of the upstream training data. However, this type of patent does not address the fundamental problem of sparse and low-quality dialect data. When faced with dialects lacking sufficient resources, the recognition and conversion effects are significantly reduced. This invention aims to solve this front-end bottleneck by generating high-quality, high-fidelity dialect data, providing a solid data foundation for such recognition and conversion technologies.

[0004] Patent CN114360561A proposes a general acoustic enhancement scheme. Such methods typically operate at the speech signal level, such as noise reduction and reverberation removal, aiming to improve speech clarity. However, this simple acoustic enhancement is far from sufficient for dialect processing. First, it is a single-modal processing method, completely ignoring visual information such as lip movements, which are crucial in dialect recognition. Second, its optimization goal is acoustic fidelity, not dialect authenticity; it cannot ensure that the enhanced speech conforms to the characteristics of a specific dialect in terms of prosody, tone, and even usage, potentially producing speech that is clear to the ear but lacks dialectal flavor. This invention, by introducing cross-modal information and an adversarial verification framework including a prosody / tone discriminator, achieves deep enhancement from acoustic clarity to the fidelity of dialect prosody and linguistic features.

[0005] Patent CN120471109A utilizes large models to expand the data volume. While such methods can generate massive amounts of diverse data, they are often an open-loop generation process, lacking effective verification of the authenticity of the generated content. Especially in scenarios with strict language rules and cultural backgrounds, such as dialects, without effective monitoring and feedback, large models are likely to generate speech content that does not conform to dialect habits or is even erroneous, causing data pollution.

[0006] The core advantage of this invention lies in constructing a multi-dimensional self-evolutionary enhancement loop. By introducing a multi-dimensional discriminator group and a dialect knowledge graph for dual constraints, it ensures that the augmented data is not only improved in quantity and diversity, but also reliably guaranteed in quality and the fidelity of dialect prosody and linguistic features, achieving a leap from simple data expansion to controllable and quality-assured data augmentation. Summary of the Invention

[0007] To overcome the shortcomings of the existing technologies, this invention systematically integrates cross-modal information, adversarial verification mechanisms, and knowledge graphs to create a novel method for generating high-fidelity, highly diverse, and highly authentic dialect speech data, fundamentally solving the data dilemma faced by dialects in the digital age.

[0008] The core subject of this invention is a dialect speech enhancement method and apparatus based on cross-modal and adversarial verification. This invention aims to address the core problems faced by dialects, especially low-resource dialects, such as sparse and low-quality speech data, and insufficient adaptability of the enhanced speech. To achieve this goal, this solution proposes a comprehensive technical framework. Firstly, this framework includes fine-grained alignment and fusion of cross-modal information, using audio features as the primary modality and dynamically adjusting the fusion weights of text and visual features to enhance speech-lip-sync consistency and semantic-speech matching, making it particularly suitable for dialect scenarios where pronunciation and lip-sync are highly correlated. A contrastive learning loss is introduced to bring matching sample pairs closer together and push away mismatched sample pairs, achieving fine-grained alignment of multimodal features at the semantic level and improving the discriminative power and stability of the fused representation. Secondly, this framework establishes a multi-dimensional collaborative adversarial verification mechanism, constructing a multi-dimensional collaborative dialect self-evolutionary adversarial kernel. The innovation of this invention lies in its mechanism, which transcends the traditional single-dimensional verification of acoustic authenticity. Instead, it uses a set of collaborative discriminators to continuously detect the authenticity of generated speech (including phonetic features and dialectal usage) from multiple orthogonal dimensions, such as acoustic fidelity, dialectal prosody and tone, and semantic consistency between content and the knowledge graph. The composite feedback results are used to back-optimize the generation model, allowing it to self-evolve and achieve deep enhancement from "clearly audible" to "sounds like" and even "speaks correctly," continuously improving the authenticity of dialectal speech. Finally, this framework is knowledge graph-driven. By constructing a dialect knowledge graph encompassing phonetic features, grammatical rules, typical vocabulary, and cultural context, and integrating it with a larger model, speech enhancement is no longer limited to the signal and text levels. Instead, it achieves a two-layer enhancement of semantics and culture, ensuring that the generated speech content conforms to the real-world usage environment and cultural background of the dialect.

[0009] To solve the above-mentioned technical problems, the specific technical solution of the present invention is as follows:

[0010] A dialect speech enhancement method based on cross-modal and adversarial verification includes the following steps:

[0011] Step 1: Multimodal data representation and knowledge embedding; Parallel feature extraction is performed on temporal data of dialect audio, lip movement video, and text transcription, and then these three features are fused; At the same time, a dialect knowledge graph KG is pre-constructed based on the dialect knowledge base;

[0012] Step 2: Using the fused features and dialect knowledge graph, synthesize audio through Transformer Decoder and vocoder;

[0013] Step 3: Multi-dimensional adversarial training and optimization; the acoustic discriminator discriminates the synthesized audio at different time scales; the prosody / tone discriminator first extracts the contour and energy information of the fundamental frequency of the speech, and then judges whether its prosody and tone patterns conform to the characteristics of the target dialect; the semantic / KG consistency discriminator first converts the generated audio into text through a pre-trained speech recognition model, and then judges whether the content of the text is consistent with the semantic and cultural context represented by the knowledge graph; the total loss function of training includes adversarial loss and auxiliary loss;

[0014] Step 4: Use the trained model to perform an enhanced speech inference task.

[0015] Step 1 is described in detail as follows:

[0016] Step 1.1: Parallel encoding of multimodal features; the dialect audio sequence, lip movement video sequence, and text transcription sequence are fed into a dedicated encoder for feature extraction to obtain acoustic representations. dynamic characteristics of mouth shape Semantic representation ;

[0017] Step 1.2: Cross-modal feature fusion; a cross-attention Transformer module is used to implement feature fusion:

[0018]

[0019] in, This indicates a cross-attention Transformer module. For joint representation, Q, K, and V represent the query vector, key vector, and value vector of the attention mechanism, respectively;

[0020] Step 1.3: Dialect Knowledge Graph Embedding; Load a pre-built dialect knowledge graph. The nodes of this graph contain the vocabulary, pronunciation rules, grammatical structure and cultural context of the dialect. A graph neural network (GNN) is used to learn the structured knowledge of the graph. The GNN updates the representation of the central node by iteratively aggregating the information of neighboring nodes. The knowledge of the entire graph or the subgraph related to the input text is encoded into a unified knowledge embedding vector.

[0021] Step 2 is described in detail below:

[0022] Step 2.1: Spectrum generation, joint representation, and dialect knowledge graph embedding based on Transformer Decoder Random noise vector The input is fed into the generator, which generates Mel spectra frame by frame. When generating the spectrum of the t-th frame, the model will simultaneously focus on the input information and the spectrum of the previous t-1 frames.

[0023] Step 2.2: GAN-based vocoder synthesis, which inversely synthesizes the Mel spectrum sequence into an audio waveform.

[0024] The acoustic discriminator adopts a multi-scale CNN structure, the prosody / tone discriminator is a long short-term memory network, and the semantic / KG consistency discriminator is a Transformer-based scorer.

[0025] The antagonistic loss It is expressed as follows:

[0026]

[0027] in, For discriminators, specifically including acoustic discriminators Prosody / Pitch Discriminator Semantic / KG Consistency Discriminator z is a random noise vector, and G is the generator;

[0028] The auxiliary loss The L1 loss is calculated using the following formula:

[0029]

[0030] in, The true Mel spectrum y and the corresponding generator G are inputs. The expected value of the joint data distribution.

[0031] Step 1 introduces a contrastive learning loss during the training process, as shown in the following formula:

[0032]

[0033] in, To compare learning loss, Let represent the acoustic characterization of the i-th sample. This represents the dynamic lip-shape feature of the i-th sample; Represents cosine similarity. Let be the temperature hyperparameter, exp be the exponential function, and N be the total number of samples in the current training batch.

[0034] The present invention also provides a dialect speech enhancement device based on cross-modal and adversarial verification, comprising:

[0035] The data acquisition and preprocessing module is responsible for processing three data sources in parallel: dialect audio, lip movement video, and text transcription. After preprocessing, these raw data are sent to the feature encoding and extraction layer.

[0036] The feature encoding and extraction layer consists of three independent encoders—audio encoder, visual / lip movement encoder, and text encoder—working in parallel to extract deep features from their respective preprocessed data, which are then fed into the cross-modal alignment and fusion module.

[0037] The cross-modal alignment and fusion module fuses multimodal data such as speech, text, and video to generate a joint representation;

[0038] The dialect knowledge graph module constructs a structured knowledge base containing vocabulary nodes, pronunciation nodes, grammar nodes, and cultural context nodes, and embeds this knowledge into its representation.

[0039] The generator receives the joint representation from the fusion module and the knowledge embedding from the knowledge graph as its input, and generates an intermediate representation of the target dialect speech through its deep network structure. The generated intermediate representation is then fed into the vocoder module.

[0040] The vocoder module converts the intermediate representation of dialect speech into the final audio waveform;

[0041] The multi-dimensional discriminator group, consisting of an acoustic discriminator, a prosody / tone discriminator, and a semantic / KG consistency discriminator, comprehensively evaluates the speech generated by the generator and feeds back the evaluation results to the generator in the form of a loss function or a semantic consistency signal, forcing the generator to continuously optimize.

[0042] The main beneficial effects of the dialect speech enhancement method and device based on cross-modal and adversarial verification provided by this invention are as follows:

[0043] 1. This invention proposes a dialect speech enhancement method across speech and image modes, which significantly improves the accuracy and naturalness of dialect generation tasks by jointly modeling dialect speech and lip movement video.

[0044] 2. This invention constructs a closed-loop enhancement framework of generation-adversarial-feedback. By introducing a multi-dimensional adversarial verification mechanism, the model is allowed to self-evolve, ensuring that the enhanced data is not only diverse but also conforms to the real usage habits of dialects.

[0045] 3. This invention innovatively integrates dialect knowledge graphs with speech generation models, achieving dual enhancement at the semantic and cultural levels. This makes the generated dialect speech content more consistent with its real language environment and cultural background, and is especially suitable for the protection and inheritance of endangered or niche dialects. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the method structure of the present invention;

[0047] Figure 2 A schematic diagram of parallel encoding of multimodal features;

[0048] Figure 3 This is a schematic diagram of GNN neighbor information aggregation;

[0049] Figure 4 A schematic diagram of the knowledge-guided speech spectrum generation model process. Detailed Implementation

[0050] Figure 1 The overall technical architecture of this invention is illustrated. This invention provides a dialect speech enhancement device based on cross-modal and adversarial verification, which systematically addresses various challenges in dialect speech data enhancement through an end-to-end deep learning framework, specifically including:

[0051] The data acquisition and preprocessing module is responsible for processing three data sources in parallel: dialect audio, lip movement video, and text transcription. After a series of preprocessing operations such as data cleaning, slicing, and scene labeling, these raw data are sent to the feature encoding and extraction layer.

[0052] Feature encoding and extraction layer: Three independent encoders—audio encoder, visual / lip movement encoder, and text encoder—work in parallel, each extracting deep features from their respective preprocessed data. These heterogeneous feature vectors are then fed into the cross-modal alignment and fusion module;

[0053] Cross-modal alignment and fusion module: Effectively fuses multimodal data such as speech, text, and video to generate a unified and information-rich joint representation.

[0054] Dialect Knowledge Graph (KG) module: By constructing a structured knowledge base containing vocabulary nodes, pronunciation nodes, grammar nodes, and cultural context nodes, and embedding these knowledge into the representation, it provides deep linguistic and cultural background constraints for speech generation.

[0055] The generator receives the joint representation from the fusion module and the knowledge embedding from the knowledge graph as input. Based on these two powerful information sources, the generator produces intermediate representations (e.g., Mel spectrograms) of the target dialect speech through its deep network structure. The resulting Mel spectrograms and other intermediate representations are then fed into the vocoder module.

[0056] Multi-dimensional Discriminator Group: This discriminator group consists of an acoustic discriminator, a prosody / tone discriminator, and a semantic / KG consistency discriminator. It comprehensively evaluates the speech generated by the generator and feeds back the evaluation results to the generator in the form of a loss function or semantic consistency signal, forcing the generator to continuously optimize and produce more realistic speech. Through multi-dimensional self-evolution enhancement loop training, the authenticity of the generated speech is ensured.

[0057] The vocoder module converts the Mel spectrum into a final audio waveform with high fidelity. This waveform, as the output, is the "high-fidelity, authentic, and scene-diverse" enhanced dialect speech that this invention aims to obtain, which can directly serve downstream applications such as speech synthesis, recognition, and dialect protection.

[0058] The specific implementation process of the dialect speech enhancement method based on cross-modal and adversarial verification provided by this invention follows the following... Figure 1 The overall block diagram shown illustrates an architecture that covers the entire process from multimodal data representation to knowledge-guided generation, and then to adversarial optimization training, specifically including the following steps:

[0059] Step 1: Multimodal data representation and knowledge embedding; the goal of this stage is to transform multi-source, heterogeneous input data into unified, information-rich feature vectors that can be understood by the generative model.

[0060] Step 1.1: Parallel encoding of multimodal features;

[0061] The input data contains three types of parallel time-series data: dialect audio sequences. lip movement video sequence and text transcription sequences These three data streams are each fed into a dedicated encoder for feature extraction, and their parallel processing mechanism is as follows: Figure 2 As shown.

[0062] The audio encoder employs a self-supervised pre-trained model, such as Wav2Vec2, to directly extract context-dependent acoustic representations from the raw audio waveform. This representation captures rich phonemic, timbre, and prosodic information. The visual encoder uses a visual model, such as a 3D convolutional neural network (3D-CNN) or Transformer, to process the lip movement video sequence and extract key dynamic features of the mouth shape. The text encoder uses a standard Transformer encoder to convert text sequences into high-dimensional semantic representations. .

[0063] Step 1.2: Cross-modal feature fusion; In order for the model to comprehensively utilize the above three types of information, it is necessary to represent the independent features. , , To achieve effective alignment and integration.

[0064] This approach employs a Cross-Attention Transformer module to achieve this goal. This module, primarily driven by audio features, dynamically integrates semantic and lip-sync information from text and visual features into the acoustic representation, generating the final joint multimodal representation. This process can be represented as:

[0065]

[0066] in This indicates a cross-attention Transformer module. For joint representation, Q, K, and V represent the query vector, key vector, and value vector of the attention mechanism, respectively. To ensure effective alignment between different modalities, a contrastive loss is introduced during training. This loss function aims to shorten the distance between matched (audio-text-video) sample pairs in the feature space while distancing mismatched sample pairs. Its calculation formula is as follows:

[0067]

[0068] in, To compare learning loss, Let represent the acoustic characterization of the i-th sample. This represents the dynamic lip-shape feature of the i-th sample; Represents cosine similarity. Let be the temperature hyperparameter, exp be the exponential function, and N be the total number of samples in the current training batch.

[0069] Step 1.3: Dialect Knowledge Graph Embedding; In parallel with multimodal feature processing, a pre-built dialect knowledge graph (KG) is loaded. The nodes of this graph contain information such as vocabulary, pronunciation rules, grammatical structure, and cultural context of the dialect. This scheme uses a graph neural network (GNN), such as GraphSAGE, to learn the structured knowledge of this graph. The GNN updates the representation of the central node by iteratively aggregating neighbor node information; its single-layer aggregation process is as follows: Figure 3 As shown, after multiple aggregations, each node obtains an embedding vector containing information about its neighborhood structure. Ultimately, the entire graph or the subgraph knowledge related to the input text is encoded into a unified knowledge embedding vector. .

[0070] Step 2: Knowledge-guided speech generation. This stage is the core generation step, which uses the fusion features extracted in the previous stage to generate high-quality intermediate speech representations.

[0071] Step 2.1: Spectrum Generation Based on Transformer Decoder; The generator in this scheme adopts an autoregressive structure similar to Transformer Decoder. Its input consists of three parts: the joint representation obtained after cross-modal fusion. Dialect knowledge graph embedding Random noise vector (Used to increase generation diversity). The generator generates Mel spectrograms frame by frame. When generating the spectrum yt of frame t, the model also considers guiding information. And the spectrum of the first t-1 frames that have already been generated. Its structure is as follows: Figure 4 As shown.

[0072] Step 2.2: GAN-based vocoder synthesis; The Mel spectrum generated by the generator is a compact acoustic representation that needs to be converted into the final audio waveform by a vocoder. This scheme uses a vocoder based on a Generative Adversarial Network (GAN), such as HiFi-GAN. This vocoder is specially trained to efficiently and faithfully inversely synthesize the Mel spectrum sequence into a natural-sounding audio waveform.

[0073] Step 3: Multi-dimensional adversarial training and optimization; To ensure that the generated dialect speech is not only clear, but also retains high fidelity in terms of dialect prosody and linguistic features, this scheme designs an adversarial training framework containing multiple professional discriminators. The goal of generator G is to generate speech that can fool all discriminators D, while the goal of the discriminators is to distinguish between real speech and generated speech as accurately as possible.

[0074] Step 3.1: Multi-dimensional discriminator group; this discriminator group consists of three functionally independent modules. Acoustic discriminator Employing a multi-scale CNN structure, it operates directly on the original waveform to determine the speech quality, realism, and presence of artifacts at different time scales. Prosody / Pitch Discriminator First, the contour and energy information of the fundamental frequency F0 of the speech are extracted. Then, a network based on Long Short-Term Memory (LSTM) is used to determine whether its prosody and tone patterns conform to the characteristics of the target dialect. Semantic / KG Consistency Discriminator First, the generated speech is converted into text using a pre-trained speech recognition model. Then, a Transformer-based scorer is used to determine whether the text content is consistent with the semantic and cultural context represented by the knowledge graph embedding.

[0075] Step 3.2: Composite Loss Function; Model training is guided by a composite loss function. The total loss of the generator G. Including adversarial losses and auxiliary losses :

[0076]

[0077] in, These are weighting coefficients; It is the sum of adversarial losses from the three discriminators, designed to make the generated results more acoustically, prosodicly, and semantically similar to real data:

[0078]

[0079] in, For discriminators, specifically including acoustic discriminators Prosody / Pitch Discriminator Semantic / KG Consistency Discriminator z is a random noise vector, and G is the generator;

[0080] It is an auxiliary L1 loss used to constrain the generated Mel spectrum to remain similar to the true Mel spectrum in the initial stage, thus accelerating model convergence.

[0081]

[0082] Given the true Mel spectrum y and the corresponding generator input Calculate the expected value of the joint data distribution.

[0083] Accordingly, each discriminator The loss is the standard GAN discriminator loss, designed to distinguish real samples. and generate samples .

[0084] Step 4: Enhance speech inference;

[0085] After model training is complete, the inference phase begins. At this stage, only the feature encoding module, knowledge embedding module, generator, and vocoder are used. The discriminator no longer participates in the computation. The system receives input dialect audio, video, and text, and through the aforementioned forward computation process, directly generates and outputs enhanced, high-fidelity, and authentic dialect speech. This output can be directly used in subsequent applications such as dialect recognition, synthesis, or preservation.

[0086] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A dialect speech enhancement method based on cross-modal and adversarial verification, characterized in that, Includes the following steps: Step 1: Multimodal data representation and knowledge embedding; Parallel feature extraction is performed on temporal data of dialect audio, lip movement video, and text transcription, and then these three features are fused; At the same time, a dialect knowledge graph KG is pre-constructed based on the dialect knowledge base; Step 2: Utilize the fused features and dialect knowledge graph to synthesize audio through Transformer Decoder and vocoder; specifically, the generator adopts the autoregressive structure of Transformer Decoder, receives the fused features and dialect knowledge graph embedding as its input, generates an intermediate representation of the target dialect speech through its deep network structure, and the generated intermediate representation is sent to the vocoder. Step 3: Multi-dimensional adversarial training and optimization; the acoustic discriminator discriminates the synthesized audio at different time scales; the prosody / tone discriminator first extracts the contour and energy information of the fundamental frequency of the speech, and then judges whether its prosody and tone patterns conform to the characteristics of the target dialect. The semantic / KG consistency discriminator first converts the generated audio into text through a pre-trained speech recognition model, and then determines whether the text content is consistent with the semantic and cultural context represented by the knowledge graph; the total loss function of training includes adversarial loss and auxiliary loss; Step 4: Use the trained model to perform an enhanced speech inference task.

2. The dialect speech enhancement method based on cross-modal and adversarial verification according to claim 1, characterized in that, Step 1 is described in detail as follows: Step 1.1: Multi-modal feature parallel encoding; the dialect audio sequence, the lip movement video sequence, and the text transcription sequence are respectively sent into a special encoder for feature extraction to obtain acoustic representation , mouth dynamic feature , semantic representation ; Step 1.2: Cross-modal feature fusion; Feature fusion is achieved using a cross-attention Transformer module: ; wherein, denotes a cross-attention Transformer module, is a joint representation, Q, K, V denote query vector, key vector, value vector of the attention mechanism, respectively; Step 1.3: Dialect Knowledge Graph Embedding; Load a pre-built dialect knowledge graph. The nodes of this graph contain the vocabulary, pronunciation rules, grammatical structure and cultural context of the dialect. Use a graph neural network (GNN) to learn the structured knowledge of the graph. The GNN updates the representation of the central node by iteratively aggregating the information of neighboring nodes. The knowledge of the entire graph or the subgraph related to the input text is encoded into a unified knowledge embedding vector.

3. The dialect speech enhancement method based on cross-modal and adversarial verification according to claim 2, characterized in that, Step 2 is described in detail below: Step 2.1: Spectrum generation, joint representation, and dialect knowledge graph embedding based on Transformer Decoder Random noise vector The input is fed into the generator, which generates Mel spectra frame by frame. When generating the spectrum of the t-th frame, the model will simultaneously focus on the input information and the spectrum of the previous t-1 frames. Step 2.2: GAN-based vocoder synthesis, which inversely synthesizes the Mel spectrum sequence into an audio waveform.

4. The dialect speech enhancement method based on cross-modal and adversarial verification according to claim 3, characterized in that, The acoustic discriminator adopts a multi-scale CNN structure, the prosody / tone discriminator is a long short-term memory network, and the semantic / KG consistency discriminator is a Transformer-based scorer.

5. A dialect speech enhancement method based on cross-modal and adversarial verification according to claim 4, characterized in that, The antagonistic loss It is expressed as follows: ; in, For discriminators, specifically including acoustic discriminators Prosody / Pitch Discriminator Semantic / KG Consistency Discriminator z is a random noise vector, and G is the generator; The auxiliary loss The L1 loss is calculated using the following formula: ; in, The true Mel spectrum y and the corresponding generator G are inputs. The expected value of the joint data distribution.

6. A dialect speech enhancement method based on cross-modal and adversarial verification according to claim 5, characterized in that, Step 1 introduces a contrastive learning loss during the training process, as shown in the following formula: ; in, To compare learning loss, Let represent the acoustic characterization of the i-th sample. This represents the dynamic lip-shape feature of the i-th sample; Represents cosine similarity. Let be the temperature hyperparameter, exp be the exponential function, and N be the total number of samples in the current training batch.

7. A dialect speech enhancement device based on cross-modal and adversarial verification, characterized in that, include: The data acquisition and preprocessing module is responsible for processing three data sources in parallel: dialect audio, lip movement video, and text transcription. After preprocessing, these raw data are sent to the feature encoding and extraction layer. The feature encoding and extraction layer consists of three independent encoders—an audio encoder, a visual / lip movement encoder, and a text encoder—that work in parallel. Each encoder extracts deep features from its corresponding preprocessed data, which are then fed into the cross-modal alignment and fusion module. The cross-modal alignment and fusion module fuses multimodal data such as speech, text, and video to generate a joint representation; The dialect knowledge graph module constructs a structured knowledge base containing vocabulary nodes, pronunciation nodes, grammar nodes, and cultural context nodes, and embeds this knowledge into its representation. The generator receives the joint representation from the fusion module and the knowledge embedding from the knowledge graph as its input, and generates an intermediate representation of the target dialect speech through its deep network structure. The generated intermediate representation is then fed into the vocoder module. The vocoder module converts the intermediate representation of dialect speech into the final audio waveform; The multi-dimensional discriminator group, consisting of an acoustic discriminator, a prosody / tone discriminator, and a semantic / KG consistency discriminator, comprehensively evaluates the speech generated by the generator and feeds back the evaluation results to the generator in the form of a loss function or a semantic consistency signal, forcing the generator to continuously optimize.

Citation Information

Patent Citations

  • Large model data enhancement method and device

    CN120471109A

  • Dialect speech recognition and conversion method and device

    CN120600003A

  • English adaptive learning system and method based on intelligent context perception

    CN121525751A