Voiceprint synthesis method and device, electronic equipment and storage medium
By using an identity-preserving generation network and a noise mapping mechanism driven by physical parameters, the problems of identity distortion and scene falsification in existing voiceprint synthesis methods are solved. This achieves the preservation of specific person identity features and high-fidelity reproduction of battlefield noise acoustic characteristics in synthesized speech, thereby improving the operator's recognition ability and stress response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-13
Smart Images

Figure CN121662053A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of defense communication technology, and in particular to a voiceprint synthesis method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, in defense communication systems, human operators are a core component ensuring the operational command and daily communication of specific personnel. Operators must accurately identify and transfer instructions from specific individuals in complex and ever-changing battlefield environments; therefore, their speech recognition capabilities are crucial. To improve operators' ability to recognize the voices of specific individuals, daily training using synthesized speech is necessary.
[0003] However, existing voiceprint synthesis methods have significant shortcomings: on the one hand, when introducing emotions or environmental noise during the synthesis process, it is easy to distort the biological characteristics of a specific person (such as formant distribution), resulting in "identity distortion"; on the other hand, randomly adding noise makes it difficult to realistically reproduce the acoustic environment of the battlefield (such as armored vehicle engines, electromagnetic interference, etc.), resulting in "fake scenes," and the synthesized training data is out of sync with the actual environment, affecting the training effectiveness. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a voiceprint synthesis method, device, electronic device and storage medium, and to propose an identity-preserving generation network and a noise mapping mechanism driven by physical parameters, which can maintain the identity features of specific persons in synthesized speech and reproduce the acoustic characteristics of battlefield noise with high fidelity, thereby providing operators with more realistic and effective training speech samples and enhancing their recognition ability and stress response in actual combat environments.
[0005] In a first aspect, embodiments of the present invention provide a voiceprint synthesis method, the method comprising: determining standardized speech features based on original speaker speech segments and constructing a battlefield noise physical parameter library; determining synthesized speech and speech quality perception evaluation indicators through an identity-preserving generation network based on standardized speech features, emotion tags, and noise parameters; determining decoupled clean identity features based on the log-Mel spectrum of synthesized speech and real speech through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training; training a voiceprint recognition model based on the clean identity features and corresponding speaker tags; and updating the voiceprint recognition model based on real-time speaker speech streams.
[0006] In an optional embodiment of this application, the steps of determining standardized speech features and constructing a battlefield noise physical parameter library based on the original speaker's speech segments include: performing speech preprocessing and basic feature extraction sequentially on the original speaker's speech segments to obtain standardized speech features; wherein, speech preprocessing includes: pre-emphasis, framing, and windowing, and standardized speech features include: Mel frequency cepstral coefficients, fundamental frequency, spectral envelope, and formant trajectory; and constructing a battlefield noise physical parameter library based on standardized speech features in a physical parameter-driven manner.
[0007] In optional embodiments of this application, the aforementioned identity preservation generation network is used for identity feature anchoring, condition generation, and multi-objective optimization.
[0008] In optional embodiments of this application, the weights of the parameters are fixed during the above-mentioned identity feature anchoring process; during the identity feature anchoring process, the generator adopts a U-Net structure and directly transmits the underlying identity information extracted by the encoder to the decoder through skip connections; the condition generation process includes: sentiment condition injection and noise physical mapping; the total loss function of multi-objective optimization is determined based on identity preservation loss, adversarial loss and reconstruction loss.
[0009] In optional embodiments of this application, the above-mentioned dual-path orthogonal attention mechanism includes: an identity attention subnetwork and a noise and emotion suppression subnetwork; the fusion weights of the dynamic gating fusion are generated by the Sigmoid function and adaptively fuse the features of the dual-path orthogonal attention mechanism; the adversarial decoupling training introduces a gradient reversal layer, and the total loss function of the adversarial decoupling training is determined based on identity preservation loss, scene classification loss and emotion classification loss.
[0010] In optional embodiments of this application, the architecture of the above-mentioned voiceprint recognition model includes: a voiceprint recognition network based on ECAPA-TDNN; the training strategy of the voiceprint recognition model includes: training using a Softmax loss function with additional angular intervals, and multi-scene hybrid training.
[0011] In an optional embodiment of this application, the step of updating the voiceprint recognition model based on the real-time speaker's speech stream includes: inputting the real-time speaker's speech stream into the voiceprint recognition model and outputting the real-time speaker's physiological state; wherein, the physiological state includes: fundamental frequency jitter rate, formant bandwidth expansion, and glottal closure quotient; performing confidence assessment based on the physiological state; and updating the voiceprint recognition model in a closed loop based on the result of the confidence assessment.
[0012] Secondly, embodiments of the present invention also provide a voiceprint synthesis device, comprising: a data preprocessing and parameter library construction module, used to determine standardized speech features based on original speaker speech segments and construct a battlefield noise physical parameter library; an identity-preserving speech synthesis module, used to determine synthesized speech and speech quality perception evaluation indicators based on standardized speech features, emotion tags, and noise parameters through an identity-preserving generation network; a scene emotion feature decoupling module, used to determine the decoupled pure identity features based on the log-Mel spectrum of synthesized speech and real speech through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training; a voiceprint recognition model training and optimization module, used to train a voiceprint recognition model based on the pure identity features and corresponding speaker tags; and a dynamic voiceprint update mechanism module, used to update the voiceprint recognition model based on real-time speaker speech streams.
[0013] Thirdly, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-described voiceprint synthesis method.
[0014] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the above-described voiceprint synthesis method.
[0015] The embodiments of the present invention bring the following beneficial effects: This invention provides a voiceprint synthesis method, apparatus, electronic device, and storage medium. It determines standardized speech features based on original speaker speech segments and constructs a battlefield noise physical parameter library. Based on standardized speech features, emotion tags, and noise parameters, it determines synthesized speech and speech quality perception evaluation indicators through an identity-preserving generation network. Based on the log-Mel spectrum of synthesized and real speech, it determines decoupled clean identity features through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training. It trains a voiceprint recognition model based on the clean identity features and corresponding speaker tags. The voiceprint recognition model is updated based on real-time speaker speech streams. This approach proposes an identity-preserving generation network and a noise mapping mechanism driven by physical parameters, which can maintain specific person identity features in synthesized speech and reproduce the acoustic characteristics of battlefield noise with high fidelity. This provides operators with more realistic and effective training speech samples, enhancing their recognition ability and stress response in real combat environments.
[0016] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0017] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A flowchart of a voiceprint synthesis method provided in an embodiment of the present invention; Figure 2 A flowchart of another voiceprint synthesis method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a voiceprint synthesis method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the overall process of a closed-loop voiceprint management system provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a voiceprint synthesis device provided in an embodiment of the present invention.
[0020] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Currently, in defense communication systems, human operators are a core component ensuring the operational command and daily communication of specific personnel. Operators must accurately identify and transfer instructions from specific individuals in complex and ever-changing battlefield environments; therefore, their speech recognition capabilities are crucial. To improve operators' ability to recognize the voices of specific individuals, daily training using synthesized speech is necessary.
[0023] However, existing voiceprint synthesis methods have significant shortcomings: on the one hand, when introducing emotions or environmental noise during the synthesis process, it is easy to distort the biological characteristics of a specific person (such as formant distribution), resulting in "identity distortion"; on the other hand, randomly adding noise makes it difficult to realistically reproduce the acoustic environment of the battlefield (such as armored vehicle engines, electromagnetic interference, etc.), resulting in "fake scenes," and the synthesized training data is out of sync with the actual environment, affecting the training effectiveness.
[0024] Based on this, the embodiments of the present invention provide a voiceprint synthesis method, device, electronic device, and storage medium, specifically providing a dynamic voiceprint synthesis method for training military operators that has high environmental robustness and emotional adaptability. It proposes an identity-preserving generative network (e.g., IPM-cGAN (IPM Conditional Generative Adversarial Network)) and a noise mapping mechanism driven by physical parameters, which can maintain specific identity features (cosine similarity ≥ 0.93) in synthesized speech and reproduce the acoustic characteristics of battlefield noise with high fidelity (PESQ (Perceptual Evaluation of Speech Quality) ≥ 3.8), thereby providing operators with more realistic and effective training speech samples and enhancing their recognition ability and stress response in actual combat environments.
[0025] To facilitate understanding of this embodiment, a voiceprint synthesis method disclosed in this embodiment of the invention will first be described in detail.
[0026] Example 1: This invention provides a voiceprint synthesis method, specifically a voiceprint synthesis method based on tone and context. (See also...) Figure 1 The flowchart shown illustrates a voiceprint synthesis method, which includes the following steps: Step S102: Determine standardized speech features based on the original speaker's speech segments and construct a database of battlefield noise physical parameters.
[0027] In this embodiment, a closed-loop voiceprint management system can be used for voiceprint synthesis. This closed-loop voiceprint management system can be an automated, cyclical system that includes data generation, model training, real-time updates, and application.
[0028] In this embodiment, the speaker can be a specific person. Data preprocessing and parameter library construction can be performed based on the original speech segments of the specific person to determine standardized speech features and construct a battlefield noise physical parameter library.
[0029] Step S104: Based on standardized speech features, emotion tags, and noise parameters, an identity-preserving generation network is used to determine the synthesized speech and speech quality perception evaluation indicators.
[0030] In this embodiment, the identity-preserving generative network can be an IPM-cGAN network, which is an improved generative adversarial network.
[0031] Generative Adversarial Networks (GANs) are deep learning models consisting of two neural networks: a generator and a discriminator. They generate realistic data through adversarial training.
[0032] The generator can take random noise as input to generate fake data, with the goal of making the fake data as close as possible to the distribution of real data in order to deceive the discriminator.
[0033] The discriminator can determine whether the input data is real or faked by the generator, and outputs a probability value (0 for false, 1 for true).
[0034] Adversarial training refers to the alternating optimization of the generator and the discriminator. The generator tries to improve the authenticity of the generated data, while the discriminator strives to improve its discrimination ability, eventually reaching a dynamic equilibrium (Nash equilibrium).
[0035] Step S106: Based on the log-Mel spectrum of synthesized speech and real speech, the decoupled clean identity features are determined through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training.
[0036] In this embodiment, a Scene-Emotion Decoupling Frontend (SEDF, a processing module used to separate identity and interference information in speech) can be used for adversarial decoupling training. Adversarial decoupling training is an advanced training strategy in Generative Adversarial Networks (GANs). Its core idea is to separate different features (such as color, shape, texture, etc.) in the data through adversarial mechanisms, enabling the generator to independently control these features, thereby improving the diversity and interpretability of the generated samples.
[0037] Step S108: Train a voiceprint recognition model based on pure identity features and corresponding speaker labels.
[0038] This embodiment can train a voiceprint recognition model based on clean identity features and corresponding speaker tags. A voiceprint recognition model is an algorithmic system used to analyze unique biometric features (such as vocal cord vibration and vocal tract structure) in speech signals. Its core objective is to determine whether a specific voice is spoken by a specific person. This model constructs an individual's "vocal fingerprint" by extracting parameters such as fundamental frequency and formants, and uses similarity comparison (such as cosine similarity) to complete identity verification.
[0039] Step S110: Update the voiceprint recognition model based on the real-time speaker's speech stream.
[0040] This embodiment employs a dynamic voiceprint update mechanism, which can dynamically update the voiceprint recognition model based on the real-time speaker's speech stream, enabling the voiceprint model to adapt to changes.
[0041] This invention provides a voiceprint synthesis method. It determines standardized speech features based on original speaker speech segments and constructs a battlefield noise physical parameter library. Based on the standardized speech features, emotion tags, and noise parameters, an identity-preserving generation network determines the synthesized speech and speech quality perception evaluation indicators. Based on the log-Mel spectrum of synthesized and real speech, a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training are used to determine the decoupled clean identity features. A voiceprint recognition model is trained based on the clean identity features and corresponding speaker tags. The voiceprint recognition model is updated based on real-time speaker speech streams. This method proposes an identity-preserving generation network and a noise mapping mechanism driven by physical parameters, which can maintain specific person identity features in synthesized speech and reproduce the acoustic characteristics of battlefield noise with high fidelity. This provides operators with more realistic and effective training speech samples, enhancing their recognition ability and stress response in real combat environments.
[0042] Example 2: This embodiment provides another voiceprint synthesis method, which is implemented based on the above embodiment. The focus is on describing the specific implementation of the voiceprint synthesis method based on tone and scene. See also... Figure 2 The flowchart shown represents another voiceprint synthesis method, which includes the following steps: Step S202: Perform speech preprocessing and basic feature extraction on the original speaker's speech segments in sequence to obtain standardized speech features; based on the standardized speech features, construct a battlefield noise physical parameter library in a physical parameter-driven manner.
[0043] See also Figure 3 The diagram shows a voiceprint synthesis method. This embodiment can perform data preprocessing and parameter library construction. Its input can be: original speaker speech segments and battlefield noise physical parameter library (containing the physical characteristics of 12 types of noise, such as shock wave attenuation coefficient and harmonic distortion).
[0044] This embodiment can perform speech preprocessing, which includes pre-emphasis, framing, and windowing.
[0045] 1. Pre-emphasis (coefficient 0.97): Boosts high-frequency components to compensate for high-frequency attenuation during speech signal production.
[0046] 2. Framing (frame length 25ms, frame shift 10ms): Continuous speech is cut into short segments, assuming that the speech signal is stable within a short period of time.
[0047] 3. Add window (Hamming window): Reduce spectrum leakage caused by framing.
[0048] This embodiment can also perform basic feature extraction, wherein the standardized speech features include: Mel frequency cepstral coefficients, fundamental frequency, spectral envelope, and formant trajectory.
[0049] 1. MFCC (Mel-Frequency Cepstral Coefficients, 39-dimensional): Speech features that simulate the auditory characteristics of the human ear and are widely used in voiceprint recognition.
[0050] 2. F0 (fundamental frequency): The basic frequency of speech, corresponding to the frequency of vocal cord vibration, and related to pitch.
[0051] 3. Spectral envelope: describes the overall shape of the speech signal spectrum and is related to the shape of the vocal organs.
[0052] 4. Formant trajectory: The peak frequency in the spectral envelope reflects the resonance characteristics of the vocal tract and is an important identification feature.
[0053] This embodiment can also construct a parameter library driven by physical parameters, which may include: 1. Armored vehicle engine: Establish speed-frequency response mapping table and measure harmonic distortion coefficient (0.05-0.15).
[0054] 2. Electromagnetic interference: Record the signal-to-noise ratio attenuation curve and pulse repetition frequency (10-100Hz).
[0055] 3. Explosion shock wave: Modeling attenuation coefficient (0.2-0.5dB / ms), reverberation time (1.2-3.5s).
[0056] Ultimately, the output of data preprocessing and parameter library construction can be a standardized speech feature library of 12 types of battlefield noise physical parameters.
[0057] Step S204: Based on standardized speech features, emotion tags, and noise parameters, an identity-preserving generation network is used to determine the synthesized speech and speech quality perception evaluation indicators.
[0058] like Figure 3 As shown, this embodiment can perform identity-preserving speech synthesis. Identity anchoring can be achieved using an IPM-cGAN network. The input to identity-preserving speech synthesis can be: preprocessed speech features + emotion tags + noise parameters.
[0059] In some embodiments, the identity preservation generation network described above is used for identity feature anchoring, condition generation, and multi-objective optimization.
[0060] In some embodiments, the weights of the parameters are fixed during the identity feature anchoring process; during the identity feature anchoring process, the generator adopts a U-Net structure and directly transmits the underlying identity information extracted by the encoder to the decoder through skip connections; the conditional generation process includes: sentiment condition injection and noise physical mapping; the total loss function of multi-objective optimization is determined based on identity preservation loss, adversarial loss and reconstruction loss.
[0061] 1. Identity Feature Anchoring: This embodiment can extract a 512-dimensional identity embedding vector, denoted as F, from the original speech using ECAPA-TDNN (a deep learning speaker recognition model) with its parameters frozen (i.e., its weights are fixed and it does not participate in the generator's training and updates). ref (Reference features) serve as the benchmark for identity verification.
[0062] The generator in this embodiment can adopt a U-Net structure (an encoder-decoder architecture with cross-layer "skip connections" that can retain detailed information of the input signal). Through skip connections, the underlying identity information extracted by the encoder is directly passed to the decoder, keeping the identity information flow independent.
[0063] 2. Condition generation process: This embodiment can perform sentiment condition injection: sentiment tags are integrated into the generation process through a conditional batch normalization layer (a technique that introduces external condition information in network normalization operations).
[0064] This embodiment can perform noise physical mapping: using a fully connected network to convert physical parameters into spectral modulation coefficients, and performing precise noise superposition in the feature map channel dimension.
[0065] 3. Multi-objective optimization: (1) Loss of identity retention (L) FA ):L FA =1-cos(F ref ,F syn (i.e., cosine similarity, used to measure the similarity of two vector directions; the closer the value is to 1, the more similar they are) ≤0.07, forced synthesis of speech feature F syn To F ref Alignment.
[0066] (2) Countermeasures against loss (L) adv The loss from the game between the generator and the discriminator causes the generated data distribution to approximate the real data distribution.
[0067] (3) Reconstruction loss (L) rec ): Calculate the L1 distance between the generated speech and the target speech to ensure content consistency.
[0068] (4) Total loss (L)total ):L total =L adv +0.8L FA +0.5L rec .
[0069] Ultimately, the output of identity-preserving speech synthesis can be: high-fidelity synthesized speech (COS (cosine similarity) ≥ 0.93, PESQ (perceptual speech quality assessment index) ≥ 3.8).
[0070] Step S206: Based on the log-Mel spectrum of synthesized speech and real speech, the clean identity features after decoupling are determined through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training.
[0071] like Figure 3 As shown, this embodiment can perform scene-emotion feature decoupling. This decoupling can be achieved through a dual-path orthogonal attention mechanism. The input for scene-emotion feature decoupling can be the log-Mel spectrum (80-dimensional) of the synthesized / real speech.
[0072] In some embodiments, the above-mentioned dual-path orthogonal attention mechanism includes: an identity attention subnetwork and a noise and sentiment suppression subnetwork; the fusion weights of the dynamic gating fusion are generated by the Sigmoid function and adaptively fuse the features of the dual-path orthogonal attention mechanism; the adversarial decoupling training introduces a gradient reversal layer, and the total loss function of the adversarial decoupling training is determined based on identity preservation loss, scene classification loss and sentiment classification loss.
[0073] 1. Bidirectional orthogonal attention mechanism: (1) IAS Subnet (Identity Attention Subnet): Adopts channel attention mechanism (such as SE module, which learns the weight of each channel through global pooling and fully connected layer) to strengthen the stable formant mode that is key to identity discrimination.
[0074] (2) NESS subnetwork (noise emotion suppression subnetwork): adopts spatial attention mechanism (learning weight graph in time-frequency dimension) to locate and suppress time-frequency regions related to emotion fluctuations and noise.
[0075] 2. Dynamic gating fusion: In this embodiment, the fusion weight g can be generated by the Sigmoid function (σ) to adaptively fuse the dual-path features.
[0076] 3. Adversarial decoupling training: This can be achieved through gradient inversion three-party game.
[0077] This embodiment can introduce a gradient inversion layer (GRL, which transmits data normally during forward propagation and multiplies the gradient by a negative constant during backward propagation to maximize the classifier loss) to construct an adversarial game between the voiceprint encoder and the scene / emotion classifier.
[0078] Loss function for adversarial decoupling training (L) total ):L total =L id (Identity recognition loss) -0.7(L) env (Scene classification loss) +L emo (Emotional classification loss) forces identity features and interfering information to be orthogonal (statistically independent) in the feature space.
[0079] Ultimately, the output of scene-emotion feature decoupling can be: the decoupled pure identity features (interference information ≤ 5%).
[0080] Step S208: Train a voiceprint recognition model based on pure identity features and corresponding speaker labels.
[0081] like Figure 3 As shown, this embodiment can be used for voiceprint recognition model training and optimization. The input for voiceprint recognition model training and optimization can be: decoupled identity features + corresponding speaker tags.
[0082] In some embodiments, the architecture of the above-mentioned voiceprint recognition model includes: a voiceprint recognition network based on ECAPA-TDNN; the training strategy of the voiceprint recognition model includes: training using a Softmax loss function with additional angular intervals, and multi-scene hybrid training.
[0083] Among them, ECAPA-TDNN is a deep learning-based voiceprint recognition model. Its core structure integrates Mel frequency cepstral coefficient (MFCC) feature extraction, time delay neural network (TDNN) architecture, and channel attention mechanism, which significantly improves the accuracy and robustness of speaker verification. It can also enhance long-term dependency modeling capabilities through dilated convolution and residual connections.
[0084] The training strategy for the ECAPA-TDNN voiceprint recognition network in this embodiment can be as follows: using AAM-Softmax (a softmax loss function with added angular intervals, which improves discriminative power by increasing inter-class intervals), with a margin of 0.2. Furthermore, multi-scene hybrid training is performed.
[0085] The performance verification of the ECAPA-TDNN voiceprint recognition network in this embodiment is as follows: cross-scene recognition accuracy: ≥95%; equal error rate (EER, the value when the false acceptance rate equals the false rejection rate, the lower the better): ≤2.5%.
[0086] Ultimately, the output of voiceprint recognition model training and optimization can be a well-trained, highly robust voiceprint recognition model.
[0087] Step S210: Update the voiceprint recognition model based on the real-time speaker's speech stream.
[0088] like Figure 3 As shown, this embodiment can perform dynamic voiceprint updates. Specifically, dynamic voiceprint updates can be performed through physiological state perception. The input for dynamic voiceprint updates can be: real-time speaker speech stream.
[0089] In some embodiments, a real-time speaker's voice stream can be input into a voiceprint recognition model, which outputs the speaker's physiological state. The physiological state includes: fundamental frequency jitter rate, formant bandwidth expansion, and glottal closure quotient. Confidence assessment is performed based on the physiological state. The voiceprint recognition model is updated in a closed loop based on the results of the confidence assessment.
[0090] 1. Real-time monitoring of physiological status: Fundamental frequency jitter: reflects minute fluctuations in the fundamental frequency, which increases when fatigued (threshold 1.2%).
[0091] Formant bandwidth expansion: The widening of the formant frequency range makes the vocal cords more blurred when fatigued (threshold 15%).
[0092] Glottal closure quotient: describes the ratio of the closed phase to the open phase of the glottis, reflecting the efficiency of sound production.
[0093] 2. Confidence assessment: This embodiment can calculate the Mahalanobis distance between the current feature and the features in the voiceprint database (a distance metric that takes into account the data covariance structure and is superior to Euclidean distance).
[0094] In this embodiment, an adaptive threshold λ = μ (mean of historical similarity) - 2σ (standard deviation of historical similarity) can be set.
[0095] 3. Closed-loop update strategy: This embodiment allows for rolling updates: high-confidence samples are retained, while low-confidence samples are archived.
[0096] Ultimately, the output of dynamic voiceprint updates can be a self-evolving dynamic voiceprint library (72-hour continuous combat accuracy ≥92%).
[0097] In summary, the voiceprint synthesis method provided by the embodiments of the present invention has the following differences and advantages compared with the traditional voiceprint synthesis methods of the prior art, mainly in the following dimensions: 1. Identity Preservation Dimension: Traditional voiceprint synthesis methods suffer from severe distortion in synthesized speech voiceprints. The IPM-cGAN network of the voiceprint synthesis method provided in this embodiment of the invention can achieve identity anchoring (COS≥0.93), and physical constraints ensure the authenticity of the synthesized data identity, thus guaranteeing safe and effective training.
[0098] 2. Noise control dimension: Traditional voiceprint synthesis methods only use simple Gaussian noise. The voiceprint synthesis method provided in this invention can reproduce battlefield noise driven by physical parameters (PESQ≥3.8), reproduce the acoustic fingerprint of real battlefield based on the physical model, and make the training more realistic.
[0099] 3. Feature decoupling: Traditional voiceprint synthesis methods suffer from mixed feature spaces. The voiceprint synthesis method provided in this invention can orthogonally separate identity from interference through dual-path attention orthogonal decomposition and adversarial training, and extract pure and robust identity features.
[0100] 4. Dynamic evolution: Traditional voiceprint synthesis methods require manual updates every six months. The voiceprint synthesis method provided in this invention can perform physiological state perception and confidence closed-loop updates, real-time perception of physiological changes and security updates, ensuring continuous and accurate recognition.
[0101] Example 3: This embodiment provides a specific example of a voiceprint synthesis method.
[0102] Example 1: Synthesis of emergency commands in a high-noise environment.
[0103] Input: Calm voice of a specific character (5 seconds) + Emergency emotional tag + Armored vehicle noise parameters (2500 rpm).
[0104] Processing: IPM-cGAN, while adding urgent intonation and engine noise, ensures the speaker characteristics (F) of the synthesized speech through an identity anchoring mechanism. syn ) and the original reference feature (F ref The cosine similarity of ) reached 0.95.
[0105] Output: High-fidelity emergency command voice, which can be used for stress training of operators in simulated high-noise environments.
[0106] 2. Example 2: Adaptive update of voiceprint under fatigue state.
[0107] Detection: Real-time analysis of a specific person's voice revealed that the fundamental frequency jitter rate increased to 1.8%, the formant bandwidth expanded by 20%, and the glottal closure quotient decreased to 0.55.
[0108] Decision: Multi-parameter fusion determines that a specific person is in a state of fatigue.
[0109] Action: The system automatically triggers the high-priority update channel, calculates the Mahalanobis distance of the new sample d=0.85, which is lower than the dynamic threshold λ (0.90), so the new voiceprint features under this "fatigue state" are safely updated into the database.
[0110] Result: The voiceprint database adapted to the physiological changes of specific individuals in a timely manner, avoiding misidentification due to vocal cord fatigue.
[0111] The method provided in the embodiments of the present invention mainly includes the following: 1. A method for decoupling voiceprint features: Deploy a dual-path orthogonal attention module (SEDF) at the front end of ECAPA-TDNN, extract identity features through channel attention, and suppress interference features through spatial attention; use gradient inversion adversarial training to make the voiceprint embedding vector unrecognizable by scene / emotion classifiers.
[0112] 2. Identity-preserving generation method: A frozen speaker encoder is embedded in the generator of the IPM-cGAN network, and the cosine similarity loss (L0.05) between the synthesized features and the original speech is calculated. FA The synthesized speech is forced to retain the original speaker's identity; battlefield environmental noise is synthesized by driving the noise injection layer through physical parameters.
[0113] 3. Dynamic voiceprint update method: Based on the fusion of multiple parameters such as fundamental frequency jitter rate and formant expansion coefficient, the speaker fatigue state is detected; the similarity distribution between the current voiceprint features and the historical voiceprint database is calculated, and when the confidence is lower than the adaptive Mahalanobis distance threshold (λ=μ-2σ), the voiceprint features are eliminated.
[0114] The method provided in this embodiment of the invention has the following irreplaceable advantages in military applications: 1. Tactical-level identity protection: The voiceprint of a specific person is anchored and constrained by biometric features throughout the synthesis / recognition process, avoiding the exposure of identity due to technical defects.
[0115] 2. Physical-level reconstruction of battlefield environment: 12 types of noise are accurately reconstructed through a physical parameter library, enabling operators to develop muscle memory and stress response during simulation training.
[0116] 3. Voiceprint combat readiness self-evolution: The "detection-update-training" closed loop enables the voiceprint database to adaptively evolve with the intensity of combat, meeting the needs of continuous combat.
[0117] The method provided in this embodiment of the invention can be applied to a closed-loop voiceprint management system. The processing flow of the closed-loop voiceprint management system can be found in [reference needed]. Figure 4 The diagram shows the overall process of a closed-loop voiceprint management system.
[0118] The method provided in this embodiment of the invention introduces an identity-preserving generative adversarial network (IPM-cGAN), a scene-emotion decoupling feature extraction front-end (SEDF), and a scene-emotion adversarial training objective (L). SEA This system, based on graph databases and dynamic voiceprint evolution management, systematically addresses the pain points in the three core stages of military voiceprint recognition: data generation, feature extraction, and model updating. Its main advantages include: 1. Solving the problem of cross-scene recognition failure: IPM-cGAN provides high-fidelity training data covering extreme environments, and the SEDF front end robustly extracts the decoupled pure identity features. The dual guarantee significantly improves the recognition accuracy in complex battlefield environments such as high noise and emotional fluctuations.
[0119] 2. Eliminate emotional blind spots: The Negative Interference Suppression Subnetwork (NESS) and adversarial training mechanism in SEDF effectively separate the influence of emotional fluctuations on voiceprint features, ensuring that the voices of specific individuals in different emotional states (such as calm, emergency, and anger) can be accurately recognized.
[0120] 3. Enhance the reliability of tactical decision-making: Dynamic voiceprint modeling can perceive the physiological state of specific individuals (such as fatigue) in real time and securely update the voiceprint database. Combined with highly environmentally robust recognition capabilities, it significantly reduces the risk of misjudgment or missed hearing of instructions caused by voiceprint changes or environmental interference, ensuring the accuracy and timeliness of combat instruction transmission and providing reliable communication support for tactical decision-making.
[0121] Example 4: Corresponding to the above method embodiments, this invention provides a voiceprint synthesis device, see [link to previous document]. Figure 5 The diagram shows a structural schematic of a voiceprint synthesis device, which includes: The data preprocessing and parameter library construction module 51 is used to determine standardized speech features based on the original speaker's speech segments and to construct a physical parameter library for battlefield noise. The identity-preserving speech synthesis module 52 is used to determine the synthesized speech and speech quality perception evaluation index based on standardized speech features, emotion tags and noise parameters through an identity-preserving generation network. The scene emotion feature decoupling module 53 is used to determine the pure identity features after decoupling based on the log-Mel spectrum of synthesized speech and real speech through a dual-path orthogonal attention mechanism, dynamic gating fusion and adversarial decoupling training. The voiceprint recognition model training and optimization module 54 is used to train the voiceprint recognition model based on pure identity features and corresponding speaker tags. The dynamic voiceprint update mechanism module 55 is used to update the voiceprint recognition model based on the real-time speaker's speech stream.
[0122] This invention provides a voiceprint synthesis device. It determines standardized speech features based on original speaker speech fragments and constructs a battlefield noise physical parameter library. Based on the standardized speech features, emotion tags, and noise parameters, it determines synthesized speech and speech quality perception evaluation indicators through an identity-preserving generation network. Based on the log-Mel spectrum of synthesized and real speech, it determines the decoupled pure identity features through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training. It trains a voiceprint recognition model based on the pure identity features and corresponding speaker tags. The voiceprint recognition model is updated based on real-time speaker speech streams. This approach proposes an identity-preserving generation network and a noise mapping mechanism driven by physical parameters, which can maintain specific person identity features in synthesized speech and reproduce the acoustic characteristics of battlefield noise with high fidelity. This provides operators with more realistic and effective training speech samples, enhancing their recognition ability and stress response in real combat environments.
[0123] The aforementioned data preprocessing and parameter library construction module is used to sequentially perform speech preprocessing and basic feature extraction on the original speaker's speech segments to obtain standardized speech features. Among them, speech preprocessing includes pre-emphasis, framing, and windowing, and standardized speech features include Mel frequency cepstral coefficients, fundamental frequency, spectral envelope, and formant trajectories. Based on the standardized speech features, a battlefield noise physical parameter library is constructed in a physical parameter-driven manner.
[0124] The aforementioned identity-preserving generative network is used for identity feature anchoring, condition generation, and multi-objective optimization.
[0125] In the above identity feature anchoring process, the weights of the parameters are fixed; in the identity feature anchoring process, the generator adopts a U-Net structure, and the low-level identity information extracted by the encoder is directly passed to the decoder through skip connections; the conditional generation process includes: sentiment condition injection and noise physical mapping; the total loss function of multi-objective optimization is determined based on identity preservation loss, adversarial loss and reconstruction loss.
[0126] The aforementioned dual-path orthogonal attention mechanism includes: an identity attention subnetwork and a noise and sentiment suppression subnetwork; the fusion weights of the dynamic gating fusion are generated by the Sigmoid function and adaptively fuse the features of the dual-path orthogonal attention mechanism; the adversarial decoupling training introduces a gradient reversal layer, and the total loss function of the adversarial decoupling training is determined based on identity preservation loss, scene classification loss and sentiment classification loss.
[0127] The architecture of the aforementioned voiceprint recognition model includes: a voiceprint recognition network based on ECAPA-TDNN; the training strategies for the voiceprint recognition model include: training using a Softmax loss function with additional angular intervals, and multi-scene hybrid training.
[0128] The aforementioned scene emotion feature decoupling module is used to input the real-time speaker's voice stream into the voiceprint recognition model and output the real-time speaker's physiological state. The physiological state includes: fundamental frequency jitter rate, formant bandwidth expansion, and glottal closure quotient. Confidence assessment is performed based on the physiological state. The voiceprint recognition model is updated in a closed loop based on the results of the confidence assessment.
[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the voiceprint synthesis device described above can be referred to the corresponding process in the aforementioned embodiments of the voiceprint synthesis method, and will not be repeated here.
[0130] Example 5: This invention also provides an electronic device for running the above-described voiceprint synthesis method; see [link to previous document]. Figure 6 The diagram shows the structure of an electronic device, which includes a memory 100 and a processor 101. The memory 100 is used to store one or more computer instructions, which are executed by the processor 101 to implement the above-mentioned voiceprint synthesis method.
[0131] Furthermore, Figure 6 The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 101, the communication interface 103 and the memory 100 connected via the bus 102.
[0132] The memory 100 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0133] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. Processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0134] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the aforementioned voiceprint synthesis method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0135] The computer program products of the voiceprint synthesis method, apparatus, electronic device and storage medium provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0136] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0137] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0138] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0139] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0140] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A voiceprint synthesis method, characterized in that, The method includes: Standardized speech features were determined based on original speaker speech fragments, and a database of battlefield noise physical parameters was constructed. Based on the standardized speech features, emotion tags, and noise parameters, synthesized speech and speech quality perception evaluation indicators are determined through an identity-preserving generation network. Based on the log-Mel spectrum of the synthesized speech and real speech, the decoupled pure identity features are determined through a dual-path orthogonal attention mechanism, dynamic gating fusion, and adversarial decoupling training. A voiceprint recognition model is trained based on the pure identity features and the corresponding speaker tags; The voiceprint recognition model is updated based on the real-time speaker's speech stream.
2. The method according to claim 1, characterized in that, The steps for determining standardized speech features and constructing a database of battlefield noise physical parameters based on original speaker speech fragments include: The original speaker's speech segments are sequentially subjected to speech preprocessing and basic feature extraction to obtain standardized speech features; wherein, the speech preprocessing includes: pre-emphasis, framing and windowing, and the standardized speech features include: Mel frequency cepstral coefficients, fundamental frequency, spectral envelope and formant trajectory; Based on the standardized speech features, a battlefield noise physical parameter library is constructed using a physical parameter-driven approach.
3. The method according to claim 1, characterized in that, The identity preservation generation network is used for identity feature anchoring, condition generation, and multi-objective optimization.
4. The method according to claim 3, characterized in that, During the identity feature anchoring process, the weights of the parameters are fixed; during the identity feature anchoring process, the generator adopts a U-Net structure and directly transmits the underlying identity information extracted by the encoder to the decoder through skip connections. The process of generating the conditions includes: emotional condition injection and noise physical mapping; The total loss function for the multi-objective optimization is determined based on identity preservation loss, adversarial loss, and reconstruction loss.
5. The method according to claim 1, characterized in that, The dual-path orthogonal attention mechanism includes: an identity attention subnetwork and a noise emotion suppression subnetwork; The fusion weights of the dynamic gating fusion are generated by the Sigmoid function and adaptively fuse the features of the dual-path orthogonal attention mechanism. The adversarial decoupling training introduces a gradient inversion layer, and the total loss function of the adversarial decoupling training is determined based on identity preservation loss, scene classification loss, and sentiment classification loss.
6. The method according to claim 1, characterized in that, The architecture of the voiceprint recognition model includes: a voiceprint recognition network based on ECAPA-TDNN; The training strategy for the voiceprint recognition model includes: training using a Softmax loss function with additional angular intervals, and multi-scene hybrid training.
7. The method according to claim 1, characterized in that, The steps of updating the voiceprint recognition model based on the real-time speaker's speech stream include: The real-time speaker's voice stream is input into the voiceprint recognition model, and the physiological state of the real-time speaker is output; wherein, the physiological state includes: fundamental frequency jitter rate, formant bandwidth expansion, and glottal closure quotient; Confidence assessment is performed based on the aforementioned physiological state; The voiceprint recognition model is updated in a closed loop based on the results of the confidence assessment.
8. A voiceprint synthesis device, characterized in that, The device includes: The data preprocessing and parameter library construction module is used to determine standardized speech features based on the original speaker's speech fragments and to construct a physical parameter library for battlefield noise. An identity-preserving speech synthesis module is used to determine synthesized speech and speech quality perception evaluation indicators through an identity-preserving generation network based on the standardized speech features, emotion tags and noise parameters. The scene emotion feature decoupling module is used to determine the decoupled pure identity features based on the log-Mel spectrum of the synthesized speech and real speech, through a dual-path orthogonal attention mechanism, dynamic gating fusion and adversarial decoupling training. The voiceprint recognition model training and optimization module is used to train the voiceprint recognition model based on the pure identity features and the corresponding speaker tags. The dynamic voiceprint update mechanism module is used to update the voiceprint recognition model based on the real-time speaker's speech stream.
9. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the voiceprint synthesis method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the voiceprint synthesis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-factor controllable speech conversion method and system based on feature decoupling
CN113327627A
Generative adversarial network optimization method and system for short voice speaker confirmation
CN114530156A
Generative adversarial network-based emotion asymmetric speaker recognition system
CN116543774A
Personalized speech synthesis method and device with noise robustness
CN118173079A
Voice anonymization method based on content-related frame-level speaker voiceprint modeling
CN119851670A