Voiceprint recognition method, device and equipment based on pulse sequence enhanced speech

By preprocessing and multi-pulse sequence enhancement of the original voice data, combined with neural network model training, the problems of insufficient generalization and stability of voiceprint recognition technology in cross-channel processing are solved, and higher recognition accuracy and feature differentiation are achieved.

CN114842853BActive Publication Date: 2025-09-23XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210320699.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2025-09-23
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

Existing voiceprint recognition technology has problems with poor generalization and stability in cross-channel processing, especially due to the diversity of natural and artificial noise, which leads to insufficient model generalization and stability.

Method used

By preprocessing the original speech data, calculating the average energy and generating a multi-pulse sequence, the speech is enhanced, combined with neural network model training, and the voiceprint recognition model is optimized to improve recognition accuracy.

Benefits of technology

The accuracy of voiceprint recognition and the generalization of the model are enhanced, the impact of noise on recognition is reduced, and the discrimination of voiceprint features and the stability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842853B_ABST
    Figure CN114842853B_ABST
Patent Text Reader

Abstract

The present invention discloses a voiceprint recognition method, apparatus, device, and storage medium based on pulse sequence-enhanced speech. The method comprises: preprocessing acquired original speech data, calculating the average energy of the original speech data, and calculating a multi-pulse sequence based on the average energy; performing speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech; extracting features from the first speech and inputting them into a preset neural network model for training to obtain a voiceprint recognition model; using the enhanced speech to be recognized as the second speech, extracting feature parameters of the second speech, and identifying the feature parameters through the voiceprint recognition model to output corresponding voiceprint features. This method can improve the accuracy of voiceprint recognition and the generalization of the voiceprint model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a voiceprint recognition method, device and equipment based on pulse sequence enhanced speech. Background Art

[0002] As technology improves and matures, voiceprint recognition technology is gradually being applied to fields like finance. However, voiceprint recognition still faces technical barriers, such as cross-channel processing. Existing solutions address these barriers by primarily enhancing the original speech with noise, performing cross-channel processing, and improving the network's expressive capabilities. However, existing original speech enhancement solutions have limitations: natural and man-made noise are diverse and complex, and few noisy datasets effectively cover them. This limits the generalization and stability of the models. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to propose a voiceprint recognition method, device and equipment based on pulse sequence enhanced speech, aiming to solve the problems of poor generalization and stability of existing voiceprint models.

[0004] To achieve the above object, the present invention provides a voiceprint recognition method based on pulse sequence enhanced speech, the method comprising:

[0005] After preprocessing the acquired original speech data, average energy of the original speech data is calculated, and a multi-pulse sequence is calculated based on the average energy;

[0006] Performing speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech;

[0007] Extracting features from the first speech and then inputting them into a preset neural network model for training to obtain a voiceprint recognition model;

[0008] The enhanced speech to be recognized is used as the second speech, characteristic parameters of the second speech are extracted, the characteristic parameters are recognized by the voiceprint recognition model, and corresponding voiceprint features are output.

[0009] Preferably, after preprocessing the acquired original speech data, calculating the average energy of the original speech data includes:

[0010] Extract the speech of the effective segment in the original speech data as the effective speech, and calculate the average energy of the effective speech.

[0011] Preferably, the calculating the average energy of the effective speech includes:

[0012] according to The weighted average sum of the effective speech lengths of the effective speech is calculated to obtain the average energy, where n is the effective speech length.

[0013] Preferably, performing voice enhancement on the original voice data includes:

[0014] The original speech data is subjected to speech enhancement of a multi-pulse sequence according to y(n)=x(n)+v(n), wherein x(n) represents the original speech and v(n) represents the multi-pulse sequence.

[0015] Preferably, the feature extraction of the first speech and inputting the feature into a preset neural network model for training to obtain a voiceprint recognition model include:

[0016] When the voiceprint recognition model satisfies the first condition during training, calculating the current EER of the voiceprint recognition model;

[0017] After adjusting the multi-pulse sequence of the original speech data based on the current EER, the voiceprint recognition model is continuously trained until a second condition is met, and the multi-pulse sequence information is recorded.

[0018] Preferably, the first condition includes that the loss function of the voiceprint recognition model tends to be stable or the loss function of the voiceprint recognition model is lower than a threshold; the second condition includes that the loss function of the voiceprint recognition model and the EER both tend to be stable.

[0019] To achieve the above-mentioned object, the present invention further provides a voiceprint recognition device based on pulse sequence enhanced speech, the device comprising:

[0020] a calculation unit, configured to calculate average energy of the acquired original speech data after preprocessing the data, and calculate a multi-pulse sequence based on the average energy;

[0021] a speech enhancement unit, configured to perform speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech;

[0022] A model training unit, configured to extract features from the first speech and then input the features into a preset neural network model for training to obtain a voiceprint recognition model;

[0023] The recognition unit is used to use the enhanced speech to be recognized as the second speech, extract feature parameters of the second speech, recognize the feature parameters through the voiceprint recognition model, and output corresponding voiceprint features.

[0024] In order to achieve the above objectives, the present invention also proposes a device, including a processor, a memory, and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the steps of a voiceprint recognition method based on pulse sequence enhanced speech as described in the above embodiment.

[0025] In order to achieve the above objectives, the present invention also proposes a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the steps of a voiceprint recognition method based on pulse sequence enhanced speech as described in the above embodiment.

[0026] Beneficial effects:

[0027] The above scheme processes speech based on multi-pulse sequences and enhances voiceprint information of certain frequencies, thereby increasing the difference between different people's voiceprints, improving the accuracy of voiceprint recognition and the generalization of the voiceprint model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0029] Figure 1 A flow chart of a voiceprint recognition method based on pulse sequence enhanced speech provided by one embodiment of the present invention.

[0030] Figure 2 A schematic diagram of the structure of a voiceprint recognition device based on pulse sequence enhanced speech provided by one embodiment of the present invention.

[0031] The realization of the objectives of the invention, the functional features and advantages will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention for which protection is sought, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0033] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of the features.

[0034] The present invention is described in detail below with reference to the embodiments.

[0035] The existing voiceprint recognition still has the following problems, including: different encoding technologies and transmission losses will have some impact on the original signal, how to compensate for or reduce these losses, the existing cross-channel processing technology cannot fully solve the problem; with the influence of different factors, the original voice may change or be distorted.

[0036] Based on this, the present invention proposes a voiceprint recognition method based on pulse sequence enhanced speech. This method minimizes the acquisition costs of natural and artificial noise, as well as the cost of cross-channel impact compensation technology, while maximizing the ability to distinguish different voiceprint features. Furthermore, it improves the generalization, stability, and accuracy of the voiceprint model to cope with increasingly variable environmental noise and technical processing differences.

[0037] Reference Figure 1 FIG2 is a flow chart of a voiceprint recognition method based on pulse sequence enhanced speech provided by an embodiment of the present invention.

[0038] In this embodiment, the method includes:

[0039] S11 , after pre-processing the acquired original speech data, calculating the average energy of the original speech data, and calculating a multi-pulse sequence according to the average energy.

[0040] The step of calculating the average energy of the original speech data after preprocessing the obtained original speech data includes:

[0041] The speech of the effective segment in the original speech data is extracted as the effective speech, and the average energy of the effective speech is calculated.

[0042] Furthermore, the calculating the average energy of the effective speech includes:

[0043] according to Calculate the weighted average sum of the effective speech lengths of the effective speech to obtain the average energy, where n is the effective speech length, x(m) is the speech data at speech point m, and w(nm) is the window function value corresponding to the speech data at that point.

[0044] In this embodiment, the input original voice data is subjected to noise reduction processing, and the effective segment of the original voice data is obtained by using webrtc, and the voice is processed according to the formula Calculate the average energy of the effective speech, and the average energy is the weighted average sum of the effective speech length. Calculate the multi-pulse sequence of the effective speech (i.e., the energy value of the multi-pulse sequence) based on the obtained average energy, so as to determine the energy size of each added pulse, so as to perform pulse sequence speech enhancement on the effective speech, while not enhancing the non-effective speech (such as silent segments, noise segments, etc.). Among them, the three elements of the pulse are amplitude, width, and phase. The amplitude value is determined according to the average energy of the speech to be enhanced, the width can be set manually, and the phase is automatically adjusted to the optimal value by the program. In S11-1, some abnormal energy values ​​(such as strong noise) are first filtered out, and then the average energy value E of the effective speech is calculated. n , and the amplitude value is set to to The average energy is used to prevent the pulse characteristics (pulse characteristics refer to the abstract characteristics such as the phase and energy of a single pulse) from having too much influence on the distinguishable voiceprint features.

[0045] S12: Perform speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech.

[0046] The performing voice enhancement on the original voice data includes:

[0047] The original speech data is subjected to speech enhancement of a multi-pulse sequence according to y(n)=x(n)+v(n), wherein x(n) represents the original speech and v(n) represents the multi-pulse sequence.

[0048] Based on the above, the generation conditions of speech in v(n) can be considered as phase and energy. The neural network model adjusts the phase, E n Adjust the energy. The purpose of En is to prevent the pulse feature from having too much influence on the distinguishable voiceprint features.

[0049] S13, extracting features from the first speech and inputting the features into a preset neural network model for training to obtain a voiceprint recognition model.

[0050] The step of extracting features from the first speech and inputting the features into a preset neural network model for training to obtain a voiceprint recognition model includes:

[0051] When the voiceprint recognition model satisfies the first condition during training, calculating the current EER of the voiceprint recognition model;

[0052] After adjusting the multi-pulse sequence of the original speech data based on the current EER, the voiceprint recognition model is continuously trained until a second condition is met, and the multi-pulse sequence information is recorded.

[0053] Furthermore, the first condition includes that the loss function of the voiceprint recognition model tends to be stable or the loss function of the voiceprint recognition model is lower than a threshold; the second condition includes that the loss function of the voiceprint recognition model and the EER both tend to be stable.

[0054] In this embodiment, features (such as MFCC features, PLP features, and FBANK features) are extracted from the first speech after speech enhancement, or wav2vec processing is performed. The extracted features are then fed into a pre-set neural network model, such as the ECAPA-TDNN, RESNET, or VGG series, for training. If the model's loss function stabilizes or its value falls below a threshold, the EER of the current model is calculated. Based on the EER, the multi-pulse sequence of the original speech data is adjusted, and model training continues. When both the loss function and the EER stabilize, the multi-pulse sequence information (including phase and energy information) and deep network information are recorded. The multi-pulse sequence information enhances the speech to be recognized. Feature parameters of the enhanced speech to be recognized are extracted and passed through the voiceprint recognition model (the model parameters are the deep network information stored after training) to calculate the voiceprint features.

[0055] Furthermore, the methods for adjusting the multi-pulse sequence of the original speech data are: (1) processing according to the Mel scale, that is, setting a pulse in each Mel frequency band, and the specific pulse frequency is optimized and calculated by the program; (2) processing according to the audio sampling rate and Nyquist theorem, that is, setting a pulse in the range of 1 to The frequency bands are fixedly allocated between the two, and there is one pulse within each frequency band. The specific pulse frequency is calculated by the program. The judgment of the stability of the loss function and the stability of EER include: (1) the loss value no longer changes, that is, the change of the loss value is less than 0.01 in more than 3 epochs; (2) the EER value no longer changes, that is, the change of EER is less than 0.01 in more than 3 epochs; (3) the loss value changes greatly, but is less than the specified threshold. In these three cases, the current stage of training can be terminated and the training after pulse adjustment can be continued or the training can be terminated (mainly because the EER has stabilized).

[0056] S14, taking the enhanced speech to be recognized as the second speech, extracting feature parameters of the second speech, recognizing the feature parameters through the voiceprint recognition model, and outputting corresponding voiceprint features.

[0057] In summary, the above solution effectively simplifies the voice enhancement process for voiceprint recognition. It automatically adjusts the energy of multi-pulse sequences based on the average energy of valid speech segments. Furthermore, it reduces the acquisition costs of natural and artificial noise, as well as the cost of cross-channel noise compensation technology, thereby improving the distinguishability of different voiceprint features.

[0058] Reference Figure 2 FIG2 is a schematic diagram showing the structure of a voiceprint recognition device based on pulse sequence enhanced speech provided by an embodiment of the present invention.

[0059] In this embodiment, the device 20 includes:

[0060] a calculation unit 21 configured to pre-process the acquired original speech data, calculate the average energy of the original speech data, and calculate a multi-pulse sequence based on the average energy;

[0061] a speech enhancement unit 22, configured to perform speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech;

[0062] A model training unit 23 is configured to extract features from the first speech and then input the features into a preset neural network model for training to obtain a voiceprint recognition model;

[0063] The recognition unit 24 is configured to take the enhanced speech to be recognized as the second speech, extract feature parameters of the second speech, recognize the feature parameters through the voiceprint recognition model, and output corresponding voiceprint features.

[0064] Furthermore, the calculation unit 21 is configured to:

[0065] The speech of the effective segment in the original speech data is extracted as the effective speech, and the average energy of the effective speech is calculated.

[0066] Among them, according to The weighted average sum of the effective speech lengths of the effective speech is calculated to obtain the average energy, where n is the effective speech length.

[0067] Furthermore, the speech enhancement unit 22 is configured to:

[0068] The original speech data is subjected to speech enhancement of a multi-pulse sequence according to y(n)=x(n)+v(n), wherein x(n) represents the original speech and v(n) represents the multi-pulse sequence.

[0069] Furthermore, the model training unit 23 includes:

[0070] A first training unit, configured to calculate a current EER of the voiceprint recognition model when the voiceprint recognition model satisfies a first condition during training;

[0071] The second training unit is configured to adjust the multi-pulse sequence of the original speech data based on the current EER and then continue to train the voiceprint recognition model until a second condition is met, and record the multi-pulse sequence information.

[0072] Among them, the first condition includes that the loss function of the voiceprint recognition model tends to be stable or the loss function of the voiceprint recognition model is lower than a threshold; the second condition includes that the loss function of the voiceprint recognition model and the EER both tend to be stable.

[0073] Each unit module of the device 20 can respectively execute the corresponding steps in the above method embodiment, so each unit module will not be described in detail here. Please refer to the description of the corresponding steps above for details.

[0074] The embodiment of the present invention further provides a device, which includes the above-mentioned voiceprint recognition device based on pulse sequence enhanced speech, wherein the voiceprint recognition device based on pulse sequence enhanced speech can adopt Figure 2 The structure of the embodiment can be executed accordingly. Figure 1 The technical solution of the method embodiment shown has similar implementation principles and technical effects. For details, please refer to the relevant records in the above embodiments and will not be repeated here.

[0075] The device includes: a mobile phone, digital camera, tablet computer, or other device with a camera function, or a device with an image processing function, or a device with an image display function. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.

[0076] Among them, the memory can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as an image playback function, etc.), etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor and the input unit with access to the memory.

[0077] The input unit can be used to receive input digital, character, or image information, and generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control. Specifically, the input unit of this embodiment includes not only a camera, but also a touch-sensitive surface (such as a touch display) and other input devices.

[0078] The display unit can be used to display information input by the user or information provided to the user and various graphical user interfaces of the device, which can be composed of graphics, text, icons, videos and any combination thereof. The display unit may include a display panel. Optionally, the display panel can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Furthermore, the touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event. The processor then provides a corresponding visual output on the display panel based on the type of touch event.

[0079] The embodiment of the present invention further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the memory in the above embodiment; or a computer-readable storage medium that exists independently and is not assembled into a device. The computer-readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement Figure 1 The voiceprint recognition method based on pulse sequence enhanced speech is shown. The computer readable storage medium can be a read-only memory, a disk or an optical disk, etc.

[0080] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For similar or identical parts between the various embodiments, reference can be made to each other. For the apparatus embodiments, device embodiments, and storage medium embodiments, since they are generally similar to the method embodiments, their descriptions are relatively simple. For relevant parts, reference can be made to the descriptions of the method embodiments.

[0081] Furthermore, in this document, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0082] While the foregoing description shows and describes preferred embodiments of the present invention, it should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments, and can be modified within the scope of the present invention by the teachings herein or by techniques or knowledge in the relevant art. Modifications and variations made by those skilled in the art without departing from the spirit and scope of the present invention are intended to be within the scope of the appended claims.

Claims

1. A voiceprint recognition method based on pulse sequence enhanced speech, characterized in that: The method comprises: After preprocessing the acquired original voice data, calculating the average energy of the original voice data, and calculating a multi-pulse sequence based on the average energy; comprising: extracting a valid segment of voice from the original voice data as valid voice, and calculating the average energy of the valid voice; wherein the valid segment of voice of the original voice data is obtained using WebRTC; Furthermore, the calculating the average energy of the effective speech includes: according to Calculating a weighted average sum of the effective speech lengths of the effective speech to obtain the average energy, wherein n is the effective speech length, x(m) is the speech data at point m of the speech, w(nm) is the window function value corresponding to the speech data at the point m, and m represents the speech point m; wherein, calculating a multi-pulse sequence of the effective speech based on the obtained average energy, thereby determining the energy of each added pulse, so as to perform pulse sequence speech enhancement on the effective speech, while not performing enhancement on the ineffective speech; Performing speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech; Furthermore, performing voice enhancement on the original voice data includes: Performing multi-pulse sequence speech enhancement on the original speech data according to y(n)=x(n)+v(n), wherein x(n) represents the original speech and v(n) represents the multi-pulse sequence; Extracting features from the first speech and then inputting them into a preset neural network model for training to obtain a voiceprint recognition model; The enhanced speech to be recognized is used as the second speech, characteristic parameters of the second speech are extracted, the characteristic parameters are recognized by the voiceprint recognition model, and corresponding voiceprint features are output.

2. The voiceprint recognition method based on pulse sequence enhanced speech according to claim 1, characterized in that: The feature extraction of the first speech is input into a preset neural network model for training to obtain a voiceprint recognition model, including: When the voiceprint recognition model satisfies the first condition during training, calculating the current EER of the voiceprint recognition model; After adjusting the multi-pulse sequence of the original speech data based on the current EER, the voiceprint recognition model is continuously trained until a second condition is met, and the multi-pulse sequence information is recorded.

3. The voiceprint recognition method based on pulse sequence enhanced speech according to claim 2, characterized in that: The first condition includes that the loss function of the voiceprint recognition model tends to be stable or the loss function of the voiceprint recognition model is lower than a threshold; the second condition includes that the loss function of the voiceprint recognition model and the EER both tend to be stable.

4. A voiceprint recognition device based on pulse sequence enhanced speech, characterized in that: The device comprises: A calculation unit is configured to calculate the average energy of the acquired original speech data after preprocessing the data, and calculate the multi-pulse sequence based on the average energy; further, the calculation unit is configured to: Extracting the effective segment of speech from the original speech data as effective speech, and calculating the average energy of the effective speech; wherein the effective segment of speech of the original speech data is obtained by using webrtc; Among them, according to Calculating a weighted average sum of the effective speech lengths of the effective speech to obtain the average energy, wherein n is the effective speech length, x(m) is the speech data at point m of the speech, w(nm) is the window function value corresponding to the speech data at the point m, and m represents the speech point m; wherein, calculating a multi-pulse sequence of the effective speech based on the obtained average energy, thereby determining the energy of each added pulse, so as to perform pulse sequence speech enhancement on the effective speech, while not performing enhancement on the ineffective speech; A speech enhancement unit is configured to perform speech enhancement on the original speech data based on the multi-pulse sequence to obtain a first speech; further, the speech enhancement unit is configured to: Performing multi-pulse sequence speech enhancement on the original speech data according to y(n)=x(n)+v(n), wherein x(n) represents the original speech and v(n) represents the multi-pulse sequence; A model training unit, configured to extract features from the first speech and then input the features into a preset neural network model for training to obtain a voiceprint recognition model; The recognition unit is used to use the enhanced speech to be recognized as the second speech, extract feature parameters of the second speech, recognize the feature parameters through the voiceprint recognition model, and output corresponding voiceprint features.

5. A device, characterized in that The method comprises a processor, a memory and a computer program stored in the memory, wherein the computer program is executed by the processor to implement the steps of a voiceprint recognition method based on pulse sequence enhanced speech as claimed in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is executed by a processor to implement the steps of a voiceprint recognition method based on pulse sequence enhanced speech as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Speech enhancement method based on attention residual learning

    CN112992121A

  • Voiceprint recognition method based on attention mechanism recurrent neural network

    CN113129897A

  • Time sequence voiceprint feature combination recognition method and device

    CN114203185A

  • Speech signal processor

    JP1994202695A