Offline personalized voiceprint learning method and speaker separation method

By performing offline personalized voiceprint learning on the device and optimizing the voiceprint recognition model using meta-learning and user feedback, the problems of low update efficiency and security risks in existing technologies are solved, achieving efficient and secure model updates and improved recognition accuracy.

CN119785801BActive Publication Date: 2025-11-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411754107.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-11-18
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Existing voiceprint recognition models cannot effectively resist new attacks or suffer from decreased recognition accuracy on electronic devices with high confidentiality requirements. Furthermore, traditional online update mechanisms are inefficient, costly, and pose security risks.

Method used

This paper provides an offline personalized voiceprint learning method. By performing meta-learning on the device using a built-in general voiceprint recognition model and personalized voiceprint learning data, a target personalized voiceprint recognition model is determined. Personalized voiceprint learning data is then constructed by combining user feedback information to achieve offline personalized training.

Benefits of technology

It improves model update efficiency, reduces costs, avoids security issues introduced by transmission, and ensures efficient and secure model updates and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785801B_ABST
    Figure CN119785801B_ABST
Patent Text Reader

Abstract

The application provides an offline personalized voiceprint learning method and a speaker separation method, relates to the technical field of voice processing, determines a built-in general voiceprint recognition model and target general voiceprint learning data on a device end, and acquires personalized voiceprint learning data; meta learning is performed on the general voiceprint recognition model by using training data and the personalized voiceprint learning data, and an initial personalized voiceprint recognition model is obtained; finally, test data is used to test the general voiceprint recognition model and the initial personalized voiceprint recognition model respectively, and a target personalized voiceprint recognition model is determined based on a first test result obtained. The method can realize offline personalized training by using the personalized voiceprint learning data and the training data built in the device end to perform meta learning and test on the general voiceprint recognition model, does not need to transmit a model update package to each device end, can greatly improve the model update efficiency, reduce the cost, and avoid the security problem caused by the transmission of the update package.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to an offline personalized voiceprint learning method and a speaker separation method. Background Technology

[0002] Voiceprint recognition technology, as an advanced biometric technology for identity verification based on voice features, relies heavily on the accuracy and timeliness of the voiceprint recognition model for its accuracy and robustness. With continuous technological advancements and evolving attack methods, voiceprint recognition models deployed on electronic devices with high security requirements may gradually reveal weaknesses in effectively resisting new attacks or experiencing decreased recognition accuracy.

[0003] Since these electronic devices cannot directly access the network, traditional online update mechanisms for voiceprint recognition models become impractical. This means that whenever a voiceprint recognition model needs to be updated to adapt to new voice features, optimize the recognition algorithm, or fix known vulnerabilities, the update package must be transmitted to each target device and installed locally. This process is not only inefficient and costly, but may also introduce new risks due to human error or security issues with the update package during transmission.

[0004] Furthermore, a prolonged lack of updates can cause voiceprint recognition models to gradually fall behind technological advancements, failing to fully utilize the latest algorithm optimizations and performance improvements, thereby impacting the overall system performance and user experience. Therefore, how to achieve efficient and secure updates to voiceprint recognition models and algorithms while ensuring confidentiality has become a pressing technical challenge. Summary of the Invention

[0005] This invention provides an offline personalized voiceprint learning method and a speaker separation method to address the deficiencies in related technologies.

[0006] This invention provides an offline personalized voiceprint learning method, applied to the device side, including:

[0007] The built-in general voiceprint recognition model and target general voiceprint learning data are determined, and personalized voiceprint learning data is obtained; the target general voiceprint learning data includes training data and test data.

[0008] Based on the training data and the personalized voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain an initial personalized voiceprint recognition model.

[0009] Based on the test data, the general voiceprint recognition model and the initial personalized voiceprint recognition model are tested respectively to obtain a first test result, and based on the first test result, the target personalized voiceprint recognition model is determined.

[0010] According to an offline personalized voiceprint learning method provided by the present invention, the first test result includes a second test result corresponding to the general voiceprint recognition model and a third test result corresponding to the initial personalized voiceprint recognition model;

[0011] The step of determining the target personalized voiceprint recognition model based on the first test result includes:

[0012] If the second test result is better than the third test result, then the general voiceprint recognition model is determined to be the target personalized voiceprint recognition model;

[0013] If the third test result is better than the second test result, then the initial personalized voiceprint recognition model is determined to be the target personalized voiceprint recognition model.

[0014] According to the present invention, an offline personalized voiceprint learning method is provided, wherein obtaining personalized voiceprint learning data includes:

[0015] Obtain sample audio;

[0016] Based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the sample speech to determine the first voiceprint recognition result, and based on the first voiceprint recognition result, speaker separation is performed on the sample speech to obtain the speaker separation result.

[0017] The system receives feedback from users regarding the speaker separation results and constructs the personalized voiceprint learning data based on the feedback and the sample speech.

[0018] According to an offline personalized voiceprint learning method provided by the present invention, the feedback information includes speakers of interest in the sample speech and the user's modification data when an error occurs in the first voiceprint recognition result.

[0019] This invention also provides an offline personalized voiceprint learning method, applied in the cloud, including:

[0020] Based on the first initial general voiceprint learning data, meta-learning is performed on the neural network model to obtain a pre-trained model. Then, based on multiple single-task general voiceprint learning data, meta-learning is performed on the pre-trained model to obtain multiple first training models.

[0021] Calculate the discrete index of the structural parameters at the same position in the plurality of first training models, and freeze the target parameters in the pre-trained model based on the discrete index;

[0022] Based on the first initial general voiceprint learning data, meta-learning is performed on the frozen model to obtain the general voiceprint recognition model.

[0023] Based on the multiple single-task general voiceprint learning data, target general voiceprint learning data is determined, and the general voiceprint recognition model and the target general voiceprint learning data are configured to the device.

[0024] According to the present invention, an offline personalized voiceprint learning method is provided, wherein the target parameter is a structural parameter whose discrete index is less than a preset threshold.

[0025] According to an offline personalized voiceprint learning method provided by the present invention, the step of determining target general voiceprint learning data based on the plurality of single-task general voiceprint learning data includes:

[0026] Based on the multiple single-task general voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain a second training model. Based on the second initial general voiceprint learning data, the general voiceprint recognition model and the second training model are tested respectively to obtain a fourth test result.

[0027] Based on the accuracy of the general voiceprint recognition model and the second training model for the multiple single-task general voiceprint learning data, and / or the fourth test result, the multiple single-task general voiceprint learning data are filtered to obtain the target general voiceprint learning data.

[0028] The present invention also provides a speaker separation method, applied at the device end, comprising:

[0029] Determine user separation configuration information; the user separation configuration information includes at least one of real-time separation mode and non-real-time separation mode.

[0030] In real-time separation mode, a real-time voice stream is received, and a target personalized voiceprint recognition model obtained based on the offline personalized voiceprint learning method described above is used to perform voiceprint recognition on the real-time voice stream to obtain a second voiceprint recognition result. Based on the second voiceprint recognition result, or based on the speaker registration information and the second voiceprint recognition result, speaker separation is performed on the real-time voice stream.

[0031] In the non-real-time separation mode, global speech is acquired, and based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the global speech to obtain a third voiceprint recognition result. Based on the third voiceprint recognition result, or based on the speaker registration information and the third voiceprint recognition result, speaker separation is performed on the global speech.

[0032] According to a speaker separation method provided by the present invention, the speaker registration information includes offline speaker voiceprint registration information and / or dynamic speaker tagging information.

[0033] According to a speaker separation method provided by the present invention, the speaker separation of the real-time speech stream is performed based on speaker registration information and the second voiceprint recognition result, and then the method further includes:

[0034] The speaker separation result is stored in a temporary voiceprint database in correspondence with the second voiceprint recognition result.

[0035] If the real-time voice stream is interrupted, the interrupted voice is received, and based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the interrupted voice to obtain a third voiceprint recognition result.

[0036] Based on the third voiceprint recognition result, the temporary voiceprint database is used to perform speaker separation on the interrupted speech.

[0037] The present invention also provides a speaker separation integrated machine, including a memory, a processor, a target personalized voiceprint recognition model stored in the memory, and a computer program stored in the processor and capable of running on the processor;

[0038] When the processor executes the computer program, it calls the target personalized voiceprint recognition model to implement the speaker separation method described above.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the offline personalized voiceprint learning method or the speaker separation method as described above.

[0040] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the offline personalized voiceprint learning method or the speaker separation method as described above.

[0041] The offline personalized voiceprint learning method and speaker separation method provided by this invention, on the device side, firstly determine the built-in general voiceprint recognition model and target general voiceprint learning data, and acquire personalized voiceprint learning data; using the training data and personalized voiceprint learning data, perform meta-learning on the general voiceprint recognition model to obtain an initial personalized voiceprint recognition model; finally, using test data, test the general voiceprint recognition model and the initial personalized voiceprint recognition model respectively to obtain a first test result, and based on the first test result, determine the target personalized voiceprint recognition model. This method can achieve offline personalized training by using personalized voiceprint learning data and the training data built into the device to perform meta-learning and testing on the general voiceprint recognition model built into the device, without needing to transmit model update packages to each device. This not only greatly improves model update efficiency and reduces costs, but also avoids security issues introduced by the transmission of update packages. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a flowchart illustrating the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0044] Figure 2 This is a schematic diagram of the discretization training process in the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0045] Figure 3 This is a schematic diagram of the process for obtaining personalized voiceprint learning data in the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0046] Figure 4 This is a schematic diagram of multiple candidate results corresponding to sample speech in the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0047] Figure 5 This is one of the schematic diagrams illustrating the feedback information and application effect in the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0048] Figure 6 This is the second illustration of feedback information and its application effect in the offline personalized voiceprint learning method applied to the device side provided by the present invention.

[0049] Figure 7 This is one of the flowcharts of the offline personalized voiceprint learning method applied to the cloud provided by the present invention.

[0050] Figure 8 This is the second flowchart of the offline personalized voiceprint learning method applied to the cloud provided by the present invention.

[0051] Figure 9 This is a schematic diagram of the overall process of the offline personalized voiceprint learning method provided by the present invention.

[0052] Figure 10 This is one of the flowcharts of the speaker separation method provided by the present invention.

[0053] Figure 11 This is the second flowchart of the speaker separation method provided by the present invention.

[0054] Figure 12This is a schematic diagram of the speaker blind segmentation effect in the speaker separation method provided by the present invention.

[0055] Figure 13 This is a schematic diagram illustrating the offline speaker voiceprint registration information and its application effect in the speaker separation method provided by the present invention.

[0056] Figure 14 This is a schematic diagram illustrating the dynamic speaker tagging information and its application effect in the speaker separation method provided by the present invention.

[0057] Figure 15 This is a schematic diagram of the existing speaker separation results before and after the interruption of the real-time speech stream in the speaker separation method provided by the present invention.

[0058] Figure 16 This is a schematic diagram illustrating the desired speaker separation results before and after a real-time speech stream interruption in the speaker separation method provided by this invention.

[0059] Figure 17 This is a schematic diagram of the process after the real-time speech stream is interrupted in the speaker separation method provided by the present invention.

[0060] Figure 18 This is one of the structural schematic diagrams of the offline personalized voiceprint learning device provided by the present invention.

[0061] Figure 19 This is the second structural schematic diagram of the offline personalized voiceprint learning device provided by the present invention.

[0062] Figure 20 This is a schematic diagram of the speaker separation system provided by the present invention.

[0063] Figure 21 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0065] When updating the voiceprint recognition model on existing electronic devices offline, it is usually necessary to transmit the update package to each target device via physical media such as a USB drive, Secure Digital Card (SD) or other removable storage media, and then install and replace it locally on the target device; or, a dedicated update server or an administrator's mobile device can be used as an intermediary to transmit the update package to the target device wirelessly via Bluetooth, near-field communication or other wireless methods.

[0066] The aforementioned offline update scheme requires individual updates for each target device, which is time-consuming and costly. Moreover, the physical medium may be lost, damaged, or tampered with during transmission.

[0067] Based on this, this embodiment of the invention provides an offline personalized voiceprint learning method.

[0068] Figure 1 This is a flowchart illustrating an offline personalized voiceprint learning method provided in an embodiment of the present invention, such as... Figure 1 As shown, this method is applied to the device side and includes:

[0069] S11, determine the built-in general voiceprint recognition model and target general voiceprint learning data, and obtain personalized voiceprint learning data; the target general voiceprint learning data includes training data and test data;

[0070] S12, Based on the training data and the personalized voiceprint learning data, perform meta-learning on the general voiceprint recognition model to obtain an initial personalized voiceprint recognition model;

[0071] S13, based on the test data, the general voiceprint recognition model and the initial personalized voiceprint recognition model are tested respectively to obtain a first test result, and based on the first test result, the target personalized voiceprint recognition model is determined.

[0072] Specifically, the offline personalized voiceprint learning method provided in this embodiment of the invention is executed by an offline personalized voiceprint learning device, which can be configured in a target device. The target device can be an electronic device that uses the target personalized voiceprint recognition model obtained by the offline personalized voiceprint learning method to perform downstream personalized voiceprint recognition tasks, without being specifically limited here.

[0073] First, step S11 is executed to determine the built-in general voiceprint recognition model M and the target general voiceprint learning data U. Both the general voiceprint recognition model M and the target general voiceprint learning data U can be determined by the cloud and configured to the target device when the target device leaves the factory.

[0074] The general voiceprint recognition model M can be a voiceprint recognition model applicable to different tasks. By inputting audio, it can output the voiceprint features in the audio, thereby distinguishing different speakers.

[0075] The target general voiceprint learning data U can be general voiceprint learning data used to assist the target device in performing meta-learning of the general voiceprint recognition model, and can include general speech data and speaker labels therein.

[0076] The target general voiceprint learning data U can be divided into training data u1 and test data u2 according to their functions. Training data u1 is used for meta-learning of the general voiceprint recognition model, and test data u2 is used to test the initial personalized voiceprint recognition model obtained from meta-learning. It is understood that training data u1 and test data u2 do not overlap.

[0077] Then, personalized voiceprint learning data V can be obtained by interacting with the user. This personalized voiceprint learning data V can be the voiceprint learning data for the target task to be performed by the target device, and can include personalized voice data and speaker tags therein.

[0078] Subsequently, a meta-learning and testing process is performed to achieve offline personalized training of the general voiceprint recognition model M on the device, such as... Figure 2 As shown.

[0079] In step S12, the training data u1 and the personalized voiceprint learning data V are used together as training samples to perform meta-learning on the general voiceprint recognition model M, resulting in the initial personalized voiceprint recognition model P0. Through meta-learning, new tasks can be learned quickly, thus adapting rapidly to new data.

[0080] In step S13, using test data u2, the general voiceprint recognition model M and the initial personalized voiceprint recognition model P0 can be tested respectively to obtain the first test result. The first test result includes the second test result corresponding to the general voiceprint recognition model M and the third test result corresponding to the initial personalized voiceprint recognition model P0. The second test result can represent the accuracy of the general voiceprint recognition model M, and the third test result can represent the accuracy of the initial personalized voiceprint recognition model P0.

[0081] Using the results of the second and third tests, the better-performing personalized voiceprint recognition model P can be selected from the general voiceprint recognition model M and the initial personalized voiceprint recognition model P0 as the target personalized voiceprint recognition model P. In other words, if the initial personalized voiceprint recognition model P0 shows improved performance on test data u2 compared to the general voiceprint recognition model M, then the initial personalized voiceprint recognition model P0 is used to replace the general voiceprint recognition model M as the target personalized voiceprint recognition model P; otherwise, the general voiceprint recognition model M continues to be used as the target personalized voiceprint recognition model P to avoid model performance degradation.

[0082] The offline personalized voiceprint learning method provided in this embodiment of the invention is applied to the device side. First, a built-in general voiceprint recognition model and target general voiceprint learning data are determined, and personalized voiceprint learning data is acquired. Then, using the training data and personalized voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain an initial personalized voiceprint recognition model. Finally, using test data, the general voiceprint recognition model and the initial personalized voiceprint recognition model are tested respectively to obtain a first test result. Based on the first test result, the target personalized voiceprint recognition model is determined. This method utilizes personalized voiceprint learning data and the device-side built-in training data to perform meta-learning and testing on the device-side built-in general voiceprint recognition model, achieving offline personalized training without needing to transmit model update packages to each device. This not only greatly improves model update efficiency and reduces costs but also avoids security issues introduced by updating package transmission.

[0083] Based on the above embodiments, the first test result includes the second test result corresponding to the general voiceprint recognition model and the third test result corresponding to the initial personalized voiceprint recognition model;

[0084] The step of determining the target personalized voiceprint recognition model based on the first test result includes:

[0085] If the second test result is better than the third test result, then the general voiceprint recognition model is determined to be the target personalized voiceprint recognition model;

[0086] If the third test result is better than the second test result, then the initial personalized voiceprint recognition model is determined to be the target personalized voiceprint recognition model.

[0087] Specifically, when determining the target personalized voiceprint recognition model, the second test result can be compared with the third test result. If the second test result is better than the third test result, that is, the accuracy of the general voiceprint recognition model on the test data is higher than the accuracy of the initial personalized voiceprint recognition model on the test data, then the performance of the general voiceprint recognition model is considered to be better than the initial personalized voiceprint recognition model. Thus, the general voiceprint recognition model is determined as the target personalized voiceprint recognition model, that is, the general voiceprint recognition model built into the device is retained for subsequent applications.

[0088] If the third test result is better than the second test result, that is, the accuracy of the initial personalized voiceprint recognition model on the test data is higher than that of the general voiceprint recognition model on the test data, then the performance of the initial personalized voiceprint recognition model is considered to be better than that of the general voiceprint recognition model. Therefore, the initial personalized voiceprint recognition model is determined as the target personalized voiceprint recognition model, that is, the general voiceprint recognition model built into the device is replaced with the initial personalized voiceprint recognition model for subsequent applications.

[0089] In this embodiment of the invention, the target personalized voiceprint recognition model can always be a high-performance voiceprint recognition model.

[0090] Based on the above embodiments, the acquisition of personalized voiceprint learning data includes:

[0091] Obtain sample audio;

[0092] Based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the sample speech to determine the first voiceprint recognition result, and based on the first voiceprint recognition result, speaker separation is performed on the sample speech to obtain the speaker separation result.

[0093] The system receives feedback from users regarding the speaker separation results and constructs the personalized voiceprint learning data based on the feedback and the sample speech.

[0094] Specifically, such as Figure 3 As shown, when acquiring personalized voiceprint learning data, sample speech can be obtained first, and then the sample speech can be input into the target personalized voiceprint recognition model. The target personalized voiceprint recognition model performs voiceprint recognition on the sample speech to obtain the first voiceprint recognition result. This first voiceprint recognition result is the voiceprint features contained in the sample speech.

[0095] Subsequently, the speaker separation results are obtained by using the first voiceprint recognition result to separate the speaker from the sample speech. It is understood that different speakers have different voiceprint features; therefore, different speakers can be identified based on the different voiceprint features in the first voiceprint recognition result. Here, the speaker separation result may include transcribed text of speech segments from different speakers in the sample speech.

[0096] Speaker separation results can include multiple candidate results, which can be sorted from highest to lowest score. The text content in the sample speech is as follows: Figure 4 In conversation 1, the speaker separation results obtained include result 1 and result 2, with result 1 having a higher accuracy score than result 2.

[0097] Finally, the device can provide a user interface to display the speaker separation results and allow users to input feedback information on the speaker separation results by marking and correcting them.

[0098] The feedback can include various types, such as speaker recognition error, voice confusion, and incomplete separation. Users can select a specific feedback type and submit detailed feedback information. The feedback information can include the user-marked error, the corrected result, and the time and context of the feedback.

[0099] Users can input feedback information on the device's display interface. The device can then receive and analyze this feedback, extracting key information, such as analyzing the error patterns and frequencies marked by the user.

[0100] Furthermore, the device can utilize the key information in this feedback and sample speech to construct personalized voiceprint learning data.

[0101] In this embodiment of the invention, based on the user's feedback on the speaker separation results, and based on the feedback information and sample speech, the constructed personalized voiceprint learning data can better meet the user's needs, thereby improving the performance of downstream tasks of the target personalized voiceprint recognition model on the device and increasing user satisfaction.

[0102] Based on the above embodiments, the feedback information includes the speakers of interest in the sample speech and the user's modification data when an error occurs in the first voiceprint recognition result.

[0103] Specifically, the feedback information can include speakers of interest from the sample speech. Since different users focus on different speakers, the user's feedback information can only label speakers of interest, leaving others unseparated, thus further improving speaker separation performance. For example, the speaker separation result might look like this: Figure 5 As shown in Session 1, after receiving the speaker of interest from the feedback information, the speaker separation result can be transformed into... Figure 5 In conversation 2, the "boss" is the speaker of interest.

[0104] The feedback information may also include the user's modified data when an error occurred in the initial voiceprint recognition result. For example, the speaker separation result may include... Figure 6 As shown in Session 1, the speaker Li Si was mistakenly identified as Zhang San. The user actively corrected Zhang San to Li Si. After receiving the modified data in the feedback information, the device can change the speaker separation result to a new Session 1. Moreover, in the subsequent Session 2, the speakers Zhang San and Li Si will also be correctly separated.

[0105] Based on the above embodiments, such as Figure 7 As shown, this embodiment of the invention also provides an offline personalized voiceprint learning method, applied in the cloud, including:

[0106] S21, based on the first initial general voiceprint learning data, meta-learning is performed on the neural network model to obtain a pre-trained model, and based on multiple single-task general voiceprint learning data, meta-learning is performed on the pre-trained model respectively to obtain multiple first training models;

[0107] S22, calculate the discrete index of the structural parameters at the same position in the plurality of first training models, and freeze the target parameters in the pre-trained model based on the discrete index;

[0108] S23, Based on the first initial general voiceprint learning data, perform meta-learning on the frozen model to obtain the general voiceprint recognition model;

[0109] S24, based on the multiple single-task general voiceprint learning data, determine the target general voiceprint learning data, and configure the general voiceprint recognition model and the target general voiceprint learning data to the device.

[0110] Specifically, the offline personalized voiceprint learning method provided in this embodiment of the invention is executed by an offline personalized voiceprint learning device, which can be configured in the cloud.

[0111] First, execute step S21, such as... Figure 8 As shown, a large amount of initial general-purpose voiceprint learning data D is used to perform meta-learning on the neural network model to obtain a pre-trained model A. The initial general-purpose voiceprint learning data D can be determined based on publicly available speech data and may include general-purpose speech data and speaker labels therein. The specific structure of the neural network model can be selected as needed and is not specifically limited here.

[0112] Multiple single-task general-purpose voiceprint learning datasets are used to perform meta-learning on the pre-trained models, resulting in multiple first-trained models. The number of single-task general-purpose voiceprint learning datasets can be set as needed, for example, it can be set to n, represented as d1, d2, ..., dn. Each single-task general-purpose voiceprint learning dataset may include a small amount of general-purpose voiceprint learning data specific to that single task.

[0113] By performing meta-learning on the pre-trained model using the general voiceprint learning data for each single task, a first training model can be obtained. That is, there is a one-to-one correspondence between the general voiceprint learning data for each single task and the first training model. Each first training model can be represented as B1, B2, ..., Bn.

[0114] Then, step S22 is executed to calculate the discrete index of the structural parameters at the same position in each of the first training models. This discrete index may include parameters such as variance and standard deviation, which are used to represent the degree of discretization of the structural parameters at the same position.

[0115] Using discrete metrics, target parameters in pre-trained model A can be selected and frozen. Here, the target parameters can be structural parameters whose discrete metrics are less than a preset threshold. This preset threshold can be set as needed and is not specifically limited here.

[0116] Here, structural parameters with discrete indices greater than or equal to a preset threshold vary significantly across different tasks and are sensitive to personalized voiceprint learning data, while target parameters with discrete indices less than the preset threshold vary less across different tasks and are not sensitive to personalized voiceprint learning data. Therefore, freezing the target parameters with discrete indices less than the preset threshold in the pre-trained model A can reduce the computational load on the device side for training the general voiceprint recognition model.

[0117] Then, step S23 is executed, using the first initial general voiceprint learning data D to perform meta-learning on the frozen model to obtain the general voiceprint recognition model M.

[0118] Finally, step S24 is executed to determine the target general voiceprint learning data using multiple single-task general voiceprint learning data. For example, multiple single-task general voiceprint learning data can be directly used as the target general voiceprint learning data, or the target general voiceprint learning data can be obtained by filtering from multiple single-task general voiceprint learning data. No specific limitation is made here.

[0119] Furthermore, the general voiceprint recognition model and the target general voiceprint learning data can be configured on the device.

[0120] The offline personalized voiceprint learning method provided in this embodiment of the invention is applied to the cloud. By training a general voiceprint recognition model in the cloud, and selecting target parameters in the pre-trained model to freeze during the training process, the frozen model is then subjected to meta-learning using the first initial general voiceprint learning data to obtain a general voiceprint recognition model. This can greatly reduce the computational load of subsequent model training on the device.

[0121] Based on the above embodiments, determining the target general voiceprint learning data based on the plurality of single-task general voiceprint learning data includes:

[0122] Based on the multiple single-task general voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain a second training model. Based on the second initial general voiceprint learning data, the general voiceprint recognition model and the second training model are tested respectively to obtain a fourth test result.

[0123] Based on the accuracy of the general voiceprint recognition model and the second training model for the multiple single-task general voiceprint learning data, and / or the fourth test result, the multiple single-task general voiceprint learning data are filtered to obtain the target general voiceprint learning data.

[0124] Specifically, when determining the target general voiceprint learning data, multiple single-task general voiceprint learning data can be used simultaneously to perform meta-learning on the general voiceprint recognition model M, resulting in a second training model W. The accuracy of the multiple single-task general voiceprint learning data is then used to filter the data based on the general voiceprint recognition model M and the second training model W, thus obtaining the target general voiceprint learning data. For example, data that the general voiceprint recognition model M cannot correctly identify but the second training model W can correctly identify can be selected to verify the effectiveness of the meta-learning.

[0125] Then, based on the second initial general voiceprint learning data, the general voiceprint recognition model M and the second trained model W are tested respectively to obtain the fourth test result. The second initial general voiceprint learning data can be general voiceprint learning data that is different from the first initial general voiceprint learning data and the general voiceprint learning data for each single task.

[0126] The fourth test result can include the fifth test result corresponding to the general voiceprint recognition model M and the sixth test result corresponding to the second trained model W. Using the fifth and sixth test results, multiple single-task general voiceprint learning data can be filtered. For example, data correctly identified by both the general voiceprint recognition model M and the second trained model W can be filtered to verify that meta-learning is invariant on general data. Finally, the target general voiceprint learning data can be obtained after filtering.

[0127] In summary, the offline personalized voiceprint learning method provided in this embodiment of the invention obtains a general voiceprint recognition model through meta-learning training in the cloud and configures it on the device. On the device, fine-tuning of the general voiceprint recognition model is then performed. The overall process is as follows: Figure 9 As shown, it includes:

[0128] The cloud can use the first initial general voiceprint learning data D and multiple single-task general voiceprint learning data to obtain a general voiceprint recognition model M, which is pre-installed on the device at the factory.

[0129] By using the general voiceprint recognition model M to filter multiple single-task general voiceprint learning data, the target general voiceprint learning data can be obtained and pre-installed on the local device at the factory.

[0130] Highly reliable personalized voiceprint learning data V is obtained through a pre-user interaction step.

[0131] On the device side, a general voiceprint recognition model M is trained using the target general voiceprint learning data U and the personalized voiceprint learning data V to obtain the target personalized voiceprint recognition model P.

[0132] Current speaker separation technologies primarily employ a "cloud + edge" approach. The device initiates a session request, uploads voice data to a cloud service for processing via network connectivity, and then transmits the speaker separation results back to the device. However, existing cloud processing solutions require devices to be connected to the internet, and users need to upload voice and other data, posing a risk of privacy breaches. Furthermore, considering the effectiveness of real-time and non-real-time separation, different algorithms are typically used, requiring separate deployments of real-time and non-real-time services in the cloud. This independent deployment approach necessitates users accessing two sets of services to achieve real-time and non-real-time separation, resulting in high deployment and maintenance costs.

[0133] Based on the above embodiments, such as Figure 10 As shown, this embodiment of the invention also provides a speaker separation method, applied to the device side, including:

[0134] S31, determine user separation configuration information; the user separation configuration information includes at least one of real-time separation mode and non-real-time separation mode;

[0135] S32, in real-time separation mode, receive real-time speech stream, and perform voiceprint recognition on the real-time speech stream based on the target personalized voiceprint recognition model obtained by the offline personalized voiceprint learning method provided in the above embodiments to obtain a second voiceprint recognition result. Based on the second voiceprint recognition result, or based on the speaker registration information and the second voiceprint recognition result, perform speaker separation on the real-time speech stream.

[0136] S33, in non-real-time separation mode, acquire global speech, and perform voiceprint recognition on the global speech based on the target personalized voiceprint recognition model to obtain a third voiceprint recognition result. Based on the third voiceprint recognition result, perform global speaker separation on the global speech.

[0137] Specifically, the speaker separation method provided in this embodiment of the invention is executed by a speaker separation system, which can be configured in a target device. The target device can be an electronic device that performs speaker separation tasks using a target personalized voiceprint recognition model obtained by an offline personalized voiceprint learning method, without being specifically limited here.

[0138] First, execute step S31 to determine the user separation configuration information. The user separation configuration information can be a separation mode pre-configured by the user, which can include at least one of real-time separation mode and non-real-time separation mode. That is, the user can configure only the real-time separation mode, only the non-real-time separation mode, or both real-time separation mode and non-real-time separation mode.

[0139] Among them, the real-time separation mode refers to the mode of directly separating the speaker after performing voiceprint recognition on the real-time speech stream, while the non-real-time separation mode refers to the mode of performing voiceprint recognition on the global speech after the real-time speech stream has been transmitted and then separating the speaker.

[0140] Perform step S32, such as Figure 11 As shown, in real-time separation mode, a real-time voice stream is received. Subsequently, voice activity detection (VAD) can be performed on the real-time voice stream to detect the presence of a voice signal.

[0141] Using the target personalized voiceprint recognition model obtained through the offline personalized voiceprint learning method provided in the above embodiments, voiceprint recognition is performed on the real-time speech stream to obtain a second voiceprint recognition result. This second voiceprint recognition result refers to the voiceprint features in the real-time speech stream.

[0142] Subsequently, the second voiceprint recognition result can be used to perform speaker segmentation on the real-time speech stream, obtaining the real-time speaker segmentation result. At this point, speaker blind segmentation can be assumed, meaning that if there are several speakers in the real-time speech stream, it will be segmented into speech fragments for those speakers. The speaker segmentation result is displayed externally in the format of Speaker 1, Speaker 2, Speaker 3, ..., Speaker N. For example... Figure 12 The conversation shown includes audio clips from Speaker 1 and Speaker 2.

[0143] Speaker separation can also be performed on real-time speech streams using speaker registration information and second voiceprint recognition results. The speaker registration information refers to the correspondence between speakers and voiceprint features.

[0144] The speaker registration information may include offline speaker voiceprint registration information and / or dynamic speaker tagging information. Offline speaker voiceprint registration information refers to the correspondence between the speaker and voiceprint features entered by the user before the session. The voiceprint features in the second voiceprint recognition result are matched with the voiceprint features in the offline speaker voiceprint registration information, and when a match is successful, the matched speaker information from the offline speaker voiceprint registration information is displayed, such as... Figure 13 Zhang San in the example. If a match is found, the default speaker blind segmentation result will continue to be displayed, such as... Figure 13 Speaker 2 in the text.

[0145] Dynamic speaker tagging information refers to speaker tagging information manually entered by the user during a conversation. After the user tags a speaker, the currently tagged speaker and that speaker in subsequent real-time audio streams are all replaced with the tagged speaker information. For example... Figure 14 As shown, if the user marks the speaker 2 on the left as Li Si's voiceprint, then the currently marked speaker 2, as well as speaker 2 in subsequent real-time audio streams, will be replaced with Li Si, as shown. Figure 14 Right side.

[0146] In real-time separation mode, it can identify and separate speech segments from multiple speakers in real time, with a response time of less than 800ms.

[0147] In step S33, under non-real-time separation mode, global speech can be obtained by caching the real-time speech stream, and then the target personalized voiceprint recognition model can be used to perform voiceprint recognition on the global speech to obtain the third voiceprint recognition result. The third voiceprint recognition result refers to the voiceprint features in the global speech.

[0148] Finally, the third voiceprint recognition result can be used to perform speaker separation on the global speech by default blind speaker segmentation, resulting in a non-real-time speaker separation result. Alternatively, speaker registration information and the third voiceprint recognition result can also be used to perform speaker separation on the global speech; no specific limitations are specified here.

[0149] Understandably, in non-real-time separation mode, asynchronous thread processing fully utilizes multi-core performance, improving program execution efficiency; one hour of global speech can yield speaker separation results within 5 seconds. Both real-time and non-real-time separation modes employ a target-specific personalized voiceprint recognition model for voiceprint recognition, significantly reducing computational load, improving efficiency, and minimizing memory usage. Furthermore, both real-time and non-real-time separation modes are configurable, allowing for flexible user adoption.

[0150] In this embodiment of the invention, real-time and non-real-time separation functions can be supported simultaneously, eliminating the need to deploy two separate engines and effectively reducing deployment costs. Furthermore, the training of the personalized voiceprint recognition model and the speaker separation operation are both performed on the device side, without requiring an internet connection, thus eliminating the risk of user privacy leakage.

[0151] Based on the above embodiments, the step of performing speaker separation on the real-time speech stream based on speaker registration information and the second voiceprint recognition result further includes:

[0152] The speaker separation result is stored in a temporary voiceprint database in correspondence with the second voiceprint recognition result.

[0153] If the real-time voice stream is interrupted, the interrupted voice is received, and based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the interrupted voice to obtain a third voiceprint recognition result.

[0154] Based on the third voiceprint recognition result, the temporary voiceprint database is used to perform speaker separation on the interrupted speech.

[0155] Specifically, in this embodiment of the invention, the method can also automatically associate the same speaker in different conversations, that is, resume the conversation, which can greatly improve the user experience.

[0156] In real-world meeting scenarios, a meeting may be interrupted due to either active or passive reasons, and the same speaker may be portrayed in different roles before and after the interruption. For example... Figure 15 As shown, assume that before the interruption, speakers A, B, and C were identified as speaker 1, speaker 2, and speaker 3. Since the session is generally treated as a new session after an interruption, if only speakers B and C speak sequentially in the new session, then speaker B will be identified as speaker 1, and speaker C as speaker 2. From the user's perspective, it would be desirable for speaker B to continue to be identified as speaker 2, and speaker C to remain speaker 3. Figure 16 .

[0157] Therefore, after speaker separation of the real-time speech stream, the speaker separation results and the second voiceprint recognition results can be stored in a temporary voiceprint database. This storage process will not be interrupted by session interruption. Thus, the temporary voiceprint database contains the voiceprint features of each speaker in the speaker separation results before and after session interruption, and may include voiceprint features of both registered and unregistered speakers. This temporary voiceprint database can be stored on a disk.

[0158] Therefore, in the event of a real-time voice stream interruption, the interrupted voice can be received and then input into the target personalized voiceprint recognition model. The target personalized voiceprint recognition model performs voiceprint recognition on the interrupted voice to obtain the third voiceprint recognition result.

[0159] After that, as Figure 17 As shown, the third voiceprint recognition result is used to match the voiceprint features in the temporary voiceprint database. If the match is successful, the speaker whose voiceprint features match in the temporary voiceprint database can be used as the speaker in the interrupted speech. If the match fails, the speaker is used as the new speaker. In this way, the effect of linking speakers before and after the interruption and resuming the conversation can be achieved.

[0160] In summary, addressing the challenge of upgrading voiceprint models and algorithms for secure, non-networked devices, this invention proposes an offline personalized voiceprint learning method and a purely offline speaker separation method, illustrating the application of the target personalized voiceprint recognition model in speaker separation tasks. The offline personalized voiceprint learning method in this invention fully considers the limitations of offline devices, using sample speech accumulated during user usage for personalized training locally. This allows for continuous iteration and improvement of voiceprint recognition performance while maintaining confidentiality. The offline personalized voiceprint learning method has low deployment costs, low maintenance costs, and is user privacy-friendly, providing an excellent solution to the problem of "difficulty in upgrading offline algorithms on the device side," and also serving as a good supplement and extension to cloud-based separation solutions. Furthermore, this offline personalized voiceprint learning method is not only applicable to the offline voiceprint field but can be extended to many other areas such as images, videos, and text, demonstrating good research prospects and practical application value. Based on the offline personalized voiceprint learning method, this invention also provides a purely offline speaker separation method, capable of both real-time and non-real-time speaker separation, and also featuring functions such as "multi-functional registration" and "session continuation."

[0161] like Figure 18 As shown, based on the above embodiments, this embodiment of the invention provides an offline personalized voiceprint learning device, applied to a device, comprising:

[0162] The built-in information acquisition module 181 is used to determine the built-in general voiceprint recognition model and the target general voiceprint learning data, and to acquire personalized voiceprint learning data; the target general voiceprint learning data includes training data and test data.

[0163] The first meta-learning module 182 is used to perform meta-learning on the general voiceprint recognition model based on the training data and the personalized voiceprint learning data to obtain an initial personalized voiceprint recognition model.

[0164] The first test module 183 is used to test the general voiceprint recognition model and the initial personalized voiceprint recognition model based on the test data, respectively, to obtain a first test result, and to determine the target personalized voiceprint recognition model based on the first test result.

[0165] Specifically, the functions of each module in the offline personalized voiceprint learning device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-mentioned method-type embodiment with the device as the execution subject, and the achieved effect is also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0166] like Figure 19 As shown, based on the above embodiments, this embodiment of the invention provides an offline personalized voiceprint learning device applied in the cloud, comprising:

[0167] The second meta-learning module 191 is used to perform meta-learning on the neural network model based on the first initial general voiceprint learning data to obtain a pre-trained model, and to perform meta-learning on the pre-trained model based on multiple single-task general voiceprint learning data to obtain multiple first training models.

[0168] The discrete index calculation module 192 is used to calculate the discrete index of the structural parameters at the same position in the plurality of first training models, and to select the target parameters in the pre-trained model for freezing based on the discrete index.

[0169] The third meta-learning module 193 is used to perform meta-learning on the frozen model based on the first initial general voiceprint learning data to obtain the general voiceprint recognition model.

[0170] The device configuration module 194 is used to determine the target general voiceprint learning data based on the multiple single-task general voiceprint learning data, and configure the general voiceprint recognition model and the target general voiceprint learning data to the device.

[0171] Specifically, the functions of each module in the offline personalized voiceprint learning device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-mentioned method-type embodiment with the cloud as the execution subject, and the achieved effect is also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0172] like Figure 20 As shown, based on the above embodiments, this embodiment of the invention provides a speaker separation system applied to a device, comprising:

[0173] The configuration information determination module 201 is used to determine user separation configuration information; the user separation configuration information includes at least one of real-time separation mode and non-real-time separation mode.

[0174] The first speaker separation module 202 is used to receive real-time speech stream in real-time separation mode, and perform speaker recognition on the real-time speech stream based on the target personalized speaker recognition model obtained by the above-mentioned offline personalized speaker learning method to obtain a second speaker recognition result. Based on the second speaker recognition result, or based on the speaker registration information and the second speaker recognition result, speaker separation is performed on the real-time speech stream.

[0175] The second speaker separation module 203 is used to acquire global speech in non-real-time separation mode, and perform voiceprint recognition on the global speech based on the target personalized voiceprint recognition model to obtain a third voiceprint recognition result. Based on the third voiceprint recognition result, or based on speaker registration information and the third voiceprint recognition result, speaker separation is performed on the global speech.

[0176] Figure 21 An example is a schematic diagram of the physical structure of a speaker separation integrated machine, such as... Figure 21 As shown, the speaker separation integrated machine may include: a processor 210, a communication interface 220, a memory 230, and a communication bus 240. The processor 210, communication interface 220, and memory 230 communicate with each other via the communication bus 240. The memory 230 may store a target personalized voiceprint recognition model, and the processor 210 may store a computer program that can run on the processor 210.

[0177] The processor 210 can call the target personalized voiceprint recognition model in the memory 230 to execute the speaker separation method provided in the above embodiments.

[0178] Furthermore, the logical instructions in the aforementioned memory 230 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0179] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the offline personalized voiceprint learning method or the speaker separation method provided by the above methods.

[0180] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the offline personalized voiceprint learning method or speaker separation method provided by the above methods.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An offline personalized voiceprint learning method, characterized in that, Applied to the cloud, including: Based on the first initial general voiceprint learning data, meta-learning is performed on the neural network model to obtain a pre-trained model. Then, based on multiple single-task general voiceprint learning data, meta-learning is performed on the pre-trained model to obtain multiple first training models. Calculate the discrete index of the structural parameters at the same position in the plurality of first training models, and freeze the target parameters in the pre-trained model based on the discrete index; Based on the first initial general voiceprint learning data, meta-learning is performed on the frozen model to obtain the general voiceprint recognition model. Based on the multiple single-task general voiceprint learning data, target general voiceprint learning data is determined, and the general voiceprint recognition model and the target general voiceprint learning data are configured to the device. The device uses personalized voiceprint learning data and training and testing data in the target general voiceprint learning data to perform meta-learning and testing on the general voiceprint recognition model to achieve offline personalized training.

2. The offline personalized voiceprint learning method according to claim 1, characterized in that, The target parameter is a structural parameter whose discrete index is less than a preset threshold.

3. The offline personalized voiceprint learning method according to claim 1, characterized in that, The determination of target general voiceprint learning data based on the multiple single-task general voiceprint learning data includes: Based on the multiple single-task general voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain a second training model. Based on the second initial general voiceprint learning data, the general voiceprint recognition model and the second training model are tested respectively to obtain a fourth test result. Based on the accuracy of the general voiceprint recognition model and the second training model for the multiple single-task general voiceprint learning data, and / or the fourth test result, the multiple single-task general voiceprint learning data are filtered to obtain the target general voiceprint learning data.

4. The offline personalized voiceprint learning method according to claim 1, characterized in that, The device is specifically used for: Acquire personalized voiceprint learning data; Based on the training data and the personalized voiceprint learning data, meta-learning is performed on the general voiceprint recognition model to obtain an initial personalized voiceprint recognition model. Based on the test data, the general voiceprint recognition model and the initial personalized voiceprint recognition model are tested respectively to obtain a first test result, and based on the first test result, the target personalized voiceprint recognition model is determined.

5. The offline personalized voiceprint learning method according to claim 4, characterized in that, The first test result includes the second test result corresponding to the general voiceprint recognition model and the third test result corresponding to the initial personalized voiceprint recognition model; The step of determining the target personalized voiceprint recognition model based on the first test result includes: If the second test result is better than the third test result, then the general voiceprint recognition model is determined to be the target personalized voiceprint recognition model; If the third test result is better than the second test result, then the initial personalized voiceprint recognition model is determined to be the target personalized voiceprint recognition model.

6. The offline personalized voiceprint learning method according to claim 4, characterized in that, The acquisition of personalized voiceprint learning data includes: Obtain sample audio; Based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the sample speech to determine the first voiceprint recognition result, and based on the first voiceprint recognition result, speaker separation is performed on the sample speech to obtain the speaker separation result. The system receives feedback from users regarding the speaker separation results and constructs the personalized voiceprint learning data based on the feedback and the sample speech.

7. The offline personalized voiceprint learning method according to claim 6, characterized in that, The feedback information includes the speakers of interest in the sample speech and the user's modification data when an error occurs in the first voiceprint recognition result.

8. A speaker separation method, characterized in that, Applied to the device side, including: Determine user separation configuration information; the user separation configuration information includes at least one of real-time separation mode and non-real-time separation mode. In real-time separation mode, a real-time voice stream is received, and a target personalized voice recognition model obtained by offline personalized training is implemented on the device side in the offline personalized voiceprint learning method as described in any one of claims 1-7. Voiceprint recognition is performed on the real-time voice stream to obtain a second voiceprint recognition result. Based on the second voiceprint recognition result, or based on the speaker registration information and the second voiceprint recognition result, speaker separation is performed on the real-time voice stream. In the non-real-time separation mode, global speech is acquired, and based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the global speech to obtain a third voiceprint recognition result. Based on the third voiceprint recognition result, or based on the speaker registration information and the third voiceprint recognition result, speaker separation is performed on the global speech.

9. The speaker separation method according to claim 8, characterized in that, The speaker registration information includes offline speaker voiceprint registration information and / or dynamic speaker tagging information.

10. The speaker separation method according to claim 8, characterized in that, The step of performing speaker separation on the real-time speech stream based on speaker registration information and the second voiceprint recognition result further includes: The speaker separation result is stored in a temporary voiceprint database in correspondence with the second voiceprint recognition result. If the real-time voice stream is interrupted, the interrupted voice is received, and based on the target personalized voiceprint recognition model, voiceprint recognition is performed on the interrupted voice to obtain a third voiceprint recognition result. Based on the third voiceprint recognition result, the temporary voiceprint database is used to perform speaker separation on the interrupted speech.

11. A speaker separation integrated machine, characterized in that, Includes a memory, a processor, a target personalized voiceprint recognition model stored in the memory, and a computer program stored in the processor and capable of running on the processor; When the processor executes the computer program, it invokes the target personalized voiceprint recognition model to implement the speaker separation method as described in any one of claims 8-10.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the offline personalized voiceprint learning method as described in any one of claims 1-7, or the speaker separation method as described in any one of claims 8-10.

Citation Information

Patent Citations

  • Method and system for training voiceprint recognition model

    CN107610709A

  • Methods and systems for updating risk control model

    CN110310206A