Speaker recognition method and device based on machine learning model, equipment and medium
Through adversarial training of machine learning models, device-related information is stripped away, the accuracy of speaker recognition is improved, the problem of insufficient recognition accuracy in traditional technologies is solved, and high-precision recognition is achieved in scenarios such as live broadcasts or online meetings.
Patent Information
- Application Number
- CN202511254947.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional speaker recognition technology lacks accuracy in complex scenarios such as live broadcasts or online meetings, especially due to insufficient generalization capabilities caused by the inconsistent distribution of training data and actual input data.
A machine learning model with adversarial training is used to extract speech features through the first machine learning model, and the second machine learning model is used to classify the collection devices. The training goal is to reduce the accuracy of the collection device classification, remove device-related information, and improve the model's generalization ability.
The accuracy of speaker recognition has been improved, especially in real-time interactive scenarios such as live broadcasts or online meetings, improving the generalization ability of the model and ensuring high-precision recognition.
Smart Images

Figure CN120808792A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a speaker recognition method, apparatus, device, and storage medium based on a machine learning model. Background Art
[0002] Speaker recognition is a technology that automatically identifies speakers by analyzing speech features in audio signals. It enables authentication, voice retrieval, and automatic generation of meeting transcripts. However, in complex scenarios like live broadcasts and online meetings, traditional speaker recognition technology often lacks the accuracy to meet practical requirements. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for speaker recognition based on a machine learning model is provided. The method includes: extracting speech features from audio data using a trained first machine learning model, the speech features being used to characterize speaker identity information; and assigning a speaker identifier to the audio data based on the speech features, wherein the first machine learning model is trained using at least a second machine learning model, the second machine learning model being configured to: classify the acquisition device of the sample audio data based on the sample speech features extracted from the sample audio data by the first machine learning model, the classification result indicating whether the acquisition device is a shared device not associated with the speaker, the training objective of the first machine learning model including reducing the accuracy of the second machine learning model's classification of the acquisition device, and the sample audio data being determined by at least one of the following: in response to the acquired first audio data including the speech of a first sample speaker and not including the speech of other speakers different from the first sample speaker, determining the first audio data as sample audio data, the first audio data being collected by a non-shared device associated with the first sample speaker, or determining the acquired second audio data as sample audio data, the second audio data being collected by a shared device not associated with the second sample speaker.
[0004] In a second aspect of the disclosure, a speaker recognition apparatus based on a machine learning model is provided. The apparatus comprises: an extraction module configured to extract speech features from audio data using a trained first machine learning model, the speech features being used to represent identity information of a speaker; and an assignment module configured to assign a speaker identification to the audio data based on the speech features, and wherein the first machine learning model is trained using at least a second machine learning model configured to: classify a capture device of sample audio data based on sample speech features extracted from the sample audio data by the first machine learning model, a result of the classification indicating whether the capture device is a shared device not associated with the speaker, a training objective of the first machine learning model comprises reducing an accuracy of the second machine learning model in classifying the capture device, and the sample audio data is determined by at least one of: determining first audio data as the sample audio data in response to the first audio data comprising speech of a first sample speaker and not comprising speech of other speakers different from the first sample speaker, the first audio data being captured by a non-shared device associated with the first sample speaker, or determining second audio data as the sample audio data, the second audio data being captured by a shared device not associated with a second sample speaker.
[0005] In a third aspect of the disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product comprises computer-executable instructions that, when executed by a processor, implement the method according to the first aspect of the disclosure.
[0008] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiment illustrated. Other objectives, features and aspects of the disclosure will become apparent from the following description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, aspects and advantages of embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters designate like elements in which: Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented; Figure 2 A flowchart illustrating an example process for speaker recognition according to some embodiments of the present disclosure is shown; Figure 3 A schematic diagram illustrating an example of model training according to some embodiments of the present disclosure is shown; Figure 4 A schematic structural block diagram of a speaker recognition device according to some embodiments of the present disclosure is shown; and Figure 5 A block diagram of an electronic device is shown in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0010] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0011] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0012] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0013] The data of the user, the acquisition and / or use of the data, etc. can be involved in the embodiments of the present disclosure. These aspects comply with the corresponding laws and regulations and relevant provisions. In the embodiments of the present disclosure, the collection, acquisition, processing, processing, forwarding, use, etc. of all data are performed on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the type of data or information that can be involved, the use range, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations. The specific notification and / or authorization manner can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this aspect.
[0014] In the present specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality basis (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. The user refuses to process personal information other than the necessary information required for the basic function, which does not affect the user's use of the basic function.
[0015] Various example implementations of the scheme will be described in detail below in combination with the drawings. It should be understood that the structure and function of each element in the environment 100 are described below only for exemplary purposes, without implying any limitation on the scope of the present disclosure.
[0016] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed in a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110. In some embodiments, the application 120 can be an application providing a single function, also known as a single-function application. In some embodiments, the application 120 can also be an integrated suite application capable of providing multiple functions to the user 140. To this end, the suite application can include multiple components respectively for implementing these functions, in which case each component can be considered to correspond to a sub-application. The suite application in the office field is sometimes also referred to as an "office application", a "collaborative office platform", etc. As an example, the integrated components in the suite application can include, but are not limited to, one or more of the following: a chat component (also known as an instant messaging (IM) component), a document component, an audio / video conference component, a mail component, a calendar component, a schedule component, a task component, etc. In some embodiments, the suite application can also be downloaded and installed in the terminal device 110 as an application.
[0017] In Figure 1In the environment 100, if the application 120 is started, the terminal device 110 may present a page 150 of the application 120 to the user 140. The page 150 may include entry controls and access portals for various components provided by the application 120, as well as interactive interfaces associated with the components, such as a conversation interface presenting chat content, an online meeting interface, a file sharing interface, and the like.
[0018] although Figure 1 While one user 140 is shown, multiple users 140 and corresponding terminal devices 110 may exist. These users 140 can interact with the application 120 via their respective terminal devices 110. The application 120 can provide these users 140 with real-time interactive scenarios, such as online meetings and live broadcasts. In such real-time interactive scenarios, one or more speakers can interact through their respective speech content.
[0019] The speaker in the real-time interactive scenario described herein may refer to the entity that utters the voice in the real-time interactive scenario. For example, if the real-time interactive scenario is an online meeting, the speaker may be a participant in the meeting. In some embodiments, a speaker may include one or more speakers, which may be specifically determined based on the terminals participating in the online meeting. For example, multiple speakers may access the online meeting through the same terminal (or using the same account). In this case, multiple speakers may be identified as the same speaker. Alternatively or additionally, in some embodiments, a speaker may correspond to one speaker.
[0020] Continue to refer Figure 1 In some embodiments, the terminal device 110 can communicate with the electronic device 130 to provide services for the application 120. For example, the terminal device 110 can send audio data collected in a real-time interactive scene to the electronic device 130. The electronic device 130 can use speaker recognition technology to identify and annotate the speaker in the audio data. Next, the electronic device 130 can return the annotation results to the terminal device 110. The terminal device 110 can then present the annotation results for reference by the user 140 using the page 150.
[0021] The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a smartbook, a media tablet, a personal digital assistant (PDA), a personal navigation device, a personal digital assistant (PDA), a digital audio player, a digital audio player, a digital camera / camcorder, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 can also be capable of supporting any type of interface to a user (such as "wearable" circuitry, etc.).
[0022] The electronic device 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and big data and artificial intelligence platform, etc. The electronic device 130 can include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0023] As briefly described above, in a complex scenario such as a live broadcast or an online conference (hereinafter also collectively referred to as a real-time interaction scenario), the recognition accuracy of traditional speaker recognition technology often fails to meet the actual demand. For example, in a speaker recognition scheme, a machine learning model can be used to process audio data collected in a real-time interaction scenario, so as to recognize the speaker in the audio data. Generally, in a real-time interaction scenario, audio data can be collected by a shared device (for example, a conference room device) or a non-shared device (for example, a personal device). However, in the training process of the machine learning model, due to the scarcity and difficulty of obtaining sample audio data from shared devices, sample audio data from non-shared devices is often used as training data. However, due to the influence of hardware parameters and the environment in which the device is located, the audio data collected by the shared device and the non-shared device for the same speech of the same user can be quite different. In this case, the training data of the machine learning model and the input data received by the model in actual application are inconsistent in data distribution. This inconsistency can result in insufficient generalization ability of the trained machine learning model in actual application, thereby affecting the speaker recognition accuracy of the machine learning model in real scenarios.
[0024] In view of this, embodiments of the present disclosure provide a speaker recognition scheme based on a machine learning model. According to the scheme, first, speech features are extracted from audio data by using a trained first machine learning model. The speech features are used to represent the identity information of a speaker. Then, a speaker identifier is assigned to the audio data based on the speech features. The first machine learning model is trained by using at least a second machine learning model, which is configured to classify a collection device of sample audio data based on sample speech features extracted from the sample audio data by the first machine learning model, and the result of the classification indicates whether the collection device is a shared device that is not associated with the speaker. The training target of the first machine learning model includes reducing the accuracy of the second machine learning model in classifying the collection device. The sample audio data is determined by using first audio data and / or second audio data. The first audio data is collected by a non-shared device associated with a first sample speaker. In embodiments of the present disclosure, after obtaining the first audio data, if the first audio data includes speech of the first sample speaker and does not include speech of other speakers, the first audio data is determined as the sample audio data. The second audio data is collected by a shared device that is not associated with a second sample speaker. In embodiments of the present disclosure, after obtaining the second audio data, the second audio data is directly determined as the sample audio data.
[0025] It will be more clearly understood through the following description that embodiments of the present disclosure use a machine learning model (e.g., the first machine learning model) to extract speech features related to a speaker from audio data and perform recognition of the speaker based on the speech features. In particular, embodiments of the present disclosure introduce a second machine learning model to form an adversarial relationship with the first machine learning model during the training process of the first machine learning model. Such adversarial training aims to gradually strip the speech features extracted by the first machine learning model of information related to the collection device (e.g., a shared device or a non-shared device), thereby obtaining independence from the collection device. In this way, the problem of insufficient model generalization ability due to inconsistency between training data and actual input data can be improved. In addition, considering that the sources of training data are different (e.g., from a shared device or a non-shared device), the focus points in the training data are also different. Embodiments of the present disclosure distinguish between training data from shared devices and non-shared devices, which helps the machine learning model (e.g., the first machine learning model) to better learn effective knowledge, thereby accelerating the convergence of the training process.
[0026] Based on the first machine learning model trained in the above manner, embodiments of the present disclosure can achieve high-precision recognition of speakers in real-time interactive scenarios (e.g., live streaming or online meetings, etc.), thereby providing effective data support for subsequent processing (e.g., summary generation and / or meeting minutes compilation, etc.).
[0027] Figure 2 A flowchart illustrating an example process 200 of speaker recognition is shown, according to some embodiments of the present disclosure, Figure 3 A schematic diagram illustrating an example 300 of model training is shown, according to some embodiments of the present disclosure. The following description is made in conjunction with Figure 1 and Figure 3 The process 200 is described. The process 200 can be implemented at the electronic device 130. At block 210, the electronic device 130 extracts speech features from audio data using a trained machine learning model (e.g., the first machine learning model 301), the speech features being used to characterize identity information of a speaker.
[0028] The audio data can be real-time data collected by an audio collection device (e.g., a shared device or a non-shared device). In embodiments of the present disclosure, a shared device can refer to a device that is not associated with any individual (e.g., a speaker) and is commonly used within a certain range. A conference room device (e.g., an audio collection device in a conference room) is an example of a shared device. A non-shared device can refer to a device that is associated with a specific individual (e.g., a speaker). A personal terminal device (e.g., a portable computer, a tablet, a mobile phone) of an individual is an example of a non-shared device.
[0029] The audio data can contain one or more raw waveform signals of speaker speech, or digital signals obtained after pre-processing (such as framing, windowing, Fourier transform, Mel frequency cepstral coefficient extraction, Mel spectrogram generation, etc.) of the raw waveform signals. In some embodiments, the audio data can be real-time data collected in a real-time interactive scenario, and the real-time interactive scenario can be participated by multiple speakers. In some embodiments, the real-time interactive scenario includes but is not limited to: an online meeting (e.g., an online audio conference or an online video conference, etc.), a network live broadcast, a telephone multi-party call, and / or a chat room, etc. In some embodiments, the audio data can be collected by a non-shared device (e.g., a terminal device of the speaker, etc.) associated with the speaker. Alternatively, in some embodiments, the audio data can be collected by a shared device (e.g., a conference room device) not associated with the speaker. In some embodiments, the audio data can be an audio segment intercepted from a real-time data stream based on a predetermined time length or other appropriate rules. Such an audio segment can be regarded as a minimum processing unit of speaker recognition.
[0030] Alternatively or additionally, in some embodiments, the audio data can be collected in a non-real-time interactive scenario. For example, the audio data can be extracted from a pre-recorded audio or video, etc. This can be determined according to actual needs, and embodiments of the present disclosure do not limit this.
[0031] The speaker can refer to a speaking entity. In some application scenarios, the speaker can be a participant in a live or online conference. The identity information of the speaker can refer to a unique identity of the speaker, which is used to distinguish different speakers. The speech feature is used to represent the identity information of the speaker, and specifically can be an acoustic attribute feature (for example, a voiceprint feature) capable of representing individual differences of the speaker. The speech feature can be encoded as a high-dimensional vector to serve as a model input of the first machine learning model 301. The first machine learning model 301 can be any model capable of realizing speech feature extraction. The first machine learning model 301 can be a model having a multi-layer neural network 3011. For example, the first machine learning model 301 includes but is not limited to a deep neural network model, a convolutional neural network model, a recurrent neural network model, or a Transformer model, etc.
[0032] The first machine learning model 301 can be trained by using a second machine learning model 302. The second machine learning model 302 is configured to classify a collection device of the sample audio data 304 based on the speech feature (hereinafter, the speech feature extracted by the first machine learning model 301 in the training process is also referred to as a sample speech feature 305) extracted from the audio data (hereinafter, the audio data provided to the first machine learning model 301 in the training process is also referred to as sample audio data 304) by the first machine learning model 301. The result 306 of the classification indicates whether the collection device is a shared device (for example, a conference room device) not associated with the speaker. The training target of the first machine learning model 301 includes reducing the accuracy of the second machine learning model in classifying the collection device, so that the second machine learning model 302 gradually fails to correctly classify the collection device.
[0033] The second machine learning model 302 can be implemented based on any classifier capable of performing a binary classification task. The second machine learning model 302 can be a classifier with a multi-layer neural network (e.g., a fully connected network) 3021. For example, the second machine learning model 302 includes, but is not limited to, a logistic regression based classifier, a support vector machine based classifier, a decision tree based classifier, and / or a neural network based classifier, etc. In some embodiments, the second machine learning model 302 can classify the sample audio data 304 based on information in the sample speech features 305 related to the capturing device and / or the ambient sound, etc. For example, after analyzing the sample speech features 305, if the second machine learning model 302 determines that the sample audio data 304 is from a conference room device such as a microphone of a conference room, the second machine learning model 302 can classify the sample audio data 304 as “captured by a shared device not associated with the speaker”; otherwise, the second machine learning model 302 can classify the sample audio data 304 as “captured by a non-shared device associated with the speaker”. For another example, after analyzing the sample speech features 305, if the second machine learning model 302 determines that the ambient sound in the sample audio data 304 is similar to the ambient sound of a conference room, the second machine learning model 302 can classify the sample audio data 304 as “captured by a shared device not associated with the speaker”; otherwise, the second machine learning model 302 can classify the sample audio data 304 as “captured by a non-shared device associated with the speaker”. And so on.
[0034] During the training of the first machine learning model 301, the training objective is to make the second machine learning model 302 unable to effectively perform the classification task described above. For example, for a binary classification scenario, the training objective aims to make the prediction confidence of the second machine learning model 302 for each class close to 50%, so that its discrimination performance approaches the level of random guessing. Thus, an adversarial relationship between the first machine learning model 301 and the second machine learning model 302 is formed. This adversarial mechanism encourages the first machine learning model 301 to tend to discard or conceal information related to the capturing device (e.g., a shared device or a non-shared device) when extracting the sample speech features 305, thereby achieving decoupling of the sample speech features 305 and the capturing device information.
[0035] In some embodiments, the training objective can be implemented with a predetermined neural network layer 307 connected between the first machine learning model 301 and the second machine learning model 302. The predetermined neural network layer 307 can be configured to change the gradient of the training loss propagated from the second machine learning model 302 to the first machine learning model 301 during the training process (e.g., the backpropagation phase) of the first machine learning model 301. In some embodiments, the predetermined neural network layer 307 can include a Gradient Reversal Layer (GRL). Alternatively, in some embodiments, the predetermined neural network layer 307 can be any other kind of neural network layer capable of changing the gradient so that the learning objective of one first machine learning model 301 is in opposition to the objective of the second machine learning model 302. In this way, the intrusion to the existing training architecture can be minimized, thus reducing the cost of modification and improving the optimization efficiency. It is noted that in addition to the predetermined neural network layer 307, in some embodiments, the training objective can be further implemented with one or more other neural network layers, which are not limited by embodiments of the present disclosure.
[0036] In some embodiments, the training process can be implemented based on a Domain-Adversarial Neural Networks (DANN) architecture. In this case, the predetermined neural network layer 307 can include a Gradient Reversal Layer. For the training loss propagated from the second machine learning model 302 to the first machine learning model 301, the Gradient Reversal Layer can be configured to reverse the gradient of the training loss.
[0037] In some embodiments, in the forward propagation phase of the training process, the Gradient Reversal Layer can directly pass the sample speech feature 305 from the first machine learning model 301 to the second machine learning model 302. The second machine learning model 302 can perform the binary classification task described above using these features. In the backpropagation phase of the training process, for the training loss propagated from the second machine learning model 302 to the first machine learning model 301, the Gradient Reversal Layer can multiply the corresponding gradient by a negative constant (e.g., -λ, where λ is an adjustable gradient reversal coefficient or learning rate coefficient), thus achieving the reversal of the gradient. Through this adversarial training, the first machine learning model 301 can effectively learn the knowledge to decouple the sample speech feature 305 from the collection device information, thus achieving the fast convergence of the model.
[0038] In some embodiments, during the training of the first machine learning model 301, the second machine learning model 302 can be synchronously trained. In this case, for the second machine learning model 302, the model parameters can be subjected to gradient descent in a regular manner. While for the first machine learning model 301, the model parameters of the first machine learning model 301 can be adjusted in a direction opposite to the gradient descent by means of the gradient reversal layer, thereby achieving the adversarial training of the first machine learning model 301 and the second machine learning model 302.
[0039] In some embodiments, the first machine learning model 301 can be trained with the second machine learning model 302 and a third machine learning model 303. In other words, the second machine learning model 302 and the third machine learning model 303 share the same speech feature extraction model (i.e., the first machine learning model 301). Unlike the second machine learning model 302, the third machine learning model 303 can be configured to identify the speaker of the sample audio data 304 based on the sample speech feature 305. The training objective of the first machine learning model 301 further includes enabling the third machine learning model 303 to correctly identify the speaker in the sample audio data 304.
[0040] The third machine learning model 303 can be implemented based on any kind of classifier capable of performing a multi-classification task. The third machine learning model 303 can be a classifier with a multi-layer neural network (e.g., a fully connected network) 3031. For example, the third machine learning model 303 includes but is not limited to a neural network-based or support vector machine-based classifier, etc. The third machine learning model 303 can output label information 308 capable of indicating the speaker of the sample audio data 304 based on the sample speech feature 305.
[0041] In this embodiment, the training process of the first machine learning model 301 is subject to double-objective constraints based on the second machine learning model 302 and the third machine learning model 303. In this way, the first machine learning model 301 can learn how to focus on extracting the sample speech feature 305 of the speaker identity representation. Such sample speech feature 305 strips the “noise” information (which is against the objective of the second machine learning model 302) related to the collection device and / or the environment, while retaining and strengthening the speech features that distinguish different individuals (which is in line with the objective of the third machine learning model 303). This is conducive to improving the generalization ability of the model while accurately identifying the speaker. In some embodiments, during the training of the first machine learning model 301, the second machine learning model 302 and the third machine learning model 303 can be synchronously trained, thereby improving the training efficiency.
[0042] In some embodiments, the training process of the first machine learning model 301 can be implemented at the electronic device 130 or any other suitable device, without limitation to the embodiments of the present disclosure. For convenience of discussion, the processing of the training data (e.g., the sample audio data 304) in the training process is described below with the example that the training process of the first machine learning model 301 is implemented at the electronic device 130.
[0043] The sample audio data 304 can be determined based on first audio data. The first audio data can be, for example, audio data collected by a non-shared device associated with a sample speaker. The sample speaker can refer to the real speaker of the sample audio data 304. For the sake of distinction below, the sample speaker here is also referred to as a first sample speaker. In this case, the electronic device 130 can obtain the first audio data. If the first audio data includes the speech of the first sample speaker and does not include the speech of other speakers, the electronic device 130 can determine the first audio data as the sample audio data 304, thereby used for training the first machine learning model 301. If the first audio data includes the speech of the first sample speaker and the speech of other speakers different from the first sample speaker (e.g., there are other people talking in the background and / or multiple people speaking at the same time, etc.), the electronic device 130 can ignore the first audio data, thereby excluding the first audio data from the training data of the first machine learning model 301. In this way, it can be avoided that the multi-person speech causes the model to be confused, thereby preventing the machine learning model from learning the wrong knowledge. Through this screening mechanism, the electronic device 130 can automatically exclude samples containing interfering speech, ensuring that the sample audio data 304 used for model training reflects the acoustic characteristics of the sample speaker as much as possible.
[0044] In some embodiments, for the sample audio data 304 determined based on the first audio data, the sample audio data 304 can be assigned with a first label and a second label. The first label can include identification information of the first sample speaker, and the second label can indicate that the sample audio data 304 is collected by a non-shared device associated with the first sample speaker. As mentioned above, in the training process of the first machine learning model 301, the second machine learning model 302 and the third machine learning model 303 can be synchronously trained. In this case, the first label can be used to indicate the first machine learning model 301 and the third machine learning model 303: “which specific speaker does the sample audio data 304 belong to”, so as to guide the respective iteration direction of the first machine learning model 301 and the third machine learning model 303 in the training process. The second label can be used to indicate the first machine learning model 301 and the second machine learning model 302: “the sample audio data 304 is collected by a non-shared device (for example, a terminal device of an individual)”, so as to guide the respective iteration direction of the first machine learning model 301 and the second machine learning model 302 in the training process.
[0045] Alternatively or additionally, in addition to the first audio data, the sample audio data 304 can also be determined based on second audio data. The second audio data can be audio data collected by a shared device not associated with the sample speaker (for example, a second sample speaker). The second sample speaker can be the same sample speaker or a different sample speaker as the first sample speaker, which can be determined according to actual needs, and embodiments of the present disclosure do not limit this. In this case, the electronic device 130 can obtain the second audio data. Then, the electronic device 130 can directly determine the second audio data as the sample audio data 304, so as to be used for training the first machine learning model 301. By directly using the audio data from the shared device as the training data of the first machine learning model 301, the deficiency of the audio data collected by the non-shared device can be made up, so as to improve the training effect of the first machine learning model 301.
[0046] In some embodiments, for the sample audio data 304 determined based on the second audio data, the sample audio data 304 can be assigned a third label, which can indicate that the sample audio data 304 is collected by a shared device not associated with the second sample speaker. As mentioned above, the second machine learning model 302 and the first machine learning model 301 can be synchronously trained in the training process of the first machine learning model 301. In this case, the third label can be used to indicate the first machine learning model 301 and the second machine learning model 302 that “the sample audio data 304 is collected by a shared device (e.g., a conference room device)”, thereby guiding the respective iteration direction of the first machine learning model 301 and the second machine learning model 302 in the training process. In some embodiments, for the sample audio data 304 determined based on the second audio data, the first label can be not assigned or assigned in a small amount according to actual needs, thereby balancing the training effect and processing efficiency.
[0047] Through the above training process, the embodiments of the present disclosure make the speech features extracted by the first machine learning model 301 strip the information related to the collection device on the basis of focusing on speaker recognition. Thus, the generalization ability of the first machine learning model 301 in the actual application process can be improved, and in particular, the speaker recognition accuracy in real-time interactive scenarios such as live broadcast or online conference can be improved.
[0048] Referring back to Figure 2 After the trained first machine learning model 301 is used to extract speech features from the audio data, in block 220, the electronic device 130 assigns a speaker identification to the audio data based on the extracted speech features.
[0049] The speaker identification can indicate the identity information of the speaker of the speech data in the real-time interactive scenario. For example, the speaker identification includes but is not limited to the user identifier (ID) of the speaker, the user name, or the anonymous code (e.g., speaker A or speaker B), etc. The manner in which the electronic device 130 determines the speaker identification can be determined according to actual needs, and the embodiments of the present disclosure do not limit this. For example, the electronic device 130 can pre-store the speech feature templates of all known speakers. On this basis, the electronic device 130 can determine the speaker through a matching mechanism and assign the corresponding speaker identification. For another example, in the case of an unknown speaker, the electronic device 130 can perform cluster analysis on the speech features extracted by the first machine learning model 301, thereby classifying the speech data with similar features into the same class. In this case, the speech data classified into the same class will be identified as the same speaker and assigned a corresponding code as the speaker identification, etc.
[0050] The embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process.Figure 4 A schematic structural block diagram of a speaker recognition apparatus 400 is shown in accordance with some embodiments of the present disclosure. The apparatus 400 can be implemented as or included in the electronic device 130. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.
[0051] Referring to Figure 4 The apparatus 400 includes an extraction module 410 and an assignment module 420. The extraction module 410 is configured to extract, from audio data, speech features for characterizing identity information of a speaker, by utilizing a trained first machine learning model. The assignment module 420 is configured to assign, for the audio data, a speaker identification based on the speech features. The first machine learning model is trained by utilizing at least a second machine learning model configured to classify, based on sample speech features extracted from sample audio data by the first machine learning model, a capturing device of the sample audio data, a result of the classification indicating whether the capturing device is a shared device not associated with the speaker. A training objective of the first machine learning model includes reducing a correctness of the classification of the capturing device by the second machine learning model. The sample audio data is determined by at least one of: determining, as the sample audio data, first audio data acquired in response to the first audio data including speech of a first sample speaker and not including speech of other speakers different from the first sample speaker, the first audio data being captured by a non-shared device associated with the first sample speaker, or determining, as the sample audio data, second audio data acquired by a shared device not associated with a second sample speaker.
[0052] In some embodiments, the training objective is implemented at least by utilizing a predetermined neural network layer connected between the first machine learning model and the second machine learning model, and the predetermined neural network layer is configured to change, during a training process of the first machine learning model, a gradient of a training loss propagated from the second machine learning model to the first machine learning model.
[0053] In some embodiments, the training process is implemented based on a domain adversarial neural network architecture, and the predetermined neural network layer includes a gradient reversal layer configured to reverse the gradient of the training loss.
[0054] In some embodiments, for the sample audio data determined based on the first audio data, the sample audio data is assigned with a first label including identification information of a first sample speaker and a second label indicating that the sample audio data is captured by a non-shared device associated with the first sample speaker.
[0055] In some embodiments, for the sample audio data determined based on the second audio data, the sample audio data is assigned a third label, the third label indicating that the sample audio data is collected by a shared device not associated with the second sample speaker.
[0056] In some embodiments, the first machine learning model is trained with the second machine learning model and a third machine learning model, the third machine learning model being configured to identify a speaker of the sample audio data based on the sample speech features, and the training target of the first machine learning model further includes making the third machine learning model correctly identify the speaker in the sample audio data.
[0057] In some embodiments, the audio data is collected in a real-time interactive scenario, and the real-time interactive scenario is participated by multiple speakers.
[0058] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented is shown. The electronic device 500 may, for example, be used to implement the electronic device 130 as shown in Figure 1 or the apparatus 400 as shown in Figure 4 It should be understood that the electronic device 500 shown is merely an example and should not be taken as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown is merely an example and should not be taken as limiting the functionality and scope of the embodiments described herein.
[0059] Referring to Figure 5 , the electronic device 500 is in the form of a general electronic device. The components of the electronic device 500 can include, but are not limited to, one or more processors 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 can be an actual or virtual processor and can perform various processing according to programs stored in the memory 520. In a multi-processor system, multiple processors perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
[0060] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media that is located either internally or externally to the electronic device 500, including, but not limited to, memory, removable storage, and non-removable storage. The memory 520 can be volatile (such as register, cache, RAM), non-volatile (such as ROM, EEPROM, flash memory), or some combination of the two. The storage 530 can be removable or non-removable and can include, but is not limited to, magnetic disks, optical disks, or tape. It will be appreciated that the memory 520 and / or the storage 530 can be set up in a raw, formatted, or used condition.
[0061] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown, a magnetic disk drive and / or an optical disk drive can be provided in the electronic device 500 for reading from or writing to a removable, non-removable, volatile, or non-volatile media such as a floppy disk, a magnetic tape, or an optical disk, for example. In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure. Figure 5
[0062] The communication unit 540 enables communications with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.
[0063] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 540, as needed, one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0064] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0065] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0066] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0067] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0068] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.
[0069] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt to a particular situation and the teachings of the present disclosure to a specific implementation, without departing from the central novel teachings of the application. The implementation(s) illustrated and described herein are meant only to serve as examples. Departures in form and detail are within the scope of the disclosure. Therefore, one skilled in the art can restructure the implementation(s) as needed, while still adhering to the principles of the present disclosure.
Claims
1. A speaker recognition method based on a machine learning model, characterized in that: include: Extracting speech features from the audio data using the trained first machine learning model, where the speech features are used to represent identity information of a speaker; as well as assigning a speaker identifier to the audio data based on the speech features, and The first machine learning model is trained using at least a second machine learning model, and the second machine learning model is configured to: classify a device that collects the sample audio data based on sample speech features extracted from the sample audio data by the first machine learning model, wherein a result of the classification indicates whether the collection device is a shared device that is not associated with the speaker, The training objective of the first machine learning model includes reducing the accuracy of the second machine learning model in classifying the acquisition device, and The sample audio data is determined by at least one of the following: In response to the acquired first audio data including the speech of a first sample speaker and not including the speech of other speakers different from the first sample speaker, the first audio data is determined as the sample audio data, the first audio data being collected by a non-shared device associated with the first sample speaker or The acquired second audio data is determined as the sample audio data, where the second audio data is collected by a shared device that is not associated with the second sample speaker.
2. The method according to claim 1, characterized in that The training objective is achieved by at least a predetermined neural network layer connected between the first machine learning model and the second machine learning model, and the predetermined neural network layer is configured to: during the training process of the first machine learning model, change the gradient of the training loss propagated from the second machine learning model to the first machine learning model.
3. The method according to claim 2, characterized in that The training process is implemented based on a domain adversarial neural network architecture, and the predetermined neural network layer includes a gradient reversal layer, which is configured to reverse the gradient of the training loss.
4. The method according to claim 1, wherein For the sample audio data determined based on the first audio data, the sample audio data is assigned a first label and a second label, the first label includes identification information of the first sample speaker, and the second label indicates that the sample audio data is collected by a non-shared device associated with the first sample speaker.
5. The method according to claim 1, wherein For the sample audio data determined based on the second audio data, a third label is assigned to the sample audio data, where the third label indicates that the sample audio data is collected by a shared device that is not associated with the second sample speaker.
6. The method according to claim 1, characterized in that The first machine learning model is trained using the second machine learning model and a third machine learning model, wherein the third machine learning model is configured to: identify a speaker of the sample audio data based on the sample speech features, and The training objective of the first machine learning model also includes enabling the third machine learning model to correctly identify the speaker in the sample audio data.
7. The method according to claim 1, characterized in that The audio data is collected in a real-time interactive scene, and the real-time interactive scene is participated by multiple speakers.
8. An audio recognition device, characterized in that: include: an extraction module configured to extract speech features from the audio data using the trained first machine learning model, wherein the speech features are used to represent the identity information of the speaker; as well as an assignment module configured to assign a speaker identifier to the audio data based on the speech feature, and The first machine learning model is trained using at least a second machine learning model, and the second machine learning model is configured to: classify the acquisition device of the sample audio data based on the sample speech features extracted from the sample audio data by the first machine learning model, and the classification result indicates whether the acquisition device is a shared device not associated with the speaker, and The training objective of the first machine learning model includes reducing the accuracy of the second machine learning model in classifying the acquisition device, and The sample audio data is obtained by at least one of the following: In response to the acquired first audio data including the speech of a first sample speaker and not including the speech of other speakers different from the first sample speaker, the first audio data is determined as the sample audio data, the first audio data being collected by a non-shared device associated with the first sample speaker or The acquired second audio data is determined as the sample audio data, where the second audio data is collected by a shared device that is not associated with the second sample speaker.
9. An electronic device, characterized in that: include: at least one processor; as well as At least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 7 when executed by the at least one processor.
10. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: The computer executable instructions are executable by a processor to implement the method according to any one of claims 1 to 7.
11. A computer program product comprising computer executable instructions, characterized in that: The computer executable instructions implement the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Far-field speaker authentication method and system based on gradient inversion layer
CN113241081A
Cross-speaker voice style modeling method and computer readable storage medium
CN114242031A
Cross-channel content-independent speaker identification method and system based on adversarial learning
CN114974260A
Overlapped voice detection method and device, electronic equipment and storage medium
CN117174111A
Adaptive environment audio classification method and system for different devices or places
CN117831516A