Real-time speaker distinguishing method, apparatus, electronic device and storage medium

WO2026199246A1PCT designated stage Publication Date: 2026-10-01GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/085105
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2026-10-01

Smart Images

  • Figure CN2025085105_01102026_PF_FP_ABST
    Figure CN2025085105_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A real-time speaker distinguishing method, an apparatus, an electronic device and a storage medium, relating to the technical field of voice processing. The method comprises: acquiring voice data of a speaker and an initial voiceprint feature set (S410); on the basis of delay information of the voice data, determining an orientation of the voice data (S420); on the basis of the orientation of the voice data, obtaining a stability evaluation value of the voice data (S430); on the basis of a cosine similarity between a voiceprint feature of the voice data and the initial voiceprint feature set, performing voiceprint recognition on the voice data, so as to output a target identifier (S440); and, when the stability evaluation value is greater than or equal to a preset threshold, updating the initial voiceprint feature set on the basis of the target identifier and the voiceprint feature of the voice data, so as to obtain an updated initial voiceprint feature set (S450). The accuracy of voiceprint feature sets and the accuracy of speaker recognition results can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Real-time speaker differentiation methods, devices, electronic equipment, and storage media Technical Field

[0001] This application relates to the field of speech processing technology, and more particularly to a real-time speaker differentiation method, apparatus, electronic device, and storage medium in the field of speech processing technology. Background Technology

[0002] To facilitate the identification of the speaker in multi-person speaking scenarios such as meetings, teaching sessions, and court hearings, a neural network model (e.g., a speaker recognition model) can be used to extract voiceprint features from the collected voice data. The extracted voiceprint features are then compared with a pre-built voiceprint feature database (hereinafter referred to as the "voiceprint feature set") to obtain the speaker corresponding to the voice data.

[0003] However, environmental interference may occur during the construction of the voiceprint feature set, resulting in a large amount of interference information in the voiceprint feature set, which reduces the accuracy of the voiceprint feature set. As a result, the speaker corresponding to the speech data determined based on the voiceprint feature set may be incorrect, leading to poor accuracy of the recognition result.

[0004] Therefore, improving the accuracy of voiceprint feature sets is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a real-time speaker differentiation method, apparatus, electronic device, and storage medium, which can improve the accuracy of voiceprint feature sets and speaker recognition results.

[0006] In a first aspect, this application provides a real-time speaker differentiation method for use in electronic devices, the method comprising:

[0007] Obtain the speaker's voice data and initial voiceprint feature set;

[0008] Determine the location of the voice data based on the time delay information of the voice data;

[0009] Based on the location of the speech data, a stability assessment value for the speech data is obtained; the stability assessment value is negatively correlated with the amount of change in the location of the speech data.

[0010] Based on the cosine similarity between the voiceprint features of the speech data and the initial voiceprint feature set, voiceprint recognition is performed on the speech data, and a target identifier is output; whereby the target identifier is used to represent the speaker's identity information;

[0011] When the stability evaluation value is greater than or equal to the preset threshold, the initial voiceprint feature set is updated based on the voiceprint features of the target identifier and the speech data to obtain the updated initial voiceprint feature set.

[0012] In this embodiment, when acquiring the speaker's voice data and an initial voiceprint feature set, the location corresponding to the voice data can first be determined using the time delay information of the voice data; then, a stability evaluation value for the voice data can be obtained using the location of the voice data; and finally, voiceprint recognition can be performed on the voice data using the cosine similarity between the voiceprint features of the voice data and the initial voiceprint feature set to obtain the target identifier corresponding to the voice data. Furthermore, when the stability evaluation value is greater than or equal to a preset threshold, the initial voiceprint feature set can be updated using the target identifier and the voiceprint features of the voice data to obtain an updated initial voiceprint feature set. Since a stability evaluation value ≥ a preset threshold indicates that the change in the location of the voice data is small, the location of the voice data is relatively stable, and the voice data will not undergo abrupt changes; or the possibility of abrupt changes in the voice data is low. Therefore, when the stability evaluation value is greater than or equal to the preset threshold, the initial voiceprint feature set can be updated using the target identifier and voiceprint features of the voice data. This can make the updated initial voiceprint feature set more stable and avoid the problem of deviation in the updated initial voiceprint feature set caused by the stability evaluation value being too small, thereby improving the accuracy of the updated initial voiceprint feature set.

[0013] Furthermore, with the updated initial voiceprint feature set being more accurate, the speaker identified based on the updated initial voiceprint feature set can be more accurate, thus improving the accuracy of speaker identification results.

[0014] In conjunction with the first aspect, in some implementations of the first aspect, the initial voiceprint feature set is updated based on the voiceprint features of the target identifier and speech data to obtain an updated initial voiceprint feature set, including:

[0015] Determine whether a target identifier exists in the initial voiceprint feature set;

[0016] When a target identifier exists in the initial voiceprint feature set, the target sub-voiceprint feature set corresponding to the target identifier is determined in the initial voiceprint feature set; based on the number of voiceprint features in the target sub-voiceprint feature set and the voiceprint features of the speech data, the initial voiceprint feature set is updated to obtain the updated initial voiceprint feature set.

[0017] When there is no target identifier in the initial voiceprint feature set, a new sub-voiceprint feature set is registered based on the voiceprint features of the speech data; the initial voiceprint feature set is updated based on the new sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0018] In this embodiment, when a target identifier exists in the initial voiceprint feature set, the initial voiceprint feature set can be updated using the number of voiceprint features in the sub-voiceprint feature set (i.e., the target sub-voiceprint feature set) corresponding to the target identifier in the initial voiceprint feature set, along with the voiceprint features of the speech data, to obtain an updated initial voiceprint feature set. Alternatively, when a target identifier does not exist in the initial voiceprint feature set, a new sub-voiceprint feature set can be registered using the voiceprint features of the speech data, and the initial voiceprint feature set can be updated using the new sub-voiceprint feature set to obtain an updated initial voiceprint feature set. By updating the initial voiceprint feature set in two different ways, there is a corresponding update method for both the presence and absence of a target identifier, rather than using a uniformly prescribed method. This avoids the problem of deviations in the updated initial voiceprint feature set caused by using a uniformly prescribed method, thereby improving the accuracy of the updated initial voiceprint feature set.

[0019] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the initial voiceprint feature set is updated based on the number of voiceprint features in the target sub-voiceprint feature set and the voiceprint features of the speech data to obtain an updated initial voiceprint feature set, including:

[0020] When the number is less than the preset number, the voiceprint features of the speech data are added to the target sub-voiceprint feature set to obtain the updated target sub-voiceprint feature set; the initial voiceprint feature set is updated based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0021] When the quantity is greater than or equal to the preset quantity, the initial voiceprint feature set is updated based on the similarity value between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set and the first voiceprint feature, respectively, to obtain the updated initial voiceprint feature set.

[0022] Among them, the first voiceprint feature represents the feature mean of the voiceprint features in the target sub-voiceprint feature set.

[0023] In this embodiment, when the number of voiceprint features in the target sub-voiceprint feature set is small, the voiceprint features of the speech data can be added to the target sub-voiceprint feature set, making the updated target sub-voiceprint feature set richer in voiceprint features and having better generalization ability. Alternatively, when the number of voiceprint features in the target sub-voiceprint feature set is large, the initial voiceprint feature set can be updated based on the similarity between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set, and the average feature values ​​of the voiceprint features in the target sub-voiceprint feature set, respectively. This can also make the updated initial voiceprint feature set have better generalization ability.

[0024] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the similarity value between the voiceprint features of the above-mentioned speech data and the first voiceprint features is the first similarity value, and the similarity value between the voiceprint features of the above-mentioned target sub-voiceprint feature set and the first voiceprint features is the second similarity value; the initial voiceprint feature set is updated based on the similarity values ​​between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set, and the first voiceprint features, respectively, to obtain the updated initial voiceprint feature set, including:

[0025] Determine the second voiceprint feature corresponding to the maximum similarity value between the first and second similarity values;

[0026] The voiceprint features and the third voiceprint features of the speech data are determined as the voiceprint features of the updated target sub-voiceprint feature set; the initial voiceprint feature set is updated based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0027] Among them, the third voiceprint feature represents the voiceprint features in the target sub-voiceprint feature set other than the second voiceprint feature.

[0028] In this embodiment, when the maximum similarity between the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set and the mean similarity of the voiceprint features in the target sub-voiceprint feature set is determined, the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set other than the voiceprint features corresponding to the maximum similarity (i.e., the third voiceprint features) are determined as the updated voiceprint features of the target sub-voiceprint feature set. That is, the updated voiceprint features of the target sub-voiceprint feature set do not include the voiceprint features corresponding to the maximum similarity. This makes the voiceprint features of the updated target sub-voiceprint feature set more representative and can represent the voiceprint features of speakers in different environments, thereby giving the updated initial voiceprint feature set better generalization.

[0029] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the above-mentioned registration of a new sub-voiceprint feature set based on the voiceprint features of speech data includes:

[0030] Determine the target feature mean of the speakerprint features in the speech data;

[0031] Register a new sub-voiceprint feature set based on the mean of the target features.

[0032] In this embodiment, when the target identifier is absent from the initial voiceprint feature set, a new sub-voiceprint feature set can be registered using the mean of the target features of the voiceprint features in the speech data. This allows the updated initial voiceprint feature set, based on the new sub-voiceprint feature set, to include the target identifier of the speech data, supplementing the speaker identity information in the updated initial voiceprint feature set. Consequently, the updated initial voiceprint feature set exhibits better generalization ability. Furthermore, registering the new sub-voiceprint feature set using the mean of the target features of the voiceprint features in the speech data enhances the generalization ability of the new sub-voiceprint feature set, further improving the generalization ability of the updated initial voiceprint feature set.

[0033] In conjunction with the first aspect and the above implementation methods, in some implementation methods of the first aspect, the above-mentioned determination of the target feature mean of the voiceprint features of speech data includes:

[0034] Determine the target number of voiceprint features from the voiceprint features of the speech data; wherein the target number is less than or equal to the total number of voiceprint features in the speech data.

[0035] The mean value of the voiceprint features of the target data is determined as the mean value of the target features.

[0036] In this embodiment, a certain number of voiceprint features are selected from the voiceprint features of the speech data, and the mean value of these certain number of voiceprint features is determined as the target feature mean value. This reduces interference information in the voiceprint features of the speech data, thereby making the determined target feature mean value more accurate. Furthermore, based on the more accurate target feature mean value, the determined new sub-voiceprint feature set can be made more accurate, thus making the updated initial voiceprint features more accurate.

[0037] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the stability evaluation value of the speech data based on the location of the speech data includes:

[0038] The location of the speech data is differentially processed to obtain the differential location.

[0039] The differential orientation is smoothed to obtain the smoothed orientation.

[0040] Based on the smoothed orientation, a stability evaluation value is obtained.

[0041] In this embodiment, when the orientation of the speech data is differentially processed, the differences between the orientations of adjacent frames can be captured, reducing redundant information and making the differential orientation more accurate. Furthermore, when the more accurate differential orientation is smoothed, interference information in the speech data can be effectively suppressed, making the smoothed orientation more accurate. This results in a more accurate stability assessment value for the speech data obtained through the smoothed orientation. Consequently, based on a more accurate stability assessment value, the updated initial voiceprint features can be made more accurate.

[0042] Combining the first aspect and the above implementation methods, in some implementation methods of the first aspect, the above-mentioned voiceprint recognition based on the cosine similarity between the voiceprint features of the speech data and the initial voiceprint feature set, and the output of the target identifier, includes:

[0043] Determine the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set; wherein, the preset voiceprint feature set includes all sub-voiceprint feature sets in the initial voiceprint feature set;

[0044] Determine whether the maximum cosine similarity among multiple cosine similarities is greater than or equal to a preset similarity.

[0045] When the maximum cosine similarity is greater than or equal to the preset similarity, the identifier corresponding to the maximum cosine similarity is determined as the target identifier corresponding to the speech data, and the target identifier is output.

[0046] In this embodiment, when the maximum cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set is greater than or equal to the preset similarity, the similarity between the voiceprint features of the speech data and the speaker's identity information can be ensured to the greatest extent, while avoiding the problem of speaker identification errors caused by excessively low maximum cosine similarity, thereby improving the accuracy of speaker identification. Furthermore, since the voice data identifier is obtained by calculating the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set, the identification process is simpler and requires less computational resources compared to using clustering to obtain speaker identity information, thus exhibiting better universality.

[0047] Secondly, this application provides a real-time speaker differentiation device configured in an electronic device, the device comprising:

[0048] The acquisition module is used to acquire the speaker's voice data and initial voiceprint feature set;

[0049] The determination module is used to determine the location of voice data based on the time delay information of the voice data;

[0050] The processing module is used to obtain a stability evaluation value for the speech data based on its location; the stability evaluation value is negatively correlated with the change in the location of the speech data.

[0051] The recognition module is used to perform voiceprint recognition on the voice data based on the voiceprint features of the voice data and the cosine similarity with the initial voiceprint feature set, and output the target identifier; wherein, the target identifier is used to represent the speaker's identity information;

[0052] The update module is used to update the initial voiceprint feature set based on the voiceprint features of the target identifier and the speech data when the stability evaluation value is greater than or equal to a preset threshold, so as to obtain the updated initial voiceprint feature set.

[0053] Thirdly, this application provides an electronic device including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the methods described in the first aspect or any possible implementation thereof.

[0054] Fourthly, this application provides a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform the method described in the first aspect or any possible implementation thereof.

[0055] Fifthly, this application provides a computer-readable storage medium storing computer program code that, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof. Attached Figure Description

[0056] Figure 1 is a schematic diagram of a meeting discussion on related technologies.

[0057] Figure 2 is a schematic diagram of a speaker recognition model based on related technologies.

[0058] Figure 3 is a schematic diagram of another speaker recognition model provided in an embodiment of this application.

[0059] Figure 4 is a flowchart illustrating a real-time speaker differentiation method provided in an embodiment of this application.

[0060] Figure 5 is another flowchart illustrating a real-time speaker differentiation method provided in an embodiment of this application.

[0061] Figure 6 is a schematic diagram of the structure of the real-time speaker differentiation device provided in the embodiment of this application.

[0062] Figure 7 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Detailed Implementation

[0063] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0064] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0065] Figure 1 is a schematic diagram of a meeting discussion on related technologies.

[0066] For example, as shown in Figure 1, Figure 1 includes user A, user B, user C, voice data 110 corresponding to user A, voice data 120 corresponding to user B, voice data 130 corresponding to user C, environmental noise data (hereinafter referred to as "noise data") 140, microphone 150, and electronic device 160.

[0067] The microphone 150 can be used to collect voice data 110, voice data 120, voice data 130 and noise data 140, and transmit the collected voice data 110, voice data 120, voice data 130 and noise data 140 to the electronic device 160 through a wireless network or wired connection. When the electronic device 160 receives the voice data 110, voice data 120, voice data 130 and noise data 140, it can process the voice data 110, voice data 120, voice data 130 and noise data 140 through its own configured neural network model (e.g., speaker recognition model) to obtain the speaker corresponding to each of the voice data 110, voice data 120 and voice data 130.

[0068] Figure 2 is a schematic diagram of a speaker recognition model based on related technologies.

[0069] For example, as shown in Figure 2, the original speaker recognition model 200 may include a signal processing module 201, a voiceprint extraction module 202, and a clustering module 203. All speech data collected by the microphone (e.g., speech data 110, speech data 120, speech data 130, and noise data 140) are used as input, and the recognition result of the speech data is used as output.

[0070] The signal processing module 201 can process all channels of the voice data received from the microphone to obtain single-channel voice data (Pulse Code Modulation, PCM). Then, it sends the single-channel voice data to the voiceprint extraction module 202.

[0071] The voiceprint extraction module 202 can be used to extract voiceprints from single-channel speech data received from the signal processing module 201, so as to extract the voiceprint features (also known as "voiceprint information") corresponding to the single-channel speech data. Then, the extracted voiceprint features are sent to the clustering module 203.

[0072] Voiceprint features may include, but are not limited to, spectral features, temporal features, frequency features, pitch, stress, vocalization order, and rhythm.

[0073] The clustering module 203 can be used to perform cluster analysis on the voiceprint features corresponding to the single-channel speech data sent by the voiceprint extraction module 202, group similar voiceprint features into one category, and obtain the speaker corresponding to each category of voiceprint features. The speaker's identity information is then output as the recognition result of the speech data.

[0074] Alternatively, the voiceprint features corresponding to the single-channel speech data extracted by the voiceprint extraction module 202 can be compared with a pre-constructed voiceprint feature set to obtain the speaker corresponding to the single-channel speech data.

[0075] However, due to environmental interference during the construction of the pre-built voiceprint feature set, there may be a lot of interference information in the pre-built voiceprint feature set, which reduces the accuracy of the pre-built voiceprint feature set. As a result, the speaker corresponding to the speech data determined based on the pre-built voiceprint feature set may be incorrect, leading to poor accuracy of the recognition result.

[0076] In related technologies, end-to-end speaker differentiation models can be used to identify the speaker corresponding to speech data. However, due to the large amount of labeled data required for training, as well as the high model complexity and poor generalization, end-to-end speaker differentiation models have poor accuracy in identifying the speaker corresponding to speech data, making them difficult to apply widely.

[0077] It should be noted that electronic device 160 can be an intelligent device with the function of recognizing the speaker of voice data, including but not limited to: personal computers, tablets, handheld devices, in-vehicle devices, wearable devices, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access electronic device, user unit, user station, mobile station, mobile station, remote station, remote electronic device, mobile device, user electronic device, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, electronic device in a 5G network or future evolved network, etc. The embodiments in this application are not limited in scope.

[0078] Therefore, in order to address the problem of poor accuracy of voiceprint feature sets, this application proposes a real-time speaker differentiation method, apparatus, electronic device, and storage medium.

[0079] The real-time speaker differentiation method provided in the embodiments of this application will be described in detail below with reference to Figures 3 to 5.

[0080] Figure 3 is a schematic diagram of another speaker recognition model provided in an embodiment of this application.

[0081] For example, as shown in Figure 3, the speaker recognition model 300 may include a signal processing module 201, a voiceprint extraction module 202, a voiceprint comparison module 204, a stability judgment module 205, and a voiceprint registration and update module 206 (which can be referred to as the "voiceprint registration / update module"). All speech data collected by the microphone (e.g., speech data 110, speech data 120, speech data 130, and noise data 140) are used as input, and the recognition result of the speech data is used as output.

[0082] The signal processing module 201 can process all channels of the voice data received from the microphone to obtain single-channel voice data (Pulse Code Modulation, PCM). This single-channel voice data is then sent to the speaker extraction module 202. Furthermore, the signal processing module 201 can also extract the spatial features (also referred to as "spatial information") of all the voice data received from the microphone. These spatial features are then sent to the stability assessment module 205.

[0083] Spatial features may include, but are not limited to, the orientation (which can represent position and direction), distance, and motion features of all speech data in three-dimensional space. These features can help the speaker recognition model 300 perceive the specific location of the speech source in all speech data. For ease of understanding, the following embodiments can be described using the direction of arrival (DOA) of the speech data.

[0084] The voiceprint extraction module 202 can be used to extract voiceprints from single-channel speech data received from the signal processing module 201, thereby extracting the voiceprint features corresponding to the single-channel speech data. Furthermore, the voiceprint features corresponding to the single-channel speech data are sent to the voiceprint comparison module 204.

[0085] Optionally, the network structure of the voiceprint extraction module 202 can employ a residual network. The residual network can include four residual blocks, each consisting of multiple convolutional layers and residual connections. The size of the feature map within each residual block is halved. Furthermore, each residual block consists of two "Basic Blocks," each containing two convolutional layers for voiceprint feature extraction. The kernel sizes of these two convolutional layers are 1×1 and 3×3, respectively; that is, if one convolutional layer has a kernel size of 1×1, the other has a kernel size of 3×3. An average pooling layer can be connected after each residual block to reduce the spatial dimension of the feature map. A fully connected layer can then be connected after the average pooling layer to normalize the reduced spatial dimension feature map, outputting the final voiceprint features.

[0086] For example, if a single-channel voice data is 1.5 seconds of voice data, the voiceprint extraction module 202 can extract voiceprint features from the 1.5 seconds of voice data and output 256-dimensional voiceprint features.

[0087] The voiceprint comparison module 204 can compare the single-channel voiceprint features received from the voiceprint extraction module 202 with the voiceprint information in the voiceprint registration / update module 206 to obtain a voiceprint comparison result. The speaker corresponding to the single-channel voiceprint features is then determined based on the comparison result. Furthermore, the speaker's identity information is output as the recognition result of the speech data. Finally, the voiceprint comparison result is sent to the voiceprint registration / update module 206.

[0088] The difference between the voiceprint comparison module 204 and the clustering module 203 is as follows:

[0089] Clustering module 203 performs clustering through a complex classification process, dividing speech data into many categories, then calculating the similarity between the voiceprint features of the speech data in multiple categories and the voiceprint features in the voiceprint feature database; it outputs the recognition result; this process consumes a large amount of central processing unit (CPU) resources, typically requiring a chip of Intel personal computer (PC) level. Moreover, as the meeting progresses and the speech data samples accumulate, the time spent by the clustering algorithm increases, making it difficult to distinguish speakers in real time. Furthermore, the clustering algorithm performs poorly in complex scenarios, such as when the speaking time is short and a clustered audio segment contains more than two speakers, resulting in inaccurate recognition.

[0090] The voiceprint comparison module 204 does not require clustering to classify speech data. Instead, it extracts voiceprint features from the speech data and calculates cosine similarity with samples in the registered voiceprint database. Simultaneously, it evaluates the stability of the speech data based on its DOA (Document of Address) information. Once the stability of the speech data meets the criteria, the voiceprint features are added to the registered voiceprint database—a continuously updating process. It consumes relatively few resources; computational resources are reduced by several times (e.g., 20 times) compared to clustering, and the required chip is a common embedded chip.

[0091] Assume that the voiceprint registration / update module 206 already includes M registered speakers (also called "registered speakers"). Among these M registered speakers, the voiceprint of the m-th registered speaker (which can be denoted as "C") is... m ") can be expressed by formula (1):

[0092] In formula (1), It can represent the first registered voiceprint feature of the m-th registered speaker. This can represent the second registered voiceprint feature of the m-th registered speaker, and so on. This can represent the nth registered voiceprint feature of the mth registered speaker. Here, m is a positive integer greater than or equal to 1.

[0093] For example, the voiceprint comparison module 204 can calculate the t-th voiceprint feature of a single channel (which can be denoted as "f"). t ") and the m-th registered speaker's C m The similarity between them is then calculated, and the average of the resulting similarities is denoted as "sim(C)". m ,f t )”). This can be expressed by formula (2):

[0094] In formula (2), cos(·) represents the cosine similarity. express with f t The cosine similarity.

[0095] For example, 1.5 seconds of speech data can be composed of the speech data from the current 60ms and the speech data from 1440ms prior to the current moment. Furthermore, speaker recognition can be performed on the acquired speech data in 60ms increments, and each 60ms of speech data can include both voiceprint features and spatial features. 't' can represent the nth 60ms increment.

[0096] When multiple sim(C) are obtained m ,f t When ), multiple sim(C) can be calculated. m ,f t The result with the highest similarity in (C) can be denoted as "max[sim(C)". m ,f t )]”), and determine the max[sim(C m ,f t Whether the value is greater than or equal to a set threshold (e.g., 95%), when max[sim(C m ,f t When C is greater than or equal to the set threshold, it can be... m The corresponding registered speaker (can be denoted as "m") * The speaker whose voiceprint feature corresponds to the single-channel speech data is identified as C. m The corresponding registered speaker's statement. Expressed by formula (3): m * =argmax m [sim(C m ,f t (3)

[0097] Furthermore, it can be done through ft For the voiceprint registration / update module 206 m * The voiceprint characteristics are updated. Alternatively, if the voiceprint registration / update module 206 does not include f... t The corresponding speaker needs to be identified through f. t Speaker registration is performed in the voiceprint registration / update module 206.

[0098] Among them, multiple sim(C m ,f t ) can be represented as sim(C1,f t ), sim(C2,f t )……sim(C m ,f t For example, when sim(C3,f) t ) is max[sim(C m ,f t )], and sim(C3,f t When the threshold is greater than or equal to the set threshold, then the registered speaker corresponding to C3 (which can be denoted as "3") can be added. * The speaker of the single-channel speech data corresponding to the voiceprint feature of that single channel is identified.

[0099] Or, when max[sim(C) m ,f t When the threshold is set, it indicates that f t Not belonging to C m The corresponding registered speaker (can be denoted as "m") * ”), that is, f t The speaker in question is not a pre-registered speaker.

[0100] The stability judgment module 205 can be used to determine whether the spatial feature is stable when it receives the spatial feature sent by the signal processing module 201. When it is determined that the spatial feature is stable, the spatial feature can be sent to the voiceprint registration / update module 206.

[0101] For example, spatial characteristics (e.g., DOA) are prone to unpredictable mutations when disturbed by the external environment. To prevent DOA mutations from affecting the voiceprint information in the voiceprint registration / update module 206, the stability judgment module 205 can perform a stability judgment on the DOA. When the DOA is relatively stable for a period of time (e.g., 2 seconds), the voiceprint information in the voiceprint registration / update module 206 is then registered or updated via the DOA.

[0102] Suppose that the stability determination module 205 continuously receives K DOAs of a single-channel voice data (e.g., the aforementioned 1.5s voice data). The DOA of this single-channel voice data (which can be denoted as "D") is... t ") can be expressed by formula (4): D t =[d t-K+1 …d t (4)

[0103] In formula (4), t represents the nth DOA.

[0104] Stability judgment module 205 obtains D through formula (4). t At that time, it is possible to target D t Perform first-order difference calculations (which can be denoted as...) ), which can be expressed by formula (5):

[0105] Stability judgment module 205 obtains the result through formula (4). At this time, the time series outlier detection (Smoothed Z-score) algorithm can be used to detect outliers. Perform smoothing filtering, and then process the smoothed filter. Perform outlier detection to obtain smoothed outlier results. When the smoothed outlier results indicate... When the percentage of outliers is greater than or equal to a preset threshold (e.g., 50%), it indicates a large number of outliers, suggesting that the DOA stability of the single-channel voice data is poor and cannot be used to register or update the voiceprint information in the voiceprint registration / update module 206. Conversely, when the smoothed outlier detection result indicates... When the percentage of outliers in the data is less than a preset threshold (e.g., 50%), it indicates that there are few outliers. It can be determined that the DOA corresponding to the voice data of this single channel has good stability and can be used to register or update the voiceprint information in the voiceprint registration / update module 206. The DOA corresponding to the voice data of this single channel is then sent to the voiceprint registration / update module 206.

[0106] It should be noted that when a speaker's position remains unchanged in the external environment (i.e., the speaker is at a fixed coordinate), the DOA is generally relatively stable with a low probability of abrupt changes, and the corresponding DOA is within the normal range. Alternatively, when a speaker's position moves slowly in the external environment, the DOA also changes slowly, and the corresponding DOA is within the normal range. However, when multiple speakers are speaking from different positions, the DOA stability is poor, and abrupt changes may occur, resulting in an outlier DOA. Furthermore, the stability assessment module 205 can be used to evaluate the stability of the DOA of speech data, filtering out speech data with more stable DOA. This prevents speech data with poor DOA from affecting the accuracy of the output of the voiceprint registration / update module 206, thereby improving the accuracy of the output of the voiceprint registration / update module 206.

[0107] The voiceprint registration / update module 206 can be used to determine whether to register or update the voiceprint feature when it receives the spatial feature sent by the stability judgment module 205 and the voiceprint comparison result sent by the voiceprint comparison module 204.

[0108] For example, when single-channel voice data is sent to the voiceprint registration / update module 206, and the speaker corresponding to the single-channel voice data is obtained through the voiceprint comparison module 204, it indicates that the speaker corresponding to the single-channel voice data belongs to a registered speaker who has already completed registration (e.g., the m-th registered speaker). The voiceprint registration / update module 206 can then use the single-channel voice data to update the registered voiceprint features of the m-th registered speaker. Wherein, the registered voiceprint features of the m-th registered speaker are...

[0109] Optionally, when updating the registered voiceprint features of the m-th registered speaker, it can be determined first whether the number of registered voiceprint features completed by the current m-th registered speaker has reached the set maximum number (e.g., 10).

[0110] When the number of registered voiceprint features of the m-th registered speaker is less than 10, the voiceprint feature extracted from the single-channel speech data can be directly determined as the (N+1)-th registered voiceprint feature of the m-th registered speaker. That is, the registered voiceprint feature of the m-th registered speaker is...

[0111] When the number of registered voiceprint features completed by the m-th registered speaker is 10, the voiceprint center of the 10 completed registered voiceprint features (which can be denoted as "P") can be calculated first, that is, the average value of the 10 registered voiceprint features (also called "feature mean") can be calculated. This can be expressed by formula (6):

[0112] In formula (6), N is the maximum number that can be set, for example, 10; it can also be 15, 20, etc., and this application embodiment does not limit this.

[0113] Furthermore, when P is calculated using formula (6), the similarity between P and each of the 10 registered voiceprint features of the m-th registered speaker can be calculated, for example, similarity m1, similarity m2... similarity m10, as well as the similarity between P and the voiceprint features of single-channel speech data, for example, similarity m11.

[0114] The similarity scores m1, m2, ..., m10 and m11 are sorted from largest to smallest or smallest to largest to obtain the largest similarity score among them (e.g., similarity m7). When the largest similarity score is determined to be m7, the voiceprint features extracted from single-channel speech data can replace the registered voiceprint features corresponding to similarity m7. That is, the updated 10 registered voiceprint features of the m-th registered speaker are the 9 pre-registered registered voiceprint features and the voiceprint features extracted from single-channel speech data, excluding the registered voiceprint features corresponding to similarity m7. This is because, in order to make the registered voiceprint features of the m-th registered speaker more diverse and generalizable, and to include multiple voiceprint features of the m-th registered speaker, thereby improving the generalization of voiceprint feature recognition by the voiceprint comparison module 204, the registered voiceprint features corresponding to the largest similarity score are replaced with the voiceprint features extracted from single-channel speech data.

[0115] For example, when single-channel voice data is sent to the voiceprint registration / update module 206, and the voiceprint comparison module 204 cannot obtain the speaker corresponding to the single-channel voice data, it indicates that the speaker corresponding to the single-channel voice data has not yet been registered. When there are continuous voice segments (e.g., 3 seconds of voice data) in the single-channel voice data, and the voiceprint features of the continuous voice segments have a high similarity, it indicates that the continuous voice segments may belong to the same speaker. Therefore, a new registered speaker can be registered for the voiceprint features of the continuous voice segments (which can be denoted as "the (m+1)th registered speaker").

[0116] Optionally, to ensure the validity of newly registered speakers, speaker registration can be performed on the continuous speech segment when its duration is greater than or equal to a preset duration (e.g., 3 seconds). For example, when the duration of the continuous speech segment is 3 seconds, it can be divided into 60ms segments to calculate 50 voiceprint features corresponding to the continuous speech segment.

[0117] Suppose that a continuous speech segment contains K consecutive speaker features (which can be denoted as "[f"). t-K+1 ,…,f t First, identify from the K voiceprint features. Each voiceprint feature, and The voiceprint features are continuous. For example, the first voiceprint feature to the second... The first voiceprint feature, or the second From the 10th voiceprint feature to the Kth voiceprint feature, or from the 10th voiceprint feature to the Kth voiceprint feature. The embodiments of this application do not limit the voiceprint features.

[0118] However, to avoid bias in some of the K voiceprint features due to delay, speaker registration can be performed using the voiceprint features with earlier temporal information. For example, the first voiceprint feature to the... Voiceprint features.

[0119] When the first voiceprint feature is assigned to the second... The voiceprint features were determined as follows When calculating the first voiceprint feature, the following steps can be performed: The voiceprint center of each voiceprint feature (which can be denoted as "C") m+1 "), and the center of that voiceprint is determined as the registered voiceprint of the new registered speaker. This is expressed by formula (7):

[0120] In formula (7), M is the original number of registered speakers (i.e., the above m). This represents the first registered voiceprint of the (M+1)th speaker.

[0121] Furthermore, when C is determined M+1 At that time, the first can be calculated sequentially. Each voiceprint feature from the first to the Kth voiceprint feature is related to C. M+1 The similarity of the feature means to C M+1 The registered voiceprint features are updated.

[0122] For example, calculate the first Voiceprint features and C M+1 The similarity when the first Voiceprint features and C M+1 When the similarity is ≥ the preset similarity (e.g., 95%), it indicates that the first... The voiceprint feature is the voiceprint feature of the (M+1)th speaker, which can be used to identify the first speaker. Each voiceprint feature is updated to C M+1 The registered voiceprint features, i.e. Conversely, when the first Voiceprint features and C M+1 When the similarity is less than the preset similarity (e.g., 95%), it indicates that the first... If the voiceprint feature is not the voiceprint feature of the (M+1)th speaker, then the voiceprint feature will not be included. Each voiceprint feature is updated to C M+1 The registered voiceprint features, i.e. And so on, up to C. M+1 When the number of registered voiceprint features reaches the maximum set number (e.g., 10), if the first... There are still voiceprint features that need to be registered among the first to the Kth voiceprint features. These can be replaced by registering voiceprint features in the manner described above, which will not be elaborated further here.

[0123] For example, when calculating the first Each voiceprint feature from the first to the Kth voiceprint feature is related to C. M+1 When considering the similarity, if the first... If most (e.g., 90%) or all of the voiceprint features from the first to the Kth voiceprint features are the voiceprint features of the (M+1)th speaker, then the (M+1)th speaker is a valid speaker and can be registered as a new registered speaker, while retaining the corresponding C for the (M+1)th speaker. M+1 Conversely, if the first If most (e.g., 90%) or all of the voiceprint features from the first to the Kth voiceprint features are not the voiceprint features of the (M+1)th speaker, then the (M+1)th speaker is an invalid speaker. Therefore, the (M+1)th speaker will not be registered, and the corresponding C address for the (M+1)th speaker can be deleted. M+1 This saves on device storage resources.

[0124] In this system, each of the K voiceprint features is temporally continuous with its preceding and following voiceprint features. For example, the temporal sequence of the 6th voiceprint feature among the K features is as follows: the timestamp of the 5th voiceprint feature is earlier than the timestamp of the 6th voiceprint feature, the timestamp of the 6th voiceprint feature is earlier than the timestamp of the 7th voiceprint feature, and the timestamps of the 5th, 6th, and 7th voiceprint features are continuous. This is because if the voiceprint features are not continuous among the K features, some voiceprint features may be ignored or lost, thus affecting the accuracy of the recognition process.

[0125] The difference between speaker recognition model 300 and end-to-end speaker discrimination models is as follows:

[0126] End-to-end speaker discrimination models directly read speech data and infer the speaker for each frame. However, their drawbacks include the need for a large model size for good performance, difficulty in obtaining training data, and poor real-time performance. Furthermore, the required computational resources increase with the accumulation of data from a conference. Sufficient inference resources and a large amount of suitable training data are needed to achieve satisfactory results.

[0127] The speaker recognition model 300 in this application uses a small ResNet18 model. Although it also needs to be trained in advance, this model does not have high requirements for data matching accuracy and can use open-source datasets, which are plentiful. The speaker recognition model 300 has good robustness. However, end-to-end model training requires a series of meeting data sessions, which are scarce in open-source datasets, leading to poor robustness of the end-to-end model.

[0128] Furthermore, compared to end-to-end models, the speaker recognition model 300 in this application can achieve better speaker discrimination performance in complex scenarios involving speaker discrimination.

[0129] Taking a scenario where the speaker is moving as an example, end-to-end models are easily affected by noise in the speaker's environment, as well as the distance and angle between the speaker and the microphone, when distinguishing speakers. This can cause abrupt changes in the collected speaker's speech data, interfering with the speaker distinction results and leading to errors in speaker distinction by the end-to-end model, thus reducing the accuracy of the speaker distinction results. However, the speaker recognition model 300 in this application can filter out relatively stable speech data from the speaker's speech data based on the DOA (Document of Audio Ability) of the speech data, avoiding interference from abrupt changes in the speaker's speech data. Therefore, even if the speaker is constantly moving during the speaking process, the model can accurately distinguish the speaker. Compared to end-to-end models, this solution improves the accuracy of speaker distinction results and enhances the robustness of the speaker recognition model 300.

[0130] Taking a scenario involving changes in speaker emotion as an example, when a speaker is emotionally agitated, the pitch, loudness, and speech rate of the speech data may change significantly, interfering with the speaker discrimination results. End-to-end models struggle to adapt to these sudden changes in speaker speech data, leading to misclassification. This makes it difficult for end-to-end models to quickly adapt to changes in speech data caused by factors such as speaker emotion and speech rate. However, the speaker recognition model 300 in this application can filter out relatively stable speaker speech data, avoiding interference from changes in speech data caused by factors such as speaker emotion and speech rate. This allows for accurate speaker discrimination even when the speaker's emotion changes. Compared to end-to-end models, this solution demonstrates strong adaptability, providing a more reliable guarantee for speaker discrimination in complex and ever-changing scenarios.

[0131] Figure 4 is a flowchart illustrating a real-time speaker differentiation method provided in an embodiment of this application. This method can be executed by the electronic device 160 shown in Figure 1.

[0132] For example, as shown in Figure 4, the real-time speaker differentiation method 400 includes the following implementation process:

[0133] S410, acquire the speaker's speech data and initial voiceprint feature set.

[0134] For example, in a multi-person speaking scenario (such as the conference scenario shown in Figure 1), to determine the speaker corresponding to the voice data, the voice data of the speakers in the conference scenario can be collected through a microphone array, and the collected voice data (hereinafter referred to as "voice data") can be input into the speaker recognition model 300 shown in Figure 3. The speaker recognition model 300 then performs speaker recognition on the voice data. Simultaneously, when acquiring the voice data, a voiceprint feature database (which can be called the "initial voiceprint feature set") can be obtained to identify the speaker corresponding to the voice data, and the initial voiceprint feature set can be updated using the voice data.

[0135] The initial voiceprint feature set may include zero, one, or more sub-voiceprint feature sets, and each sub-voiceprint feature set may include one or more registered voiceprint features. It should be understood that the initial voiceprint feature set may represent a registered voiceprint database.

[0136] For example, the initial voiceprint feature set includes sub-voiceprint feature set A and sub-voiceprint feature set B. Sub-voiceprint feature set A includes 3 registered voiceprint features, and sub-voiceprint feature set B includes 10 registered voiceprint features. Alternatively, the initial voiceprint feature set may not include sub-voiceprint feature sets.

[0137] S420 determines the location of voice data based on the time delay information of the voice data.

[0138] For example, when acquiring speech data through a microphone array, the time at which each microphone in the array acquires speech data may differ. The time delay estimate of the speech data (which can be called "time delay information") can be calculated using the generalized cross-correlation (GCC) function; or, cepstral analysis can be used to calculate the time delay information of the speech data.

[0139] When the time delay information of the speech data is calculated, the location of the sound source corresponding to the time delay information can be quickly looked up in a preset mapping table. For example, d t When multiple d values ​​are obtained from the voice data t At this time, the final location of the voice data can be determined (hereinafter referred to as "location"), for example, D t .

[0140] The correspondence between time delay information and orientation in the preset mapping table can be pre-calibrated, and this application embodiment does not limit this.

[0141] S430 obtains a stability assessment value for the voice data based on its location.

[0142] The stability assessment value is negatively correlated with the change in the location of the speech data. The larger the stability assessment value, the smaller the change in the location of the speech data, indicating that the speech data is more stable; conversely, the smaller the stability assessment value, the larger the change in the location of the speech data, indicating that the speech data is less stable.

[0143] For example, when voice data is acquired, the location corresponding to the voice data can be determined, and the stability of the voice data can be evaluated by the determined location of the voice data to obtain the stability evaluation value of the voice data.

[0144] Optionally, when determining the stability evaluation value of the speech data, the orientation of the speech data can be input into the stability judgment module first. The stability judgment module performs differential processing on the orientation of the speech data (as shown in the above formula (5)) to obtain the orientation after differential processing (which can be called "differential orientation"). ).

[0145] Furthermore, once the differential orientation is obtained, the Smoothed Z-score algorithm can be used to... Perform smoothing processing (i.e., the smoothing filter mentioned above) to obtain the smoothed orientation (which can be called the "smoothed orientation").

[0146] Once the smoothed orientation is obtained, outlier detection can be performed on the smoothed orientation to obtain the proportion of outliers in the smoothed orientation. Then, the stability of the speech data can be evaluated by using this proportion of outliers and the total number of smoothed orientations to obtain the stability evaluation value of the speech data.

[0147] For example, if the total number of smoothed orientations is 100 and the number of outliers is 10, then... The corresponding stability assessment value is 1 - 10% = 90%.

[0148] Outliers can represent excessively large or small directional values ​​in the smoothed directional data. Furthermore, the directional data in the voice data can be one or more, and this embodiment does not limit this.

[0149] In this embodiment, when the orientation of the speech data is differentially processed, the differences between the orientations of adjacent frames can be captured, reducing redundant information and making the differential orientation more accurate. Furthermore, when the more accurate differential orientation is smoothed, interference information in the speech data can be effectively suppressed, making the smoothed orientation more accurate. This results in a more accurate stability assessment value for the speech data obtained through the smoothed orientation. Consequently, based on a more accurate stability assessment value, the updated initial voiceprint features can be made more accurate.

[0150] S440 performs voiceprint recognition on the speech data based on the voiceprint features of the speech data and the cosine similarity with the initial voiceprint feature set, and outputs the target identifier.

[0151] The target identifier is used to represent the speaker's identity information.

[0152] For example, when voice data is acquired, voiceprint recognition can be performed on the voice data by comparing the voiceprint features of the voice data with the cosine similarity of the initial voiceprint feature set, to obtain the identifier (which can be called the "target identifier") corresponding to the speaker's identity information of the voice data. The speaker corresponding to the target identifier is then identified as the speaker of the voice data.

[0153] For example, if the speaker corresponding to the target identifier is user A, then user A is the speaker of the voice data corresponding to the target identifier.

[0154] The target identifier may include, but is not limited to, name, number, identifier, and code.

[0155] Optionally, when performing voiceprint recognition on the acquired speech data, the speech data can first be input into the voiceprint extraction module to extract the voiceprint features of the speech data. Once the voiceprint features of the speech data are extracted, they can be input into the voiceprint comparison module, which then calculates the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set.

[0156] For example, first calculate the cosine similarity between each voiceprint feature in each sub-voiceprint feature set and the voiceprint features in the speech data, then calculate the mean similarity of at least one cosine similarity (i.e., the above sim(C)). m ,f t The cosine similarity between the voiceprint features of the speech data and the cosine similarity between each sub-voiceprint feature set is denoted as ).

[0157] Furthermore, when multiple cosine similarities are calculated, they can be sorted in ascending or descending order to obtain the maximum cosine similarity among the multiple cosine similarities.

[0158] When the maximum cosine similarity among multiple cosine similarities is obtained, in order to avoid the problem of speaker identification error caused by the maximum cosine similarity being too small, it is possible to further determine whether the maximum cosine similarity is greater than or equal to a preset similarity (e.g., 95%).

[0159] When the maximum cosine similarity (e.g., 97%) is greater than 95%, it indicates that the voiceprint features in the sub-voiceprint feature set corresponding to the maximum cosine similarity have a high degree of agreement with the voiceprint features of the speech data. Therefore, the identifier of the sub-voiceprint feature set corresponding to the maximum cosine similarity can be determined as the target identifier of the speech data, and the determined target identifier of the speech data can be output.

[0160] When the maximum cosine similarity (e.g., 30%) is less than 95%, it indicates that the voiceprint features in the sub-voiceprint feature set corresponding to the maximum cosine similarity have a low degree of match with the voiceprint features of the speech data. Therefore, it can be concluded that the identifier of the sub-voiceprint feature set corresponding to the maximum cosine similarity is not the target identifier of the speech data, that is, the speaker's identity information corresponding to the speech data does not exist in the preset voiceprint feature set. In this case, a reminder message can be generated to remind the user to add the speaker's identity information corresponding to the speech data to the preset voiceprint feature set.

[0161] The preset voiceprint feature set may include all sub-voiceprint feature sets in the initial voiceprint feature set. That is, the sub-voiceprint feature sets in the preset voiceprint feature set may be completely the same as the sub-voiceprint feature sets in the initial voiceprint feature set, or the sub-voiceprint feature sets in the preset voiceprint feature set may include other sub-voiceprint feature sets in addition to the sub-voiceprint feature sets in the initial voiceprint feature set. This application embodiment does not limit this.

[0162] It should be noted that the preset similarity can be determined according to the speaker's recognition needs. The preset similarity can be 95%, 98%, or 92%, etc., and this application embodiment does not limit it.

[0163] S420 and S440 can be executed simultaneously or sequentially, and this application embodiment does not limit this.

[0164] In this embodiment, when the maximum cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set is greater than or equal to the preset similarity, the similarity between the voiceprint features of the speech data and the speaker's identity information can be ensured to the greatest extent, while avoiding the problem of speaker identification errors caused by excessively low maximum cosine similarity, thereby improving the accuracy of speaker identification. Furthermore, since the voice data identifier is obtained by calculating the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set, the identification process is simpler and requires less computational resources compared to using clustering to obtain speaker identity information, thus exhibiting better universality.

[0165] S450: When the stability evaluation value is greater than or equal to the preset threshold, the initial voiceprint feature set is updated based on the voiceprint features of the target identifier and the speech data to obtain the updated initial voiceprint feature set.

[0166] For example, when the stability assessment value of the voice data is obtained, it can be determined whether the stability assessment value is greater than or equal to a preset threshold (e.g., 60%).

[0167] When the stability assessment value of the voice data (e.g., 80%) is greater than 60%, it indicates that the directional change of the voice data is small, and the voice data is relatively stable. In the voiceprint registration / update module, the initial voiceprint feature set can be updated using the target identifier and voiceprint features of the voice data to obtain an updated initial voiceprint feature set. Furthermore, to ensure the accuracy of speaker differentiation in the voice data, the updated initial voiceprint feature set can replace the unupdated initial voiceprint feature set (i.e., the initial voiceprint feature set in S410) to distinguish speakers from the voice data. It should be understood that once the updated initial voiceprint feature set is obtained, it can be used to replace the unupdated initial voiceprint feature set, and the number of replacements is not limited, for example, two or three times. This embodiment does not impose such a limitation.

[0168] When the stability assessment value of the speech data (e.g., 20%) is less than 60%, it indicates that the speech data has a large degree of directional variation and is unstable. In this case, updating the initial voiceprint feature set using the target identifier and voiceprint features of the speech data may lead to deviations in the updated initial voiceprint feature set, affecting its accuracy. Therefore, this speech data can be discarded, and the initial voiceprint feature set can be skipped by not updating it using the target identifier and voice data. Instead, new speech data can be acquired and used as new input for speaker differentiation.

[0169] In method 400 as shown in Figure 4, when acquiring the speaker's speech data and initial voiceprint feature set, the location corresponding to the speech data can first be determined through the time delay information of the speech data; then, the stability evaluation value of the speech data can be obtained through the location of the speech data; and finally, voiceprint recognition can be performed on the speech data through the cosine similarity between the voiceprint features of the speech data and the initial voiceprint feature set to obtain the target identifier corresponding to the speech data. Furthermore, when the stability evaluation value is greater than or equal to a preset threshold, the initial voiceprint feature set can be updated using the target identifier of the speech data and the voiceprint features of the speech data to obtain an updated initial voiceprint feature set. Since when the stability evaluation value is ≥ the preset threshold, it indicates that the change in the location of the speech data is small, the location of the speech data is relatively stable, and the speech data will not undergo abrupt changes; or the possibility of abrupt changes in the speech data is low. Therefore, when the stability evaluation value is greater than or equal to the preset threshold, the initial voiceprint feature set can be updated using the target identifier and voiceprint features of the voice data. This can make the updated initial voiceprint feature set more stable and avoid the problem of deviation in the updated initial voiceprint feature set caused by the stability evaluation value being too small, thereby improving the accuracy of the updated initial voiceprint feature set.

[0170] Furthermore, with the updated initial voiceprint feature set being more accurate, the speaker identified based on the updated initial voiceprint feature set can be more accurate, thus improving the accuracy of speaker identification results.

[0171] For example, when the stability assessment value of the voice data (e.g., 80%) is greater than 60%, it can be determined whether there is a target identifier corresponding to the voice data in the initial voiceprint feature set.

[0172] When a target identifier corresponding to the speech data exists in the initial voiceprint feature set, a sub-voiceprint feature set corresponding to the target identifier (which can be called the "target sub-voiceprint feature set") can be determined from the initial voiceprint feature set. Furthermore, when the target sub-voiceprint feature set corresponding to the target identifier is determined, the number of voiceprint features included in that target sub-voiceprint feature set can be obtained.

[0173] For example, if the sub-voiceprint feature set corresponding to the target identifier is C, then the sub-voiceprint feature set C is the target sub-voiceprint feature set corresponding to the target identifier.

[0174] Furthermore, the initial voiceprint feature set is updated based on the number of voiceprint features included in the target sub-voiceprint feature set and the voiceprint features of the speech data, resulting in an updated initial voiceprint feature set.

[0175] Among them, the voiceprint features included in the target sub-voiceprint feature set are all voiceprint features of the target identifier in different environments, which have good generalization ability.

[0176] Optionally, when the number of voiceprint features included in the target sub-voiceprint feature set is obtained, it can be determined whether the number is less than a preset number (e.g., 10).

[0177] When the number (e.g., 2) is less than 10, the voiceprint features of the speech data can be directly added to the target sub-voiceprint feature set (e.g., the sub-voiceprint feature set is C) to obtain the updated target sub-voiceprint feature set. Furthermore, since the initial voiceprint feature set is also updated at the same time as the target sub-voiceprint feature set is updated, the updated initial voiceprint feature set can be obtained.

[0178] For example, the sub-voiceprint feature set C before the update includes voiceprint feature C1 and voiceprint feature C2 (i.e., the above). When the voiceprint features (e.g., voiceprint feature M) of the speech data are added to the unupdated sub-voiceprint feature set C, the updated sub-voiceprint feature set C includes voiceprint features C1, C2, and M, as described above.

[0179] Alternatively, when the number (e.g., 10) = 10, the mean value of all voiceprint features in the target sub-voiceprint feature set (which can be called the "first voiceprint feature", denoted as "P") can be determined by the above formula (6). The cosine similarity value between P and the voiceprint features of the speech data (which can be called the "first similarity value"), and the cosine similarity value between P and each voiceprint feature in the target sub-voiceprint feature set (which can be called the "second similarity value") are calculated respectively.

[0180] When the first similarity value and the second similarity value are obtained, the initial voiceprint feature set can be updated using the first similarity value and the second similarity value to obtain the updated initial voiceprint feature set.

[0181] It should be noted that the preset number can be determined according to the update requirements of the initial voiceprint feature set. The preset number can be 10, 15 or 12, etc., and this application embodiment does not limit it.

[0182] In this embodiment, when the number of voiceprint features in the target sub-voiceprint feature set is small, the voiceprint features of the speech data can be added to the target sub-voiceprint feature set, making the updated target sub-voiceprint feature set richer in voiceprint features and having better generalization ability. Alternatively, when the number of voiceprint features in the target sub-voiceprint feature set is large, the initial voiceprint feature set can be updated based on the similarity between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set, and the average feature values ​​of the voiceprint features in the target sub-voiceprint feature set, respectively. This can also make the updated initial voiceprint feature set have better generalization ability.

[0183] Furthermore, the first and second similarity values ​​are first sorted from largest to smallest or smallest to largest to obtain the maximum similarity between the first and second similarity values. For example, when the first similarity is greater than the second similarity, the first similarity is the maximum similarity between the first and second similarity values.

[0184] Next, the voiceprint feature corresponding to the maximum similarity value (which can be called the "second voiceprint feature") is determined. The voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set other than the second voiceprint feature (which can be called the "third voiceprint feature") are determined as the updated voiceprint features of the target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0185] For example, if the voiceprint feature corresponding to the maximum similarity value is the voiceprint feature of the speech data, then the voiceprint feature of the target sub-voiceprint feature set remains unchanged.

[0186] For example, if the voiceprint feature corresponding to the maximum similarity value is voiceprint feature M1 in the target sub-voiceprint feature set, then voiceprint feature M1 can be filtered out and replaced with the voiceprint feature of the speech data to obtain the updated target sub-voiceprint feature set.

[0187] In this embodiment, when the maximum similarity between the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set and the mean similarity of the voiceprint features in the target sub-voiceprint feature set is determined, the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set other than the voiceprint features corresponding to the maximum similarity (i.e., the third voiceprint features) are determined as the updated voiceprint features of the target sub-voiceprint feature set. That is, the updated voiceprint features of the target sub-voiceprint feature set do not include the voiceprint features corresponding to the maximum similarity. This makes the voiceprint features of the updated target sub-voiceprint feature set more representative and can represent the voiceprint features of speakers in different environments, thereby giving the updated initial voiceprint feature set better generalization.

[0188] For example, if there is no target identifier corresponding to the speech data in the initial voiceprint feature set, it means that the speaker's identity information corresponding to the speech data has not been registered in the initial voiceprint feature set. In order to enrich the speaker's identity information in the initial voiceprint feature set, a new sub-voiceprint feature set can be registered through the speech data; and the new sub-voiceprint feature set is added to the initial voiceprint feature set to update the initial voiceprint feature set and obtain the updated initial voiceprint feature set.

[0189] In this embodiment, when a target identifier exists in the initial voiceprint feature set, the initial voiceprint feature set can be updated using the number of voiceprint features in the sub-voiceprint feature set (i.e., the target sub-voiceprint feature set) corresponding to the target identifier in the initial voiceprint feature set, along with the voiceprint features of the speech data, to obtain an updated initial voiceprint feature set. Alternatively, when a target identifier does not exist in the initial voiceprint feature set, a new sub-voiceprint feature set can be registered using the voiceprint features of the speech data, and the initial voiceprint feature set can be updated using the new sub-voiceprint feature set to obtain an updated initial voiceprint feature set. By updating the initial voiceprint feature set in two different ways, there is a corresponding update method for both the presence and absence of a target identifier, rather than using a uniformly prescribed method. This avoids the problem of deviations in the updated initial voiceprint feature set caused by using a uniformly prescribed method, thereby improving the accuracy of the updated initial voiceprint feature set.

[0190] Optionally, when there is no target identifier corresponding to the speech data in the initial voiceprint feature set, the mean value of the voiceprint features of the speech data (which can be called the "target feature mean") can be determined by the above formula (6).

[0191] For example, selecting a target number of voiceprint features from the voiceprint features of speech data (e.g., the above). (Each voiceprint feature) is calculated using formula (6). The mean value of each voiceprint feature is determined and set as the target mean value.

[0192] in, The number of individual voiceprint features is less than or equal to the total number of voiceprint features in the speech data, and The voiceprint features are temporally continuous. Furthermore, the number of voiceprint features in the speech data is a positive integer greater than or equal to 1.

[0193] In this embodiment, a certain number of voiceprint features are selected from the voiceprint features of the speech data, and the mean value of these certain number of voiceprint features is determined as the target feature mean value. This reduces interference information in the voiceprint features of the speech data, thereby making the determined target feature mean value more accurate. Furthermore, based on the more accurate target feature mean value, the determined new sub-voiceprint feature set can be made more accurate, thus making the updated initial voiceprint features more accurate.

[0194] Furthermore, when the target feature mean of the speech data is determined, this target feature mean can be registered as the first voiceprint feature of a newly added sub-voiceprint feature set (which can be called the "new sub-voiceprint feature set") in the initial voiceprint feature set. For example, as mentioned above...

[0195] Then, by extracting the voiceprint features from the speech data... Voiceprint features other than individual voiceprint features and C M+1 The similarity of the feature means for C M+1 The voiceprint features in the data are updated until the update is complete, resulting in the updated initial voiceprint feature set.

[0196] For example, the initial voiceprint feature set includes sub-voiceprint feature set C and sub-voiceprint feature set D. After adding a new sub-voiceprint feature set E, the updated initial voiceprint feature set includes sub-voiceprint feature set C, sub-voiceprint feature set D, and sub-voiceprint feature set E.

[0197] For example, if the initial voiceprint feature set does not include the sub-voiceprint feature set, adding a new sub-voiceprint feature set E will update the initial voiceprint feature set to include the sub-voiceprint feature set E.

[0198] In this embodiment, when the target identifier is absent from the initial voiceprint feature set, a new sub-voiceprint feature set can be registered using the mean of the target features of the voiceprint features in the speech data. This allows the updated initial voiceprint feature set, based on the new sub-voiceprint feature set, to include the target identifier of the speech data, supplementing the speaker identity information in the updated initial voiceprint feature set. Consequently, the updated initial voiceprint feature set exhibits better generalization ability. Furthermore, registering the new sub-voiceprint feature set using the mean of the target features of the voiceprint features in the speech data enhances the generalization ability of the new sub-voiceprint feature set, further improving the generalization ability of the updated initial voiceprint feature set.

[0199] Figure 5 is another flowchart illustrating a real-time speaker differentiation method provided in an embodiment of this application.

[0200] For example, as shown in Figure 5, the method 500 includes the following implementation process:

[0201] S501, Obtain the speaker's speech data and initial voiceprint feature set.

[0202] For example, in a multi-person speaking scenario (as shown in Figure 1, a meeting scenario), the speech data of the speakers in the meeting scenario can be collected through a microphone, and the collected speech data (hereinafter referred to as "speech data") can be input into the speaker recognition model 300 as shown in Figure 3. The speaker recognition model 300 then performs speaker recognition on the speech data. At the same time, an initial voiceprint feature set can also be obtained when the speech data is acquired.

[0203] S502 performs differential processing on the azimuth of the voice data calculated from the time delay information of the voice data to obtain the differential azimuth; performs smoothing processing on the differential azimuth to obtain the smoothed azimuth; and obtains the stability evaluation value based on the smoothed azimuth.

[0204] For example, when acquiring voice data, the location of the voice data can be calculated first using the time delay information of the voice data; then the location of the voice data (i.e., D) can be... t The input is sent to the stability judgment module, which performs differential processing on the azimuth of the speech data (as shown in formula (5) above) to obtain the differential azimuth (i.e., Then, the Smoothed Z-score algorithm is used to... The orientation is smoothed by performing a smoothing process.

[0205] Furthermore, outlier detection is performed on the smoothed orientation to obtain the proportion of outliers in the smoothed orientation; then, the stability of the speech data is evaluated by using this proportion of outliers and the total number of smoothed orientations to obtain the stability evaluation value of the speech data.

[0206] S503, determine whether the stability assessment value is greater than or equal to 60%. If yes, proceed to S507; otherwise, continue to S501.

[0207] For example, when the stability assessment value of the voice data is obtained, it can be determined whether the stability assessment value is greater than or equal to a preset threshold (e.g., 60%).

[0208] S504, determine the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set.

[0209] For example, when acquiring speech data, the cosine similarity between each voiceprint feature in each sub-voiceprint feature set and the voiceprint features in the speech data can be calculated first, and then the mean similarity of at least one cosine similarity (i.e., the aforementioned sim(C)) can be calculated. m ,f t The cosine similarity between the voiceprint features of the speech data and the cosine similarity between each sub-voiceprint feature set is denoted as ).

[0210] S505, determine whether the maximum cosine similarity among multiple cosine similarities is greater than or equal to 95%. If so, proceed to S506.

[0211] For example, when obtaining the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set, the multiple cosine similarities can be sorted in ascending or descending order to obtain the maximum cosine similarity among the multiple cosine similarities. Then, it is determined whether the maximum cosine similarity is greater than or equal to 95%.

[0212] S506, determine the identifier corresponding to the maximum cosine similarity as the target identifier corresponding to the speech data; output the target identifier.

[0213] For example, when the maximum cosine similarity is ≥95% obtained through S505, it can be said that the voiceprint features in the sub-voiceprint feature set corresponding to the maximum cosine similarity have a high degree of agreement with the voiceprint features of the speech data. Therefore, the identifier corresponding to the maximum cosine similarity can be determined as the target identifier of the speech data and the target identifier can be output.

[0214] For example, when the maximum cosine similarity obtained through S505 is <95%, it can be said that the voiceprint features in the sub-voiceprint feature set corresponding to the maximum cosine similarity have a low degree of agreement with the voiceprint features of the speech data. Therefore, it can be said that the identifier corresponding to the maximum cosine similarity is not the target identifier of the speech data, that is, the speaker identity information corresponding to the speech data does not exist in the preset voiceprint feature set.

[0215] S507, Determine whether a target identifier exists in the initial voiceprint feature set. If not, proceed to S508; if yes, proceed to S510.

[0216] For example, when a stability evaluation value of ≥60% is obtained through S503 and a target identifier of the voice data is obtained through S506, it can be determined whether the target identifier exists in the initial voiceprint feature set.

[0217] For example, when the stability evaluation value obtained through S503 is <60%, it indicates that the location of the speech data changes significantly and the speech data is unstable. The speech data can be discarded, and S501 can be executed to obtain the speaker's speech data and the initial voiceprint feature set.

[0218] S508, determine the target feature mean of the top 50% of voiceprint features in the speech data; based on the target feature mean, register a new sub-voiceprint feature set.

[0219] For example, when a target identifier for which no speech data exists is obtained in the initial voiceprint feature set via S507, the top 50% of voiceprint features in the speech data are selected (e.g., the above). The target number of voiceprint features is calculated using formula (6), and the mean value of the top 50% of the voiceprint features (i.e., the target feature mean) is calculated. This target feature mean is then registered as a new sub-voiceprint feature set within the initial voiceprint feature set (e.g., the aforementioned...). The first voiceprint feature is used to register the speaker corresponding to the voice data as a new speaker.

[0220] S509, the new sub-voiceprint feature set is added to the initial voiceprint feature set to update the initial voiceprint feature set, resulting in the updated initial voiceprint feature set.

[0221] For example, when a new sub-voiceprint feature set is obtained, this new sub-voiceprint feature set can also be added to the initial voiceprint feature set to update the initial voiceprint feature set, resulting in an updated initial voiceprint feature set. Furthermore, when the updated initial voiceprint feature set is obtained, it can be used to replace the initial voiceprint feature set in S501.

[0222] S510, determine the target sub-voiceprint feature set corresponding to the target identifier in the initial voiceprint feature set.

[0223] For example, when a target identifier containing speech data is obtained in the initial voiceprint feature set via S507, a target sub-voiceprint feature set corresponding to the target identifier can be determined in the initial voiceprint feature set (e.g., the sub-voiceprint feature set is C).

[0224] S511, determine whether the number of voiceprint features in the target sub-voiceprint feature set is less than 10. If yes, execute S512; otherwise, continue to execute S514.

[0225] For example, when the target sub-voiceprint feature set corresponding to the target identifier is obtained, the number of voiceprint features in the target sub-voiceprint feature set can be obtained. And it can be determined whether the number is less than 10 (i.e., the preset number mentioned above).

[0226] S512, add the voiceprint features of the speech data to the target sub-voiceprint feature set to obtain the updated target sub-voiceprint feature set.

[0227] For example, when the number obtained through S511 is less than 10, the voiceprint features of the speech data can be directly added to the target sub-voiceprint feature set to obtain the updated target sub-voiceprint feature set.

[0228] S513, the initial voiceprint feature set is updated based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0229] For example, when the updated target sub-voiceprint feature set is obtained, the initial voiceprint feature set is also updated accordingly, resulting in an updated initial voiceprint feature set. Furthermore, when the updated initial voiceprint feature set is obtained, it can be used to replace the initial voiceprint feature set in S501.

[0230] S514, determine the maximum similarity value between the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set and the mean value of the voiceprint features in the target sub-voiceprint feature set.

[0231] For example, when the number of voiceprint features obtained through S511 is ≥10, the maximum similarity value between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set, and the average feature value of the voiceprint features in the target sub-voiceprint feature set can be determined.

[0232] S515, the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set, excluding the maximum similarity value, are determined as the updated voiceprint features of the target sub-voiceprint feature set.

[0233] For example, when the maximum similarity value is determined, the voiceprint features of the speech data and the voiceprint features of the target sub-voiceprint feature set other than the maximum similarity value are determined as the updated voiceprint features of the target sub-voiceprint feature set.

[0234] Furthermore, when the updated target sub-voiceprint feature set is obtained, the initial voiceprint feature set is also updated accordingly, resulting in an updated initial voiceprint feature set. Moreover, when the updated initial voiceprint feature set is obtained, it can replace the initial voiceprint feature set in S501.

[0235] In summary, to improve the accuracy of speaker recognition results when performing speaker identification on speech data, a stability assessment value can be obtained from the location of the speech data. When the stability assessment value indicates that the location of the speech data is relatively stable, the pre-registered voiceprint feature set (i.e., the initial voiceprint feature set) is updated using the target identifier obtained after voiceprint recognition and the voiceprint features of the speech data. This makes the updated initial voiceprint feature set more accurate. Furthermore, using the more accurate updated initial voiceprint feature set for speaker identification improves the accuracy and stability of the speaker identification results. Moreover, speaker identification on speech data is performed by calculating the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set to obtain the speaker identifier. Compared to using clustering to obtain speaker identity information, the identification process is simpler, requires less computational resources, has better universality, and is easier to widely apply.

[0236] It should be noted that all steps in Figure 5 are described in detail in the corresponding embodiments in Figures 2 to 4, and will not be repeated here.

[0237] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values ​​or scenarios exemplified. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of this application.

[0238] The real-time speaker differentiation method provided by the embodiments of this application has been described in detail above with reference to Figures 1 to 5; the device embodiments of this application will be described in detail below with reference to Figures 6 and 7. It should be understood that the device in the embodiments of this application can execute the various methods of the foregoing embodiments of this application, that is, the specific working process of the various products below can be referred to the corresponding process in the foregoing method embodiments.

[0239] Figure 6 is a schematic diagram of the structure of the real-time speaker differentiation device provided in the embodiment of this application.

[0240] For example, as shown in FIG6, the real-time speaker differentiation device 600 is configured in an electronic device and includes:

[0241] The acquisition module 610 is used to acquire the speaker's voice data and initial voiceprint feature set;

[0242] The determination module 620 is used to determine the location of the voice data based on the time delay information of the voice data.

[0243] The processing module 630 is used to obtain a stability evaluation value of the speech data based on the orientation of the speech data; wherein the stability evaluation value is negatively correlated with the change in the orientation of the speech data.

[0244] The recognition module 640 is used to perform voiceprint recognition on the voice data based on the voiceprint features of the voice data and the cosine similarity of the initial voiceprint feature set, and output the target identifier; wherein, the target identifier is used to represent the speaker's identity information;

[0245] The update module 650 is used to update the initial voiceprint feature set based on the voiceprint features of the target identifier and the speech data when the stability evaluation value is greater than or equal to a preset threshold, so as to obtain the updated initial voiceprint feature set.

[0246] In one possible implementation, the update module 650 is specifically used for:

[0247] Determine whether a target identifier exists in the initial voiceprint feature set;

[0248] When a target identifier exists in the initial voiceprint feature set, the target sub-voiceprint feature set corresponding to the target identifier is determined in the initial voiceprint feature set; based on the number of voiceprint features in the target sub-voiceprint feature set and the voiceprint features of the speech data, the initial voiceprint feature set is updated to obtain the updated initial voiceprint feature set.

[0249] When there is no target identifier in the initial voiceprint feature set, a new sub-voiceprint feature set is registered based on the voiceprint features of the speech data; the initial voiceprint feature set is updated based on the new sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0250] In one possible implementation, the update module 650 is specifically used for:

[0251] When the number is less than the preset number, the voiceprint features of the speech data are added to the target sub-voiceprint feature set to obtain the updated target sub-voiceprint feature set; the initial voiceprint feature set is updated based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0252] When the quantity is greater than or equal to the preset quantity, the initial voiceprint feature set is updated based on the similarity value between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set and the first voiceprint feature, respectively, to obtain the updated initial voiceprint feature set.

[0253] Among them, the first voiceprint feature represents the feature mean of the voiceprint features in the target sub-voiceprint feature set.

[0254] In one possible implementation, the similarity value between the voiceprint features of the aforementioned speech data and the first voiceprint features is a first similarity value, and the similarity value between the voiceprint features of the aforementioned target sub-voiceprint feature set and the first voiceprint features is a second similarity value; the update module 650 is specifically used for:

[0255] Determine the second voiceprint feature corresponding to the maximum similarity value between the first and second similarity values;

[0256] The voiceprint features and the third voiceprint features of the speech data are determined as the voiceprint features of the updated target sub-voiceprint feature set; the initial voiceprint feature set is updated based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

[0257] Among them, the third voiceprint feature represents the voiceprint features in the target sub-voiceprint feature set other than the second voiceprint feature.

[0258] In one possible implementation, the update module 650 is specifically used for:

[0259] Determine the target feature mean of the speakerprint features in the speech data;

[0260] Register a new sub-voiceprint feature set based on the mean of the target features.

[0261] In one possible implementation, the update module 650 is specifically used for:

[0262] Determine the target number of voiceprint features from the voiceprint features of the speech data; wherein the target number is less than or equal to the total number of voiceprint features in the speech data.

[0263] The mean value of the voiceprint features of the target data is determined as the mean value of the target features.

[0264] In one possible implementation, the determining module 620 is specifically used for:

[0265] Determine the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set; wherein, the preset voiceprint feature set includes all sub-voiceprint feature sets in the initial voiceprint feature set;

[0266] Determine whether the maximum cosine similarity among multiple cosine similarities is greater than or equal to a preset similarity.

[0267] When the maximum cosine similarity is greater than or equal to the preset similarity, the identifier corresponding to the maximum cosine similarity is determined as the target identifier corresponding to the speech data, and the target identifier is output.

[0268] In one possible implementation, the processing module 630 is specifically used for:

[0269] The location of the speech data is differentially processed to obtain the differential location.

[0270] The differential orientation is smoothed to obtain the smoothed orientation.

[0271] Based on the smoothed orientation, a stability evaluation value is obtained.

[0272] It should be noted that the aforementioned device 600 is embodied in the form of a functional module. The term "module" here can be implemented in software and / or hardware, without specific limitations.

[0273] For example, a "module" can be a software program, hardware circuit, or a combination of both that implements the above functions. Hardware circuits may include application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or combined processors) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0274] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0275] Figure 7 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application.

[0276] For example, as shown in FIG7, the electronic device 700 includes a memory 710 and a processor 720, wherein the memory 710 stores executable program code 7101, and the processor 720 is used to call and execute the executable program code 7101 to perform a real-time speaker differentiation method.

[0277] This application can divide electronic devices into functional modules based on the above method examples. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0278] When each functional module is divided according to its corresponding function, the electronic device may include: an acquisition module, a determination module, a processing module, an identification module, and an update module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0279] The electronic device provided in this application is used to execute the above-described real-time speaker differentiation method, and thus can achieve the same effect as the above-described implementation method.

[0280] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of relevant program code and data by the electronic device.

[0281] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.

[0282] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described in the foregoing embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives, magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0283] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a real-time speaker differentiation method as described in the above embodiments.

[0284] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or module. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip execute a real-time speaker differentiation method in the above embodiments.

[0285] The electronic devices, computer-readable storage media, computer program products or chips provided in this application are all used to perform the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0286] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0287] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0288] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A real-time speaker differentiation method, wherein, Applied to electronic devices, the method includes: Obtain the speaker's voice data and initial voiceprint feature set; Based on the time delay information of the voice data, the location of the voice data is determined; Based on the location of the speech data, a stability evaluation value for the speech data is obtained; wherein, the stability evaluation value is negatively correlated with the amount of change in the location of the speech data; Based on the cosine similarity between the voiceprint features of the speech data and the initial voiceprint feature set, voiceprint recognition is performed on the speech data, and a target identifier is output; wherein, the target identifier is used to represent the speaker's identity information; When the stability evaluation value is greater than or equal to a preset threshold, the initial voiceprint feature set is updated based on the target identifier and the voiceprint features of the speech data to obtain the updated initial voiceprint feature set.

2. The method according to claim 1, wherein, The process of updating the initial voiceprint feature set based on the target identifier and the voiceprint features of the speech data to obtain an updated initial voiceprint feature set includes: Determine whether the target identifier exists in the initial voiceprint feature set; When the target identifier exists in the initial voiceprint feature set, a target sub-voiceprint feature set corresponding to the target identifier is determined in the initial voiceprint feature set; based on the number of voiceprint features in the target sub-voiceprint feature set and the voiceprint features of the speech data, the initial voiceprint feature set is updated to obtain the updated initial voiceprint feature set. When the target identifier is not present in the initial voiceprint feature set, a new sub-voiceprint feature set is registered based on the voiceprint features of the speech data; the initial voiceprint feature set is updated based on the new sub-voiceprint feature set to obtain the updated initial voiceprint feature set.

3. The method according to claim 2, wherein, The update process, which involves applying the number of voiceprint features in the target sub-voiceprint feature set and the voiceprint features in the speech data to the initial voiceprint feature set to obtain the updated initial voiceprint feature set, includes: When the number is less than the preset number, the voiceprint features of the speech data are added to the target sub-voiceprint feature set to obtain an updated target sub-voiceprint feature set; the updated target sub-voiceprint feature set is then used to update the initial voiceprint feature set to obtain the updated initial voiceprint feature set. When the quantity is greater than or equal to the preset quantity, the initial voiceprint feature set is updated based on the similarity value between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set and the first voiceprint feature, respectively, to obtain the updated initial voiceprint feature set. Wherein, the first voiceprint feature represents the feature mean of the voiceprint features in the target sub-voiceprint feature set.

4. The method according to claim 3, wherein, The similarity value between the voiceprint features of the speech data and the first voiceprint features is a first similarity value, and the similarity value between the voiceprint features of the target sub-voiceprint feature set and the first voiceprint features is a second similarity value; the update processing is performed on the initial voiceprint feature set based on the similarity values ​​between the voiceprint features of the speech data, the voiceprint features of the target sub-voiceprint feature set, and the first voiceprint features, to obtain the updated initial voiceprint feature set, including: Determine the second voiceprint feature corresponding to the maximum similarity value between the first similarity value and the second similarity value; The voiceprint features and the third voiceprint features of the speech data are determined as the voiceprint features of the updated target sub-voiceprint feature set; the updated voiceprint feature set is then processed based on the updated target sub-voiceprint feature set to obtain the updated initial voiceprint feature set. The third voiceprint feature refers to the voiceprint features in the target sub-voiceprint feature set other than the second voiceprint feature.

5. The method according to any one of claims 2 to 4, wherein, Registering a new sub-voiceprint feature set based on the voiceprint features of the speech data includes: Determine the target feature mean of the voiceprint features of the speech data; Based on the mean of the target features, register the new sub-voiceprint feature set.

6. The method according to claim 5, wherein, Determining the target feature mean of the voiceprint features of the speech data includes: A target number of voiceprint features are determined from the voiceprint features of the speech data; wherein the target number is less than or equal to the total number of voiceprint features of the speech data; The mean value of the voiceprint features of the target data is determined as the mean value of the target features.

7. The method according to any one of claims 1 to 4, wherein, The process of obtaining a stability evaluation value for the voice data based on its location includes: The location of the voice data is differentially processed to obtain the differential location. The differential orientation is smoothed to obtain a smoothed orientation. The stability evaluation value is obtained based on the smoothed orientation.

8. The method according to any one of claims 1 to 4, wherein, The method of performing voiceprint recognition on the voice data based on the cosine similarity between the voiceprint features of the voice data and the initial voiceprint feature set, and outputting a target identifier, includes: Determine the cosine similarity between the voiceprint features of the speech data and each sub-voiceprint feature set in the preset voiceprint feature set; wherein, the preset voiceprint feature set includes all sub-voiceprint feature sets in the initial voiceprint feature set; Determine whether the maximum cosine similarity among multiple cosine similarities is greater than or equal to a preset similarity. When the maximum cosine similarity is greater than or equal to the preset similarity, the identifier corresponding to the maximum cosine similarity is determined as the target identifier corresponding to the speech data, and the target identifier is output.

9. A real-time speaker differentiation device, wherein, Configured in an electronic device, the device includes: The acquisition module is used to acquire the speaker's voice data and initial voiceprint feature set; The determination module is used to determine the location of the voice data based on the time delay information of the voice data; The processing module is used to obtain a stability evaluation value of the voice data based on the location of the voice data; wherein the stability evaluation value is negatively correlated with the amount of change in the location of the voice data; The recognition module is used to perform voiceprint recognition on the voice data based on the voiceprint features of the voice data and the cosine similarity of the initial voiceprint feature set, and output a target identifier; wherein the target identifier is used to represent the speaker's identity information; The update module is used to update the initial voiceprint feature set based on the target identifier and the voiceprint features of the speech data when the stability evaluation value is greater than or equal to a preset threshold, so as to obtain the updated initial voiceprint feature set.

10. An electronic device, wherein, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 8.