Speaker Log Generation Method, Device, Computer Equipment and Readable Storage Medium
By sharpening and clustering the speech signal similarity matrix, a speaker log that meets the preset conditions is generated, which solves the problem that background noise and unknown number of speakers affect the speaker log accuracy in the prior art, and achieves higher speaker log accuracy.
Patent Information
- Application Number
- CN202210124036.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-02-10
AI Technical Summary
In the prior art, when generating speaker logs, due to the influence of background noise and unknown number of speakers, the speaker labels obtained by clustering are of poor accuracy, which in turn affects the accuracy of speaker logs.
By obtaining the similarity matrix corresponding to the speech signal and sharpening it with the first preset parameter, a target clustering matrix that meets the preset clustering conditions is obtained. Then the target clustering matrix is clustered according to the number of speakers, and the speaker tag is obtained, and the voice signal is finally divided according to the tag to generate the speaker log.
Effectively avoid the influence of background noise and unknown number of speakers on clustering results, and improve the accuracy of speaker logs.
Smart Images

Figure CN114446284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech detection, and more particularly, to a method, apparatus, computer device, and readable storage medium for generating a speaker log. Background Art
[0002] Speaker logging is a common task in the field of speech detection. Different from the speech recognition task of judging 'whensays what', the speaker logging task needs to judge 'who says when'. Speaker logging differentiates the speaking segments of multiple speakers in a speech signal, and its accuracy is a prerequisite for the success of many speech detection tasks. For example, in the single-channel customer service / customer recording quality inspection task, it is first necessary to identify the speech segments of the customer service, then identify the content of the speech segments, and finally make a violation judgment. The existing speaker log generation technology directly clusters the similarity matrix generated based on the speech signal to obtain speaker labels, and then uses the speaker labels to segment the speech signal to obtain the corresponding speaker log. Affected by background noise and the situation of 'unknown number of speakers', the accuracy of the speaker labels obtained by clustering is poor, which in turn affects the accuracy of the speaker log. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art, embodiments of the present invention provide a method, apparatus, computer device, and readable storage medium for generating a speaker log. The specific solutions are as follows:
[0004] In a first aspect, embodiments of the present invention provide a method for generating a speaker log, the method comprising:
[0005] Obtaining a similarity matrix corresponding to a speech signal;
[0006] Determining a target clustering matrix and the number of speakers according to a first preset parameter and the similarity matrix, wherein the target clustering matrix is obtained by sharpening the similarity matrix, and the first preset parameter constrains the sharpening process, and the target clustering matrix obtained under the constraint of the first preset parameter satisfies a preset clustering condition;
[0007] Clustering the target clustering matrix according to the number of speakers to obtain speaker labels;
[0008] Segmenting the speech signal according to the speaker labels to generate a speaker log.
[0009] In a possible implementation manner, the step of obtaining a similarity matrix corresponding to a speech signal includes:
[0010] Dividing the speech signal into multiple speech segments;
[0011] Input each of the speech segments into a voiceprint detection model to obtain a speaker feature vector corresponding to each speech segment;
[0012] Obtain a similarity matrix corresponding to the language signal based on all the speaker feature vectors.
[0013] In a possible implementation, the step of obtaining a similarity matrix corresponding to the language signal based on all the speaker feature vectors includes:
[0014] Calculate a similarity coefficient between each speaker feature vector and other speaker feature vectors;
[0015] Obtain a similarity matrix corresponding to the language signal based on all the similarity coefficients.
[0016] In a possible implementation, the step of determining a target clustering matrix and the number of speakers based on the first preset parameter and the similarity matrix includes:
[0017] Use a plurality of second preset parameters to respectively constrain the sharpening process of the similarity matrix to obtain a plurality of to-be-determined clustering matrices, where the second preset parameter is determined by the dimension of the similarity matrix;
[0018] Determine a target clustering matrix from the plurality of to-be-determined clustering matrices according to the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing, and use the second preset parameter corresponding to the target clustering matrix as the first preset parameter;
[0019] Determine the number of speakers according to the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing.
[0020] In a possible implementation, the step of determining a target clustering matrix from all the to-be-determined clustering matrices according to the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing includes:
[0021] For each to-be-determined clustering matrix, perform a difference processing on the eigenvalues determined by it to obtain a maximum eigenvalue difference corresponding to the to-be-determined clustering matrix, where the maximum eigenvalue difference characterizes the clustering ease of the to-be-determined clustering matrix;
[0022] Determine a target clustering matrix from the plurality of to-be-determined clustering matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each to-be-determined clustering matrix.
[0023] In a possible implementation, the step of determining the target clustering matrix from multiple said pending clustering matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each said pending clustering matrix includes:
[0024] For each said pending clustering matrix, calculate the ratio of the corresponding second preset parameter to the maximum eigenvalue difference;
[0025] Take the pending clustering matrix with the minimum ratio as the target clustering matrix.
[0026] In a possible implementation, the step of segmenting the speech signal according to the speaker label to generate a speaker log includes:
[0027] Train a speaker recognition model according to the speaker label;
[0028] Input the speech signal into the speaker recognition model to obtain the speaker log.
[0029] In a second aspect, an embodiment of the present invention provides a speaker log generation device, and the device includes:
[0030] An acquisition module, configured to acquire a similarity matrix corresponding to a speech signal;
[0031] A determination module, configured to determine a target clustering matrix and the number of speakers according to a first preset parameter and the similarity matrix, wherein the target clustering matrix is obtained by sharpening the similarity matrix, and the first preset parameter constrains the sharpening process, and the target clustering matrix obtained under the constraint of the first preset parameter satisfies a preset clustering condition;
[0032] A clustering module, configured to cluster the target clustering matrix according to the number of speakers to obtain speaker labels;
[0033] A generation module, configured to segment the speech signal according to the speaker labels to generate a speaker log.
[0034] In a third aspect, an embodiment of the present invention provides a computer device, and the computer device includes: a memory and a processor, where the memory is used to store a computer program; the processor is configured to execute the speaker log generation method as described in the first aspect when calling the computer program.
[0035] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the speaker log generation method as described in the first aspect.
[0036] Compared with the prior art, a method, device, computer device and readable storage medium for generating a speaker log provided by an embodiment of the present invention first obtains a similarity matrix corresponding to a voice signal; then, according to a first preset parameter and the similarity matrix, determines a target clustering matrix and the number of speakers, where the target clustering matrix is obtained by sharpening the similarity matrix, and the first preset parameter constrains the sharpening process, and the target clustering matrix obtained under the constraint of the first preset parameter satisfies a preset clustering condition; then, clusters the target clustering matrix according to the number of speakers to obtain speaker labels; finally, segments the voice signal according to the speaker labels to generate a speaker log. Since the embodiment of the present invention uses the first preset parameter to constrain the sharpening process of the similarity matrix, obtains a target clustering matrix that satisfies the preset clustering condition, and then clusters the target clustering matrix, thereby avoiding the influence of background noise and the situation of "unknown number of speakers" on the accuracy of the speaker labels obtained by clustering, and further improving the accuracy of the speaker log. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic diagram of an example of a speaker log provided by an embodiment of the present invention;
[0039] Figure 2 It is a flowchart of a method for generating a speaker log provided by an embodiment of the present invention;
[0040] Figure 3 It is a flowchart of a method for obtaining a similarity matrix corresponding to a voice signal provided by an embodiment of the present invention;
[0041] Figure 4 It is a flowchart of another method for obtaining a similarity matrix corresponding to a voice signal provided by an embodiment of the present invention;
[0042] Figure 5 It is a flowchart of a method for determining a target clustering matrix and the number of speakers provided by an embodiment of the present invention;
[0043] Figure 6 It is a flowchart of another method for determining a target clustering matrix and the number of speakers provided by an embodiment of the present invention;
[0044] Figure 7Schematic flowchart of a method for generating a speaker log based on speaker tags provided by an embodiment of the present invention;
[0045] Figure 8 Block diagram of a speaker log generation device provided by an embodiment of the present invention;
[0046] Figure 9 Schematic block diagram of the structure of a computer device provided by an embodiment of the present invention.
[0047] Icons: 200 - Speaker log generation device; 201 - Acquisition module; 202 - Determination module; 203 - Clustering module; 204 - Generation module; 300 - Computer device; 310 - Memory; 320 - Processor. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein generally may be arranged and designed in a variety of different configurations.
[0049] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0050] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0051] In addition, terms such as "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance.
[0052] It should be noted that the features in the embodiments of the present invention may be combined with each other without conflict.
[0053] Speaker Diarization, also known as speaker separation, voiceprint segmentation and clustering, such as Figure 1As shown, it detects the start and end time points of different speakers' speeches from a continuous multi-speaker speech signal, segments the segments of each speaker, and solves the problem of "who says when". Based on the speaker log, the structured management of the audio data stream can be completed, which has wide application value.
[0054] In a possible implementation, the speaker log generation technology directly clusters the similarity matrix generated based on the speech signal to obtain speaker labels, and then uses the speaker labels to segment the speech signal to obtain the corresponding speaker log. Due to the existence of background noise in the speech signal and the situation of unknown number of speakers, the clustering effect of the similarity matrix is poor, and the accuracy of the speaker labels obtained by clustering is not high, which in turn affects the accuracy of the speaker log.
[0055] In view of this, the embodiments of the present invention provide a speaker log generation method, device, computer device and readable storage medium to avoid the background noise and the situation of "unknown number of speakers" from affecting the accuracy of the speaker labels obtained by clustering, thereby improving the accuracy of the speaker log, which will be described in detail below.
[0056] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a speaker log generation method provided by the embodiments of the present invention. The speaker log generation method described in the embodiments of the present invention includes steps S101 to S104.
[0057] S101, obtain the similarity matrix corresponding to the speech signal.
[0058] In the embodiments of the present invention, the speech signal can be extracted from audio data (such as the recording of a telephone conference, etc.) that records the speech content of multiple speakers speaking alternately. Generally, the audio data is input into a voice activity detection model (Voice Activity Detection, VAD), and all the outputs of the VAD model are spliced to obtain the speech signal. Among them, the VAD model can perform voice activity detection on the audio data, and divide the audio data frame by frame into two categories: speech with someone speaking and speech without someone speaking. The speech without someone speaking includes, but is not limited to, pure silence, ambient sound, music, etc., and the speech signal belongs to the speech with someone speaking.
[0059] A voice signal can be regarded as composed of multiple voice segments. The value of any element in the similarity matrix corresponding to the voice signal represents the possibility that two voice segments are from the same speaker. Ideally, when two voice segments are from the same speaker, the value of the corresponding element in the similarity matrix is 1, and when two voice segments are not from the same speaker, the value of the corresponding element in the similarity matrix is 0. However, due to the existence of background noise, the values of the elements in the similarity matrix are distributed in the interval (0, 1).
[0060] S102. Determine a target clustering matrix and the number of speakers according to a first preset parameter and the similarity matrix.
[0061] Among them, the target clustering matrix is obtained by sharpening the similarity matrix. The first preset parameter constrains the sharpening process. The target clustering matrix obtained under the constraint of the first preset parameter meets the preset clustering conditions.
[0062] In the embodiment of the present invention, sharpening the similarity matrix means retaining the values of several elements in the matrix and setting the values of the remaining elements to 0. The implementation method can be to retain the values of the larger several elements in each row of the similarity matrix and set the values of the remaining elements to 0. It can be understood that the result of sharpening the similarity matrix is not unique. For example, for the similarity matrix Performing sharpening processing can obtain Three results, which respectively correspond to "retaining the value of the larger 1 element in each row of elements and setting the values of the remaining elements to 0", "retaining the values of the larger 2 elements in each row of elements and setting the values of the remaining elements to 0", and "retaining the values of the larger 3 elements in each row of elements and setting the values of the remaining elements to 0".
[0063] Since some of the values of the elements in the similarity matrix after sharpening processing are 0, it is possible to directly know the voice segments that are not from the same speaker based on these elements with a value of 0. That is, the process of sharpening the similarity matrix can be regarded as a process of removing the interference of background noise. However, the corresponding clustering accuracy and clustering ease of different sharpening results are different. It is manifested that the fewer the elements with values retained in each row of the similarity matrix and the more the elements with values set to 0, the larger the corresponding number of clusters. Since the number of clusters is proportional to the clustering accuracy and inversely proportional to the clustering ease, the sharpening result that can ensure both clustering accuracy and clustering ease needs to be used as the target clustering matrix.
[0064] The first preset parameter refers to the number of elements with values retained in each row of the target clustering matrix. Using the first preset parameter to constrain the process of sharpening the similarity matrix, a target clustering matrix is obtained, and then the number of speakers is determined based on the target clustering matrix.
[0065] S103. Cluster the target clustering matrix according to the number of speakers to obtain speaker labels.
[0066] In an embodiment of the present invention, the K-Means clustering algorithm can be selected to cluster the target clustering matrix. When using the K-Means clustering algorithm, the initial clustering center points are selected according to the number of speakers. After multiple iterations, the speech segments from different speakers are classified into the corresponding speakers, and the corresponding speaker identities are marked to obtain speaker labels. That is, according to the speaker labels of the speech segments, it can be known which speaker the speech segment corresponds to.
[0067] S104. Segment the speech signal according to the speaker labels to generate a speaker log.
[0068] In an embodiment of the present invention, performing steps S101 to S103 is equivalent to performing a first segmentation on the speech information. However, the segmentation result may be ambiguous near the segmentation line. It is necessary to perform a more refined segmentation on the speech signal according to the result of the first segmentation. The training sample data is obtained from the speech signal according to the speaker labels, and then the speech signal is segmented again using the model obtained based on the training sample data to obtain a segmentation result with less ambiguity as the speaker log.
[0069] The beneficial effect of the method provided in the above embodiment of the present invention is that the sharpening process of the similarity matrix is constrained by the first preset parameter to obtain a target clustering matrix that meets the preset clustering conditions, and then the target clustering matrix is clustered, thereby avoiding the influence of background noise on the accuracy of the speaker labels obtained by clustering, and further improving the accuracy of the speaker log.
[0070] Based on Figure 2 , a specific implementation manner for obtaining the similarity matrix corresponding to the speech signal is provided in an embodiment of the present invention. Please refer to Figure 3 , Figure 3 is a schematic flowchart of a method for obtaining the similarity matrix corresponding to the speech signal provided in an embodiment of the present invention. Step S101 includes sub-steps S101-1 to S101-3.
[0071] S101-1. Divide the speech signal into multiple speech segments.
[0072] In an embodiment of the present invention, the speech signal can be segmented according to a fixed segmentation duration to obtain multiple speech segments, and there is an overlapping part between the speech segments. For example, the segmentation duration can be 1.5 seconds, that is, the length of each speech segment is 1.5 seconds, and the overlapping part between two adjacent speech segments is 0.5 s.
[0073] S101-2. Input each speech segment into the voiceprint detection model to obtain the speaker feature vector corresponding to each speech segment.
[0074] In the embodiments of the present invention, voices of different speakers can present different voiceprint features on the voiceprint spectrogram based on physical attributes of the voice (such as voice quality, voice length, voice intensity, and pitch, etc.). A speaker feature vector refers to a vector that can represent the voiceprint features of a speaker's voice.
[0075] A voiceprint detection model can refer to a model used to extract speaker feature vectors in a speech segment. Specifically, it can include, but is not limited to, the I-Vector model and the X-Vector model.
[0076] S101-3. Obtain a similarity matrix corresponding to the language signal according to all speaker feature vectors.
[0077] In the embodiments of the present invention, by quantifying the similarity degree of the voiceprint features represented by the speaker feature vectors corresponding to each speech segment, the values of the elements in the similarity matrix are obtained. For example, a speech signal is divided into speech segment 1, speech segment 2, and speech segment 3. Among them, by quantifying the similarity degree of the voiceprint features represented by the speaker feature vectors corresponding to speech segment 1 and speech segment 3, and the quantization value is 0.6, then the value of the element in the first row and the third column of the similarity matrix is 0.6, that is, there is a 60% possibility that speech segment 1 and speech segment 3 come from the same speaker.
[0078] Based on Figure 3 , the embodiments of the present invention provide a specific implementation manner for determining a similarity matrix corresponding to a language signal according to speaker feature vectors. Please refer to Figure 4 , Figure 4 is a schematic flowchart of another method for obtaining a similarity matrix corresponding to a speech signal provided by the embodiments of the present invention. Sub-step S101-3 includes sub-steps S101-3-1 to S101-3-2.
[0079] S101-3-1. Calculate the similarity coefficient between each speaker feature vector and other speaker feature vectors.
[0080] In the embodiments of the present invention, the similarity degree of the voiceprint features represented by each speaker feature vector can be quantified by calculating the similarity coefficient between each speaker feature vector and other speaker feature vectors. Among them, the calculation methods of the similarity coefficient include the Probabilistic Linear Discriminant Analysis (PLDA) and the cosine distance method.
[0081] S101-3-2. Obtain a similarity matrix corresponding to the language signal according to all similarity coefficients.
[0082] In an embodiment of the present invention, the similarity matrix is an n-order matrix, and the dimension n is determined by the number of speech segments. For example, if a speech signal is divided into speech segment 1, speech segment 2, and speech segment 3, the similarity matrix corresponding to the speech signal is a 3-order matrix. Among them, the value of the element in the first row and the second column is the similarity coefficient between the speaker feature vector corresponding to speech segment 1 and the speaker feature vector corresponding to speech segment 2, and the value of the element in the second row and the third column is the similarity coefficient between the speaker feature vector corresponding to speech segment 2 and the speaker feature vector corresponding to speech segment 3.
[0083] Based on Figure 2 , an embodiment of the present invention provides a specific implementation manner for determining the target clustering matrix and the number of speakers. Please refer to Figure 5 , Figure 5 is a schematic flowchart of a method for determining the target clustering matrix and the number of speakers provided by an embodiment of the present invention. Step S102 includes sub-steps S102-1 to S102-3.
[0084] S102-1, using multiple second preset parameters to respectively constrain the sharpening process of the similarity matrix, and obtaining multiple to-be-determined clustering matrices. The second preset parameters are determined by the dimension of the similarity matrix.
[0085] In an embodiment of the present invention, the second preset parameter refers to the number of elements whose values are retained in each row of the matrix when sharpening the similarity matrix. According to the dimension of the similarity matrix, the number of second preset parameters can be determined, and the second preset parameter is greater than 1 and less than the dimension of the similarity matrix. For example, if the similarity matrix is a 5-order matrix, that is, the dimension is 5, the second preset parameters are 2, 3, and 4. The processes of sharpening the similarity matrix are respectively constrained by different second preset parameters to obtain multiple to-be-determined clustering matrices. The elements with retained values and the elements with values of 0 in different to-be-determined clustering matrices are different, and the corresponding clustering accuracy and clustering ease are also different.
[0086] S102-2, according to the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing, determining the target clustering matrix from multiple to-be-determined clustering matrices, and using the second preset parameter corresponding to the target clustering matrix as the first preset parameter.
[0087] In an embodiment of the present invention, it is necessary to select a to-be-determined clustering matrix that can ensure both clustering accuracy and clustering ease from multiple to-be-determined clustering matrices as the target clustering matrix. The clustering accuracy and clustering ease of the to-be-determined clustering matrix can be judged by obtaining the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing. For example, a to-be-determined clustering matrix obtained by sharpening the similarity matrix is Perform binarization and symmetrization processing on the to-be-determined clustering matrix, and the result obtained is The Laplacian matrix corresponding to this result is Then, based on the three eigenvalues -1, -0.36, and 1.36 of this Laplacian matrix, judge the to-be-determined clustering matrix in terms of clustering accuracy and clustering ease, and traverse the similarity matrix corresponding to all to-be-determined clustering matrices, determine the clustering accuracy and clustering ease of each to-be-determined clustering matrix, and use the second preset parameter corresponding to the to-be-determined clustering matrix selected as the target clustering matrix as the first preset parameter.
[0088] S102-3, Determine the number of speakers according to the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing.
[0089] In the embodiments of the present invention, since when using the K-Means clustering method to obtain speaker labels, the initial clustering center points need to be selected according to the number of speakers. If the "number of speakers" is unknown, the number of speakers can be determined by obtaining the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing. Specifically, it can be achieved according to the maximum difference between eigenvalues. For example, the target clustering matrix is The eigenvalues determined by the Laplacian matrix corresponding to the result after binarization and symmetrization processing of the target clustering matrix are arranged from small to large as (-1, -0.36, 1.36). After normalizing them and taking the difference, the result obtained is (0.47, 1.26). It can be seen that the maximum difference between eigenvalues is 1.26, which is the difference between the second eigenvalue and the third eigenvalue. Therefore, the number of speakers can be determined to be 2.
[0090] Based on Figure 5 , the embodiments of the present invention provide a specific implementation manner for determining the target clustering matrix from multiple to-be-determined clustering matrices. Please refer to Figure 6 , Figure 6 is a flowchart of another method for determining the target clustering matrix and the number of speakers provided by the embodiments of the present invention. Sub-step S102-2 includes sub-steps S102-2-1 to S102-2-2.
[0091] S102-2-1, For each to-be-determined clustering matrix, perform difference processing on the eigenvalues determined by it to obtain the maximum eigenvalue difference corresponding to this to-be-determined clustering matrix. The maximum eigenvalue difference characterizes the clustering ease of this to-be-determined clustering matrix.
[0092] In an embodiment of the present invention, the maximum eigenvalue difference refers to the maximum difference between the eigenvalues of the Laplacian matrix corresponding to the to-be-clustered matrix after binarization and symmetrization processing. The value of the maximum eigenvalue difference is directly proportional to the clustering ease of the to-be-clustered matrix, and the greater the clustering ease of the to-be-clustered matrix.
[0093] S102-2-2. Determine a target clustering matrix from multiple to-be-clustered matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each to-be-clustered matrix.
[0094] In an embodiment of the present invention, when sharpening the similarity matrix, the fewer the elements with the middle values retained in each row of the similarity matrix and the more the elements with the values set to 0, the greater the corresponding number of clusters. And the number of clusters is directly proportional to the clustering accuracy. It can be understood that the second preset parameter is inversely proportional to the clustering accuracy of the to-be-clustered matrix. The to-be-clustered matrix that can ensure both clustering accuracy and clustering ease can be determined according to the second preset parameter and the maximum eigenvalue difference corresponding to each to-be-clustered matrix as the target clustering matrix.
[0095] Specifically, the implementation process of sub-step S102-2-2 is as follows:
[0096] First, for each to-be-clustered matrix, calculate the ratio of the corresponding second preset parameter to the maximum eigenvalue difference.
[0097] Then, use the to-be-clustered matrix with the smallest ratio as the target clustering matrix.
[0098] In an embodiment of the present invention, since the second preset parameter is inversely proportional to the clustering accuracy of the to-be-clustered matrix, the maximum eigenvalue difference is directly proportional to the clustering ease, and the clustering effect of the to-be-clustered matrix is affected by the clustering accuracy and the clustering ease. The second preset parameter can be used as the numerator and the maximum eigenvalue difference as the denominator to calculate the ratio of the second preset parameter to the maximum eigenvalue difference as an index for evaluating the clustering effect of the to-be-clustered matrix. The smaller the index value, the better the clustering effect. Therefore, the to-be-clustered matrix with the best clustering effect (i.e., the smallest index value) is used as the target clustering matrix. For example, the values of the second preset parameter determined according to the dimension of the similarity matrix are 2 and 3. When the second preset parameter is 2, the maximum eigenvalue difference of the corresponding to-be-clustered matrix is 1.26, and the ratio of the second preset parameter to the maximum eigenvalue difference is 1.58. When the second preset parameter is 3, the maximum eigenvalue difference of the corresponding to-be-clustered matrix is 1.5, and the ratio of the second preset parameter to the maximum eigenvalue difference is 2. Since 1.58 is less than 2, the to-be-clustered matrix corresponding to the second preset parameter of 2 can be used as the target clustering matrix.
[0099] Based on Figure 2, an embodiment of the present invention provides a specific implementation manner for segmenting a voice signal according to a speaker label to generate a speaker log. Please refer to Figure 7 , Figure 7 is a schematic flowchart of a method for generating a speaker log based on a speaker label provided by an embodiment of the present invention. Step S104 includes sub-steps S104-1 to S104-2.
[0100] S104-1, training a speaker recognition model according to the speaker label.
[0101] In an embodiment of the present invention, training sample data can be obtained from the voice signal according to the speaker label. Generally, several voice segments are extracted from the voice signal at a segmentation duration smaller than that in sub-step S101-1, and the speaker label of each voice segment is determined, so as to obtain the training sample data. Then, the training sample data is input into a deep learning network for training to obtain a speaker recognition model.
[0102] S104-2, inputting the voice signal into the speaker recognition model to obtain a speaker log.
[0103] In an embodiment of the present invention, the voice signal is secondarily segmented by the speaker recognition model to obtain a speaker log.
[0104] To execute the corresponding steps in the above embodiments and various possible implementation manners, the following gives an implementation manner of a speaker log generation device 200. Please refer to Figure 8 , Figure 8 shows a block diagram of the speaker log generation device 200 provided by an embodiment of the present invention. It should be noted that the basic principle and the technical effects generated by the speaker log generation device 200 provided by an embodiment of the present invention are the same as those in the above embodiments. For the sake of brief description, the embodiments of the present invention do not mention them.
[0105] The speaker log generation device 200 includes an acquisition module 201, a determination module 202, a clustering module 203, and a generation module 204.
[0106] The acquisition module 201 is configured to acquire a similarity matrix corresponding to the voice signal.
[0107] The determination module 202 determines a target clustering matrix and the number of speakers according to a first preset parameter and the similarity matrix. Among them, the target clustering matrix is obtained by sharpening the similarity matrix, and the first preset parameter constrains the sharpening process. The target clustering matrix obtained under the constraint of the first preset parameter satisfies a preset clustering condition.
[0108] The clustering module 203 is configured to cluster the target clustering matrix according to the number of speakers to obtain a speaker label.
[0109] A generating module 204, configured to segment a voice signal according to a speaker label to generate a speaker log.
[0110] As an implementation manner, the obtaining module 201 is specifically configured to divide the voice signal into multiple voice segments; input each voice segment into a voiceprint detection model to obtain a speaker feature vector corresponding to each voice segment; and obtain a similarity matrix corresponding to the language signal according to all the speaker feature vectors.
[0111] As an implementation manner, when the obtaining module 201 is used to obtain a similarity matrix corresponding to the language signal according to all the speaker feature vectors, it is further specifically configured to calculate a similarity coefficient between each speaker feature vector and other speaker feature vectors; and obtain a similarity matrix corresponding to the language signal according to all the similarity coefficients.
[0112] As an implementation manner, the determining module 202 is specifically configured to use multiple second preset parameters to respectively constrain the sharpening process of the similarity matrix to obtain multiple to-be-determined clustering matrices, where the second preset parameters are determined by the dimension of the similarity matrix; determine a target clustering matrix from the multiple to-be-determined clustering matrices according to the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing, and use the second preset parameter corresponding to the target clustering matrix as the first preset parameter; and determine the number of speakers according to the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing.
[0113] As an implementation manner, when the determining module 202 is used to determine a target clustering matrix from the multiple to-be-determined clustering matrices according to the eigenvalues of the Laplacian matrix corresponding to each to-be-determined clustering matrix after binarization and symmetrization processing, it is further specifically configured to perform a difference processing on the eigenvalues determined by each to-be-determined clustering matrix to obtain a maximum eigenvalue difference corresponding to the to-be-determined clustering matrix, where the maximum eigenvalue difference characterizes the clustering ease of the to-be-determined clustering matrix; and determine a target clustering matrix from the multiple to-be-determined clustering matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each to-be-determined clustering matrix.
[0114] As an implementation manner, when the determining module 202 is used to determine a target clustering matrix from the multiple to-be-determined clustering matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each to-be-determined clustering matrix, it is further specifically configured to calculate a ratio of the second preset parameter corresponding to each to-be-determined clustering matrix to the maximum eigenvalue difference; and use the to-be-determined clustering matrix with the minimum ratio as the target clustering matrix.
[0115] As an implementation manner, the generating module 204 is specifically configured to train a speaker recognition model according to the speaker label; and input the voice signal into the speaker recognition model to obtain a speaker log.
[0116] Further, please refer to Figure 9 , Figure 9 which is a schematic block diagram of a computer device 300 provided by an embodiment of the present invention. The computer device 300 may include a memory 310 and a processor 320.
[0117] Among them, the processor 320 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program for the speaker log generation method provided by the foregoing method embodiments.
[0118] The memory 310 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 310 may exist independently and be connected to the processor 320 through a communication bus. The memory 310 may also be integrated with the processor 320. Among them, the memory 310 is used to store machine-executable instructions for executing the solution of the present application. The processor 320 is used to execute the machine-executable instructions stored in the memory 310 to implement the foregoing method embodiments.
[0119] Since the computer device 300 provided by the embodiment of the present invention is another implementation form of the speaker log generation method provided by the foregoing method embodiments, the technical effects that can be obtained therefrom can refer to the above method embodiments and will not be elaborated herein.
[0120] The embodiment of the present invention also provides a readable storage medium containing computer-executable instructions, and the computer-executable instructions can be used to perform related operations in the speaker log generation method provided by the foregoing method embodiments when executed.
[0121] In summary, for a speaker log generation method, device, computer device, and readable storage medium provided by an embodiment of the present invention, first, a similarity matrix corresponding to a voice signal is obtained; then, according to a first preset parameter and the similarity matrix, a target clustering matrix and the number of speakers are determined, where the target clustering matrix is obtained by sharpening the similarity matrix, and the first preset parameter constrains the sharpening process, and the target clustering matrix obtained under the constraint of the first preset parameter satisfies a preset clustering condition; then, clustering is performed on the target clustering matrix according to the number of speakers to obtain speaker labels; finally, the voice signal is segmented according to the speaker labels to generate a speaker log. Since the embodiment of the present invention uses the first preset parameter to constrain the sharpening process of the similarity matrix, a target clustering matrix that satisfies the preset clustering condition is obtained, and then clustering is performed on the target clustering matrix, thereby avoiding the influence of background noise on the accuracy of the speaker labels obtained by clustering, and further improving the accuracy of the speaker log.
[0122] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for generating a speaker log, characterized in that, the method includes: Obtaining a similarity matrix corresponding to a speech signal; wherein, the value of any element in the similarity matrix corresponding to the speech signal represents the possibility that each speech segment comes from the same speaker, and the speech segments are obtained by dividing the speech signal according to a fixed segmentation duration; Using a plurality of second preset parameters to respectively constrain the sharpening process of the similarity matrix to obtain a plurality of to-be-determined clustering matrices, and the second preset parameters are determined by the dimension of the similarity matrix; For each of the to-be-determined clustering matrices, performing a difference processing on the eigenvalues determined by it to obtain the maximum eigenvalue difference corresponding to the to-be-determined clustering matrix, and the maximum eigenvalue difference represents the clustering ease of the to-be-determined clustering matrix; According to the second preset parameter and the maximum eigenvalue difference corresponding to each of the to-be-determined clustering matrices, determining a target clustering matrix from the plurality of to-be-determined clustering matrices; Determining the number of speakers according to the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing; Clustering the target clustering matrix according to the number of speakers to obtain speaker labels; Segmenting the speech signal according to the speaker labels to generate a speaker log.
2. The method according to claim 1, characterized in that, the step of obtaining a similarity matrix corresponding to a speech signal includes: Dividing the speech signal into a plurality of speech segments; Inputting each of the speech segments into a voiceprint detection model to obtain a speaker feature vector corresponding to each of the speech segments; According to all the speaker feature vectors, obtaining a similarity matrix corresponding to the speech signal.
3. The method according to claim 2, characterized in that, the step of obtaining a similarity matrix corresponding to the speech signal according to all the speaker feature vectors includes: Calculating the similarity coefficient between each of the speaker feature vectors and other speaker feature vectors; According to all the similarity coefficients, obtaining a similarity matrix corresponding to the speech signal.
4. The method according to claim 1, characterized in that, the step of determining a target clustering matrix from the plurality of to-be-determined clustering matrices according to the second preset parameter and the maximum eigenvalue difference corresponding to each of the to-be-determined clustering matrices includes: For each of the to-be-determined clustering matrices, calculating the ratio of the corresponding second preset parameter to the maximum eigenvalue difference; Taking the to-be-determined clustering matrix with the smallest ratio as the target clustering matrix.
5. The method according to claim 1, characterized in that, the step of segmenting the speech signal according to the speaker labels to generate a speaker log includes: Training a speaker recognition model according to the speaker labels; Inputting the speech signal into the speaker recognition model to obtain the speaker log.
6. A device for generating a speaker log, characterized in that, the device includes: An acquisition module, configured to acquire a similarity matrix corresponding to a speech signal; wherein, the value of any element in the similarity matrix corresponding to the speech signal represents the possibility that each speech segment comes from the same speaker, and the speech segments are obtained by segmenting the speech signal according to a fixed segmentation duration; A determination module, configured to respectively use a plurality of second preset parameters to constrain the sharpening process of the similarity matrix, and obtain a plurality of to-be-determined clustering matrices, where the second preset parameters are determined by the dimension of the similarity matrix; For each of the to-be-determined clustering matrices, perform a difference processing on the eigenvalues determined by it to obtain the maximum eigenvalue difference corresponding to the to-be-determined clustering matrix, and the maximum eigenvalue difference represents the clustering ease of the to-be-determined clustering matrix; According to the second preset parameter and the maximum eigenvalue difference corresponding to each of the to-be-determined clustering matrices, determine a target clustering matrix from the plurality of to-be-determined clustering matrices; According to the eigenvalues of the Laplacian matrix corresponding to the target clustering matrix after binarization and symmetrization processing, determine the number of speakers; A clustering module, configured to cluster the target clustering matrix according to the number of speakers to obtain speaker labels; A generation module, configured to segment the speech signal according to the speaker labels to generate a speaker log.
7. A computer device, characterized in that, it includes: A memory and a processor, the memory is used to store a computer program; the processor is used to execute the speaker log generation method according to any one of claims 1 to 5 when calling the computer program.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, the computer program, when executed by a processor, implements the speaker log generation method according to any one of claims 1 - 5.
Citation Information
Patent Citations
Contract signing method and device, computer equipment and storage medium
CN109872233A
Speaker segmentation clustering method and device, equipment and storage medium
CN113870890A