Speech clustering method and apparatus, storage medium, and electronic device

By processing speech features through a multi-layer neural network model, speaker category clustering without the need to register voiceprint information is achieved, solving the problem of cumbersome speech category clustering operations in existing technologies and improving clustering efficiency.

CN116013315BActive Publication Date: 2026-04-17HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
Filing Date
2022-11-29
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing speaker category recognition algorithms require users to register their voiceprint information in advance and cannot perform voice category clustering for users who have not registered their voiceprint information, resulting in cumbersome operation.

Method used

By determining the speech feature encoding sequence and label vector of the target speech, and using a multi-layer neural network model to calculate the high-dimensional feature vector and probability value of the speech features, clustering of different human speakers is achieved, avoiding the registration process of voiceprint information.

Benefits of technology

It simplifies the clustering process for different speaking groups, effectively identifies users who have not registered their voiceprint information, and improves the efficiency of voice category clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013315B_ABST
    Figure CN116013315B_ABST
Patent Text Reader

Abstract

This application discloses a speech clustering method, apparatus, storage medium, and electronic device, relating to the field of smart home technology. The speech clustering method includes: determining an encoded sequence of speech features of an acquired target speech; determining a label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech; inputting the label vector and the encoded sequence into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model, wherein the high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs, and the first neural network model includes a multi-layer network encoder; inputting the high-dimensional feature vector and the label of the high-dimensional feature vector into a second neural network model to obtain a target probability value output by the second neural network model, wherein the target probability value is used to represent the probability that the target speech belongs to the same category as other speech, and the second neural network model includes a multi-layer network encoder, and the other speech is speech that has already undergone speech category clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a speech clustering method, apparatus, storage medium, and electronic device. Background Technology

[0002] Existing speaker category recognition algorithms rely on users pre-registering their voiceprint information and then using that information to identify the speaker, without clustering the identity categories. Furthermore, registering voiceprint information requires significant storage space, and it's impossible to cluster the speech categories of users who haven't registered their voiceprints. Summary of the Invention

[0003] This invention provides a speech clustering method, apparatus, storage medium, and electronic device to at least solve the problem of cumbersome speech category clustering operations in related technologies.

[0004] According to an embodiment of the present invention, a speech clustering method is provided, comprising: determining an encoded sequence of speech features of an acquired target speech; determining a label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech; inputting the label vector and the encoded sequence into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model, wherein the high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs, and the first neural network model includes a multi-layer network encoder; inputting the high-dimensional feature vector and the label of the high-dimensional feature vector into a second neural network model to obtain a target probability value output by the second neural network model, wherein the target probability value is used to represent the probability that the target speech belongs to the same category as other speech, and the second neural network model includes a multi-layer network encoder, wherein the other speech is speech that has already undergone speech category clustering.

[0005] According to an embodiment of the present invention, a speech clustering apparatus is provided, comprising: a first determining module, configured to determine an encoded sequence of speech features of an acquired target speech; a second determining module, configured to determine a label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech; a first input module, configured to input the label vector and the encoded sequence into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model, wherein the high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs, and the first neural network model includes a multi-layer network encoder; and a second input module, configured to input the high-dimensional feature vector and the label of the high-dimensional feature vector into a second neural network model to obtain a target probability value output by the second neural network model, wherein the target probability value is used to represent the probability that the target speech belongs to the same category as other speech, and the second neural network model includes a multi-layer network encoder, wherein the other speech is speech that has already undergone speech category clustering.

[0006] In an exemplary embodiment, the first determining module includes: a first transformation unit, configured to perform speech feature transformation on the target speech to obtain a speech feature sequence of the target speech; and a first input unit, configured to input the speech feature sequence into a first encoder to obtain the encoded sequence output by the first encoder, wherein the first encoder includes an encoder of a cluster of hyperdimensional neural networks.

[0007] In one exemplary embodiment, the second determining module includes: a first marking unit, configured to mark the encoding sequence according to the continuity of the encoding sequence to obtain the label vector.

[0008] In an exemplary embodiment, the first input module includes a third input unit, configured to input the label vector and the encoded sequence into a multi-layer network encoder in the first neural network model, so as to distinguish the voiceprint information of the target speech in the high-dimensional mapping space of the multi-layer network encoder and obtain the high-dimensional feature vector.

[0009] In an exemplary embodiment, the apparatus further includes: a first analysis module, configured to cluster the high-dimensional feature vectors and their labels before inputting them into a second neural network model to obtain the target probability value output by the second neural network model, thereby obtaining the categories of the high-dimensional feature vectors; and a first labeling module, configured to label the categories of the high-dimensional feature vectors to obtain the labels of the high-dimensional feature vectors.

[0010] In an exemplary embodiment, the apparatus further includes a third determining module, configured to input the high-dimensional feature vector and the label of the high-dimensional feature vector into a second neural network model, obtain the target probability value output by the second neural network model, and determine that the target speech belongs to the same category as other speech if the target probability value is greater than a preset probability value.

[0011] In an exemplary embodiment, the apparatus further includes: a fourth determining module, configured to determine multiple speech feature encoding sequences of the acquired multiple speech sounds before inputting the label vectors and the encoding sequences into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model; a fifth determining module, configured to determine label vectors of the multiple speech feature encoding sequences to obtain multiple label vectors, wherein each label vector is used to represent the continuity of each speech sound; and a first training module, configured to train a first original neural network model using the multiple label vectors and the multiple speech feature encoding sequences to obtain the first neural network model.

[0012] In one exemplary embodiment, the apparatus further includes a second training module, configured to train a second original neural network model with multiple high-dimensional feature vectors of multiple speech and multiple labels of the high-dimensional feature vectors before inputting the high-dimensional feature vectors and labels of the high-dimensional feature vectors into a second neural network model to obtain the target probability value output by the second neural network model, thereby obtaining the second neural network model.

[0013] According to yet another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.

[0014] According to yet another embodiment of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0015] This invention involves determining the encoded sequence of speech features of the acquired target speech; determining the label vector of the encoded sequence, where the label vector represents the continuity of the target speech; inputting the label vector and the encoded sequence into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model, where the high-dimensional feature vector represents the category to which the object uttering the target speech belongs; the first neural network model includes a multi-layer network encoder; inputting the high-dimensional feature vector and its label into a second neural network model to obtain a target probability value output by the second neural network model, where the target probability value represents the probability that the target speech belongs to the same category as other speech, the second neural network model includes a multi-layer network encoder, and the other speech is speech that has already undergone speech category clustering. Because of the high-dimensional feature vector in the above method, the feature discrimination of the speaker can be determined, thereby clustering the speaker's category. Pre-registration of voiceprint information is not required. This simplifies the process of clustering speaker categories. Therefore, it solves the problem of the cumbersome clustering operation for speech categories in related technologies. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the hardware environment for a speech clustering method according to an embodiment of this application;

[0019] Figure 2 This is a flowchart of a speech clustering method according to an embodiment of the present invention;

[0020] Figure 3 This is a structural block diagram of a speech clustering device according to an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] According to one aspect of the embodiments of this application, a voice clustering method is provided. This interaction method for smart home devices is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and smart residential ecosystems. Optionally, in this embodiment, the above-mentioned voice clustering method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0024] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0025] This embodiment provides an image processing method. Figure 2 This is a flowchart of a speech clustering method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0026] Step S202: Determine the encoding sequence of the speech features of the acquired target speech;

[0027] Step S204: Determine the label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech;

[0028] Step S206: Input the label vector and the encoding sequence into the first neural network model to obtain the high-dimensional feature vector output by the first neural network model. The high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs. The first neural network model includes a multi-layer network encoder.

[0029] Step S208: Input the high-dimensional feature vector and the label of the high-dimensional feature vector into the second neural network model to obtain the target probability value output by the second neural network model. The target probability value is used to represent the probability that the target speech belongs to the same category as other speech. The second neural network model includes a multi-layer network encoder, and the other speech is speech that has been clustered into speech categories.

[0030] Optionally, this embodiment may be applied to scenarios involving clustering of different human speech patterns.

[0031] Optionally, the target speech can be any speech segment acquired in a home setting without a clearly identifiable speaker. For example, acquiring a control command to "turn on the air conditioner".

[0032] Optionally, both the first neural network model and the second neural network model may include multiple neural network encoders.

[0033] The entity performing the above steps may be a terminal, a server, a specific processor set in the terminal or server, or a processor or processing device set up relatively independently of the terminal or server, but is not limited to these.

[0034] Through the above steps, the following steps are taken: First, the encoded sequence of the acquired target speech features is determined. Then, a label vector is determined for the encoded sequence, where the label vector represents the continuity of the target speech. The label vector and the encoded sequence are input into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model. This high-dimensional feature vector represents the category to which the target speech belongs. The first neural network model includes a multi-layer encoder. Finally, the high-dimensional feature vector and its label are input into a second neural network model to obtain a target probability value output by the second neural network model. This target probability value represents the probability that the target speech belongs to the same category as other speech. The second neural network model includes a multi-layer encoder, and the other speech is speech that has already undergone speech category clustering. Because of the high-dimensional feature vector in this method, the speaker's feature discrimination can be determined, thereby clustering the speaker's category. Pre-registration of voiceprint information is not required, thus simplifying the process of clustering speaker categories. Therefore, this method solves the problem of cumbersome speech category clustering operations in related technologies.

[0035] In one exemplary embodiment, determining the encoded sequence of speech features of the acquired target speech includes:

[0036] S1, Perform speech feature transformation on the target speech to obtain the speech feature sequence of the target speech;

[0037] S2, input the speech feature sequence into the first encoder to obtain the encoded sequence output by the first encoder, wherein the first encoder includes an encoder of a cluster of hyperdimensional neural networks.

[0038] Optionally, the speech feature transformation can be performed using frequency domain features (fbank) or MFCC. In speech recognition and speaker identification, the most commonly used speech feature is the Mel-spectral coefficient (MFCC). The MFCC extraction process includes preprocessing, fast Fourier transform, Mei filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction. For example, the speech feature sequence corresponding to the target speech Xi can be {y(i,0), y(i,1),…}. The speech feature encoding sequence is processed by the neural network encoder Enc to obtain {{Enc(y(0,0)), Enc(y(0,1)),…}.

[0039] In one exemplary embodiment, determining the label vector of the encoded sequence includes:

[0040] S1, mark the encoding sequence according to the continuity of the encoding sequence to obtain the label vector.

[0041] Optionally, the target speech feature encoding sequence can be input in chronological order according to the speaking time of the same speaker, in which case the output feature vector is marked as 1; or it can be input in random order according to the speaking time of different speakers, in which case the output feature vector is marked as 0.

[0042] In one exemplary embodiment, the label vector and the encoded sequence are input into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model, including:

[0043] S1, the label vector and the encoding sequence are input into the multi-layer network encoder in the first neural network model to distinguish the voiceprint information of the target speech in the high-dimensional mapping space of the multi-layer network encoder and obtain the high-dimensional feature vector.

[0044] Optionally, the training function of the first neural network model can be: -\sum{\frac{exp{W1*z(label1)}}{exp{W1*z(label1)}+exp{W0*z(label0)}}}, where \sum represents the summation of all speech data samples, and W1 and W0 represent the linear network parameters for the label feature vector. This objective function continuously decreases during neural network training, representing that through the iteration of all network parameters, the multi-layer network encoder can distinguish speech data as much as possible in the high-dimensional mapping space. Moreover, this distinction can reflect personal information because each speech segment is an independent speaker. Thus, the final high-dimensional feature vector z can extract speaker identity information.

[0045] In one exemplary embodiment, before inputting the high-dimensional feature vector and its label into the second neural network model to obtain the target probability value output by the second neural network model, the method further includes:

[0046] S1, cluster the high-dimensional feature vectors to obtain the categories of the high-dimensional feature vectors;

[0047] S2, label the category of the high-dimensional feature vector to obtain the label of the high-dimensional feature vector.

[0048] Optionally, the category of the high-dimensional feature vector is used to represent the category to which the voiceprint information in the target speech belongs. For example, if speaker A and speaker B have different voiceprint information, the resulting categories will also be different. The label can be used for numerical representation, such as {I, I = 0, 1, 2, ...}.

[0049] In one exemplary embodiment, after inputting the high-dimensional feature vector and its label into a second neural network model to obtain the target probability value output by the second neural network model, the method further includes:

[0050] S1, if the target probability value is greater than the preset probability value, determine that the target speech belongs to the same category as other speech.

[0051] Optionally, the target probability value can be a comparison between two speakers. For example, comparing the voiceprint of speaker A with that of speaker B, the output is the probability that A and B belong to the same person. If the probability value is greater than a preset probability value, then A and B are determined to belong to the same person. If it is less than the preset probability value, then A and B do not belong to the same person. In the absence of B, A can be classified into a separate category.

[0052] In an exemplary embodiment, before inputting the label vector and the encoded sequence into the first neural network model to obtain the high-dimensional feature vector output by the first neural network model, the method further includes:

[0053] S1, determine the multiple speech feature coding sequences of the acquired multiple speech;

[0054] S2, determine the label vectors of multiple speech feature coding sequences to obtain multiple label vectors, where each label vector is used to represent the continuity of each speech;

[0055] S3, using multiple label vectors and multiple speech feature encoding sequences to train the first original neural network model, thus obtaining the first neural network model.

[0056] Optionally, multiple speech data can be speech data without a definite speaker, used as sample data for training the network model. For example, denoted as {X0,X1,…}, where Xi (i=0,1,…) represents a speech segment with a unique speaker. Each Xi speech data is subjected to speech feature transformation. For example, fbank or mfcc obtains the speech feature sequence {y(i,0),y(i,1),…} corresponding to Xi, then the overall speech feature sequence is {{y(0,0),y(0,1),…},{y(1,0),y(1,1),…},…}. These feature sequences are first passed through the neural network encoder Enc to obtain the speech feature encoding sequence {{Enc(y(0,0)),Enc(y(0,1)),…},{Enc(y(1,0)),Enc(y(1,1)),…},…}.

[0057] In one exemplary embodiment, before inputting the high-dimensional feature vector and its label into the second neural network model to obtain the target probability value output by the second neural network model, the method further includes:

[0058] S1, using multiple high-dimensional feature vectors and labels of multiple high-dimensional feature vectors from multiple speech samples to train a second original neural network model, thus obtaining the second neural network model.

[0059] Optionally, multiple high-dimensional feature vectors and multiple labels of these high-dimensional feature vectors are input into the second original network model for training. The training function is: -\sum{\frac{exp{W1*z(label1)}}{exp{W1*z(label1)}+exp{W0*z(label0)}}}, where \sum represents the summation over all speech data samples, and W1 and W0 represent the linear network parameters for the label feature vectors. This objective function continuously decreases during neural network training, indicating that through the iteration of all network parameters, the multi-layer network encoder can distinguish speech data as much as possible in the high-dimensional mapping space. Moreover, this distinction can reflect personal information because each speech segment represents an independent speaker. Thus, the final high-dimensional feature vector z can extract speaker identity information.

[0060] The present invention will now be described in conjunction with specific embodiments:

[0061] This embodiment identifies the category of speech based on unsupervised contrastive learning. Optionally, the principle of unsupervised contrastive learning includes category clustering. That is, the goal of contrastive learning is to place all similar objects in adjacent regions in the feature space, while dissimilar objects are placed in non-adjacent regions. Contrastive learning is typically applied in self-supervised learning of image representations, for example, the contrastive learning framework SimCLR.

[0062] Optionally, the training method of the network model based on contrastive learning in this embodiment includes the following steps:

[0063] S1, a contrastive learning feature extraction model for massive speech data: Collect a large amount of speech data without clear speaker labels, denoted as {X0,X1,…}, where Xi (i=0,1,…) represents a speech segment with a unique speaker. Perform speech feature transformation on each Xi speech data, for example, fbank or mfcc to obtain the speech feature sequence {y(i,0),y(i,1),…} corresponding to Xi. Then the overall speech feature sequence is {{y(0,0),y(0,1),…},{y(1,0),y(1,1),…},…}. These feature sequences are first passed through a neural network encoder Enc to obtain the speech feature encoding sequence {{Enc(y(0,0)),Enc(y(0,1)),…},{Enc(y(1,0)),Enc(y(1,1)),…},…}.

[0064] These encoded sequences are passed through a cluster of high-dimensional neural network encoders {R(l), l=0,1,…}, and the output results are labeled with 0 and 1. The output feature vectors of samples input to R(l) in chronological order are labeled with 1, and the output feature vectors of samples input to R(l) in random order are labeled with 0. For example, R(2)(Enc(y(0,0)),Enc(y(0,1))), R(2)(Enc(y(0,6)),Enc(y(0,7))), R(3)(Enc(y(1))), Enc(y(0,6)), Enc(y(0,7))), Enc(y(0,1))), Enc(y(0,6)), Enc(y(0,7))), ...). The labels of R(2)(Enc(y(1,3)),Enc(y(1,4))) are 1, while the labels of R(2)(Enc(y(0,5)),Enc(y(1,0))), R(2)(Enc(y(1,6)),Enc(y(0,7))), and R(3)(Enc(y(1,2)),Enc(y(1,5)),Enc(y(1,4))) are 0. This results in a large number of feature vectors with 0 and 1 labels {z(l,i)=R(l)(E The neural network module currently obtained includes Enc and R. The input is the speech feature sequence y and the generated 0 and 1 labels. The neural network training objective function is: -\sum{\frac{exp{W1*z(label1)}}{exp{W1*z(label1)}+exp{W0*z(label0)}}}, where \sum is used to represent the summation of all data samples, and W1 and W0 are used to represent the linear network parameters for the label feature vector. This objective function continuously decreases during neural network training, which means that through the iteration of all network parameters, the multi-layer network encoder can distinguish the speech data as much as possible in the high-dimensional mapping space. Moreover, this distinction can reflect personal information because each speech is an independent speaker. In this way, the final high-dimensional feature vector z can extract speaker identity information.

[0065] S2, Speaker identity feature extraction from training data: Using the high-dimensional speaker information feature vector z obtained in S1, large-scale cluster analysis is performed to obtain the category of each z. Based on the clustering results, speaker labels {I, I = 0, 1, 2, ...} are set for all training set speech data. This way, z and its labels are obtained, which can be used to train the speaker discriminator.

[0066] S3, Speaker identification algorithm: Using the labels and speaker feature vectors z obtained in S2, train the speaker identification model P(z,I), which can output a score on the known speaker category for the input speaker features;

[0067] S4. By calculating P(z,I) on a large amount of test data, the speaker category discrimination threshold T(I) is determined. That is, if the score on a certain category is greater than T(I), the speech is judged to belong to that speaker category. Finally, a speaker feature encoding network and an identity discriminator obtained by training the input audio are obtained. Combining the two can perform speaker identity discrimination on new speech. The proposed high-dimensional speaker feature extractor based on contrastive learning can better enhance the speaker information extraction capability and improve the accuracy of the feature extractor by utilizing massive amounts of unlabeled data.

[0068] This embodiment introduces contrastive learning, which can more effectively mine the effective high-dimensional features for speaker differentiation in massive speech training data, thereby greatly improving the feature extraction capability for non-registered speaker recognition and thus improving the accuracy of speaker discrimination.

[0069] This embodiment also provides a speech clustering device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0070] Figure 3 This is a structural block diagram of a speech clustering device according to an embodiment of the present invention, such as... Figure 3 As shown, the device includes:

[0071] The first determining module 32 is used to determine the encoded sequence of the speech features of the acquired target speech;

[0072] The second determining module 34 is used to determine the label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech;

[0073] The first input module 36 is used to input the label vector and the encoding sequence into the first neural network model to obtain the high-dimensional feature vector output by the first neural network model. The high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs. The first neural network model includes a multi-layer network encoder.

[0074] The second input module 38 is used to input the high-dimensional feature vector and the label of the high-dimensional feature vector into the second neural network model to obtain the target probability value output by the second neural network model. The target probability value is used to represent the probability that the target speech belongs to the same category as other speech. The second neural network model includes a multi-layer network encoder, and the other speech is speech that has been clustered into speech categories.

[0075] In one exemplary embodiment, the first determining module described above includes:

[0076] The first transformation unit is used to transform the target speech into speech features to obtain a speech feature sequence of the target speech.

[0077] The first input unit is used to input the above-mentioned speech feature sequence into the first encoder to obtain the above-mentioned encoded sequence output by the first encoder, wherein the first encoder includes an encoder of a cluster of hyperdimensional neural networks.

[0078] In one exemplary embodiment, the second determining module described above includes:

[0079] The first tagging unit is used to tag the above-mentioned encoding sequence according to the continuity of the above-mentioned encoding sequence to obtain the above-mentioned tag vector.

[0080] In one exemplary embodiment, the first input module described above includes:

[0081] The third input unit is used to input the label vector and the encoding sequence into the multi-layer network encoder in the first neural network model, so as to distinguish the voiceprint information of the target speech in the high-dimensional mapping space of the multi-layer network encoder and obtain the high-dimensional feature vector.

[0082] In one exemplary embodiment, the above-described apparatus further includes:

[0083] The first analysis module is used to cluster the high-dimensional feature vectors and their labels before inputting them into the second neural network model to obtain the target probability value output by the second neural network model, thereby obtaining the category of the high-dimensional feature vectors.

[0084] The first labeling module is used to label the categories of the high-dimensional feature vectors and obtain the labels of the high-dimensional feature vectors.

[0085] In one exemplary embodiment, the above-described apparatus further includes:

[0086] The third determining module is used to input the high-dimensional feature vector and the label of the high-dimensional feature vector into the second neural network model, obtain the target probability value output by the second neural network model, and determine that the target speech belongs to the same category as other speech if the target probability value is greater than the preset probability value.

[0087] In one exemplary embodiment, the above-described apparatus further includes:

[0088] The fourth determining module is used to determine multiple speech feature encoding sequences of multiple speech sounds before inputting the above label vector and the above encoding sequence into the first neural network model to obtain the high-dimensional feature vector output by the first neural network model.

[0089] The fifth determining module is used to determine the label vectors of the multiple speech feature coding sequences mentioned above, thereby obtaining multiple label vectors, wherein each of the above label vectors is used to represent the continuity of each of the above speech segments;

[0090] The first training module is used to train a first original neural network model using multiple of the aforementioned label vectors and multiple of the aforementioned speech feature encoding sequences, thereby obtaining the aforementioned first neural network model.

[0091] In one exemplary embodiment, the above-described apparatus further includes:

[0092] The second training module is used to input the aforementioned high-dimensional feature vectors and their labels into the second neural network model to obtain the target probability value output by the second neural network model. Before this, the module trains the second original neural network model with the multiple high-dimensional feature vectors and their labels from the aforementioned speech to obtain the second neural network model.

[0093] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A speech clustering method, characterized in that, include: Determine the encoding sequence of the speech features of the acquired target speech, wherein the encoding sequence is obtained by an encoder of a high-dimensional neural network; Determine the label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech; The label vector and the encoded sequence are input into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model. The high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs. The first neural network model includes a multi-layer network encoder. The high-dimensional feature vector and its label are input into the second neural network model to obtain the target probability value output by the second neural network model. The target probability value is used to represent the probability that the target speech belongs to the same category as other speech. The second neural network model includes a multi-layer network encoder, and the other speech is speech that has been clustered into speech categories. The process of inputting the label vector and the encoded sequence into a first neural network model to obtain a high-dimensional feature vector output by the first neural network model includes: inputting the label vector and the encoded sequence into a multi-layer network encoder in the first neural network model to distinguish the voiceprint information of the target speech in the high-dimensional mapping space of the multi-layer network encoder, thereby obtaining the high-dimensional feature vector.

2. The method according to claim 1, characterized in that, Determine the encoded sequence of speech features of the acquired target speech, including: The target speech is subjected to speech feature transformation to obtain the speech feature sequence of the target speech; The speech feature sequence is input into a first encoder to obtain the encoded sequence output by the first encoder, wherein the first encoder includes an encoder of a cluster of hyperdimensional neural networks.

3. The method according to claim 1, characterized in that, Determining the label vector of the encoded sequence includes: The coding sequence is labeled according to its continuity to obtain the label vector.

4. The method according to claim 1, characterized in that, Before inputting the high-dimensional feature vector and its label into the second neural network model to obtain the target probability value output by the second neural network model, the method further includes: Cluster the high-dimensional feature vectors to obtain the categories of the high-dimensional feature vectors; The category of the high-dimensional feature vector is labeled to obtain the label of the high-dimensional feature vector.

5. The method according to claim 1, characterized in that, After inputting the high-dimensional feature vector and its label into the second neural network model to obtain the target probability value output by the second neural network model, the method further includes: If the target probability value is greater than the preset probability value, the target speech is determined to belong to the same category as other speech.

6. The method according to claim 1, characterized in that, Before inputting the label vector and the encoded sequence into the first neural network model to obtain the high-dimensional feature vector output by the first neural network model, the method further includes: Determine multiple speech feature encoding sequences from the acquired speech samples; A label vector is determined for a plurality of speech feature encoding sequences, resulting in a plurality of label vectors, wherein each label vector is used to represent the continuity of each speech; The first original neural network model is obtained by training a first neural network model using multiple label vectors and multiple speech feature encoding sequences.

7. The method according to claim 6, characterized in that, Before inputting the high-dimensional feature vector and its label into the second neural network model to obtain the target probability value output by the second neural network model, the method further includes: The second original neural network model is obtained by training multiple high-dimensional feature vectors of the multiple speech sounds and the labels of the multiple high-dimensional feature vectors.

8. A speech clustering device, characterized in that, include: The first determining module is used to determine the encoding sequence of the speech features of the acquired target speech, wherein the encoding sequence is obtained by an encoder of a hyperdimensional neural network; The second determining module is used to determine the label vector of the encoded sequence, wherein the label vector is used to represent the continuity of the target speech; The first input module is used to input the label vector and the encoded sequence into the first neural network model to obtain a high-dimensional feature vector output by the first neural network model. The high-dimensional feature vector is used to represent the category to which the object uttering the target speech belongs. The first neural network model includes a multi-layer network encoder. The second input module is used to input the high-dimensional feature vector and the label of the high-dimensional feature vector into the second neural network model to obtain the target probability value output by the second neural network model. The target probability value is used to represent the probability that the target speech belongs to the same category as other speech. The second neural network model includes a multi-layer network encoder, and the other speech is speech that has been clustered into speech categories. The device is further configured to input the label vector and the encoded sequence into a multi-layer network encoder in the first neural network model, so as to distinguish the voiceprint information of the target speech in the high-dimensional mapping space of the multi-layer network encoder and obtain the high-dimensional feature vector.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 7.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Speaker separation method and device, electronic equipment and storage medium

    CN111524527A

  • Voice detection method and device, electronic equipment and computer readable storage medium

    CN114333771A