Voice recommendation, model training method, device, electronic device and storage medium
The importance characteristics of speech are extracted through the importance characterization model of multi-task training, which solves the problem of poor voice recommendation in the prior art, and achieves end-to-end speech recommendation with high reliability and accuracy.
Patent Information
- Application Number
- CN202510026166.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing voice recommendation methods rely on multiple links such as voice transcription and intention recognition, resulting in poor recommendation results, and are prone to missed recommendations or incorrect recommendations.
The importance characterization model obtained by multi-task training is used to extract the importance characteristics of candidate speech and recommend the speech based on these characteristics. This model combines text content features, speaker features and importance features to recommend them in an end-to-end manner.
End-to-end voice recommendation is realized, avoiding the influence of multiple links cascades, and improving the reliability and accuracy of recommendations. Multiple measurements of importance characteristics ensure comprehensiveness and reliability of recommendations.
Smart Images

Figure CN119479699B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and in particular, to a speech recommendation, model training method, device, electronic device, and storage medium. Background Art
[0002] With the development of information technology, a vast amount of information is presented in the form of speech, and automated speech recommendation methods have emerged as the times require.
[0003] In current speech recommendation methods, most of them first transcribe speech into text, and then perform text analysis such as intent recognition on the text, and select important speech according to the results of the text analysis and recommend it to relevant personnel. The recommendation effect of the above method is affected by the effects of each link such as speech transcription and intent recognition. Once the effect of a certain link is not good, it is very likely that there will be missed recommendations or mis-recommendations, affecting the user experience. Summary of the Invention
[0004] The present invention provides a speech recommendation, model training method, device, electronic device, and storage medium to solve the defect of poor speech recommendation effect in related technologies.
[0005] The present invention provides a speech recommendation method, including:
[0006] Obtaining candidate speech;
[0007] Inputting the candidate speech into an importance characterization model to obtain the importance feature of the candidate speech output by the importance characterization model based on the text content feature and / or speaker feature of the candidate speech; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature;
[0008] Performing speech recommendation based on the importance feature of the candidate speech.
[0009] According to the speech recommendation method provided by the present invention, the step of inputting the candidate speech into an importance characterization model to obtain the importance feature of the candidate speech output by the importance characterization model based on the text content feature and / or speaker feature of the candidate speech includes:
[0010] Extracting the text content feature of the candidate speech based on the text content extraction network in the importance characterization model and the speech feature of the candidate speech, and / or, extracting the speaker feature of the candidate speech based on the speaker extraction network in the importance characterization model and the speech feature of the candidate speech;
[0011] Extract the importance feature of the candidate speech based on the importance extraction network in the importance characterization model, the speech feature, and the text content feature and / or the speaker feature of the candidate speech.
[0012] According to a speech recommendation method provided by the present invention, performing speech recommendation based on the importance feature of the candidate speech includes:
[0013] Cluster the importance features of each candidate speech to obtain at least one speech cluster;
[0014] Determine a candidate speech from each speech cluster as a deduplicated speech respectively, and perform speech recommendation based on the importance feature of the deduplicated speech.
[0015] According to a speech recommendation method provided by the present invention, performing speech recommendation based on the importance feature of the deduplicated speech includes:
[0016] Perform speech recommendation based on the similarity between the importance feature of the deduplicated speech and the importance features of each registered speech cluster;
[0017] Each of the registered speech clusters is obtained by clustering the importance features of each registered speech.
[0018] According to a speech recommendation method provided by the present invention, performing speech recommendation based on the similarity between the importance feature of the deduplicated speech and the importance features of each registered speech cluster includes:
[0019] When the maximum value of the similarity between the importance feature of the deduplicated speech and the importance features of each registered speech cluster is greater than a preset threshold, and the registered speech cluster corresponding to the maximum value is a recommended speech cluster, perform speech recommendation for the deduplicated speech.
[0020] According to a speech recommendation method provided by the present invention, inputting the candidate speech into the importance characterization model to obtain the importance feature of the candidate speech output by the importance characterization model based on the text content feature and / or the speaker feature of the candidate speech includes:
[0021] Input each speech segment of the candidate speech into the importance characterization model to obtain the importance feature of each speech segment output by the importance characterization model based on the speech feature and / or the text content feature of each speech segment;
[0022] Based on the importance features of each speech segment, determine the importance feature of the candidate speech.
[0023] The present invention also provides a model training method, including:
[0024] Determine training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features;
[0025] Based on the training samples for the multiple tasks, perform model training to obtain an importance characterization model, where the importance characterization model is used to output the importance features of a candidate speech based on the text content features and / or speaker features of the candidate speech.
[0026] According to a model training method provided by the present invention, the performing model training based on the training samples for the multiple tasks to obtain an importance characterization model includes:
[0027] Based on the text content extraction network in the initial model, extract sample text content features from the speech features of the training samples for the speech recognition task, and / or, based on the speaker extraction network in the initial model, extract sample speaker features from the speech features of the training samples for the speaker recognition task;
[0028] Based on the importance extraction network in the initial model, extract sample importance features from the speech features of the training samples for the importance classification task, as well as the sample text content features and / or sample speaker features of the training samples for the importance classification task;
[0029] Based on the speech recognition loss determined by the sample text content features and / or the speaker recognition loss determined by the sample speaker features, and the importance loss determined by the sample importance features, perform parameter iteration on the initial model to obtain the importance characterization model.
[0030] The present invention also provides a speech recommendation device, including:
[0031] A speech acquisition unit for acquiring a candidate speech;
[0032] A feature extraction unit for inputting the candidate speech into the importance characterization model to obtain the importance features of the candidate speech output by the importance characterization model based on the text content features and / or speaker features of the candidate speech; the importance characterization model is obtained through training for multiple tasks, where the multiple tasks include a speech recognition task based on the text content features and / or a speaker recognition task based on the speaker features, and an importance classification task based on the importance features;
[0033] A speech recommendation unit for performing speech recommendation based on the importance features of the candidate speech.
[0034] The present invention also provides a model training device, including:
[0035] A sample acquisition unit, configured to determine training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features;
[0036] A model training unit, configured to perform model training based on the training samples of multiple tasks to obtain an importance characterization model, where the importance characterization model is used to output the importance features of a candidate speech based on the text content features and / or speaker features of the candidate speech.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the speech recommendation method or the model training method as described in any one of the above.
[0038] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the speech recommendation method or the model training method as described in any one of the above.
[0039] The present invention also provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the speech recommendation method or the model training method as described in any one of the above.
[0040] The speech recommendation, model training method, device, electronic device, and storage medium provided by the present invention extract the importance features of a candidate speech through an importance characterization model obtained by training multiple tasks, and perform speech recommendation based on this. First, end-to-end speech recommendation is achieved, and there is no situation in the prior art where multiple links are cascaded and the accuracy of each link will affect the speech recommendation effect, which can improve the reliability of speech recommendation; second, the importance features can reflect the importance of the candidate speech from aspects such as text content and / or speaker information, and whether it is important, making the measurement of importance more comprehensive and reliable, and further ensuring the accuracy of speech recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0042] Figure 1 It is a schematic flowchart of speech recommendation in related technologies.
[0043] Figure 2It is a schematic flowchart of the voice recommendation method provided by the present invention.
[0044] Figure 3 It is a schematic structural diagram of the importance characterization model provided by the present invention.
[0045] Figure 4 It is a schematic flowchart of the model training method provided by the present invention.
[0046] Figure 5 It is a schematic structural diagram of the voice recommendation device provided by the present invention.
[0047] Figure 6 It is a schematic structural diagram of the model training device provided by the present invention.
[0048] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] With the development of information technology, a vast amount of information is presented in the form of voice. In order to mine important information, relevant personnel usually need to collect a large amount of voice data by various means and then screen it. Due to the uncertainty of the collection target and the diversity of collection means, the voice data obtained is often large in quantity and complex in content, which brings challenges to relevant personnel for information mining.
[0051] To facilitate relevant personnel to perform information mining based on voice data, an automated voice recommendation method has emerged as the times require.
[0052] Figure 1 It is a schematic flowchart of voice recommendation in the related art. As Figure 1 shown, in the current voice recommendation method, most of them first transcribe the voice into text, and then perform text analysis such as intent recognition on the text, and select important voices according to the results of the text analysis and recommend them to relevant personnel.
[0053] However, due to the large amount of voice data to be recommended and the possible existence of many duplicate voices, directly performing voice transcription will result in a waste of a large amount of computing power and reduce the voice recommendation efficiency. Moreover, limited by the implementation accuracy of each link such as voice transcription and intent recognition, each link in voice recommendation will affect the recommendation effect, thus leading to the problem of missing important voices or misrecommending unimportant voices. Only when the accuracy of each link such as voice transcription and intent recognition is very high can a better voice recommendation be achieved.
[0054] In view of the above problems, an embodiment of the present invention provides a voice recommendation method. Figure 2 It is a schematic flowchart of the voice recommendation method provided by the present invention. As Figure 2 shown, the method includes:
[0055] Step 210, obtain candidate voices.
[0056] Here, the candidate voices are the voices that are pre-collected and can be used for voice recommendation.
[0057] The candidate voices can be obtained through a sound pickup device. Here, the sound pickup device can be a smart phone, a tablet computer, or also a smart appliance such as a speaker, a TV, and an air conditioner. After the sound pickup device picks up the candidate voices through a microphone array, it can also amplify and denoise the candidate voices. In addition, the candidate voices can be voice segments formed after the sound pickup ends, or voice streams during the real-time sound pickup process. The embodiment of the present invention does not make specific limitations on this.
[0058] It can be understood that the number of candidate voices can be one or multiple, and the embodiment of the present invention does not make specific limitations on this.
[0059] Step 220, input the candidate voices into an importance characterization model to obtain the importance characteristics of the candidate voices output by the importance characterization model based on the text content characteristics and / or speaker characteristics of the candidate voices; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content characteristics and / or a speaker recognition task based on the speaker characteristics, as well as an importance classification task based on the importance characteristics.
[0060] Specifically, after obtaining the candidate voices, the candidate voices can be input into the importance characterization model, and the importance characterization model is used to represent the importance of the input candidate voices at the voice level, thereby obtaining the importance characteristics of the candidate voices output by the importance characterization model. Here, the importance characteristics of the candidate voices are the characteristics used to characterize the importance of the candidate voices in the voice recommendation scenario.
[0061] Here, the importance representation model is a neural network model that has been pre-trained to extract the importance features of the input speech. When training the importance representation model, a combined training method with multiple tasks can be adopted. The multiple tasks mentioned here include the importance classification task, as well as the speech recognition task and / or the speaker recognition task.
[0062] Among them, for the importance classification task, a classification layer is connected after the importance features extracted from the speech to achieve the classification of whether the speech is important. For the speech recognition task, a fully connected layer for speech recognition is connected after the text content features extracted from the speech to achieve speech recognition. For the speaker recognition task, a classification layer is connected after the speaker features extracted from the speech to achieve the classification of the speaker to which the speech belongs.
[0063] When training the importance representation model, it can be trained by combining the importance classification task and the speech recognition task. Through this training, the importance representation model can extract the text content features of the input candidate speech and, based on the text content features of the candidate speech, extract the importance features of the candidate speech. The obtained importance features integrate the text content in the candidate speech and can reflect the importance of the candidate speech from two aspects: the text content and whether it is important.
[0064] Or, when training the importance representation model, it can be trained by combining the importance classification task and the speaker recognition task. Through this training, the importance representation model can extract the speaker features of the input candidate speech and, based on the speaker features of the candidate speech, extract the importance features of the candidate speech. The obtained importance features integrate the speaker information in the candidate speech and can reflect the importance of the candidate speech from two aspects: the speaker information and whether it is important.
[0065] Or, when training the importance representation model, it can be trained by combining the importance classification task, the speech recognition task, and the speaker recognition task. Through this training, the importance representation model can extract the text content features of the input candidate speech, as well as the speaker features of the candidate speech, and, based on the text content features and speaker features of the candidate speech, extract the importance features of the candidate speech. The obtained importance features integrate the text content and speaker information in the candidate speech and can reflect the importance of the candidate speech from three aspects: the text content, the speaker information, and whether it is important.
[0066] That is, the importance characterization model trained through multiple tasks can extract the importance features of candidate voices. The extracted importance features can reflect the importance of candidate voices from aspects such as text content and / or speaker information, and whether it is important. Compared with the related art that only considers whether a candidate voice is important from the intention, the measurement of the importance of candidate voices is more comprehensive and reliable.
[0067] Moreover, in the embodiments of the present invention, the importance characterization model is an end-to-end model. By inputting a candidate voice, its importance features can be obtained, and there is no situation in the related art where multiple links are cascaded and the accuracy of each link will affect the voice recommendation effect. Thus, the reliability of voice recommendation can be greatly enhanced.
[0068] Step 230: Perform voice recommendation based on the importance features of the candidate voice.
[0069] Specifically, after obtaining the importance features of the candidate voice, voice recommendation can be performed based on the importance features. For example, the importance features of the candidate voice can be used to classify whether the candidate voice is important, and thus the important candidate voices can be used as the voices to be recommended for recommendation. Another example is that the distance between the importance features of the candidate voice and the pre-set recommended importance features can be calculated. When the distance is less than the preset threshold, the candidate voice is used as the voice to be recommended for recommendation. When the distance is greater than or equal to the preset threshold, the candidate voice is not recommended. The embodiments of the present invention do not make specific limitations in this regard.
[0070] The method provided by the embodiments of the present invention extracts the importance features of candidate voices through an importance characterization model trained through multiple tasks, and performs voice recommendation based on this. On the one hand, it realizes end-to-end voice recommendation, and there is no situation in the prior art where multiple links are cascaded and the accuracy of each link will affect the voice recommendation effect, which can improve the reliability of voice recommendation. On the other hand, the importance features can reflect the importance of candidate voices from aspects such as text content and / or speaker information, and whether it is important, making the measurement of importance more comprehensive and reliable, and further ensuring the accuracy of voice recommendation.
[0071] Based on the above embodiments, in step 220, the inputting the candidate voice into the importance characterization model to obtain the importance features of the candidate voice output by the importance characterization model based on the text content features and / or speaker features of the candidate voice includes:
[0072] Extracting the text content features of the candidate voice based on the text content extraction network in the importance characterization model and the voice features of the candidate voice, and / or, extracting the speaker features of the candidate voice based on the speaker extraction network in the importance characterization model and the voice features of the candidate voice;
[0073] Based on the importance extraction network in the importance characterization model, the speech features, and the text content features and / or the speaker features, extract the importance features of the candidate speech.
[0074] Specifically, the importance characterization model may include a text content extraction network and / or a speaker extraction network. Among them, the text content extraction network is a network for extracting text content features, and the text content extraction network may be trained based on the speech recognition task during the multi-task training of the importance characterization model; the speaker extraction network is a network for extracting speaker features, and the speaker extraction network may be trained based on the speaker recognition task during the multi-task training of the importance characterization model.
[0075] When inputting the candidate speech into the importance characterization model, the importance characterization model may extract the speech features of the candidate speech. For example, the importance characterization model may further include a pre-trained speech model, and the speech features of the candidate speech are extracted through the pre-trained speech model. Here, the pre-trained speech model may be a CPC (Contrastive Predictive Coding) neural network, a wav2vec2.0 neural network, etc., and the embodiments of the present invention do not make specific limitations thereto.
[0076] After obtaining the speech features of the candidate speech, the importance characterization model may extract the text content features of the candidate speech from the speech features of the candidate speech based on the text content extraction network included therein; and / or, the importance characterization model may extract the speaker features of the candidate speech from the speech features of the candidate speech based on the speaker extraction network included therein.
[0077] Here, both the text content extraction network and the speaker extraction network may be network structures including convolutional neural networks, and both may further extract high-level features of corresponding tasks from the speech features of the candidate speech. The high-level features extracted by the text content extraction network are the text content features, and the high-level features extracted by the speaker extraction network are the speaker features.
[0078] In addition, the importance characterization model further includes an importance extraction network, and the input of the importance extraction network is connected to the output of the text content extraction network and / or the speaker extraction network. The importance extraction network may be trained by combining the speech recognition task and / or the speaker recognition task, and the importance classification task during the multi-task training of the importance characterization model.
[0079] After obtaining the speech features of the candidate speech, as well as the text content features and / or speaker features of the candidate speech, the importance characterization model can extract the importance features of the candidate speech based on the importance extraction network included therein, in combination with the speech features of the candidate speech, as well as the text content features and / or speaker features of the candidate speech.
[0080] Here, the importance extraction network can also be a network structure including a convolutional neural network, and the importance extraction network can further extract high-level features from the speech features of the candidate speech, as well as the text content features and / or speaker features, that is, obtain importance features that can fuse text content and / or speaker information.
[0081] Based on any of the above embodiments, Figure 3 is a schematic structural diagram of the importance characterization model provided by the present invention. As Figure 3 shown, the importance characterization model can include Figure 3 the part boxed by the dashed line in
[0082] That is, the importance characterization model includes a pre-trained model, a text content extraction network, an importance extraction network, and a speaker extraction network.
[0083] The text content extraction network takes the speech features of the candidate speech as input and the text content features of the candidate speech as output. When performing multi-task training on the importance representation model, a speech recognition fully-connected layer can be connected after the text content extraction network. The speech recognition fully-connected layer takes the text content features as input and the speech recognition text as output, and constructs the loss of the speech recognition task based on the speech recognition text and the text label of the corresponding speech, as part of the loss of the importance representation model.
[0084] The speaker extraction network takes the speech features of the candidate speech as input and the speaker features of the candidate speech as output. When performing multi-task training on the importance representation model, a speaker classification layer can be connected after the speaker extraction network. The speaker classification layer takes the speaker features as input and the speaker classification result as output, and constructs the loss of the speaker recognition task based on the speaker classification result and the speaker label of the corresponding speech, as part of the loss of the importance representation model.
[0085] The importance extraction network takes the speech features, text content features, and speaker features of the candidate speech as input and the importance features of the candidate speech as output. That is, the input of the importance extraction network is respectively connected to the output of the pre-trained model, the output of the text content extraction network, and the output of the speaker extraction network. When performing multi-task training on the importance representation model, an important / unimportant classification layer can be connected after the importance extraction network. The important / unimportant classification layer takes the importance features as input and the importance classification result as output, and the importance classification result here reflects whether the speech itself is important or not. Based on the importance classification result and the importance label of the corresponding speech, the loss of the importance classification task can be constructed, as part of the loss of the importance representation model.
[0086] Correspondingly, when performing joint training on multiple tasks for the importance representation model, the loss function applied in the joint training can be expressed as the following formula:
[0087] Loss = CTCLoss + CELoss spk + CELoss important-uninmportant
[0088] In the formula, Loss is the loss function of the joint training, and α, β, and γ are all pre-set parameters. For example, α can be set to 0.30, β to 0.30, and γ to 0.40. CTCLoss is the loss of the speech recognition task, and CELoss spk is the loss of the speaker recognition task, and CELoss important-uninmportant is the loss of the importance classification task.
[0089] The entire importance representation model can be iteratively refined based on the loss of joint training until the loss of joint training stabilizes or reaches the maximum number of iterations, at which point the training of the importance representation model is completed. It can be understood that Figure 3 the fully connected layer for speech recognition, the speaker classification layer, and the important / unimportant classification layer in
[0090] Based on any of the above embodiments, in step 230, the performing speech recommendation based on the importance features of the candidate speech includes:
[0091] Clustering the importance features of each candidate speech to obtain at least one speech cluster;
[0092] Determining a candidate speech from each speech cluster as a deduplicated speech, and performing speech recommendation based on the importance features of the deduplicated speech.
[0093] Specifically, due to the non-fixed nature of the collection target and the diversity of collection means, there are often a large number of speeches with duplicate content among the massive candidate speeches. If not screened, it will lead to the situation where speeches with duplicate content are repeatedly recommended, affecting the recommendation experience.
[0094] To address this issue, in the embodiments of the present invention, for the importance features of each extracted candidate speech, feature clustering can be performed. Feature clustering for the importance features here aims to divide each candidate speech into several speech clusters, and the importance features of the candidate speeches within the same speech cluster are similar to each other, while the importance features of the candidate speeches in different speech clusters are different. The algorithm for feature clustering here can be K-means, hierarchical clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), etc. For example, the importance features of each candidate speech can be clustered based on AP clustering (Affinity Propagation Clustering), and the clustering threshold can be set to 0.9999 or other values for AP clustering. It can be understood that the higher the value of the clustering threshold here, the greater the degree of duplication of the candidate speeches within the clustered speech clusters.
[0095] After completing the clustering, at least one speech cluster can be obtained, and each speech cluster includes at least one candidate speech.
[0096] For each voice cluster, a candidate voice can be selected from the voice cluster as the representative of the voice cluster, that is, as the deduplicated voice of the voice cluster, and based on the importance feature of the deduplicated voice, it is determined whether to recommend the deduplicated voice as a recommended voice.
[0097] It can be understood that each voice cluster can have a deduplicated voice, and the reason for selecting only one deduplicated voice from one voice cluster is that the candidate voices belonging to the same voice cluster are candidate voices with highly similar importance features, that is, the candidate voices belonging to the same voice cluster are almost all duplicate voices. In this case, if it is determined for each candidate voice in the same voice cluster whether to be recommended as a recommended voice, it is very likely to cause the problem of duplicate recommendation. If only one candidate voice is selected from one voice cluster as the deduplicated voice, the deduplication operation for duplicate voices can be realized, thereby avoiding the problem of duplicate recommendation and greatly reducing the workload required for voice recommendation based on importance features.
[0098] Furthermore, one candidate voice is selected from each voice cluster respectively. Specifically, one candidate voice can be randomly selected from each voice cluster as the deduplicated voice, or one candidate voice with the longest duration can be selected from each voice cluster as the deduplicated voice. The text content contained in the deduplicated voice selected in this way is the richest in the entire voice cluster. The embodiments of the present invention do not make specific limitations on this.
[0099] In the method provided by the embodiments of the present invention, through clustering the importance features of each candidate voice, the deduplication operation for candidate voices is realized, thereby avoiding the situation of duplicate recommendation and improving the efficiency of voice recommendation.
[0100] Based on any of the above embodiments, in step 230, the performing voice recommendation based on the importance feature of the deduplicated voice includes:
[0101] Performing voice recommendation based on the similarity between the importance feature of the deduplicated voice and the importance features of each registered voice cluster;
[0102] Each of the registered voice clusters is obtained by clustering the importance features of each registered voice.
[0103] Specifically, in order to implement voice recommendation based on importance features, registered voices can be collected in advance. Here, the registered voice is the reference voice for voice recommendation. The registered voice can be defaulted to be an important voice that needs to be recommended, or the registered voice can be defaulted to be an unimportant voice that is not recommended.
[0104] After obtaining the registered voice, the importance features of the registered voice can be extracted, and feature clustering can be performed on the importance features of the registered voice, thereby obtaining at least one voice cluster. Here, the voice cluster composed of the registered voice is denoted as the registered voice cluster. It can be understood that the method for extracting the importance features of the registered voice and the method for performing feature clustering on the importance features of the registered voice are the same as the method for extracting the importance features of the candidate voice and the method for performing feature clustering on the importance features of the candidate voice in the above embodiments, and will not be elaborated here.
[0105] After obtaining the registered voice cluster, the importance features of the registered voice cluster can be determined. Here, for any registered voice cluster, the importance features of the registered voice cluster can be the mean of the importance features of all the registered voices in the registered voice cluster. It can be understood that the importance features of the registered voice cluster can reflect the common importance features of a class of registered voices. For example, for a registered voice cluster, the importance features of the registered voice cluster can be marked as:
[0106]
[0107] In the formula, is the importance feature of the kth registered voice cluster. The kth registered voice cluster contains a total of M registered voices, where is the importance feature of the mth registered voice among the M registered voices.
[0108] After completing deduplication and obtaining the deduplicated voice, the importance features of the deduplicated voice can be compared with the importance features of each registered voice cluster, thereby determining whether to recommend the deduplicated voice as a recommended voice. Here, the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster can be calculated. It can be understood that the higher the similarity, the more similar the deduplicated voice is to the registered voices in the registered voice cluster, and the higher the probability of determining whether to recommend according to the recommended or non-recommended labels corresponding to the registered voices in the registered voice cluster. The lower the similarity, the more different the deduplicated voice is from the registered voices in the registered voice cluster, and the lower the probability of determining whether to recommend according to the recommended or non-recommended labels corresponding to the registered voices in the registered voice cluster.
[0109] Based on any of the above embodiments, in step 230, the voice recommendation based on the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster includes:
[0110] When the maximum value of the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster is greater than a preset threshold, and the registered voice cluster corresponding to the maximum value is a recommended voice cluster, the deduplicated voice is recommended for voice.
[0111] Specifically, in the case where there are multiple registered voice clusters, the duplicate-removed voice needs to calculate the similarity of the importance features with each registered voice cluster respectively, and thus multiple similarities can be obtained. The similarity calculation here can be to calculate the cosine similarity between the importance features, or can be implemented by other similarity calculation methods. The embodiments of the present invention do not make specific limitations on this.
[0112] In the case of obtaining multiple similarities, the maximum value among the multiple similarities can be obtained therefrom. The maximum value reflects the similarity of the importance features between the duplicate-removed voice and the registered voice cluster that is most similar to the duplicate-removed voice. After obtaining the maximum value, the size between the maximum value and a preset threshold can be compared.
[0113] Here, the preset threshold is a similarity threshold preset for reflecting whether the duplicate-removed voice belongs to the registered voice cluster. If the maximum value is greater than the preset threshold, it can be determined that the duplicate-removed voice and the registered voice in the registered voice cluster corresponding to the maximum value belong to the same type of voice. In this case, if the registered voice in the registered voice cluster corresponding to the maximum value is an important voice recommended by default, it can also be determined that the duplicate-removed voice is a recommended voice and is recommended;
[0114] If the maximum value is greater than the preset threshold, and the registered voice in the registered voice cluster corresponding to the maximum value is an unimportant voice that is not recommended, it can also be determined that the duplicate-removed voice is an unimportant voice and is not recommended.
[0115] In addition, if the maximum value is less than or equal to the preset threshold, it is determined that the duplicate-removed voice does not belong to any registered voice. In this case, it can be default that the duplicate-removed voice is an unimportant voice and is not recommended.
[0116] Based on any of the above embodiments, in step 220, the inputting the candidate voice into the importance characterization model to obtain the importance features of the candidate voice output by the importance characterization model based on the text content features and / or speaker features of the candidate voice includes:
[0117] Inputting each voice segment of the candidate voice into the importance characterization model to obtain the importance features of each voice segment output by the importance characterization model based on the voice features and / or text content features of each voice segment;
[0118] Determining the importance features of the candidate voice based on the importance features of each voice segment.
[0119] Specifically, considering that the duration of some candidate voices is relatively long, if the entire candidate voice is directly input into the importance characterization model for importance feature extraction, it will take a relatively long time. In the embodiments of the present invention, the candidate voice can be segmented, that is, the candidate voice is divided into multiple voice segments, and each voice segment here is a part of the candidate voice. For example, the duration of each voice segment can be between 0.5 seconds and 20 seconds.
[0120] Subsequently, the candidate voice can be input into the importance characterization model in units of voice segments, that is, each voice segment of the candidate voice can be input into the importance characterization model. Thus, the importance characterization model can perform importance feature extraction in units of voice segments, and then output the importance features of each voice segment.
[0121] In this case, the importance features of the entire candidate voice can be determined by combining the importance features of each voice segment in the candidate voice. For example, the importance features of the candidate voice can be determined by the following formula:
[0122]
[0123] In the formula, is the importance feature of the i-th candidate voice. The i-th candidate voice is divided into L voice segments, and the importance feature of the i-th voice segment is . That is, the importance feature of the i-th candidate voice is the mean value of the importance features of all voice segments.
[0124] In the method provided by the embodiments of the present invention, segmenting the candidate voice into voice segments to extract importance features can improve the extraction efficiency of importance features, and further improve the voice recommendation efficiency.
[0125] Based on any of the above embodiments, Figure 4 is a schematic flowchart of the model training method provided by the present invention. As Figure 4 shown, the method includes:
[0126] Step 410, determining training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features.
[0127] Step 420, performing model training based on the training samples of multiple tasks to obtain an importance characterization model, where the importance characterization model is used to output the importance features of the candidate voice based on the text content features and / or speaker features of the candidate voice.
[0128] The model training method provided by the embodiments of the present invention is applied to the model training of the importance characterization model. Here, the importance characterization model is used to extract the importance features of speech, and the importance features of speech can be used as an index to judge whether the speech is important or as an index to judge whether to perform speech recommendation. The embodiments of the present invention do not make specific limitations on this.
[0129] In the process of training the importance characterization model, multiple task training samples can be determined first. Here, the multiple tasks include the importance classification task, and also include the speech recognition task and / or the speaker recognition task.
[0130] Among them, for the importance classification task, after the importance features extracted for the speech, a classification layer is connected to realize the classification of whether the speech is important. For the speech recognition task, after the text content features extracted for the speech, a fully connected layer for speech recognition is connected to realize speech recognition. For the speaker recognition task, after the speaker features extracted for the speech, a classification layer is connected to realize the classification of the speaker to which the speech belongs.
[0131] For the importance classification task, the speech recognition task, and the speaker recognition task, their respective training samples can be determined separately. Among them, for the importance classification task, its training samples can include the speech labeled with the tag of whether it is important; for the speech recognition task, its training samples can include the speech labeled with the corresponding text; for the speaker recognition task, its training samples can include the speech labeled with the speaker tag.
[0132] When training the importance characterization model, it can be trained by combining the importance classification task and the speech recognition task. Through this training, the importance characterization model can extract the text content features of the input candidate speech, and can extract the importance features of the candidate speech based on the text content features of the candidate speech. The obtained importance features integrate the text content in the candidate speech and can reflect the importance of the candidate speech from two aspects: the text content and whether it is important.
[0133] Or, when training the importance characterization model, it can be trained by combining the importance classification task and the speaker recognition task. Through this training, the importance characterization model can extract the speaker features of the input candidate speech, and can extract the importance features of the candidate speech based on the speaker features of the candidate speech. The obtained importance features integrate the speaker information in the candidate speech and can reflect the importance of the candidate speech from two aspects: the speaker information and whether it is important.
[0134] Alternatively, when training the importance characterization model, it can be trained in combination with the importance classification task, the speech recognition task, and the speaker recognition task. Through this training, the importance characterization model can extract the text content features of the input candidate speech, as well as the speaker features of the candidate speech, and can extract the importance features of the candidate speech based on the text content features and speaker features of the candidate speech. The obtained importance features integrate the text content and speaker information in the candidate speech and can reflect the importance of the candidate speech from three aspects: text content, speaker information, and whether it is important.
[0135] That is, the importance characterization model obtained through training with multiple tasks can extract the importance features of the candidate speech, and the extracted importance features can reflect the importance of the candidate speech from aspects such as text content and / or speaker information, and whether it is important. Compared with the related art that only considers whether the candidate speech is important from the intention, the measurement of the importance of the candidate speech is more comprehensive and reliable.
[0136] The method provided by the embodiments of the present invention trains the importance characterization model through multiple tasks, providing a basis for extracting the importance features of the candidate speech and thus enabling end-to-end speech recommendation. There is no situation in the prior art where multiple links are cascaded and the accuracy of each link will affect the speech recommendation effect, which can improve the reliability of speech recommendation. Second, the importance characterization model obtained through training with multiple tasks can reflect the importance of the candidate speech from aspects such as text content and / or speaker information, and whether it is important when extracting the importance features, making the measurement of importance more comprehensive and reliable, and further ensuring the accuracy of speech recommendation.
[0137] Based on any of the above embodiments, in step 420, the training of the model to obtain the importance characterization model based on the training samples of multiple tasks includes:
[0138] Extracting the sample text content features from the speech features of the training samples of the speech recognition task based on the text content extraction network in the initial model, and / or extracting the sample speaker features from the speech features of the training samples of the speaker recognition task based on the speaker extraction network in the initial model;
[0139] Extracting the sample importance features from the speech features of the training samples of the importance classification task, as well as the sample text content features and / or sample speaker features of the training samples of the importance classification task based on the importance extraction network in the initial model;
[0140] Based on the speech recognition loss determined by the sample text content features and / or the speaker recognition loss determined by the sample speaker features, and the importance loss determined by the sample importance features, perform parameter iteration on the initial model to obtain the importance characterization model.
[0141] Specifically, the initial model is the importance characterization model before training. The initial model can be understood as the model after parameter initialization, and the initial model can include a text content extraction network and / or a speaker extraction network. Among them, the text content extraction network is the network used to extract text content features. In the multi-task training of the importance characterization model, the text content extraction network can be used to support the training of the speech recognition task; the speaker extraction network is the network used to extract speaker features. In the multi-task training of the importance characterization model, the speaker extraction network can be used to support the training of the speaker recognition task.
[0142] In specific implementation, for the text content extraction network, the speech features of the speech in the training samples of the speech recognition task can be used as the training input of the text content extraction network. The text content extraction network can extract the sample text content features from the speech features.
[0143] For the speaker extraction network, the speech features of the speech in the training samples of the speaker recognition task can be used as the training input of the speaker extraction network. The speaker extraction network can extract the sample speaker features from the speech features.
[0144] In addition, the initial model can also include an importance extraction network, and the input of the importance extraction network is connected to the output of the text content extraction network and / or the speaker extraction network. The importance extraction network can support the training of the importance classification task in the multi-task training of the importance characterization model.
[0145] In specific implementation, for the importance extraction network, the speech features of the speech in the training samples of the importance classification task, as well as the sample text content features and / or sample speaker features of the speech in the training samples of the importance classification task, can be used as the training input of the importance extraction network. The importance extraction network can extract the sample importance features from the speech features, as well as the sample text content features and / or sample speaker features.
[0146] After obtaining the sample text content features for the speech recognition task and / or the sample speaker features for the speaker recognition task, and the sample importance features for the importance classification task respectively, iterative training of the initial model parameters can be performed based on the above features, thereby obtaining the importance characterization model.
[0147] Specifically, for the sample text content features of the speech recognition task, the speech recognition text can be obtained through the fully connected layer of speech recognition, and the speech recognition text can be compared with the text label in the training sample to obtain the loss of the speech recognition task; for the sample speaker features of the speaker recognition task, the speaker classification result can be obtained through the classification layer of speaker recognition, and the speaker classification result can be compared with the speaker label in the training sample to obtain the loss of the speaker recognition task; for the sample importance features of the importance classification task, the classification result of whether it is important or not can be obtained through the importance classification layer, and the classification result of whether it is important or not can be compared with the label of whether it is important or not in the training sample to obtain the loss of the importance classification task.
[0148] Subsequently, the losses of the speech recognition task, the speaker recognition task, and the importance classification task can be combined to construct a loss function for the initial model, and based on this, parameter iteration can be performed on the initial model to obtain the importance characterization model.
[0149] Based on any of the above embodiments, a speech recommendation method may include two stages: model training and model application.
[0150] Among them, the model training stage may include the following steps:
[0151] First, collect training samples:
[0152] The training samples here include unlabeled business data and labeled speech data.
[0153] Among them, the unlabeled business data is in the form of speech and the effective duration is not less than 1000 hours; the labeled speech data may include open-source speech recognition data, speaker recognition data, and business data labeled as important / unimportant. For example, the speech recognition data may include data in major languages such as Chinese and English, and the speaker data requires no less than 10,000 individuals, etc.
[0154] Secondly, preprocess the collected training samples:
[0155] Specifically, it can be to filter out invalid data such as silent noise, leave the effective speech segments, and randomly split the effective speech segments into single segments with a duration between 0.5 seconds and 20 seconds, etc.
[0156] Then, perform pre-training:
[0157] The preprocessed unlabeled training samples can be randomly input into the pre-training model one batch (batch) of data at a time to train the pre-training model. The pre-training model obtained by this training can achieve speech feature extraction.
[0158] Subsequently, training for multiple tasks is carried out:
[0159] The pre-trained model can be connected to the networks of multiple tasks. For example, Figure 3 in [case], the outputs of the pre-trained network are respectively connected to the inputs of the text content extraction network, the importance extraction network, and the speaker extraction network. In this case, three batches of labeled training samples can be selected respectively, that is, speech recognition data, speaker recognition data, and business data labeled as important / unimportant, and sent into the pre-trained model. Then, the sample text content features, sample speaker features, and sample importance features are respectively extracted through the networks corresponding to the speech recognition task, the speaker recognition task, and the importance classification task.
[0160] Based on the sample text content features, sample speaker features, and sample importance features, the loss value of joint training is calculated, and the parameters of the model including the pre-trained model, the text content extraction network, the importance extraction network, and the speaker extraction network are iterated based on the loss value, thereby obtaining the importance characterization model.
[0161] Among them, the model application stage can include a registration stage and a recommendation stage. The registration stage includes the following steps:
[0162] First, registration voices are collected:
[0163] Here, the registration voices can include voices that are defaulted to be important and need to be recommended, as well as voices that are defaulted to be unimportant and not recommended. The registration voices can overlap with the business data labeled as important / unimportant in the training samples collected in the model training stage.
[0164] Second, preprocessing is performed on the registration voices:
[0165] Specifically, it can be to filter out invalid data such as silent noise, leave valid voice segments, and randomly split the valid voice segments into single segments with a duration between 0.5 seconds and 20 seconds, etc.
[0166] Then, the importance characterization model obtained by training is loaded, the importance features of the valid voice segments in the registration voices are respectively extracted, and the mean value of the importance features of the valid voice segments in the same registration voice is calculated as the importance feature of the registration voice.
[0167] Subsequently, clustering is performed on the importance features of the registration voices, thereby obtaining multiple registration voice clusters. The importance feature of each registration voice cluster is the mean value of the importance features of all registration voices under the registration voice cluster.
[0168] The recommendation stage includes the following steps:
[0169] First, obtain candidate voices to be recommended, and perform preprocessing on the candidate voices to obtain valid voice segments of the candidate voices.
[0170] Subsequently, load the importance characterization model obtained through training, extract importance features from the valid voice segments in the candidate voices respectively, and calculate the average value of the importance features of the valid voice segments in the same candidate voice as the importance feature of this candidate voice.
[0171] Next, perform clustering on the importance features of each candidate voice, thereby obtaining at least one voice cluster, and select the candidate voice with the longest valid voice segment from each voice cluster as the deduplicated voice.
[0172] Finally, set a preset threshold, and calculate the similarity between the importance feature of each deduplicated voice and the importance feature of each registered voice cluster obtained in the registration stage. For each deduplicated voice, compare the maximum value of the similarity with the preset threshold. If the maximum value is greater than or equal to the preset threshold, and the registered voice in the registered voice cluster corresponding to the maximum value is an important voice that needs to be recommended, then recommend this deduplicated voice as a recommended voice; otherwise, regard this deduplicated voice as an unimportant voice and do not recommend this deduplicated voice.
[0173] The voice recommendation method provided by the embodiments of the present invention proposes an importance characterization model obtained through training on multiple tasks. When characterizing the importance features, this model integrates information such as text content, speaker information, and importance, and can effectively improve the robustness of the importance features. And since this importance characterization model realizes the mapping from voice to importance features, there is no need to perform the process of voice transcription and then intent recognition during this process, and it also avoids the situation where multiple links are cascaded and the accuracy of each link will affect the voice recommendation effect.
[0174] In addition, in the voice recommendation method provided by the embodiments of the present invention, clustering is performed on the importance features of each candidate voice, thereby obtaining at least one voice cluster, and the candidate voice with the longest valid voice segment is selected from each voice cluster as the deduplicated voice, thereby realizing the deduplication operation for candidate voices and avoiding the waste of computing power on duplicate voices in voice recommendation.
[0175] Next, a description is given of the voice recommendation device provided by the present invention. The voice recommendation device described below can be correspondingly referred to the voice recommendation method described above.
[0176] Figure 5 is a schematic structural diagram of the voice recommendation device provided by the present invention, as Figure 5 shown, this device includes:
[0177] A voice acquisition unit 510, configured to acquire candidate voices;
[0178] A feature extraction unit 520, configured to input the candidate speech into the importance characterization model, and obtain the importance feature of the candidate speech output by the importance characterization model based on the text content feature and / or speaker feature of the candidate speech; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature;
[0179] A voice recommendation unit 530, configured to perform voice recommendation based on the importance feature of the candidate speech.
[0180] The voice recommendation device provided by the embodiment of the present invention extracts the importance feature of the candidate speech through an importance characterization model trained through multiple tasks, and performs voice recommendation based on this. On the one hand, it realizes end-to-end voice recommendation, and there is no situation where multiple links are cascaded in the prior art, and the accuracy of each link will affect the voice recommendation effect, which can improve the reliability of voice recommendation; on the other hand, the importance feature can reflect the importance of the candidate speech from aspects such as text content and / or speaker information, and whether it is important, making the measurement of importance more comprehensive and reliable, and further ensuring the accuracy of voice recommendation.
[0181] Based on any of the above embodiments, the feature extraction unit is specifically configured to:
[0182] Based on the text content extraction network in the importance characterization model and the speech feature of the candidate speech, extract the text content feature of the candidate speech, and / or, based on the speaker extraction network in the importance characterization model and the speech feature of the candidate speech, extract the speaker feature of the candidate speech;
[0183] Based on the importance extraction network in the importance characterization model, the speech feature, and the text content feature and / or the speaker feature, extract the importance feature of the candidate speech.
[0184] Based on any of the above embodiments, the voice recommendation unit is specifically configured to:
[0185] Cluster the importance features of each candidate speech to obtain at least one speech cluster;
[0186] Determine a candidate speech from each speech cluster as a deduplicated speech, and perform voice recommendation based on the importance feature of the deduplicated speech.
[0187] Based on any of the above embodiments, the voice recommendation unit is specifically configured to:
[0188] Performing voice recommendation based on the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster;
[0189] Each of the registered voice clusters is obtained by clustering the importance features of each registered voice.
[0190] Based on any of the above embodiments, the voice recommendation unit is specifically configured to:
[0191] When the maximum value of the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster is greater than a preset threshold, and the registered voice cluster corresponding to the maximum value is the recommended voice cluster, performing voice recommendation on the deduplicated voice.
[0192] Based on any of the above embodiments, the feature extraction unit is specifically configured to:
[0193] Inputting each voice segment of the candidate voice into the importance characterization model, and obtaining the importance features of each voice segment output by the importance characterization model based on the voice features and / or text content features of each voice segment;
[0194] Based on the importance features of each voice segment, determining the importance feature of the candidate voice.
[0195] The model training device provided by the present invention will be described below. The model training device described below can be correspondingly referred to the model training method described above.
[0196] Figure 6 is a schematic structural diagram of the model training device provided by the present invention, as Figure 6 shown. The device includes:
[0197] A sample acquisition unit 610, configured to determine training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features;
[0198] A model training unit 620, configured to perform model training based on the training samples of multiple tasks to obtain an importance characterization model, where the importance characterization model is used to output the importance features of the candidate voice based on the text content features and / or speaker features of the candidate voice.
[0199] The model training device provided by the embodiment of the present invention trains the importance characterization model through multiple tasks, providing a basis for extracting the importance features of candidate voices and performing voice recommendation based on this. Thus, end-to-end voice recommendation can be achieved, and there is no situation in the prior art where multiple links are cascaded and the accuracy of each link will affect the voice recommendation effect, which can improve the reliability of voice recommendation. Second, the importance characterization model obtained through training with multiple tasks can reflect the importance of candidate voices from aspects such as text content and / or speaker information, as well as whether it is important when extracting importance features, making the measurement of importance more comprehensive and reliable, and further ensuring the accuracy of voice recommendation.
[0200] Based on any of the above embodiments, the model training unit is specifically configured to:
[0201] Extract sample text content features from the voice features of the training samples of the speech recognition task based on the text content extraction network in the initial model, and / or extract sample speaker features from the voice features of the training samples of the speaker recognition task based on the speaker extraction network in the initial model;
[0202] Extract sample importance features from the voice features of the training samples of the importance classification task, as well as the sample text content features and / or sample speaker features of the training samples of the importance classification task based on the importance extraction network in the initial model;
[0203] Perform parameter iteration on the initial model based on the speech recognition loss determined by the sample text content features and / or the speaker recognition loss determined by the sample speaker features, and the importance loss determined by the sample importance features, to obtain the importance characterization model.
[0204] Figure 7 An example of the physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the voice recommendation method, which includes:
[0205] Obtain candidate voices;
[0206] Input the candidate speech into the importance characterization model to obtain the importance feature of the candidate speech output by the importance characterization model based on the text content feature and / or speaker feature of the candidate speech; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature;
[0207] Perform speech recommendation based on the importance feature of the candidate speech.
[0208] Alternatively, the processor 710 may call the logic instructions in the memory 730 to execute a model training method, which includes:
[0209] Determine the training samples for multiple tasks, where the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature;
[0210] Perform model training based on the training samples of multiple tasks to obtain an importance characterization model, which is used to output the importance feature of the candidate speech based on the text content feature and / or speaker feature of the candidate speech.
[0211] In addition, when the logic instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0212] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech recommendation method provided by the above-mentioned various methods, and the method includes:
[0213] Obtain a candidate speech;
[0214] Input the candidate speech into the importance characterization model to obtain the importance features of the candidate speech output by the importance characterization model based on the text content features and / or speaker features of the candidate speech; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content features and / or a speaker recognition task based on the speaker features, as well as an importance classification task based on the importance features;
[0215] Perform speech recommendation based on the importance features of the candidate speech.
[0216] Alternatively, a computer can execute the model training method provided by each of the above methods. This method includes:
[0217] Determine the training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, as well as an importance classification task based on importance features;
[0218] Perform model training based on the training samples of multiple tasks to obtain an importance characterization model, which is used to output the importance features of the candidate speech based on the text content features and / or speaker features of the candidate speech.
[0219] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the speech recommendation method provided by each of the above methods. This method includes:
[0220] Obtain a candidate speech;
[0221] Input the candidate speech into the importance characterization model to obtain the importance features of the candidate speech output by the importance characterization model based on the text content features and / or speaker features of the candidate speech; the importance characterization model is trained through multiple tasks, and the multiple tasks include a speech recognition task based on the text content features and / or a speaker recognition task based on the speaker features, as well as an importance classification task based on the importance features;
[0222] Perform speech recommendation based on the importance features of the candidate speech.
[0223] Alternatively, when the computer program is executed by a processor, it is configured to execute the model training method provided by each of the above methods. This method includes:
[0224] Determine the training samples for multiple tasks, where the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, as well as an importance classification task based on importance features;
[0225] Based on training samples of multiple tasks, model training is performed to obtain an importance characterization model, and the importance characterization model is used to output the importance feature of the candidate speech based on the text content feature and / or speaker feature of the candidate speech.
[0226] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0227] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0228] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice recommendation method, characterized in that: include: Obtain candidate voices; Inputting the candidate speech into an importance characterization model to obtain an importance feature of the candidate speech output by the importance characterization model based on text content features and / or speaker features of the candidate speech; The importance representation model is obtained through training of multiple tasks, wherein the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature, wherein the importance feature integrates the text content and / or speaker information in the candidate speech and is used to represent the importance of the candidate speech in the speech recommendation scenario from the aspects of text content and / or speaker information, and whether the text content and / or speaker information are important. Based on the importance features of the candidate voices, voice recommendations are performed; The step of inputting the candidate speech into an importance characterization model to obtain an importance feature of the candidate speech output by the importance characterization model based on a text content feature and / or a speaker feature of the candidate speech comprises: Extracting text content features of the candidate speech based on the text content extraction network in the importance representation model and the speech features of the candidate speech, and / or extracting speaker features of the candidate speech based on the speaker extraction network in the importance representation model and the speech features of the candidate speech; Importance features of the candidate speech are extracted based on the importance extraction network in the importance representation model, the speech features, and the text content features and / or the speaker features.
2. The voice recommendation method according to claim 1, characterized in that: The voice recommendation based on the importance feature of the candidate voices includes: Clustering the importance features of each candidate speech to obtain at least one speech cluster; A candidate voice is determined from each voice cluster as a deduplicated voice, and voice recommendation is performed based on the importance features of the deduplicated voice.
3. The voice recommendation method according to claim 2, characterized in that: The performing voice recommendation based on the importance feature of the deduplicated voice includes: Performing voice recommendation based on the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster; The registered speech clusters are obtained by clustering the importance features of the registered speech.
4. The voice recommendation method according to claim 3, characterized in that: The voice recommendation based on the similarity between the importance features of the deduplicated voice and the importance features of each registered voice cluster includes: When the maximum value of the similarities between the importance features of the deduplicated speech and the importance features of each registered speech cluster is greater than a preset threshold, and the registered speech cluster corresponding to the maximum value is a recommended speech cluster, speech recommendation is performed on the deduplicated speech.
5. The voice recommendation method according to any one of claims 1 to 4, characterized in that: The step of inputting the candidate speech into an importance characterization model to obtain an importance feature of the candidate speech output by the importance characterization model based on a text content feature and / or a speaker feature of the candidate speech comprises: Inputting each speech segment of the candidate speech into an importance characterization model, and obtaining importance features of each speech segment output by the importance characterization model based on speech features and / or text content features of each speech segment; Based on the importance features of the speech segments, the importance features of the candidate speech are determined.
6. A model training method, characterized in that: include: Determining training samples for a plurality of tasks, wherein the plurality of tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features; Based on the training samples of multiple tasks, model training is performed to obtain an importance characterization model, wherein the importance characterization model is used to output the importance feature of the candidate voice based on the text content feature and / or speaker feature of the candidate voice, wherein the importance feature integrates the text content and / or speaker information in the candidate voice, and is used to characterize the importance of the candidate voice in the voice recommendation scenario from the aspects of text content and / or speaker information, and whether the candidate voice is important; The importance representation model includes a text content extraction network and / or a speaker extraction network, wherein the text content extraction network is used to extract the text content features of the candidate speech based on the speech features of the candidate speech, and the speaker extraction network is used to extract the speaker features of the candidate speech based on the speech features of the candidate speech. The importance representation model also includes an importance extraction network, which is used to extract the importance features of the candidate speech based on the speech features, the text content features and / or the speaker features.
7. The model training method according to claim 6, characterized in that: The training samples based on multiple tasks are used to perform model training to obtain an importance representation model, including: Extracting sample text content features from speech features of training samples of the speech recognition task based on a text content extraction network in the initial model, and / or extracting sample speaker features from speech features of training samples of the speaker recognition task based on a speaker extraction network in the initial model; Extracting sample importance features from speech features of the training samples of the importance classification task, and sample text content features and / or sample speaker features of the training samples of the importance classification task based on the importance extraction network in the initial model; Based on the speech recognition loss determined by the sample text content features and / or the speaker recognition loss determined by the sample speaker features, and the importance loss determined by the sample importance features, the initial model is iterated on parameters to obtain the importance representation model.
8. A voice recommendation device, characterized in that: include: A speech acquisition unit, used for acquiring candidate speech; A feature extraction unit, configured to input the candidate speech into an importance characterization model, and obtain an importance feature of the candidate speech output by the importance characterization model based on a text content feature and / or a speaker feature of the candidate speech; The importance representation model is obtained through training of multiple tasks, wherein the multiple tasks include a speech recognition task based on the text content feature and / or a speaker recognition task based on the speaker feature, and an importance classification task based on the importance feature, wherein the importance feature integrates the text content and / or speaker information in the candidate speech and is used to represent the importance of the candidate speech in the speech recommendation scenario from the aspects of text content and / or speaker information, and whether the text content and / or speaker information are important. A voice recommendation unit, used for making voice recommendations based on the importance features of the candidate voices; The feature extraction unit is specifically used for: Extracting text content features of the candidate speech based on the text content extraction network in the importance representation model and the speech features of the candidate speech, and / or extracting speaker features of the candidate speech based on the speaker extraction network in the importance representation model and the speech features of the candidate speech; Importance features of the candidate speech are extracted based on the importance extraction network in the importance representation model, the speech features, and the text content features and / or the speaker features.
9. A model training device, characterized in that: include: A sample acquisition unit, used to determine training samples for multiple tasks, wherein the multiple tasks include a speech recognition task based on text content features and / or a speaker recognition task based on speaker features, and an importance classification task based on importance features; A model training unit, configured to perform model training based on training samples of multiple tasks to obtain an importance characterization model, wherein the importance characterization model is configured to output an importance feature of the candidate voice based on the text content feature and / or speaker feature of the candidate voice, wherein the importance feature integrates the text content and / or speaker information in the candidate voice and is configured to characterize the importance of the candidate voice in a voice recommendation scenario from the aspects of the text content and / or speaker information and whether the candidate voice is important; The importance representation model includes a text content extraction network and / or a speaker extraction network, wherein the text content extraction network is used to extract the text content features of the candidate speech based on the speech features of the candidate speech, and the speaker extraction network is used to extract the speaker features of the candidate speech based on the speech features of the candidate speech. The importance representation model also includes an importance extraction network, which is used to extract the importance features of the candidate speech based on the speech features, the text content features and / or the speaker features.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the voice recommendation method as described in any one of claims 1 to 5, or the model training method as described in claim 6 or 7 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice recommendation method according to any one of claims 1 to 5, or the model training method according to claim 6 or 7 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voice recommendation method according to any one of claims 1 to 5, or the model training method according to claim 6 or 7 is implemented.
Citation Information
Patent Citations
Music recommendation method and device, electronic equipment and storage medium
CN113297412A