Model training method, voice interaction method, server and medium
By filtering and labeling target voice requests in vehicle voice interaction, the problem of poor training effect of vehicle voice interaction models in the existing technology is solved, and efficient model training and performance improvement is achieved.
Patent Information
- Application Number
- CN202510320373.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
The existing technology is difficult to effectively train the vehicle voice interaction model, resulting in poor execution of the vehicle voice interaction function and affecting the user's driving experience.
By acquiring multiple first voice requests, the processing result of each first voice request is determined using the pre-trained second voice request processing model, the target voice request is filtered out, and used as training samples and tags, and the second voice request processing model is trained.
The screening of multiple first voice requests is realized, and the high-quality voice requests are retained, the demand for manual annotation is reduced, the labeling efficiency is improved, and the full training and performance improvement of the second voice request processing model is ensured.
Smart Images

Figure CN120183387A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice interaction, and particularly relates to a training method for a voice request processing model for voice interaction, a voice interaction method, a server, and a computer-readable storage medium. Background Art
[0002] In the related art, a vehicle can implement an in-vehicle voice interaction function through a pre-trained natural language processing model, enabling a user to control the vehicle to perform corresponding actions through voice commands, such as turning on the air conditioner. However, the performance of the model is related to the quantity and quality of the samples used during model training, and it is difficult to obtain high-quality samples suitable for in-vehicle voice interaction. Therefore, both the training effect and performance of the model are difficult to meet expectations, resulting in a poor execution effect of the in-vehicle voice interaction function implemented based on the model, ultimately affecting the user's driving experience in the vehicle. Summary of the Invention
[0003] The present application provides a training method for a voice request processing model for voice interaction, a voice interaction method, a server, and a computer-readable storage medium.
[0004] A training method for a voice request processing model for voice interaction provided by an embodiment of the present application includes:
[0005] Obtain a plurality of first voice requests;
[0006] Determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models;
[0007] Determine a target voice request among the plurality of first voice requests according to the first processing results;
[0008] Train a pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0009] Thus, in the embodiments of the present application, in the case of obtaining a plurality of first voice requests with unknown tags and unknown quality, a plurality of first processing results of each first voice request can be determined through a pre-trained first voice request processing model, and a target voice request can be determined from the plurality of first voice requests through the plurality of first processing results of each first voice request. Thus, to a certain extent, voice requests with lower quality among the plurality of first voice requests can be removed and voice requests with higher quality, that is, the target voice request, can be retained, thereby realizing the screening of the plurality of first voice requests. Moreover, since the first processing result of the first voice request is determined according to the first voice request processing model, the first processing result of the target voice request can, to a certain extent, be used as the tag of the target voice request, thereby completing the annotation of the target voice request. Therefore, the manual participation link in the annotation process of the target voice request is reduced, and the annotation efficiency of the target voice request is improved. At the same time, the target voice request and the first processing result of the target voice request can be used as the training sample and sample tag of the second voice request processing model, so as to train the second voice request processing model, enabling the training sample and sample tag of the second voice request processing model to be obtained in an efficient manner, ensuring the acquisition efficiency, thereby realizing the sufficient training of the second voice request processing model, and further ensuring the performance of the second voice request processing model.
[0010] In some embodiments of the present application, determining the target voice request among the plurality of first voice requests according to the first processing result includes:
[0011] Determining a plurality of second voice requests among the plurality of first voice requests according to the plurality of first processing results of the first voice request;
[0012] Determining the target voice request among the plurality of second voice requests according to the first processing result of the second voice request.
[0013] Thus, in the embodiments of the present application, a plurality of second voice requests among the plurality of first voice requests can be determined according to the plurality of first processing results of the first voice request, and the target voice request among the plurality of second voice requests can be determined according to the first processing result of the second voice request, so that the target voice request can be determined based on a two-stage screening method, ensuring the effectiveness and reliability of the target voice request to a certain extent.
[0014] In some embodiments of the present application, determining a plurality of second voice requests among the plurality of first voice requests according to the plurality of first processing results of the first voice request includes:
[0015] If the similarity between any two of the multiple first processing results of the first voice request is greater than or equal to a preset similarity threshold, the first voice request is determined as the second voice request.
[0016] Thus, in the embodiment of the present application, among multiple first voice requests, a first voice request in which the similarity between any two first processing results is greater than or equal to the preset similarity threshold can be determined as the second voice request, and the quality of the second voice request can be guaranteed. Therefore, the quality of the finally determined target voice request can also be guaranteed, thereby improving the model training effect and model performance of the second voice request processing model to a certain extent.
[0017] In some embodiments of the present application, the first processing result includes a first text prediction. Determining a target voice request among multiple second voice requests according to the first processing result of the second voice request includes:
[0018] Determining the target voice request according to a third voice request and a preset voice request discard probability, where the third voice request is a voice request among multiple second voice requests in which the first text prediction matches a first preset text.
[0019] Thus, in the embodiment of the present application, the target voice request can be determined according to the third voice request and the preset voice request discard probability, so that some third voice requests can be used as target voice requests to participate in model training, and the model performance can be guaranteed to a certain extent.
[0020] In some embodiments of the present application, the first processing result includes a second text prediction, and each second voice request corresponds to at least one preset scenario information. Determining a target voice request among multiple second voice requests according to the first processing result of the second voice request includes:
[0021] When the second text prediction of the second voice request matches a second preset text, or the scenario information corresponding to the second voice request is target scenario information, the second voice request is determined as the target voice request, where the second preset text and the target scenario information correspond to a voice request with an incorrect prediction by the second voice request processing model.
[0022] Thus, in the embodiments of the present application, when the second text prediction of the second voice request matches the second preset text, or the scenario information corresponding to the second voice request is the target scenario information, the second voice request can be determined as the target voice request, so that the target voice request for training the second voice request processing model can be related to the voice request with incorrect prediction by the second voice request processing model. Furthermore, based on the training of the target voice request, the defects in the performance of the second voice request processing model can be compensated, and the performance of the second voice request processing model can be steadily improved through training.
[0023] In some embodiments of the present application, the obtaining of the multiple first voice requests includes:
[0024] Performing natural language processing on multiple fourth voice requests according to the second voice request processing model to determine a second processing result for each of the fourth voice requests;
[0025] Determining the first voice request among the multiple fourth voice requests according to the second processing result.
[0026] Thus, in the embodiments of the present application, natural language processing can be performed on multiple fourth voice requests according to the second voice request processing model to determine a second processing result for each fourth voice request, and to determine the first voice request among the multiple fourth voice requests according to the second processing result, thereby ensuring the quality of the first voice request to a certain extent.
[0027] In some embodiments of the present application, the second processing result includes a third text prediction and a prediction probability of the third text prediction. The determining of the first voice request among the multiple fourth voice requests according to the second processing result includes:
[0028] When the third text prediction of the fourth voice request does not match the third preset text, determining the fourth voice request as the first voice request; and / or
[0029] When the third text prediction of the fourth voice request matches the third preset text and the prediction probability of the third text prediction of the fourth voice request is less than or equal to a preset probability threshold, determining the fourth voice request as the first voice request.
[0030] Thus, in the embodiments of the present application, the fourth voice request can be determined as the first voice request when the third text prediction of the fourth voice request does not match the third preset text, and / or the fourth voice request can be determined as the first voice request when the third text prediction of the fourth voice request matches the third preset text and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold, thereby further ensuring the effectiveness of the first voice request.
[0031] An embodiment of the present application provides a voice interaction method, including:
[0032] Obtain a current voice request;
[0033] Perform natural language processing on the current voice request according to a third voice request processing model to determine a third processing result of the current voice request, where the third voice request processing model is trained by the above-mentioned training method of the voice request processing model for voice interaction;
[0034] Determine a vehicle control instruction according to the third processing result;
[0035] Send the vehicle control instruction to the vehicle to perform the voice interaction.
[0036] In this way, in the embodiment of the present application, the third voice request processing model determined through training can be used to perform natural language processing on the obtained current voice request, so as to determine the third processing result of the current voice request, and the vehicle control instruction can be determined through the third processing result of the current voice request and sent to the vehicle, thereby performing voice interaction, thus realizing in-vehicle voice interaction, and to a certain extent, ensuring the stable execution of in-vehicle voice interaction. Moreover, for the model training process, in the case of obtaining multiple first voice requests with unknown labels and unknown quality, the multiple first processing results of each first voice request can be determined through the pre-trained first voice request processing model, and the target voice request can be determined from the multiple first voice requests through the multiple first processing results of each first voice request. Thus, to a certain extent, the voice requests with lower quality in the multiple first voice requests can be eliminated and the voice requests with higher quality, that is, the target voice requests, can be retained, thereby realizing the screening of the multiple first voice requests. And because the first processing result of the first voice request is determined according to the first voice request processing model, the first processing result of the target voice request can be used as the label of the target voice request to a certain extent, so that the annotation of the target voice request can be completed. Therefore, the manual participation link in the annotation process of the target voice request is reduced, and the annotation efficiency of the target voice request is improved. At the same time, the target voice request and the first processing result of the target voice request can be used as the training samples and sample labels of the second voice request processing model, so as to train the second voice request processing model, so that the training samples and sample labels of the second voice request processing model can be obtained in an efficient manner, and the acquisition efficiency is guaranteed. Thus, the sufficient training of the second voice request processing model can be realized, and further, the performance of the second voice request processing model is guaranteed.
[0037] An embodiment of the present application provides a server, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the training method of the above-mentioned voice request processing model for voice interaction is implemented, or the above-mentioned voice interaction method is implemented.
[0038] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by one or more processors, the training method of the above-mentioned voice request processing model for voice interaction is implemented, or the above-mentioned voice interaction method is implemented.
[0039] The server and the computer-readable storage medium provided by the embodiment of the present application can, in the case of obtaining a plurality of first voice requests with unknown tags and unknown quality, determine a plurality of first processing results of each first voice request through a pre-trained first voice request processing model, and determine a target voice request from the plurality of first voice requests through the plurality of first processing results of each first voice request. Thus, to a certain extent, the voice requests with lower quality in the plurality of first voice requests can be eliminated and the voice requests with higher quality, that is, the target voice requests, can be retained, thereby realizing the screening of the plurality of first voice requests. And, because the first processing result of the first voice request is determined according to the first voice request processing model, the first processing result of the target voice request can, to a certain extent, be used as the tag of the target voice request. Thus, the annotation of the target voice request can be completed. Therefore, the manual participation link in the annotation process of the target voice request is reduced, and the annotation efficiency of the target voice request is improved. At the same time, the target voice request and the first processing result of the target voice request can be used as the training samples and sample tags of the second voice request processing model, so as to train the second voice request processing model, so that the training samples and sample tags of the second voice request processing model can be obtained in an efficient manner, and the acquisition efficiency is guaranteed. Thus, the sufficient training of the second voice request processing model can be realized, and further, the performance of the second voice request processing model is guaranteed.
[0040] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0042] Figure 1 is a schematic flowchart of a training method of a voice request processing model for voice interaction in some embodiments of the present application;
[0043] Figure 2 This is a schematic flowchart of a training method for a voice request processing model for voice interaction in some embodiments of the present application;
[0044] Figure 3 This is a schematic flowchart of a training method for a voice request processing model for voice interaction in some embodiments of the present application;
[0045] Figure 4 This is a schematic diagram of an application scenario in some embodiments of the present application;
[0046] Figure 5 This is a schematic flowchart of a voice interaction method in some embodiments of the present application. Detailed implementation manners
[0047] The following details the implementation manners of the present application. The examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the implementation manners of the present application, and should not be construed as a limitation to the implementation manners of the present application.
[0048] In the in-vehicle voice interaction scenario, to ensure that the voice commands of the user can be accurately responded to by the vehicle, the vehicle can capture the audio signal in the cockpit space. For example, after the user says the sentence "Turn on the air conditioner", the vehicle can upload the captured audio signal to the cloud. After receiving the audio signal sent by the vehicle, the cloud can call the pre-trained recognition model to recognize the audio signal, so as to determine the text corresponding to the audio signal, and then can perform operations such as slot recognition and vehicle control command generation based on the text, and send the finally generated vehicle control command to the vehicle, so that the vehicle can perform corresponding operations based on the received command, such as turning on the vehicle air conditioner.
[0049] It can be understood that to improve the voice interaction instructions in the vehicle, the recognition model in the cloud can be updated frequently. Specifically, the recognition model is greatly affected by factors such as the in-vehicle front end and vehicle models. For example, when the front end is updated or a new vehicle model is added, the recognition model often needs to be updated and optimized synchronously. Also, with the update of the in-vehicle speech product set and the adjustment of online BadCases or internal test BadCases, the cloud recognition model also needs to be updated to optimize and solve cases.
[0050] It can also be understood that each update of the cloud recognition model involves a large amount of consumption of human resources, such as model fine-tuning, model self-testing, online regression testing, etc.
[0051] Based on the above possible problems, please refer to Figure 1, an embodiment of the present application provides a method for training a voice request processing model for voice interaction, including:
[0052] 01: Obtain a plurality of first voice requests;
[0053] 02: Determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models;
[0054] 03: Determine a target voice request among the plurality of first voice requests according to the first processing results;
[0055] 04: Train a pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0056] An embodiment of the present application provides a training device for a voice request processing model for voice interaction. The training method for the voice request processing model for voice interaction in the embodiment of the present application can be implemented by the training method for the voice request processing model for voice interaction in the embodiment of the present application. Specifically, the training device includes an acquisition module, a processing result determination module, a voice request determination module, and a training module. Among them, the acquisition module is used to obtain a plurality of first voice requests. The processing result determination module is used to determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models. The voice request determination module is used to determine a target voice request among the plurality of first voice requests according to the first processing results. The training module is used to train a pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0057] An embodiment of the present application further provides a server, which includes a memory and a processor. The training method for the voice request processing model for voice interaction in the embodiment of the present application can be implemented by the server in the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain a plurality of first voice requests, and to determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models, and to determine a target voice request among the plurality of first voice requests according to the first processing results, and to train a pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0058] Specifically, an embodiment of the present application proposes a solution for automatically updating, optimizing, and iteratively testing a model using online semi-supervised desensitized data. Through a complete closed-loop Pipeline, it can achieve continuous online data collection and processing, regular model iteration, and automated testing, thereby reducing manual input and realizing automated model updates.
[0059] Specifically, in the embodiment of the present application, the server can obtain a plurality of first voice requests, and when obtaining these plurality of first voice requests, call a plurality of pre-trained second voice request processing models, so that these plurality of second voice request processing models respectively perform natural language processing on each first voice request. Furthermore, for each first voice request, one first voice request corresponds to a plurality of first processing results.
[0060] Next, according to the plurality of first processing results of each first voice request, determine the target voice requests that can ultimately be used for model training from all the obtained first voice requests.
[0061] Finally, training can be performed based on the target voice requests and one or more first processing results of the target voice requests. Or rather, use the target voice requests as samples and one or more first processing results of the target voice requests as sample labels, thereby constructing a training set to train the pre-set second voice request processing model.
[0062] In this way, in the embodiment of the present application, when obtaining a plurality of first voice requests with unknown labels and unknown quality, the plurality of first processing results of each first voice request can be determined through the pre-trained first voice request processing model, and the target voice requests can be determined from the plurality of first voice requests through the plurality of first processing results of each first voice request. Thus, to a certain extent, the voice requests with lower quality in the plurality of first voice requests can be eliminated and the voice requests with higher quality, that is, the target voice requests, can be retained, thereby realizing the screening of the plurality of first voice requests. Moreover, because the first processing results of the first voice requests are determined according to the first voice request processing model, the first processing results of the target voice requests can, to a certain extent, be used as the labels of the target voice requests, thereby completing the annotation of the target voice requests. Therefore, the manual participation link in the annotation process of the target voice requests is reduced, and the annotation efficiency of the target voice requests is improved. At the same time, the target voice requests and the first processing results of the target voice requests can be used as the training samples and sample labels of the second voice request processing model, so as to train the second voice request processing model, enabling the training samples and sample labels of the second voice request processing model to be obtained in an efficient manner, ensuring the acquisition efficiency, and thus realizing the full training of the second voice request processing model, and further ensuring the performance of the second voice request processing model.
[0063] In an example, the first voice request is online data, or rather, for the voice request obtained by the vehicle during the voice interaction between the vehicle and the user, after processing the voice request such as desensitization, etc., the first voice request in the embodiment of the present application can be obtained.
[0064] In one example, the second voice processing model is a large language model (LLM), such as open-source large language models like Paraformer, SenseVoice, and Whisper v3.
[0065] In one example, the server can call the three pre-trained large language models, Paraformer, SenseVoice, and Whisper v3, to perform speech recognition on each obtained first voice request, and thus can obtain three speech recognition results for each first voice request, or rather, can obtain three processing results for each first voice request.
[0066] In one example, the first processing result of the first voice request refers to the text recognition result of the first voice request. Furthermore, for any first voice request, if the multiple first processing results of the first voice request are the same text, or if the multiple first processing results of a voice request are multiple texts with the same or similar semantics, then the first voice request can be used as the target voice request. Conversely, if the multiple first processing results of the first voice request are not the same text, or rather, the multiple first processing results of the first voice request are not multiple texts with the same or similar semantics, then the first voice request cannot be used as the target voice request.
[0067] In one example, based on the target voice request and any one of the multiple first processing results of the target voice request, they can be used as the voice request sample and the sample label of this voice request sample respectively, thereby constructing a training set and training the second voice request processing model.
[0068] Please refer to Figure 2 , in some embodiments of the present application, step 03 includes:
[0069] 030: Determine multiple second voice requests among the multiple first voice requests according to the multiple first processing results of the first voice request;
[0070] 031: Determine the target voice request among the multiple second voice requests according to the first processing result of the second voice request.
[0071] The voice request determination module in the embodiment of the present application is further configured to determine multiple second voice requests among the multiple first voice requests according to the multiple first processing results of the first voice request, and to determine the target voice request among the multiple second voice requests according to the first processing result of the second voice request.
[0072] The processor according to the embodiment of the present application is further configured to determine multiple second voice requests among the multiple first voice requests based on multiple first processing results of the first voice requests, and to determine a target voice request among the multiple second voice requests based on the first processing result of the second voice request.
[0073] Specifically, in the embodiment of the present application, the server can determine the target voice request from all the obtained first voice requests through a two-stage screening step.
[0074] Specifically, after the server determines multiple first processing results of each first voice request through multiple pre-trained first voice request processing models, for any one first voice request, the server can determine whether the first voice request can be used as a second voice request based on the multiple first processing results of the first voice request, thereby completing the first-stage screening.
[0075] Then, after the server screens out the second voice requests from all the first voice requests, for each second voice request, the server can determine whether the second voice request can be used as the target voice request based on the first processing result of the second voice request, thereby completing the second-stage screening.
[0076] In this way, in the embodiment of the present application, multiple second voice requests among the multiple first voice requests can be determined based on the multiple first processing results of the first voice requests, and the target voice request among the multiple second voice requests can be determined based on the first processing result of the second voice request, so that the target voice request can be determined based on a two-stage screening method, which to a certain extent ensures the effectiveness and reliability of the target voice request.
[0077] In some embodiments of the present application, step 030 includes:
[0078] Determining a target voice request according to a third voice request and a preset voice request discard probability, where the third voice request is a voice request among the multiple second voice requests whose first text prediction matches a first preset text.
[0079] The voice request determination module according to the embodiment of the present application is further configured to determine a target voice request according to a third voice request and a preset voice request discard probability, where the third voice request is a voice request among the multiple second voice requests whose first text prediction matches a first preset text.
[0080] The processor according to the embodiment of the present application is further configured to determine a target voice request according to a third voice request and a preset voice request discard probability, where the third voice request is a voice request among the multiple second voice requests whose first text prediction matches a first preset text.
[0081] Specifically, in the embodiment of the present application, after the server determines multiple first processing results of each first voice request through multiple pre-trained first voice request processing models, for any one first voice request, the server can determine whether each first voice request processing model outputs the same or similar prediction results according to the multiple first processing results of the first voice request, so as to determine whether the first voice request can be used as a second voice request.
[0082] For example, if the number of first voice request processing models is 3, the server can call these 3 first voice request processing models to perform natural language processing on the first voice request Q respectively, so as to obtain 3 first processing results of the first voice request, and let these 3 first processing results be R1, R2 and R3 in sequence.
[0083] Next, for any one first voice request, the server can calculate the similarity between every two of the 3 first processing results of the first voice request, such as the similarity S1 between R1 and R2, the similarity S2 between R1 and R3, and the similarity S3 between R2 and R3.
[0084] Then, if any one of S1, S2 and S3 is greater than the pre-set similarity threshold, it can be considered that the 3 first processing results obtained after the 3 first voice request processing models perform natural language processing on the first voice request Q are similar or the same. In other words, different voice request processing models can output the same or similar answers for the same first voice request Q, indicating that the first voice request Q can be mapped to a standard answer, so the quality of the first voice request Q is relatively high.
[0085] In an example, the first processing result is a text prediction result. In other words, the first voice request processing model can perform speech recognition on the first voice request to predict the text corresponding to the first voice request. Furthermore, for the text prediction results of multiple first voice request processing models for the same first voice request, the server can calculate the similarity between every two of these multiple text prediction results, such as cosine similarity, to determine whether these multiple text prediction results are semantically close and the same.
[0086] In an example, the pre-set similarity threshold is 95%.
[0087] In this way, in the embodiment of the present application, among multiple first voice requests, the first voice request whose similarity between any two first processing results is greater than or equal to the pre-set similarity threshold can be determined as the second voice request, and the quality of the second voice request can be guaranteed. Therefore, the quality of the finally determined target voice request can also be guaranteed, thereby improving the model training effect and model performance of the second voice request processing model to a certain extent.
[0088] In some embodiments of the present application, the first processing result includes a first text prediction. Furthermore, step 031 includes:
[0089] For a third voice request among multiple second voice requests whose first text prediction matches a first preset text, a target voice request is determined according to a preset voice request discard probability and the third voice request.
[0090] The voice request determination module according to the embodiment of the present application is further configured to determine a target voice request according to a preset voice request discard probability and a third voice request for a third voice request among multiple second voice requests whose first text prediction matches a first preset text.
[0091] The processor according to the embodiment of the present application is further configured to determine a target voice request according to a preset voice request discard probability and a third voice request for a third voice request among multiple second voice requests whose first text prediction matches a first preset text.
[0092] Specifically, in the embodiment of the present application, for the second voice request in the first voice request, the server can determine whether the text prediction of the second voice request is similar to or the same as a preset text. If so, it indicates that the second voice request is a high-frequency voice request with a large quantity and frequently triggered by users. Therefore, the server can perform a random discard process on the second voice request according to the preset voice request discard probability, and use the finally undiscarded second voice request as the target voice request.
[0093] Specifically, in the embodiment of the present application, the first processing result is a text prediction result, that is, the first text prediction. Further, it can be understood that the first voice request processing model can perform speech recognition on the first voice request to predict the text corresponding to the first voice request.
[0094] Further, in the embodiment of the present application, the first preset text can be understood as a text that appears frequently in the in-vehicle voice interaction scenario, such as "turn on the air conditioner", "close the window", etc.
[0095] Even further, in the embodiment of the present application, after the server determines that the first text prediction of the second voice request matches the first preset text through text matching or other means, the second voice request is randomly discarded according to the preset voice request discard probability.
[0096] For example, when the discard probability of the voice request is 0.9, the server can process the second voice request based on a preset program or code, etc., so that the second voice request has a 0.9 probability of being discarded and a 0.1 probability of not being discarded. For example, when there are 100 second voice requests in which the first text prediction matches the first preset text (i.e., the third voice request) among all the second voice requests, then among these 100 second voice requests, 10 second voice requests may be used as target voice requests, and 90 second voice requests may be discarded.
[0097] It can be understood that when there are multiple third voice requests, due to the preset voice request discard probability, some of the multiple third voice requests can be used as target voice requests to participate in model training.
[0098] It can also be understood that when the first voice request is online data, or in other words, for the voice request obtained by the vehicle during the voice interaction between the vehicle and the user, the voice request can be processed such as desensitization to obtain the first voice request in the embodiment of the present application.
[0099] Further, in the embodiment of the present application, the first preset text can be understood as texts that appear frequently in the in-vehicle voice interaction scenario, such as "turn on the air conditioner", "close the window", etc.
[0100] Furthermore, it can be understood that since most of the voice requests in the online data may correspond to texts that appear frequently, such as "turn on the air conditioner", "close the window", etc. In other words, when multiple second voice requests are determined from multiple first voice requests, most of the text predictions of the second voice requests may be these texts that appear frequently. Therefore, based on the setting of the voice request discard probability in the embodiment of the present application, a part of the third voice requests in the second voice requests whose "text predictions are these texts that appear frequently" can participate in model training. Thus, after model training, it can be applicable to the processing of "voice requests corresponding to 'these texts that appear frequently'".
[0101] In this way, in the embodiment of the present application, the target voice request can be determined according to the third voice request and the preset voice request discard probability, so that some of the third voice requests can be used as target voice requests to participate in model training, and the model performance can be guaranteed to a certain extent.
[0102] In some embodiments of the present application, the first processing result includes a second text prediction, and each second voice request corresponds to at least one preset scenario information. Furthermore, step 031 includes:
[0103] In the case where the second text prediction of the second voice request matches the second preset text, or the scenario information corresponding to the second voice request is the target scenario information, the second voice request is determined as the target voice request, where the second preset text and the target scenario information correspond to the voice request for which the second voice request processing model makes an incorrect prediction.
[0104] The voice request determination module according to the embodiment of the present application is further configured to determine the second voice request as the target voice request in the case where the second text prediction of the second voice request matches the second preset text, or the scenario information corresponding to the second voice request is the target scenario information, where the second preset text and the target scenario information correspond to the voice request for which the second voice request processing model makes an incorrect prediction.
[0105] The processor according to the embodiment of the present application is further configured to determine the second voice request as the target voice request in the case where the second text prediction of the second voice request matches the second preset text, or the scenario information corresponding to the second voice request is the target scenario information, where the second preset text and the target scenario information correspond to the voice request for which the second voice request processing model makes an incorrect prediction.
[0106] Specifically, in the embodiment of the present application, the second preset text and the target scenario information corresponding to the voice request can be determined according to the voice request for which the second voice request processing model makes an incorrect prediction. For example, in the case where the weather is rainy and the vehicle is in a high-speed driving state (such as the vehicle speed is greater than 60 kilometers per hour), if the user utters the voice command "Play the songs that should be listened to on a rainy day", and the vehicle is unable to determine the text "Play the songs that should be listened to on a rainy day" according to the voice request based on the second voice request processing model deployed locally or in the cloud, then the text "Play the songs that should be listened to on a rainy day" corresponding to the voice request will be used as the second preset text, and the scenarios "rainy day" and "high speed" corresponding to the voice request will be used as the target scenario information.
[0107] It can be understood that, in the embodiment of the present application, the first voice request processing model can perform voice recognition on the first voice request to recognize the text prediction of the first voice request, that is, the second text prediction.
[0108] It can also be understood that the acquisition method of the scenario information can be set according to the actual situation. For example, in one example, when the vehicle sends a voice request to the server, it can actively send the scenario corresponding to the voice request. For example, when the vehicle speed is higher than 60 kilometers per hour and the current weather is rainy, when the vehicle sends a voice request to the server, it can also synchronously send the scenario information of the voice request as "rainy day" and "high speed". Therefore, if a certain voice request is not correctly recognized
[0109] Furthermore, for a voice request for which the prediction of the second voice request processing model is incorrect, the server can screen each second voice request according to the correct text corresponding to the voice request, that is, the second preset text, and in combination with the scenario information corresponding to the voice request, that is, the target scenario information, so as to screen out second voice requests whose scenario information is the same as that of the "voice request for which the prediction of the second voice request processing model is incorrect", and second voice requests whose second text prediction is the same as / similar to the "second preset text of the voice request for which the prediction of the second voice request processing model is incorrect", so as to use these screened second voice requests as target voice requests to train the second voice request processing model. Thus, the defects in the performance of the second voice request processing model can be compensated, and the performance of the second voice request processing model can be improved.
[0110] In this way, in the embodiment of the present application, when the second text prediction of the second voice request matches the second preset text, or the scenario information corresponding to the second voice request is the target scenario information, the second voice request can be determined as the target voice request, so that the target voice request for training the second voice request processing model can be related to the voice request for which the prediction of the second voice request processing model is incorrect. Furthermore, based on the training of the target voice request, the defects in the performance of the second voice request processing model can be compensated, and the performance of the second voice request processing model can be steadily improved through training.
[0111] Please refer to Figure 3 , in some embodiments of the present application, step 01 includes:
[0112] 010: Perform natural language processing on multiple fourth voice requests according to the second voice request processing model to determine the second processing result of each fourth voice request;
[0113] 011: Determine the first voice request among the multiple fourth voice requests according to the second processing result.
[0114] The acquisition module in the embodiment of the present application is further configured to perform natural language processing on multiple fourth voice requests according to the second voice request processing model to determine the second processing result of each fourth voice request, and determine the first voice request among the multiple fourth voice requests according to the second processing result.
[0115] The processor in the embodiment of the present application is further configured to perform natural language processing on multiple fourth voice requests according to the second voice request processing model to determine the second processing result of each fourth voice request, and determine the first voice request among the multiple fourth voice requests according to the second processing result.
[0116] Specifically, in the embodiments of the present application, for the speech request processing model to be trained or updated, or rather, for the second speech request processing model, when the server obtains multiple fourth speech requests, it can call the second speech request processing model to perform natural language processing on each fourth speech request, so as to determine the natural language processing result of each fourth speech request, that is, the second processing result.
[0117] Then, for each fourth speech request, based on the second processing result of each fourth speech request, among all the obtained fourth speech requests, some fourth speech requests are used as speech request samples available for candidates, that is, the above-mentioned first speech requests.
[0118] In one example, the second speech request processing model is applicable to the speech recognition task. Furthermore, the second speech request processing model can perform speech recognition on the fourth speech request. Therefore, the second processing result of the fourth speech request can be understood as the speech recognition result of the fourth speech request, such as the text recognition result of the fourth speech request.
[0119] In one example, the second processing result of the fourth speech request is used to indicate whether the fourth speech request corresponds to a text, or rather, to indicate whether the second speech request processing model can recognize the text corresponding to the fourth speech request.
[0120] Furthermore, for a certain fourth speech request, if the second speech request processing model can recognize the text corresponding to the fourth speech request, the server can use the fourth speech request as the first speech request for subsequent processing. On the contrary, if the second speech request processing model fails to recognize the text corresponding to the fourth speech request, the server can discard or delete the fourth speech request.
[0121] In this way, in the embodiments of the present application, natural language processing can be performed on multiple fourth speech requests according to the second speech request processing model to determine the second processing result of each fourth speech request, and the first speech request among the multiple fourth speech requests can be determined according to the second processing result, thereby ensuring the quality of the first speech request to a certain extent.
[0122] In some embodiments of the present application, the second processing result includes the third text prediction and the prediction probability of the third text prediction. Furthermore, step 011 includes:
[0123] When the third text prediction of the fourth speech request does not match the third preset text, determining the fourth speech request as the first speech request; and / or
[0124] When the third text prediction of the fourth voice request matches the third preset text, and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold, the fourth voice request is determined as the first voice request.
[0125] The acquisition module according to the embodiment of the present application is further configured to determine the fourth voice request as the first voice request when the third text prediction of the fourth voice request does not match the third preset text, and / or is configured to determine the fourth voice request as the first voice request when the third text prediction of the fourth voice request matches the third preset text, and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold.
[0126] The processor according to the embodiment of the present application is further configured to determine the fourth voice request as the first voice request when the third text prediction of the fourth voice request does not match the third preset text, and / or is configured to determine the fourth voice request as the first voice request when the third text prediction of the fourth voice request matches the third preset text, and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold.
[0127] Specifically, in the embodiment of the present application, the second voice request processing model is applicable to the speech recognition task. Further, the second voice request processing model can perform speech recognition on the fourth voice request. Therefore, the second processing result of the fourth voice request can be understood as the text prediction of the fourth voice request, that is, the above-mentioned third text prediction.
[0128] In addition, in the embodiment of the present application, the third preset text refers to texts that are preset and have a relatively high occurrence frequency in the in-vehicle voice interaction scenario, such as "turn on the air conditioner", "close the window", etc.
[0129] Further, in the embodiment of the present application, when the third text prediction of the fourth voice request does not match the third preset text, it indicates that the fourth voice request may not be directed to texts with a relatively high occurrence frequency in the in-vehicle voice interaction scenario, such as "turn on the air conditioner", "close the window", etc. Therefore, it is a relatively rare voice request for the model. If the fourth voice request is used as model training data, it will help improve the model performance.
[0130] Similarly, if the third text prediction of the fourth voice request matches the third preset text, and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold, it indicates that the fourth voice request may not be directed to texts with a relatively high occurrence frequency in the in-vehicle voice interaction scenario, such as "turn on the air conditioner", "close the window", etc. Therefore, it is a relatively rare voice request for the model. If the fourth voice request is used as model training data, it will help improve the model performance.
[0131] In contrast, if the third text prediction of the fourth voice request matches the third preset text, and the prediction probability of the third text prediction of the fourth voice request is greater than the preset probability threshold, it indicates that the fourth voice request can directly point to texts that appear frequently in in-vehicle voice interaction scenarios such as "turn on the air conditioner" and "close the window". Therefore, for the model, it is a relatively common voice request. Furthermore, the fourth voice request can be discarded or deleted.
[0132] In one example, the similarity between the third text prediction of the fourth voice request and the third preset text can be calculated to determine whether the third text prediction of the fourth voice request matches the third preset text. For example, when the cosine similarity between the third text prediction of the fourth voice request and the third preset text is greater than 0.7, it is considered that the third text prediction of the fourth voice request matches the third preset text; otherwise, it does not match.
[0133] In one example, the prediction probability of the third text prediction refers to the confidence level of the third text prediction.
[0134] In this way, in the implementation manner of the present application, the fourth voice request can be determined as the first voice request when the third text prediction of the fourth voice request does not match the third preset text, and / or when the third text prediction of the fourth voice request matches the third preset text and the prediction probability of the third text prediction of the fourth voice request is less than or equal to the preset probability threshold, thereby further ensuring the effectiveness of the first voice request.
[0135] To more clearly illustrate the training method of the voice request processing model for voice interaction in the implementation manner of the present application, please refer to Figure 4 , Figure 4 which is a schematic diagram of the application scenario in some implementation manners of the present application. Specifically, as Figure 5 shown, in the implementation manner of the present application, first, a model with available performance (i.e., the second voice request processing model) can be trained through internal training data such as online collected data, self-owned data, open-source data, etc. and put online.
[0136] Next, after the model is put online, the previous online data of the vehicle (such as voice requests) is desensitized to obtain the fourth voice request.
[0137] Then, according to the natural language processing of the model for the fourth voice request, the second processing result of the fourth voice request is determined, and according to the second processing result of the fourth voice request, the first voice request among the obtained multiple fourth voice requests is determined (corresponding to Figure 5 the online semi-supervised data in
[0138] Then, based on the data processing module, from all the first voice requests, the available annotated data is determined, that is, the target voice request and the first processing result of the target voice request.
[0139] Then, using the target voice request and the first processing result of the target voice request, combined with other training data such as the above internal training data, the above self-owned data, etc., the model is trained such as fine-tuning, iterative optimization processing, etc.
[0140] After that, the trained model is evaluated using the internal test set and the open-source test set.
[0141] Finally, if the evaluation result meets the expectation, the trained model is put online.
[0142] It should be noted that the data processing module may include three large speech recognition models: Paraformer, SenseVoice, and Whisper v3. Already, the data processing module can use these three models to decode the online semi-supervised data, and retain the data with the similarity of the three recognition results greater than 95%, that is, the second voice request.
[0143] In addition, the data processing module includes a Filter component to achieve filtering. It can be understood that the data received online by mobile phones is concentrated on common instructions such as "open the window" and "turn on the air conditioner". Therefore, the high-frequency instructions can be filtered through the Filter component, or in other words, the second voice requests that match the speech prediction results with common instructions such as "open the window" and "turn on the air conditioner" are filtered to retain some of such second voice requests in different vehicle models and different scenarios.
[0144] In addition, the data processing module can also retain the second voice requests with the same text or the same scenario as the BadCase according to the pre-obtained BadCase during the data processing to ensure that the model after iteration improves on this BadCase.
[0145] Please refer to Figure 5 , corresponding to the above training method of the voice request processing model for voice interaction, the embodiment of the present application also provides a voice interaction method, including:
[0146] 05: Obtain the current voice request;
[0147] 06: Perform natural language processing on the current voice request according to the third voice request processing model to determine the third processing result of the current voice request, where the third voice request processing model is trained by the above training method of the voice request processing model for voice interaction;
[0148] 07: Determine a vehicle control command according to the third processing result;
[0149] 08: Send the vehicle control command to the vehicle to perform voice interaction.
[0150] The embodiment of the present application provides a voice interaction device. The voice interaction method for voice interaction in the embodiment of the present application can be implemented by the training of the voice request processing model for voice interaction in the embodiment of the present application. Specifically, the training device includes an acquisition module, a processing result determination module, a voice request determination module, and a training module. Among them, the acquisition module is used to acquire a plurality of first voice requests. The processing result determination module is used to determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models. The voice request determination module is used to determine the target voice request among the plurality of first voice requests according to the first processing results. The training module is used to train the pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0151] The embodiment of the present application also provides a server, which includes a memory and a processor. The training method of the voice request processing model for voice interaction in the embodiment of the present application can be implemented by the server in the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to acquire a plurality of first voice requests, and to determine a plurality of first processing results of the first voice requests according to a plurality of pre-trained second voice request processing models, and to determine the target voice request among the plurality of first voice requests according to the first processing results, and to train the pre-set second voice request processing model according to the target voice request and the first processing result of the target voice request.
[0152] Specifically, in the embodiment of the present application, when a user in the vehicle cockpit wants to control the vehicle to perform actions such as turning off the air conditioner and opening the window at the current moment, and thus utters statements such as "turn off the air conditioner" and "open the window", the vehicle can capture the corresponding sound through a sound collection component such as a microphone to obtain the current voice request, and forward the current voice request to the server communicatively connected to the vehicle.
[0153] Next, when the server trains the third voice request processing model through the above training method for voice interaction to obtain the trained third voice request processing model, or when training the second voice request processing model and determining the trained second voice request processing model as the third voice request processing model, the server can call the third voice request processing model to perform natural language processing on the current voice request and determine the third processing result of the current voice request. For example, when the third voice request processing model is applicable to the speech recognition task, the server performs speech recognition on the current voice request through the third voice request processing model to obtain the speech recognition result of the current voice request, such as text prediction.
[0154] Next, the server can determine the vehicle control instruction corresponding to the current voice request and available for implementing the current voice request according to the third processing result of the current voice request.
[0155] Finally, the server can send the vehicle control instruction to the vehicle, so that the vehicle can perform actions such as turning on the air conditioner and closing the window according to the received vehicle control instruction, thereby realizing voice interaction with the user.
[0156] It should be understood that the running process and training process of the voice request processing model can refer to the foregoing content. To avoid repetition, it will not be elaborated again here.
[0157] It can also be understood that the process of determining the vehicle control instruction according to the third processing result of the current voice request can be set according to the actual situation. For example, in one example, if the third voice request processing model is applicable to the speech recognition task, the third processing result of the current voice request is the text prediction of the current voice request. Therefore, the server can perform natural language processing such as slot recognition and application programming interface prediction on the text prediction of the current voice request to determine the corresponding vehicle control instruction, such as an instruction for setting the operating state of vehicle components, so that the vehicle can adjust the operating state of vehicle components after receiving this instruction.
[0158] Thus, in the embodiments of the present application, the third voice request processing model determined through training can perform natural language processing on the obtained current voice request, so as to determine the third processing result of the current voice request, and the vehicle control instruction can be determined based on the third processing result of the current voice request and sent to the vehicle, thereby performing voice interaction, realizing in-vehicle voice interaction, and ensuring the stable execution of in-vehicle voice interaction to a certain extent. Moreover, for the model training process, in the case of obtaining multiple first voice requests with unknown labels and unknown quality, the first voice request processing model that has been pre-trained can be used to determine multiple first processing results of each first voice request, and based on the multiple first processing results of each first voice request, the target voice request can be determined from the multiple first voice requests. Thus, to a certain extent, the voice requests with lower quality among the multiple first voice requests can be eliminated and the voice requests with higher quality, that is, the target voice requests, can be retained, thereby realizing the screening of the multiple first voice requests. And because the first processing result of the first voice request is determined according to the first voice request processing model, the first processing result of the target voice request can be used as the label of the target voice request to a certain extent, thereby completing the annotation of the target voice request. Therefore, the manual participation link in the annotation process of the target voice request is reduced, and the annotation efficiency of the target voice request is improved. At the same time, the target voice request and the first processing result of the target voice request can be used as the training samples and sample labels of the second voice request processing model, so as to train the second voice request processing model, ensuring that the training samples and sample labels of the second voice request processing model can be obtained in an efficient manner, and the acquisition efficiency is guaranteed. Thus, the full training of the second voice request processing model can be realized, and further, the performance of the second voice request processing model is guaranteed.
[0159] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, which when executed by one or more processors, implements the above-mentioned training method of the voice request processing model for voice interaction or the above-mentioned voice interaction method.
[0160] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0161] Any process or method description shown in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0162] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for training a voice request processing model for voice interaction, characterized in that: include: Obtaining multiple first voice requests; Determining multiple first processing results of the first voice request according to multiple pre-trained second voice request processing models; determining a target voice request among the plurality of first voice requests according to the first processing result; A preset second voice request processing model is trained according to the target voice request and the first processing result of the target voice request.
2. The method according to claim 1, characterized in that The step of determining a target voice request among the plurality of first voice requests according to the first processing result includes: Determining, according to the plurality of first processing results of the first voice requests, a plurality of second voice requests among the plurality of first voice requests; A target voice request among multiple second voice requests is determined according to the first processing result of the second voice request.
3. The method according to claim 2, characterized in that The determining, according to the plurality of first processing results of the first voice requests, a plurality of second voice requests among the plurality of first voice requests comprises: If, among the multiple first processing results of the first voice request, the similarity between any two of the first processing results is greater than or equal to a preset similarity threshold, the first voice request is determined to be the second voice request.
4. The method according to claim 2, characterized in that: The first processing result includes a first text prediction, and determining a target voice request among the plurality of second voice requests according to the first processing result of the second voice request includes: The target voice request is determined according to a third voice request and a preset voice request discard probability, wherein the third voice request is a voice request in which the first text prediction matches the first preset text among a plurality of the second voice requests.
5. The method according to claim 2, characterized in that: The first processing result includes a second text prediction, each of the second voice requests corresponds to at least one preset scene information, and determining a target voice request among the plurality of second voice requests according to the first processing result of the second voice request includes: When the second text prediction of the second voice request matches the second preset text, or the scene information corresponding to the second voice request is the target scene information, the second voice request is determined as the target voice request, wherein the second preset text, the target scene information correspond to the voice request that is incorrectly predicted by the second voice request processing model.
6. The method according to claim 1, characterized in that The obtaining of multiple first voice requests includes: performing natural language processing on a plurality of fourth voice requests according to the second voice request processing model to determine a second processing result for each of the fourth voice requests; Determine the first voice request among multiple fourth voice requests according to the second processing result.
7. The method according to claim 6, characterized in that The second processing result includes a third text prediction and a prediction probability of the third text prediction, and determining the first voice request among a plurality of fourth voice requests according to the second processing result includes: If the third text prediction of the fourth voice request does not match the third preset text, determining the fourth voice request as the first voice request; and / or When the third text prediction of the fourth voice request matches the third preset text and the prediction probability of the third text prediction of the fourth voice request is less than or equal to a preset probability threshold, the fourth voice request is determined as the first voice request.
8. A voice interaction method, characterized in that: include: Get the current voice request; performing natural language processing on the current voice request according to a third voice request processing model to determine a third processing result of the current voice request, wherein the third voice request processing model is trained by the method according to any one of claims 1 to 7; Determining a vehicle control instruction according to the third processing result; The vehicle control instruction is sent to the vehicle to perform the voice interaction.
9. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 8 is implemented.