Voice interaction method, server, and computer-readable storage medium

The large language model receives and processes the user's voice requests in the vehicle, generates guidance information to help the user adjust voice commands, solves the problem of voice interaction failure caused by the user's unclear expression during the first use, and improves the success rate and user experience of the interaction.

CN118136014BActive Publication Date: 2025-07-01GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410413440.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-07
Publication Date
2025-07-01
Estimated Expiration
2044-04-07

AI Technical Summary

Technical Problem

When users use the vehicle for the first time, due to their unfamiliarity with the vehicle and voice interaction functions, they may express vague or incomplete voice commands, which makes it difficult for the vehicle to understand and perform corresponding functions, resulting in the failure of voice interaction.

Method used

Receive the current voice request forwarded by the vehicle through a large language model. If the corresponding vehicle control command is not determined, the target guidance information is generated to guide the user to adjust the voice request and feedback it to the user.

Benefits of technology

Effectively guide users to adjust voice requests to make them clearer and more accurate, ensure successful voice interaction, improve user experience, and improve the reliability and adaptability of the model through training and update of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118136014B_ABST
    Figure CN118136014B_ABST
Patent Text Reader

Abstract

The present application discloses a voice interaction method, a server, and a computer-readable storage medium. The method includes: receiving a current voice request forwarded by a vehicle; in the case where the corresponding vehicle control instruction cannot be determined according to the current voice request, determining, according to a large language model and the current voice request, target guiding information for guiding a user to adjust the current voice request; and feeding back the target guiding information to guide the user to complete the voice interaction. In this way, the server of the present application can feed back the target guiding information, and the user adjusts the current voice request based on the target guiding information, making the target guiding information relevant to the current voice request, ensuring the reliability of the target guiding information, determining the target guiding information through the large language model, making the target guiding information credible, and enabling the update of each training round of the large language model to be performed based on the prediction results of the previous training round and the current training round, ensuring the model training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and particularly relates to a voice interaction method, a server, and a computer-readable storage medium. Background Art

[0002] When a user starts using a vehicle for the first time, since they are relatively unfamiliar with the vehicle and its voice interaction function, they may express some vague and incomplete voice commands. When the vehicle processes these voice commands, it is difficult to understand the actual meaning of the voice commands, and thus cannot execute the corresponding functions, resulting in a voice interaction failure. Summary of the Invention

[0003] This application provides a voice interaction method, a server, and a computer-readable storage medium.

[0004] An embodiment of this application provides a voice interaction method, including:

[0005] Receiving a current voice request forwarded by a vehicle;

[0006] In the case where a vehicle control command corresponding to the current voice request cannot be determined according to the current voice request, determining target guidance information for guiding the user to adjust the current voice request according to a large language model and the current voice request, the large language model being capable of generating guidance information according to a voice request, and each training round of the large language model being trained based on the prediction result of the first guidance information in the previous training round and the prediction result of the second guidance information in the current training round;

[0007] Feeding back the target guidance information to guide the user to complete the voice interaction.

[0008] In the voice interaction method provided by an embodiment of this application, the server can receive a current voice request forwarded by a vehicle, and in the case where a vehicle control command corresponding to the current voice request cannot be determined according to the current voice request, determine target guidance information for guiding the user to adjust the current voice request according to a pre-trained large language model and the current voice request, and feed back the target guidance information to guide the user to complete the voice interaction.

[0009] Thus, in the embodiments of the present application, when the server fails to determine the vehicle control instruction corresponding to the current voice request, it can determine the target guidance information for guiding the user to adjust the voice request and feedback the target guidance information, so that the user can adjust the current voice request based on the guidance of the target guidance information. Furthermore, when the user is unclear about how to express the voice request for controlling the vehicle, the target guidance information can guide the user to adjust the voice request, thus ensuring the user experience of using the vehicle and the vehicle voice interaction function. The embodiments of the present application can determine the target guidance information according to the current voice request, so that the target guidance information can be related to the current voice request, and further, to a certain extent, ensure the adaptation of the target guidance information to the user's current usage needs, and the reliability of the target guidance information can be guaranteed. The embodiments of the present application can determine the target guidance information through the large language model, so the credibility of the target guidance information can be guaranteed to a certain extent. In addition, the embodiments of the present application enable the update of the large language model in each training round to be based on the first guidance information prediction result of the large language model in the previous training round and the second guidance information prediction result of the large language model in the current training round, realizing model training based on the previous training round and the current training round, and the model training effect can be guaranteed to a certain extent.

[0010] In some embodiments of the present application, the training steps of the large language model include:

[0011] Obtain a voice request sample and a guidance information label corresponding to the voice request sample;

[0012] Train a pre-determined first reference model according to the voice request sample and the guidance information label to determine the large language model.

[0013] Thus, in the embodiments of the present application, the server can train the first reference model with the obtained voice request sample and the guidance information label corresponding to the voice request sample to determine the large language model, so that the large language model can reliably output the corresponding guidance information according to the voice request.

[0014] In some embodiments of the present application, the training of the pre-determined first reference model according to the voice request sample and the guidance information label to determine the large language model includes:

[0015] Train the first reference model according to the voice request sample and the guidance information label to obtain a second reference model;

[0016] Determine a prediction result of the first guidance information of the second reference model for the voice request sample in the previous training round and a prediction result of the second guidance information of the second reference model for the voice request sample in the current training round according to the second reference model and the voice request sample;

[0017] Train the second reference model according to the prediction result of the first guidance information and the prediction result of the second guidance information to obtain the large language model.

[0018] In this way, in the embodiment of the present application, the server can train the second reference model and thus obtain the large language model according to the prediction result of the first guidance information of the second reference model for the voice request sample in the previous training round and the prediction result of the second guidance information of the second reference model for the voice request sample in the current training round, so that the training of the second reference model can be based on the prediction results for the voice request sample before and after each update, and the training effect of the second reference model can be guaranteed to a certain extent.

[0019] In some embodiments of the present application, the prediction result of the first guidance information includes a first probability that the second reference model determines the guidance information label according to the voice request sample in the previous training round, and the prediction result of the second guidance information includes a second probability that the second reference model determines the guidance information label according to the voice request sample in the current training round. Training the second reference model according to the prediction result of the first guidance information and the prediction result of the second guidance information to obtain the large language model includes:

[0020] Train the second reference model according to the similarity between the first probability and the second probability to obtain the large language model.

[0021] In this way, in the embodiment of the present application, the server can train the second reference model to obtain the large language model based on the first probability that the second reference model predicts the guidance information label corresponding to the voice request sample through the voice request sample in the previous training round and the second probability that the second reference model predicts the guidance information label through the voice request sample in the current training round, so that the training of the large language model can depend on its own prediction probabilities for the guidance information label corresponding to the voice request sample in the previous and current training rounds, thereby realizing self-play and update of the large language model, and enabling the large language model to be aligned with the obtained voice request sample and guidance information label to a certain extent, and the performance of the large language model can be guaranteed.

[0022] In some embodiments of the present application, the first guidance information prediction result includes a third probability of the guidance information predicted by the reference model in the previous training round based on the voice request sample, and the second guidance information prediction result includes a fourth probability of the guidance information predicted by the reference model in the current training round based on the voice request sample. Training the second reference model to obtain the large language model according to the similarity between the first probability and the second probability includes:

[0023] Training the second reference model to obtain the large language model according to the similarity between the first probability and the second probability, and the similarity between the third probability and the fourth probability.

[0024] In this way, in the embodiments of the present application, the server can use the third probability of the guidance information predicted by the second reference model in the previous training round through the voice request sample, and the fourth probability of the guidance information predicted by the second reference model in the current training round through the voice request sample to train the second reference model in the current training round, thereby avoiding to a certain extent the situation where the prediction result of the second reference model becomes too divergent due to continuous training, and ensuring the training effect of the second reference model.

[0025] In some embodiments of the present application, when the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining the target guidance information for guiding the user to adjust the current voice request according to the large language model and the current voice request includes:

[0026] When the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model, determining the target guidance information according to the large language model and the current voice request.

[0027] In this way, in the embodiments of the present application, the server can confirm the vehicle control instruction of the current voice request through the large language model, and when the large language model fails to predict the vehicle control instruction of the current voice request, determine the target guidance information corresponding to the current voice request based on the large language model, so that the prediction of the vehicle control instruction of the current voice request and the prediction of the target guidance information of the current voice request can both be realized based on the large language model, avoiding the situation where a module for the vehicle control instruction corresponding to the voice request and a module for predicting the target guidance information need to be separately set in the server, and reducing the load of the server.

[0028] In some embodiments of the present application, in the case where a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, determining the target guidance information based on the large language model and the current voice request includes:

[0029] In the case where a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, determining the target guidance information based on the current voice request, the large language model, and a pre-configured prompt information template.

[0030] Thus, in the embodiments of the present application, the large language model can perform natural language understanding and processing on the current voice request based on the indication of the prompt information template, so as to generate the target guidance information of the current voice request, ensuring the stable operation of the large language model.

[0031] In some embodiments of the present application, the prompt information template includes a first prompt information sub-template and a second prompt information sub-template. In the case where a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, determining the target guidance information based on the current voice request, the large language model, and the pre-configured prompt information template includes:

[0032] In the case where a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request, the large language model, and the first prompt information sub-template, determining the target guidance information based on the current voice request, the large language model, and the second prompt information sub-template.

[0033] Thus, in the embodiments of the present application, the large language model can confirm the correspondence between the current voice request and the vehicle control instruction based on the first sub-prompt information template, the second sub-prompt information template in the prompt information template, and the current voice request, and in the case where it is confirmed that the current voice request corresponds to a vehicle control instruction, confirm the target guidance information corresponding to the current voice request, realizing the inference of the large language model based on the prompt information template, so that the inference accuracy of the large language model can be guaranteed to a certain extent.

[0034] An embodiment of the present application provides a server, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.

[0035] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by one or more processors, implements the above-mentioned voice interaction method.

[0036] For the server and computer-readable storage medium provided by the embodiment of the present application, when it is impossible to determine the vehicle control instruction corresponding to the current voice request, the target guidance information for guiding the user to adjust the voice request can be determined and fed back, so that the user can adjust the current voice request based on the guidance of the target guidance information. Furthermore, when the user is unclear about how to express the voice request for controlling the vehicle, the target guidance information can guide the user to adjust the voice request, thus ensuring the user's experience of using the vehicle and the vehicle voice interaction function. The embodiment of the present application can determine the target guidance information according to the current voice request, so that the target guidance information can be related to the current voice request, and then to a certain extent, the adaptability of the target guidance information to the user's current usage requirements can be ensured, and the reliability of the target guidance information can be guaranteed. The embodiment of the present application can determine the target guidance information through the large language model, so the credibility of the target guidance information can be guaranteed to a certain extent. In addition, in the embodiment of the present application, the update of the large language model in each training round can be based on the first guidance information prediction result of the large language model in the previous training round and the second guidance information prediction result of the large language model in the current training round, realizing model training based on the previous training round and the current training round, and the model training effect can be guaranteed to a certain extent.

[0037] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0039] Figure 1 is a schematic flowchart of a voice interaction method in some embodiments of the present application;

[0040] Figure 2 is a schematic flowchart of a voice interaction method in some embodiments of the present application;

[0041] Figure 3 is a schematic flowchart of a voice interaction method in some embodiments of the present application;

[0042] Figure 4 is a schematic flowchart of a voice interaction method in some embodiments of the present application;

[0043] Figure 5 is a schematic flow chart of a voice interaction method in some embodiments of the present application;

[0044] Figure 6 is a schematic flow chart of a voice interaction method in some embodiments of the present application. Specific embodiments

[0045] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.

[0046] When a user first uses or just starts using a vehicle and the voice interaction function carried by the vehicle, since the user is relatively unfamiliar with the objects and functions that can be controlled by voice in the vehicle, the user may utter some voice commands with ambiguous semantics, chaotic or missing sentence components, such as "Adjust the air conditioner", "Turn on 'that thing for adjusting the temperature'", etc.

[0047] It can be understood that these voice commands have one or more problems such as unclear intention, missing (or ambiguous) components, incorrect collocation of sentence components, etc. For example, it can be considered that the collocation relationship between 'adjust' and 'air conditioner' in "Adjust the air conditioner" is inappropriate. It can also be considered that the intention corresponding to "Adjust the air conditioner" includes the increase and decrease of the cooling temperature / heating temperature of the air conditioner and the switching of the operation mode of the air conditioner, etc. Therefore, the intention of "Adjust the air conditioner" is unclear.

[0048] Therefore, due to one or more problems such as unclear intention, missing (or ambiguous) components, incorrect collocation of sentence components, etc. in these voice commands, it is difficult to determine the specific vehicle control commands for implementing or completing these voice commands when processing these voice commands, thereby resulting in the failure of the voice interaction with the user this time.

[0049] To improve this situation, some solutions are to play specific information through timed dynamic advertisement examples to teach users to use specific functions of the vehicle. However, this solution is difficult to meet the current usage needs of users.

[0050] In some other solutions, the user is requested to clarify these voice commands. That is to say, when the user utters a voice command with ambiguous semantics or incomplete sentence components, a counter-question is initiated to the user. For example, when the voice command uttered by the user is "Adjust the air conditioner", the vehicle can play a voice of "Failed to understand your meaning. Please be more specific" to request the user to clarify the true intention of this voice command.

[0051] However, when asking the user for clarification, the interaction with the user is usually based on fixed statements, such as "I didn't understand what you meant. Please be more specific." as mentioned above. In this case, if the user is using the vehicle and its voice interaction function for the first time or just starting to use it, and is thus unfamiliar with the objects and functions that can be controlled by voice in the vehicle, then when faced with a rhetorical question (or clarification request) from the vehicle, the user may not know how to clarify, or rather, may not know "how to formulate a voice command to enable the vehicle to meet the current requirements", resulting in the need for the user to have multiple rounds of conversations and clarifications with the vehicle before being able to control the vehicle through a correct (or rather, understandable by the vehicle) voice command.

[0052] Based on the above possible problems, please refer to Figure 1 , an embodiment of the present application provides a voice interaction method, including:

[0053] 01: Receive the current voice request forwarded by the vehicle;

[0054] 02: In the case where a vehicle control command corresponding to the current voice request cannot be determined according to the current voice request, determine target guidance information for guiding the user to adjust the current voice request according to the large language model and the current voice request. The large language model can generate guidance information according to the voice request, and each training round of the large language model is based on the prediction result of the first guidance information in the previous training round and the prediction result of the second guidance information in the current training round;

[0055] 03: Feedback the target guidance information to guide the user to complete the voice interaction.

[0056] An embodiment of the present application provides a voice interaction device. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the voice interaction device includes a receiving module, a determining module, and an interaction module. Among them, the receiving module is used to receive the current voice request forwarded by the vehicle. The determining module is used to determine target guidance information for guiding the user to adjust the current voice request according to the large language model and the current voice request in the case where a vehicle control command corresponding to the current voice request cannot be determined according to the current voice request. The large language model can generate guidance information according to the voice request, and each training round of the large language model is based on the prediction result of the first guidance information in the previous training round and the prediction result of the second guidance information in the current training round. The interaction module is used to feedback the target guidance information to guide the user to complete the voice interaction.

[0057] An embodiment of the present application also provides a server, which includes a memory and a processor. The voice interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory. The processor is configured to receive the current voice request forwarded by the vehicle, and in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determine the target guidance information for guiding the user to adjust the current voice request according to the large language model and the current voice request. The large language model can generate guidance information according to the voice request. Each training round of the large language model is based on the prediction result of the first guidance information in the previous training round and the prediction result of the second guidance information in the current training round, and is configured to feedback the target guidance information to guide the user to complete the voice interaction.

[0058] Specifically, in the embodiment of the present application, when the user desires to control the vehicle to perform a specific function by voice at the current moment, thereby triggering the current voice request, the vehicle can obtain the current voice request and forward the current voice request to the server. The server receives the current voice request and determines the vehicle control instruction corresponding to the current voice request according to the algorithm or model implemented by the preset code. In the case where one or more of the problems such as unclear intention, missing (or fuzzy) components, and incorrect collocation of sentence components exist in the current voice request, resulting in the inability to determine the vehicle control instruction corresponding to the current voice request, the pre-trained large language model (LLM) that can generate guidance information through the voice request can be called, and the current voice request is input into the large language model, so that the large language model infers the prompt information for the current voice request and infers the target guidance information for guiding the user to adjust the current voice request. Thus, the server can feedback the target guidance information to the user, so that the user can adjust the current voice request through the target guidance information to an effective voice request that does not have any of the problems such as unclear intention, missing (or fuzzy) components, and incorrect collocation of sentence components, and then can control the vehicle to perform a specific function through the effective voice request.

[0059] It can be understood that the voice request in the embodiment of the present application can be understood as a voice instruction having one or more of the problems such as unclear intention, missing (or fuzzy) components, and incorrect collocation of sentence components, such as "Adjust the air conditioner", "Turn on 'that thing for adjusting the temperature'", etc.

[0060] It can also be understood that the server of the embodiment of the present application also performs natural language understanding on the voice request based on the preset program, code, or pre-completed natural language processing model, etc., and generates a vehicle control instruction through the natural language understanding result.

[0061] Further, in the case where a vehicle control instruction corresponding to the current voice request cannot be generated, or rather, in the case where a vehicle control instruction for implementing the current voice request cannot be generated (for example, because the intention of "adjusting the air conditioner" is not clear and thus the corresponding vehicle control instruction cannot be generated), the server can call a pre-trained large language model and input the current voice request into the large language model to generate target guidance information for guiding the user to adjust the current voice request through the large language model.

[0062] Exemplarily, when the user wants to adjust the air conditioner temperature by voice at the current moment, and the current voice request triggered thereby is "adjust the air conditioner", the vehicle forwards "adjust the air conditioner" to the server. The server determines the corresponding vehicle control instruction for "adjust the air conditioner".

[0063] However, when the server fails to determine a vehicle control instruction for implementing "adjust the air conditioner", or rather, when there are multiple vehicle control instructions for implementing "adjust the air conditioner", such as an instruction to increase the cooling temperature of the air conditioner, an instruction to decrease the heating temperature of the air conditioner, and an instruction to adjust the air blowing mode of the air conditioner, etc., but the instruction required by the user among these instructions cannot be determined, the server can determine the target guidance information for guiding the user to adjust "adjust the air conditioner" according to the pre-trained large language model, such as "I don't understand what you mean. You can be more specific, such as: adjust the air conditioner to 20 degrees, adjust the air volume of the air conditioner to the third gear, open the air conditioner page."

[0064] It can be understood that since the target guidance information is generated by the large language model according to the current voice request, the target guidance information is related to the current voice request and can, to a certain extent, guide the user to adjust and clarify the current voice request.

[0065] It can also be understood that the large language model in the embodiment of the present application has been correspondingly trained before being called by the server, so as to output the corresponding target guidance information when receiving the input of the current voice request.

[0066] Further, in the embodiment of the present application, during the process of the server training the large language model, the model training of the current training round can be completed by using the first guidance information inference result generated by the large language model that has not been trained yet in the previous training round and the second guidance information inference result generated in the current training round.

[0067] Therefore, in the embodiments of the present application, the update of each training round of the large language model can be based on the prediction results of its own guidance information in the previous training round and the current training round. For example, it can be trained by any one or more of the similarity degree, difference degree, and change situation between the prediction results of the guidance information in the previous training round and the prediction results of the guidance information in the current training round, so that the large language model in the current training round can rely on the output results of the large language model in the previous training round.

[0068] It can be understood that "the training of the large language model based on the prediction results of the guidance information in the previous and current training rounds" can be understood to a certain extent as: the "large language model in the current training round" conducts self-play and update through the "large language model in the previous training round".

[0069] It can also be understood that after the server generates the target guidance information for the current voice request through the pre-trained large language model, it can send the target guidance information to the vehicle. When the vehicle receives the target guidance information, it can play the target guidance information based on the vehicle display component or the sound playback component to guide the user to adjust or clarify the current voice request expressed previously, thereby adjusting the current voice request for which the vehicle control instruction cannot be determined into a valid voice request for which the vehicle control instruction can be determined, and enabling the vehicle to execute the vehicle control instruction corresponding to the valid voice request, and finally completing the voice interaction.

[0070] Optionally, in some embodiments of the present application, the server can also send the target guidance information to the user's terminal so that the user can adjust the current voice request through the target guidance information played by the terminal.

[0071] In summary, in the embodiments of the present application, when the server fails to determine the vehicle control instruction corresponding to the current voice request, it can determine the target guidance information for guiding the user to adjust the voice request, and feedback the target guidance information, so that the user can adjust the current voice request based on the guidance of the target guidance information. Furthermore, when the user is unclear about how to express the voice request for controlling the vehicle, the target guidance information can guide the user to adjust the voice request, thereby ensuring the user's experience of using the vehicle and the vehicle voice interaction function. The embodiments of the present application can determine the target guidance information according to the current voice request, so that the target guidance information can be related to the current voice request, and further, to a certain extent, ensure the adaptation of the target guidance information to the user's current usage needs, and the reliability of the target guidance information can be guaranteed. The embodiments of the present application can determine the target guidance information through the large language model, so the credibility of the target guidance information can be guaranteed to a certain extent. In addition, in the embodiments of the present application, the update of the large language model in each training round can be based on the prediction result of the first guidance information of the large language model in the previous training round and the prediction result of the second guidance information of the large language model in the current training round, realizing model training based on the previous training round and the current training round, and the model training effect can be guaranteed to a certain extent.

[0072] Please refer to Figure 2 , in some embodiments of the present application, the training steps of the large language model include:

[0073] 04: Obtain voice request samples and the guidance information labels corresponding to the voice request samples;

[0074] 05: Train a pre-determined first reference model according to the voice request samples and the guidance information labels to determine the large language model.

[0075] The voice interaction device of the embodiments of the present application further includes a sample acquisition module and a first model training module. Among them, the sample acquisition module is used to obtain voice request samples and the guidance information labels corresponding to the voice request samples. The first model training model is used to train a pre-determined first reference model according to the voice request samples and the guidance information labels to determine the large language model.

[0076] The processor of the embodiments of the present application is further used to obtain voice request samples and the guidance information labels corresponding to the voice request samples, and to train a pre-determined first reference model according to the voice request samples and the guidance information labels to determine the large language model.

[0077] Specifically, to ensure that the large language model can reliably predict the guiding information corresponding to the voice request after receiving the voice request, the server in the implementation manner of this application can target a large language model with certain natural language processing capabilities that has not been trained yet, that is, the first reference model, and use the collected voice request samples and the guiding information labels of the voice request samples to perform corresponding training on the first reference model, so that the first reference model can be applicable to the downstream task of guiding information reasoning.

[0078] It can be understood that the voice request samples in the implementation manner of this application can be obtained through methods such as collecting online voice interaction logs or manual construction. And the guiding information labels corresponding to the voice request samples can be obtained through manual annotation.

[0079] It can also be understood that the specific form of the voice request samples and the specific form of the guiding information labels corresponding to the voice request samples are both content that can be set according to the actual situation. For example, in one example, the voice request sample is "Adjust the air conditioner", and the guiding information label corresponding to this voice request is "Sorry, I didn't understand what you meant. If you want to adjust the temperature, you can say: Adjust the air conditioner to twenty degrees. If you want to adjust the air volume, you can say: Adjust the air volume to the third gear."

[0080] It can be understood that the first reference model in the implementation manner of this application is a large language model with certain natural language processing capabilities that has not been trained yet and is relatively unfamiliar with the downstream task of guiding information reasoning. In response to this, the first reference model can be trained in the way of supervised fine-tuning (SFT). Specifically, the server can input the voice request samples into the first reference model so that the first reference model outputs the prediction results corresponding to the voice request samples. Furthermore, the server can determine the difference between this prediction result and the guiding information label corresponding to the voice request sample based on the prediction result corresponding to the voice request sample and the guiding information label corresponding to the voice request sample, and update the parameters in the first reference model such as weights and biases based on this difference.

[0081] It can also be understood that when the trained first reference model faces the input of voice request samples, it can output an output result that is semantically similar to "the guiding information label corresponding to the voice request sample".

[0082] For example, in one example, the voice request sample received by the trained first reference model is "Adjust the air conditioner", and the corresponding guidance information label for this voice request is "Sorry, I didn't understand what you meant. If you want to adjust the temperature, you can say: Adjust the air conditioner to twenty degrees. If you want to adjust the air volume, you can say: Adjust the air volume to the third gear." Then, the guidance information prediction result generated by the trained first reference model for this voice request sample can be "Sorry, I didn't understand what you meant. It is recommended that you say, for example: Adjust the air conditioner to twenty degrees, Adjust the air volume to the third gear", or it can also be "I can't understand what you mean. You can speak more standardly, for example: Adjust the air conditioner to twenty degrees, Adjust the air volume to the third gear."

[0083] In addition, it can also be clearly understood that in the implementation manner of the present application, the number of voice request samples obtained by the server is multiple. Correspondingly, the server also obtains the guidance information label corresponding to each voice request sample respectively. Therefore, the server in the implementation manner of the present application can complete the training of the first reference model through the obtained multiple voice request samples and the guidance information label corresponding to each voice request sample.

[0084] In this way, in the implementation manner of the present application, the server can train the first reference model through the obtained voice request samples and the guidance information label corresponding to the voice request samples to determine the large language model, so that the large language model can reliably output the corresponding guidance information according to the voice request.

[0085] Please refer to Figure 3 , in some implementation manners of the present application, step 05 includes:

[0086] 050: Train the first reference model according to the voice request sample and the guidance information label to obtain the second reference model;

[0087] 051: According to the second reference model and the voice request sample, determine the first guidance information prediction result of the second reference model for the voice request sample in the previous training round and the second guidance information prediction result in the current training round;

[0088] 052: Train the second reference model according to the first guidance information prediction result and the second guidance information prediction result to obtain the large language model.

[0089] The first model training module in the implementation manner of the present application is also used to train the first reference model according to the voice request sample and the guidance information label to obtain the second reference model, and is used to determine the first guidance information prediction result of the second reference model for the voice request sample in the previous training round and the second guidance information prediction result in the current training round according to the second reference model and the voice request sample, and is used to train the second reference model according to the first guidance information prediction result and the second guidance information prediction result to obtain the large language model.

[0090] The processor according to the embodiment of the present application is further configured to train a first reference model based on the voice request sample and the guidance information label to obtain a second reference model, and to determine, according to the second reference model and the voice request sample, the first guidance information prediction result of the second reference model for the voice request sample in the previous training round and the second guidance information prediction result in the current training round, and to train the second reference model based on the first guidance information prediction result and the second guidance information prediction result to obtain a large language model.

[0091] Specifically, to reduce the annotation difficulty of samples during the training process and to ensure the robustness of the large language model, the embodiment of the present application sets steps of training a first reference model to obtain a second reference model, and steps of training a second reference model to obtain a large language model.

[0092] It should be noted that in the step of training the second reference model to obtain a large language model, the second reference model can complete the parameter update of itself in the current training round according to the first guidance information prediction result generated by itself for the voice request sample in the previous training round and the second guidance information prediction result generated by itself for the voice request sample in the current training round.

[0093] It can be understood that based on the guidance information prediction results of the second reference model for the same voice request sample in two consecutive training rounds, the change situation of the prediction results of the second reference model for the same voice request before and after the update can be determined.

[0094] Therefore, in some embodiments of the present application, both the first guidance information prediction result and the second guidance information include the guidance information predicted by the second reference model according to the voice request sample. Furthermore, the server can update the large language model in the current training round according to the first distance between the "guidance information output by the second reference model for the voice request sample in the previous training round" and the "guidance information label corresponding to the voice request sample", and according to the second distance between the "guidance information output by the second reference model for the voice request sample in the current training round" and the "guidance information label corresponding to the voice request sample".

[0095] In this way, in the embodiment of the present application, the server can train the second reference model based on the first guidance information prediction result of the second reference model for the voice request sample in the previous training round and the second guidance information prediction result of the second reference model for the voice request sample in the current training round, and thus obtain a large language model, so that the training of the second reference model can be carried out based on the prediction results for the voice request sample before and after each update, and the training effect of the second reference model can be guaranteed to a certain extent.

[0096] In some embodiments of the present application, the first guidance information prediction result includes the first probability of the second reference model determining the guidance information label according to the speech request sample in the previous training round, and the second guidance information prediction result includes the second probability of the second reference model determining the guidance information label according to the speech request sample in the current training round. Accordingly, step 052 includes:

[0097] Training the second reference model to obtain a large language model according to the similarity between the first probability and the second probability.

[0098] The second model update module according to the embodiments of the present application is further configured to train the second reference model to obtain a large language model according to the similarity between the first probability and the second probability.

[0099] The processor according to the embodiments of the present application is further configured to train the second reference model to obtain a large language model according to the similarity between the first probability and the second probability.

[0100] It can be understood that in the step of training the first reference model to obtain the second reference model, the training objective may include "enabling the large language model to predict the guidance information corresponding to the speech request". And in the step of training the second reference model to obtain the large language model, the training objective may include "enabling the large language model to stably predict the guidance information 'similar to the guidance information label' when receiving the speech request".

[0101] Optionally, in the embodiments of the present application, the guidance information labels corresponding to each speech request sample may have the same characteristics, such as the same language style. Among them, the language style can be understood as representations such as detailed, concise, and humorous.

[0102] For example, taking the guidance information label corresponding to "adjust the air conditioner" as an example, in one example, when the guidance information label is expressed in a "detailed" language style, it can be "Sorry, I didn't understand what you meant. If you want to adjust the temperature, you can say: Adjust the air conditioner to twenty degrees. If you want to adjust the air volume, you can say: Adjust the air volume to the third gear". And in another example, when expressed in a "concise" language style, it can be "I can't understand what you mean. You can speak more standardly, such as: Adjust the air conditioner to twenty degrees, Adjust the air volume to the third gear".

[0103] It can be understood that in conventional supervised training, the probability of the neural network model predicting the label through the sample in the early stage of training is usually low, while the probability of the neural network model predicting the label through the sample in the later stage of training is usually high.

[0104] Therefore, in the embodiments of the present application, for a data pair (x, y) composed of a certain voice request sample x and its corresponding guiding information label y, the server can update the parameters of the second reference model in the current training based on the probability p1 that the second reference model predicted the guiding information label y through the voice request sample x in the previous training round, and the probability p2 that the second reference model predicted the guiding information label y through the voice request sample x in the current training round.

[0105] It can be understood that the parameter update of each training round of the second reference model can be based on the probabilities p1 and p2 of predicting the guiding information label y through the voice request sample x in the previous training round and the current training round. Therefore, under the training objective of "enabling the large language model to stably predict guiding information'similar to the guiding information label' when receiving a voice request", as the number of training rounds increases, the probability p that the second reference model predicts the guiding information label y through the voice request sample x can also continuously increase.

[0106] It can also be understood that in the process of updating the second reference model according to p1 and p2, the parameter update of the second reference model in the current training round can be carried out based on the training objectives of "taking the maximum of the similarity degree between p1 and p2" and "taking the maximum of p1", so that the probability that the "second reference model after parameter update in the current training round" predicts the guiding information label y through the voice request sample x is higher than the probability that the "second reference model before parameter update in the current training round" predicts the guiding information label y through the voice request sample x.

[0107] In addition, it should be noted that the embodiments of the present application can update the second reference model based on the similarity degree between p1 and p2. Therefore, in the process of updating the second reference model, the probabilities p that the second reference model predicts the guiding information label y through the voice request sample x in two consecutive training rounds can also continuously increase. Furthermore, the similarity degree between p1 and p2 continuously increases until it stabilizes. Furthermore, after the second reference model is trained, the second reference model (or the large language model) that receives the voice request can stably predict guiding information'similar to the guiding information label'.

[0108] For example, before the training of the second reference model, the second reference model (or the first reference model that has been trained) can output three guiding information prediction results, namely res1, res2, and res3, when receiving the voice request input of "adjust the air conditioner".

[0109] Among them, res1 may include: "I can't understand what you mean. Please speak more standardly. For example: Set the air conditioner to 20 degrees and adjust the air volume to the third gear." res2 may include: "Sorry, I didn't understand what you meant. It is recommended that you say, for example: Set the air conditioner to 20 degrees and adjust the air volume to the third gear." res3 may include: "Sorry, I didn't understand what you meant. If you want to adjust the temperature, you can say: Set the air conditioner to 20 degrees. If you want to adjust the air volume, you can say: Adjust the air volume to the third gear."

[0110] Further, if during the training process of the first reference model, res3 is the guiding information label corresponding to "Adjust the air conditioner", after the training of the second reference model, the trained second reference model (or, the large language model) can, when receiving the voice request input of "Adjust the air conditioner", output the probability of res3 higher than the probabilities of res1 and res2. Therefore, in the inference link after the model is deployed, the trained second reference model (or, the large language model) can output res3 based on "Adjust the air conditioner".

[0111] In this way, in the implementation manner of this application, the server can predict the first probability of the guiding information label corresponding to the voice request sample by the second reference model in the previous training round, and the second probability of the guiding information label predicted by the second reference model through the voice request sample in the current training round, and train the second reference model to obtain the large language model, so that the training of the large language model can depend on its own prediction probabilities of the guiding information label corresponding to the voice request sample in the previous and current two training rounds, thereby realizing the self-play and update of the large language model, and enabling the large language model to be aligned with the obtained voice request sample and guiding information label to a certain extent, and the performance of the large language model can be guaranteed.

[0112] In some implementation manners of this application, the first guiding information prediction result includes the third probability of the guiding information predicted by the reference model according to the voice request sample in the previous training round, and the second guiding information prediction result includes the fourth probability of the guiding information predicted by the reference model according to the voice request sample in the current training round. Furthermore, the step of training the second reference model to obtain the large language model according to the similarity degree between the first probability and the second probability includes:

[0113] Training the second reference model to obtain the large language model according to the similarity degree between the first probability and the second probability, and the similarity degree between the third probability and the fourth probability.

[0114] The second model training module in the implementation manner of this application is further configured to train the second reference model to obtain the large language model according to the similarity degree between the first probability and the second probability, and the similarity degree between the third probability and the fourth probability.

[0115] The processor according to the embodiment of the present application is further configured to train a second reference model to obtain a large language model according to the similarity between the first probability and the second probability, and the similarity between the third probability and the fourth probability.

[0116] Specifically, in order to avoid the situation that in the process of self-play of the second reference model, there is a large difference between the model after the current training round and the model in the previous round, resulting in the divergence of the prediction results, or even the regression of the model performance. Therefore, in the embodiment of the present application, for the voice request sample x and the guiding information label y corresponding to the voice request sample x, the server can combine the probability p1 that the second reference model predicts the guiding information label y through the voice request sample x in the previous training round, the probability p3 that the second reference model predicts the guiding label y' through the voice request sample x in the previous training round, the probability p2 that the second reference model predicts the guiding information label y through the voice request sample x in the current training round, and the probability p4 that the second reference model predicts the above guiding label y' through the voice request sample x in the current training round, to update the model of the second reference model in the current training round.

[0117] It can be understood that in the case of updating the second reference model based on the similarity between the third probability p3 and the fourth probability p4, the higher the similarity between the third probability p3 and the fourth probability p4, the more consistent the prediction results of the second reference model in the previous and current training rounds. Therefore, in the case of taking "the maximum similarity between the third probability p3 and the fourth probability p4" as the training target, it is possible to avoid the situation that there is a large difference between the model after the current training round and the model in the previous round, resulting in the divergence of the prediction results, or in other words, the regression of the model performance.

[0118] In this way, in the embodiment of the present application, the server can train the second reference model in the current training round through the third probability predicted by the second reference model for the guiding information through the voice request sample in the previous training round, and the fourth probability predicted by the second reference model for this guiding information through the voice request sample in the current training round, so as to avoid to a certain extent the situation that the prediction results of the second reference model are too divergent due to continuous training, and the training effect of the second reference model can be guaranteed.

[0119] Optionally, to more clearly illustrate the embodiment of the present application, to more clearly illustrate the embodiment of the present application, please refer to Figure 4 , Figure 4Schematic diagram of the training process of the large language model in some embodiments of the present application. Specifically, for the voice request sample "Adjust the air conditioner", and the corresponding guiding information label of the voice request sample, that is, "Sorry, I didn't understand what you meant. If you want to adjust the temperature, you can say: Adjust the air conditioner to twenty degrees. If you want to adjust the air volume, you can say: Adjust the air volume to the third gear", then: For the first training round, the server can copy the second reference model Model to obtain Reference Model1 and Policy Model.

[0120] Next, input the voice request sample "Adjust the air conditioner" into Reference Model1 to obtain the guiding information prediction Res1 output by ReferenceModel1, that is, "I don't understand what you mean. You can speak more standardly, such as: Adjust the air conditioner to twenty degrees, Adjust the air volume to the third gear".

[0121] Furthermore, let the voice request sample "Adjust the air conditioner" be X, let the above-mentioned guiding information label corresponding to the voice request sample "Adjust the air conditioner" be Chosen, and let the guiding information prediction Res1 output by Reference Model1 through X be Rejected. Then the server can determine the probability Prop1 that Reference Model1 predicts Chosen through X, and determine Prop2 that Reference Model1 predicts Rejected through X.

[0122] At the same time, the server can also input X into Policy Model to determine the probability Prop3 that Policy Model predicts Chosen through X, and determine Prop4 that Policy Model predicts Rejected through X.

[0123] Furthermore, complete the parameter update of Policy Model in the first training round according to the similarity between Prop1 and Prop3, and the similarity between Prop2 and Prop4 to obtain Policy Model1.

[0124] Subsequently, for the second training round, the server can copy Policy Model1 to obtain PolicyModel2, and input X into Policy Model1 to obtain the guiding information prediction Res2 output by Policy Model1, that is, "Sorry, I didn't understand what you meant. I suggest you say, such as: Adjust the air conditioner to twenty degrees, Adjust the air volume to the third gear".

[0125] Then, based on the probability Prop2 that the Policy Model predicts Chosen through X in the first training round, and the determined Prop4 that the Policy Model predicts Rejected through X, and according to the probability Prop5 that Policy Model1 predicts Chosen through X and the probability Prop6 that Policy Model1 predicts Rejected through X in the second training round (i.e., the current training round), the parameter update of Policy Model1 in the second training round is completed through the similarity between Prop3 and Prop5 and the similarity between Prop4 and Prop6 to obtain Policy Model3 (not shown in Figure 4 ), and the value of Chosen is updated from Res1 to Res2 to perform the training of the third training round and subsequent multiple training rounds, and finally the large language model of the embodiment of the present application is obtained.

[0126] It can be understood that when the server copies Policy Model1 to obtain Policy Model2, the copied Policy Model2 can be regarded as a new Reference Model, so the copied Policy Model2 can be set as Reference Model2.

[0127] Optionally, in some embodiments of the present application, the training objective of the second reference model includes: making the probability of Chosen in the Policy Model continuously greater than the probability of Rejected, and making the probability of Chosen in the reference model gradually less than Rejected.

[0128] In addition, it can also be understood that in the embodiment of the present application, "the labels used in the training process of the second reference model" are "the labels used in the training process of the first reference model". In other words, the embodiment of the present application enables the labels in the previous model training stage to be carried over to the next model training stage. Therefore, it can be understood that the embodiment of the present application can avoid the situation where, after the training of the first reference model is completed and before the training of the second reference model starts, the annotator needs to construct training samples and perform sample annotation, thereby reducing the difficulty of sample acquisition, sample annotation, and sample annotation cost in the model training process.

[0129] Optionally, in some embodiments of the present application, the server is for a basic model applicable to general natural language processing tasks but unfamiliar with voice request processing tasks in the vehicle field. Therefore, the server in the embodiments of the present application can also perform corresponding pre-training and / or knowledge injection on the basic model so that the basic model can understand the semantics of unique words, phrases, and sentences in the vehicle field, thereby obtaining the above-mentioned first reference model.

[0130] In some embodiments of the present application, step 02 includes:

[0131] In the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model, the target guidance information is determined according to the large language model and the current voice request.

[0132] The determination module in the embodiments of the present application is further configured to determine the target guidance information according to the large language model and the current voice request in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model.

[0133] The processor in the embodiments of the present application is further configured to determine the target guidance information according to the large language model and the current voice request in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model.

[0134] Specifically, to reduce the number of module deployments and make full use of the trainable parameters in the large language model, the large language model in the embodiments of the present application also learns the knowledge of "outputting the vehicle control instruction corresponding to the voice request according to the received voice request input" during the training process. Furthermore, the server in the embodiments of the present application can call the large language model to predict the vehicle control instruction for the current voice request.

[0135] Further, if the large language model can predict the corresponding vehicle control instruction according to the current voice request, the server can feedback the vehicle control instruction to the vehicle so that the vehicle performs corresponding operations according to the vehicle control instruction.

[0136] Furthermore, if there is a semantic ambiguity in the current voice request, resulting in the large language model being unable to confirm the vehicle control instruction corresponding to the current voice request, the large language model will also predict the guidance information according to the current voice request, thereby obtaining the target guidance information, and feedback the target guidance information to the user to guide the user to adjust the current voice request.

[0137] Thus, in the embodiments of the present application, the server can confirm the vehicle control instruction of the current voice request through the large language model, and in the case where the large language model fails to predict the vehicle control instruction of the current voice request, determine the target guidance information corresponding to the current voice request based on the large language model, so that the prediction of the vehicle control instruction of the current voice request and the prediction of the target guidance information of the current voice request can both be realized based on the large language model, avoiding the situation where modules for the vehicle control instruction corresponding to the voice request and modules for predicting the target guidance information need to be separately set in the server, and reducing the load of the server.

[0138] In some embodiments of the present application, the step of determining the target guidance information according to the large language model and the current voice request in the case where the vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model includes:

[0139] In the case where the vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, determine the target guidance information according to the current voice request, the large language model, and the pre-configured prompt information template.

[0140] The determination module in the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model, and the pre-configured prompt information template in the case where the vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model.

[0141] The processor in the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model, and the pre-configured prompt information template in the case where the vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model.

[0142] To ensure that the large language model can predict the guidance information corresponding to the current voice request when receiving the current voice request, in the embodiments of the present application, the pre-configured prompt (prompt or instruction) information template can be input into the large language model at the same time when the current voice request is input into the large language model.

[0143] It can be understood that the prompt information template can be understood as a natural language text segment, including "natural language text describing the downstream task that the model needs to process", and also including "explanatory text describing how the model should reason about the input information".

[0144] It is also understandable that the prompt information template in the embodiments of the present application is content that can be set according to actual situations. Exemplarily, in one example, the prompt information template may include: "Suppose you are an intelligent voice assistant, and recommend relevant annotation instructions to the user according to the instructions provided by the user."

[0145] Among them, in the example of the above prompt information template, "instruction" can be understood as the voice request in the embodiments of the present application.

[0146] Optionally, in some embodiments of the present application, after the server combines the prompt information template and the current voice request to form prompt information, the server inputs the prompt information into the large language model, so that the large language model completes the inference work of the current voice request according to the prompt information template and the current voice request in the prompt information.

[0147] In this way, in the embodiments of the present application, the large language model can perform natural language understanding and processing on the current voice request based on the indication of the prompt information template, so as to generate the target guidance information of the current voice request, ensuring the stable operation of the large language model.

[0148] In some embodiments of the present application, the prompt information template includes a first prompt information sub-template and a second prompt information sub-template. In the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model, the step of determining the target guidance information according to the current voice request, the large language model and the pre-configured prompt information template includes:

[0149] In the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first prompt information sub-template, determine the target guidance information according to the current voice request, the large language model and the second prompt information sub-template.

[0150] The determination module in the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model and the second prompt information sub-template in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first prompt information sub-template.

[0151] The processor in the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model and the second prompt information sub-template in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first prompt information sub-template.

[0152] Specifically, to ensure the reliable reasoning of the large language model in the embodiments of the present application, the prompt information template in the embodiments of the present application can be divided into two parts. One part is the first sub-prompt information template that can be understood as being used to prompt the large language model to recognize the "vehicle control instruction corresponding to the current voice request". The other part is the second sub-prompt information template that can be understood as being used to prompt the large language model to generate the "target guidance information corresponding to the current voice request".

[0153] In one example, the prompt information template including the first sub-prompt information template and the second sub-prompt information template is as follows: "Suppose you are an intelligent voice assistant. First step, according to the instruction provided by the user, select the API that can execute or provide information, and extract the key information in the instruction for the API to search for relevant information. The returned result includes, API name: api, key information: arguments. Second step: If the API is unclear, recommend relevant labeled instructions to the user. If it is not unclear, directly return the API result."

[0154] Furthermore, it can be understood that the "First step, according to the instruction provided by the user, select the API that can execute or provide information, and extract the key information in the instruction for the API to search for relevant information. The returned result includes, API name: api, key information: arguments" in the above example can be the first sub-prompt information template, while the "If the API is unclear, recommend relevant labeled instructions to the user. If it is not unclear, directly return the API result" in the above example can be the second sub-prompt information template.

[0155] Exemplarily, please refer to Figure 5 , Figure 5 which is a schematic flow diagram of the voice interaction method in some embodiments of the present application. That is, when the current voice request forwarded by the vehicle received by the server is "adjust the air conditioner", the server inputs "adjust the air conditioner" and the above prompt information template into the large language model. When the large language model confirms that the API corresponding to "adjust the air conditioner" is unclear based on the first sub-prompt information template, the large language model can infer the recommended voice request corresponding to "adjust the air conditioner" (i.e., "relevant labeled instructions") based on the second sub-prompt information template.

[0156] Furthermore, in the case of recommended voice requests such as "adjust the air conditioner to 20 degrees", "adjust the air volume to gear 3", and "open the air conditioner page" inferred, the large language model can complete and perfect these recommended voice requests to form a complete target guidance information, such as "I don't understand what you mean. You can speak more standardly, such as: adjust the air conditioner to 20 degrees, adjust the air conditioner air volume to gear 3, open the air conditioner page".

[0157] Furthermore, after the target guidance information is sent to the vehicle, the vehicle can play the target guidance information to guide the user to complete the adjustment of the current voice request "adjust the air conditioner" based on the target guidance information.

[0158] Optionally, in some embodiments of the present application, the server can train the model based on the prompt information template, samples, and sample labels during the pre-training process of the large language model, and control the large language model to perform corresponding inference work through the prompt information template and voice requests during the post-training inference process. Therefore, to ensure the reliability of the large language model, the prompt information templates used in the pre-training process and the prompt information templates used in the post-training inference process of the embodiments of the present application can be the same.

[0159] In this way, in the embodiments of the present application, the large language model can confirm the correspondence between the current voice request and the vehicle control instruction based on the first sub-prompt information template, the second sub-prompt information template, and the current voice request in the prompt information template, and when it is confirmed that the current voice request corresponds to the vehicle control instruction, confirm the target guidance information corresponding to the current voice request, realizing the inference of the large language model based on the prompt information template. Therefore, the inference accuracy of the large language model can be guaranteed to a certain extent.

[0160] Optionally, please refer to Figure 6 , Figure 6 which is a schematic flowchart of the voice interaction method in some embodiments of the present application. That is, in the embodiments of the present application, to ensure that the large language model can recognize whether the current voice request corresponds to the vehicle control instruction, corresponding downstream tasks and prompts are set so that the large language model can distinguish between "voice requests corresponding to vehicle control instructions" and "voice requests not corresponding to vehicle control instructions" and perform natural language processing on them separately.

[0161] Specifically, as Figure 6 shown, the server can train a basic model with certain natural language processing capabilities based on the second voice request sample, the API label corresponding to the second voice request sample, and the prompt corresponding to the second voice request sample, so that the basic model can distinguish between "voice requests with an API label of unclear" and "voice requests with other APIs outside the label of unclear".

[0162] It can be understood that the prompt corresponding to the second voice request sample is "Suppose you are an intelligent voice assistant. You can, according to the instructions provided by the user, select the API that can execute this information, and extract the key information in this instruction for searching relevant information using the API. The returned result includes: API name: api, key information: arguments. Instruction: Help me open the window. Output: {API: WindowsOpen, arguments: {drivce: window}}". Also, the "Instruction: Help me open the window. Output: {API: WindowsOpen, arguments: {drivce: window}}" in this prompt can be understood as an example.

[0163] Also, the server can train the basic model based on the third voice request sample, the vehicle control instruction label corresponding to the third voice request sample, and the prompt corresponding to the second voice request sample, so that the basic model can infer the vehicle control instruction corresponding to the "voice request of different APIs".

[0164] It can be understood that the prompt corresponding to the third voice request sample is "Suppose you are an intelligent voice assistant. You can generate a standard user instruction according to the function point, API, and key information arguments provided by the user. Input: {API: WindowsOpen, Arguments: {device: window}}. Output: Standard instruction: Open the window, Open the window for a while". Also, the "Input: {API: WindowsOpen, Arguments: {device: window}}. Output: Standard instruction: Open the window, Open the window for a while" in this prompt can be understood as an example.

[0165] In addition, the server can train the basic model based on the "fourth voice request sample that does not correspond to the vehicle control instruction", the recommended voice request label corresponding to the fourth voice request sample, and the prompt corresponding to the second voice request sample, so that when the basic model faces a voice request with an API of unclear, it can infer the recommended voice request corresponding to (or, semantically similar to) this voice request (that is, Figure 6 the'relevant standard instruction' in

[0166] It can be understood that the prompt corresponding to the fourth voice request sample is "Suppose you are an intelligent voice assistant. You can provide a standard instruction for the user according to the incomplete or unclear instruction provided by the user and give the user options. Input: window. Output: Standard instruction: Open the window, Close the window". Also, the "Input: window. Output: Standard instruction: Open the window, Close the window" in this prompt can be understood as an example.

[0167] It can also be understood that after completing the training of the basic model based on the above-mentioned second voice request sample, third voice request sample, and fourth voice request sample, and the labels and prompts corresponding to these three voice request samples respectively, the first reference model in the implementation manner of this application can be obtained.

[0168] Optionally, in some implementation manners of this application, the basic model can be understood as a model suitable for conventional natural language processing tasks in the general field but relatively unfamiliar with voice request processing tasks in the vehicle field. Therefore, the server in the implementation manner of this application can also perform corresponding pre-training and / or knowledge injection on the basic model before the training shown in Figure 6 so that the basic model can understand the semantics of unique words, phrases, and sentences in the vehicle field, thereby improving the adaptability of the basic model to voice request processing tasks in the vehicle field.

[0169] This application also provides a computer-readable storage medium storing a computer program, which when executed by one or more processors, implements the above-mentioned voice interaction method.

[0170] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials, or characteristics described in connection with the implementation manners or examples are included in at least one implementation manner or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same implementation manner or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more implementation manners or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0171] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred implementation manner of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0172] Although the above-mentioned implementation manners of this application have been shown and described, it can be understood that the above-mentioned implementation manners are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above-mentioned implementation manners within the scope of this application.

Claims

1. A voice interaction method, characterized in that: include: Receive the current voice request forwarded by the vehicle; In a case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining target guidance information for guiding a user to adjust the current voice request according to the large language model and the current voice request, wherein the large language model is capable of generating guidance information according to the voice request, and each training round of the large language model is trained based on a first guidance information prediction result of a previous training round and a second guidance information prediction result of a current training round; Feedback the target guidance information to guide the user to complete the voice interaction; The first guide information prediction result includes a first probability of the guide information label determined by the second reference model according to the voice request sample in the previous training round, and a third probability of the guide label prediction determined by the second reference model according to the voice request sample in the previous training round, the second guide information prediction result includes a second probability of the guide information label determined by the second reference model according to the voice request sample in the current training round, and a fourth probability of the guide label prediction determined by the second reference model according to the voice request sample in the current training round, the second reference model is obtained by training a predetermined first reference model, and the training steps of the large language model include: According to the similarity between the first probability and the second probability, and the similarity between the third probability and the fourth probability, the second reference model is trained to obtain the large language model.

2. The method according to claim 1, characterized in that The training steps of the large language model include: Acquire a voice request sample and a guidance information tag corresponding to the voice request sample; According to the voice request sample and the guide information label, a predetermined first reference model is trained to determine the large language model.

3. The method according to claim 2, characterized in that The step of training a predetermined first reference model according to the voice request sample and the guide information label to determine the large language model comprises: The first reference model is trained according to the voice request sample and the guide information label to obtain a second reference model.

4. The method according to claim 1, characterized in that The method of determining target guidance information for guiding a user to adjust the current voice request according to the large language model and the current voice request when a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request includes: When a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, the target guidance information is determined based on the large language model and the current voice request.

5. The method according to claim 4, characterized in that The method of determining the target guidance information according to the large language model and the current voice request when a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model includes: When a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request and the large language model, the target guidance information is determined based on the current voice request, the large language model and a pre-configured prompt information template.

6. The method according to claim 5, characterized in that The prompt information template includes a first prompt information sub-template and a second prompt information sub-template, and when a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request and the large language model, determining the target guidance information according to the current voice request, the large language model and a pre-configured prompt information template includes: When a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request, the large language model and the first prompt information sub-template, the target guidance information is determined based on the current voice request, the large language model and the second prompt information sub-template.

7. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Neural network linguistic model training method and device, equipment and storage medium

    CN110379416A

  • Refrigerator and intelligent refrigerator system

    CN115077158A

  • Chinese character pronunciation conversion method, electronic equipment and storage medium

    CN116484806A

  • Voice interaction method, server and computer readable storage medium

    CN117373456A