Voice interaction method, server, and computer-readable storage medium
By receiving voice requests forwarded by the vehicle, using the vehicle function knowledge graph and large language model to determine the guidance information, the problem of unclear user expression is solved, and the success rate and user experience of voice interaction are improved.
Patent Information
- Application Number
- CN202410445051.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-04-12
AI Technical Summary
When users use the vehicle for the first time, due to their unfamiliarity with the vehicle and voice interaction functions, they may express vague or incomplete voice commands, which makes it difficult for the vehicle to understand and perform corresponding functions, resulting in the failure of voice interaction.
By receiving voice requests forwarded by the vehicle, the target guidance information is determined using the pre-constructed vehicle function knowledge graph and large language model to guide the user to adjust the voice request and then complete the voice interaction.
It effectively solves the problem of user unclear voice commands, improves the success rate of voice interaction, and improves the user's experience of vehicle voice interaction functions.
Smart Images

Figure CN118116382B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and particularly to a voice interaction method, a server, and a computer-readable storage medium. Background Art
[0002] When a user starts using a vehicle for the first time, since they are relatively unfamiliar with the vehicle and its voice interaction function, they may express some vague and incomplete voice commands. When the vehicle processes these voice commands, it is difficult to understand the actual meaning of the voice commands, and thus cannot execute the corresponding functions, resulting in a failed voice interaction. Summary of the Invention
[0003] This application provides a voice interaction method, a server, and a computer-readable storage medium.
[0004] An embodiment of this application provides a voice interaction method, including:
[0005] Receiving a current voice request forwarded by a vehicle;
[0006] Determining target vehicle function knowledge information according to the current voice request and a pre-constructed vehicle function knowledge graph;
[0007] In the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining target guidance information for guiding the user to adjust the current voice request according to a large language model, the current voice request, and the target vehicle function knowledge information, the large language model being pre-trained and capable of determining guidance information according to a voice request and vehicle function knowledge information;
[0008] Feeding back the target guidance information to guide the user to complete the voice interaction.
[0009] In the voice interaction method provided by the embodiment of this application, the server can receive a voice request forwarded by a vehicle, determine target vehicle function knowledge information according to the current voice request and a pre-constructed vehicle function knowledge graph, and in the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determine target guidance information for guiding the user to adjust the current voice request according to the current voice request, the target vehicle function knowledge information, and a pre-trained large language model, and feed back the target guidance information to enable the user to complete the voice interaction through the guidance of the target guidance information.
[0010] Thus, in the embodiments of the present application, when the server fails to determine the corresponding vehicle control instruction through the voice request expressed by the user, it can determine the target guidance information for guiding the user to adjust the voice request, and feedback the target guidance information to the user to guide the user to adjust the voice request, thereby completing the voice interaction, and the user experience of using the vehicle and the vehicle voice interaction function is guaranteed. Moreover, the embodiments of the present application can determine the target vehicle function knowledge according to the current voice request and the pre-constructed vehicle function knowledge graph, so that the large language model can predict the target guidance information corresponding to the current voice request according to the target vehicle function knowledge, and the credibility of the target guidance information can be guaranteed.
[0011] In some embodiments of the present application, the determining the target vehicle function knowledge information according to the current voice request and the pre-constructed vehicle function knowledge graph includes:
[0012] Search the vehicle function knowledge graph according to the current voice request to determine the target vehicle function knowledge information semantically related to the current voice request.
[0013] Thus, in the embodiments of the present application, the server can search the vehicle function knowledge graph through the current voice request to obtain the target vehicle function knowledge information semantically related to the current voice request, so that the target vehicle function knowledge information can be matched to a certain extent with the current usage requirements of the user, and the credibility of the target guidance information determined by the large language model, the target vehicle function knowledge information and the current voice request can be guaranteed.
[0014] In some embodiments of the present application, the searching the vehicle function knowledge graph according to the current voice request to determine the target vehicle function knowledge information semantically related to the current voice request includes:
[0015] Determine the key information of the current voice request according to the current voice request;
[0016] Search the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information.
[0017] Thus, the embodiments of the present application can determine and search the vehicle function knowledge graph according to the key information of the current voice request to determine the target vehicle function knowledge information, so that during the process of searching for the target vehicle function knowledge information, the interference of noise in the current voice request can be reduced, and thus the accuracy of the target vehicle function knowledge information can be guaranteed.
[0018] In some embodiments of the present application, the vehicle function knowledge graph is constructed based on the pre-determined vehicle devices and the functions that the vehicle devices can perform. Searching the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information includes:
[0019] Searching the vehicle function knowledge graph according to the key information to determine the target device and target function corresponding to the key information so as to determine the target vehicle function knowledge information.
[0020] In this way, in the embodiments of the present application, the server can search the vehicle function knowledge graph through the key information of the current voice request to determine the target device and target function corresponding to the key information, so as to obtain the target vehicle function knowledge information, enabling the target vehicle function knowledge information to represent the vehicle devices and vehicle functions related to the current voice request.
[0021] In some embodiments of the present application, the vehicle function knowledge graph includes nodes of multiple levels. Searching the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information includes:
[0022] Searching the vehicle function knowledge graph according to the key information to obtain the nodes of each level corresponding to the key information so as to determine the target vehicle function knowledge information.
[0023] In this way, in the embodiments of the present application, for the vehicle function knowledge graph including nodes of multiple levels, the server can use the key information of the current voice request to search the vehicle function knowledge graph to obtain the nodes of each level corresponding to the key information to determine the target vehicle function knowledge.
[0024] In some embodiments of the present application, determining the key information of the current voice request according to the current voice request includes:
[0025] Determining the key information of the current voice request according to the current voice request and the pre-trained named entity extraction model.
[0026] In this way, in the embodiments of the present application, the server can call the pre-trained named entity extraction model to determine the key information of the current voice request, ensuring the accuracy and reliability of the key information to a certain extent.
[0027] In some embodiments of the present application, in the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, according to the large language model, the current voice request, and the target vehicle function knowledge information, determining target guidance information for guiding the user to adjust the current voice request includes:
[0028] Determining current prompt information according to the target vehicle function knowledge information and a pre-configured prompt information template;
[0029] In the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining the target guidance information according to the large language model, the current voice request, and the current prompt information.
[0030] In this way, in the embodiments of the present application, the server can determine the current prompt information according to the target vehicle function knowledge information and the pre-configured prompt information template, and determine the target guidance information according to the large language model, the current prompt information, and the current voice request, so that the large language model can infer the target guidance information corresponding to the current voice request based on the assistance of the current prompt information, thus ensuring the accuracy and reliability of the target guidance information.
[0031] In some embodiments of the present application, the current prompt information includes a first sub-prompt information and a second sub-prompt information. In the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining the target guidance information according to the large language model, the current voice request, and the current prompt information includes:
[0032] In the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model, and the first sub-prompt information, determining the target guidance information according to the current voice request, the large language model, and the second sub-prompt information.
[0033] In this way, in the embodiments of the present application, the large language model can determine the correspondence between the current voice request and the vehicle control instruction based on the current voice request and the first sub-prompt information and the second sub-prompt information in the current prompt information, and determine the target guidance information in the case where a vehicle control instruction corresponding to the current voice request cannot be determined.
[0034] The embodiments of the present application provide a server, including a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the above-mentioned voice interaction method is implemented.
[0035] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by one or more processors, implements the above-mentioned voice interaction method.
[0036] The server and computer-readable storage medium provided by the embodiments of the present application can determine target guidance information for guiding the user to adjust the voice request and feedback the target guidance information to the user in the case where the corresponding vehicle control instruction cannot be determined through the voice request expressed by the user, so as to guide the user to adjust the voice request, thereby completing the voice interaction, and the user's experience of using the vehicle and the vehicle voice interaction function is guaranteed. In addition, the embodiments of the present application can determine target vehicle function knowledge according to the current voice request and the pre-constructed vehicle function knowledge graph, so that the large language model can predict the target guidance information corresponding to the current voice request according to the target vehicle function knowledge, and the credibility of the target guidance information can be guaranteed.
[0037] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0039] Figure 1 is a schematic flow chart of the voice interaction method in some embodiments of the present application;
[0040] Figure 2 is a schematic diagram of the vehicle function knowledge graph in some embodiments of the present application;
[0041] Figure 3 is a schematic flow chart of the voice interaction method in some embodiments of the present application;
[0042] Figure 4 is a schematic flow chart of the voice interaction method in some embodiments of the present application;
[0043] Figure 5 is a schematic flow chart of the voice interaction method in some embodiments of the present application;
[0044] Figure 6 is a schematic flow chart of the voice interaction method in some embodiments of the present application;
[0045] Figure 7 is a schematic flow chart of the voice interaction method in some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and should not be construed as a limitation on the embodiments of the present application.
[0047] When a user first uses a vehicle and the voice interaction function carried by the vehicle, since the user is unfamiliar with the objects and functions that can be controlled by voice in the vehicle, some voice commands with ambiguous semantics, chaotic or missing sentence components may be expressed, such as "Adjust the air conditioner", "Turn on 'that thing for adjusting the temperature'", etc.
[0048] It can be understood that these voice commands have one or more of the problems such as unclear intent, missing (or ambiguous) components, and incorrect collocations of sentence components. For example, it can be considered that the collocation relationship between 'adjust' and 'air conditioner' in "Adjust the air conditioner" is inappropriate. It can also be considered that the intent corresponding to "Adjust the air conditioner" includes increasing or decreasing the cooling / heating temperature of the air conditioner and switching the operation mode of the air conditioner, etc. Therefore, the intent of "Adjust the air conditioner" is unclear.
[0049] Therefore, when processing these voice commands, since these voice commands have one or more problems such as unclear intent, missing (or ambiguous) components, and incorrect collocations of sentence components, it is difficult to determine the specific vehicle control commands for implementing or completing these voice commands, which in turn leads to the failure of the voice interaction with the user this time.
[0050] Furthermore, to improve this situation, some solutions are to play timed dynamic advertisement examples, that is, to regularly play predetermined audio-visual data to teach users to use specific functions of the vehicle. However, this solution is difficult to meet the current usage needs of users.
[0051] In some other solutions, the user is requested to clarify these voice commands. That is to say, when the user expresses a voice command with ambiguous semantics and incomplete sentence components, a counter-question is sent to the user. For example, when the voice command spoken by the user is "Adjust the air conditioner", the vehicle can play a voice such as "I didn't understand what you meant. Please be more specific" to request the user to clarify the true intent of this voice command.
[0052] However, when asking the user for clarification, interactions with the user are usually based on fixed statements, such as "I didn't understand what you meant. Please be more specific." In this case, if the user is using the vehicle and its voice interaction function for the first time or just starting to use it, and thus is unfamiliar with the objects and functions that can be controlled by voice in the vehicle, when faced with a rhetorical question (or, a clarification request) from the vehicle, the user may not know how to clarify, or rather, may not know "how to formulate a voice command to make the vehicle execute the corresponding function", resulting in the need for the user to have multiple rounds of conversations and clarifications with the vehicle before being able to control the vehicle through correct (or, understandable by the vehicle) voice commands.
[0053] Furthermore, to mitigate the negative impacts when asking the user for clarification using fixed statements, some solutions use fixed statements containing examples to ask the user for clarification, such as "I didn't understand what you meant. Please be more specific. For example, to adjust the seat, you can say: Turn on the seat heating." This statement includes an example "To adjust the seat, you can say: Turn on the seat heating." to enable the user to clarify based on this example.
[0054] However, although such statements containing examples can play a certain guiding role, there may be situations where "it misses the mark", such as when the user says "Adjust the air conditioner" and the vehicle plays "I understand what you mean. Please be more specific. For example, to adjust the seat, you can say: Turn on the seat heating." It is understandable that the example "To adjust the seat, you can say: Turn on the seat heating." is significantly different from the user's actual intention, and the guiding role it plays is relatively limited.
[0055] Based on the above possible problems, please refer to Figure 1 , an embodiment of the present application provides a voice interaction method, including:
[0056] 01: Receive the current voice request forwarded by the vehicle;
[0057] 02: Determine the target vehicle function knowledge information according to the current voice request and the pre-constructed vehicle function knowledge graph;
[0058] 03: In the case where the vehicle control command corresponding to the current voice request cannot be determined according to the current voice request, determine the target guiding information for guiding the user to adjust the current voice request according to the large language model, the current voice request, and the target vehicle function knowledge information. The large language model is pre-trained and can determine the guiding information according to the voice request and the vehicle function knowledge information;
[0059] 04: Feedback the target guiding information to guide the user to complete the voice interaction.
[0060] Embodiments of the present application provide a voice interaction device. The voice interaction method of the embodiments of the present application can be implemented by the voice interaction device of the embodiments of the present application. Specifically, the voice interaction device includes a receiving module, a knowledge information determination module, a guidance information determination module, and an interaction module. The receiving module is used to receive the current voice request forwarded by the vehicle. The knowledge information determination module is used to determine the target vehicle function knowledge information according to the current voice request and the pre-constructed vehicle function knowledge graph. The guidance information determination module is used to determine the target guidance information for guiding the user to adjust the current voice request according to the large language model, the current voice request, and the target vehicle function knowledge information when the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request. The large language model is pre-trained and can determine the guidance information according to the voice request and the vehicle function knowledge information. The interaction module is used to feedback the target guidance information to guide the user to complete the voice interaction.
[0061] Embodiments of the present application also provide a server, which includes a memory and a processor. The voice interaction method of the embodiments of the present application can be implemented by the server of the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is used to receive the current voice request forwarded by the vehicle, and to determine the target vehicle function knowledge information according to the current voice request and the pre-constructed vehicle function knowledge graph, and to determine the target guidance information for guiding the user to adjust the current voice request according to the large language model, the current voice request, and the target vehicle function knowledge information when the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request. The large language model is pre-trained and can determine the guidance information according to the voice request and the vehicle function knowledge information, and to feedback the target guidance information to guide the user to complete the voice interaction.
[0062] Specifically, in the embodiments of the present application, when the user desires to control the vehicle to perform a specific operation by voice and thus triggers the corresponding current voice request at the current moment, the vehicle can receive and forward the current voice request to the server. The server receives the current voice request forwarded by the vehicle and determines the target vehicle function knowledge according to the current voice request and the pre-constructed vehicle function knowledge graph (Knowledge Graph).
[0063] Subsequently, the server can process the current voice request accordingly to determine the vehicle control instruction corresponding to the current voice request. Since there are semantic defects in the current voice request, such as unclear intention, missing components, incorrect collocations of sentence components, etc., the server fails to determine the vehicle control instruction corresponding to the current voice request. Therefore, the server can call the pre-trained large language model (LLM) based on the current voice request and the target vehicle function knowledge, so that the large language model can determine the target guidance information for guiding the user to adjust the current voice request with the assistance of the target vehicle function knowledge.
[0064] Further, in the case of obtaining the target guidance information through the large language model, the server can send the target guidance information to the vehicle. When the vehicle receives the target guidance information, it can display and / or play the target guidance information through the image display component and / or the sound playback component to guide the user to adjust the current voice request, such as adjusting it to a valid voice request for which the vehicle can determine the corresponding vehicle control instruction, so that the vehicle and the server can complete the voice interaction with the user through this valid voice request.
[0065] Exemplarily, when the user wants to adjust the fragrance of the vehicle air conditioning system to emit "fragrance A" (hereinafter referred to as "fragrance A"), and thus the triggered current voice request is "adjust fragrance", the server can receive "adjust fragrance" obtained and forwarded by the vehicle and, in combination with the vehicle function knowledge graph, determine that the target vehicle function knowledge is "{air conditioning -> fragrance -> [fragrance A, fragrance B]}".
[0066] In the case where the server fails to determine the vehicle control instruction corresponding to "adjust fragrance", the server can use both "adjust fragrance" and "{air conditioning -> fragrance -> [fragrance A, fragrance B]}" as the input to the large language model, so that the large language model outputs the target guidance information, that is, "Sorry, I didn't understand what you meant. If you want to turn on the fragrance, you can say: Turn on fragrance A, Turn on fragrance B", thereby guiding the user to adjust "adjust fragrance" to "turn on fragrance A".
[0067] It can be understood that the voice requests in the embodiments of the present application can be understood as the above-mentioned voice information or text information with semantic defects, such as "adjust air conditioning", "adjust fragrance", and "turn on temperature", etc.
[0068] It can also be understood that in the case where the voice request received by the server does not have semantic defects or other problems and thus the vehicle control instruction corresponding to the voice request can be determined, the server can send the vehicle control instruction corresponding to the voice request to the vehicle, so that the vehicle can perform corresponding actions and functions according to the vehicle control instruction.
[0069] It is also understandable that the vehicle function knowledge graph in the embodiments of the present application can be understood as a graph representing the functions that the in-vehicle system can execute. Exemplarily, please refer to Figure 2 , Figure 2 which is a schematic diagram of the vehicle function knowledge graph in some embodiments of the present application. That is, as Figure 2 shown, some sub-graphs of the vehicle function knowledge graph in the embodiments of the present application may include an air conditioner, air conditioner adjustable function parameters, fragrance (a sub-function of the air conditioner), and fragrance adjustable function parameters.
[0070] Optionally, in some embodiments of the present application, the vehicle function knowledge graph is used to describe the main function points, secondary function points, and the relationships between function points in the vehicle. Exemplarily, taking Figure 2 as an example, Figure 2 the "air conditioner" in
[0071] is a primary function point, "air volume", "temperature", and "fragrance" are all secondary function points subordinate to the "air conditioner", and "scent A", "scent B", "scent C", "scent D", and "scent E" are all tertiary function points subordinate to the "fragrance".
[0072] It can also be understood that the function knowledge information (or rather, the target vehicle function knowledge information) in the embodiments of the present application can be regarded as one kind or a part of the prompt, and thus can assist the large language model in reasoning about the "guidance information for guiding the user to adjust the voice request".
[0073] It can be understood that without considering the vehicle function knowledge information, or rather, when the input parameters of the large language model do not include the vehicle function knowledge information, the large language model will complete the prediction of the guidance information based on the knowledge learned during its training process. However, after training the large language model using the pre-received training corpus and deploying the trained large language model, due to vehicle function updates or additions, the voice requests triggered by users may not be covered by the training corpus, resulting in the large language model outputting incorrect prediction results when receiving such voice requests.
[0074] Therefore, to improve the above situation, in the embodiments of the present application, before predicting the guiding information corresponding to the current voice request through the large language model, the current voice request and the pre-constructed vehicle function knowledge graph are used to determine the target vehicle function knowledge, and then the target vehicle function knowledge and the current voice request are input into the large language model together, so that the large language model can complete the reasoning function of the target guiding information with the help of external knowledge, thereby improving the situation that the large language model may infer incorrect results when performing reasoning based on the learned knowledge before being updated again.
[0075] It can also be understood that in the case where the server determines the target guiding information based on the large language model, the server can send the target guiding information to the vehicle, and the vehicle can play the target guiding information, so that the user can complete the adjustment of the current voice request based on the currently playing target guiding information and complete the voice interaction with the vehicle and the server.
[0076] Optionally, in some embodiments of the present application, the server can feedback the target guiding information to the user's terminal so that the user can complete the adjustment of the current voice request through the target guiding information played by the terminal.
[0077] In summary, in the embodiments of the present application, when the server fails to determine the corresponding vehicle control instruction through the voice request expressed by the user, the server can determine the target guiding information for guiding the user to adjust the voice request, and feedback the target guiding information to the user to guide the user to adjust the voice request, thereby completing the voice interaction, and the user's experience of using the vehicle and the vehicle voice interaction function is guaranteed. In addition, the embodiments of the present application can determine the target vehicle function knowledge according to the current voice request and the pre-constructed vehicle function knowledge graph, so that the large language model can predict the target guiding information corresponding to the current voice request according to the target vehicle function knowledge, and the credibility of the target guiding information can be guaranteed.
[0078] Moreover, the embodiments of the present application can determine the target guiding information according to the current voice request, so that the target guiding information can be related to the current voice request, thereby enabling the target guiding information to match the user's current usage requirements to a certain extent, and further ensuring the guiding role of the target guiding information. In addition, since the target guiding information is determined based on the large language model, the credibility of the target guiding information can be guaranteed.
[0079] Please refer to Figure 3 , in some embodiments of the present application, step 02 includes:
[0080] 020: Search the vehicle function knowledge graph according to the current voice request to determine the target vehicle function knowledge information semantically related to the current voice request.
[0081] The knowledge information determination module of the embodiment of the present application is further configured to search the vehicle function knowledge graph according to the current voice request, and determine the target vehicle function knowledge information related to the semantics of the current voice request.
[0082] The processor of the embodiment of the present application is further configured to search the vehicle function knowledge graph according to the current voice request, and determine the target vehicle function knowledge information related to the semantics of the current voice request.
[0083] Specifically, to ensure the credibility of the inference result output by the large language model, the server of the embodiment of the present application can use the current voice request to search the vehicle function knowledge graph, so as to obtain a sub-graph related to the semantics of the current voice request in the vehicle function knowledge graph, that is, the target vehicle function knowledge information.
[0084] For example, let the current voice request be "adjust the fragrance", and please refer to again Figure 2 the vehicle function knowledge graph shown. The server of the embodiment of the present application can use "adjust the fragrance" to search the vehicle function knowledge graph, and the obtained target vehicle function knowledge information may include: {air conditioner -> fragrance -> [fragrance A, fragrance B]}.
[0085] It can be understood that since the target vehicle function knowledge is related to the semantics of the current voice request, the degree of relevance between the target vehicle function knowledge and the user's actual needs (or, true intention) can be ensured. Furthermore, when the large language model determines the target guidance information based on the target vehicle function knowledge information, the degree of relevance between the target guidance information and the user's actual needs (or, true intention) can also be ensured, and the guiding effect of the target guidance information on the user to adjust the current voice request can also be ensured.
[0086] In this way, in the embodiment of the present application, the server can search the vehicle function knowledge graph through the current voice request to obtain the target vehicle function knowledge information related to the semantics of the current voice request, so that the target vehicle function knowledge information can be matched to a certain extent with the user's current usage needs, and the credibility of the target guidance information determined by the large language model, the target vehicle function knowledge information, and the current voice request can be ensured.
[0087] Please refer to Figure 4 , in some embodiments of the present application, step 020 includes:
[0088] 0200: Determine the key information of the current voice request according to the current voice request;
[0089] 0201: Search the vehicle function knowledge graph according to the key information, and determine the target vehicle function knowledge information.
[0090] The knowledge information determination module of the embodiment of the present application is further configured to determine the key information of the current voice request according to the current voice request, and to search the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information.
[0091] The processor of the embodiment of the present application is further configured to determine the key information of the current voice request according to the current voice request, and to search the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information.
[0092] Specifically, to ensure the credibility of the target vehicle function knowledge information and to reduce the interference of noise information in the current voice request, the server of the embodiment of the present application can also perform corresponding natural language processing on the current voice request to determine the key information of the current voice request, and search the vehicle function knowledge graph through the key information to obtain the target vehicle function knowledge information.
[0093] For example, if the current voice request is "adjust the fragrance", the server can process "adjust the fragrance" based on an algorithm or model implemented in advance by code, and thus obtain the key information "fragrance".
[0094] Furthermore, the server can search the vehicle function knowledge graph according to "fragrance", and the target vehicle function knowledge information obtained thereby may include: {air conditioner -> fragrance -> [scent A, scent B]}
[0095] In this way, the embodiment of the present application can determine and search the vehicle function knowledge graph according to the key information of the current voice request to determine the target vehicle function knowledge information, so that during the process of searching for the target vehicle function knowledge information, the interference of noise in the current voice request can be reduced, and thus the accuracy of the target vehicle function knowledge information can be ensured.
[0096] In some embodiments of the present application, the vehicle function knowledge graph is constructed according to pre-determined vehicle devices and functions that the vehicle devices can perform. Further, step 0201 includes:
[0097] Search the vehicle function knowledge graph according to the key information to determine the target device and target function corresponding to the key information to determine the target vehicle function knowledge information.
[0098] The knowledge information determination module of the embodiment of the present application is further configured to search the vehicle function knowledge graph according to the key information to determine the target device and target function corresponding to the key information to determine the target vehicle function knowledge information.
[0099] The processor according to the embodiment of the present application is further configured to search the vehicle function knowledge graph according to the key information, determine the target device and target function corresponding to the key information to determine the target vehicle function knowledge information.
[0100] Specifically, as Figure 2 shown, the vehicle function knowledge graph in the embodiment of the present application represents each device of the vehicle and the functions that each device can perform. Therefore, the server in the embodiment of the present application can search for the target device and target function corresponding to the key information through the key information of the current voice request.
[0101] In one example, let the current voice request be "adjust the fragrance", and let the key information of "adjust the fragrance" be "fragrance". Then the server can retrieve the vehicle function knowledge graph as Figure 2 shown.
[0102] It can be understood that in the Figure 2 shown vehicle function knowledge graph, the key information "fragrance" corresponds to the Figure 2 node "fragrance" in it. At the same time, there is a pointing relationship between the node "fragrance" and the node "air conditioner". Therefore, the target devices corresponding to the key information "fragrance" are the air conditioner and the fragrance.
[0103] At the same time, in the Figure 2 shown vehicle function knowledge graph, the key information "fragrance" has a pointing relationship with the Figure 2 nodes "scent A", "scent B", "scent C", "scent D" and "scent E" in it. Therefore, the target functions corresponding to the key information "fragrance" are "scent A", "scent B", "scent C", "scent D" and "scent E".
[0104] Thus, the target vehicle function knowledge information may include: {air conditioner -> fragrance -> [scent A, scent B, scent C, scent D, scent E]}.
[0105] Optionally, to prevent the large language model from taking too long to infer the guiding information due to the need to process long input information, in some embodiments of the present application, the server can trim the target vehicle function knowledge information. For example, for the above {air conditioner -> fragrance -> [scent A, scent B, scent C, scent D, scent E]}, it can be trimmed to {air conditioner -> fragrance -> [scent A, scent B]}, so that the information length of the target vehicle function knowledge information is reduced, and thus the length of the input information of the large language model is also reduced.
[0106] Thus, in the embodiments of the present application, the server can search the vehicle function knowledge graph through the key information of the current voice request to determine the target device and target function corresponding to the key information, so as to obtain the target vehicle function knowledge information, enabling the target vehicle function knowledge information to represent the vehicle devices and vehicle functions related to the current voice request.
[0107] In some embodiments of the present application, the vehicle function knowledge graph includes nodes of multiple levels, and thus step 0201 includes:
[0108] Search the vehicle function knowledge graph according to the key information to obtain nodes of each level corresponding to the key information, so as to determine the target vehicle function knowledge information.
[0109] The knowledge information determination module in the embodiments of the present application is further configured to search the vehicle function knowledge graph according to the key information to obtain nodes of each level corresponding to the key information, so as to determine the target vehicle function knowledge information.
[0110] The processor in the embodiments of the present application is further configured to search the vehicle function knowledge graph according to the key information to obtain nodes of each level corresponding to the key information, so as to determine the target vehicle function knowledge information.
[0111] Specifically, to ensure the reliability and effectiveness of the target vehicle function knowledge, the server in the embodiments of the present application can also search at least one node of each level corresponding to the "key information of the current voice request" during the process of searching the target vehicle function knowledge graph, thereby forming a vehicle function knowledge sub-graph corresponding to the key information, that is, the target vehicle function knowledge. In addition, it can be understood that the Knowledge Graph can be used to describe nodes (Points) and the edges (Edges) between nodes.
[0112] To illustrate the embodiments of the present application more clearly, please refer to again Figure 2 . That is, Figure 2 "Air conditioner" in
[0113] is a first-level function point, "Air volume", "Temperature", and "Fragrance" are all second-level function points subordinate to "Air conditioner", and "Fragrance A", "Fragrance B", "Fragrance C", "Fragrance D", and "Fragrance E" are all third-level function points subordinate to "Fragrance". Figure 2 Furthermore, when the current voice request is "Adjust the fragrance", and the key information of "Adjust the fragrance" includes "Fragrance", the server retrieves the vehicle function knowledge graph as shown in
[0114] through "Fragrance" and determines the second-level function point "Fragrance" in the vehicle function knowledge graph that matches the key information "Fragrance".Figure 2 The vehicle function knowledge graph shown includes a total of three levels of nodes, namely primary function points, secondary function points, and tertiary function points. Therefore, the server can search upstream of the secondary function point "aromatherapy" to determine the primary function point, and search downstream of the secondary function point "aromatherapy" to determine the tertiary function.
[0115] Therefore, by searching upstream of the secondary function point "aromatherapy", the primary function point "air conditioning" pointed to by the secondary function point "aromatherapy" can be obtained. By searching downstream of the secondary function point "aromatherapy", the tertiary function points "scent A", "scent B", "scent C", "scent D", and "scent E" pointed to by the secondary function point "aromatherapy" can be obtained.
[0116] Thus, the primary function point, secondary function point, and tertiary function point have all been searched, and thus the vehicle function knowledge subgraph corresponding to the key information "aromatherapy", or the target vehicle function knowledge, can be constructed, that is: {air conditioning -> aromatherapy -> [scent A, scent B, scent C, scent D, scent E]}.
[0117] Similarly, if the current voice request is "adjust air conditioning", and the key information of "adjust air conditioning" is "air conditioning", then based on the key information "air conditioning", search the Figure 2 vehicle function knowledge graph shown, then the primary function point "air conditioning" can be obtained, the downstream secondary function points of the primary function point "air conditioning", namely "aromatherapy", "air volume", and "temperature", and the downstream tertiary function points of the secondary function point "aromatherapy", namely "scent A", "scent B", "scent C", "scent D", and "scent E", and thus the vehicle function knowledge subgraph corresponding to the key information "air conditioning", or the target vehicle function knowledge, can be obtained, that is: {air conditioning -> [aromatherapy, air volume, temperature] -> [scent A, scent B, scent C, scent D, scent E]}.
[0118] It can be understood that when the target vehicle function knowledge information includes nodes of each level corresponding to the key information of the current voice request in the vehicle function knowledge graph, the target vehicle function knowledge information is relatively rich and can reliably represent the knowledge in the vehicle function knowledge graph related to the current voice request.
[0119] Optionally, to prevent the large language model from taking too long to infer the guiding information due to the need to process long input information, in some embodiments of the present application, the server can trim the target vehicle function knowledge information.
[0120] For example, taking the above {air conditioner -> aroma diffuser -> [scent A, scent B, scent C, scent D, scent E]} as an example, the server can trim {air conditioner -> aroma diffuser -> [scent A, scent B, scent C, scent D, scent E]} to {air conditioner -> aroma diffuser -> [scent A, scent B]}, reducing the information length of the target vehicle function knowledge information, and thus also reducing the length of the input information of the large language model.
[0121] In this way, in the embodiment of the present application, for a vehicle function knowledge graph including nodes with multiple levels, the server can utilize the key information of the current voice request to search the vehicle function knowledge graph to obtain nodes at each level corresponding to the key information to determine the target vehicle function knowledge.
[0122] In some embodiments of the present application, step 0200 includes:
[0123] Determine the key information of the current voice request according to the current voice request and a pre-trained named entity extraction model.
[0124] The knowledge information determination module in the embodiment of the present application is further configured to determine the key information of the current voice request according to the current voice request and a pre-trained named entity extraction model.
[0125] The processor in the embodiment of the present application is further configured to determine the key information of the current voice request according to the current voice request and a pre-trained named entity extraction model.
[0126] Specifically, for the accuracy of the target vehicle function knowledge information and the target guidance information, the server in the embodiment of the present application can call a pre-trained NER (Named Entity Recognition) extraction model to perform named entity recognition and extraction on the current voice request, thereby obtaining the named entities in the current voice request, that is, the key information of the current voice request.
[0127] For example, for "adjust the air conditioner" and "adjust the aroma diffuser", the server in the embodiment of the present application can use the NER extraction model to extract the entities in the current voice request, so as to obtain the key information "air conditioner" of "adjust the air conditioner" and the key information "aroma diffuser" of "adjust the aroma diffuser".
[0128] In this way, in the embodiment of the present application, the server can call a pre-trained named entity extraction model to determine the key information of the current voice request, ensuring the accuracy and reliability of the key information to a certain extent.
[0129] Optionally, in some embodiments of the present application, the NER extraction model is used to extract three named entities in the voice request, namely Function (function), Device (hardware), and Action (action).
[0130] For example, in one example, for the voice request "Adjust the fragrance", the NER extraction model can extract the Action "Adjust" and the Function "Fragrance" in "Adjust the fragrance".
[0131] Another example, in another example, for the voice request "Adjust the air conditioner", the NER extraction model can extract the Action "Adjust" and the Device "Air conditioner" in "Adjust the air conditioner".
[0132] Another example, in yet another example, for the voice request "Adjust the air volume", the NER extraction model can extract the Action "Adjust" and the Argument "Air volume" in "Adjust the air conditioner".
[0133] Optionally, in some embodiments of the present application, if the NER extraction model extracts the current voice request and the named recognition obtained by the extraction includes one of Argument, Function, and Device, the server can use "one of Argument, Function, and Device" as the key information to search the vehicle function knowledge graph.
[0134] Optionally, if the named recognition obtained by the extraction only includes Action, the key information of the current voice request is empty or null.
[0135] Optionally, if the named recognition obtained by the extraction includes at least two of Argument, Function, and Device, if the Device is not empty, the Device is used as the key information to search the vehicle function knowledge graph, and if the Device is empty, the Argument is used as the key information to search the vehicle function knowledge graph.
[0136] Please refer to Figure 5 , in some embodiments of the present application, step 03 includes:
[0137] 030: Determine the current prompt information according to the target vehicle function knowledge information and the pre-configured prompt information template;
[0138] 031: In the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determine the target guidance information according to the large language model, the current voice request, and the current prompt information.
[0139] The guidance information determination module of the embodiment of the present application is further configured to determine the current prompt information according to the target vehicle function knowledge information and the pre-configured prompt information template, and to determine the target guidance information according to the large language model, the current voice request, and the current prompt information when the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request.
[0140] The processor of the embodiment of the present application is further configured to determine the current prompt information according to the target vehicle function knowledge information and the pre-configured prompt information template, and to determine the target guidance information according to the large language model, the current voice request, and the current prompt information when the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request.
[0141] Specifically, to ensure the accuracy and reliability of the guidance information prediction result of the large language model, the server of the embodiment of the present application may also input the pre-configured prompt (prompt or instruction) information template, the current voice request, and the target vehicle function knowledge information into the large language model together, so that the large language model can understand and reason about the current voice request and the target vehicle function knowledge information according to the natural language understanding method and reasoning method indicated or implied by the prompt information template, and then output the target guidance information.
[0142] At the same time, to ensure that the large language model can accurately understand the prompt information template, the target vehicle function knowledge information, and the current voice request, the server of the embodiment of the present application may also construct the current prompt information based on the target vehicle function knowledge information and the prompt information template.
[0143] Exemplarily, in one example, the prompt information template includes: "Suppose you are an intelligent voice assistant, and recommend relevant marked instructions to the user according to the instructions provided by the user. Knowledge sub-graph:".
[0144] Moreover, the current voice request is "Adjust the fragrance", and the target vehicle function knowledge information corresponding to the current voice request is {Air conditioner -> Fragrance -> [Fragrance A, Fragrance B]}.
[0145] Furthermore, the current prompt information determined according to the prompt information template and the target vehicle function knowledge information may include: "Suppose you are an intelligent voice assistant, and recommend relevant marked instructions to the user according to the instructions provided by the user. Knowledge sub-graph: Air conditioner -> Fragrance -> [Fragrance A, Fragrance B].".
[0146] It can be understood that the "instructions provided by the user" in the above example can be understood as the current voice request, and the "knowledge sub-graph" can be understood as the target vehicle function knowledge information determined by the current voice request and the vehicle function knowledge graph.
[0147] Thus, in the embodiments of the present application, the server can determine the current prompt information according to the target vehicle function knowledge information and the pre-configured prompt information template, and determine the target guidance information according to the large language model, the current prompt information and the current voice request, so that the large language model can infer the target guidance information corresponding to the current voice request with the assistance of the current prompt information, thus ensuring the accuracy and reliability of the target guidance information.
[0148] In some embodiments of the present application, the current prompt information includes the first sub-prompt information and the second sub-prompt information. Further, step 031 includes:
[0149] In the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first sub-prompt information, determine the target guidance information according to the current voice request, the large language model and the second sub-prompt information.
[0150] The guidance information determination module of the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model and the second sub-prompt information in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first sub-prompt information.
[0151] The processor of the embodiments of the present application is further configured to determine the target guidance information according to the current voice request, the large language model and the second sub-prompt information in the case where the vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the large language model and the first sub-prompt information.
[0152] Specifically, to reduce the operating load of the server and reuse the trainable parameters in the large language model, the large language model of the embodiments of the present application not only has the ability to "predict the target guidance information corresponding to the current voice request", but also has the ability to "predict the vehicle control instruction corresponding to the current voice request".
[0153] Therefore, the server of the embodiments of the present application can determine the vehicle control instruction corresponding to the current voice request based on the large language model, and determine the target guidance information based on the current voice request and the current prompt information in the case where the vehicle control instruction corresponding to the current voice request cannot be determined.
[0154] Exemplarily, in some embodiments of the present application, the pre-configured prompt information template may include: "If you are an intelligent voice assistant, first, according to the instructions provided by the user, select an API that can execute or provide information, and extract the key information in the instructions for the API to search for relevant information. The return results include, API name: api, key information: arguments. Second step: If the API is unclear, recommend relevant annotation instructions to the user. Knowledge subgraph: "
[0155] And, the current voice request is "Adjust the fragrance", and the target vehicle function knowledge information corresponding to the current voice request is {Air conditioner -> Fragrance -> [Fragrance A, Fragrance B]}.
[0156] Furthermore, the current prompt information determined according to the prompt information template and the target vehicle function knowledge information may include: "If you are an intelligent voice assistant, first, according to the instructions provided by the user, select an API that can execute or provide information, and extract the key information in the instructions for the API to search for relevant information. The return results include, API name: api, key information: arguments. Second step: If the API is unclear, recommend relevant annotation instructions to the user. Knowledge subgraph: Air conditioner -> Fragrance -> [Fragrance A, Fragrance B]."
[0157] It should be noted that "First, according to the instructions provided by the user, select an API that can execute or provide information, and extract the key information in the instructions for the API to search for relevant information. The return results include, API name: api, key information: arguments" in the above example can be understood as the first sub-prompt information.
[0158] And, "Second step: If the API is unclear, recommend relevant annotation instructions to the user. Knowledge subgraph: Air conditioner -> Fragrance -> [Fragrance A, Fragrance B]" in the above example can be understood as the second sub-prompt information.
[0159] Therefore, after the server inputs the current voice request and the current prompt information into the large language model, the large language model can predict the API (Application Interface) corresponding to the current voice request according to the first sub-prompt information in the current voice request and the current prompt information to determine the vehicle control instruction.
[0160] Further, when the large language model predicts that the API corresponding to the current voice request is Unclear, so the current voice request does not correspond to a vehicle control instruction, the large language model predicts target guiding information for guiding the user to adjust the current voice request based on the current voice request and the first sub-guiding information in the current prompting information, such as: "Sorry, I didn't understand what you meant. If you want to turn on the fragrance, you can say: Turn on fragrance A, Turn on fragrance B".
[0161] Thus, in the embodiment of the present application, the large language model can determine the correspondence between the current voice request and the vehicle control instruction based on the current voice request, and the first sub-guiding information and the second sub-guiding information in the current prompting information, and determine the target guiding information when the vehicle control instruction corresponding to the current voice request cannot be determined.
[0162] Moreover, it can be understood that since both the prediction of the vehicle control instruction of the current voice request and the prediction of the target guiding information can be realized based on the large language model, there is no need to set up a module dedicated to the prediction of the vehicle control instruction and a module dedicated to the prediction of the target guiding information in the server, and the number of modules in the server is reduced, so that the operating load of the server is reduced.
[0163] At the same time, when the large language model includes a large number of training parameters, enabling the large language model to have the capabilities of "predicting the target guiding information corresponding to the current voice request" and "predicting the vehicle control instruction corresponding to the current voice request" can, to a certain extent, ensure the full use of the large number of parameters in the large language model.
[0164] Optionally, please refer back to Figure 2 and refer to Figure 6 and Figure 7 ., Figure 6 and Figure 7 are all flow diagrams of the voice interaction method in some embodiments of the present application. That is, as Figure 6 shown, the current voice request triggered by the user at the current moment is "Adjust the fragrance", and the vehicle obtains the current voice request and forwards or reports the current voice request to the server.
[0165] After receiving the current voice request forwarded by the vehicle, the server can, as Figure 7 shown, extract the key information of the current voice request to extract the key information of the current voice request ("Adjust the fragrance"), that is, {"function: fragrance"}.
[0166] Subsequently, the server can, based on the key information of the current voice request, for example, as Figure 2Retrieve and perform subgraph matching on the vehicle function knowledge graph shown to obtain a knowledge subgraph and target vehicle function information corresponding to the key information of the current voice request, that is: air conditioner -> fragrance -> [Fragrance A, Fragrance B].
[0167] Then, as Figure 6 shown, the server can construct the current prompt message based on the target vehicle function information and the pre-configured prompt message template. The pre-configured prompt message template may include: "Suppose you are an intelligent voice assistant. First step, according to the instruction provided by the user, select an API that can execute or provide information, and extract the key information in the instruction for the API to search for relevant information. The returned results include, API name: api, key information: arguments. Second step: If the API is unclear, recommend relevant labeled instructions to the user. Knowledge subgraph:".
[0168] Furthermore, the current prompt message determined according to the prompt message template and the target vehicle function knowledge information may include: "Suppose you are an intelligent voice assistant. First step, according to the instruction provided by the user, select an API that can execute or provide information, and extract the key information in the instruction for the API to search for relevant information. The returned results include, API name: api, key information: arguments. Second step: If the API is unclear, recommend relevant labeled instructions to the user. Knowledge subgraph: air conditioner -> fragrance -> [Fragrance A, Fragrance B].".
[0169] Then, the server can input the current prompt message and the current voice request into the large language model together, so that the large language model predicts vehicle control instructions and target guidance information.
[0170] Furthermore, due to the semantic defect in the current voice request ("adjust fragrance"), the large language model fails to determine the corresponding vehicle control instruction based on the current voice request ("adjust fragrance") and "First step, according to the instruction provided by the user, select an API that can execute or provide information, and extract the key information in the instruction for the API to search for relevant information. The returned results include, API name: api, key information: arguments" in the current prompt message.
[0171] Furthermore, the large language model can predict the target guidance information for guiding the user to adjust the current voice request based on the current voice request ("adjust fragrance") and "Second step: If the API is unclear, recommend relevant labeled instructions to the user. Knowledge subgraph: air conditioner -> fragrance -> [Fragrance A, Fragrance B]" in the current prompt message, such as: "Sorry, I didn't understand what you meant. If you want to turn on the fragrance, you can say: Turn on Fragrance A, Turn on Fragrance B".
[0172] It can be understood that in the solution of determining the guiding information only through a neural network model (such as a large language model) without considering the vehicle function knowledge graph and the target vehicle function knowledge, since the training process of the neural network model is usually carried out through pre-collected corpus data, or rather, through static knowledge. And when the neural network model is put into practice, once there are new and updated function points, if the neural network model is retrained to align the model with the new and updated function points, there will be a problem of high training cost.
[0173] If fine-tuning (such as supervised fine-tuning, SFT) is selected for the neural network model, after the fine-tuning training, the neural network model may still reason according to the previously learned knowledge, and then may reason out answers that do not match the "new and updated function points", that is, there is model hallucination.
[0174] Based on this, the embodiment of the present application is based on a pre-constructed external knowledge base, that is, the vehicle function knowledge graph, and recalls the target vehicle function knowledge information corresponding to the current voice request from the vehicle function knowledge graph, so that the pre-trained large language model can determine the target guiding information with the assistance of the target vehicle function knowledge information, ensuring the credibility and guiding role of the target guiding information.
[0175] At the same time, the embodiment of the present application can also update the voice request in the voice request data in the case of new and updated function points, so that the large language model can complete the reasoning work based on the external knowledge provided by the vehicle function knowledge graph, thus avoiding the situation that the large language model needs to be updated immediately in the case of new and updated function points. The performance of the large language model can be aligned with the new knowledge (that is, the new and updated function points) through the voice request data, and then the user needs can be met in real time and efficiently.
[0176] Moreover, compared with the method of setting up an external database and storing and maintaining "standard voice requests corresponding to vehicle control instructions" in the external database, so as to input the standard voice requests similar to the current voice request in the external database into the large language model when predicting the vehicle control instructions and target guidance information of the current voice request, this method using the external database usually restricts the "standard voice requests similar to the current voice request" to avoid searching for too many "standard voice requests similar to the current voice request", which may lead to noise in the input information of the large language model. Furthermore, due to the limited number of standard voice requests, some functional points related to the current voice request may not be input into the large language model.
[0177] In addition, in this method using the external database, the "standard voice requests similar to the current voice request" are usually determined based on similarity retrieval, and there may be errors in similarity retrieval. For example, when the current voice request is "adjust the air conditioner", both "turn on the air conditioner" and "adjust the seat" are similar to "adjust the air conditioner", which makes it difficult for the large language model to accurately complete the prediction of the vehicle control instructions and target guidance information corresponding to "adjust the air conditioner" through "turn on the air conditioner" and "adjust the seat".
[0178] Compared with this method using the external database, the implementation mode of this application is based on the vehicle function knowledge graph and vehicle function knowledge recall, and can directly input the obtained target vehicle function knowledge (or knowledge sub-graph) that matches into the large language model, ensuring the conciseness, integrity and accuracy of the external knowledge input into the large language model, and avoiding the situation of too long input information input into the large language model.
[0179] In addition, it can also be understood that when there are updates or additions to functional points, the vehicle function knowledge graph can be updated and maintained to achieve the output control of the large language model, thereby reducing the training and maintenance costs of the large language model.
[0180] The implementation mode of this application also provides a computer-readable storage medium storing a computer program, which, when executed by one or more processors, implements the above-mentioned voice interaction method.
[0181] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0182] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment or portion of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0183] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A voice interaction method, characterized in that: include: Receive the current voice request forwarded by the vehicle; Determine target vehicle function knowledge information according to the current voice request and a pre-built vehicle function knowledge graph; In the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, target guidance information for guiding a user to adjust the current voice request is determined according to the large language model, the current voice request and the target vehicle function knowledge information, wherein the large language model is pre-trained and can determine guidance information according to the voice request and the vehicle function knowledge information, and in the case where a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, the current language request corresponds to a text with semantic defects; Feedback the target guidance information to guide the user to complete the voice interaction; In the case that a vehicle control instruction corresponding to the current voice request cannot be determined according to the current voice request, determining target guidance information for guiding the user to adjust the current voice request according to the large language model, the current voice request and the target vehicle function knowledge information, including: Determine current prompt information according to the target vehicle function knowledge information and a pre-configured prompt information template, wherein the prompt information template conforms to a natural language understanding method, and the current prompt information includes first sub-prompt information and second sub-prompt information; When a vehicle control instruction corresponding to the current voice request cannot be determined based on the current voice request, the large language model and the first sub-prompt information, the target guidance information is determined based on the current voice request, the large language model and the second sub-prompt information.
2. The method according to claim 1, characterized in that The determining the target vehicle function knowledge information according to the current voice request and the pre-built vehicle function knowledge graph includes: According to the current voice request, the vehicle function knowledge graph is searched to determine the target vehicle function knowledge information that is semantically related to the current voice request.
3. The method according to claim 2, characterized in that The step of searching the vehicle function knowledge graph according to the current voice request to determine the target vehicle function knowledge information semantically related to the current voice request includes: Determining key information of the current voice request according to the current voice request; The vehicle function knowledge graph is searched according to the key information to determine the target vehicle function knowledge information.
4. The method according to claim 3, characterized in that The vehicle function knowledge graph is constructed according to the predetermined vehicle equipment and the functions that the vehicle equipment can perform, and the vehicle function knowledge graph is searched according to the key information to determine the target vehicle function knowledge information, including: The vehicle function knowledge graph is searched according to the key information, and the target device and target function corresponding to the key information are determined to determine the target vehicle function knowledge information.
5. The method according to claim 3, characterized in that: The vehicle function knowledge graph includes nodes at multiple levels, and searching the vehicle function knowledge graph according to the key information to determine the target vehicle function knowledge information includes: The vehicle function knowledge graph is searched according to the key information to obtain nodes of each level corresponding to the key information, so as to determine the target vehicle function knowledge information.
6. The method according to claim 3, characterized in that The determining, according to the current voice request, key information of the current voice request includes: The key information of the current voice request is determined according to the current voice request and a pre-trained named entity extraction model.
7. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Fuzzy semantic guiding method and device based on vehicle
CN116861919A
Voice interaction method, server and computer readable storage medium
CN117373456A
Voice interaction method, server and computer readable storage medium
CN117558277A