Voice interaction method, server and computer readable storage medium

By introducing a search mechanism for similar voice requests and application interface lists into the large language model, the data volume and inference pressure of the large language model in complex scenarios is solved, and the model's reasoning ability and user experience are improved.

CN120375818APending Publication Date: 2025-07-25SHANGHAI XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430658.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the application scenarios of complex and diversified functions, the data volume and inference pressure increase, resulting in increased usage costs and decreased recognition accuracy, and poor user experience.

Method used

By determining the list of similar voice requests and similar application interfaces, using the RAG model for encoding processing and vector retrieval, the data processing volume and inference burden of large language models are reduced, and the model structure is not required when adding functional interfaces.

Benefits of technology

It improves the reasoning ability of large language models, accurately understands user needs, reduces the pressure of model iteration and calculation frequency, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375818A_ABST
    Figure CN120375818A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method, a server and a computer readable storage medium. The method comprises the following steps: determining a similar voice request list according to received voice requests; then, according to the voice request, determining a similar application interface list; and then, according to the voice request, the similar voice request list and the similar application interface list, determining a natural language processing result. And finally, completing voice interaction according to a natural language processing result. Therefore, through the determined similar voice request list and the similar application interface list, the big language model can improve the model reasoning ability, so that the voice requests are subjected to reasoning efficiently, the user requirements are understood accurately, and the user experience is improved. Moreover, by pre-determining the similar voice request list and the similar application interface list, the data volume needing to be processed by the large language model can be reduced, and the reasoning burden of the large language model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction, and particularly to a voice interaction method, a server, and a computer-readable storage medium. Background Art

[0002] In the related art, when a user issues a voice request, the large language model can directly output a natural language processing result that meets the user's needs according to the text recognition result of the voice request. However, the complexity of the application scenarios and the diversification of the functions of the large language model have increased the data volume and inference pressure of the large language model, resulting in an increase in the usage cost of the large language model. Summary of the Invention

[0003] This application provides a voice interaction method, a server, and a computer-readable storage medium.

[0004] An embodiment of this application provides a voice interaction method, which includes:

[0005] Determine a list of similar voice requests according to the received voice request;

[0006] Determine a list of similar application interfaces according to the voice request;

[0007] Determine a natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces;

[0008] Complete the voice interaction according to the natural language processing result.

[0009] In this way, the server determines a list of similar voice requests according to the received voice request. Then, the server determines a list of similar application interfaces according to the voice request. Next, the server determines a natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces. Finally, the server completes the voice interaction according to the natural language processing result. In this way, through the determined list of similar voice requests and the list of similar application interfaces, the large language model can improve the model inference ability, and then efficiently infer the voice request, accurately understand the user's needs, thereby improving the user experience. Moreover, by pre-determining the list of similar voice requests and the list of similar application interfaces, the data volume that the large language model needs to process can be reduced, and the inference burden of the large language model can be reduced. In addition, when enriching the model function diversity through update and iteration, there is no need to update and iterate the large language model, reducing the model iteration pressure.

[0010] In some embodiments, the determining a list of similar voice requests according to the received voice request includes:

[0011] Based on a preset model, encode the voice request to determine a target vector corresponding to the voice request;

[0012] According to the target vector, determine the list of similar voice requests from a pre-constructed voice request database, where the voice request database includes candidate voice requests, first candidate application interfaces corresponding to the candidate voice requests, and first application interface parameters of the first candidate application interfaces that have been filled.

[0013] In this way, the server encodes the voice request based on a preset model to determine a target vector corresponding to the voice request. Then, the server determines the list of similar voice requests from a pre-constructed voice request database according to the target vector, where the voice request database includes candidate voice requests, first candidate application interfaces corresponding to the candidate voice requests, and first application interface parameters of the first candidate application interfaces that have been filled. In this way, through encoding processing and vector retrieval, a list of similar voice requests related to the voice request can be efficiently retrieved from the voice request database for subsequent natural language processing. Moreover, when adding a new large language model function interface, only the candidate voice requests and related parameter information need to be added to the voice request database, without modifying the structure of the large language model.

[0014] In some embodiments, the determining the list of similar application interfaces according to the voice request includes:

[0015] According to the target vector, determine the list of similar application interfaces from a pre-constructed application interface database, where the application interface database includes second candidate application interfaces, second application interface parameters to be filled corresponding to the second candidate application interfaces, application interface descriptions of the second candidate application interfaces, and interface parameter descriptions of the second application interface parameters.

[0016] In this way, the server determines the list of similar application interfaces from a pre-constructed application interface database according to the target vector, where the application interface database includes second candidate application interfaces, second application interface parameters to be filled corresponding to the second candidate application interfaces, application interface descriptions of the second candidate application interfaces, and interface parameter descriptions of the second application interface parameters. In this way, through encoding processing and vector retrieval, a list of similar application interfaces related to the voice request can be efficiently retrieved from the application interface database for subsequent natural language processing. Moreover, when adding a new large language model function interface, only the second candidate application interfaces and related parameter information need to be added to the application interface database, without modifying the structure of the large language model.

[0017] In some embodiments, determining the natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces includes:

[0018] Based on a preset prompt template, splice the voice request, the list of similar voice requests, and the list of similar application interfaces to determine a target prompt;

[0019] Based on a preset large language model, determine the natural language processing result according to the target prompt.

[0020] In this way, based on the preset prompt template, the server splices the voice request, the list of similar voice requests, and the list of similar application interfaces to determine the target prompt. Then, based on the preset large language model, the server determines the natural language processing result according to the target prompt. In this way, by splicing the user's current voice request, the list of similar voice requests, and the list of similar application interfaces through a preset template, a prompt containing multi-dimensional information is formed, providing a more complete decision-making basis for the large language model.

[0021] In some embodiments, the list of similar voice requests includes a similar voice request, a first similar application interface corresponding to the similar voice request, and first similar application interface parameters of the filled first similar application interface. The determining the natural language processing result based on the preset large language model and according to the target prompt includes:

[0022] Based on the preset large language model, determine the similarity degree between the voice request and the similar voice request according to the target prompt, where the similarity degree includes a first similarity degree, a second similarity degree, and a third similarity degree, the first similarity degree is greater than the second similarity degree, and the second similarity degree is greater than the third similarity degree;

[0023] Determine the natural language processing result according to the similarity degree and the target prompt.

[0024] In this way, based on the preset large language model, the server determines the similarity degree between the voice request and the similar voice request according to the target prompt, where the similarity degree includes a first similarity degree, a second similarity degree, and a third similarity degree, the first similarity degree is greater than the second similarity degree, and the second similarity degree is greater than the third similarity degree. Then, the server determines the natural language processing result according to the similarity degree and the target prompt. In this way, through multi-dimensional comparison of the voice request and the similar voice request by the preset large language model, the similarity degree between the voice request and the similar voice request is determined, and based on different similarity degrees, an appropriate processing method is selected to determine the natural language processing result.

[0025] In some embodiments, determining the natural language processing result according to the similarity degree and the target prompt word includes:

[0026] When the similarity degree is the first similarity degree, determining the first similarity application interface and the first similarity application interface parameters as the natural language processing result.

[0027] In this way, when the similarity degree is the first similarity degree, the server determines the first similarity application interface and the first similarity application interface parameters as the natural language processing result. Thus, when the semantic matching degree between the voice request and the historical similar voice request reaches the first similarity degree, the first similarity application interface and the first similarity application interface parameters are directly reused, bypassing the complex reasoning process of the large language model, achieving fast response, and thus improving the user experience. Moreover, by reusing the first similarity application interface and the first similarity application interface parameters, the calculation frequency of the large language model can be significantly reduced, and the overall throughput can be improved.

[0028] In some embodiments, determining the natural language processing result according to the similarity degree and the target prompt word includes:

[0029] When the similarity degree is the second similarity degree, determining a first target parameter value of the first similarity application interface parameters according to the voice request;

[0030] Performing parameter filling processing on the first similarity application interface parameters according to the first target parameter value;

[0031] Determining the natural language processing result according to the first similarity application interface and the result of the parameter filling processing.

[0032] In this way, when the similarity degree is the second similarity degree, the server determines a first target parameter value of the first similarity application interface parameters according to the voice request. Then, the server performs parameter filling processing on the first similarity application interface parameters according to the first target parameter value. Finally, the server determines the natural language processing result according to the first similarity application interface and the result of the parameter filling processing. Thus, when the similarity degree is the second similarity degree, the first similarity application interface parameters are reused, and the first target parameter value to be filled into the first similarity application interface parameters is inferred to determine the natural language processing result, thereby reducing the inference task amount of the large language model.

[0033] In some embodiments, the similarity application interface list includes a second similarity application interface, second similarity application interface parameters to be filled corresponding to the second similarity application interface, a similarity application interface description, and a similarity interface parameter description. Determining the natural language processing result according to the similarity degree and the target prompt word includes:

[0034] When the similarity degree is the third similarity degree, determine a second target parameter value of the second similar application interface parameter according to the voice request;

[0035] Perform parameter filling processing on the second similar application interface parameter according to the second target parameter value;

[0036] Determine the natural language processing result according to the second similar application interface and the result of the parameter filling processing;

[0037] In this way, when the similarity degree is the third similarity degree, according to the voice request, determine the second target parameter value of the second similar application interface parameter. Then, the server performs parameter filling processing on the second similar application interface parameter according to the second target parameter value. Finally, the server determines the natural language processing result according to the second similar application interface and the result of the parameter filling processing. In this way, when the similarity degree is the third similarity degree, based on the pre-determined list of similar application interfaces, the voice request is inferred, reducing the inference task volume of the large language model and improving the response time of the large language model, thereby improving the user experience.

[0038] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.

[0039] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0040] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:

[0042] Figure 1 is one of the flow diagrams of the voice interaction method of some embodiments of the present application;

[0043] Figure 2 is another flow diagram of the voice interaction method of some embodiments of the present application;

[0044] Figure 3 is yet another flow diagram of the voice interaction method of some embodiments of the present application;

[0045] Figure 4 It is the fourth flow schematic diagram of the voice interaction method of some embodiments of the present application;

[0046] Figure 5 It is the fifth flow schematic diagram of the voice interaction method of some embodiments of the present application;

[0047] Figure 6 It is the sixth flow schematic diagram of the voice interaction method of some embodiments of the present application;

[0048] Figure 7 It is the seventh flow schematic diagram of the voice interaction method of some embodiments of the present application;

[0049] Figure 8 It is the eighth flow schematic diagram of the voice interaction method of some embodiments of the present application;

[0050] Figure 9 It is the flow schematic diagram of the voice request processing of some embodiments of the present application. Specific Embodiments

[0051] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.

[0052] When a user issues a voice request, traditional natural language processing solutions rely on large language models to directly output natural language processing results that meet the user's needs based on the text recognition results of the voice request. However, with the increasingly complex application scenarios and diverse functions of large language models, such as in intelligent vehicle cockpits, various instructions need to be recognized and processed, such as adjusting the seat, controlling the air conditioner, and playing music, resulting in a sharp increase in the amount of data required by the large language model and an increase in the inference pressure. This not only significantly increases the training and deployment costs of the large language model but also makes it difficult to ensure the recognition accuracy and efficiency of the large language model in different scenarios, resulting in a poor user experience.

[0053] Based on the above problems, please refer to Figure 1 , some embodiments of the present application provide a voice interaction method, the method including:

[0054] 01: Determine a list of similar voice requests according to the received voice request;

[0055] 02: Determine a list of similar application interfaces according to the voice request;

[0056] 03: Determine the natural language processing result based on the voice request, the list of similar voice requests, and the list of similar application interfaces;

[0057] 04: Complete the voice interaction according to the natural language processing result.

[0058] The embodiment of the present application also provides a server, including a memory and a processor. The training method of the voice recognition model in the embodiment of the present application can be implemented by the server in the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is configured to determine a list of similar voice requests according to the received voice request. And determine a list of similar application interfaces according to the voice request. The processor is configured to determine the natural language processing result based on the voice request, the list of similar voice requests, and the list of similar application interfaces. And complete the voice interaction according to the natural language processing result.

[0059] The embodiment of the present application also provides a voice interaction device. The training method of the voice recognition model in the embodiment of the present application can be implemented by the voice interaction device in the embodiment of the present application. Specifically, the voice interaction device includes a determination module and an interaction module. The determination module is configured to determine a list of similar voice requests according to the received voice request. And determine a list of similar application interfaces according to the voice request. And determine the natural language processing result based on the voice request, the list of similar voice requests, and the list of similar application interfaces. The interaction module is configured to complete the voice interaction according to the natural language processing result.

[0060] Specifically, the voice request refers to an instruction issued by the user in a voice manner. For example, "turn on the air conditioner", "navigate to the company", "play music", and "adjust the seat position forward by 30%", etc. It should be noted that the voice request needs to be subjected to voice recognition processing to convert the voice signal emitted by the user into a text form, that is, to recognize the words, phrases, and sentence structures in the voice and convert them into text for subsequent processing.

[0061] The similar voice request list refers to other retrieved voice requests that are semantically similar to the user's voice request, including similar voice requests, first similar application interfaces corresponding to the similar voice requests, and first similar application interface parameters of the first similar application interfaces that have been filled. For example, if the user's voice request is "adjust the seat position forward 30 percent", then the similar voice request list may be [Similar Voice Request A: "Adjust the seat position all the way forward", first similar application interface a1: "ControlSet", first similar application interface parameter a2: "'function':'Position', 'set_type':'Adjust to', 'device':'Seat', 'value':'Front'"; Similar Voice Request B: "Adjust the seat back a little", first similar application interface b1: "ControlSet", first similar application interface parameter b2: "'function':'Position', 'set_type':'Adjust back', 'device':'Seat', 'value':'A little'"; Similar Voice Request C: "Change the sound to private mode", first similar application interface c1: "VoiceOpen", first similar application interface parameter c2: "'target_function':'Private mode', 'device':'Headrest speaker'"].

[0062] The list of similar application interfaces refers to the executable application interfaces retrieved that are related to the user's voice request, including the second similar application interface, the second similar application interface parameters to be filled corresponding to the second similar application interface, the description of the similar application interface, and the description of the similar interface parameters. For example, if the user's voice request is "Adjust the seat position forward by 30%", then the list of similar application interfaces may include ["Second similar application interface D: ControlOpen", "Second similar application interface parameter d1: device, function, target_function, and seat_position", "Description of the similar application interface d2: Open relevant in-vehicle devices or functions, and can control multiple devices and various functions in the vehicle", and "Description of the similar interface parameter d3: device (string, optional) is used to specify the in-vehicle device and is only provided when a specific function within the device needs to be specified; function (string, optional) is used to specify the function related to the in-vehicle device and is only provided when a specific function option within the device needs to be specified; target_function (string, optional) is used to specify different modes or different functions of the function related to the in-vehicle device and is only provided when a specific function option within the device needs to be specified; seat_position (enumeration value, optional) is used to specify the seat position in the vehicle, and the value can be optionally one of ['Driver's seat', 'Passenger seat', 'Front row', 'Left side of the second row', 'Right side of the second row', 'Second row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row']"].

[0063] It should be noted that according to the voice request, determining the list of similar voice requests and the list of similar application interfaces are both implemented by the Retrieval-Augmented Generation (RAG) model.

[0064] The natural language processing result refers to the structured output generated after the large language model parses and reasons the user's voice request, including the application program interface and the filled application program interface parameters.

[0065] After receiving the user's voice request, the server will input the voice request into the RAG model. Then, the RAG model can determine the list of similar voice requests and the list of similar application interfaces according to the voice request. Subsequently, the large language model determines the natural language processing result based on the voice request, the list of similar voice requests, and the list of similar application interfaces. Finally, the server calls the corresponding application program interface according to the natural language processing result to execute relevant program instructions.

[0066] In summary, in the voice interaction method and server provided by the embodiments of the present application, the server determines a list of similar voice requests based on the received voice request. Next, the server determines a list of similar application interfaces based on the voice request. Then, the server determines the natural language processing result based on the voice request, the list of similar voice requests, and the list of similar application interfaces. Finally, the server completes the voice interaction based on the natural language processing result. In this way, through the determined list of similar voice requests and the list of similar application interfaces, the large language model can improve the model inference ability, and then efficiently infer the voice request, accurately understand the user's needs, thereby improving the user experience. Moreover, by pre-determining the list of similar voice requests and the list of similar application interfaces, the amount of data that the large language model needs to process can be reduced, and the inference burden of the large language model can be reduced. In addition, when enriching the diversity of model functions through update and iteration, there is no need to update and iterate the large language model, reducing the model iteration pressure.

[0067] Please refer to Figure 2 , in some embodiments, step 01 (determining a list of similar voice requests based on the received voice request) includes:

[0068] 011: Based on a preset model, encode the voice request to determine a target vector corresponding to the voice request;

[0069] 012: According to the target vector, determine a list of similar voice requests from the pre-constructed voice request database.

[0070] In some embodiments, the determining module is further configured to encode the voice request based on a preset model to determine a target vector corresponding to the voice request. And according to the target vector, determine a list of similar voice requests from the pre-constructed voice request database.

[0071] In some embodiments, the processor is further configured to encode the voice request based on a preset model to determine a target vector corresponding to the voice request. And according to the target vector, determine a list of similar voice requests from the pre-constructed voice request database.

[0072] Specifically, the preselected model refers to the RAG model, and the RAG model can retrieve a list of similar voice requests similar to the user's voice request from the pre-set similar voice request library.

[0073] Encoding processing refers to converting the target voice request training data into vectors so that it can be understood and processed by machine learning models. In some embodiments, the encoding processing of voice requests is implemented based on the BGE model. The BGE model refers to an embedding model based on a graph structure that can convert user voice requests into vector representations and retrieve application interface description vectors semantically similar to the user instructions from the application interface database. In addition, methods such as word embedding and sequence encoding can also be used to convert text into high-dimensional vectors while preserving the semantic information of the text.

[0074] The target vector refers to the vector representation corresponding to the voice request after encoding processing. It can be used as the basis for retrieving a list of similar voice requests, calculating the similarity with candidate voice requests in a pre-constructed voice request database, selecting similar candidate voice requests, and thus determining the list of similar voice requests. In this way, retrieving based on vector similarity can accurately match candidate voice requests related to the voice request and avoid mis-matching.

[0075] Candidate voice requests refer to the voice requests pre-stored in the voice request database, such as "Adjust the seat position all the way forward" and "Adjust the seat back a little". Candidate voice requests serve as the candidate set for the RAG model to retrieve similar voice requests and can help the model find voice requests semantically similar to the user instructions.

[0076] The first candidate application interface refers to the application interface corresponding to the candidate voice request, that is, the application interface that implements the user requirements expressed by the corresponding candidate voice request, such as "ControlSet" and "VoiceOpen".

[0077] The first application interface parameter refers to the filled application interface parameter corresponding to the first candidate application interface, such as "{'function': 'position','set_type': 'adjust to', 'device':'seat', 'value': 'front most'}.

[0078] The voice request database refers to a database that stores candidate voice requests, first candidate application interfaces corresponding to the candidate voice requests, and first application interface parameters of the first candidate application interfaces that have been filled in, and can provide a data basis for retrieving similar voice requests. The voice request database may include data information such as [similar voice request A: "Adjust the seat position to the farthest forward", first similar application interface a1: "ControlSet", first similar application interface parameter a2: "'function':'Position', 'set_type':'Adjust to', 'device':'Seat', 'value':'Front'"], [similar voice request B: "Adjust the seat a little bit back", first similar application interface b1: "ControlSet", first similar application interface parameter b2: "'function':'Position', 'set_type':'Adjust back', 'device':'Seat', 'value':'A little bit'"] and [similar voice request C: "Change the sound to private mode", first similar application interface c1: "VoiceOpen", first similar application interface parameter c2: "'target_function':'Private mode', 'device':'Headrest speaker'"], etc. It should be noted that the construction of the voice request database is based on the training data of the large language model.

[0079] First, the user's voice request is encoded to determine the target vector, and then the target vector is similarly calculated with the candidate voice requests in the voice request database to obtain a list of similar voice requests.

[0080] In this way, the server encodes the voice request based on the preset model and determines the target vector corresponding to the voice request. Then, the server determines a list of similar voice requests from a pre-built voice request database based on the target vector, wherein the voice request database includes a candidate voice request, a first candidate application interface corresponding to the candidate voice request, and a first application interface parameter of the first candidate application interface that has been filled. In this way, through encoding processing and vector retrieval, a list of similar voice requests related to the voice request can be efficiently retrieved from the voice request database for subsequent natural language processing. Moreover, when adding a large language model function interface, it is only necessary to add the candidate voice request and related parameter information to the voice request database without modifying the large language model structure.

[0081] See also Figure 3 In some implementations, step 02 (determining a similar application interface list according to the voice request) includes:

[0082] 021: According to the target vector, determine a list of similar application interfaces from a pre-built application interface database.

[0083] In some embodiments, the determination module is further configured to determine a list of similar application interfaces from a pre-constructed application interface database according to the target vector.

[0084] In some embodiments, the processor is further configured to determine a list of similar application interfaces from a pre-constructed application interface database according to the target vector.

[0085] Specifically, the application interface database refers to a database that stores application interfaces and their corresponding parameters, providing a data basis for retrieving similar application interfaces, including the second candidate application interface, the second application interface parameters to be filled corresponding to the second candidate application interface, the application interface description of the second candidate application interface, and the interface parameter description of the second application interface parameters. The data information that the application interface database may include is ["Second similar application interface D: ControlOpen", "Second similar application interface parameter d1: device, function, target_function, and seat_position", "Similar application interface description d2: Open relevant in-vehicle devices or functions, can control multiple devices and various functions in the vehicle", "Similar interface parameter description d3: device (string, optional) is used to specify the in-vehicle device, only provided when specific functions within the device need to be specified; function (string, optional) is used to specify the functions related to the in-vehicle device, only provided when specific function options within the device need to be specified; target_function (string, optional) is used to specify different modes or different functions of the functions related to the in-vehicle device, only provided when specific function options within the device need to be specified; seat_position (enumeration value, optional) is used to specify the seat position in the vehicle, and the value can be selected from one of ['driver', 'passenger', 'front row', 'left side of the second row', 'right side of the second row', 'second row', 'all', 'left side of the third row', 'right side of the third row', 'third row']"], ["Second similar application interface E: ControlClose", "Second similar application interface parameter e1: device, function, target_function, and seat_position", "Similar application interface description e2: Open relevant in-vehicle devices or functions, can control multiple devices and various functions in the vehicle", "Similar interface parameter description e3: device (string, optional) is used to specify the in-vehicle device, only provided when specific functions within the device need to be specified; function (string, optional) is used to specify the functions related to the in-vehicle device, only provided when specific function options within the device need to be specified; target_function (string, optional) is used to specify different modes or different functions of the functions related to the in-vehicle device, only provided when specific function options within the device need to be specified;seat_position (enumeration value, optional) is used to specify the seat position in the vehicle, and the value can be selected from one of ['Driver's seat', 'Passenger seat', 'Front row', 'Left side of the second row', 'Right side of the second row', 'Second row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row'] and ["Second similar application interface F: ControlSet", "Second similar application interface parameter f1: device, function, target_function, set_type, and seat_position", "Similar application interface description f2: Adjust relevant devices or functions in the vehicle, and multiple devices and various functions in the vehicle can be adjusted", "Similar interface parameter description f3: device (string, optional) is used to specify the device in the vehicle and is only provided when a specific function within the device needs to be specified; function (string, optional) is used to specify the function related to the device in the vehicle and is only provided when a specific function option within the device needs to be specified; target_function (string, optional) is used to specify different modes or different functions of the function related to the device in the vehicle and is only provided when a specific function option within the device needs to be specified; set_type (enumeration value, optional) specifies the type to be adjusted, and the value can be selected from one of ['Adjust up', 'Adjust down', 'Adjust higher', 'Adjust lower', 'Adjust left', 'Adjust right', 'Adjust to', 'Adjust larger', 'Adjust smaller']; seat_position (enumeration value, optional) is used to specify the seat position in the vehicle, and the value can be selected from one of ['Driver's seat', 'Passenger seat', 'Front row', 'Left side of the second row', 'Right side of the second row', 'Second row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row']. It should be noted that the construction of the application interface database is based on the training data of the large language model.;

[0086] By retrieving similar voice requests and similar application interfaces, the RAG model can help the large language model reduce the inference pressure, enabling the large language model to more efficiently handle natural language understanding tasks in complex scenarios. Moreover, the introduction of the RAG model can reduce the pressure of iterative training of the large language model, thereby reducing the training cost and human input. In addition, the RAG model can assist the large language model in handling unseen tasks and enhancing the generalization ability of the large language model.

[0087] Calculate the similarity between the target vector and the application interface description of the second candidate application interface in the pre-constructed application interface database to obtain a list of similar application interfaces.

[0088] In this way, the server determines a list of similar application interfaces from the pre-constructed application interface database according to the target vector. The application interface database includes second candidate application interfaces, second application interface parameters to be filled corresponding to the second candidate application interfaces, application interface descriptions of the second candidate application interfaces, and interface parameter descriptions of the second application interface parameters. In this way, through encoding processing and vector retrieval, a list of similar application interfaces related to the voice request can be efficiently retrieved from the application interface database for subsequent natural language processing. Moreover, when adding new large language model function interfaces, only the second candidate application interfaces and related parameter information need to be added to the application interface database, without modifying the structure of the large language model.

[0089] Please refer to Figure 4 , in some embodiments, step 03 (determining the natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces) includes:

[0090] 031: Based on a preset prompt template, splice the voice request, the list of similar voice requests, and the list of similar application interfaces to determine the target prompt;

[0091] 032: Based on a preset large language model, determine the natural language processing result according to the target prompt.

[0092] In some embodiments, the determination module is further configured to splice the voice request, the list of similar voice requests, and the list of similar application interfaces based on a preset prompt template to determine the target prompt. And based on a preset large language model, determine the natural language processing result according to the target prompt.

[0093] In some embodiments, the processor is also configured to splice the voice request, the list of similar voice requests, and the list of similar application interfaces based on a preset prompt template to determine the target prompt. And based on a preset large language model, determine the natural language processing result according to the target prompt.

[0094] Specifically, the preset prompt template refers to a pre-defined text template in natural language processing tasks, which is used to guide the large language model to generate or understand text. In some embodiments, the preset prompt template may be "You are a vehicle-mounted voice assistant (Assistant), responsible for listening to all conversations between users (User) in the vehicle. Please infer the current user's intention based on the conversation context and give the corresponding executable API and ARGUMENTS.

[0095] The following will give the API and similar queries related to the user instruction. Please:

[0096] 1. For similar queries, directly provide the results based on the similar queries.

[0097] 2. For partially similar queries, provide the results based on the similar queries and the similar APIs.

[0098] 3. For no similar queries, use the similar APIs to fill in the corresponding slots and provide the results.

[0099] The output format of the Assistant is:

[0100] RESULT: {'API': 'api1', 'ARGUMENTS': [{'arg1': 'value1', 'arg2': 'value2',... 'argn': 'valuen'}]}

[0101] <List of similar voice requests>

[0102] <List of similar application interfaces>

[0103] <Voice request>".

[0104] The pre-trained large language model refers to the large language model that has been trained before performing natural language processing tasks. In the context of the embodiments of this application, the large language model is used to refer to the pre-trained large language model.

[0105] Concatenate the voice request, the list of similar voice requests, and the list of similar application interfaces according to the preset prompt template to generate a complete text, which is the target prompt. For example, if the user's voice request is "Adjust the seat position forward by 30%", then the target prompt may be "You are a vehicle-mounted voice assistant (Assistant), responsible for listening to all conversations between users (User) in the vehicle. Please infer the current user's intention based on the conversation context and give the corresponding executable API and ARGUMENTS.

[0106] The following will give the APIs and similar queries related to the user's instructions. Please:

[0107] 1. For similar queries, directly provide the results based on the similar queries.

[0108] 2. For partially similar queries, provide the results based on the similar queries and the similar APIs.

[0109] 3. For no similar queries, use the similar APIs to fill in the corresponding slots and provide the results.

[0110] The output format of the Assistant is:

[0111] RESULT: {'API': 'api1', 'ARGUMENTS': [{'arg1': 'value1', 'arg2': 'value2',... 'argn': 'valu en'}}]

[0112] List of similar application interfaces:

[0113] ControlOpen

[0114] ControlClose

[0115] ControlSet

[0116]

ControlOpen

[0117] Description: It can open the in-vehicle related devices or functions and control multiple devices and various functions in the vehicle.

[0118] ARGUMENTS:

[0119] "device": (string, optional) Specifies the in-vehicle device and is provided only when a specific function within the device needs to be specified.

[0120] "function": (string, optional) Specifies the function related to the in-vehicle device and is provided only when a specific function option within the device needs to be specified.

[0121] "target_function": (string, optional) Specifies different modes or different functions of the in-vehicle device related functions and is provided only when a specific function option within the device needs to be specified

[0122] "seat_position": (enumeration value, optional) Specifies the seat position in the vehicle. The value can be one of ['Driver', 'Passenger', 'Front Row', 'Left Side of Second Row', 'Right Side of Second Row', 'Second Row', 'All', 'Left Side of Third Row', 'Right Side of Third Row', 'Third Row'].

[0123]

ControlClose

[0124] Description: It can open the in-vehicle related devices or functions and control multiple devices and various functions in the vehicle.

[0125] ARGUMENTS:

[0126] "device": (string, optional) Specifies the in-vehicle device and is provided only when a specific function within the device needs to be specified.

[0127] "function": (string, optional) Specifies the function related to the in-vehicle device, provided only when specific function options within the device need to be specified.

[0128] "target_function": (string, optional) Specifies different modes or functions of the function related to the in-vehicle device, provided only when specific function options within the device need to be specified

[0129] "seat_position": (enumeration value, optional) Specifies the seat position in the vehicle, and the value can be one of ['Driver', 'Passenger', 'Front Row', 'Left of Second Row', 'Right of Second Row', 'Second Row', 'All', 'Left of Third Row', 'Right of Third Row', 'Third Row'].

[0130]

ControlSet

[0131] Description: Adjusts the in-vehicle related devices or functions, and can adjust multiple devices and various functions within the vehicle.

[0132] ARGUMENTS:

[0133] "device": (string, optional) Specifies the in-vehicle device, provided only when specific functions within the device need to be specified.

[0134] "function": (string, optional) Specifies the function related to the in-vehicle device, provided only when specific function options within the device need to be specified.

[0135] "target_function": (string, optional) Specifies different modes or functions of the function related to the in-vehicle device, provided only when specific function options within the device need to be specified

[0136] "set_type": (enumeration value, optional) Specifies the type to be adjusted, and the value can be one of ['Adjust Up', 'Adjust Down', 'Adjust Higher', 'Adjust Lower', 'Adjust Left', 'Adjust Right', 'Adjust To', 'Adjust Larger', 'Adjust Smaller'].

[0137] "seat_position": (enumeration value, optional) Specifies the seat position in the vehicle, and the value can be one of ['Driver', 'Passenger', 'Front Row', 'Left of Second Row', 'Right of Second Row', 'Second Row', 'All', 'Left of Third Row', 'Right of Third Row', 'Third Row'].

[0138] Similar voice request list:

[0139] 1. Move the seat position forward to the end ControlSet{'function': 'position','set_type': 'adjust to', 'device':'seat', 'value': 'the frontmost'}

[0140] 2. Move the seat back a little ControlSet{'function': 'position','set_type': 'adjust back', 'device':'seat', 'value': 'a little'}

[0141] 3. Change the sound to private mode VoiceOpen{'target_function': 'private mode', 'device': 'headrest speaker'}

[0142] 4. Set the broadcast to private mode VoiceOpen{'target_function': 'private mode', 'device': 'headrest speaker'}

[0143] 5. Adjust the sound to driving enjoyment mode VoiceOpen{'target_function': 'driving enjoyment mode', 'device': 'headrest speaker'}

[0144] Voice request:

[0145] Adjust the seat position forward by thirty percent.

[0146] Next, based on a pre - set large - language model, natural language processing is performed according to the target prompt, and the natural language processing result is determined.

[0147] In this way, based on a pre - set prompt template, the server splices and processes the voice request, the list of similar voice requests, and the list of similar application interfaces to determine the target prompt. Next, based on the pre - set large - language model, the server determines the natural language processing result according to the target prompt. In this way, by splicing the user's current voice request, the list of similar voice requests, and the list of similar application interfaces through a pre - set template, a prompt containing multi - dimensional information is formed, providing a more complete decision - making basis for the large - language model.

[0148] Please refer to Figure 5 , in some embodiments, the list of similar voice requests includes similar voice requests, the first similar application interfaces corresponding to the similar voice requests, and the first similar application interface parameters of the first similar application interfaces that have been filled. Step 032 (Based on a pre - set large - language model, according to the target prompt, determine the natural language processing result) includes:

[0149] 0321: Based on a pre - set large - language model, according to the target prompt, determine the similarity degree between the voice request and the similar voice requests;

[0150] 0322: Determine the natural language processing result according to the similarity degree and the target prompt word.

[0151] In some embodiments, the determining module is further configured to determine the similarity degree between the voice request and the similar voice request based on a preset large language model according to the target prompt word, and determine the natural language processing result according to the similarity degree and the target prompt word.

[0152] In some embodiments, the processor is further configured to determine the similarity degree between the voice request and the similar voice request based on a preset large language model according to the target prompt word, and determine the natural language processing result according to the similarity degree and the target prompt word.

[0153] Specifically, the similarity degree refers to the semantic similarity between the user's voice request and the similar voice request. The large language model will calculate the similarity degree between the user's voice request and the similar voice request according to their contents, and classify it into three levels: the first similarity degree, the second similarity degree, and the third similarity degree. Among them, the first similarity degree is used to indicate that the user's voice request is exactly the same as or very similar to the similar voice request. The second similarity degree is used to indicate that there are some differences between the user's voice request and the similar voice request, but the overall intention is the same. The third similarity degree is used to indicate that there are large differences between the user's voice request and the similar voice request, and a list of similar application interfaces is required to assist the large language model in reasoning about the user's voice request. It should be noted that the similarity degree here is different from the similarity mentioned in determining the list of similar voice requests and the list of similar application interfaces according to the voice request. Here, it is the semantic similarity directly obtained by the large language model, and the previous similarity is the vector similarity obtained by the RAG model.

[0154] Based on a preset large language model, calculate the semantic similarity between the voice request and the similar voice request according to the target prompt word. Then, determine the final natural language processing result according to the similarity degree and the target prompt word.

[0155] In this way, based on a preset large language model, the server determines the similarity degree between the voice request and the similar voice request according to the target prompt word, where the similarity degree includes the first similarity degree, the second similarity degree, and the third similarity degree, and the first similarity degree is greater than the second similarity degree, and the second similarity degree is greater than the third similarity degree. Then, the server determines the natural language processing result according to the similarity degree and the target prompt word. In this way, through multi-dimensional comparison of the voice request and the similar voice request by the preset large language model, the similarity degree between the voice request and the similar voice request is determined, and based on different similarity degrees, an appropriate processing method is selected to determine the natural language processing result.

[0156] Please refer to Figure 6, in some embodiments, step 0322 (determine the natural language processing result according to the similarity degree and the target prompt word) includes:

[0157] 03221: When the similarity degree is the first similarity degree, determine the first similar application interface and the first similar application interface parameters as the natural language processing result.

[0158] In some embodiments, the determination module is further configured to, when the similarity degree is the first similarity degree, determine the first similar application interface and the first similar application interface parameters as the natural language processing result.

[0159] In some embodiments, the processor is further configured to, when the similarity degree is the first similarity degree, determine the first similar application interface and the first similar application interface parameters as the natural language processing result.

[0160] Specifically, when the semantic similarity between the voice request and the similar voice request judged by the large language model is the first similarity degree, the natural language processing result of the similar voice request can be directly used as the final natural language processing result, thereby simplifying the inference process of the large language model and improving the efficiency.

[0161] Thus, when the similarity degree is the first similarity degree, the server determines the first similar application interface and the first similar application interface parameters as the natural language processing result. In this way, when the semantic matching degree between the voice request and the historical similar voice request reaches the first similarity degree, the first similar application interface and the first similar application interface parameters are directly reused, bypassing the complex inference process of the large language model, achieving fast response, and thus enhancing the user experience. Moreover, by reusing the first similar application interface and the first similar application interface parameters, the calculation frequency of the large language model can be significantly reduced, and the overall throughput can be improved.

[0162] Please refer to Figure 7 , in some embodiments, step 0322 (determine the natural language processing result according to the similarity degree and the target prompt word) includes:

[0163] 03222: When the similarity degree is the second similarity degree, determine the first target parameter value of the first similar application interface parameter according to the voice request;

[0164] 03223: Perform parameter filling processing on the first similar application interface parameter according to the first target parameter value;

[0165] 03224: Determine the natural language processing result according to the first similar application interface and the result of the parameter filling processing.

[0166] In some embodiments, the determination module is further configured to, when the similarity level is the second similarity level, determine a first target parameter value of the first similar application interface parameter according to the voice request. And according to the first target parameter value, perform parameter filling processing on the first similar application interface parameter. And according to the first similar application interface and the result of the parameter filling processing, determine the natural language processing result.

[0167] In some embodiments, the processor is further configured to, when the similarity level is the second similarity level, determine a first target parameter value of the first similar application interface parameter according to the voice request. And according to the first target parameter value, perform parameter filling processing on the first similar application interface parameter. And according to the first similar application interface and the result of the parameter filling processing, determine the natural language processing result.

[0168] Specifically, when the semantic similarity between the voice request and the similar voice requests determined by the large language model is the second similarity level, the server determines a first target parameter value of the first similar application interface parameter according to the voice request. In some embodiments, the server may, according to the voice request, determine the most user - demand - compliant similar voice request from the list of similar voice requests, and then determine the first similar application interface parameter associated with the similar voice request as the first similar application interface parameter that needs to be processed for parameter filling. Thus, based on the first similar application interface parameter, perform slot recognition on the voice request to determine the first target parameter value, and perform parameter filling processing to determine the result of the parameter filling processing that most meets the user's needs.

[0169] In some embodiments, the server may also first perform slot recognition on the voice request to determine the first target parameter value, and then perform parameter filling processing on all the first similar application interface parameters in the list of similar voice requests to determine the result of the parameter filling processing that most meets the user's needs.

[0170] Then, according to the first similar application interface and the result of the parameter filling processing, determine the final natural language processing result.

[0171] Slot recognition refers to extracting specific information fragments from the user's input, which are usually called "slots". Slots are usually the key information necessary to complete a certain task or request, such as time, location, object, etc. Taking the user's voice request "What's the temperature tomorrow" as an example, the slot information obtained through slot recognition includes ["tomorrow" - Date], that is, the slot information includes the slot value and the slot type. Among them, "tomorrow" is the slot value, and Date is the slot type. Taking the user's voice request "Navigate to Address A" as an example, the slot information obtained through slot recognition is ["Address A" - Place], where "Address A" is the slot value and Place is the slot type.

[0172] Parameter filling processing refers to, on the basis of slot recognition, associating and mapping these slots with the parameters of the application programming interface. That is to say, parameter filling processing binds the slot values and the application interface parameters to ensure that the issued instructions can be correctly executed and the user's needs can be met.

[0173] In this way, when the similarity level is the second similarity level, the server determines the first target parameter value of the first similar application interface parameter according to the voice request. Then, the server performs parameter filling processing on the first similar application interface parameter according to the first target parameter value. Finally, the server determines the natural language processing result according to the first similar application interface and the result of the parameter filling processing. In this way, when the similarity level is the second similarity level, the first similar application interface parameter is reused, and the first target parameter value to be filled into the first similar application interface parameter is inferred to determine the natural language processing result, thereby reducing the inference task volume of the large language model.

[0174] Please refer to Figure 8 , in some embodiments, the list of similar application interfaces includes a second similar application interface, the second similar application interface parameters to be filled corresponding to the second similar application interface, the description of the similar application interface, and the description of the similar interface parameters. Step 0322 (determining the natural language processing result according to the similarity level and the target prompt word) includes:

[0175] 03225: When the similarity level is the third similarity level, determine the second target parameter value of the second similar application interface parameter according to the voice request;

[0176] 03226: Perform parameter filling processing on the second similar application interface parameter according to the second target parameter value;

[0177] 03227: Determine the natural language processing result according to the second similar application interface and the result of the parameter filling processing.

[0178] In some embodiments, the determination module is further configured to, when the similarity level is the third similarity level, determine a second target parameter value of the second similar application interface parameter according to the voice request. And perform parameter filling processing on the second similar application interface parameter according to the second target parameter value. And determine the natural language processing result according to the second similar application interface and the result of the parameter filling processing.

[0179] In some embodiments, the processor is further configured to, when the similarity level is the third similarity level, determine a second target parameter value of the second similar application interface parameter according to the voice request. And perform parameter filling processing on the second similar application interface parameter according to the second target parameter value. And determine the natural language processing result according to the second similar application interface and the result of the parameter filling processing.

[0180] Specifically, when the semantic similarity between the voice request and the similar voice request determined by the large language model is the third similarity level, the server determines a second target parameter value of the second similar application interface parameter according to the voice request. In some embodiments, the server may, according to the voice request, determine the most user - demand - compliant similar application interface from the list of similar application interfaces, and then determine the second similar application interface parameter associated with the similar application interface as the second similar application interface parameter that needs to be processed for parameter filling. Thus, based on the second similar application interface parameter, perform slot recognition on the voice request to determine the second target parameter value, and perform parameter filling processing to determine the result of the parameter filling processing that most meets the user's needs.

[0181] In some embodiments, the server may also first perform slot recognition on the voice request to determine the second target parameter value, and then perform parameter filling processing on all the second similar application interface parameters in the list of similar application interfaces to determine the result of the parameter filling processing that most meets the user's needs.

[0182] Then, determine the final natural language processing result according to the second similar application interface and the result of the parameter filling processing.

[0183] Please refer to Figure 9 , Figure 9 which is a schematic diagram of the voice request processing flow of the voice interaction method provided by the embodiments of this application. Among them, after receiving the user's voice request, retrieve according to the RAG model to determine the list of similar voice requests and the list of similar application interfaces. Subsequently, splice the voice request, the list of similar voice requests, and the list of similar application interfaces to determine the target prompt word. And input the target prompt word into the large language model, and the large language model performs reasoning to determine the final natural language processing result.

[0184] Thus, in the case where the similarity degree is the third similarity degree, according to the voice request, the second target parameter value of the second similar application interface parameter is determined. Then, the server performs parameter filling processing on the second similar application interface parameter according to the second target parameter value. Finally, the server determines the natural language processing result according to the second similar application interface and the result of the parameter filling processing. In this way, when the similarity degree is the third similarity degree, reasoning is performed on the voice request based on the pre-determined list of similar application interfaces, reducing the inference task volume of the large language model and improving the response time of the large language model, thereby enhancing the user experience.

[0185] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method as described above are implemented.

[0186] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution media, etc.

[0187] In the description of this specification, the descriptions with reference to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in combination with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0188] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment or part of executable instructions including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.

[0189] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A voice interaction method, characterized in that, The method includes: Determining a list of similar voice requests according to the received voice request; Determining a list of similar application interfaces according to the voice request; Determining a natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces; Completing the voice interaction according to the natural language processing result.

2. The method according to claim 1, wherein The determining a list of similar voice requests according to the received voice request includes: Performing encoding processing on the voice request based on a preset model to determine a target vector corresponding to the voice request; Determining the list of similar voice requests from a pre-constructed voice request database according to the target vector, where the voice request database includes candidate voice requests, first candidate application interfaces corresponding to the candidate voice requests, and first application interface parameters of the first candidate application interfaces that have been filled.

3. The method according to claim 2, wherein The determining a list of similar application interfaces according to the voice request includes: Determining the list of similar application interfaces from a pre-constructed application interface database according to the target vector, where the application interface database includes second candidate application interfaces, second application interface parameters to be filled corresponding to the second candidate application interfaces, application interface descriptions of the second candidate application interfaces, and interface parameter descriptions of the second application interface parameters.

4. The method according to claim 1, wherein The determining a natural language processing result according to the voice request, the list of similar voice requests, and the list of similar application interfaces includes: Performing splicing processing on the voice request, the list of similar voice requests, and the list of similar application interfaces based on a preset prompt template to determine a target prompt; Determining the natural language processing result based on a preset large language model according to the target prompt.

5. The method according to claim 4, wherein The list of similar voice requests includes similar voice requests, first similar application interfaces corresponding to the similar voice requests, and first similar application interface parameters of the first similar application interfaces that have been filled. The determining the natural language processing result based on a preset large language model according to the target prompt includes: Determining the similarity degree between the voice request and the similar voice requests based on the preset large language model according to the target prompt, where the similarity degree includes a first similarity degree, a second similarity degree, and a third similarity degree, the first similarity degree is greater than the second similarity degree, and the second similarity degree is greater than the third similarity degree; Determining the natural language processing result according to the similarity degree and the target prompt.

6. The method according to claim 5, characterized in that, The determining the natural language processing result according to the similarity degree and the target prompt includes: In the case where the similarity degree is the first similarity degree, determining the first similar application interface and the first similar application interface parameters as the natural language processing result.

7. The method according to claim 5, wherein The determining the natural language processing result according to the similarity degree and the target prompt includes: When the similarity degree is the second similarity degree, determine a first target parameter value of the first similar application interface parameter according to the voice request; Perform parameter filling processing on the first similar application interface parameter according to the first target parameter value; Determine the natural language processing result according to the first similar application interface and the result of the parameter filling processing; 8. The method according to claim 5, characterized in that, The list of similar application interfaces includes a second similar application interface, a second similar application interface parameter to be filled corresponding to the second similar application interface, a description of the similar application interface, and a description of the similar interface parameter. Determining the natural language processing result according to the similarity degree and the target prompt word includes: When the similarity degree is the third similarity degree, determine a second target parameter value of the second similar application interface parameter according to the voice request; Perform parameter filling processing on the second similar application interface parameter according to the second target parameter value; Determine the natural language processing result according to the second similar application interface and the result of the parameter filling processing; 9. A vehicle, characterized in that, The vehicle includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the method described in any one of claims 1-8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method described in any one of claims 1-8 are implemented.