Voice interaction method and device, and server

By constructing a pre-defined knowledge database to obtain compressed vehicle technical elements and application interface data, the problem of increased latency in generating large language models in vehicles was solved, resulting in faster natural language processing and an improved user experience.

CN121747547APending Publication Date: 2026-03-27GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The increased latency of large language models in generating structured semantic representations in vehicles leads to a poor user experience, mainly due to the lengthy vehicle reference knowledge being retrieved.

Method used

By constructing a pre-defined knowledge database, compressed data of vehicle technical elements and application interfaces related to voice requests are obtained and processed based on a pre-defined large language model, thereby reducing the length of input data and improving computing speed.

Benefits of technology

It reduces the time required for voice interaction, improves the user experience, and ensures the accuracy and speed of natural language processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747547A_ABST
    Figure CN121747547A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method, a voice interaction device and a server. The method comprises the following steps: determining first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle; searching first vehicle technical element knowledge compressed data corresponding to the first vehicle technical element data and first application program interface knowledge compressed data corresponding to the first application program interface data in a preset knowledge database; and obtaining a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compressed data and the first application program interface knowledge compressed data. In this way, the first vehicle technical element knowledge compressed data and the first application program interface knowledge compressed data replace the uncompressed vehicle technical element knowledge data and the uncompressed application program interface data to be input into the preset large language model, the calculation speed of the preset large language model can be increased, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a voice interaction method, a voice interaction device and a server. BACKGROUND

[0002] In the related art, a user voice request can be parsed into a structured semantic representation by a large language model. However, the retrieved vehicle reference knowledge is relatively lengthy, which can easily increase the time delay of the large language model in generating the structured semantic representation, thereby affecting the user experience. SUMMARY

[0003] The present application provides a voice interaction method, a voice interaction device and a server.

[0004] The present application provides a voice interaction method, a voice interaction device and a server. determining first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle; finding, in a preset knowledge database, first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and first application program interface knowledge compression data corresponding to the first application program interface data, wherein the preset knowledge database includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, and a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; obtaining a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data.

[0005] In this way, in the present application, according to the voice request forwarded by the vehicle, the first vehicle technical element data and the first application program interface data related to the voice request can be obtained, and based on the preset knowledge database, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data can be obtained. Compared with the case where the uncompressed vehicle technical element knowledge data and the uncompressed application program interface data are input into the preset large language model, inputting the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data into the preset large language model instead of the uncompressed vehicle technical element knowledge data and the uncompressed application program interface data can improve the calculation speed of the preset large language model in obtaining the natural language processing result, reduce the time consumption of voice interaction, and improve the user experience to a certain extent.

[0006] In some embodiments, the preset knowledge database comprises an application program interface knowledge compression library and a vehicle technical element knowledge compression library, the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data are searched from the preset knowledge database, comprising: searching the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data from the vehicle technical element knowledge compression library; searching the first application program interface knowledge compression data corresponding to the first application program interface data from the application program interface knowledge compression library.

[0007] In this way, the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data is searched from the vehicle technical element knowledge compression library, and the first application program interface knowledge compression data corresponding to the first application program interface data is searched from the application program interface knowledge compression library. In this way, different search ranges can be divided according to the first vehicle technical element data and the first application program interface data, the search efficiency is improved, the matching accuracy of the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data is guaranteed, and it is ensured that the compression data input into the preset language model is completely consistent with the voice request intention.

[0008] In some embodiments, the natural language processing result is obtained based on the preset large language model, the voice request, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data, comprising: inputting the voice request and a preset prompt information template into the preset large language model for feature extraction processing, and outputting voice request feature data corresponding to the voice request and a preset prompt information compression template corresponding to the preset prompt information template; splicing the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data to determine first splicing data; inputting the first splicing data into the preset large language model to output the natural language processing result.

[0009] Thus, the voice request and the preset prompt information template are input into the preset large language model for feature extraction processing, and voice request feature data corresponding to the voice request and a preset prompt information compression template corresponding to the preset prompt information template are output; the preset prompt information compression template, the voice request feature data, first vehicle technical element knowledge compression data, and first application program interface knowledge compression data are subjected to splicing processing to determine first splicing data; and the first splicing data is input into the preset large language model to output a natural language processing result. In this way, by compressing the voice request and the preset prompt information template, the input format of the preset large language model can be unified, and the length of the input data is further reduced, so that the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data are subjected to splicing processing according to the preset prompt information compression template to obtain structured first splicing data, and it is ensured that the subsequent target large language model can output a natural language processing result consistent with the voice request, the first vehicle technical element data, and the first application program interface data according to the first splicing data.

[0010] In some embodiments, the method further comprises: inputting a plurality of second vehicle technical element data and a plurality of second application program interface data into the preset large language model to output second vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and second application program interface knowledge compression data corresponding to each of the second application program interface data; constructing the preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compression data corresponding to each of the second application program interface data.

[0011] Thus, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the preset large language model, and second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data are output; and the preset knowledge database is constructed according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compression data corresponding to each second application program interface data. In this way, by inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model, the comprehensiveness of the preset knowledge database can be ensured, and standardized second vehicle technical element knowledge compression data and second application program interface knowledge compression data are generated, so as to realize the standardized construction of the preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compression data corresponding to each second application program interface data, thereby providing a structural basis for subsequent voice interaction retrieval.

[0012] In some embodiments, the inputting of the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model to output second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data comprises: inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model to output vehicle technical element coding data of each second vehicle technical element data and application program interface data coding data of each second application program interface data; performing preset data compression processing on the vehicle technical element coding data of each second vehicle technical element data and the application program interface data coding data of each second application program interface data, respectively, to determine the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and the second application program interface knowledge compression data corresponding to each second application program interface data.

[0013] Thus, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the preset large language model, and vehicle technical element code data of each second vehicle technical element data and application program interface code data of each second application program interface data are output; the vehicle technical element code data of each second vehicle technical element data and the application program interface code data of each second application program interface data are respectively subjected to preset data compression processing, and second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data are determined. In this way, the plurality of second vehicle technical element data and the plurality of second application program interface data can be converted into code data through encoding processing, ensuring that the information of each piece of data is completely mapped into a structured vector without information loss, and facilitating subsequent accurate extraction of the features of the data according to the structured vector; and then the data is simplified through compression processing to remove redundant features, laying a foundation for subsequent database construction and retrieval of the preset large language model.

[0014] In some embodiments, the method further comprises: According to the voice request sample, determining third vehicle technical element data and third application program interface data related to the voice request sample; Searching for third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and third application program interface knowledge compression data corresponding to the third application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, and a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; Based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data, and the third application program interface knowledge compression data, updating the preset large language model to determine an updated preset large language model.

[0015] Accordingly, third vehicle technical element data and third application program interface data related to the voice request sample are determined according to the voice request sample; third vehicle technical element knowledge compressed data corresponding to the third vehicle technical element data and third application program interface knowledge compressed data corresponding to the third application program interface data are searched from a preset knowledge database, wherein the preset knowledge database includes a plurality of vehicle technical element data, vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data, and application program interface knowledge compressed data corresponding to each application program interface data; and the preset large language model is updated based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data, and the third application program interface knowledge compressed data, to determine an updated preset large language model. In this way, the inference error of the preset large language model can be quantified based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data, and the third application program interface knowledge compressed data, to update the parameters of the preset large language model based on the error and determine the updated preset large language model.

[0016] In some embodiments, updating the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data, and the third application program interface knowledge compressed data to determine an updated preset large language model comprises: obtaining a prediction result of the voice request sample based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data, and the third application program interface knowledge compressed data; determining a target loss function value according to the prediction result of the voice request sample and a voice request label of the voice request sample; updating the preset large language model according to the target loss function value to determine an updated preset large language model.

[0017] Thus, based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data, and the third application program interface knowledge compression data, a prediction result of the voice request sample is obtained; according to the prediction result of the voice request sample and a voice request label of the voice request sample, a target loss function value is determined; and according to the target loss function value, the preset large language model is updated to determine an updated preset large language model. In this way, the prediction result of the voice request sample and the voice request label of the voice request sample are subjected to deviation calculation, the model inference deviation is accurately identified, and the deviation is quantified by outputting the target loss function value, so as to update the preset large language model based on the target loss function, improve the parameter adjustment efficiency of the subsequent preset large language model, ensure the stability of model updating, and as new functions and new application program interfaces of the vehicle increase, the present embodiment only needs to supplement the corresponding sample and label to realize model optimization without retraining the entire model, thereby reducing the model maintenance and training cost to a certain extent.

[0018] In some embodiments, the preset knowledge database includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data, and the method further includes: inputting a plurality of second vehicle technical element data and a plurality of second application program interface data into the updated preset large language model to output fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each of the second application program interface data; updating the preset knowledge database according to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each of the second application program interface data.

[0019] Thus, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the updated preset large language model, and fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each second application program interface data are output; and the preset knowledge database is updated according to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each second application program interface data. In this way, based on the new preset large language model, the vehicle technical element data and the application program interface data in the preset knowledge database, i.e., the plurality of second vehicle technical element data and the plurality of second application program interface data, are re-compressed, the knowledge compression data corresponding to the vehicle technical element data and the application program interface data is obtained to cover the original knowledge compression data, and the updating of the preset knowledge database is realized.

[0020] The embodiment of the present application provides a voice interaction device, and the device comprises a control module. The control module is configured to determine first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle. The control module is further configured to search, in a preset knowledge database, first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and first application program interface knowledge compression data corresponding to the first application program interface data, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, and a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data. The control module is further configured to obtain a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data.

[0021] The embodiment of the present application provides a server comprising a memory and a processor, wherein the memory stores a computer program, and the computer program is executed by the processor to implement the steps of the above method.

[0022] The embodiment of the present application provides a computer readable storage medium storing a computer program, and the computer program is executed by one or more processors to implement the steps of the above method.

[0023] The server and the computer readable storage medium provided by the embodiments of the present application determine first vehicle technical element data and first application program interface data related to the voice request forwarded by the vehicle according to the voice request forwarded by the vehicle; find first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and first application program interface knowledge compression data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; and obtain a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data. In this way, according to the voice request forwarded by the vehicle, the first vehicle technical element data and the first application program interface data related to the voice request can be obtained, and based on the preset knowledge database, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data can be obtained to reduce the length of the input preset large language model data, improve the calculation speed of the preset large language model to obtain the natural language processing result, reduce the time consumption of voice interaction, and improve the user experience to a certain extent.

[0024] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the description of the embodiments, taken in conjunction with the following drawings in which: Figure 1 is one of the flow diagrams of the voice interaction method of some embodiments of the present application; Figure 2 is a retrieval flow diagram of some embodiments of the present application; Figure 3 is the second flow diagram of the voice interaction method of some embodiments of the present application; Figure 4 is the third flow diagram of the voice interaction method of some embodiments of the present application; Figure 5 is the fourth flow diagram of the voice interaction method of some embodiments of the present application; Figure 6 is the fifth flow diagram of the voice interaction method of some embodiments of the present application; Figure 7 is a construction scheme diagram of the preset knowledge database of some embodiments of the present application; Figure 8 This is a flowchart of a voice interaction method according to certain embodiments of this application, number six. Figure 9 This is the seventh flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 10 This is the eighth flowchart of a voice interaction method according to certain embodiments of this application; Figure 11 This is a schematic diagram illustrating the update strategy of some embodiments of this application. Detailed Implementation

[0026] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.

[0027] In related technologies, a large language model (LLM) can be used to process vehicle technical element data, user voice requests, and application programming interfaces (APIs) that are semantically similar to user voice requests, to obtain structured semantic representations, such as natural language understanding (NLU) results, so as to control the execution of corresponding vehicle cockpit functions based on the NLU results.

[0028] However, the retrieved vehicle technical element data is often quite lengthy, resulting in a large overall length of the large language model data, which consists of vehicle technical element data, user voice requests, and application programming interfaces (APIs) that are semantically similar to user voice requests. This leads to a significant increase in the latency of the large language model generating the first token, which not only makes the actual deployment of the model difficult but also causes users to experience a noticeable delay while waiting for the system response, seriously affecting the user's interactive experience.

[0029] Based on the above issues, please refer to Figure 1 This application provides a voice interaction method, the method including: 01: Determine the first vehicle technical element data and the first application programming interface data related to the voice request forwarded by the vehicle; 02: search for first vehicle technical element knowledge compressed data corresponding to the first vehicle technical element data and first application program interface knowledge compressed data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data; 03: obtain a natural language processing result based on the preset large language model, the voice request, the first vehicle technical element knowledge compressed data and the first application program interface knowledge compressed data.

[0030] The voice interaction method of the embodiment of the application can be implemented by the voice interaction device of the embodiment of the application. Specifically, the voice interaction device comprises a control module. The control module determines first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle. The control module is also configured to search for first vehicle technical element knowledge compressed data corresponding to the first vehicle technical element data and first application program interface knowledge compressed data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data. The control module is also configured to obtain a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compressed data and the first application program interface knowledge compressed data.

[0031] The voice interaction method of the embodiment of the application can be implemented by the voice interaction device of the embodiment of the application. Specifically, the voice interaction device comprises a control module. The control module determines first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle. The control module is also configured to search for first vehicle technical element knowledge compressed data corresponding to the first vehicle technical element data and first application program interface knowledge compressed data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data. The control module is also configured to obtain a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compressed data and the first application program interface knowledge compressed data.

[0032] Specifically, the voice request refers to a natural language instruction issued by a user through a vehicle-mounted voice collection device, which requires the vehicle to perform a specific function, such as "kinetic energy recovery standard instruction" and the like.

[0033] The first vehicle technical element data refers to data that can represent vehicle expertise extracted from the voice request after obtaining the voice request forwarded by the vehicle, such as vehicle function or component information, for providing clear vehicle domain knowledge reference for subsequent preset large language models.

[0034] The first application program interface data refers to an API with high semantic similarity to the voice request, for providing a standardized interface matching the function of the first voice request for subsequent preset large language models.

[0035] The preset knowledge database is a pre-constructed static database, including a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data.

[0036] The first vehicle technical element knowledge compression data is a simplified knowledge representation corresponding to the first vehicle technical element data, which can be obtained by searching in the preset knowledge data according to the first vehicle technical element data.

[0037] The first application program interface knowledge compression data is a compressed knowledge representation corresponding to the first application program interface data, which can be obtained by searching in the preset knowledge data according to the first vehicle technical element data.

[0038] In one example, the voice request is searched by an Aho-Corasick (AC) automaton, which can match the technical element generalization words corresponding to the voice request, and obtain the first vehicle technical element data through the mapping relationship data between the technical element generalization words and the original technical element words, to convert the user's natural language instruction into vehicle professional knowledge understandable by the LLM, and solve the problem of user instruction ambiguity affecting the accuracy of subsequent knowledge input.

[0039] In Figure 2 In the AC automaton search process shown, according to the user instruction of "kinetic energy recovery standard mode", i.e. the voice request, the AC automaton can retrieve "kinetic energy", "kinetic energy recovery", and "standard mode" generalization words from the Entity library and the Entity generalization word library, i.e. the pre-obtained vehicle technical element library and vehicle technical element generalization word library, and by deduplicating the original technical element words corresponding to the generalization words, the unique technical element words matching the user instruction, i.e. the first vehicle technical element data, can be obtained.

[0040] It should be noted that after deduplication, if there are multiple technical element words matching the user instruction, the multiple technical element words are mapped to vehicle technical element data through the technical element word-vehicle technical element mapping relationship data to obtain the mapped vehicle technical element data, providing a mapping basis for subsequent construction or updating of the preset knowledge database.

[0041] By retrieving the voice request through a Retrieval-Augmented Generation (RAG) model, the voice request can be converted into a vector to match the first several API description vectors in the API library that have the highest semantic similarity with the voice request vector, and the multiple API description vectors are mapped to APIs, i.e., the first application program interface data is obtained.

[0042] Before generating NLU through the preset large language model, the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data can be obtained by querying the preset knowledge database.

[0043] It can be understood that in the related art, the length of the vehicle technical element data and the application program interface data input into the preset large language model is usually long, which can easily increase the time delay of the preset large language model in generating natural language processing results, thereby affecting the speed of voice interaction response and user experience.

[0044] According to the static mapping relationship between the preset knowledge database, the vehicle technical element data, and the vehicle technical element knowledge compression data, the static mapping relationship between the application program interface data and the application program interface knowledge compression data, and the input first vehicle technical element data and first application program interface data, the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data are obtained, the process of real-time compression of lengthy knowledge is skipped, the quick retrieval of compressed knowledge is realized, and the time consumption is reduced.

[0045] By inputting the voice request, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data into the preset large language model, the preset large language model can understand the intent of the voice request based on the simplified compressed knowledge, avoid processing lengthy knowledge data, quickly complete semantic analysis, and generate structured natural language processing results, so as to realize voice interaction according to the natural language processing results.

[0046] In summary, in the embodiments of the present application, according to the voice request forwarded by the vehicle, the first vehicle technical element data and the first application program interface data related to the voice request can be obtained, and based on the preset knowledge database, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data are obtained. Compared with the case where the uncompressed vehicle technical element knowledge data and the uncompressed application program interface data are input into the preset large language model, inputting the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data into the preset large language model instead of the uncompressed vehicle technical element knowledge data and the uncompressed application program interface data can improve the calculation speed of the preset large language model to obtain the natural language processing result, reduce the time consumption of voice interaction, and improve the user experience to a certain extent.

[0047] Please refer to Figure 3 In some embodiments, the preset knowledge database includes an application program interface knowledge compression library and a vehicle technical element knowledge compression library, and step 02 (finding the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data in the preset knowledge database) includes: 021: finding the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data in the vehicle technical element knowledge compression library; 022: finding the first application program interface knowledge compression data corresponding to the first application program interface data in the application program interface knowledge compression library.

[0048] In some embodiments, the control module finds the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data in the vehicle technical element knowledge compression library. The control module is also configured to find the first application program interface knowledge compression data corresponding to the first application program interface data in the application program interface knowledge compression library.

[0049] In some embodiments, the processor is also configured to find the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data in the vehicle technical element knowledge compression library. The processor is also configured to find the first application program interface knowledge compression data corresponding to the first application program interface data in the application program interface knowledge compression library.

[0050] Specifically, the preset knowledge database includes an application program interface knowledge compression library and a vehicle technical element knowledge compression library, wherein the vehicle technical element knowledge compression library includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data; and the application program interface knowledge compression library includes a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data.

[0051] For the vehicle technical element data, the vehicle technical element knowledge compression library can be called to quickly locate the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data according to the first vehicle technical element data by traversing all mappings in the vehicle technical element knowledge compression library.

[0052] In one example, the first vehicle technical element data includes "kinetic energy recovery" and "standard mode", and the vehicle technical element knowledge compression library can be called to traverse the key-value mappings in the vehicle technical element knowledge compression library with "kinetic energy recovery" and "standard mode" as keys, respectively, to obtain the simplified vector corresponding to "kinetic energy recovery", i.e., the corresponding first vehicle technical element knowledge compression data, which supports the "X-pedal driving mode and supports level adjustment". The application program interface data is similar to the vehicle technical element data, and the first application program interface knowledge compression data corresponding to the first application program interface data can be quickly located by referring to the above steps, which will not be described here.

[0053] According to different data types, different search ranges can be divided to improve search efficiency and ensure the matching accuracy of the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data, so as to ensure that the compressed data input into the preset language model is completely consistent with the voice request intention.

[0054] In this way, the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data in the vehicle technical element knowledge compression library is found, and the first application program interface knowledge compression data corresponding to the first application program interface data in the application program interface knowledge compression library is found. In this way, according to the first vehicle technical element data and the first application program interface data, different search ranges can be divided to improve search efficiency and ensure the matching accuracy of the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data, so as to ensure that the compressed data input into the preset language model is completely consistent with the voice request intention.

[0055] Please refer to Figure 4 In some embodiments, step 03 (obtaining a natural language processing result based on a preset large language model, a voice request, first vehicle technical element knowledge compression data, and first application program interface knowledge compression data) comprises: 031: input the voice request and the preset prompt information template into the preset large language model for feature extraction processing, and output the voice request feature data corresponding to the voice request and the preset prompt information compression template corresponding to the preset prompt information template; 032: splice the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data to determine first spliced data; 033: input the first spliced data into the preset large language model to output a natural language processing result.

[0056] In some embodiments, the control module is further configured to input the voice request and the preset prompt information template into the preset large language model for feature extraction processing, to output voice request feature data corresponding to the voice request and a preset prompt information compression template corresponding to the preset prompt information template. The control module is further configured to splice the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data to determine first spliced data. The control module is further configured to input the first spliced data into the preset large language model to output a natural language processing result.

[0057] In some embodiments, the processor is further configured to input the voice request and the preset prompt information template into the preset large language model for feature extraction processing, to output voice request feature data corresponding to the voice request and a preset prompt information compression template corresponding to the preset prompt information template. The processor is further configured to splice the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data to determine first spliced data. The processor is further configured to input the first spliced data into the preset large language model to output a natural language processing result.

[0058] Specifically, the preset prompt information template is a pre-set text framework, which can be used to standardize the organization logic of the input information of the preset large language model.

[0059] The feature extraction processing can be used to further compress the length of the data input into the preset large language model.

[0060] The splicing processing is used to splice the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data.

[0061] The first spliced data refers to input data with unified structure, which can be obtained by splicing the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data.

[0062] Based on the preset large language model, the voice request and the preset prompt information template can be subjected to feature extraction processing, the length of the data input into the preset large language model is further compressed, the voice request feature data corresponding to the voice request and the preset prompt information compression template corresponding to the preset prompt information template are obtained, and the input data format of the preset large language model is unified.

[0063] By calling the preset prompt information compression template, the first vehicle technical element knowledge compression data and the first vehicle technical element knowledge compression data can be respectively filled in the preset prompt information compression template, and structured input data, i.e., the first splicing data, is formed.

[0064] Compared with inputting lengthy knowledge data into the preset large language model, the first splicing data obtained by the present application is used to replace the original input data and is input into the preset large language model, which can ensure stable and accurate output of the natural language processing result, reduce the inference time delay of the preset large language model, for example, compress more than 1000 tokens into less than 300 tokens, and further improve the calculation speed of the preset large language model in obtaining the natural language processing result, thereby improving the user experience to a certain extent.

[0065] In this way, the voice request and the preset prompt information template are input into the preset large language model for feature extraction processing, and the voice request feature data corresponding to the voice request and the preset prompt information compression template corresponding to the preset prompt information template are output. The preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data are subjected to splicing processing to determine the first splicing data. The first splicing data is input into the preset large language model, and the natural language processing result is output. In this way, by compressing the voice request and the preset prompt information template, the input format of the preset large language model can be unified, the length of the input data is further reduced, and the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data are subjected to splicing processing according to the preset prompt information compression template to obtain structured first splicing data, so that the subsequent target large language model can output the natural language processing result consistent with the voice request, the first vehicle technical element data, and the first application program interface data according to the first splicing data.

[0066] Please refer to Figure 5 In some embodiments, the method further comprises: 04: inputting a plurality of second vehicle technical element data and a plurality of second application program interface data into the preset large language model, outputting second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and second application program interface knowledge compression data corresponding to each second application program interface data; 05: constructing a preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compressed data corresponding to each second application program interface data.

[0067] In some embodiments, the control module is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into a preset large language model, and output the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data and the second application program interface knowledge compressed data corresponding to each second application program interface data. The control module is further configured to construct the preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compressed data corresponding to each second application program interface data.

[0068] In some embodiments, the processor is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into a preset large language model, and output the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data and the second application program interface knowledge compressed data corresponding to each second application program interface data. The processor is further configured to construct the preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compressed data corresponding to each second application program interface data.

[0069] Specifically, the second vehicle technical element data is the original vehicle technical element basic data for constructing the preset knowledge database, and needs to include all vehicle function technical elements that can be controlled by voice, to ensure that the preset knowledge database is comprehensive, so that in the process of voice interaction, the first vehicle technical element knowledge compressed data corresponding to the first vehicle technical element data related to the voice request forwarded by the vehicle can be found in the preset knowledge database.

[0070] The second application program interface data is the original API basic data for constructing the preset knowledge database, which is a collection of all APIs for implementing vehicle function control, and needs to include the name, parameter range, function description and other information of each API, to ensure that the preset knowledge database is comprehensive, so that in the process of voice interaction, the first application program interface knowledge compressed data corresponding to the first application program interface data related to the voice request forwarded by the vehicle can be found in the preset knowledge database.

[0071] It can be understood that constructing the preset knowledge database is to obtain vehicle technical element knowledge compressed data corresponding to the vehicle technical element data, construct a mapping relationship between the vehicle technical element data and the vehicle technical element knowledge compressed data, and obtain program interface knowledge compressed data corresponding to the program interface data, and construct a mapping relationship between the program interface data and the program interface knowledge compressed data. The originally lengthy vehicle technical element data and application program interface data are converted into feature vectors, the matching of the feature vectors is used to replace the matching of long texts in the related technology, the overhead of LLM real-time calculation is reduced, and the retrieval speed of LLM is improved.

[0072] Therefore, the preset knowledge database construction needs to be realized by the preset large language model with semantic understanding and coding ability.

[0073] Based on the preset large language model, by performing feature extraction processing on the pre-existing multiple second vehicle technical element data and multiple second application program interface data, the redundant description of the vehicle technical element data and the application program interface data can be removed, and they are converted into compact and semantic feature vectors, that is, the second vehicle technical element knowledge compressed data corresponding to each second vehicle technical element data, and the second application program interface knowledge compressed data corresponding to each second application program interface data are obtained.

[0074] The preset knowledge database includes an application program interface knowledge compressed library and a vehicle technical element knowledge compressed library. For vehicle technical element data, the vehicle technical element knowledge compressed library can be constructed according to the multiple second vehicle technical element data, the multiple second vehicle technical element knowledge compressed data, and the mapping data of each second vehicle technical element data-second vehicle technical element knowledge compressed data. For API, the application program interface knowledge compressed library can be constructed according to the multiple second application program interface data, the multiple second application program interface knowledge compressed data, and the mapping data of each second application program interface data-second application program interface knowledge compressed data, to provide a structural basis for subsequent retrieval.

[0075] Thus, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the preset large language model, and second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data are output; and the preset knowledge database is constructed according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compression data corresponding to each second application program interface data. In this way, by inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model, the comprehensiveness of the preset knowledge database can be ensured, and standardized second vehicle technical element knowledge compression data and second application program interface knowledge compression data are generated, so as to realize the standardized construction of the preset knowledge database according to the plurality of second vehicle technical element data and the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the second application program interface knowledge compression data corresponding to each second application program interface data, thereby providing a structural basis for subsequent voice interaction retrieval.

[0076] Referring to Figure 6 In some embodiments, step 04 (inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model, and outputting second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data) comprises: 041: inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the preset large language model, and outputting vehicle technical element coding data of each second vehicle technical element data and application program interface data coding data of each second application program interface data; 042: performing preset data compression processing on the vehicle technical element coding data of each second vehicle technical element data and the application program interface data coding data of each second application program interface data respectively, to determine second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data.

[0077] In some embodiments, the control module is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into a preset large language model to output vehicle technical element coding data of each second vehicle technical element data and application program interface data coding data of each second application program interface data. The control module is further configured to perform preset data compression processing on the vehicle technical element coding data of each second vehicle technical element data and the application program interface data coding data of each second application program interface data respectively, to determine second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data.

[0078] In some embodiments, the processor is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into a preset large language model to output vehicle technical element coding data of each second vehicle technical element data and application program interface data coding data of each second application program interface data. The processor is further configured to perform preset data compression processing on the vehicle technical element coding data of each second vehicle technical element data and the application program interface data coding data of each second application program interface data respectively, to determine second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data.

[0079] Specifically, the vehicle technical element coding data refers to a structured vector obtained by performing semantic coding on the second vehicle technical element data, and is used to improve the accuracy of compressing the second vehicle technical element data.

[0080] The application program interface data coding data refers to a structured vector obtained by performing semantic coding on the second application program interface data, and is used to improve the accuracy of compressing the second application program interface data.

[0081] Inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into a preset large language model to output second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data includes two steps of coding processing and compression processing.

[0082] Firstly, in the encoding process, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the preset large language model, and the semantic encoding of each second vehicle technical element data and each second application program interface data can be performed, so as to convert the plurality of second vehicle technical element data and the plurality of second application program interface data into semantic vector data, i.e., the vehicle technical element encoding data of each second vehicle technical element data and the application program interface data encoding data of each second application program interface data, thereby providing uniform standard data for subsequent compression processing.

[0083] Understandably, compared with directly extracting feature data of vehicle technical element data and application program interface data, converting the vehicle technical element data and the application program interface data into encoding data can ensure that the information of each piece of data is completely mapped into a structured vector without information loss, and the feature of the data can be accurately extracted according to the structured vector.

[0084] Then, in the compression process, the vehicle technical element encoding data of each second vehicle technical element data and the application program interface data encoding data of each second application program interface data are respectively subjected to preset data compression processing, so as to perform standardized compression operation on the encoding data, remove redundant features such as repeatedly appearing features, retain semantic feature information, and finally convert the encoding data into second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and second application program interface knowledge compression data corresponding to each second application program interface data, so as to simplify the data and lay a foundation for subsequent database construction and retrieval of the preset large language model.

[0085] In this way, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the preset large language model, and the vehicle technical element encoding data of each second vehicle technical element data and the application program interface data encoding data of each second application program interface data are output; the vehicle technical element encoding data of each second vehicle technical element data and the application program interface data encoding data of each second application program interface data are respectively subjected to preset data compression processing, and the second vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and the second application program interface knowledge compression data corresponding to each second application program interface data are determined. In this way, the plurality of second vehicle technical element data and the plurality of second application program interface data can be converted into encoding data through the encoding process, so as to ensure that the information of each piece of data is completely mapped into a structured vector without information loss, and the feature of the data can be accurately extracted according to the structured vector; and the data is simplified through the compression process, and the redundant features are removed, thereby laying a foundation for subsequent database construction and retrieval of the preset large language model.

[0086] In some examples, please refer toFigure 7 The construction scheme of the preset knowledge database is based on a vehicle knowledge base and an API knowledge base, and vehicle knowledge (second vehicle technical element data) and API knowledge (second application program interface data) can be obtained from the vehicle knowledge base and the API knowledge base, respectively. The vehicle knowledge base has n pieces of vehicle knowledge, the API knowledge base has m pieces of API knowledge, and the length of each piece of knowledge is L.

[0087] First, all vehicle knowledge and all API knowledge in the vehicle knowledge base and the API knowledge base are input into the LLM. Then, embedding is performed on all vehicle knowledge and all API by the LLM to perform semantic analysis and numerical conversion, and (m+n) pieces of L*d encoding data can be obtained, where d is the dimension of each token in the LLM. Then, the mean value of the (m+n) pieces of L*d feature vectors is obtained, that is, the 1*d feature vector of each piece of knowledge is obtained. Finally, the key-value mapping relationship between the vehicle knowledge and the feature vector of each vehicle knowledge, and the key-value mapping relationship between the API knowledge and the feature vector of each API knowledge are constructed, and the preset knowledge database is constructed according to the key-value mapping relationship, so that the AC automatic machine and the RAG model can quickly find the feature vector according to the preset knowledge database in the process of voice interaction.

[0088] Please refer to Figure 8 In some embodiments, the method further comprises: 06: determining third vehicle technical element data and third application program interface data related to the voice request sample according to the voice request sample; 07: searching for third vehicle technical element knowledge compressed data corresponding to the third vehicle technical element data and third application program interface knowledge compressed data corresponding to the third application program interface data in the preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, and a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data; 08: updating the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data, and the third application program interface knowledge compressed data, and determining an updated preset large language model.

[0089] In some embodiments, the control module is further configured to determine, according to the voice request sample, third vehicle technical element data and third application program interface data related to the voice request sample. The control module is further configured to search, in a preset knowledge database, third vehicle technical element knowledge compressed data corresponding to the third vehicle technical element data and third application program interface knowledge compressed data corresponding to the third application program interface data, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data. The control module is further configured to update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data and the third application program interface knowledge compressed data, and determine an updated preset large language model.

[0090] In some embodiments, the processor is further configured to determine, according to the voice request sample, third vehicle technical element data and third application program interface data related to the voice request sample. The processor is further configured to search, in a preset knowledge database, third vehicle technical element knowledge compressed data corresponding to the third vehicle technical element data and third application program interface knowledge compressed data corresponding to the third application program interface data, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compressed data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compressed data corresponding to each application program interface data. The processor is further configured to update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compressed data and the third application program interface knowledge compressed data, and determine an updated preset large language model.

[0091] Specifically, the voice request sample is a standardized data set for training the updated LLM, covering various vehicle function scenarios such as basic function control, new function calling, fault consultation, etc.

[0092] The third vehicle technical element data is vehicle technical element information related to the sample content extracted from the voice request sample, similar to the first vehicle technical element data.

[0093] The third application program interface data is API information related to the sample content extracted from the voice request sample, similar to the first application program interface data.

[0094] The third vehicle technical element knowledge compressed data is a concise knowledge representation corresponding to the third vehicle technical element data, which can be obtained by searching in the preset knowledge data according to the third vehicle technical element data.

[0095] The third application program interface knowledge compression data corresponds to a compressed knowledge representation of the third application program interface data, and can be obtained by searching in the preset knowledge data according to the third vehicle technical element data.

[0096] It can be understood that the updating mechanism of the preset language model is similar to the voice interaction process, and the processes of obtaining the third vehicle technical element data and the third application program interface data, and obtaining the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data can refer to the related content of the voice interaction described above, and will not be described herein.

[0097] Based on the preset large language model, the inference error of the LLM can be quantified according to the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, so as to update the parameters of the preset large language model based on the error, and determine the updated preset large language model.

[0098] In this way, according to the voice request sample, the third vehicle technical element data and the third application program interface data related to the voice request sample are determined; the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data are searched in the preset knowledge database, wherein the preset knowledge database includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, the preset large language model is updated, and the updated preset large language model is determined. In this way, according to the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, the inference error of the preset large language model can be quantified, so as to update the parameters of the preset large language model based on the error, and determine the updated preset large language model.

[0099] Please refer to Figure 9 In some embodiments, step 08 (updating the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, and determining the updated preset large language model) includes: 081: obtaining a prediction result of the voice request sample based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data; 082: Determine a target loss function value according to the prediction result of the voice request sample and the voice request label of the voice request sample; 083: Update the preset large language model according to the target loss function value to determine an updated preset large language model.

[0100] In some embodiments, the control module is further configured to obtain a prediction result of the voice request sample based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data, and the third application program interface knowledge compression data. The control module is further configured to determine a target loss function value according to the prediction result of the voice request sample and a voice request label of the voice request sample. The control module is further configured to update the preset large language model according to the target loss function value to determine an updated preset large language model.

[0101] In some embodiments, the processor is further configured to obtain a prediction result of the voice request sample based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data, and the third application program interface knowledge compression data. The processor is further configured to determine a target loss function value according to the prediction result of the voice request sample and a voice request label of the voice request sample. The processor is further configured to update the preset large language model according to the target loss function value to determine an updated preset large language model.

[0102] Specifically, the prediction result of the voice request sample is a semantic analysis result of the voice request sample by the current preset large language model, which is used for error analysis.

[0103] The voice request label of the voice request sample is a standard correct result corresponding to the voice request sample, such as an artificial annotation or a benchmark answer generated based on vehicle function logic.

[0104] The target loss function value is a numerical value quantifying the deviation between the prediction result and the standard correct result, which can be used to optimize the preset large language model.

[0105] The deviation between the prediction result of the voice request sample and the voice request label of the voice request sample is calculated, and the severity of the deviation is obtained by comparing the API in the prediction result and the voice request label, so as to accurately identify the model reasoning deviation. Then, the target loss function value is output to quantify the deviation, so as to improve the parameter adjustment efficiency of the subsequent preset large language model.

[0106] The target loss function value is input into a parameter update module of the preset large language model, and then a back propagation algorithm is started to update each parameter value in the preset large language model until the target loss function value is lower than a preset threshold value, and the update is stopped to obtain an updated preset large language model.

[0107] Further, as new vehicle functions and new APIs increase, the embodiments of the present application only need to supplement the corresponding samples and labels to realize model optimization without retraining the entire model, thereby reducing the model maintenance and training costs to a certain extent.

[0108] In this way, the prediction result of the voice request sample is obtained based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data, and the third application program interface knowledge compression data; the target loss function value is determined according to the prediction result of the voice request sample and the voice request label of the voice request sample; and the preset large language model is updated according to the target loss function value to determine the updated preset large language model. In this way, the deviation calculation is performed on the prediction result of the voice request sample and the voice request label of the voice request sample, the model inference deviation is accurately identified, and the deviation is quantified by outputting the target loss function value to update the preset large language model based on the target loss function, improve the parameter adjustment efficiency of the subsequent preset large language model, ensure the stability of the model update, and as new vehicle functions and new APIs increase, the embodiments of the present application only need to supplement the corresponding samples and labels to realize model optimization without retraining the entire model, thereby reducing the model maintenance and training costs to a certain extent.

[0109] Please refer to Figure 10 In some embodiments, the preset knowledge database includes a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data, and the method further includes: 09: inputting the plurality of second vehicle technical element data and the plurality of second application program interface data into the updated preset large language model to output fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each second application program interface data; 010: updating the preset knowledge database according to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each second application program interface data.

[0110] In some embodiments, the control module is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into the updated preset large language model, and output fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each of the second application program interface data. The control module is further configured to update the preset knowledge database according to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each of the second application program interface data.

[0111] In some embodiments, the processor is further configured to input the plurality of second vehicle technical element data and the plurality of second application program interface data into the updated preset large language model, and output fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each of the second application program interface data. The processor is further configured to update the preset knowledge database according to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each of the second application program interface data.

[0112] Specifically, the second vehicle technical element data and the second application program interface data are data already existing in the preset knowledge database.

[0113] The fourth vehicle technical element knowledge compression data is a concise knowledge representation corresponding to the second vehicle technical element data obtained based on the new preset large language model, and can be different from the second vehicle technical element knowledge compression data.

[0114] The fourth application program interface knowledge compression data is compressed knowledge corresponding to the second application program interface data obtained based on the new preset large language model, and can be different from the second application program interface knowledge compression data.

[0115] It can be understood that the preset knowledge database is updated as the preset large language model is updated, and the logic of updating the preset knowledge database is similar to the logic of constructing the preset knowledge database. The updating of the preset knowledge database is based on the new preset large language model, and the steps of updating the preset knowledge database in the embodiments of the present application can refer to the related contents of constructing the preset knowledge database described above, which will not be described here.

[0116] Thus, the plurality of second vehicle technical element data and the plurality of second application program interface data are input into the updated preset large language model, fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each second application program interface data are output, and the preset knowledge database is updated according to the plurality of second vehicle technical element data, the fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, the plurality of second application program interface data, and the fourth application program interface knowledge compression data corresponding to each second application program interface data. In this way, based on the new preset large language model, the vehicle technical element data and the application program interface data in the preset knowledge database are re-compressed, that is, the plurality of second vehicle technical element data and the plurality of second application program interface data, the knowledge compression data corresponding to the vehicle technical element data and the application program interface data can be obtained to cover the original knowledge compression data, and the updating of the preset knowledge database is realized.

[0117] The process of model updating is similar to the voice interaction process. The following takes Figure 11 as an example to explain the updating strategy of the embodiments of the present application: Firstly, based on the user instruction ( "kinetic energy recovery standard mode" ), that is, driving the AC automatic machine and RAG retrieval, the knowledge features related to the instruction, that is, the third vehicle technical element data and the third application program interface data, and the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data, are extracted from the vehicle knowledge base and the API knowledge base respectively. Then, these knowledge features are input into the LLM ( preset large language model) together with the system prompt ( preset prompt information template) and the user instruction. Next, the LLM generates the target output ( such as NLU result) based on the input knowledge features, system prompt and user instruction. At the same time, the matching degree of the output and the label ( label) is combined to adjust the model parameters of the LLM itself through the loss back mechanism, and the semantic understanding and generation ability are optimized. Finally, after the model is updated, the second vehicle technical element data, the fourth vehicle technical element knowledge compression data corresponding to each second vehicle technical element data, and the fourth application program interface knowledge compression data corresponding to each second application program interface data are re-acquired based on the new LLM model to update the vehicle knowledge base and the API knowledge base ( preset knowledge database).

[0118] The present application also provides a computer readable storage medium having a computer program stored thereon. When the computer program processor is executed, the steps of the voice interaction method as described above are realized.

[0119] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable storage medium can include any technical elements or devices capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0120] In the description of the present specification, the description referring to the terms "specifically", "further", "particularly", "it can be understood that", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not intend to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0121] Any process or method descriptions in flow charts or described elsewhere herein can be understood as representing one or more modules, segments, or portions of code that include one or more executable instructions for implementing specific logic functions or steps in the processes. The scope of preferred embodiments of the present application includes additional implementation in which the functions are performed in different orders, in substantially simultaneous fashion, or in reverse order, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0122] Although the embodiments of the present application have been shown and described above, it can be understood that the above-described embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.

Claims

1. A voice interaction method, characterized in that, The method comprises: determining first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle; finding first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and first application program interface knowledge compression data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data, vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data, and application program interface knowledge compression data corresponding to each application program interface data; obtaining a natural language processing result based on a preset large language model, the voice request, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data.

2. The method of claim 1, wherein, The preset knowledge database comprises an application program interface knowledge compression library and a vehicle technical element knowledge compression library, and the finding of the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and the first application program interface knowledge compression data corresponding to the first application program interface data in the preset knowledge database comprises: finding the first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data in the vehicle technical element knowledge compression library; finding the first application program interface knowledge compression data corresponding to the first application program interface data in the application program interface knowledge compression library.

3. The method of claim 1, wherein, The obtaining of the natural language processing result based on the preset large language model, the voice request, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data comprises: inputting the voice request and a preset prompt information template into the preset large language model for feature extraction processing, and outputting voice request feature data corresponding to the voice request and a preset prompt information compression template corresponding to the preset prompt information template; performing splicing processing on the preset prompt information compression template, the voice request feature data, the first vehicle technical element knowledge compression data, and the first application program interface knowledge compression data to determine first splicing data; inputting the first splicing data into the preset large language model to output the natural language processing result.

4. The method of claim 1, wherein, The method further comprises: inputting a plurality of second vehicle technical element data and a plurality of second application program interface data into the preset large language model to output second vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and second application program interface knowledge compression data corresponding to each of the second application program interface data; constructing the preset knowledge database according to the plurality of second vehicle technical element data, the second vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, the plurality of second application program interface data, and the second application program interface knowledge compression data corresponding to each of the second application program interface data.

5. The method of claim 4, wherein, The method further comprises: According to the voice request sample, determine the third vehicle technical element data and the third application program interface data related to the voice request sample; Find the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data in the preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; 6. The method of claim 1, wherein, Update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, and determine the updated preset large language model. The method further comprises: According to the voice request sample, determine the third vehicle technical element data and the third application program interface data related to the voice request sample; Find the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data in the preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; 7. The method of claim 6, wherein, Update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, and determine the updated preset large language model. The method further comprises: According to the voice request sample, determine the third vehicle technical element data and the third application program interface data related to the voice request sample; Find the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data in the preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; 8. The method of claim 7, wherein, Update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, and determine the updated preset large language model. The method further comprises: According to the voice request sample, determine the third vehicle technical element data and the third application program interface data related to the voice request sample; Find the third vehicle technical element knowledge compression data corresponding to the third vehicle technical element data and the third application program interface knowledge compression data corresponding to the third application program interface data in the preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data; Update the preset large language model based on the preset large language model, the voice request sample, the third vehicle technical element knowledge compression data and the third application program interface knowledge compression data, and determine the updated preset large language model. The plurality of second vehicle technical element data and the plurality of second application program interface data are input into the updated preset large language model, and fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data and fourth application program interface knowledge compression data corresponding to each of the second application program interface data are output. According to the plurality of second vehicle technical element data and the fourth vehicle technical element knowledge compression data corresponding to each of the second vehicle technical element data, and the plurality of second application program interface data and the fourth application program interface knowledge compression data corresponding to each of the second application program interface data, the preset knowledge database is updated.

9. A voice interaction device, characterized by The device comprises: The control module is configured to determine first vehicle technical element data and first application program interface data related to a voice request forwarded by a vehicle, and find first vehicle technical element knowledge compression data corresponding to the first vehicle technical element data and first application program interface knowledge compression data corresponding to the first application program interface data in a preset knowledge database, wherein the preset knowledge database comprises a plurality of vehicle technical element data and vehicle technical element knowledge compression data corresponding to each vehicle technical element data, a plurality of application program interface data and application program interface knowledge compression data corresponding to each application program interface data, and a natural language processing result is obtained based on a preset large language model, the voice request, the first vehicle technical element knowledge compression data and the first application program interface knowledge compression data.

10. A server, characterized by The device comprises a memory and a processor, and the memory stores a computer program which is executed by the processor to implement the method of any one of claims 1-8.