Model training method of large language model, server and storage medium

By determining the associated application interface list and target prompt words in the large language model and adopting constraint-guided training, the problem that large language models are difficult to generate correct vehicle control instructions in the vehicle cockpit scenario is solved, and the model's reasoning ability and user experience are improved.

CN120183389APending Publication Date: 2025-06-20SHANGHAI XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510474860.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the vehicle cockpit scenario, it is difficult for existing large language models to generate vehicle control instructions that can be executed correctly, resulting in poor user experience.

Method used

By determining the associated application interface list, target prompt word and first target data pair, a constraint-guided training method is used to optimize the large language model so that it can generate correct vehicle control instructions.

Benefits of technology

Improve the reasoning ability of the large language model, ensuring that voice requests for correctly executed vehicle control instructions cannot be generated before training, and correct and executable instructions can be obtained, thereby improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183389A_ABST
    Figure CN120183389A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method of a large language model, a server and a computer readable storage medium. The method comprises the steps that an associated application interface list is determined according to original training data, the original training data comprises an original voice request, and a preset reasoning thinking chain and a preset reasoning result corresponding to the original voice request, and based on the preset reasoning thinking chain and the preset reasoning result, an associated application interface list which can be executed by a vehicle can be generated; and a vehicle control instruction corresponding to the original voice request. Then, based on a preset cue word template, according to the associated application interface list and the original voice request, determining a target cue word; then, determining a first target data pair according to the target prompt word, a preset reasoning thinking chain and a preset reasoning result; and finally, training the to-be-trained large language model according to the first target data pair. In this way, by conducting constraint guiding type training on the large language model, the reasoning ability of the large language model is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training, and particularly to a model training method, a server, and a computer-readable storage medium for a large language model. Background Art

[0002] In the related art, in the vehicle cockpit scenario, when a user issues a voice request, the large language model can directly output a preset inference result that meets the user's needs based on the text recognition result of the voice request, and then generate a corresponding vehicle control instruction that meets the user's needs. However, for some voice requests, according to the preset inference result obtained by the large language model, it is often impossible to generate a vehicle control instruction that can be correctly executed, resulting in a poor user experience. Summary of the Invention

[0003] The present application provides a model training method, a server, and a computer-readable storage medium for a large language model.

[0004] An embodiment of the present application provides a model training method for a large language model, the method comprising:

[0005] Determine an associated application interface list according to the original training data, wherein the original training data includes an original voice request, a preset inference thought chain corresponding to the original voice request, and a preset inference result, and a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated based on the preset inference thought chain and the preset inference result;

[0006] Determine a target prompt word based on a preset prompt word template according to the associated application interface list and the original voice request;

[0007] Determine a first target data pair according to the target prompt word, the preset inference thought chain, and the preset inference result;

[0008] Train the large language model to be trained according to the first target data pair.

[0009] In this way, the server determines an associated application interface list based on the original training data. The original training data includes an original voice request, a preset inference thinking chain corresponding to the original voice request, and a preset inference result. Based on the preset inference thinking chain and the preset inference result, a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated. Then, based on a preset prompt template, the server determines a target prompt based on the associated application interface list and the original voice request. Next, the server determines a first target data pair based on the target prompt, the preset inference thinking chain, and the preset inference result. Finally, the server trains the large language model to be trained based on the first target data pair. In this way, the original training data includes the original voice request, the corresponding preset inference thinking chain, and the preset inference result. Among them, based on the preset inference thinking chain and the preset inference result, a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated. During the training of the large language model, through the above-mentioned preset inference thinking chain and preset inference result, a first target data pair corresponding to the original voice request is generated. Furthermore, based on the first target data pair, a constraint-guided training is performed on the large language model, so that the large language model can perform reasoning on voice requests that cannot generate correctly executable vehicle control instructions before training and obtain correct and executable vehicle control instructions, thereby improving the reasoning ability of the large language model and enhancing the user experience.

[0010] In some embodiments, the determining the associated application interface list according to the original training data includes:

[0011] Encoding the original voice request based on a preset model to determine a target vector;

[0012] Determining the associated application interface list from a pre-constructed application interface database according to the target vector, where the application interface database includes candidate application interfaces, application interface parameters to be filled corresponding to the candidate application interfaces, application interface descriptions of the candidate application interfaces, and interface parameter descriptions of the application interface parameters.

[0013] In this way, based on a preset model, the server encodes the original voice request to determine a target vector. Then, based on the target vector, the server determines an associated application interface list from a pre-constructed application interface database, where the application interface database includes candidate application interfaces, application interface parameters to be filled corresponding to the candidate application interfaces, application interface descriptions of the candidate application interfaces, and interface parameter descriptions of the application interface parameters. In this way, through encoding processing and vector retrieval, a similar application interface list related to the original voice request can be efficiently retrieved from the application interface database, thereby enhancing the inference ability of the large language model. Moreover, when adding new function interfaces to the large language model, only the candidate application interfaces and related parameter information need to be added to the application interface database, without modifying the structure of the large language model.

[0014] In some embodiments, determining the target prompt based on the preset prompt template, the associated application interface list, and the original voice request includes:

[0015] Performing splicing processing on the preset prompt template, the associated application interface list, and the original voice request to determine the target prompt.

[0016] In this way, the server performs splicing processing on the preset prompt template, the associated application interface list, and the original voice request to determine the target prompt. In this way, based on the preset prompt template, the associated application interface list and the original voice request are spliced to form a prompt containing multi-dimensional information, thereby providing a more complete decision-making basis for the large language model and guiding the inference direction of the large language model.

[0017] In some embodiments, determining the first target data pair according to the target prompt, the preset inference thinking chain, and the preset inference result includes:

[0018] Based on the large language model to be trained, according to the target prompt, the preset inference thinking chain, and the preset inference result, determining the first target voice request and the first negative sample sub-data in the first target data pair;

[0019] Based on the base large language model, according to the target prompt corresponding to the first target voice request, the preset inference thinking chain, and the preset inference result, determining the first positive sample sub-data in the first target data pair, where the parameter scale of the base large language model is larger than that of the large language model to be trained;

[0020] Determining the first target data pair according to the first positive sample sub-data and the first negative sample sub-data.

[0021] Thus, based on the large language model to be trained, the server determines the first target voice request and the first negative sample sub-data in the first target data pair according to the target prompt, the preset reasoning thought chain, and the preset reasoning result. Then, based on the base large language model, the server determines the first positive sample sub-data in the first target data pair according to the target prompt, the preset reasoning thought chain, and the preset reasoning result corresponding to the first target voice request, where the parameter scale of the base large language model is larger than that of the large language model to be trained. Finally, the server determines the first target data pair according to the first positive sample sub-data and the first negative sample sub-data. In this way, the first target data pair constructed by the first negative sample sub-data and the first positive sample sub-data can train the large language model to be trained, making the logical reasoning path of the large language model to be trained conform to the preset reasoning thought chain and improving the accuracy of the logical reasoning path and the reasoning result of the large language model to be trained.

[0022] In some embodiments, the determining, based on the large language model to be trained, of the first target voice request and the first negative sample sub-data in the first target data pair according to the target prompt, the preset reasoning thought chain, and the preset reasoning result includes:

[0023] Based on the large language model to be trained, determine a first reasoning result and a first thinking logic chain according to the target prompt;

[0024] According to the preset reasoning thought chain and the preset reasoning result, perform screening processing on the first reasoning result and the first thinking logic chain to determine a first error output result, where the first error output result includes a first error reasoning result and / or a first error reasoning logic chain;

[0025] Determine the original voice request corresponding to the first error output result as the first target voice request;

[0026] Determine the first target voice request and the first error output result as the first negative sample sub-data.

[0027] Thus, based on the large language model to be trained, the server determines the first inference result and the first thinking logic chain according to the target prompt. Then, the server screens and processes the first inference result and the first thinking logic chain according to the preset inference thinking chain and the preset inference result, and determines the first error output result, which includes the first incorrect preset inference result and / or the first incorrect inference logic chain. Then, the server determines the original voice request corresponding to the first error output result as the first target voice request. Finally, the server determines the first target voice request and the first error output result as the first negative sample sub-data. In this way, by determining the first error output result and the first target voice request corresponding to the first error output result as the first negative sample sub-data, the weak points of the large language model to be trained can be targeted for training, enhancing the inference ability of the large language model to be trained. In addition, by comparing the preset inference thinking chain and the preset inference result, the specific breakpoints in the inference process of the large language model to be trained can be located, eliminating the need for manual annotation and reducing the data annotation cost.

[0028] In some embodiments, the determining the first positive sample sub-data in the first target data pair based on the base large language model, according to the target prompt corresponding to the first target voice request, the preset inference thinking chain, and the preset inference result, includes:

[0029] Based on the base large language model, perform natural language processing on the first target voice request according to the target prompt to determine a second inference result and a second thinking logic chain;

[0030] When both the second inference result and the second thinking logic chain are correct, determine the second inference result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data.

[0031] Thus, based on the base large language model, the server performs natural language processing on the first target voice request according to the target prompt to determine a second inference result and a second thinking logic chain. Then, when both the second inference result and the second thinking logic chain are correct, the server determines the second inference result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data. In this way, only when both the generated second inference result and the second logic chain pass the verification are they marked as positive samples, ensuring data reliability. Moreover, by comparing the above-generated first negative sample sub-data and first positive sample sub-data, the large language model to be trained can distinguish between correct and incorrect inference paths, enhancing the inference ability of the large language model to be trained.

[0032] In some embodiments, the training the large language model to be trained according to the first target data pair includes:

[0033] Based on a preset loss function, perform direct preference optimization training on the large language model to be trained according to the first target data pair.

[0034] In this way, based on the preset loss function, the server performs direct preference optimization training on the large language model to be trained according to the first target data pair. Thus, by performing direct preference optimization training based on the preset loss function and using the first target data, the large language model to be trained can learn how to accurately understand and process voice requests issued by users in the vehicle cockpit scenario and generate corresponding vehicle control instructions, thereby improving the user experience.

[0035] In some embodiments, the method further includes:

[0036] Based on the optimized large language model obtained through the preference optimization training, perform natural language processing on the original training data according to the target prompt word, and determine a third inference result and a third thinking logic chain corresponding to the third inference result;

[0037] According to the preset inference thinking chain and the preset inference result, perform screening processing on the third inference result and the third thinking logic chain to determine a second error output result, where the second error output result includes a second error inference result and / or a second error inference logic chain;

[0038] Determine the original voice request corresponding to the second error output result as the second target voice request;

[0039] Determine the second target voice request and the second error output result as the second negative sample sub-data;

[0040] Based on the optimized large language model, perform secondary natural language processing on the second target voice request according to the target prompt word to obtain a fourth inference result and a fourth thinking logic chain;

[0041] When both the fourth inference result and the fourth thinking logic chain are correct, determine the fourth inference result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub-data;

[0042] Determine the second target data pair according to the second positive sample sub-data and the second negative sample sub-data;

[0043] Based on the preset loss function, perform direct preference optimization training on the optimized large language model according to the second target data pair.

[0044] Thus, based on the optimized large language model obtained through preference-optimized training, the server performs natural language processing on the original training data according to the target prompt word to determine the third inference result and the third thinking logic chain corresponding to the third inference result. Then, the server performs screening processing on the third inference result and the third thinking logic chain according to the preset inference thinking chain and the preset inference result to determine the second error output result, which includes the second incorrect inference result and / or the second incorrect inference logic chain. Then, the server determines the original voice request corresponding to the second error output result as the second target voice request. Subsequently, the server determines the second target voice request and the second error output result as the second negative sample sub-data. Then, based on the optimized large language model, the server performs secondary natural language processing on the second target voice request according to the target prompt word to obtain the fourth inference result and the fourth thinking logic chain. Subsequently, when both the fourth inference result and the fourth thinking logic chain are correct, the server determines the fourth inference result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub-data. Then, the server determines the second target data pair according to the second positive sample sub-data and the second negative sample sub-data. Finally, based on the preset loss function, the server performs direct preference optimization training on the optimized large language model according to the second target data pair. In this way, by using the preset loss function and performing direct preference optimization training with the second target data obtained based on the optimized large language model, the large language model to be trained can learn how to accurately understand and process the voice requests issued by users in the vehicle cockpit scenario and generate corresponding vehicle control instructions, thereby improving the user experience. Moreover, by continuously detecting and correcting the errors of the model and using the second target data for direct preference optimization training again, the performance of the large language model is continuously optimized and the user experience is improved.

[0045] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.

[0046] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0047] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0049] Figure 1 is one of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0050] Figure 2 is the second of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0051] Figure 3 is the third of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0052] Figure 4 is the fourth of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0053] Figure 5 is the fifth of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0054] Figure 6 is the sixth of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0055] Figure 7 is the seventh of the flow diagrams of the model training method of the large language model in some embodiments of the present application;

[0056] Figure 8 is the flow diagram of the model training process in some embodiments of the present application;

[0057] Figure 9 is the eighth of the flow diagrams of the model training method of the large language model in some embodiments of the present application. Detailed Embodiments

[0058] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.

[0059] In the voice interaction scenario of a vehicle cockpit, the existing technology usually adopts an end-to-end large language model (LLM). When a user issues a voice request, the large language model can quickly recognize the voice and convert it into text, and then based on this text information, output a preset inference result that meets the user's needs, and further generate corresponding vehicle control instructions that meet the user's needs. For example, when a user issues a composite instruction such as "open the window and set the air conditioner to 25°C", the ASR module translates the voice into text and inputs it into the LLM, and the model directly outputs structured instructions such as "<window control, opening degree = 10101%>; <air conditioner control, mode = cooling, temperature = 25°C>" based on semantic understanding.

[0060] However, when the original voice request involves unclear instructions, local accents, non-standard expressions, or knowledge in a professional field, the large language model may have difficulty accurately capturing the user's intention due to parsing deviations, and then the generated inference logic cannot be effectively converted into correct vehicle control instructions. In this way, the vehicle may not be able to perform operations as expected by the user, which not only reduces the fluency of the human-machine interaction experience but also may reduce the user's trust in the vehicle intelligent system. For example, when a user says "it's too hot inside the car, lower the temperature inside the car", there are multiple ways to lower the temperature, including turning on the air conditioner, opening the window, and turning on the seat ventilation, etc. The large language model cannot determine which way to use. In this way, the user may not be satisfied with the inference result generated by the large language model.

[0061] Based on the above problems, please refer to Figure 1 , an embodiment of the present application provides a model training method for a large language model, and the method includes:

[0062] 011: Determine an associated application interface list according to the original training data;

[0063] 012: Based on a preset prompt word template, determine a target prompt word according to the associated application interface list and the original voice request;

[0064] 013: Determine a first target data pair according to the target prompt word, a preset inference thinking chain, and a preset inference result;

[0065] 014: Train the large language model to be trained according to the first target data pair.

[0066] The embodiments of the present application also provide a server, including a memory and a processor. The method for training the speech recognition model according to the embodiments of the present application can be implemented by the server according to the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is used to determine an associated application interface list according to the original training data. And based on a preset prompt word template, according to the associated application interface list and the original voice request, determine the target prompt word. The processor is used to determine the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result. And according to the first target data pair, train the large language model to be trained.

[0067] The embodiments of the present application also provide a model training device. The method for training the speech recognition model according to the embodiments of the present application can be implemented by the model training device according to the embodiments of the present application. Specifically, the model training device includes a determination module and a training module. The determination module is used to determine an associated application interface list according to the original training data. And based on a preset prompt word template, according to the associated application interface list and the original voice request, determine the target prompt word. And according to the target prompt word, the preset inference thinking chain, and the preset inference result, determine the first target data pair. The training module is used to train the large language model to be trained according to the first target data pair.

[0068] Specifically, the original training data refers to the basic data set used for training the large language model, including the original voice request, the preset inference thinking chain, and the preset inference result.

[0069] Among them, the original voice request refers to the instructions sent by the user in a voice manner collected from the log files in the large language model. For example, "turn on the air conditioner", "navigate to the company", "play music", and "adjust the seat position forward by 30%", etc. It should be noted that the original voice request needs to be processed by speech recognition to convert the voice signal sent by the user into text form, that is, to recognize the words, phrases, and sentence structures in the voice and convert them into text for subsequent processing.

[0070] The preset inference thinking chain refers to a series of logical reasoning steps from the original voice request to the final vehicle control instruction for each original voice request, which can guide the complete inference process of the large language model from semantic understanding to instruction generation. In some embodiments, the preset inference thinking chain is obtained by pre-disassembling through manual or rule engines, which can meet the needs and preferences of users.

[0071] The preset inference result refers to a standardized instruction corresponding to the preset inference thinking chain, which can meet the needs and preferences of users and can be directly executed by the vehicle electronic control unit, usually represented in structured data (such as JSON or XML), and is used to control the vehicle to perform specific operations.

[0072] The list of similar application interfaces refers to the executable application interfaces retrieved that are related to the original voice request, including similar application interfaces, the corresponding similar application interface parameters to be filled for the similar application interfaces, the descriptions of the similar application interfaces for the similar application interfaces, and the descriptions of the similar interface parameters for the similar application interface parameters. For example, if the user's voice request is "Adjust the seat position forward by 30%", then the list of similar application interfaces may include ["Similar application interface D: ControlOpen", "Description of similar application interface d1: Open relevant in-vehicle devices or functions, and can control multiple devices and various functions in the vehicle", "Similar application interface parameter d2: device, function, target_function, and seat_position", and "Similar application interface parameter d3: device (string, optional), used to specify the in-vehicle device, and is only provided when it is necessary to specify a specific function within the device; function (string, optional), used to specify the function related to the in-vehicle device, and is only provided when it is necessary to specify a specific function option within the device; target_function (string, optional), used to specify different modes or different functions of the function related to the in-vehicle device, and is only provided when it is necessary to specify a specific function option within the device; seat_position (enumeration value, optional), used to specify the seat position in the vehicle, and the value can be selected from one of ['Driver's seat', 'Passenger seat', 'Front row', 'Left side of the row', 'Right side of the row', 'Row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row']"].

[0073] It should be noted that according to the original training data, it is determined that the list of similar application interfaces is implemented by a Retrieval-Augmented Generation (RAG) model.

[0074] The target prompt refers to a guiding input dynamically generated based on a preset prompt template, the associated application interface list, and the original voice request, which can provide additional context information to help the large language model more accurately infer the user's intention from the original voice request and generate corresponding vehicle control instructions.

[0075] The first target data pair refers to the training sample pair generated based on the target prompt, the preset reasoning thought chain, and the preset reasoning result, which is used to supervise the large language model to learn the deterministic reasoning path from the original voice request to the executable instruction, ensuring that the result inferred and generated by the large language model simultaneously satisfies logical correctness (i.e., satisfies the preset reasoning thought chain) and physical executability (i.e., satisfies the preset reasoning result).

[0076] The large language model to be trained refers to a large language model that has been pre-trained on general data and needs to optimize parameters through specific domain data to adapt to vertical scenarios. For the convenience of description, in the embodiments of this application, the large language model is used to refer to the large language model to be trained.

[0077] Based on the RAG model, the server first determines a list of associated application interfaces related to the original voice request according to the original voice request in the original training data. Among them, the original training data also includes a preset inference thinking chain and a preset inference result corresponding to the original voice request. The preset inference thinking chain analyzes the user's intention through a logical deduction path, and the preset inference result is a vehicle control instruction that conforms to the vehicle execution specification generated according to this path.

[0078] Next, according to the preset prompt template, the list of associated application interfaces and the original voice request are fused to construct a target prompt.

[0079] Subsequently, the server determines the first target data pair according to the target prompt, the preset inference thinking chain, and the preset inference result.

[0080] Finally, the server adopts the Direct Preference Optimization (DPO) framework to optimize the parameters of the large language model according to the first target data pair and the relevant contrast loss function, so that the large language model can output a more accurate vehicle control instruction preference standard.

[0081] In summary, in the model training method and server of the large language model provided by the embodiment of the present application, the server determines an associated application interface list according to the original training data, where the original training data includes an original voice request, a preset inference thinking chain corresponding to the original voice request, and a preset inference result. Based on the preset inference thinking chain and the preset inference result, a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated. Then, based on the preset prompt word template, the server determines a target prompt word according to the associated application interface list and the original voice request. Then, the server determines a first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result. Finally, the server trains the large language model to be trained according to the first target data pair. In this way, the original training data includes the original voice request, the corresponding preset inference thinking chain, and the preset inference result, where based on the preset inference thinking chain and the preset inference result, a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated. During the training of the large language model, through the above-mentioned preset inference thinking chain and preset inference result, a first target data pair corresponding to the original voice request is generated, and then the large language model is subjected to constraint-guided training according to the first target data pair, so that the large language model can reason about voice requests that cannot generate vehicle control instructions that can be correctly executed before training and obtain correct and executable vehicle control instructions, thereby improving the inference ability of the large language model and improving the user experience.

[0082] Please refer to Figure 2 , in some embodiments, step 011 (determining an associated application interface list according to the original training data) includes:

[0083] 0111: Based on a preset model, encode the original voice request to determine a target vector;

[0084] 0112: According to the target vector, determine an associated application interface list from a pre-constructed application interface database.

[0085] In some embodiments, the determining module is further configured to encode the original voice request based on a preset model to determine a target vector. And according to the target vector, determine an associated application interface list from a pre-constructed application interface database.

[0086] In some embodiments, the processor is further configured to encode the original voice request based on a preset model to determine a target vector. And according to the target vector, determine an associated application interface list from a pre-constructed application interface database.

[0087] Specifically, the pre-selection model refers to the RAG model, which can retrieve a list of associated application interfaces similar to the original voice request from a pre-set application interface database.

[0088] Encoding processing refers to converting the target voice request training data into vectors so that it can be understood and processed by machine learning models. In some embodiments, the encoding processing of voice requests is implemented based on the BGE model, which refers to an embedding model based on a graph structure and can convert the original voice request into a vector representation and retrieve application interface description vectors semantically similar to the user instruction from the application interface database. In addition, methods such as word embedding and sequence encoding can also be used to convert text into high-dimensional vectors and retain the semantic information of the text.

[0089] The target vector refers to the vector representation corresponding to the original voice request after encoding processing, which can be used as the basis for retrieving the list of associated application interfaces, calculate the similarity with the interface parameter descriptions in the pre-constructed application interface database, select similar candidate application interfaces, and thus determine the list of associated application interfaces. In this way, retrieval based on vector similarity can accurately match candidate application interfaces related to the original voice request and avoid mis-matching.

[0090] Candidate application interfaces refer to the application interfaces pre-stored in the application interface database and can be called in the vehicle, such as "Media_Voice_Open" and "Headrest_Voice_Model_Open", etc.

[0091] Application interface description refers to the structured metadata definition of candidate application interfaces, which is used to clarify the functions, parameter rules, constraint conditions, and usage scenarios of candidate application interfaces. That is, the application interface description details the functions, uses, and usage methods of application interfaces, and can help large language models understand the operations that each application interface can perform and how these operations correspond to the user's voice request.

[0092] Application interface parameters refer to the parameter values that must or can be optionally provided corresponding to candidate application interfaces. Application interface parameters determine the specific behavior of application interface calls, such as target_function and seat_position, etc.

[0093] Interface parameter description refers to the metadata description of each interface parameter, including data type, value range, whether it is mandatory, dependency relationship, etc. The interface parameter description is a detailed explanation of the application interface parameters, which explains the meaning, usage, and possible value range of each interface parameter, and can help the large language model learn how to select appropriate parameter values according to the content of the voice request and generate correct instructions that can be executed by the vehicle. For example, "target_function: (enumeration value, mandatory) values are ['sound']" or "seat_position: (enumeration value, optional) values can be selected from ['driver's seat', 'passenger seat', 'front row', 'left side of the second row', 'right side of the second row', 'second row', 'all', 'left side of the third row', 'right side of the third row', 'third row']".

[0094] The application interface database refers to a database that stores application interfaces and their corresponding parameters, and can provide a data basis for retrieving similar application interfaces, including candidate application interfaces, application interface parameters to be filled corresponding to the candidate application interfaces, application interface descriptions of the candidate application interfaces, and interface parameter descriptions of the application interface parameters. The data information that the application interface database may include is ["Similar application interface D: ControlOpen", "Similar application interface parameter d1: device, function, target_function, and seat_position", "Similar application interface description d2: Open relevant in-vehicle devices or functions, and can control multiple devices and various functions in the vehicle", "Similar interface parameter description d3: device (string, optional) is used to specify the in-vehicle device, and is provided only when a specific function within the device needs to be specified; function (string, optional) is used to specify the function related to the in-vehicle device, and is provided only when a specific function option within the device needs to be specified; target_function (string, optional) is used to specify different modes or different functions of the function related to the in-vehicle device, and is provided only when a specific function option within the device needs to be specified; seat_position (enumeration value, optional) is used to specify the seat position in the vehicle, and the value can be selected from one of ['Driver', 'Passenger', 'Front row', 'Left side of the row', 'Right side of the row', 'Row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row']"], ["Similar application interface E: Media_Voice_Open", "Similar application interface parameter e1: target_function and seat_position", "Similar application interface description e2: MediaVoiceOpen turns on the sound or unmutes, where MediaVoice represents the media sound, and in-vehicle media such as music, video, radio, etc. can all be understood as sound, which is different from MediaEffect", "Similar interface parameter description e3: target_function (enumeration value, required), the value is ['Sound']; seat_position (enumeration value, optional), the value can be selected from ['Driver', 'Passenger', 'Front row', 'Left side of the second row', 'Right side of the second row', 'Second row', 'All', 'Left side of the third row', 'Right side of the third row', 'Third row']"], and ["Similar application interface F: Headrest_Voice_Model_Open", "Similar application interface parameter f1: target_function, function, seat_position, and device", "Similar application interface description f2: HeadrestVoiceModelOpen turns on the in-vehicle headrest audio or the driver's audio mode, where HeadrestVoice represents the headrest audio and Model represents the mode."The functions of broadcasting and voice are set to private and driving enjoyment, which means the headrest speakers are set to a certain audio mode." "Similar interface parameter description f3: target_function (enumeration value, required), the optional values are ['Intelligent Mode', 'Full Vehicle Playback Mode', 'Standard Mode', 'Driving Enjoyment Mode', 'Private Enjoyment Mode']; seat_position (enumeration value, optional), the value is ['Driver's Seat'], this parameter only supports the driver's seat, and when no relevant information such as the driver's seat is mentioned, it is not output; device (enumeration value, required), the value is ['Headrest Speakers']".

[0095] By retrieving similar application interfaces, the RAG model can help large language models reduce inference pressure, enabling large language models to more efficiently handle natural language understanding tasks in complex scenarios. Moreover, the introduction of the RAG model can reduce the pressure of iterative training of large language models, thereby reducing training costs and human input. In addition, the RAG model can assist large language models in handling unseen tasks and enhance the generalization ability of large language models.

[0096] First, encode the user's voice request to determine the target vector. Then, calculate the similarity between the target vector and the application interface descriptions of candidate application interfaces in the pre-constructed application interface database to obtain a list of similar application interfaces.

[0097] In this way, based on a preset model, the server encodes the original voice request to determine the target vector. Then, the server determines an associated application interface list from the pre-constructed application interface database according to the target vector. The application interface database includes candidate application interfaces, application interface parameters to be filled corresponding to the candidate application interfaces, application interface descriptions of the candidate application interfaces, and interface parameter descriptions of the application interface parameters. Thus, through encoding processing and vector retrieval, a list of similar application interfaces related to the original voice request can be efficiently retrieved from the application interface database, thereby enhancing the inference ability of the large language model. And when adding new function interfaces to the large language model, only need to add candidate application interfaces and related parameter information to the application interface database, without modifying the structure of the large language model.

[0098] Please refer to Figure 3 , in some embodiments, step 012 (based on a preset prompt template, determine the target prompt according to the associated application interface list and the original voice request), includes:

[0099] 0121: Concatenate and process the preset prompt template, the associated application interface list, and the original voice request to determine the target prompt.

[0100] In some embodiments, the determination module is further configured to splice and process a preset prompt template, an associated application interface list, and an original voice request to determine a target prompt.

[0101] In some embodiments, the processor is further configured to splice and process a preset prompt template, an associated application interface list, and an original voice request to determine a target prompt.

[0102] Specifically, the preset prompt template refers to a pre-defined text template in natural language processing tasks, which is used to guide the large language model to generate or understand text. In some embodiments, the preset prompt template may be "You are a vehicle voice assistant. According to the user's instructions, combined with the APIs and their parameters given in the

associated application interface list

vehicle knowledge

[0103] 1. First, you should think according to the user's instructions and

vehicle knowledge

[0104] 2. According to the thinking logic, judge whether the following APIs can help you complete the task. If there is no API that can help you complete the task, then output "none".

[0105] 3. If there is a suitable API in the list that can be used, then please first select one of them, and then fill in the parameters of the selected API in turn.

[0106] 4. Note that if the parameter is a required parameter, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values.

[0107]

Your output format

[0108] Thinking logic:

[0109] 1. Understand the user's instructions:

[0110] 2. Refer to the vehicle knowledge:

[0111] 3. Select a suitable API:

[0112] 4. Fill in the API parameters:

[0113] Is there a suitable API: yes or no

[0114] Selected API: {"API": "api1", "ARGUMENTS": {"arg1": "value1", "arg2": "value2",...,"argn": "valuen"}}

[0115] <Vehicle knowledge>

[0116] <List of associated application interfaces>

[0117] <User instruction>. Among them, in some embodiments, the preset prompt word template may not contain content related to <Vehicle knowledge>.

[0118] The voice request, the list of similar voice requests, and the list of similar application interfaces are spliced according to a preset prompt word template to generate a complete text, that is, the target prompt word. For example, if the user's voice request is "Cancel mute", then the target prompt word may be "You are a vehicle-mounted voice assistant. According to the user's instruction, combined with the APIs and their parameters given in the

Candidate list of APIs

Vehicle knowledge

[0119] 1. First, you should think based on the user instruction and

Vehicle knowledge

[0120] 2. Judge whether the following APIs can help you complete the task according to the thinking logic. If there is no API that can help you complete the task, then output "None"

[0121] 3. If there is a suitable API in the list that can be used, then please first select one of them, and then fill in the parameters of the selected API in turn

[0122] 4. Note that if it is a required parameter, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values

[0123]

Your output format

[0124] Thinking logic:

[0125] 1. Understand the user instruction:

[0126] 2. Refer to the vehicle knowledge:

[0127] 3. Select a suitable API:

[0128] 4. Fill in the API parameters:

[0129] Is there a suitable API: Yes or No

[0130] Selected API: {"API": "api1", "ARGUMENTS": {"arg1": "value1", "arg2": "value2",...,"argn": "valuen"}}

[0131]

Vehicle knowledge

[0132] Sharing mode = Full vehicle playback = Standard mode

[0133]

List of available APIs

[0134]

Media_Voice_Open

[0135] Name: Media_Voice_Open

[0136] Description: MediaVoiceOpen turns on the sound or unmutes, where MediaVoice represents the media sound. In-vehicle media such as music, video, radio, etc. can all be understood as sound, which is different from MediaEffect.

[0137] ARGUMENTS:

[0138] "target_function": (enumeration value, required) The value is ['sound']

[0139] "seat_position": (enumeration value, optional) The optional values are ['driver's seat', 'passenger seat', 'front row', 'left side of the second row', 'right side of the second row', 'second row', 'all', 'left side of the third row', 'right side of the third row', 'third row']

[0140]

Headrest_Voice_Model_Open

[0141] Name: Headrest_Voice_Model_Open

[0142] Description: HeadrestVoiceModelOpen turns on the in-vehicle headrest audio or the driver's seat audio mode, where HeadrestVoice represents the headrest audio and Model represents the mode. Broadcasting, setting the sound to private or driving enjoyment means setting the headrest audio to a certain audio mode.

[0143] ARGUMENTS:

[0144] "target_function": (enumeration value, required) The optional values are ['intelligent mode', 'full vehicle playback mode','standard mode', 'driving enjoyment mode', 'private enjoyment mode']

[0145] "seat_position": (enumeration value, optional) The value is ['driver's seat']. This parameter only supports the driver's seat and will not be output when no relevant information such as the driver's seat is mentioned.

[0146] "device": (enumeration value, required) The value is ['headrest audio']

[0147] <|im_start|>user

[0148] Unmute

[0149] <|im_end|>”.

[0150] In this way, the server splices and processes the preset prompt word template, the associated application interface list, and the original voice request to determine the target prompt word. In this way, based on the preset prompt word template, the associated application interface list and the original voice request are spliced to form a prompt word containing multi-dimensional information, thereby providing a more complete decision-making basis for the large language model and guiding the inference direction of the large language model.

[0151] Please refer to Figure 4 , in some embodiments, step 013 (determining the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result) includes:

[0152] 0131: Based on the large language model to be trained, determine the first target voice request and the first negative sample sub-data in the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result;

[0153] 0132: Based on the base large language model, determine the first positive sample sub-data in the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result corresponding to the first target voice request;

[0154] 0133: Determine the first target data pair according to the first positive sample sub-data and the first negative sample sub-data.

[0155] In some embodiments, the determination module is further configured to determine the first target voice request and the first negative sample sub-data in the first target data pair based on the large language model to be trained according to the target prompt word, the preset inference thinking chain, and the preset inference result. And based on the base large language model, determine the first positive sample sub-data in the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result corresponding to the first target voice request. And determine the first target data pair according to the first positive sample sub-data and the first negative sample sub-data.

[0156] In some embodiments, the processor is further configured to determine the first target voice request and the first negative sample sub-data in the first target data pair based on the large language model to be trained according to the target prompt word, the preset inference thinking chain, and the preset inference result. And based on the base large language model, determine the first positive sample sub-data in the first target data pair according to the target prompt word, the preset inference thinking chain, and the preset inference result corresponding to the first target voice request. And determine the first target data pair according to the first positive sample sub-data and the first negative sample sub-data.

[0157] Specifically, the first target voice request refers to the voice request that can be used to train the inference ability of the large language model determined after preliminary processing of the original voice request during the model training process. It should be noted that only when there is a deviation between the output (the first inference result and the first thinking logic chain) of the large language model to be trained and the preset standards (the preset inference thinking chain and the preset inference result), the corresponding original voice request will be marked as the first target voice request.

[0158] The first negative sample sub-data refers to the results where the first inference result and the first thinking logic chain generated by the large language model to be trained do not meet the preset standards, including logical errors, format errors, or potential safety hazards in the output. Not meeting the preset standards means that the first thinking logic chain generated by the large language model to be trained does not conform to the preset inference thinking chain, or the first inference result generated does not conform to the preset inference result.

[0159] The base large language model refers to a pre-trained large language model with a larger parameter scale and stronger performance compared to the large language model to be trained, which can be used as a "teacher model" to generate high-quality first positive sample sub-data to guide the large language model to be trained to learn.

[0160] The first positive sample sub-data refers to the results that meet the expected standards generated by the base large language model, which can be used as the "standard answer" for the large language model to be trained to learn.

[0161] First, determine whether the inference results (including the first inference result and the first thinking logic chain) generated by the large language model to be trained meet the preset standards. If not, mark the corresponding original voice request as the first target voice request and use the generated inference results as the first negative sample sub-data.

[0162] Subsequently, relying on the base large language model with a more massive parameter scale, according to the target prompt word, the preset inference thinking chain, and the preset inference result corresponding to the first target voice request, generate the first positive sample sub-data, which meets the preset standards and serves as the standard answer for the large language model to be trained to learn.

[0163] Finally, combine the first positive sample sub-data and the first negative sample sub-data to form the first target data pair. The first target data pair can be used to train the large language model to be trained, helping the large language model to be trained learn how to extract intentions from voice requests and generate correct vehicle control instructions based on these intentions.

[0164] Thus, based on the large language model to be trained, the server determines the first negative sample sub-data in the first target voice request and the first target data pair according to the target prompt, the preset reasoning thought chain, and the preset reasoning result. Then, based on the base large language model, the server determines the first positive sample sub-data in the first target data pair according to the target prompt, the preset reasoning thought chain, and the preset reasoning result corresponding to the first target voice request, where the parameter scale of the base large language model is larger than that of the large language model to be trained. Finally, the server determines the first target data pair according to the first positive sample sub-data and the first negative sample sub-data. In this way, the first target data pair constructed by the first negative sample sub-data and the first positive sample sub-data can train the large language model to be trained, making the logical reasoning path of the large language model to be trained conform to the preset reasoning thought chain and improving the accuracy of the logical reasoning path and reasoning result of the large language model to be trained.

[0165] Please refer to Figure 5 , in some embodiments, step 0131 (based on the large language model to be trained, according to the target prompt, the preset reasoning thought chain, and the preset reasoning result, determining the first negative sample sub-data in the first target voice request and the first target data pair) includes:

[0166] 01311: Based on the large language model to be trained, according to the target prompt, determine the first reasoning result and the first thinking logic chain;

[0167] 01312: According to the preset reasoning thought chain and the preset reasoning result, perform screening processing on the first reasoning result and the first thinking logic chain to determine the first incorrect output result;

[0168] 01313: Determine the original voice request corresponding to the first incorrect output result as the first target voice request;

[0169] 01314: Determine the first target voice request and the first incorrect output result as the first negative sample sub-data.

[0170] In some embodiments, the determination module is further configured to, based on the large language model to be trained, according to the target prompt, determine the first reasoning result and the first thinking logic chain. And according to the preset reasoning thought chain and the preset reasoning result, perform screening processing on the first reasoning result and the first thinking logic chain to determine the first incorrect output result. The determination module is further configured to determine the original voice request corresponding to the first incorrect output result as the first target voice request. And determine the first target voice request and the first incorrect output result as the first negative sample sub-data.

[0171] In some embodiments, the processor is further configured to determine a first inference result and a first thinking logic chain based on the large language model to be trained and according to the target prompt. And screen the first inference result and the first thinking logic chain according to the preset inference thinking chain and the preset inference result to determine a first error output result. The processor is further configured to determine the original voice request corresponding to the first error output result as the first target voice request. And determine the first target voice request and the first error output result as the first negative sample sub-data.

[0172] Specifically, the first inference result and the first thinking logic chain refer to the natural language processing results and the reasoning process obtained by guiding the large language model to be trained to perform natural language processing on the original voice request based on the target prompt.

[0173] The screening process refers to comparing the model output (i.e., the first inference result and the first thinking logic chain) with the standard answer (i.e., the preset inference thinking chain and the preset inference result) through preset rules to identify incorrect inference paths. In some embodiments, the preset rules include logical chain consistency and inference result compliance. Among them, logical chain consistency refers to whether the first thinking logic chain generated by the large language model to be trained completely covers the key nodes of the preset thinking chain. For example, if the logic of the preset thinking chain is "parse instruction → confirm device → calculate parameter → verify range → generate application interface", and the first thinking logic chain is "parse instruction → directly generate application interface", then it is considered that the preset thinking chain and the first thinking logic chain are inconsistent in terms of logical chain. Inference result compliance refers to verifying whether the first inference result conforms to the preset inference result. For example, if the output format of the preset inference result is JSON and the output format of the first inference result is XML, then it is considered that the first inference result does not conform to the preset inference result.

[0174] The first error output result refers to the content generated by the large language model that is determined to deviate from the preset standard (i.e., the preset inference thinking chain and the preset inference result) after the screening process, including the first incorrect preset inference result or the first incorrect inference logic chain.

[0175] First, based on the target prompt, guide the large language model to be trained to perform natural language processing on the original voice request to generate a first inference result and a first thinking logic chain.

[0176] Next, screen the generated first preset result and the first thinking logic chain respectively with the preset inference thinking chain and the preset inference result. If there is a deviation between the results generated by the large language model to be trained (the first preset result and the first thinking logic chain) and the preset standard (the preset inference thinking chain and the preset inference result), then mark these results as the first error output result, and the first error output result includes incorrect inference results and / or incorrect inference logic chains.

[0177] Subsequently, mark the original voice request corresponding to the first error output result as the first target voice request.

[0178] Finally, combine the first target voice request and the first error output result to form the first negative sample sub-data.

[0179] For example, the original voice request is "Unmute", and the preset reasoning thought chain and preset reasoning result are "Thought logic chain:

[0180] 1. Understand the user's instruction: The user hopes not to be muted, that is, to turn on the media sound.

[0181] 2. Refer to vehicle knowledge: Vehicle knowledge does not directly involve sound control, but according to the user's instruction, relevant APIs need to be used to turn on the sound.

[0182] 3. Select the appropriate API: The Media_Voice_Open API allows turning on the sound or unmuting, which meets the user's needs.

[0183] 4. Fill in the API parameters:

[0184] - target_function: sound, because the user hopes to turn on the sound.

[0185] - seat_position: Since the user did not specify a specific seat, this parameter can be omitted.

[0186] Is there a suitable API: Yes

[0187] Selected API: {"API": "Media_Voice_Open", "ARGUMENTS": {"target_function": "sound"}}”.

[0188] The first reasoning result and the first thought logic chain to be generated by the large language model to be trained may be "Thought logic chain:

[0189] 1. Understand the user's instruction: The user hopes to unmute.

[0190] 2. Refer to vehicle knowledge: Sharing mode = full vehicle playback = standard mode, indicating that the current vehicle is set to the standard mode of full vehicle playback.

[0191] 3. Select the appropriate API: According to the user's need to unmute, the Media_Voice_Open API can meet the need because it can control the mute state of the in-vehicle media sound.

[0192] 4. Fill in API parameters: The user hopes to unmute, so they select 'Sound' for target_function and 'All' for seat_position.

[0193] Is there a suitable API: Yes

[0194] Selected API: {"API":"Media_Voice_Open","ARGUMENTS":{"target_function":"Sound","seat_position":"All"}}.

[0195] Then, by comparing the first inference result and the first thinking logic chain generated by the large language model to be trained with the preset inference thinking chain and the preset inference result respectively, it can be found that the logic chains are inconsistent and the inference results are non-compliant.

[0196] Therefore, the original voice request 'Unmute' is determined as the first target voice request, and the above first inference result, the above first thinking logic chain, and the first target voice request are determined as the first negative sample sub-data.

[0197] In this way, based on the large language model to be trained, the server determines the first inference result and the first thinking logic chain according to the target prompt word. Then, the server screens and processes the first inference result and the first thinking logic chain according to the preset inference thinking chain and the preset inference result to determine the first error output result, which includes the first incorrect preset inference result and / or the first incorrect inference logic chain. Then, the server determines the original voice request corresponding to the first error output result as the first target voice request. Finally, the server determines the first target voice request and the first error output result as the first negative sample sub-data. In this way, by determining the first error output result and the first target voice request corresponding to the first error output result as the first negative sample sub-data, it is possible to conduct targeted training on the weak points of the large language model to be trained and enhance the inference ability of the large language model to be trained. In addition, by comparing the preset inference thinking chain and the preset inference result, it is possible to locate the specific breakpoints in the inference process of the large language model to be trained, without the need for manual annotation, which can reduce the data annotation cost.

[0198] Please refer to Figure 6 , in some embodiments, step 0132 (based on the base large language model, according to the target prompt word, the preset inference thinking chain, and the preset inference result corresponding to the first target voice request, determine the first positive sample sub-data in the first target data pair), includes:

[0199] 01321: Based on the base large language model, according to the target prompt, perform natural language processing on the first target voice request to determine the second reasoning result and the second thinking logic chain;

[0200] 01322: When both the second reasoning result and the second thinking logic chain are correct, determine the second reasoning result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data.

[0201] In some embodiments, the determination module is further configured to perform natural language processing on the first target voice request based on the base large language model according to the target prompt to determine the second reasoning result and the second thinking logic chain. And when both the second reasoning result and the second thinking logic chain are correct, determine the second reasoning result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data.

[0202] In some embodiments, the processor is further configured to perform natural language processing on the first target voice request based on the base large language model according to the target prompt to determine the second reasoning result and the second thinking logic chain. And when both the second reasoning result and the second thinking logic chain are correct, determine the second reasoning result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data.

[0203] Specifically, based on the target prompt, guide the base large language model to perform natural language processing on the first target voice request to generate the second reasoning result and the second thinking logic chain.

[0204] Next, perform screening processing on the second reasoning thinking chain and the second reasoning result respectively based on the preset reasoning thinking chain and the preset reasoning result. If the results generated by the base large language model (the second preset result and the second thinking logic chain) meet the preset criteria (the preset reasoning thinking chain and the preset reasoning result), then determine the second reasoning result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data.

[0205] For example, the original voice request is "Unmute", and the preset reasoning thinking chain and the preset reasoning result are "Thinking logic chain:

[0206] 1. Understand the user's instruction: The user hopes not to be muted, that is, to turn on the media sound.

[0207] 2. Refer to vehicle knowledge: Vehicle knowledge does not directly involve sound control, but according to the user's instruction, relevant APIs need to be used to turn on the sound.

[0208] 3. Select the appropriate API: The Media_Voice_Open API allows turning on the sound or unmuting, which meets the user's needs.

[0209] 4. Fill in the API parameters:

[0210] - target_function: Sound, because the user hopes to turn on the sound.

[0211] - seat_position: Since the user did not specify a specific seat, this parameter can be omitted.

[0212] Is there a suitable API: Yes

[0213] Selected API: {"API": "Media_Voice_Open", "ARGUMENTS": {"target_function": "Sound"}}”.

[0214] The second inference result and the second thinking logic chain generated by the base large language model may be "Thinking logic chain:

[0215] 1. Understand the user's instruction: The user hopes not to mute, that is, to turn on the media sound.

[0216] 2. Refer to vehicle knowledge: Vehicle knowledge does not directly involve sound control, but according to the user's instruction, relevant APIs need to be used to turn on the sound.

[0217] 3. Select a suitable API: The Media_Voice_Open API allows turning on the sound or unmuting, which meets the user's needs.

[0218] 4. Fill in the API parameters:

[0219] - target_function: Sound, because the user hopes to turn on the sound.

[0220] - seat_position: Since the user did not specify a specific seat, this parameter can be omitted.

[0221] Is there a suitable API: Yes

[0222] Selected API: {"API": "Media_Voice_Open", "ARGUMENTS": {"target_function": "Sound"}}”.

[0223] Then, comparing the second inference result and the second thinking logic chain generated by the base large language model with the preset inference thinking chain and the preset inference result respectively, it can be found that the logic chains are consistent and the inference results are compliant.

[0224] Therefore, the above second inference result, the above second thinking logic chain, and the first target voice request are determined as the first positive sample sub-data.

[0225] It should be noted that in some embodiments, if the base large language model cannot generate the correct second reasoning thought chain and second reasoning result either, then the correct second reasoning thought chain and second reasoning result are generated through manual intervention or by using a large language model with a larger parameter scale.

[0226] In this way, based on the base large language model, the server performs natural language processing on the first target voice request according to the target prompt word to determine the second reasoning result and the second thinking logic chain. Then, when both the second reasoning result and the second thinking logic chain are correct, the server determines the second reasoning result, the second thinking logic chain, and the first target voice request as the first positive sample sub-data. In this way, only when both the generated second reasoning result and the second logic chain pass the verification are they marked as positive samples, ensuring data reliability. And through the comparison of the first negative sample sub-data and the first positive sample sub-data generated above, the large language model to be trained can distinguish correct and incorrect reasoning paths, improving the reasoning ability of the large language model to be trained.

[0227] Please refer to Figure 7 , in some embodiments, step 014 (training the large language model to be trained according to the first target data pair) includes:

[0228] 0141: Based on a preset loss function, directly perform preference optimization training on the large language model to be trained according to the first target data pair.

[0229] In some embodiments, the model training module is also used to directly perform preference optimization training on the large language model to be trained based on a preset loss function according to the first target data pair.

[0230] In some embodiments, the processor is also used to directly perform preference optimization training on the large language model to be trained based on a preset loss function according to the first target data pair.

[0231] Specifically, the preset loss function refers to the DPO loss function, which can be used to measure the difference between the prediction of the large language model and the true value. By minimizing the loss function, the large language model can learn how to better predict or generate the required output. Specifically, the DPO loss function guides the learning process of the large language model to be trained by comparing the difference between the prediction of the large language model to be trained and the prediction of the base large language model. The calculation formula of the DPO loss function is: Among them, represents the expected value of the first target data pair (x, y w , y l ). X is the first target voice request, y W is the first positive sample sub-data, and y l is the first negative sample sub-data. First, calculate the prediction distribution π of the base large language model θ (y W |x) and the log-likelihood ratio between the reference distribution π ref (y W |x), as well as the log-likelihood ratio between the prediction distribution π of the large language model to be trained θ (y l |x) and the reference distribution π ref (y l |x). Then, these two log-likelihood ratios are weighted (with weight β) and subtracted, and the result is mapped through the activation function σ, where σ is the sigmoid function, mapping the result to the interval (0,1). In addition, β is a hyperparameter used to adjust the weights between different terms. By adjusting the value of β, the trade-off between privacy protection and performance of the model can be controlled.

[0232] Based on a preset loss function, the server performs direct preference optimization training on the large language model to be trained according to the first target data pair.

[0233] Please refer to Figure 8 , Figure 8 , which is a schematic diagram of the model training process of the large language model provided by the embodiment of the present application. Among them, after obtaining the original voice request, retrieve according to the RAG model to determine the associated application interface list. Subsequently, splice the original voice request and the associated application interface list to determine the target prompt word. And use the large language model to be trained and the base large language model to perform natural language processing on the first target voice request for which the large language model to be trained cannot infer the correct inference result or correct inference thought chain respectively, to determine the first target data pair. Furthermore, use this first target data pair to train the large language model to be trained.

[0234] In this way, based on a preset loss function, the server performs direct preference optimization training on the large language model to be trained according to the first target data pair. In this way, through direct preference optimization training based on a preset loss function and using the first target data, the large language model to be trained can learn how to accurately understand and process the voice requests sent by users in the vehicle cockpit scenario and generate corresponding vehicle control instructions, thereby improving the user experience.

[0235] Please refer to Figure 9 , in some embodiments, the method further includes:

[0236] 015: Based on the optimized large language model obtained through preference optimization training, perform natural language processing on the original training data according to the target prompt word to determine the third inference result and the third thinking logic chain corresponding to the third inference result;

[0237] 016: According to the preset reasoning thought chain and preset reasoning result, screen and process the third reasoning result and the third thinking logic chain to determine the second error output result;

[0238] 017: Determine the original voice request corresponding to the second error output result as the second target voice request;

[0239] 018: Determine the second target voice request and the second error output result as the second negative sample sub-data;

[0240] 019: Based on the optimized large language model, according to the target prompt word, perform secondary natural language processing on the second target voice request to obtain the fourth reasoning result and the fourth thinking logic chain;

[0241] 020: When both the fourth reasoning result and the fourth thinking logic chain are correct, determine the fourth reasoning result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub-data;

[0242] 021: Determine the second target data pair according to the second positive sample sub-data and the second negative sample sub-data;

[0243] 022: Based on the preset loss function, directly optimize and train the optimized large language model according to the second target data pair.

[0244] In some embodiments, the determination module is further configured to, based on the optimized large language model obtained through preference optimization training, perform natural language processing on the original training data according to the target prompt word to determine the third reasoning result and the third thinking logic chain corresponding to the third reasoning result. And according to the preset reasoning thought chain and preset reasoning result, screen and process the third reasoning result and the third thinking logic chain to determine the second error output result. And determine the original voice request corresponding to the second error output result as the second target voice request. The determination module is further configured to determine the second target voice request and the second error output result as the second negative sample sub-data. And based on the optimized large language model, perform secondary natural language processing on the second target voice request according to the target prompt word to obtain the fourth reasoning result and the fourth thinking logic chain. And when both the fourth reasoning result and the fourth thinking logic chain are correct, determine the fourth reasoning result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub-data. The determination module is further configured to determine the second target data pair according to the second positive sample sub-data and the second negative sample sub-data. The model training module is further configured to, based on the optimized large language model, perform secondary natural language processing on the second target voice request according to the target prompt word to obtain the fourth reasoning result and the fourth thinking logic chain.

[0245] In some embodiments, the processor is further configured to perform natural language processing on the original training data based on the optimized large language model obtained through preference-optimized training, according to the target prompt, to determine a third reasoning result and a third thinking logic chain corresponding to the third reasoning result. And according to the preset reasoning thinking chain and preset reasoning result, perform screening processing on the third reasoning result and the third thinking logic chain to determine a second error output result. And determine the original voice request corresponding to the second error output result as the second target voice request. The processor is further configured to determine the second target voice request and the second error output result as the second negative sample sub-data. And based on the optimized large language model, according to the target prompt, perform secondary natural language processing on the second target voice request to obtain a fourth reasoning result and a fourth thinking logic chain. And in the case where both the fourth reasoning result and the fourth thinking logic chain are correct, determine the fourth reasoning result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub-data. The processor is further configured to determine a second target data pair according to the second positive sample sub-data and the second negative sample sub-data. And based on the preset loss function, perform direct preference optimization training on the optimized large language model according to the second target data pair.

[0246] Specifically, the optimized large language model refers to the large language model obtained by performing direct preference optimization training on the large language model to be trained based on the first target data pair.

[0247] The second error output result refers to the content generated by the large language model that is determined to deviate from the preset standard (i.e., the preset reasoning thinking chain and preset reasoning result) after the screening process, including the second incorrect preset reasoning result or the second incorrect reasoning logic chain.

[0248] The second target voice request refers to the voice request that can be used to train the reasoning ability of the optimized large language model determined after the original voice request is preliminarily processed during the model training process.

[0249] The second negative sample sub-data refers to the result where the first reasoning result and the first thinking logic chain generated by the optimized large language model do not meet the preset standard, including logical errors, format errors, or potential safety hazards in the output, etc.

[0250] The second positive sample sub-data refers to the result that meets the expected standard generated by the optimized large language model and can be used as the "standard answer" for the optimized large language model to learn.

[0251] The second target data pair refers to the training sample pair generated according to the target prompt, the preset reasoning thinking chain, and the preset reasoning result, which is used to supervise the large language model to learn the deterministic reasoning path from the original voice request to the executable instruction, ensuring that the result generated by the large language model reasoning simultaneously meets logical correctness (i.e., meets the preset reasoning thinking chain) and physical executability (i.e., meets the preset reasoning result).

[0252] It should be noted that the specific method for determining the second target data pair is similar to that for determining the first target data pair as described above. The difference is that both the second positive sample sub - data and the second negative sample sub - data are determined by the optimized large - language model, which will not be elaborated here.

[0253] In this way, based on the optimized large - language model obtained through preference - optimized training, the server performs natural language processing on the original training data according to the target prompt word to determine the third inference result and the third thinking logic chain corresponding to the third inference result. Then, according to the preset inference thinking chain and the preset inference result, the server filters and processes the third inference result and the third thinking logic chain to determine the second error output result, which includes the second wrong inference result and / or the second wrong inference logic chain. Then, the server determines the original voice request corresponding to the second error output result as the second target voice request. Subsequently, the server determines the second target voice request and the second error output result as the second negative sample sub - data. Then, based on the optimized large - language model, the server performs secondary natural language processing on the second target voice request according to the target prompt word to obtain the fourth inference result and the fourth thinking logic chain. Subsequently, when both the fourth inference result and the fourth thinking logic chain are correct, the server determines the fourth inference result, the fourth thinking logic chain, and the second target voice request as the second positive sample sub - data. Then, the server determines the second target data pair according to the second positive sample sub - data and the second negative sample sub - data. Finally, based on the preset loss function, the server performs direct preference - optimized training on the optimized large - language model according to the second target data pair. In this way, by using the preset loss function and performing direct preference - optimized training with the second target data obtained based on the optimized large - language model, the large - language model to be trained can learn how to accurately understand and process the voice requests sent by users in the vehicle cockpit scenario and generate corresponding vehicle control instructions, thereby improving the user experience. Moreover, by continuously detecting and correcting the errors of the model and using the second target data for direct preference - optimized training again, the performance of the large - language model is continuously optimized and the user experience is improved.

[0254] This application also provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method as described above are implemented.

[0255] It can be understood that a computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc.

[0256] In the description of this specification, the descriptions referring to terms such as "specifically", "furthermore", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0257] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment or part of executable request code including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.

[0258] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A model training method for a large language model, characterized in that: The method comprises: Determine a list of associated application interfaces according to original training data, wherein the original training data includes an original voice request, a preset reasoning chain and a preset reasoning result corresponding to the original voice request, and a vehicle control instruction corresponding to the original voice request that can be executed by the vehicle can be generated based on the preset reasoning chain and the preset reasoning result; Based on a preset prompt word template, according to the associated application interface list and the original voice request, determining a target prompt word; Determining a first target data pair according to the target prompt word, the preset reasoning thought chain and the preset reasoning result; The large language model to be trained is trained according to the first target data pair.

2. The method according to claim 1, characterized in that The step of determining the associated application interface list according to the original training data includes: Based on a preset model, encoding the original voice request to determine a target vector; According to the target vector, the associated application interface list is determined from a pre-built application interface database, wherein the application interface database includes candidate application interfaces, application interface parameters to be filled in corresponding to the candidate application interfaces, application interface descriptions of the candidate application interfaces, and interface parameter descriptions of the application interface parameters.

3. The method according to claim 1, characterized in that The step of determining the target prompt word based on the preset prompt word template according to the associated application interface list and the original voice request includes: The preset prompt word template, the associated application interface list and the original voice request are concatenated to determine the target prompt word.

4. The method according to claim 1, characterized in that: The step of determining a first target data pair according to the target prompt word, the preset reasoning thought chain and the preset reasoning result includes: Based on the large language model to be trained, determining the first negative sample sub-data in the first target voice request and the first target data pair according to the target prompt word, the preset reasoning thinking chain and the preset reasoning result; Based on the base large language model, determining the first positive sample sub-data in the first target data pair according to the target prompt word corresponding to the first target voice request, the preset reasoning thinking chain and the preset reasoning result, wherein the parameter scale of the base large language model is greater than that of the large language model to be trained; The first target data pair is determined according to the first positive sample sub-data and the first negative sample sub-data.

5. The method according to claim 4, characterized in that The determining, based on the large language model to be trained, according to the target prompt word, the preset reasoning thinking chain and the preset reasoning result, the first negative sample sub-data in the first target voice request and the first target data pair includes: Based on the large language model to be trained and according to the target prompt word, determining a first reasoning result and a first thinking logic chain; According to the preset reasoning thinking chain and the preset reasoning result, the first reasoning result and the first thinking logic chain are screened and processed to determine a first error output result, wherein the first error output result includes a first error pre-reasoning result and / or a first error reasoning logic chain; Determine the original voice request corresponding to the first error output result as the first target voice request; The first target voice request and the first error output result are determined as the first negative sample sub-data.

6. The method according to claim 5, characterized in that The method of determining the first positive sample data in the first target data pair based on the base large language model according to the target prompt word corresponding to the first target voice request, the preset reasoning thinking chain and the preset reasoning result includes: Based on the base large language model, according to the target prompt word, natural language processing is performed on the first target voice request to determine a second reasoning result and a second thinking logic chain; When the second inference result and the second thinking logic chain are both correct, the second inference result, the second thinking logic chain and the first target voice request are determined as the first positive sample data.

7. The method according to claim 1, characterized in that The training of the large language model to be trained according to the first target data pair includes: Based on a preset loss function, direct preference optimization training is performed on the large language model to be trained according to the first target data pair.

8. The method according to claim 7, characterized in that The method further comprises: Based on the optimized large language model obtained through the preference optimization training, and according to the target prompt word, natural language processing is performed on the original training data to determine a third reasoning result and a third thinking logic chain corresponding to the third reasoning result; According to the preset reasoning thinking chain and the preset reasoning result, the third reasoning result and the third thinking logic chain are screened and processed to determine a second error output result, wherein the second error output result includes an error second error reasoning result and / or a second error reasoning logic chain; Determine the original voice request corresponding to the second erroneous output result as a second target voice request; Determine the second target voice request and the second error output result as second negative sample sub-data; Based on the optimized large language model, and according to the target prompt word, performing secondary natural language processing on the second target voice request to obtain a fourth reasoning result and a fourth thinking logic chain; When the fourth inference result and the fourth thinking logic chain are both correct, determining the fourth inference result, the fourth thinking logic chain and the target voice request as second positive sample data; Determine the second target data pair according to the second positive sample sub-data and the second negative sample sub-data; Based on the preset loss function, direct preference optimization training is performed on the optimized large language model according to the second target data pair.

9. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Training method of retrieval enhancement model, voice interaction method, server and medium

    CN120783774A

  • Voice interaction method and device, equipment and storage medium

    CN120895035A

  • Voice interaction method and device, equipment and storage medium

    CN120954403A