Model training method, server and computer readable storage medium
By generating target training data to simulate human thinking, training a large language model is solved, and the problem that the large language model cannot understand user voice requests in complex cockpit scenarios is improved, the adaptability and accuracy of the model is improved, and the user experience is enhanced.
Patent Information
- Application Number
- CN202510499072.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, large language models cannot accurately understand user voice requests and selected application interfaces in complex cockpit scenarios, resulting in poor user experience and may cause security risks.
By determining the associated application interface list and preset prompt word template, generating target training data, simulating human thinking methods to train large language models, gradually reasoning and understanding the relationship between user voice requests and application interfaces.
It improves the adaptability and accuracy of large language models in complex cockpit scenarios, enhances the model's understanding and reasoning ability, reduces misunderstandings and errors, and provides a good user experience.
Smart Images

Figure CN120356464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large language model training, and particularly to a model training method, a server, and a computer-readable storage medium for a large language model. Background Art
[0002] In the related art, based on the speech recognition result of a complete user voice request, and an application interface retrieved by retrieval-augmented generation technology and similar to the speech recognition result, they are spliced into a new prompt, and then input into the large language model to output a natural language processing result, so as to realize the user's control of the vehicle through voice interaction. However, in this way, only directly giving the natural language result as a label to train the large language model. In a complex cockpit scenario, even if the large language model is given a detailed application interface description and application interface parameters, the large language model cannot completely extract slots according to the user voice request and the selected application interface, but only corresponds the user voice request with the natural language result, and does not fully understand the given application interface. The large language model cannot accurately understand the user voice request, thus affecting the user experience. Summary of the Invention
[0003] The present application provides a model training method, a server, and a computer-readable storage medium for a large language model.
[0004] An embodiment of the present application provides a model training method for a large language model, the method comprising:
[0005] Determining an associated application interface list according to target voice request training data;
[0006] Based on a preset prompt template, determining target training data according to the associated application interface list and the target voice request training data;
[0007] Training the large language model according to the target training data.
[0008] In this way, the server determines an associated application interface list according to the target voice request training data. Then, based on the preset prompt template, the server determines the target training data according to the associated application interface list and the target voice request training data. Finally, the server trains the large language model according to the target training data. In this way, by generating target training data including the model logical thinking chain, the large language model is trained in terms of thinking logic, enabling the large language model to simulate the human thinking mode, gradually reason and understand the relationship between the user voice request and the application interface, thereby improving the adaptability of the large language model to various scenarios and enhancing the accuracy and generalization of the large language model.
[0009] In some embodiments, the determining an associated application interface list according to target voice request training data includes:
[0010] Encode the target voice request training data to determine a target vector;
[0011] According to the target vector, determine the associated application interface list from a pre-constructed application interface database, where the application interface database includes application interfaces and interface parameters corresponding to the application interfaces.
[0012] In this way, the server encodes the target voice request training data to determine a target vector. Then, the server determines the associated application interface list from the pre-constructed application interface database according to the target vector, where the application interface database includes application interfaces and interface parameters corresponding to the application interfaces. In this way, through encoding processing and vector retrieval, the associated application interface list related to the target voice request training data can be efficiently retrieved from the application interface database, improving the retrieval efficiency.
[0013] In some embodiments, determining the target training data based on the preset prompt template, according to the associated application interface list and the target voice request training data, includes:
[0014] According to the first preset sub-prompt in the preset prompt template, the associated application interface list, and the target voice request training data, determine a first target sub-prompt, where the first target sub-prompt is configured to guide a target large language model to determine a target thinking logic chain according to the associated application interface list and the target voice request training data;
[0015] According to the second preset sub-prompt in the preset prompt template, the associated application interface list, and the target voice request training data, determine a second target sub-prompt, where the second target sub-prompt is configured to guide a target large language model to determine the suitability of the associated application interface list for the target voice request training data according to the target thinking logic chain;
[0016] According to the third preset sub-prompt in the preset prompt template, the associated application interface list, and the target voice request training data, determine a third target sub-prompt, where the third target sub-prompt is configured to guide a target large language model to determine a natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data;
[0017] Determine the target training data according to the target voice request training data, the first target sub-prompt, the second target sub-prompt, and the third target sub-prompt.
[0018] In this way, the server determines the first target sub-prompt word according to the first preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The first target sub-prompt word is configured to guide the target large language model to determine the target thinking logic chain according to the associated application interface list and the target voice request training data. Then, the server determines the second target sub-prompt word according to the second preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The second target sub-prompt word is configured to guide the target large language model to determine the adaptability of the associated application interface list to the target voice request training data according to the target thinking logic chain. Then, the server determines the third target sub-prompt word according to the third preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The third target sub-prompt word is configured to guide the target large language model to determine the natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data. Finally, the server determines the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word. In this way, through the sub-prompt words in the preset prompt word template and the target voice request training data, the target sub-prompt words that can effectively guide the target large language model to think are generated, enabling the large language model to gradually reason and understand the relationship between the user's voice request and the application interface.
[0019] In some embodiments, the determining the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word includes:
[0020] Based on the large language model, according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, determine the first training data to be processed;
[0021] Based on a preset model, determine the validity of the first training data to be processed;
[0022] Within the first preset validity determination count threshold, when it is determined that the first training data to be processed is valid, determine the first training data to be processed as the target training data.
[0023] Thus, based on the large language model, the server determines the first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data. Then, based on the preset model, the server determines the validity of the first training data to be processed. Finally, within the threshold of the first preset validity determination times, when the server determines that the first training data to be processed is valid, the first training data to be processed is determined as the target training data. In this way, by using the preset model to judge the validity of the first training data to be processed, the accuracy of the final target training data can be guaranteed, and the use of incorrect or low-quality training data can be avoided.
[0024] In some embodiments, the determining, based on the large language model, the first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data includes:
[0025] According to the target voice request training data, vehicle knowledge associated with the target voice request training data is obtained from a pre-constructed vehicle knowledge base;
[0026] According to the vehicle knowledge and the first target sub-prompt word, the target thinking logic chain is determined;
[0027] According to the second target sub-prompt word, the suitability of the associated application interfaces in the associated application interface list for the target voice request training data is determined;
[0028] According to the suitability and the third target sub-prompt word, the natural language processing result is determined;
[0029] According to the target voice request training data, the target thinking logic chain, the preset prompt word template, and the natural language processing result, the first training data to be processed is determined.
[0030] In this way, the server trains data based on the target voice request and obtains vehicle knowledge associated with the target voice request training data from a pre-constructed vehicle knowledge base. Then, the server determines the target thinking logic chain based on the vehicle knowledge and the first target sub-prompt word. Next, the server determines the adaptability of the associated application interfaces in the associated application interface list to the target voice request training data according to the second target sub-prompt word. Subsequently, the server determines the natural language processing result based on the adaptability and the third target sub-prompt word. Finally, the server determines the first training data to be processed based on the target voice request training data, the target thinking logic chain, the preset prompt word template, and the natural language processing result. In this way, by obtaining vehicle knowledge, the large language model can accurately understand the relationship between the user's voice request and the cockpit scenario. Moreover, through the target sub-prompt words, the target large language model can be effectively guided to perform logical thinking training, enabling the large language model to gradually reason and understand the relationship between the user's voice request and the application interface.
[0031] In some embodiments, the associated application interface list includes associated interface parameters corresponding to the associated application interface, and determining the natural language processing result according to the adaptability and the third target sub-prompt word includes:
[0032] When the associated application interface is applicable to the target voice request training data, perform slot recognition on the target voice request training data to determine slot entities;
[0033] Perform application interface parameter filling according to the associated interface parameters and the slot entities to determine the application interface parameter filling result;
[0034] Determine the natural language processing result according to the associated application interface and the application interface parameter filling result.
[0035] In this way, when the associated application interface is applicable to the target voice request training data, the server performs slot recognition on the target voice request training data to determine slot entities. Then, the server performs application interface parameter filling according to the associated interface parameters and the slot entities to determine the application interface parameter filling result. Finally, the server determines the natural language processing result according to the associated application interface and the application interface parameter filling result. In this way, through slot recognition and application interface prediction, the large language model can accurately understand the user's request and select the most appropriate application interface to perform the task, thereby reducing misunderstandings and errors and providing a good user experience.
[0036] In some embodiments, the method further includes:
[0037] In the case where it is determined that the first training data to be processed is invalid within the first preset validity determination times threshold, based on an alternative large language model, according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, determine the second training data to be processed, where the parameter scale of the alternative large language model is larger than that of the large language model;
[0038] Based on a preset model, perform a discrimination process on the second training data to be processed to determine the validity of the second training data to be processed;
[0039] In the case where it is determined that the second training data to be processed is valid within the second preset validity determination times threshold, determine the second training data to be processed as the target training data.
[0040] In this way, in the case where it is determined that the first training data to be processed is invalid within the first preset validity determination times threshold, based on an alternative large language model, the server determines the second training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, where the parameter scale of the alternative large language model is larger than that of the large language model. Then, based on a preset model, the server performs a discrimination process on the second training data to be processed to determine the validity of the second training data to be processed. Finally, in the case where it is determined that the second training data to be processed is valid within the second preset validity determination times threshold, the server determines the second training data to be processed as the target training data. In this way, when the first training data to be processed cannot meet the validity requirements of the preset model, using an alternative large language model with a larger parameter scale can generate higher-quality second training data to be processed.
[0041] In some embodiments, the method further includes:
[0042] In the case where it is determined that the second training data to be processed is invalid within the second preset validity determination times threshold, based on a preset rule, perform a correction process on the second training data to be processed to determine the target training data.
[0043] In this way, in the case where it is determined that the second training data to be processed is invalid within the second preset validity determination times threshold, based on a preset rule, the server performs a correction process on the second training data to be processed to determine the target training data. In this way, by correcting the second training data to be processed according to a preset rule, the accuracy of the final target training data can be ensured, and the use of incorrect or low-quality training data can be avoided.
[0044] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above-mentioned method is implemented.
[0045] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0046] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the embodiments of the present application. Description of the Drawings
[0047] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0048] Figure 1 is one of the flow diagrams of the model training method of some embodiments of the present application;
[0049] Figure 2 is the second of the flow diagrams of the model training method of some embodiments of the present application;
[0050] Figure 3 is the third of the flow diagrams of the model training method of some embodiments of the present application;
[0051] Figure 4 is the fourth of the flow diagrams of the model training method of some embodiments of the present application;
[0052] Figure 5 is the fifth of the flow diagrams of the model training method of some embodiments of the present application;
[0053] Figure 6 is the sixth of the flow diagrams of the model training method of some embodiments of the present application;
[0054] Figure 7 is the seventh of the flow diagrams of the model training method of some embodiments of the present application;
[0055] Figure 8 is the eighth of the flow diagrams of the model training method of some embodiments of the present application;
[0056] Figure 9 is the flow diagram of the target training data generation of some embodiments of the present application. Detailed Embodiments
[0057] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.
[0058] In the existing intelligent cockpit voice interaction technology, the system usually based on the speech recognition result of the user's voice request, matches the interfaces semantically similar to the speech recognition result from a predefined application interface database through retrieval-augmented generation technology (such as "turn off media sound" or "adjust headrest audio mode"), and concatenates the descriptions and parameters of these application interfaces into structured prompt words to input into a large language model (LLM), and finally outputs a natural language understanding result to execute vehicle control.
[0059] However, this solution exposes significant defects in complex scenarios: the model is only trained through the mechanical mapping of "voice request - result", lacking in-depth understanding of the functions of application interfaces and the logical reasoning ability of user intentions. For example, when the user's voice request is "make the rear passengers unable to hear the navigation prompt sound", the model may mechanically select the "media volume adjustment" application interface and reduce the global volume instead of calling the "headrest audio private mode" application interface and accurately setting seat_position = "driver's seat". This rigid learning mode causes the model to be unable to disassemble complex voice requests (such as "turn on the air conditioner first and then turn off the rear air outlet"), misfill parameters (such as mapping "private" to volume instead of audio mode), or ignore implicit requirements (such as the cockpit common sense of "inside the vehicle = all"). The root cause is that the existing technology directly uses the natural language understanding result as the training label, making the model become a "keyword matching tool" rather than an intelligent agent with context awareness and step-by-step reasoning ability, ultimately resulting in fragmented user experience, deviation in function execution, and even potential safety hazards (such as accidentally touching the window control).
[0060] Based on the above problems, please refer to Figure 1 , embodiments of the present application provide a model training method for a large language model, the method comprising:
[0061] 01: Determine an associated application interface list according to the target voice request training data;
[0062] 02: Based on a preset prompt word template, determine target training data according to the associated application interface list and the target voice request training data;
[0063] 03: Train the large language model according to the target training data.
[0064] The embodiments of the present application also provide a server, including a memory and a processor. The model training method of the large language model in the embodiments of the present application can be implemented by the server in the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is used to determine an associated application interface list according to the target voice request training data. And based on a preset prompt word template, determine target training data according to the associated application interface list and the target voice request training data. And train the large language model according to the target training data.
[0065] The embodiments of the present application also provide a model training device. The model training method of the large language model in the embodiments of the present application can be implemented by the model training device in the embodiments of the present application. Specifically, the model training device includes a determination module and a model training module. The determination module is used to determine an associated application interface list according to the target voice request training data. The determination module is also used to determine target training data based on a preset prompt word template according to the associated application interface list and the target voice request training data. The model training module is used to train the large language model according to the target training data.
[0066] Specifically, a large language model refers to a model with a large number of parameters and broad language understanding capabilities, which can be trained through a large amount of text data, can understand and generate natural language, and can perform various language tasks. In the embodiments of the present application, by using the natural language processing ability of the large language model, the target voice request training data is processed to generate target training data that can improve the understanding and reasoning ability of the large language model in complex cockpit scenarios, and the large language model itself is trained according to the target training data to strengthen the understanding and reasoning ability of the large language model. It should be noted that the large language model without being trained by the target training data can output a lot of Chain of Thought (COT) data according to the target voice request training data, but not all of these generated COT data can meet the user's needs. To enable the large language model to accurately output COT data that meets the user's needs, it is necessary to select suitable COT data from the large amount of output COT data as the target training data to train the large language model and strengthen the large language model so that the large language model can accurately output COT data that meets the user's needs.
[0067] COT data is a strategy used for reasoning in the field of natural language processing. The core idea of COT data is to decompose problems and simulate the step-by-step reasoning process of human thinking, thereby helping large language models better understand complex tasks and give more accurate answers. Specifically, the COT method decomposes complex tasks into a series of subtasks or intermediate steps, and each step provides more detailed reasoning information to help the model draw the final correct conclusion through reasoning. Through step-by-step reasoning, the thinking mode of COT not only improves the accuracy of solving complex problems but also enhances the interpretability of large language models.
[0068] The target voice request training data refers to a piece of voice data randomly obtained from the voice request sample data originally used for training large language models, and the text training data obtained by converting this voice data into text through a speech recognition model. Among them, these voice request sample data usually come from real user voice inputs, representing various requests that users may send to in-vehicle voice assistants. For example, the voice request of the user to the in-vehicle voice assistant "Turn off the sound in the car".
[0069] The associated application interface list refers to the set of application program interfaces related to the target voice request training data, including associated application interfaces and associated interface parameters corresponding to the associated application interfaces. The determination of the associated application interface list is based on the content of the voice request, that is, according to the user voice request, the system identifies the related application interfaces. For example, for the instruction "Turn off the sound in the car", the associated application interfaces may include turning off the headrest audio or the main driver audio mode in the car, as well as turning on the sound or canceling the mute.
[0070] The preset prompt template refers to a template used to guide large language models to perform reasoning and generate Chain of Thought (COT) data.
[0071] The target training data refers to the COT data used for actual training of large language models according to the preset prompt template and after being processed, including the input target voice request training data, the final output result, and the intermediate reasoning steps required to achieve this result. Training large language models with target training data can improve the understanding and reasoning abilities of large language models in complex cockpit scenarios. That is, by constructing COT data, the large language model simulates the human thinking mode, first selects the corresponding application interface, then selects the appropriate application interface according to the user voice request and the application interface description, and fills in the corresponding slots step by step according to the selected application interface to obtain the final natural language processing result, thereby enhancing the understanding and reasoning abilities of the large language model in complex cockpit scenarios.
[0072] First, based on the training data of the target voice request, a list of associated application interfaces is retrieved through retrieval-augmented generation technology. For example, for the user instruction "Turn off the sound in the car", the RAG model retrieves the list of associated application interfaces "Associated Application Interface 1: Headrest_Voice_Model_Close and corresponding parameters; Associated Application Interface 2: Media_Voice_Open and corresponding parameters".
[0073] Next, the server constructs a COT prompt based on a preset prompt template, in combination with the list of associated application interfaces and the training data of the target voice request. And based on the constructed COT prompt, the target training data is determined.
[0074] Finally, the server trains the large language model based on the target training data. During the training process, the large language model will learn how to gradually reason and generate correct natural language processing results according to the user voice request and the list of associated application interfaces.
[0075] In summary, in the model training method and server of the large language model provided by the embodiment of the present application, the server determines the list of associated application interfaces according to the training data of the target voice request. Next, based on the preset prompt template, the server determines the target training data according to the list of associated application interfaces and the training data of the target voice request. Finally, the server trains the large language model according to the target training data. In this way, by generating the target training data including the model logical thinking chain, the large language model is trained in terms of thinking logic, enabling the large language model to simulate the human thinking mode, gradually reason and understand the relationship between the user voice request and the application interface, thereby improving the adaptability of the large language model to various scenarios and enhancing the accuracy and generalization of the large language model.
[0076] Please refer to Figure 2 , in some embodiments, step 01 (determining the list of associated application interfaces according to the training data of the target voice request) includes:
[0077] 011: Perform encoding processing on the training data of the target voice request to determine the target vector;
[0078] 012: Determine the list of associated application interfaces from the pre-constructed application interface database according to the target vector.
[0079] In some embodiments, the determining module is further configured to perform encoding processing on the training data of the target voice request to determine the target vector. And determine the list of associated application interfaces from the pre-constructed application interface database according to the target vector.
[0080] In some embodiments, the processor is further configured to encode the target voice request training data to determine a target vector, and based on the target vector, determine an associated application interface list from a pre-constructed application interface database.
[0081] Specifically, the encoding process refers to converting the target voice request training data into a vector based on the BGE model, enabling it to be understood and processed by a machine learning model. The BGE model is an embedding model based on a graph structure that can convert a user voice request into a vector representation and retrieve an application interface description vector semantically similar to the user instruction from the application interface database. In some embodiments, methods such as word embedding and sequence encoding can also be used to convert text into a high-dimensional vector while preserving the semantic information of the text.
[0082] The target vector refers to the vector representation corresponding to the target voice request training data after encoding processing, which can be used as a basis for retrieving the associated application interface list, calculating the similarity with the interface description vectors in the pre-constructed application interface database, and thus determining the associated application interfaces. In this way, retrieving based on vector similarity can more accurately match the application interfaces related to the target voice request training data and avoid mis-matching.
[0083] The application interface database refers to a database that stores application interfaces and their corresponding parameters, providing a data basis for retrieving associated application interfaces, which includes the descriptions of application interfaces and the corresponding application interface parameter information. The data information that the application interface database may include is "Headrest_Voice_Model_Close [Application interface description: HeadrestVoiceModelClose closes the in-vehicle headrest audio or the driver's seat audio mode, where HeadrestVoice represents the headrest audio and Model represents the mode. Broadcasting and setting the sound to private or driver's enjoyment means setting the headrest audio to a certain audio mode. ARGUMENTS: "target_function": (enumeration value, required) the value can be selected from ['intelligent mode', 'full vehicle playback mode','standard mode', 'driver's enjoyment mode', 'private enjoyment mode']; "seat_position": (enumeration value, optional) the value is ['driver's seat'], this parameter only supports the driver's seat, and when no relevant information about the driver's seat is mentioned, it is not output; "device": (enumeration value, required) the value is ['headrest audio']]", "Media_Voice_Open [Application interface description: MediaVoiceOpen turns on the sound or unmutes, where MediaVoice represents the media sound. In-vehicle media such as music, video, radio, etc. can all be understood as sound, which is different from MediaEffect. ARGUMENTS: "target_function": (enumeration value, required) the value is ['sound']; "seat_position": (enumeration value, optional) the value can be selected from ['driver's seat', 'passenger seat', 'front row', 'left side of the second row', 'right side of the second row','second row', 'all', 'left side of the third row', 'right side of the third row', 'third row']]".
[0084] First, based on the BGE model, the server encodes the target voice request training data to determine the target vector. For example, converting "Turn off the sound in the vehicle" into a vector.
[0085] Next, based on the BGE model, according to the target vector, a list of candidate application interfaces similar to it is retrieved from the pre-constructed application interface database. In some embodiments, the number of candidate application interfaces in the candidate application interface list can be set. For example, the two most similar candidate application interfaces are selected. After determining the list of candidate application interfaces based on the BGE model and the target vector, then based on the retrieval enhancement generation technology, similar associated application interfaces are selected from the list of candidate application interfaces to generate a list of associated application interfaces. Continuing with the above example, according to the user voice request "Turn off the sound in the vehicle", the determined list of associated application interfaces is "Headrest_Voice_Model_Close" and "Media_Voice_Open".
[0086] In this way, the server encodes the target voice request training data to determine the target vector. Then, based on the target vector, the server determines an associated application interface list from a pre-constructed application interface database, where the application interface database includes application interfaces and corresponding interface parameters. Thus, through encoding processing and vector retrieval, the associated application interface list related to the target voice request training data can be efficiently retrieved from the application interface database, improving the retrieval efficiency.
[0087] Please refer to Figure 3 , in some embodiments, step 02 (determining the target training data based on a preset prompt word template, the associated application interface list, and the target voice request training data) includes:
[0088] 021: Determine the first target sub-prompt word according to the first preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data;
[0089] 022: Determine the second target sub-prompt word according to the second preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data;
[0090] 023: Determine the third target sub-prompt word according to the third preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data;
[0091] 024: Determine the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word.
[0092] In some embodiments, the determination module is further configured to determine the first target sub-prompt word according to the first preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. And determine the second target sub-prompt word according to the second preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The determination module is further configured to determine the third target sub-prompt word according to the third preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. And determine the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word.
[0093] In some embodiments, the processor is further configured to determine a first target sub-prompt word according to a first preset sub-prompt word, an associated application interface list, and target voice request training data in a preset prompt word template. And determine a second target sub-prompt word according to a second preset sub-prompt word, an associated application interface list, and target voice request training data in the preset prompt word template. The processor is further configured to determine a third target sub-prompt word according to a third preset sub-prompt word, an associated application interface list, and target voice request training data in the preset prompt word template. And determine target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word.
[0094] Specifically, the first preset sub-prompt word refers to a prompt word template in the preset prompt word template used to guide the target large language model to determine the target thinking logic chain according to the associated application interface list and the target voice request training data. By processing with the associated application interface list and the target voice request training data, the first target sub-prompt word can be generated. The first target sub-prompt word can guide the target large language model to determine the target thinking logic chain according to the associated application interface list and the target voice request training data. In some embodiments, the first preset sub-prompt word may be "You are a vehicle-mounted voice assistant. According to the user's instructions, combined with the APIs and their parameters given in the
associated application interface list
[0095] First, you should think according to the user's instructions and give the thinking logic;
[0096]
Available associated application interface list
[0097] User voice request: ".
[0098] The determined first target sub-prompt word may be "You are a vehicle-mounted voice assistant. According to the user voice request 'Turn off the sound in the car', combined with the APIs and their parameters given in the
associated application interface list
[0099] First, you should think according to the user's instructions and give the thinking logic;
[0100]
Available associated application interface list
[0101]
Headrest_Voice_Model_Close
[0102]
Media_Voice_Open
[0103] Description: HeadrestVoiceModelClose turns off the in-vehicle headrest audio or the driver's seat audio mode. Here, HeadrestVoice represents the headrest audio, and Model represents the mode. Broadcasting and setting the sound to private or driver-exclusive means setting the headrest audio to a certain audio mode.
[0104] ARGUMENTS:
[0105] "target_function": (enumeration value, required) The value can be selected from ['Intelligent Mode', 'Full Vehicle Playback Mode', 'Standard Mode', 'Driver-Exclusive Mode', 'Private Mode'].
[0106] "seat_position": (enumeration value, optional) The value is ['Driver's Seat']. This parameter only supports the driver's seat and is not output when no relevant information about the driver's seat is mentioned.
[0107] "device": (enumeration value, required) The value is ['Headrest Audio'].
[0108]
Media_Voice_Open
[0109] Description: MediaVoiceOpen turns on the sound or unmutes it. Here, MediaVoice represents the media sound. In-vehicle media such as music, videos, and radio can all be understood as sound, which is different from MediaEffect.
[0110] ARGUMENTS:
[0111] "target_function": (enumeration value, required) The value is ['Sound'].
[0112] "seat_position": (enumeration value, optional) The value can be selected from ['Driver's Seat', 'Passenger Seat', 'Front Row', 'Left Side of the Second Row', 'Right Side of the Second Row', 'Second Row', 'All', 'Left Side of the Third Row', 'Right Side of the Third Row', 'Third Row'];
[0113] User voice request: "Turn off the sound in the vehicle."
[0114] The second preset sub - prompt refers to the prompt template in the preset prompt template used to guide the target large - language model to determine the adaptability of the associated application interface list to the target voice request training data according to the target thinking logic chain. By processing it with the associated application interface list and the target voice request training data, a second target sub - prompt can be generated. The second target sub - prompt can guide the target large - language model to determine the adaptability of the associated application interface list to the target voice request training data according to the target thinking logic chain. In some embodiments, the second preset sub - prompt may be "You are a vehicle voice assistant. According to the user's instructions, combined with the APIs and their parameters given in
associated application interface list
[0115] According to the thinking logic, judge whether the following APIs can help you complete the task. If there is no API that can help you complete the task, then output 'None'.
[0116]
List of available associated application interfaces
[0117] User voice request: ".
[0118] Replace the corresponding positions in the second preset sub - prompt according to the user voice request and the associated application interface list, which is the second target sub - prompt, and will not be elaborated here.
[0119] The third preset sub - prompt refers to the prompt template in the preset prompt template used to guide the target large - language model to determine the natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data. By processing it with the associated application interface list and the target voice request training data, a third target sub - prompt can be generated. The third target sub - prompt can guide the target large - language model to determine the natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data. In some embodiments, the third preset sub - prompt may be "You are a vehicle voice assistant. According to the user's instructions, combined with the APIs and their parameters given in
associated application interface list
[0120] If there is a suitable API in the list that can be used, then please first select 1 of them, and then fill in the parameters of the selected API in turn;
[0121] Note that if it is a required parameter in the parameters, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values;
[0122]
List of available associated application interfaces
[0123] User voice request: ".
[0124] The corresponding positions in the third preset sub-prompt word are replaced according to the user's voice request and the associated application interface list, which is the third target sub-prompt word, and will not be elaborated here.
[0125] First, according to the first preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data, the first target sub-prompt word is determined. This first target sub-prompt word is used to guide the target large language model to determine the target thinking logic chain according to the associated application interface list and the target voice request training data.
[0126] Next, according to the second preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data, the second target sub-prompt word is determined. This second target sub-prompt word is used to guide the target large language model to determine the adaptability of the associated application interface list to the target voice request training data according to the target thinking logic chain.
[0127] Then, according to the third preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data, the third target sub-prompt word is determined. This target sub-prompt word is used to guide the target large language model to determine the natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data.
[0128] Finally, according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word, the target training data is determined. In some embodiments, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word are concatenated to generate "You are a vehicle-mounted voice assistant. According to the user's instructions, combined with the APIs and their parameters given in the
associated application interface list
[0129] 1. First, you should think according to the user's instructions and give the thinking logic
[0130] 2. According to the thinking logic, judge whether the following APIs can help you complete the task. If there is no API that can help you complete the task, then output "None"
[0131] 3. If there is a suitable API in the list that can be used, then please first select 1 of them, and then fill in the parameters of the selected API in turn
[0132] 4. Note that if a parameter is required, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values
[0133]
Your output format
[0134] Thinking logic: xxx
[0135] Is there a suitable API: Yes or No
[0136] Selected API: {"API": "api1", "ARGUMENTS": {"arg1": "value1", "arg2": "value2",...,"argn": "valuen"}};
[0137]
List of available associated application interfaces
[0138] User voice request: ”.
[0139] In this way, the server determines the first target sub-prompt word based on the first preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The first target sub-prompt word is configured to guide the target large language model to determine the target thinking logic chain based on the associated application interface list and the target voice request training data. Then, the server determines the second target sub-prompt word based on the second preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The second target sub-prompt word is configured to guide the target large language model to determine the adaptability of the associated application interface list to the target voice request training data based on the target thinking logic chain. Then, the server determines the third target sub-prompt word based on the third preset sub-prompt word in the preset prompt word template, the associated application interface list, and the target voice request training data. The third target sub-prompt word is configured to guide the target large language model to determine the natural language processing result based on the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data. Finally, the server determines the target training data based on the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word. In this way, through the sub-prompt words in the preset prompt word template and the target voice request training data, target sub-prompt words that can effectively guide the target large language model to think are generated, enabling the large language model to gradually reason and understand the relationship between the user voice request and the application interface.
[0140] Please refer to Figure 4 , in some embodiments, step 024 (determining the target training data based on the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word) includes:
[0141] 0241: Based on the large language model, determine the first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data;
[0142] 0242: Determine the validity of the first training data to be processed based on a preset model;
[0143] 0243: When it is determined that the first training data to be processed is valid within the first preset validity determination times threshold, determine the first training data to be processed as the target training data.
[0144] In some embodiments, the determination module is further configured to determine the first training data to be processed based on a large language model according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, and determine the validity of the first training data to be processed based on a preset model. The determination module is further configured to determine the first training data to be processed as the target training data when it is determined that the first training data to be processed is valid within the first preset validity determination times threshold.
[0145] In some embodiments, the processor is further configured to determine the first training data to be processed based on a large language model according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, and determine the validity of the first training data to be processed based on a preset model. The processor is further configured to determine the first training data to be processed as the target training data when it is determined that the first training data to be processed is valid within the first preset validity determination times threshold.
[0146] Specifically, the first training data to be processed refers to the initial training data generated by a large language model according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word. The first training data to be processed is the basis for subsequent verification and screening, including correct and incorrect results.
[0147] The preset model refers to a pre-trained model for evaluating the validity of training data, which can identify and judge whether the generated training data to be processed meets the expectations, such as whether it contains correct semantic information, whether it can meet user requirements, and whether it meets specific format requirements, etc.
[0148] The first preset validity determination times threshold refers to the preset number limit for determining the validity of the first training data to be processed. Within this number threshold, as long as the preset model determines that the first training data to be processed is valid, it is considered that the first training data to be processed is valid.
[0149] First, based on a large language model, generate the first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data.
[0150] Next, based on a preset model, determine the validity of the first training data to be processed. The preset model analyzes the first training data to be processed and determines whether it conforms to the user's voice request and the associated application interface list and whether it can correctly complete the task.
[0151] Finally, within the threshold of the first preset validity determination count, if the first training data to be processed is valid, it is determined as the final target training data. If the first training data to be processed is invalid, the first training data to be processed is regenerated, and the above steps are repeated until the preset validity determination count threshold is reached.
[0152] Continuing with the above example, based on the large language model, according to the first target sub-prompt, the second target sub-prompt, the third target sub-prompt, and the target voice request training data, generate the prompt "You are a vehicle-mounted voice assistant. According to the user's instructions, combined with the APIs and their parameters given in the
associated application interface list
[0153] 1. First, you should think according to the user's instructions and give the thinking logic
[0154] 2. According to the thinking logic, determine whether the following APIs can help you complete the task. If there is no API that can help you complete the task, then output "None"
[0155] 3. If there is a suitable API in the list that can be used, then please first select 1 of them, and then fill in the parameters of the selected API in turn
[0156] 4. Note that if a parameter is required, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values
[0157]
Your output format
[0158] Thinking logic: xxx
[0159] Is there a suitable API: Yes or No
[0160] Selected API: {"API": "api1", "ARGUMENTS": {"arg1": "value1", "arg2": "value2",...,"argn": "valuen"}};
[0161]
List of available associated application interfaces
[0162]
Headrest_Voice_Model_Close
[0163]
Media_Voice_Open
[0164] Description: HeadrestVoiceModelClose turns off the in-vehicle headrest audio or the driver's seat audio mode. Here, HeadrestVoice represents the headrest audio, and Model represents the mode. Broadcasting and setting the sound to private or driver-exclusive means setting the headrest audio to a certain audio mode.
[0165] ARGUMENTS:
[0166] "target_function": (enumeration value, required) The value can be selected from ['Intelligent Mode', 'Full Vehicle Playback Mode', 'Standard Mode', 'Driver-Exclusive Mode', 'Private Mode']
[0167] "seat_position": (enumeration value, optional) The value is ['Driver's Seat']. This parameter only supports the driver's seat. When no information related to the driver's seat is mentioned, it is not output.
[0168] "device": (enumeration value, required) The value is ['Headrest Audio']
[0169]
Media_Voice_Open
[0170] Description: MediaVoiceOpen turns on the sound or unmutes it. Here, MediaVoice represents the media sound. In-vehicle media such as music, videos, and radio can all be understood as sound, which is different from MediaEffect.
[0171] ARGUMENTS:
[0172] "target_function": (enumeration value, required) The value is ['Sound']
[0173] "seat_position": (enumeration value, optional) The value can be selected from ['Driver's Seat', 'Passenger Seat', 'Front Row', 'Left Side of the Second Row', 'Right Side of the Second Row', 'Second Row', 'All', 'Left Side of the Third Row', 'Right Side of the Third Row', 'Third Row'];
[0174] User voice request: "Turn off the sound in the vehicle". Determine that the natural language processing result is "Thought logic: The user wants to turn on the sound in the vehicle. According to [Vehicle Knowledge], in-vehicle = all. Therefore, it is necessary to turn on the sound of the entire vehicle. From the API list, Media_Voice_Open can be used to turn on or unmute the in-vehicle sound, which exactly meets the requirement.
[0175] Is there a suitable API: Yes
[0176] Selected API: {"API": "Media_Voice_Open", "ARGUMENTS": {"target_function": "Sound", "seat_position": "All"}}. Subsequently, it is determined that the first training data to be processed is "User voice request: [Turn off the sound in the car]; Target prompt word: []; Thinking logic: [The user wants to turn on the sound in the car. According to
Vehicle Knowledge
[0177] Next, the above first training data to be processed is verified according to the preset model to determine the validity of the first training data to be processed.
[0178] If the first training data to be processed is valid, then the first training data to be processed is used as the target training data, and the transformation process of this target voice request training data is ended. If the first training data to be processed is invalid, that is, the user's needs cannot be met or there are other problems, then new first training data to be processed is generated according to the large language model, and this new first training data to be processed is verified based on the preset model until the generated first training data to be processed is valid, or the generation times of the first training data to be processed reach the first preset validity determination times threshold.
[0179] In this way, based on the large language model, the server determines the first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word and the target voice request training data. Then, based on the preset model, the server determines the validity of the first training data to be processed. Finally, within the first preset validity determination times threshold, when the server determines that the first training data to be processed is valid, the first training data to be processed is determined as the target training data. In this way, by judging the validity of the first training data to be processed through the preset model, the accuracy of the final target training data can be guaranteed, and the use of incorrect or low-quality training data can be avoided.
[0180] Please refer to Figure 5 , in some embodiments, step 0241 (Based on the large language model, according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word and the target voice request training data, determine the first training data to be processed), includes:
[0181] 02411: Acquire vehicle knowledge associated with the target voice request training data from a pre-constructed vehicle knowledge base according to the target voice request training data;
[0182] 02412: Determine the target thinking logic chain according to the vehicle knowledge and the first target sub-prompt word;
[0183] 02413: Determine the adaptability of the associated application interface in the associated application interface list to the target voice request training data according to the second target sub-prompt word;
[0184] 02414: Determine the natural language processing result according to the adaptability and the third target sub-prompt word;
[0185] 02415: Determine the first training data to be processed according to the target voice request training data, the target thinking logic chain, the preset prompt word template, and the natural language processing result.
[0186] In some embodiments, the model training device further includes an acquisition module, and the acquisition module is used to acquire vehicle knowledge associated with the target voice request training data from a pre-constructed vehicle knowledge base according to the target voice request training data. The determination module is further used to determine the target thinking logic chain according to the vehicle knowledge and the first target sub-prompt word. And determine the adaptability of the associated application interface in the associated application interface list to the target voice request training data according to the second target sub-prompt word. The determination module is further used to determine the natural language processing result according to the adaptability and the third target sub-prompt word. And determine the first training data to be processed according to the target voice request training data, the target thinking logic chain, the preset prompt word template, and the natural language processing result.
[0187] In some embodiments, the processor is further used to acquire vehicle knowledge associated with the target voice request training data from a pre-constructed vehicle knowledge base according to the target voice request training data. And determine the target thinking logic chain according to the vehicle knowledge and the first target sub-prompt word. And determine the adaptability of the associated application interface in the associated application interface list to the target voice request training data according to the second target sub-prompt word. The processor is further used to determine the natural language processing result according to the adaptability and the third target sub-prompt word. And determine the first training data to be processed according to the target voice request training data, the target thinking logic chain, the preset prompt word template, and the natural language processing result.
[0188] Specifically, the vehicle knowledge base refers to a pre - constructed database that includes various knowledge information related to vehicles, such as vehicle configurations, functions, and operation methods. The vehicle knowledge base provides prior knowledge for the large - language model, helping the model better understand user instructions and vehicle functions, thus generating more accurate natural language processing results. The vehicle knowledge base may include "vehicle identification number", "brand model", "engine specifications", "transmission type", "sharing mode = full - vehicle playback = standard mode", and "inside the vehicle = all", etc.
[0189] The target thinking logic chain is the thinking process by which the large - language model infers the final answer based on the user's voice request and vehicle knowledge. The target thinking logic chain can help the large - language model clearly display its reasoning process, thereby improving the transparency and interpretability of the model. For example, when the user's voice request is "turn off the sound inside the vehicle", the target thinking logic chain may include the following steps: First, according to vehicle knowledge, determine that "inside the vehicle" refers to "all seats". Second, according to the user's instruction, determine that the operation to be performed is "turn off the sound". Third, from the API list, select the API that can turn off the sound (for example, Media_Voice_Open). Fourth, according to the parameter definition of the API, determine the parameter values to be filled (for example, target_function = "sound", seat_position = "all").
[0190] First, based on the target voice request training data, obtain the vehicle knowledge associated with this instruction from the pre - constructed vehicle knowledge base. For example, for the instruction "turn off the sound inside the vehicle", obtain the vehicle knowledge "inside the vehicle = all".
[0191] Next, based on the vehicle knowledge and the first target sub - prompt word, determine the target thinking logic chain. For example, based on the vehicle knowledge "inside the vehicle = all" and the sentence "you should think according to the user's instruction and give the thinking logic" in the first target sub - prompt word, determine the target thinking logic chain "turn off the sound of the whole vehicle".
[0192] Then, according to the second target sub - prompt word, evaluate the suitability of the APIs in the associated application interface list for the target voice request training data. For example, according to the sentence "judge whether the following APIs can help you complete the task according to the thinking logic. If there is no API that can help you complete the task, then output 'none'" in the second target sub - prompt word, evaluate the suitability of the two APIs "Headrest_Voice_Model_Close" and "Media_Voice_Open", and select the "Media_Voice_Open" API.
[0193] Subsequently, based on the API adaptability and the third target sub-prompt, determine the natural language processing result. For example, according to the third target sub-prompt "If there is a suitable API in the list that can be used, then please first select 1 of them, and then fill in the parameters of the selected API in sequence; note that if it is a required parameter in the parameters, then you must fill it in; if it is an enumerated value, then you can only select from the listed candidate values", fill in the parameters "target_function='voice', seat_position='all'" of the "Media_Voice_Open" API to generate the NLU result "Media_Voice_Open(target_function='voice', seat_position='all')".
[0194] Finally, based on the target voice request training data, the target thinking logic chain, the preset prompt template, and the natural language processing result, determine the first training data to be processed. The first training data to be processed includes information such as the target voice request training data, the target thinking logic chain, the selected associated application interface, the associated interface parameters, and the natural language processing result.
[0195] In this way, the server obtains vehicle knowledge associated with the target voice request training data from the pre-constructed vehicle knowledge base according to the target voice request training data. Then, the server determines the target thinking logic chain according to the vehicle knowledge and the first target sub-prompt. Next, the server determines the adaptability of the associated application interface in the list of associated application interfaces to the target voice request training data according to the second target sub-prompt. Subsequently, the server determines the natural language processing result according to the adaptability and the third target sub-prompt. Finally, the server determines the first training data to be processed according to the target voice request training data, the target thinking logic chain, the preset prompt template, and the natural language processing result. In this way, by obtaining vehicle knowledge, the large language model can accurately understand the relationship between the user's voice request and the cockpit scenario. And, through the target sub-prompt, the target large language model can be effectively guided to perform logical thinking training, enabling the large language model to gradually reason and understand the relationship between the user's voice request and the application interface.
[0196] Please refer to Figure 6 , in some embodiments, the list of associated application interfaces includes associated interface parameters corresponding to the associated application interfaces. Step 02414 (determine the natural language processing result according to the adaptability and the third target sub-prompt) includes:
[0197] 024141: In the case where the associated application interface is applicable to the target voice request training data, perform slot recognition on the target voice request training data to determine slot entities;
[0198] 024142: Fill the application interface parameters according to the associated interface parameters and slot entities, and determine the result of filling the application interface parameters.
[0199] 024143: Determine the natural language processing result according to the associated application interface and the result of filling the application interface parameters.
[0200] In some embodiments, the determination module is further configured to perform slot recognition on the target voice request training data to determine slot entities when the associated application interface is applicable to the target voice request training data. And fill the application interface parameters according to the associated interface parameters and slot entities, and determine the result of filling the application interface parameters. And determine the natural language processing result according to the associated application interface and the result of filling the application interface parameters.
[0201] In some embodiments, the processor is further configured to perform slot recognition on the target voice request training data to determine slot entities when the associated application interface is applicable to the target voice request training data. And fill the application interface parameters according to the associated interface parameters and slot entities, and determine the result of filling the application interface parameters. And determine the natural language processing result according to the associated application interface and the result of filling the application interface parameters.
[0202] Specifically, slot recognition refers to extracting specific information segments from the user's input, and these information segments are usually called "slots". Slots are usually the key information necessary to complete a certain task or request, such as time, location, object, etc. Taking the user's voice request "What's the temperature tomorrow" as an example, the slot information obtained by performing slot recognition can include ["tomorrow" - Date], that is, the slot information includes the slot value and the slot type, where "tomorrow" is the slot value and Date is the slot type. Taking the user's voice request "Navigate to Address A" as an example, the slot information obtained by performing slot recognition is ["Address A" - Place], where "Address A" is the slot value and Place is the slot type.
[0203] Slot entities refer to the named entities obtained by performing slot recognition on the voice request. Such as the above slot information ["tomorrow" - Date] and slot information ["Address A" - Place], etc.
[0204] When the associated application interface is applicable to the target voice request training data, perform slot recognition on the target voice request training data to determine slot entities. For example, for the instruction "Turn off the sound in the car", the slot entities "sound" and "in the car" are recognized.
[0205] Next, based on the associated interface parameters and slot entities, perform application interface parameter filling to determine the result of application interface parameter filling. For example, according to the parameter "target_function" of the "Media_Voice_Open" API and the identified slot entity "voice", fill the parameter as "target_function ='voice'".
[0206] Finally, based on the associated application interface and the result of application interface parameter filling, determine the natural language processing result. For example, according to the "Media_Voice_Open" API and the application interface parameter filling result "target_function ='voice'", generate the NLU result "Media_Voice_Open(target_function ='voice')".
[0207] In this way, when the associated application interface is applicable to the target voice request training data, the server performs slot recognition on the target voice request training data to determine the slot entity. Next, the server performs application interface parameter filling based on the associated interface parameters and the slot entity to determine the result of application interface parameter filling. Finally, the server determines the natural language processing result based on the associated application interface and the result of application interface parameter filling. In this way, through slot recognition and application programming interface prediction, the large language model can accurately understand the user's request and select the most appropriate application interface to execute the task, thereby reducing misunderstandings and errors and providing a good user experience.
[0208] Please refer to Figure 7 , in some embodiments, the method further includes:
[0209] 0244: In the case where it is determined that the first training data to be processed is invalid within the first preset validity determination count threshold, based on the alternative large language model, determine the second training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data;
[0210] 0245: Based on the preset model, perform a discrimination process on the second training data to be processed to determine the validity of the second training data to be processed;
[0211] 0246: In the case where it is determined that the second training data to be processed is valid within the second preset validity determination count threshold, determine the second training data to be processed as the target training data.
[0212] In some embodiments, the determination module is further configured to, when it is determined that the first training data to be processed is invalid within the first preset validity determination frequency threshold, determine the second training data to be processed based on the alternative large language model according to the first target sub-prompt, the second target sub-prompt, the third target sub-prompt, and the target voice request training data. And based on the preset model, perform discrimination processing on the second training data to be processed to determine the validity of the second training data to be processed. And when it is determined that the second training data to be processed is valid within the second preset validity determination frequency threshold, determine the second training data to be processed as the target training data.
[0213] In some embodiments, the processor is further configured to, when it is determined that the first training data to be processed is invalid within the first preset validity determination frequency threshold, determine the second training data to be processed based on the alternative large language model according to the first target sub-prompt, the second target sub-prompt, the third target sub-prompt, and the target voice request training data. And based on the preset model, perform discrimination processing on the second training data to be processed to determine the validity of the second training data to be processed. And when it is determined that the second training data to be processed is valid within the second preset validity determination frequency threshold, determine the second training data to be processed as the target training data.
[0214] Specifically, the parameter scale of the alternative large language model is larger than that of the large language model. When it is determined that the first training data to be processed is invalid within the first preset validity determination frequency threshold, that is, when the large language model with a smaller parameter scale cannot upload accurate COT data, by using the alternative large language model with richer prior knowledge and stronger reasoning ability, more accurate thought chains can be generated, thus generating natural language processing results that can meet user needs. And if the second training data to be processed generated by the alternative large language model with a large parameter scale is valid, it can be used to distill the large language model, enabling the large language model to learn more accurate knowledge and reasoning methods, thereby improving the generalization ability and accuracy of the large language model, and enabling the large language model to perform correct reasoning in more complex scenarios. In some embodiments, the parameter scale of the large language model can be 7B, and the parameter scale of the alternative large language model can be 72B.
[0215] When the first training data to be processed is invalid within the first preset validity determination frequency threshold, the alternative large language model is used to generate the second training data to be processed. The parameter scale of the alternative large language model is larger than that of the large language model used for the first training data to be processed. It should be noted that the process of the alternative large language model determining the second training data to be processed is the same as that of the large language model determining the first training data to be processed, and will not be elaborated here.
[0216] Next, use the preset model to discriminate the second training data to be processed and determine its validity. The preset model will analyze the second training data to be processed and determine whether the second training data to be processed conforms to the user voice request and the associated application interface list and whether it can correctly complete the task.
[0217] Finally, within the second preset validity determination count threshold, if the second training data to be processed is valid, it is determined as the final target training data. If it is invalid, the second training data to be processed is regenerated, and the above steps are repeated.
[0218] In this way, within the first preset validity determination count threshold, when it is determined that the first training data to be processed is invalid, based on the alternative large language model, the server determines the second training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data, where the parameter scale of the alternative large language model is larger than that of the large language model. Then, based on the preset model, the server discriminates the second training data to be processed and determines the validity of the second training data to be processed. Finally, within the second preset validity determination count threshold, when the server determines that the second training data to be processed is valid, the second training data to be processed is determined as the target training data. In this way, when the first training data to be processed cannot meet the validity requirements of the preset model, using an alternative large language model with a larger parameter scale can generate higher-quality second training data to be processed.
[0219] Please refer to Figure 8 , in some embodiments, the method further includes:
[0220] 0247: Within the second preset validity determination count threshold, when it is determined that the second training data to be processed is invalid, based on the preset rule, correct the second training data to be processed and determine the target training data.
[0221] In some embodiments, the determination module is used to, within the second preset validity determination count threshold, when it is determined that the second training data to be processed is invalid, based on the preset rule, correct the second training data to be processed and determine the target training data.
[0222] In some embodiments, the processor is further used to, within the second preset validity determination count threshold, when it is determined that the second training data to be processed is invalid, based on the preset rule, correct the second training data to be processed and determine the target training data.
[0223] Specifically, the preset rule refers to the rule for correcting invalid data. In some embodiments, the preset rule is formulated according to expert knowledge or experience, such as judging whether the application interface selection is correct and whether the parameter filling is reasonable.
[0224] Please refer to Figure 9 , Figure 9 which is a schematic flow diagram for generating target training data. When the second training data to be processed (generated by the alternative large language model) is invalid within the second preset validity determination count threshold, the second training data to be processed is corrected using a preset rule. The second training data to be processed corrected by the preset rule is the final target training data.
[0225] In this way, within the second preset validity determination count threshold, when it is determined that the second training data to be processed is invalid, based on the preset rule, the server corrects the second training data to be processed to determine the target training data. In this way, by correcting the second training data to be processed using the preset rule, the accuracy of the final target training data can be ensured, and the use of incorrect or low-quality training data can be avoided.
[0226] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the model training method of the large language model as described above are implemented.
[0227] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution media, etc.
[0228] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials, or characteristics described in combination with the embodiments or examples are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0229] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of executable code that includes one or more steps for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0230] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A model training method for a large language model, characterized in that, The method includes: Determine an associated application interface list according to the target voice request training data; Based on a preset prompt template, determine target training data according to the associated application interface list and the target voice request training data; Train the large language model according to the target training data.
2. The method according to claim 1, characterized in that The determining the associated application interface list according to the target voice request training data includes: Perform encoding processing on the target voice request training data to determine a target vector; According to the target vector, determine the associated application interface list from a pre-constructed application interface database, where the application interface database includes application interfaces and interface parameters corresponding to the application interfaces.
3. The method according to claim 1, wherein The determining the target training data based on the preset prompt template, according to the associated application interface list and the target voice request training data, includes: According to the first preset sub-prompt word in the preset prompt template, the associated application interface list, and the target voice request training data, determine a first target sub-prompt word, where the first target sub-prompt word is configured to guide the target large language model to determine a target thinking logic chain according to the associated application interface list and the target voice request training data; According to the second preset sub-prompt word in the preset prompt template, the associated application interface list, and the target voice request training data, determine a second target sub-prompt word, where the second target sub-prompt word is configured to guide the target large language model to determine the adaptability of the associated application interface list to the target voice request training data according to the target thinking logic chain; According to the third preset sub-prompt word in the preset prompt template, the associated application interface list, and the target voice request training data, determine a third target sub-prompt word, where the third target sub-prompt word is configured to guide the target large language model to determine a natural language processing result according to the associated application interface list and the target voice request training data when the associated application interface list adapts to the target voice request training data; Determine the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word.
4. The method according to claim 3, characterized in that, The determining the target training data according to the target voice request training data, the first target sub-prompt word, the second target sub-prompt word, and the third target sub-prompt word includes: Based on the large language model, determine first training data to be processed according to the first target sub-prompt word, the second target sub-prompt word, the third target sub-prompt word, and the target voice request training data; Based on a preset model, determine the validity of the first training data to be processed; Within a first preset validity determination count threshold, when it is determined that the first training data to be processed is valid, determine the first training data to be processed as the target training data.
5. The method according to claim 4, characterized in that, Based on the large language model, determining the first training data to be processed according to the first target sub-prompt, the second target sub-prompt, the third target sub-prompt, and the target voice request training data, includes: According to the target voice request training data, obtaining vehicle knowledge associated with the target voice request training data from a pre-constructed vehicle knowledge base; Determining the target thinking logic chain according to the vehicle knowledge and the first target sub-prompt; Determining the adaptability of the associated application interfaces in the associated application interface list to the target voice request training data according to the second target sub-prompt; Determining the natural language processing result according to the adaptability and the third target sub-prompt; Determining the first training data to be processed according to the target voice request training data, the target thinking logic chain, the preset prompt template, and the natural language processing result.
6. The method according to claim 5, wherein The associated application interface list includes associated interface parameters corresponding to the associated application interfaces. The determining the natural language processing result according to the adaptability and the third target sub-prompt includes: When the associated application interface is applicable to the target voice request training data, performing slot recognition on the target voice request training data to determine slot entities; Performing application interface parameter filling according to the associated interface parameters and the slot entities to determine the application interface parameter filling result; Determining the natural language processing result according to the associated application interface and the application interface parameter filling result.
7. The method according to claim 4, wherein The method further includes: When it is determined that the first training data to be processed is invalid within the first preset validity determination times threshold, based on an alternative large language model, determining the second training data to be processed according to the first target sub-prompt, the second target sub-prompt, the third target sub-prompt, and the target voice request training data, where the parameter scale of the alternative large language model is larger than that of the large language model; Performing discrimination processing on the second training data to be processed based on a preset model to determine the validity of the second training data to be processed; When it is determined that the second training data to be processed is valid within the second preset validity determination times threshold, determining the second training data to be processed as the target training data.
8. The method according to claim 7, characterized in that, The method further includes: When it is determined that the second training data to be processed is invalid within the second preset validity determination times threshold, performing correction processing on the second training data to be processed based on a preset rule to determine the target training data.
9. A server, characterized in that, The server includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the method according to any one of claims 1-8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.