Model training method, vehicle control method, device and related equipment
By training a lightweight student model and building a target model to handle ambiguous or incomplete user instructions, the problem of inaccurate vehicle response is solved and accurate user interaction is achieved.
Patent Information
- Application Number
- CN202511030037.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-25
AI Technical Summary
With the diversification of vehicle functions, users need to remember a large number of operating instructions. Ambiguous or incomplete instructions result in low vehicle response accuracy or no response, and existing technologies make it difficult to achieve accurate user interaction.
By training the lightweight student model, building a target model, using basic instructions to generate fuzzy instructions and multi-round dialogue texts, and building a training set, the training model has the ability to understand intent and structure control instructions, thereby improving response accuracy.
It enables accurate response to control commands with incomplete or unclear user input, improving the efficiency and accuracy of vehicle user interaction.
Smart Images

Figure CN120523921B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a model training method, a vehicle control method, a device and related equipment. Background Art
[0002] With the rapid development of vehicle intelligence, modern vehicles now have over hundreds of controllable functions, ranging from basic air conditioning to complex driver assistance systems. This diverse functionality offers the potential for enhancing the driving experience and safety, but it also brings new challenges, particularly regarding the complexity of user-vehicle interactions.
[0003] In traditional vehicle control, users typically operate specific functions through physical buttons or simple voice commands. However, as the number of functions increases, users need to remember a large number of commands and operating methods, which undoubtedly increases the learning and usage burden. At the same time, in actual use, users often use vague or incomplete commands, such as "cool me down" or "play music." These vague or incomplete commands make it difficult for the system to respond accurately, for example, resulting in low response accuracy or no response. Summary of the Invention
[0004] The embodiments of the present application provide a model training method, a vehicle control method, an apparatus and related equipment. By training a lightweight student model, a target model is obtained that has the ability to understand the intention of the control instructions input by the user and obtain structured control instructions. The target model is configured on the vehicle, and the vehicle can obtain complete structured control instructions based on the target model, which facilitates accurate response to the control instructions input by the user.
[0005] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0006] Get basic instructions;
[0007] Performing intent recognition on the basic instruction to obtain at least one intent and multiple slots corresponding to each intent;
[0008] For each of the intentions, processing the multiple slots corresponding to the intention according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intention;
[0009] Generating, based on each of the fuzzy instructions and the dialogue state record, a first multi-round dialogue text corresponding to each of the fuzzy instructions, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes a plurality of question-answer pairs;
[0010] Constructing a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of conversation text, wherein the training set includes a plurality of training samples, and each of the training samples includes one of the fuzzy instructions and the first multiple rounds of conversation text corresponding to the fuzzy instruction;
[0011] The training set is used to train the student model to obtain a target model, and the target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
[0012] In a second aspect, an embodiment of the present application provides a vehicle control method, the method comprising:
[0013] receiving target control instructions;
[0014] Inputting the target control instruction and historical multi-round dialogue text into a target model to obtain a structured control instruction, wherein the target model is obtained according to the model training method described in the first aspect, and the historical multi-round dialogue text is dialogue text obtained through interaction with the user before the target control instruction;
[0015] The vehicle is controlled according to the structured control instructions.
[0016] In a third aspect, an embodiment of the present application provides a model training device, comprising:
[0017] A first acquisition module is used to acquire basic instructions;
[0018] an identification module, configured to perform intent recognition on the basic instruction to obtain at least one intent and a plurality of slots corresponding to each intent;
[0019] A second acquisition module is configured to process, for each of the intentions, a plurality of slots corresponding to the intention according to the first fuzzy instruction generation strategy to obtain a plurality of fuzzy instructions corresponding to the intention;
[0020] a first generating module configured to generate, based on each of the fuzzy instructions and a dialogue state record, a first multi-round dialogue text corresponding to each of the fuzzy instructions, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes a plurality of question-answer pairs;
[0021] A construction module is configured to construct a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of dialogue text, wherein the training set includes a plurality of training samples, and each of the training samples includes one of the fuzzy instructions and the first multiple rounds of dialogue text corresponding to the fuzzy instruction;
[0022] The training module is used to train the student model using the training set to obtain a target model, and the target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
[0023] In a fourth aspect, an embodiment of the present application provides a vehicle control device, the device comprising:
[0024] A receiving module, used for receiving target control instructions;
[0025] an acquisition module, configured to input the target control instruction and historical multi-round dialogue text into a target model to obtain a structured control instruction, wherein the target model is obtained according to the model training method described in the first aspect, and the historical multi-round dialogue text is dialogue text obtained through interaction with the user before the target control instruction;
[0026] A control module is used to control the vehicle according to the structured control instructions.
[0027] In a fifth aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the model training method described in the first aspect or the vehicle control method described in the second aspect.
[0028] In a sixth aspect, an embodiment of the present application provides a vehicle, which includes the electronic device described in the fifth aspect.
[0029] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the model training method as described in the first aspect or the vehicle control method as described in the second aspect is implemented.
[0030] In an eighth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the model training method as described in the first aspect, or the electronic device executes the vehicle control method as described in the second aspect.
[0031] In this embodiment, fuzzy instructions are constructed based on basic instructions, and the first multiple rounds of dialogue texts are constructed according to the fuzzy instructions, so as to construct a training set to train the student model, so that the trained target model has the ability to understand the intention of the input instructions and obtain structured control instructions. The student module can use a lightweight model to configure the trained target model on the vehicle. The vehicle can obtain complete structured control instructions according to the target model, which is convenient for accurately responding to the control instructions input by the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0033] Figure 1 This is a flow chart of the model training method provided in the embodiment of the present application;
[0034] Figure 2 This is a flow chart of a vehicle control method provided by an embodiment of the present application;
[0035] Figure 3 This is another flowchart of the model training method provided in the embodiment of the present application;
[0036] Figure 4 Schematic diagram of the structure of the model training device provided in the embodiment of the present application;
[0037] Figure 5 is a schematic structural diagram of a vehicle control device provided in an embodiment of the present application;
[0038] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0040] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0041] In each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0042] In order to solve the problems of the prior art, the embodiments of the present application provide a model training method, a vehicle control method, an apparatus and related equipment. The model training method provided by the embodiments of the present application is first introduced below.
[0043] Figure 1 FIG1 shows a flow chart of a model training method provided by an embodiment of the present application. Figure 1 As shown, the model training method provided in the embodiment of the present application includes the following steps 101 to 106, wherein:
[0044] Step 101: Get basic instructions.
[0045] Basic instructions refer to natural language instructions used to control the vehicle. Basic instructions can be understood as instructions with clear intentions or completeness. Basic instructions must include clear control objects and control parameters. For example, the instruction "Please fully open the window on the main driver's side of the vehicle". In this instruction, the control object "the window on the main driver's side of the vehicle" is clearly referred to, and the control parameter "fully open" is clearly stated. This instruction can be used as a basic instruction. For another example, for the instruction "Please open the window on the main driver's side of the vehicle", since this instruction only states that the window is opened, the degree of opening is uncertain, that is, the control parameters are uncertain, this instruction cannot be used as a basic instruction.
[0046] The basic instructions can be text instructions or voice instructions, which are not limited here.
[0047] Step 102: perform intent recognition on the basic instruction to obtain at least one intent and multiple slots corresponding to each intent.
[0048] Perform intent recognition on basic instructions to obtain at least one intent. Intent recognition is to determine what the user wants to do. Intent refers to the action or goal expressed by the user through natural language. For example, the basic instruction "Please fully open the window on the main driver's side of the vehicle at 8 am every day" has the intention of opening the window on the main driver's side. This basic instruction includes one intent; for another example, the basic instruction "Please fully open the window on the main driver's side of the vehicle, and then fully open the sunroof of the vehicle" has two intentions, namely opening the window on the main driver's side and opening the sunroof.
[0049] Each intent has multiple slots. For example, for the intent to open the driver's window, "Driver's Window" is a slot for the control object, "Fully Open" is a slot for the control parameter, and "Everyday at 8:00 AM" is a slot for the time. Each slot has a specific function and is used to determine a specific type of information.
[0050] Step 103: For each of the intentions, multiple slots corresponding to the intention are processed according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intention.
[0051] A fuzzy instruction refers to an instruction that is not clearly stated or is incomplete. Exemplarily, the first fuzzy instruction generation strategy can be used to delete at least one of the multiple slots corresponding to the intention. By deleting key semantic components, the fuzzy expression of the user in the real scene can be simulated to construct a robust training sample for the student model. That is, for each intent, one or more of the multiple slots corresponding to the intent can be deleted to obtain the corresponding fuzzy instruction. For example, for the intention of opening the window on the main driver's side mentioned above, the control object slot can be deleted to obtain the fuzzy instruction "fully open the vehicle at 8 am every day", or the control object slot and the time slot can be deleted to obtain the fuzzy instruction "fully open the vehicle".
[0052] Step 104: Generate a first multi-round dialogue text corresponding to each fuzzy instruction based on each fuzzy instruction and the dialogue state record, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes multiple question-answer pairs.
[0053] The first multi-round dialogue text may include M question-answer pairs, where M can be selected based on actual circumstances, for example, 4 or 5, and is not limited here.
[0054] The conversation state record includes the user intent state, system response state, slot filling information, and interaction history sequence for each conversation turn in the conversation context. The user intent state for each conversation turn indicates whether the user's input information is complete, which slots are missing, and whether the intent is clear. The system response state records whether a user's counter-question receives a clear response. The slot filling information records whether the slot corresponding to each intent has been filled, and if so, the slot value. The interaction history sequence includes the encoding of each conversation turn and the timing of each conversation turn. The conversation state record can be viewed as a continuously updated conversation "memory board," updated after each conversation turn. The conversation state record allows users to determine which slots have been filled and which slots are still unclear. Slot filling is also a crucial component of natural language understanding, similar to sequence labeling, used to extract key information from user input. Slot filling is the process of completing information to translate user intent into clear instructions.
[0055] The first multiple rounds of dialogue texts may be generated by manual annotation or by using a large language model, which is not limited here.
[0056] Step 105: construct a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of dialogue texts, wherein the training set includes multiple training samples, and each of the training samples includes one of the fuzzy instructions and the first multiple rounds of dialogue texts corresponding to the fuzzy instruction.
[0057] Step 106: Use the training set to train the student model to obtain a target model. The target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
[0058] The student model is trained using training samples from the training set until the iteration stopping condition is met. The student model obtained from the last round of training is used as the target model. The student model can use a lightweight neural network model to avoid large resource consumption and facilitate deployment on the vehicle.
[0059] In the above description, a structured control instruction refers to an instruction with a preset format. The preset format can be set according to actual circumstances. For example, the preset format is "Please adjust XX to YY," where XX is the control object and YY is the control parameter. Other preset formats are also possible and are not limited here. A structured control instruction includes a clear control object and control parameters. The target model understands the intent of the input instruction (which may lack the control object and control parameters) and fills in the control object and / or control parameters to obtain a structured control instruction.
[0060] By training the student model with fuzzy instructions and the corresponding first multiple rounds of dialogue texts, the student model can perform logical reasoning based on the first multiple rounds of dialogue texts to fill in the missing slots in the fuzzy instructions and obtain structured control instructions. Through continuous training, the student model's ability to understand intentions can be improved, and the accuracy of slot filling can be improved, thereby making the output structured control instructions more accurate.
[0061] In this embodiment, fuzzy instructions are constructed based on basic instructions, and the first multiple rounds of dialogue texts are constructed according to the fuzzy instructions, so as to construct a training set to train the student model, so that the trained target model has the ability to understand the intention of the input instructions and obtain structured control instructions. The vehicle can obtain complete structured control instructions according to the target model, so as to accurately respond to control instructions input by the user with incomplete or unclear intention expressions.
[0062] In one embodiment of the present application, the first fuzzy instruction generation strategy includes weights of the plurality of slots corresponding to the intention;
[0063] The method of training the student model using the training set to obtain a target model includes:
[0064] After multiple rounds of training of the student model using the training set, if the student model meets a preset iteration stop condition, the student model obtained from the last round of training is used as the target model;
[0065] If the student model does not meet the iteration stopping condition, inputting the verification samples in the verification set into the student model to obtain a second structured control instruction corresponding to each verification sample, wherein the verification sample includes a verification fuzzy instruction and multiple rounds of dialogue text corresponding to the verification fuzzy instruction;
[0066] If there is a second structured control instruction that does not meet the preset condition, adjusting the weights of the plurality of slots corresponding to the intention in the first fuzzy instruction generation strategy to obtain a second fuzzy instruction generation strategy, wherein the preset condition is determined according to a preset compliance requirement;
[0067] The second fuzzy instruction generation strategy is used as the first fuzzy instruction generation strategy, and the step of processing multiple slots corresponding to each intention according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intention is executed.
[0068] In this embodiment, the iteration stop condition can be set according to actual conditions, for example, the total number of model training times reaches the maximum number of training times, or the student model's intention recognition error rate for the constructed fuzzy instructions is lower than the first set threshold (such as 5%), or the student model's recognition accuracy for high-frequency slots such as temperature and seat position in the verification set exceeds the second set threshold (such as 95%), and the rate of change in the past N rounds (such as N is 3) is less than the third set threshold (such as 1%), etc.
[0069] The first fuzzy instruction generation strategy also includes a slot fuzzy strategy, which is used to delete the slots in the intent whose weight is less than a first threshold. During initialization, the weights of multiple slots corresponding to the intent are the same. The first threshold can be set according to actual conditions and is not limited here. If the weight of a slot is greater than the first threshold, the slot will be retained and will not be deleted during fuzzy processing.
[0070] After the student model is trained for multiple rounds using the training set, for example, 5 rounds of training, the number of training times can be set according to actual conditions and is not limited here. If the student model meets the preset iterative stop condition, the student model obtained in the last round of training is used as the target model; if the student model does not meet the preset iterative stop condition, the student model needs to continue to be trained. In this case, the verification sample in the verification set is input into the student model to obtain a second structured control instruction corresponding to each verification sample. The verification sample includes a verification fuzzy instruction and multiple rounds of dialogue text corresponding to the verification fuzzy instruction, wherein the verification fuzzy instruction can also be obtained based on the basic instruction. The training samples in the training set are divided into two parts, one part is used to train the student model, and the other part is used as a verification sample to verify the training effect of the student model.
[0071] If there is a second structured control instruction that does not meet the preset conditions, the weights of the multiple slots corresponding to the intention in the first fuzzy instruction generation strategy are adjusted to obtain a second fuzzy instruction generation strategy. For example, the weight of the slots whose false positive rate is higher than the average false positive rate is increased, or the weight of the slots whose false positive rate is higher than the preset false positive rate threshold is increased, where the false positive rate refers to the probability of a slot being incorrectly identified in multiple rounds of training, and the average false positive rate refers to the average value of the probabilities of all slots being incorrectly identified in multiple rounds of training. The first fuzzy instruction generation strategy is updated using the second fuzzy instruction generation strategy, and the updated first fuzzy instruction generation strategy is used to regenerate the training set to train the student model, forming a closed-loop iterative training until the iteration termination condition is met and the target model is obtained.
[0072] In this embodiment, the first fuzzy instruction generation strategy is adjusted by the output results of the student model on the verification set to generate a new training set, so as to dynamically adjust the training sample distribution according to the output results of the student model, realize the purpose of co-evolution of data and model, and gradually improve the robustness of the student model in fuzzy instruction parsing.
[0073] In the above, the preset condition is determined according to at least one of the following:
[0074] a value range of a control parameter corresponding to the control object in the second structured control instruction;
[0075] a control object in the second structured control instruction;
[0076] Format requirements of the second structured control instruction.
[0077] The second structured control instruction is verified for format legality, parameter range, safety boundary and other protocols through preset conditions to determine whether there is a slot filling error or excessive ambiguity. If so, the first fuzzy instruction generation strategy is adjusted to generate a new training set, thereby dynamically adjusting the training sample distribution according to the output results of the student model, realizing the co-evolution of data and model, and gradually improving the robustness of the student model in parsing fuzzy instructions.
[0078] In another embodiment of the present application, the weights of the plurality of slots corresponding to the intent in the first fuzzy instruction generation strategy are adjusted to obtain a second fuzzy instruction generation strategy, including:
[0079] Counting the number of slots that are incorrectly recognized and the number of slots that are missing in the second structured control instruction in each round of training;
[0080] Increase the weight of the missing slots, and increase the weight of the slots whose misjudgment rate is higher than the average misjudgment rate, to obtain a second fuzzy instruction generation strategy, where the misjudgment rate refers to the probability of a slot being incorrectly identified in multiple rounds of training, and the average misjudgment rate refers to the average value of the probabilities of all slots being incorrectly identified in multiple rounds of training.
[0081] Specifically, the trained student model will fill in the missing slots in the verification fuzzy instructions based on the multi-round dialogue texts corresponding to the verification fuzzy instructions to obtain the second structured control instructions. Referring to the standard instructions, the missing slots in the second structured control instructions can be determined, as well as the incorrect slots can be identified. Among them, the verification fuzzy instructions are obtained based on the standard instructions. The standard instructions must include clear control objects and control parameters. The standard instructions and the basic instructions can be the same instructions or different instructions, which are not limited here.
[0082] Count the missing slots in the second structured control instruction in each round of training. For the missing slots, the weight of the slot can be increased. For example, if the parameter slot of "temperature" is often missing, it means that the slot has a greater impact on the performance of the student model. The student model is highly sensitive to the slot, and the weight of the slot needs to be increased. The degree of increase can be set according to the actual situation. For example, it can be increased to the first threshold to avoid the slot from being deleted during fuzzy processing (fuzzy processing refers to generating fuzzy instructions using the first fuzzy instruction generation strategy). There is no limitation here.
[0083] Count the slots that are misidentified in the second structured control instruction in each round of training. If the parameter slot "temperature" has a recognition error once in 5 rounds of training, the misjudgment rate is 20%. Count the misjudgment rate of each slot and calculate the average misjudgment rate to obtain the average misjudgment rate. For each slot, compare the misjudgment rate of the slot with the average misjudgment rate. If the misjudgment rate is higher than the average misjudgment rate, it means that the slot is highly sensitive and the weight of the slot needs to be increased. In this way, when the slots with weights less than the first threshold in the intention are deleted according to the slot fuzzy strategy, the slots with weights greater than or equal to the first threshold can be avoided from being deleted. If the misjudgment rate is lower than the average misjudgment rate, it means that the slot is less sensitive and can be deleted during fuzzy processing.
[0084] It should be noted that in the process of training the student model, the student model is constantly learning, and the error rate of a slot with a weight greater than or equal to the first threshold may be lower than the average error rate, indicating that the student model is less sensitive to the slot and has a strong understanding of it. When performing fuzzy processing, the slot can be deleted to enhance the student model's adaptability to extreme scenarios. In this case, the weight of the slot can be reduced, for example, the weight of the slot can be reduced to a value less than the first threshold.
[0085] In this embodiment, the weights of the slots are adjusted by counting the incorrectly identified slots and missing slots in the second structured control instructions in each round of training, thereby adjusting the first fuzzy instruction generation strategy. For slots with low sensitivity of the student model, higher weights will be set to reduce their missing probability in fuzzy processing, so as to ensure the integrity of key semantic information.
[0086] In one embodiment of the present application, generating the first multi-round dialogue text corresponding to each fuzzy instruction according to each fuzzy instruction and the dialogue state record includes:
[0087] For each of the fuzzy instructions, a pre-trained large language model (LLM) is used to generate first multiple-round dialogue text corresponding to the fuzzy instruction based on the dialogue state record and a constraint condition, where the constraint condition includes at least one of the following:
[0088] (1) Structural limitations
[0089] Restrictions: Each round of dialogue sample must contain a question-answer pair consisting of a ut (user speech) and a st (system counter-question or feedback). In other words, each round of dialogue text includes a question-answer pair consisting of the user speech text and the system feedback text.
[0090] Purpose: Ensure the standardized structure of training samples to facilitate supervised training and context state modeling of downstream models.
[0091] (2) Semantic constraints
[0092] Constraints: Each round of user speech ut must revolve around the original fuzzy intent and must not deviate from the main semantic path. That is, the semantic similarity between the user speech text in the question-answer pair in each round and the fuzzy instruction is greater than a second threshold.
[0093] Purpose: To prevent topic drift during the generated multiple rounds of question-and-answer sessions. For example, if the user originally intended to "open a window," the conversation might generate content about "playing music," affecting the clarity of the instruction chain.
[0094] (3) Protocol verification
[0095] Restrictions: The control objects and slot parameters involved in all generated content must comply with the value range and structure specifications in the vehicle control protocol document. That is to say, the control objects included in each round of question and answer pairs belong to the objects supported by the vehicle, and the control parameters of the control objects included in each round of question and answer pairs are within the control parameter value range set for the control objects; for example, the control parameters of the control object "air conditioning" should be between 15 degrees and 35 degrees, and 15 degrees to 35 degrees is the value range of the control parameters set for the control object "air conditioning".
[0096] Purpose: To ensure that the generated training samples are executable when deployed in the future, and to avoid functional errors or hardware risks caused by illegal parameters.
[0097] (4) Limitation on the number of rounds
[0098] Constraints: Each conversation must contain a maximum of 3-4 question-answer rounds (i.e., a maximum of 4 question-answer pairs) to avoid generating lengthy and ineffective interaction chains. In other words, the number of question-answer pairs in the first round of conversation must fall within a pre-set range.
[0099] Purpose: Limiting the length of conversations can prevent distractions during student model training and improve the efficiency of modeling key supplementary questions and answers.
[0100] (5) Security policy restrictions
[0101] Restriction: No high-risk operation instructions that violate functional safety rules are allowed, such as "open the door while driving" or "automatically unlock." In other words, each round of question and answer must comply with the preset vehicle functional safety usage rules.
[0102] Purpose: To comply with functional safety standards such as ISO 26262, prevent training data from misleading the model into generating high-risk operations, and ensure system security after model deployment.
[0103] (6) Style control restrictions
[0104] Constraints: The system's language style must be consistent, for example, starting each question with "Please ask..." and avoiding expressions like "You want to..." In other words, the text format of each question-answer pair must be the same in each round.
[0105] Purpose: Unifying the speech style can enhance the model's ability to model system roles, help converge the language style during training, and improve generation quality.
[0106] In the above, the LLM is pre-trained so that the LLM has the ability to generate the first multiple rounds of dialogue texts based on the dialogue state records, constraints and fuzzy instructions.
[0107] In this embodiment, the first multiple rounds of conversation texts are generated by the LLM large model, which can improve the efficiency of obtaining the first multiple rounds of conversation texts. In addition, LLM generates the first multiple rounds of conversation texts in accordance with the restriction conditions, which can ensure that the data generation format is standardized, the semantics are reasonable, the protocol is legal, the security is compliant, and the style is unified, providing high-quality training samples for the training of the student model.
[0108] In one embodiment of the present application, a method for generating first-round dialogue texts by supplementing questions and answers is further provided, that is, asking questions and receiving user answers according to target slots, thereby generating question-answer pairs. Specifically, generating first-round dialogue texts corresponding to each fuzzy instruction according to each fuzzy instruction and the dialogue state record includes:
[0109] For each of the fuzzy instructions, slot identification is performed on the fuzzy instruction according to the dialogue state record to obtain a target slot, and first multi-round dialogue texts corresponding to the fuzzy instruction are obtained according to at least one of the following:
[0110] (1) generating a first question instruction for the target slot (e.g., a missing slot) based on the user's historical preferences, receiving a first answer instruction, and constructing a question-answer pair based on the first question instruction and the first answer instruction;
[0111] For example, the fuzzy instruction: "Turn on the air conditioner" lacks key information such as "control object slot (which row of air conditioners)" and "parameter slot (temperature)". The dialogue state record is used to determine the filling status of the missing slots in the fuzzy instruction. The air conditioning area, temperature, and wind speed in the dialogue state record are all empty. In other words, the fuzzy instruction cannot be filled in the slot through the dialogue state record. In this case, for the target slot, the first question instruction can be generated based on the user's historical preferences. For example, if it has been set to 22°C many times in the history, a default completion suggestion can be made to generate the first question instruction "Is it set to 22°C?" The user can answer "yes" or "no" or "20°C" as the first answer instruction. The first question instruction and the first answer instruction constitute a question-answer pair. In the process of filling the target slot according to the user's historical preferences, multiple rounds of dialogue can also be conducted to generate multiple question-answer pairs;
[0112] (2) Generate a second question instruction based on the current state of the vehicle and the target slot, receive a second answer instruction, and construct a question-answer pair based on the second question instruction and the second answer instruction.
[0113] The current state of the vehicle may refer to the driving state of the vehicle, for example, "driving" or "not driving", or may refer to the usage state of a seat in the vehicle, for example, no one is sitting in the "passenger seat" and the "passenger seat" is not in use.
[0114] For example, if it is detected that only the driver's seat is in use in the vehicle, questions about the passenger seat can be automatically omitted.
[0115] For another example, if the current temperature in the vehicle is 22°C, the second question instruction generated is "The air conditioner is set to 22°C, does it need to be adjusted?" The user can answer "yes" or "no" or "set to 20°C" as the second answer instruction.
[0116] In the process of filling the target slot according to the current status of the vehicle, multiple rounds of dialogue can be conducted to generate multiple question-answer pairs;
[0117] (3) Generate a third question instruction only according to the target slot, receive a third answer instruction, and construct a question-answer pair according to the third question instruction and the third answer instruction.
[0118] In this case, the third question instruction is generated directly according to the target slot. For example, for the fuzzy instruction: "Turn on the air conditioner", the third question instruction is directly generated: "Which row of air conditioners needs to be set?", "Is it set to 22 degrees as last time?", "What is the wind speed?" These three questions can be output at one time and wait for the user to answer, or they can be output in sequence, that is, one question instruction is output each time, and after receiving the user's clear answer, the next question instruction is output. In this way, multiple question and answer pairs are constructed.
[0119] For example, a user enters the ambiguous command "Let me cool down a bit." The system (which can be understood as the device constructing the first round of multi-round dialogue text) identifies the intention to turn on the air conditioner through semantic analysis, but lacks a temperature parameter slot. The conversation state record reveals that the temperature parameter slot is missing, generating the first round of questioning: "Do you want to set a specific temperature?" The user responds, "I'll set it to 22 degrees." The conversation state record is updated, and the temperature parameter slot is filled. The system then checks to see if there are any missing control object slots, such as the air conditioning zone (driver's / passenger's). The system then asks a further question, generating the second round of questioning: "Which row do you want to turn on the air conditioner?" The user responds, "Front row." The conversation state record is updated, and the control object slot is filled. In this process, each round of questioning and user response constitutes a question-answer pair. Multiple question-answer pairs are used to generate the first round of multi-round dialogue text.
[0120] In the above, question instructions are generated for the target slot and reply instructions are received, thereby constructing the first multi-round dialogue text corresponding to the fuzzy instructions. In this process, logical deduction and dynamic questioning are adopted to improve the quality of the first multi-round dialogue text, thereby providing high-quality training samples for the training of the student model.
[0121] In one embodiment of the present application, a method for constructing a multi-round dialogue text is further provided. Specifically, before the student model is trained using the training set to obtain the target model, the method further includes:
[0122] Generalizing the basic instruction to obtain a plurality of counter-instructions, wherein the plurality of counter-instructions include instructions having the same intention as the basic instruction and instructions having the opposite intention to the basic instruction;
[0123] For each of the confrontation instructions, generating a second multi-round dialogue text corresponding to the confrontation instruction;
[0124] Each of the adversarial instructions and the corresponding second-round dialogue texts are added to the training set as training samples.
[0125] In this embodiment, when generalizing basic instructions, the control objects and control parameters of the basic instructions can be generalized to obtain multiple adversarial instructions. For example, by generalizing the control objects in a diversified manner (such as generalizing from "air conditioning" to "refrigeration system"), the student model's ability to understand multiple expressions can be improved.
[0126] By analyzing the basic instructions, fuzzy trigger points (such as fuzzy references, abstract adjectives, user habitual phrases, etc.) are determined, the fuzzy trigger points are generalized, and fuzziness is generalized as the dominant feature to obtain fuzzy generalized instructions with potential ambiguity. The fuzzy generalized instructions include instructions with the same intention as the basic instructions, as well as instructions with the opposite intention to the basic instructions.
[0127] When generating the second multiple rounds of dialogue texts according to the confrontation instructions, the method of generating the first multiple rounds of dialogue texts according to the fuzzy samples in the above embodiment can be adopted, which will not be described in detail here.
[0128] In the generalization of this embodiment, an "intention reversal" strategy is introduced to generate generalized instructions with the semantic antonym of the basic instructions as the goal (that is, to obtain instructions that are opposite to the intention of the basic instructions), thereby introducing semantic ambiguity interference to the student model during the training stage, and improving the robustness of the student model in fuzzy instruction discrimination and intention recovery.
[0129] Furthermore, by not only generating adversarial commands that reverse intent but also constructing potentially ambiguous fuzzy generalization commands, the student model's generalization ability can be improved in borderline samples and with highly confusing semantics. This allows for a more systematic simulation of the ambiguous expressions and ambiguous behaviors seen in real user interactions, leading to higher generalization and robustness in fuzzy command understanding tasks.
[0130] Figure 2 The flow chart of the vehicle control method provided in the embodiment of the present application is as follows: Figure 2 The vehicle control method provided in the embodiment of the present application can be applied to a vehicle and includes the following steps 201 to 203, wherein:
[0131] Step 201: Receive a target control instruction.
[0132] Target control commands can be user voice or text commands captured by the vehicle. Target control commands are commands described in natural language. Target control commands can be vague or incomplete in their intent, such as "turn up a bit," which lacks a clear control target or parameters.
[0133] In step 202, the target control command and historical multi-turn dialogue text are input into a target model to obtain a structured control command. The historical multi-turn dialogue text is the dialogue text obtained through interaction with the user before the target control command. The target model is pre-configured on the vehicle. The target model is capable of interpreting the input target control command's intent to obtain a structured control command.
[0134] The historical multi-round dialogue texts meet at least one of the following conditions:
[0135] (1) Collecting historical conversations within a recent preset duration (e.g., 60 seconds) can avoid introducing outdated semantic information;
[0136] (2) Retain at most the most recent J rounds of valid questions and answers (for example, J can be 4 or 5) to reduce the burden on the student model to process redundant information while ensuring the integrity of the intent-taking logic.
[0137] (3) Only the dialogue texts that are semantically related to the target control instruction intent are retained, which can be filtered using embedding similarity or intent label matching.
[0138] For example, when the user says "drive a little further", the vehicle (specifically, it can be an on-board device configured on the vehicle) will look back to see whether the previous rounds of conversation text mentioned keywords such as windows, sunroof, air conditioning, etc. to determine the context-dependent content of the current "a little further" command.
[0139] If at least one of the control object and control parameters included in the target control instruction is unclear, the target model can fill in the missing slots in the target control instruction based on historical multi-round dialogue texts to obtain a structured control instruction.
[0140] A structured control instruction is an instruction with a preset format. This preset format can be set based on actual circumstances. For example, the preset format is "Please adjust XX to YY," where XX is the control object and YY is the control parameter. Other preset formats are also possible and are not limited here. A structured control instruction includes a clear control object and control parameters. These control objects and control parameters are obtained by the target model understanding the intent of the input target control instruction and filling in the control object slot and / or control parameter slot.
[0141] Step 203: Control the vehicle according to the structured control instruction.
[0142] Since the control object and control parameters in the structured control instruction are clear, the structured control instruction can be used to control the vehicle. For example, if the structured control instruction is "please open the sunroof completely", the vehicle will execute the action of opening the sunroof completely.
[0143] It should be noted that the target control instruction can also be an instruction with clear intention or completeness. In this case, the target control instruction includes a clear control object and clear control parameters. For example, the target control instruction is "turn on the air conditioner of the main driver's seat to 20℃", the control object "the air conditioner of the main driver's seat" is clear and unambiguous, and the control parameter "turn on the air conditioner to 20℃" is clear and unambiguous.
[0144] In this embodiment, the target model can be used to understand the intention of the target control instruction, and the target slots (such as missing slots) in the target control instruction can be filled after logical reasoning based on historical multi-round dialogue texts to obtain structured control instructions. The structured control instructions are then used to control the vehicle, and accurate responses can be given to control instructions with incomplete or unclear intention expressions input by the user.
[0145] It should be noted that in some cases, such as when the historical multi-round dialogue text contains too little information, the target model cannot fill the target slot in the target control instruction based on the historical multi-round dialogue text, or fill it correctly. In this case, the structured control instruction cannot be executed or legally executed in the vehicle. To ensure that the structured control instruction can be legally executed in the vehicle, the structured control instruction can be subjected to protocol verification, which includes at least one of the following:
[0146] (1) Structural integrity verification:
[0147] Check structured control instructions for missing or redundant slot fields;
[0148] Verify whether the structured control instructions meet the preset protocol template definition.
[0149] (2) Value range validity check:
[0150] Check whether the slot values of each slot in the structured control instructions are within the range specified by the protocol (for example, the temperature should be between 16°C and 30°C);
[0151] If the slot value is illegal, it is marked as "illegal instruction".
[0152] (3) Context consistency check:
[0153] For relative instructions such as "raise it a little bit more", it is necessary to verify whether the target is reasonable in combination with historical status.
[0154] (4) Vehicle current state constraint verification:
[0155] If the vehicle is currently traveling at high speed, the protocol will prohibit the execution of certain instructions (such as "open the door"), which is a scene-sensitive constraint.
[0156] After the above verification, the protocol verification result is obtained, which includes: whether the verification passes; and the specific reasons when it fails, such as field missing, value exceeding the limit, security conflict, etc.
[0157] If the protocol verification result indicates that the structured instruction contains potential risks or illegal content, security adjustments must be made to obtain new structured control instructions to ensure that the new structured control instructions meet the actual executable standards. Security adjustments include at least one of the following:
[0158] (1) Risk instruction shielding mechanism:
[0159] If the structured control instruction involves a dangerous operation (such as "open the trunk while driving"), the vehicle will automatically refuse to execute it and generate alternative feedback (such as "this operation is not available in the current state").
[0160] (2) Parameter fallback strategy:
[0161] If a parameter in a structured control command exceeds the permitted range, the vehicle will attempt to roll it back to the last known safe value or the protocol default value. For example, if the temperature is set to 50°C, it will automatically roll back to 22°C.
[0162] (3) Ambiguous instruction clarification mechanism (user interaction):
[0163] The vehicle can generate a counter-question, such as: "You are currently driving, are you sure you want to unlock the door?"
[0164] After completing the above adjustments, the system outputs new structured control instructions. The new structured control instructions comply with driving safety and parameter value compliance to ensure implementation.
[0165] For example, the vehicle collects the most recent conversation context between the user and the vehicle in real time. It recommends retaining semantically related question-and-answer pairs from the past minute or the last five rounds as input for the student model to interpret the current command intent. The structured control command X output by the student model undergoes a protocol validator validation, examining the protocol for structural integrity, value range compliance, context consistency, and operational state restrictions, resulting in result Y. If any non-compliance is detected, the system will provide safety feedback and adjustments based on the risk level and parameter conflict type. Through strategies such as command rejection, parameter rollback, or interactive clarification, the system ultimately outputs a protocol-compliant, safe, and controllable structured control command, which is then sent to the vehicle system for execution.
[0166] The following is an example of the model training method provided in the embodiments of the present application.
[0167] The model training method provided in the embodiment of the present application can be composed of five agents, forming a closed-loop process of training set generation - model training - feedback optimization:
[0168] 1. Semantic Parsing Agent (SPA) (also known as the semantic parsing module): parses the intent and parameters in the seed query;
[0169] 2. Fuzzy Injection Agent (FIA) (also known as Fuzzy Injection Module): constructs instructions with missing components and ambiguous parameters
[0170] 3. Multi-turn Evolution Agent (MTA) (also known as Multi-turn Evolution Module): Builds continuous context and forms multi-turn dialogue samples
[0171] 4. Adversarial Generation Agent (AGA) (also known as Adversarial Generation Module): Diversify and generalize control objects / parameters
[0172] 5. Student Model and Feedback Mechanism (Student + Protocol Validator): Use positive and negative chain training + feedback parameter adjustment to achieve protocol alignment and model optimization.
[0173] Among them, the semantic parsing agent:
[0174] Input: The user's natural language instructions, also known as basic instructions. Basic instructions are instructions that express clear and complete meanings.
[0175] Output: extracted intent and related parameters.
[0176] Function: Analyze the semantic structure of user instructions and identify the user's true intention and required parameter information.
[0177] Fuzzy injection agent:
[0178] Input: Basic explicit instructions.
[0179] Output: Constructed fuzzy or incomplete instructions (i.e. fuzzy instructions).
[0180] Function: By deleting key semantic components such as control objects and parameter slots, it simulates the user's fuzzy expressions in real scenarios and constructs robust training samples for the model.
[0181] Multi-round evolution agent:
[0182] Input: fuzzy instructions, dialogue state matrix (i.e. the dialogue state information mentioned above).
[0183] Output: A sequence of multi-round contextual interactive instructions.
[0184] Function: Dynamically generates multiple rounds of supplementary questions and answers (i.e., the first round of dialogue text) based on historical dialogue text and vehicle status information, which is used to train the student model to have the ability to inherit intent across rounds.
[0185] Adversarial Generation Agent:
[0186] Input: Basic instructions.
[0187] Output: Statements that control object and parameter generalization.
[0188] Function: By generalizing the control objectives in a diversified manner (for example, from "air conditioning" to "refrigeration system"), the model's ability to understand multiple expressions is improved.
[0189] Student Model:
[0190] Input: training set.
[0191] Output: Structured control instructions.
[0192] Function: Learn to recognize ambiguous instructions, perform execution intent analysis, and perform protocol mapping in supervised fine-tuning.
[0193] Feedback module (protocol compliance verification):
[0194] Input: Structured control instructions output by the student model.
[0195] Output: Protocol verification results, fed back to the fuzzy injection agent.
[0196] Function: Ensure that the generated structured control instructions comply with vehicle protocol specifications. If not, adjust the data generation strategy (for example, increase the training ratio of the control subject).
[0197] In the above, the "dialogue state matrix" refers to the structured representation used by the multi-round evolution agent to represent the current dialogue context state. It usually includes:
[0198] User Intent State for each round: Indicates whether the current user input is complete, which slots are missing, and whether the intent is clear;
[0199] System Feedback State: records the system's counter-questions and clarifications to the user in previous rounds, and whether it received a clear response;
[0200] Slot Filling Matrix: records in a structured manner whether the slot corresponding to each intent has been filled and its current value;
[0201] History Turn Embedding: For example, each round of dialogue (user input and system response) is encoded into a time series matrix in the form of a vector or code.
[0202] This matrix is similar to a continuously updated conversation "memory board". It is updated after each round of conversation, allowing the system to determine which information has been obtained and which is still unclear, supporting complex interactive behaviors such as dynamic questioning, turn control, and intention inheritance.
[0203] Figure 3 This is another flow chart of the model training method provided in the embodiment of the present application. Figure 3 As shown, the Semantic Parsing Agent (SPA) first extracts user intent from a seed query and generates basic instructions. Subsequently, the Fuzzy Injection Agent (FIA) constructs fuzzy instructions (as negative chain instructions) by removing key components of the basic instructions. The Adversarial Generation Agent (AGA) generalizes the control objects and control parameters of the basic instructions to generate diverse adversarial instructions (as positive chain instructions). Based on the fuzzy and adversarial instructions, a multi-turn dialogue with contextual dependencies is constructed, generating training samples covering a variety of fuzzy scenarios. These training samples are used to train the student model. During training, the structured control instructions output are further verified for compliance by a protocol verifier. Based on the verification results, errors in the student model's parsing process are fed back to the Fuzzy Injection Agent, which dynamically updates the slot sensitivity (i.e., the slot weights) and continuously optimizes the fuzzy strategy (i.e., the first fuzzy instruction generation strategy) within the Fuzzy Injection Agent. This ultimately forms a closed-loop distillation framework integrating data generation, model training, and protocol feedback.
[0204] The model training method provided in the embodiment of the present application has the following specific advantages:
[0205] 1. Multi-agent collaborative distillation to improve data quality and diversity:
[0206] The introduction of multiple agents (semantic parsing, fuzzy injection, multi-round evolution, and adversarial generation) collaboratively generates high-quality training data, enabling an end-to-end chain construction from intent recognition, fuzzy simulation, and multi-round reasoning. Compared with single-agent or data augmentation methods, this approach has stronger expression generalization capabilities and abnormal instruction coverage, significantly improving the model's robustness in complex scenarios.
[0207] 2. Dynamic feedback-driven distillation optimization mechanism:
[0208] By combining protocol constraints with feedback from student model misjudgments, the fuzzy injection strategy is dynamically adjusted to achieve precise sampling of positive and negative chain data. Through a closed-loop feedback mechanism, the training data distribution is automatically adjusted to address model weaknesses (such as ambiguity in control objects and slot omissions), forming a model-data co-evolution mechanism and addressing the bottleneck of the traditional distillation "static generation and single-round training" model.
[0209] 3. Strong structured protocol alignment capabilities, ensuring secure and controllable deployment:
[0210] All generated data is strictly aligned with the vehicle control protocol structure, supporting fine-grained control function mapping and parameter limit checking, avoiding the typical problem of "legitimate command generation but illegal execution." Compared to the open LLM generation strategy, this application ensures the implementation and security of generated commands in real vehicle control systems, meeting the deployment requirements of high-security application scenarios.
[0211] Figure 4 The schematic diagram of the structure of the model training device provided in the embodiment of the present application is shown. Figure 4 As shown, the model training device 400 includes:
[0212] A first acquisition module 401 is used to acquire basic instructions;
[0213] Identification module 402, configured to perform intent recognition on the basic instruction to obtain at least one intent and multiple slots corresponding to each intent;
[0214] A second acquisition module 403 is configured to process, for each of the intents, multiple slots corresponding to the intent according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intent;
[0215] A first generating module 404 is configured to generate a first multi-round dialogue text corresponding to each fuzzy instruction based on each fuzzy instruction and a dialogue state record, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes a plurality of question-answer pairs;
[0216] A construction module 405 is configured to construct a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of conversation text, wherein the training set includes a plurality of training samples, each of which includes one of the fuzzy instructions and the first multiple rounds of conversation text corresponding to the fuzzy instruction;
[0217] The training module 406 is used to train the student model using the training set to obtain a target model. The target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
[0218] In one embodiment of the present application, the first fuzzy instruction generation strategy includes weights of the plurality of slots corresponding to the intention;
[0219] The training module 406 includes:
[0220] A first acquisition submodule is configured to, after performing multiple rounds of training on the student model using the training set, use the student model obtained in the last round of training as the target model if the student model meets a preset iteration stopping condition;
[0221] a second acquisition submodule, configured to input verification samples in a verification set into the student model if the student model does not satisfy the iteration stopping condition, and obtain a second structured control instruction corresponding to each verification sample, wherein the verification sample includes a verification fuzzy instruction and multiple rounds of dialogue text corresponding to the verification fuzzy instruction;
[0222] a first adjustment submodule, configured to adjust the weights of the plurality of slots corresponding to the intent in the first fuzzy instruction generation strategy to obtain a second fuzzy instruction generation strategy if there is a second structured control instruction that does not meet a preset condition, wherein the preset condition is determined according to a preset compliance requirement;
[0223] The second adjustment submodule is configured to use the second fuzzy instruction generation strategy as the first fuzzy instruction generation strategy and trigger the second acquisition module to execute it.
[0224] In one embodiment of the present application, the first adjustment submodule includes:
[0225] a statistical unit, configured to count incorrectly identified slots and missing slots in the second structured control instruction in each round of training;
[0226] An adjustment unit is used to increase the weight of the missing slot and the weight of the slot whose error rate is higher than the average error rate, to obtain a second fuzzy instruction generation strategy, wherein the error rate refers to the probability of a slot being incorrectly identified in multiple rounds of training, and the average error rate refers to the average value of the probabilities of all slots being incorrectly identified in multiple rounds of training.
[0227] In one embodiment of the present application, the first generating module 404 includes:
[0228] The first generation submodule is configured to generate, for each of the fuzzy instructions, a first multi-round dialogue text corresponding to the fuzzy instruction using a pre-trained large language model (LLM) according to the dialogue state record and a constraint condition, wherein the constraint condition includes at least one of the following:
[0229] Each round of dialogue text includes a question-answer pair, which consists of the user's speech text and the system's feedback text;
[0230] The semantic similarity between the user speech text and the fuzzy instruction in each round of the question-answer pair is greater than a second threshold;
[0231] The control object included in each round of the question-answer pair belongs to the object supported by the vehicle, and the control parameter of the control object included in each round of the question-answer pair is within the control parameter value range set for the control object;
[0232] The number of question-answer pairs in the first round of dialogue texts is within a preset range;
[0233] The questions and answers in each round must comply with the preset vehicle function safety usage rules;
[0234] The text expression format of the question and answer pairs in each round is the same.
[0235] In one embodiment of the present application, the first generating module 404 includes:
[0236] The second generating submodule is configured to perform slot identification on each of the fuzzy instructions according to the dialogue state record to obtain a target slot, and obtain first multi-round dialogue texts corresponding to the fuzzy instructions according to at least one of the following:
[0237] generating a first question instruction for the target slot according to the user's historical preferences, receiving a first answer instruction, and constructing a question-answer pair according to the first question instruction and the first answer instruction;
[0238] generating a second question instruction according to the current state of the vehicle and the target slot, receiving a second answer instruction, and constructing a question-answer pair according to the second question instruction and the second answer instruction;
[0239] A third question instruction is generated only according to the target slot, and a third reply instruction is received, and a question-answer pair is constructed according to the third question instruction and the third reply instruction.
[0240] In one embodiment of the present application, the model training device 400 further includes:
[0241] a generalization module, configured to generalize the basic instruction to obtain a plurality of counter-instructions, wherein the plurality of counter-instructions include instructions having the same intention as the basic instruction and instructions having the opposite intention to the basic instruction;
[0242] A second generating module, for each of the confrontation instructions, generates a second multi-round dialogue text corresponding to the confrontation instruction;
[0243] An adding module is used to add each of the confrontation instructions and the corresponding second multiple rounds of dialogue text as training samples to the training set.
[0244] The model training device 400 provided in the embodiment of the present application can implement each process implemented in the aforementioned model training method embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.
[0245] In one embodiment of the present application, the first fuzzy instruction generation strategy is used to delete at least one slot among the multiple slots corresponding to the intention.
[0246] In one embodiment of the present application, the first fuzzy instruction generation strategy also includes a slot fuzzy strategy, and the slot fuzzy strategy is used to delete the slots in the intention whose weight is less than a first threshold.
[0247] Figure 5 FIG. 1 shows a schematic diagram of the structure of a vehicle control device provided in an embodiment of the present application. Figure 5 As shown, the vehicle control device 500 includes:
[0248] Receiving module 501, used for receiving target control instructions;
[0249] An acquisition module 502 is configured to input the target control instruction and historical multi-round dialogue text into a target model to obtain a structured control instruction, wherein the target model is obtained using the model training method in the above embodiment, and the historical multi-round dialogue text is dialogue text obtained through interaction with the user before the target control instruction is obtained;
[0250] The control module 503 is configured to control the vehicle according to the structured control instructions.
[0251] The vehicle control device 500 provided in the embodiment of the present application can implement the various processes implemented in the aforementioned vehicle control method embodiment and achieve the same technical effects. To avoid repetition, they will not be described here.
[0252] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.
[0253] The electronic device may include a processor 601 and a memory 602 storing computer program instructions.
[0254] Specifically, the processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0255] Memory 602 may include a large-capacity memory for data or instructions. By way of example and not limitation, memory 602 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, memory 602 is a non-volatile solid-state memory.
[0256] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect or the second aspect of the present disclosure.
[0257] The processor 601 implements any one of the information auditing methods in the above embodiments by reading and executing computer program instructions stored in the memory 602 .
[0258] In one example, the electronic device may further include a communication interface 603 and a bus 610. Figure 6 As shown, the processor 601, the memory 602, and the communication interface 603 are connected via a bus 610 and communicate with each other.
[0259] The communication interface 603 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0260] Bus 610 includes hardware, software, or both that couples components of the information audit method or verification device to each other. By way of example, and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industrial Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 610 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0261] In addition, in conjunction with the model training method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the model training methods or vehicle control methods in the above embodiments is implemented.
[0262] An embodiment of the present application may provide a vehicle, wherein when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes any one of the model training methods or vehicle control methods in the above embodiments.
[0263] An embodiment of the present application also provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes any one of the model training methods or vehicle control methods in the above embodiments.
[0264] It should be understood that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described as examples. However, the method process of the present application is not limited to the specific steps described. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0265] The functional blocks shown in the block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they may be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments may be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or communication link. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memory, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. Code segments may be downloaded via a computer network such as the Internet or an intranet.
[0266] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0267] Aspects of the present disclosure have been described above with reference to flowcharts and / or block diagrams of methods, devices (systems) according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that execution of these instructions by the processor of the computer or other programmable data processing device enables the implementation of the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0268] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Get basic instructions; Performing intent recognition on the basic instruction to obtain at least one intent and multiple slots corresponding to each intent; For each of the intentions, processing the multiple slots corresponding to the intention according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intention; Generating, based on each of the fuzzy instructions and the dialogue state record, a first multi-round dialogue text corresponding to each of the fuzzy instructions, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes a plurality of question-answer pairs; Constructing a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of conversation text, wherein the training set includes a plurality of training samples, and each of the training samples includes one of the fuzzy instructions and the first multiple rounds of conversation text corresponding to the fuzzy instruction; The training set is used to train the student model to obtain a target model, and the target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
2. The model training method according to claim 1, characterized in that The first fuzzy instruction generation strategy includes weights of the plurality of slots corresponding to the intention; The method of training the student model using the training set to obtain a target model includes: After multiple rounds of training of the student model using the training set, if the student model meets a preset iteration stop condition, the student model obtained from the last round of training is used as the target model; If the student model does not meet the iteration stopping condition, inputting the verification samples in the verification set into the student model to obtain a second structured control instruction corresponding to each verification sample, wherein the verification sample includes a verification fuzzy instruction and multiple rounds of dialogue text corresponding to the verification fuzzy instruction; If there is a second structured control instruction that does not meet the preset condition, the weights of the plurality of slots corresponding to the intention in the first fuzzy instruction generation strategy are adjusted to obtain a second fuzzy instruction generation strategy, wherein the preset condition is determined according to a preset compliance requirement; The second fuzzy instruction generation strategy is used as the first fuzzy instruction generation strategy, and the step of processing multiple slots corresponding to each intention according to the first fuzzy instruction generation strategy to obtain multiple fuzzy instructions corresponding to the intention is executed.
3. The model training method according to claim 2, characterized in that The step of adjusting the weights of the plurality of slots corresponding to the intention in the first fuzzy instruction generation strategy to obtain a second fuzzy instruction generation strategy includes: Counting the number of slots that are incorrectly recognized and the number of slots that are missing in the second structured control instruction in each round of training; Increase the weight of the missing slots, and increase the weight of the slots whose misjudgment rate is higher than the average misjudgment rate, to obtain a second fuzzy instruction generation strategy, where the misjudgment rate refers to the probability of a slot being incorrectly identified in multiple rounds of training, and the average misjudgment rate refers to the average value of the probabilities of all slots being incorrectly identified in multiple rounds of training.
4. The model training method according to claim 1, characterized in that Generating first multiple-round dialogue texts corresponding to each fuzzy instruction according to each fuzzy instruction and the dialogue state record includes: For each of the fuzzy instructions, a pre-trained large language model is used to generate first multiple-round dialogue texts corresponding to the fuzzy instruction based on the dialogue state record and constraints, where the constraints include at least one of the following: Each round of dialogue text includes a question-answer pair, which consists of the user's speech text and the system's feedback text; The semantic similarity between the user speech text and the fuzzy instruction in each round of the question-answer pair is greater than a second threshold; The control object included in each round of the question-answer pair belongs to the object supported by the vehicle, and the control parameter of the control object included in each round of the question-answer pair is within the control parameter value range set for the control object; The number of question-answer pairs in the first round of dialogue texts is within a preset range; The questions and answers in each round must comply with the preset vehicle function safety usage rules; The text expression format of the question and answer pairs in each round is the same.
5. The model training method according to claim 1, characterized in that Generating first multiple-round dialogue texts corresponding to each fuzzy instruction according to each fuzzy instruction and the dialogue state record includes: For each of the fuzzy instructions, slot identification is performed on the fuzzy instruction according to the dialogue state record to obtain a target slot, and first multi-round dialogue texts corresponding to the fuzzy instruction are obtained according to at least one of the following: generating a first question instruction for the target slot according to the user's historical preferences, receiving a first answer instruction, and constructing a question-answer pair according to the first question instruction and the first answer instruction; generating a second question instruction according to the current state of the vehicle and the target slot, receiving a second answer instruction, and constructing a question-answer pair according to the second question instruction and the second answer instruction; A third question instruction is generated only according to the target slot, and a third reply instruction is received, and a question-answer pair is constructed according to the third question instruction and the third reply instruction.
6. The model training method according to claim 1, characterized in that Before the student model is trained using the training set to obtain the target model, the method further includes: Generalizing the basic instruction to obtain a plurality of counter-instructions, wherein the plurality of counter-instructions include instructions having the same intention as the basic instruction and instructions having the opposite intention to the basic instruction; For each of the confrontation instructions, generating a second multi-round dialogue text corresponding to the confrontation instruction; Each of the adversarial instructions and the corresponding second-round dialogue texts are added to the training set as training samples.
7. The model training method according to claim 1, characterized in that The first fuzzy instruction generation strategy is used to delete at least one slot among the multiple slots corresponding to the intention.
8. The model training method according to claim 2, characterized in that The first fuzzy instruction generation strategy also includes a slot fuzzy strategy, which is used to delete slots in the intention whose weight is less than a first threshold.
9. A vehicle control method, characterized in that: The method comprises: receiving target control instructions; Inputting the target control instruction and historical multi-round dialogue text into a target model to obtain a structured control instruction, wherein the target model is obtained by the model training method according to any one of claims 1 to 8, and the historical multi-round dialogue text is dialogue text obtained by interacting with the user before the target control instruction; The vehicle is controlled according to the structured control instructions.
10. A model training device, characterized in that: The device comprises: A first acquisition module is used to acquire basic instructions; an identification module, configured to perform intent recognition on the basic instruction to obtain at least one intent and a plurality of slots corresponding to each intent; A second acquisition module is configured to process, for each of the intentions, a plurality of slots corresponding to the intention according to the first fuzzy instruction generation strategy to obtain a plurality of fuzzy instructions corresponding to the intention; a first generating module configured to generate, based on each of the fuzzy instructions and a dialogue state record, a first multi-round dialogue text corresponding to each of the fuzzy instructions, wherein the dialogue state record is used to record the dialogue context state of the basic instruction, and the first multi-round dialogue text includes a plurality of question-answer pairs; A construction module is configured to construct a training set based on each of the fuzzy instructions and the corresponding first multiple rounds of dialogue text, wherein the training set includes a plurality of training samples, and each of the training samples includes one of the fuzzy instructions and the first multiple rounds of dialogue text corresponding to the fuzzy instruction; The training module is used to train the student model using the training set to obtain a target model, and the target model has the ability to understand the intention of the input instructions and obtain structured control instructions.
11. A vehicle control device, characterized in that: The device comprises: A receiving module, used for receiving target control instructions; an acquisition module, configured to input the target control instruction and historical multi-round dialogue text into a target model to obtain a structured control instruction, wherein the target model is obtained according to the model training method according to any one of claims 1 to 8, and the historical multi-round dialogue text is dialogue text obtained through interaction with the user before the target control instruction; A control module is used to control the vehicle according to the structured control instructions.
12. An electronic device, characterized in that: include: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the model training method as described in any one of claims 1 to 8, or when the processor executes the computer program instructions, it implements the vehicle control method as described in claim 9.
13. A vehicle, characterized in that: The vehicle includes the electronic device according to claim 12.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the model training method as described in any one of claims 1 to 8, or, when executed by a processor, implement the vehicle control method as described in claim 9.
15. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the model training method as described in any one of claims 1 to 8, or the electronic device executes the vehicle control method as described in claim 9.
Citation Information
Patent Citations
Generation method and device of intention slot position recognition model, and electronic equipment
CN117520793A
Generative intention recognition method based on prior information and related equipment
CN120046620A