Large language model training method, voice interaction method and server
By using the training method of a large language model in the vehicle assistant, using vehicle knowledge information to generate training data, and performing incremental pre-training of the base model, the existing vehicle assistant's training cost and poor scalability are solved, and an efficient voice interaction experience is achieved.
Patent Information
- Application Number
- CN202510220173.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing vehicle assistants process user voice requests through multiple serial models, requiring a large amount of labeled data for training, which is costly and poorly scalable.
The training method of a large language model is adopted, and the training data is generated based on the preset mask language model, and the base model is incrementally pre-trained to understand user instructions without labeling data.
It reduces the cost of model training, improves the understanding ability of large language models, realizes intelligent and convenient voice interaction experience, and enhances user experience.
Smart Images

Figure CN120071899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction, and particularly to a training method for a large language model for vehicle voice interaction, a voice interaction method, a server, and a computer-readable storage medium. Background Art
[0002] In related technologies, in-vehicle assistants use multiple serial models to separately perform slot extraction processing, intent recognition processing, and slot parameter filling processing on user voice requests to understand and meet user needs. However, in-vehicle assistants including multiple serial models require a large amount of labeled data for training, resulting in high training costs. Summary of the Invention
[0003] The present application provides a training method for a large language model for vehicle voice interaction, a voice interaction method, a server, and a computer-readable storage medium.
[0004] An embodiment of the present application provides a training method for a large language model for vehicle voice interaction, the method including:
[0005] Obtaining vehicle knowledge information, where the vehicle knowledge information includes function point information of vehicle function points and association information between the vehicle function points;
[0006] Based on a preset masked language model, generating first training data according to the vehicle knowledge information;
[0007] Training a base model according to the first training data to obtain the large language model.
[0008] In this way, the server obtains vehicle knowledge information, where the vehicle knowledge information includes function point information of vehicle function points and association information between the vehicle function points. Then, based on the preset masked language model, the server generates first training data according to the vehicle knowledge information. Finally, the server trains the base model according to the first training data to obtain the large language model. In this way, incremental pre-training is performed on the base model according to the generated first training data, improving the understanding ability of the large language model, enabling the model to accurately understand user instructions and provide users with an intelligent and convenient interaction experience without labeled data, thereby reducing the cost of model training.
[0009] In some embodiments, the generating the first training data based on the preset masked language model according to the vehicle knowledge information includes:
[0010] Performing random masking processing on the vehicle knowledge information based on the preset masked language model to generate the first training data, where the first training data includes the vehicle knowledge information and tokenized knowledge information, and the tokenized knowledge information is vehicle knowledge information with some words masked.
[0011] In this way, based on a pre-set masked language model, the server performs random masking on vehicle knowledge information to generate first training data. The first training data includes vehicle knowledge information and tokenized knowledge information, where the tokenized knowledge information is vehicle knowledge information with some words masked. In this way, by training the base model with the first training data generated by randomly masking vehicle knowledge information, the base model can learn rich vehicle knowledge and accurately understand the instruction intent of the user. Through the random masking process, it can help the base model learn the relationships between words and context information, thereby improving the language understanding ability of the base model.
[0012] In some embodiments, the method further includes:
[0013] Configuring a first prompt;
[0014] Performing supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data.
[0015] In this way, the server configures a first prompt. Then, the server performs supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data. In this way, through the second training data for supervised fine-tuning training, it can strengthen the large language model's understanding ability of cross-domain context and improve the large language model's cross-vertical domain multi-round understanding ability.
[0016] In some embodiments, the second training data includes single-round conversation data and multi-round conversation data. The performing supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data includes:
[0017] Performing intent understanding training on the large language model based on the first prompt according to the single-round conversation data;
[0018] Performing semantic understanding training on the large language model based on the first prompt according to the multi-round conversation data.
[0019] Thus, the second training data includes single-round conversation data and multi-round conversation data. Based on the first prompt, the large language model is trained for intent understanding according to the single-round conversation data in the second training data. Then, based on the first prompt, the server trains the large language model for semantic understanding according to the multi-round conversation data in the second training data. In this way, through the training with single-round conversation data, the large language model can learn the correspondence between user intents and application interface calls, thereby improving the accuracy of intent understanding. Through the training with multi-round conversation data, the large language model can learn the influence of context information on user intents, thereby enhancing the semantic understanding ability and accurately understanding the true intents of users.
[0020] An embodiment of the present application provides a voice interaction method. The voice interaction method is based on the large language model trained by the above training method. The method includes:
[0021] Obtain the current voice request;
[0022] Based on the large language model, determine a vehicle control instruction corresponding to the current voice request according to the current voice request;
[0023] Send the vehicle control instruction to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction.
[0024] Thus, the current voice request is obtained. Then, based on the large language model, the server determines a vehicle control instruction corresponding to the current voice request according to the current voice request. Finally, the vehicle control instruction is sent to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction. In this way, the large language model can understand the multi-round conversation content of the user and achieve intent understanding and semantic understanding of the user's voice request according to the context information, thereby accurately understanding the user's intent and converting the user's intent into a vehicle control instruction, thus realizing the function of voice control of the vehicle and natural and smooth multi-round interaction, and enhancing the user experience.
[0025] In some embodiments, the large language model is configured with a second prompt. The step of, based on the large language model, determining a vehicle control instruction corresponding to the current voice request according to the current voice request includes:
[0026] Based on the second prompt, guide the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request.
[0027] In this way, based on the second prompt, the server guides the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request. In this way, the second prompt can guide the large language model to reason according to the pre-designed thinking process, so as to accurately understand the user's intention and output the corresponding vehicle control instruction. Moreover, the second prompt can help the large language model accurately understand the context information, thereby improving the semantic understanding ability and realizing natural and smooth multi-round interaction.
[0028] In some embodiments, the guiding the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request based on the second prompt includes:
[0029] Based on the second prompt, guiding the large language model to determine a target voice request according to the current voice request and the historical voice request;
[0030] Determining a target intention according to the target voice request;
[0031] Determining a natural language processing result according to the target intention;
[0032] Determining the vehicle control instruction according to the natural language processing result.
[0033] In this way, based on the second prompt, the server guides the large language model to determine a target voice request according to the current voice request and the historical voice request. Then, the server determines the target intention according to the target voice request. Next, the server determines the natural language processing result according to the target intention. Finally, the server determines the vehicle control instruction according to the natural language processing result. In this way, the second prompt can guide the large language model to reason according to the pre-designed thinking process, so as to accurately understand the user's intention and output the corresponding vehicle control instruction. Moreover, the second prompt can help the large language model accurately understand the context information, thereby improving the semantic understanding ability and realizing natural and smooth multi-round interaction.
[0034] In some embodiments, the guiding the large language model to determine a target voice request according to the current voice request and the historical voice request based on the second prompt includes:
[0035] When the current voice request is associated with the historical voice request, performing a complement processing on the current voice request according to the historical voice request to determine the target voice request;
[0036] When the current voice request is not associated with the historical voice request, determining the current voice request as the target voice request.
[0037] Thus, when the current voice request is associated with the historical voice request, the server completes the current voice request according to the historical voice request to determine the target voice request. Or when the current voice request is not associated with the historical voice request, the server determines the current voice request as the target voice request. In this way, by completing the current voice request, the user's intention can be accurately understood, avoiding ambiguity and misunderstanding. Moreover, by associating with the historical voice request, multi-round interaction can be achieved, making the conversation smooth and natural, thereby enhancing the user experience.
[0038] In some embodiments, determining the natural language processing result according to the target intention includes:
[0039] Performing slot recognition on the target voice request to determine a slot recognition result;
[0040] Performing application programming interface prediction according to the slot recognition result and the target intention to determine a predicted application interface;
[0041] According to the slot recognition result and the predicted application interface, selecting the predicted application interface to perform application programming interface parameter filling to determine the natural language processing result.
[0042] Thus, the server performs slot recognition on the target voice request to determine a slot recognition result. Then, the server performs application programming interface prediction according to the slot recognition result and the target intention to determine a predicted application interface. Finally, the server selects the predicted application interface according to the slot recognition result and the predicted application interface to perform application programming interface parameter filling to determine the natural language processing result. In this way, according to the result of slot recognition, selecting the predicted application programming interface to perform application programming interface parameter filling, the server can understand and respond to the target voice request, directly output the execution result to the vehicle to complete the voice interaction, thereby reducing the latency of the in-vehicle system, improving the response speed to user commands, and providing a convenient and natural interaction experience.
[0043] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.
[0044] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0045] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. Description of the Drawings
[0046] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0047] Figure 1 is one of the flow diagrams of the training method according to some embodiments of the present application;
[0048] Figure 2 is another flow diagram of the training method according to some embodiments of the present application;
[0049] Figure 3 is yet another flow diagram of the training method according to some embodiments of the present application;
[0050] Figure 4 is still another flow diagram of the training method according to some embodiments of the present application;
[0051] Figure 5 is one of the flow diagrams of the voice interaction method according to some embodiments of the present application;
[0052] Figure 6 is another flow diagram of the voice interaction method according to some embodiments of the present application;
[0053] Figure 7 is yet another flow diagram of the voice interaction method according to some embodiments of the present application;
[0054] Figure 8 is the flow diagram of the large language model processing the current voice request according to some embodiments of the present application;
[0055] Figure 9 is still another flow diagram of the voice interaction method according to some embodiments of the present application;
[0056] Figure 10 is yet another flow diagram of the voice interaction method according to some embodiments of the present application. Detailed Embodiments
[0057] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the embodiments of the present application, and should not be construed as a limitation to the embodiments of the present application.
[0058] Currently, users interact through in-vehicle assistants, achieving convenient vehicle control. In related technologies, in-vehicle assistants usually use multiple serial models to process users' voice requests to complete the understanding and execution of users' voice requests, including slot extraction processing, intent recognition processing, and slot parameter filling processing. Slot extraction processing refers to using a named entity recognition model to identify key information in a voice request, such as the operation object, time, location, etc., and marking it as a slot. Intent recognition processing refers to using an intent classification model to identify the user's intent based on slot information, such as opening the window, playing music, setting navigation, etc. Slot parameter filling processing refers to using rules or templates to generate executable APIs and parameters based on the intent and slot information, such as "open the left rear window", etc.
[0059] However, each step of the above in-vehicle assistant for processing voice requests requires a separate training of a model, which requires building and annotating a large amount of data, consuming a large amount of manpower and resources. Moreover, when new functions need to be added or existing functions need to be modified, the relevant models need to be retrained, with a long iteration cycle and poor scalability. In addition, this solution mainly targets the vehicle control vertical domain, with insufficient support for vertical domains such as navigation and multimedia search, limited understanding ability, and poor user experience.
[0060] Based on the above problems, please refer to Figure 1 , the embodiments of the present application provide a training method for a large language model for vehicle voice interaction, and the method includes:
[0061] 011: Obtain vehicle knowledge information;
[0062] 012: Based on a preset masked language model, generate first training data according to the vehicle knowledge information;
[0063] 013: Train a base model according to the first training data to obtain a large language model.
[0064] The embodiments of the present application also provide a server, including a memory and a processor. The training method for the large language model for vehicle voice interaction in the embodiments of the present application can be implemented by the server in the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain vehicle knowledge information. And based on a preset masked language model, generate first training data according to the vehicle knowledge information. And train a base model according to the first training data to obtain a large language model.
[0065] The embodiments of the present application also provide a model training device. The training method of the large language model for vehicle voice interaction in the embodiments of the present application can be implemented by the model training device in the embodiments of the present application. Specifically, the model training device includes an acquisition module, a generation module, and a training module. The acquisition module is used to acquire vehicle knowledge information. The generation module is used to generate first training data based on a preset masked language model and the vehicle knowledge information. The training module is used to train a base model according to the first training data to obtain a large language model.
[0066] Specifically, the vehicle knowledge information includes the vehicle function points of the vehicle and the association information between the vehicle function points. The association information includes the hierarchical relationship between the function points and the applicable scenarios of the function points, etc. The vehicle function points of the vehicle include opening / closing the window, turning on / off the air conditioner, navigation, music playing, etc. In some embodiments, the acquisition sources of the vehicle knowledge information include user manuals, vehicle configuration information, in-vehicle system data, etc.
[0067] The hierarchical relationship between the vehicle function points refers to the inclusion and being included relationship between the vehicle function points in the intelligent vehicle and their positions in the functional hierarchical structure. This hierarchical relationship can help the large language model clearly understand the functional scope and mutual relationship of each vehicle function point. For example, the hierarchical relationship between the function points "vehicle control > window control > window switch". Vehicle control is the top-level vehicle function point module, including sub-function modules such as window control and seat control. Window control includes more specific operation function points such as window switch and window lifting. Through this hierarchical relationship, the large language model can easily understand that the window switch function point belongs to vehicle control and can further understand what operations it can perform, such as opening the window and closing the window.
[0068] The applicable scenario of the function point refers to the situation in which a certain vehicle function point will be used. Understanding the applicable scenario of the vehicle function point can help the large language model accurately perform speech recognition and understanding of the user's speech request.
[0069] The preset masked language model (Masked Language Model, MLM) refers to a deep learning technology for pre-training language models, aiming to let the model learn the semantic representation of language. The core idea of MLM is to randomly mask a part of the words in the input text and train the large language model to predict the masked words, so that the large language model can learn the grammar and semantic information of the language.
[0070] The first training data refers to the data generated by a preset masked language model through masking vehicle knowledge information, that is, the pre-training task data of the MLM (Masked Language Model) constructed using vehicle knowledge data. Specifically, some words in the vehicle knowledge data are randomly masked (represented by the [mask] symbol) as the content for the large language model to learn. The training objective of the MLM is to let the model predict what the masked word is. For example, when masking the word "window" in "Open the window", the model needs to predict what "window" is. In this way, the large language model can better understand vehicle-related language, such as "Open the window", "Navigate to XX location", etc.
[0071] The base model refers to a model with broad language understanding ability learned during the pre-training stage, usually trained using a large amount of text data, such as books, news, web pages, etc. The base model is the foundation for subsequent specific task models, such as text classification, machine translation, question answering systems, etc. Through incremental pre-training on the base model, the large language model can better understand user instructions. For example, when the user says "Open the window", the model can understand that "window" refers to a functional point on the vehicle and can perform corresponding operations. Compared with training a large language model from scratch, incremental pre-training can greatly reduce the need for training data and improve training efficiency.
[0072] Extract information about vehicle functional points from user manuals or other data sources, including the names of functional points, descriptions of functional points, and hierarchical relationships between functional points, etc.
[0073] Next, based on the preset masked language model, convert the vehicle knowledge information into training data. For example, randomly replace the names of functional points with masks (such as the [mask] symbol) to let the large language model learn and understand these functions.
[0074] Finally, use the generated first training data to train the base model so that the model can learn and understand vehicle knowledge. After training is completed, the base model will possess certain vehicle knowledge and be able to recognize and understand vehicle-related vocabulary and concepts.
[0075] In summary, the present application provides a training method and a server for a large language model for vehicle voice interaction. The server obtains vehicle knowledge information, which includes function point information of vehicle function points and association information between vehicle function points. Then, based on a preset masked language model, the server generates first training data according to the vehicle knowledge information. Finally, the server trains a base model according to the first training data to obtain a large language model. In this way, incremental pre-training of the base model according to the generated first training data can improve the model's understanding ability, so that the model can accurately understand user instructions without labeled data and provide users with an intelligent and convenient interaction experience.
[0076] Please refer to Figure 2 , in some embodiments, step 012 (generating first training data based on a preset masked language model according to vehicle knowledge information) includes:
[0077] 0121: Based on a preset masked language model, perform random masking processing on the vehicle knowledge information to generate first training data.
[0078] In some embodiments, the generating module is further configured to perform random masking processing on the vehicle knowledge information based on a preset masked language model to generate first training data.
[0079] In some embodiments, the processor is further configured to perform random masking processing on the vehicle knowledge information based on a preset masked language model to generate first training data.
[0080] Specifically, during the random masking processing for constructing the first training data, some words at randomly selected positions in the text are blocked so that they cannot be directly seen by the model. The purpose of this is to train the large language model's ability to predict the blocked content, thereby learning the context information in the language.
[0081] The tokenized knowledge information refers to the vehicle knowledge information after random masking processing during the construction of the first training data, where the blocked part is replaced by special tokens such as [mask]. The tokenized knowledge information includes vehicle knowledge information with some words blocked, and these blocked words need to be predicted by the model according to the context information. The tokenized knowledge information is a part of the first training data, which randomly blocks the original vehicle knowledge information to form a new data form, including the blocked content and the unblocked content. The large language model needs to predict the blocked content according to the unblocked content and the context information, thereby learning the knowledge in the input text and improving the model's semantic understanding ability.
[0082] In some embodiments, it is also necessary to structurally process vehicle knowledge information. Structural processing refers to the process of converting unstructured or semi-structured data into structured data. Structured data usually has a clear format and fields, such as data in a table or a database. Structural processing can improve the storability, processability, and analyzability of data, and help improve the accuracy and understanding ability of the model. For example, converting "Open the window" into the format of "{Operation: Open, Control: Window}". Subsequently, mask the "Control" part in "{Operation: Open, Control: Window}" to become "{Operation: Open, Control: [MASK]}".
[0083] Use a preset masked language model to randomly mask the vehicle knowledge information collected from user manuals or other relevant materials. The masked language model will randomly select some words for masking and replace them with the [mask] symbol. Through the masking process, the large language model is forced to learn the content of the masked part, thereby accurately understanding the meaning of the entire text and enhancing the understanding of vehicle knowledge.
[0084] Then, combine the original vehicle knowledge information and the masked and tokenized knowledge information to generate the first training data for model training.
[0085] Finally, use the first training data to train the large language model, enabling the large language model to learn various types of vehicle knowledge and understand the content of the masked part.
[0086] For example, the descriptive text of the vehicle function "The vehicle has a window control function and can open, close the window, and adjust the opening degree of the window." After random masking, it may become: "The vehicle has a window control function and can open, close [mask], and adjust the [mask] opening degree." Then, based on the training data obtained above, the knowledge that the large language model can learn includes: First, "window" and "opening degree" are the masked parts. Second, "window" and "opening degree" are associated with operations such as "open", "close", and "adjust". Third, the large language model can understand user instructions, such as "Open the window" or "Close the window halfway".
[0087] In this way, based on a pre-set masked language model, the server randomly masks the vehicle knowledge information to generate first training data, which includes the vehicle knowledge information and tokenized knowledge information. The tokenized knowledge information is the vehicle knowledge information with some words masked. In this way, by using the first training data generated by randomly masking the vehicle knowledge information to train the base model, the base model can learn rich vehicle knowledge and accurately understand the user's instruction intention. Through the random masking process, it can help the base model learn the relationships between words and context information, thereby improving the language understanding ability of the base model.
[0088] Please refer to Figure 3 , in some embodiments, the method further includes:
[0089] 014: Configure a first prompt;
[0090] 015: Perform supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data.
[0091] In some embodiments, the model training device further includes a configuration module, and the configuration module is further used to configure the first prompt. The training module is further used to perform supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data.
[0092] In some embodiments, the processor is further used to configure the first prompt. And perform supervised fine-tuning training on the large language model according to the first prompt and pre-configured second training data.
[0093] Specifically, the first prompt refers to a prompt that guides the large language model to perform intention understanding and semantic understanding, including task description, user-side input, and output content format. Among them, the task description is used to describe the task that the large language model needs to complete. The user-side input refers to the data information input into the large language model. The first prompt may be "Suppose you are an in-vehicle voice assistant (Assistant), responsible for listening to all conversations between users in the vehicle. When you determine that the conversation content is an instruction sent to you, combine the conversation context, infer the current user's intention, and give the corresponding executable API and ARGUMENTS".
[0094] The second training data refers to the data mined from online data that can be used for supervised fine-tuning training of the large language model, including single-round conversation data and multi-round conversation data. The second training data can be used to learn the user's intention understanding and the corresponding relationship between the user's intention and API+ARGUMENTS, as well as the understanding of multi-round semantics.
[0095] Supervised fine-tuning training refers to using a pre-trained large language model, combined with a first prompt and second training data, to further train the model so that it can better understand and execute specific tasks, that is, to improve the large language model's ability to understand user instructions, enabling it to accurately identify the user's intent and provide corresponding APIs and ARGUMENTS.
[0096] Configure the first prompt and the second training data. Then, input the first prompt and the second training data into the large language model for supervised fine-tuning training. The large language model will learn how to understand the user's intent based on the user input and context information and output corresponding APIs and parameters. In this way, leveraging the powerful semantic understanding ability of the large language model, end-to-end instruction understanding is achieved without distinguishing vertical domains, improving the scalability and usability of the model. Also, through training with multi-round dialogue data, the model can understand context information, improving the accuracy and fluency of multi-round interactions. Additionally, without distinguishing slots, intents, etc., directly outputting APIs and ARGUMENTS simplifies the model structure and improves efficiency.
[0097] For example, assume the user wants to open the window and adjust the temperature, and the following conversation occurs: "User: Open the window; Model: Okay, the window is open. User: Close it a little more. ; Model: Okay, the window is closed halfway. User: Adjust the temperature; Model: Do you want to increase or decrease the temperature?" Here, the large language model first understands the user's intent based on the user's instruction "Open the window" and executes the corresponding API call to open the window. Then, based on the user's subsequent instruction "Close it a little more", the model understands that the user's intent is to continue closing the window and executes the corresponding API call. Finally, based on the user's instruction "Adjust the temperature", the model asks the user whether they want to increase or decrease the temperature to further execute the corresponding operation.
[0098] In this way, the server configures the first prompt. Then, the server performs supervised fine-tuning training on the large language model according to the first prompt and the pre-configured second training data. Thus, through the second training data for supervised fine-tuning training, the ability of the large language model to understand cross-domain context can be strengthened, and the cross-vertical domain multi-round understanding ability of the large language model can be improved.
[0099] Please refer to Figure 4 , in some embodiments, the second training data includes single-round dialogue data and multi-round dialogue data, and step 015 (performing supervised fine-tuning training on the large language model according to the first prompt and the pre-configured second training data) includes:
[0100] 0151: Based on the first prompt, perform intent understanding training on the large language model according to the single-round dialogue data;
[0101] 0152: Based on the first prompt, semantic understanding training is performed on the large language model according to multi-turn dialogue data.
[0102] In some embodiments, the training module is further configured to perform intent understanding training on the large language model based on the first prompt according to single-turn dialogue data, and perform semantic understanding training on the large language model based on the first prompt according to multi-turn dialogue data.
[0103] In some embodiments, the processor is further configured to perform intent understanding training on the large language model based on the first prompt according to single-turn dialogue data, and perform semantic understanding training on the large language model based on the first prompt according to multi-turn dialogue data.
[0104] Specifically, the second training data includes two types of data, namely single-turn dialogue data and multi-turn dialogue data.
[0105] Among them, the single-turn dialogue data is used to learn user intent understanding. The single-turn dialogue data includes a single instruction proposed by the user and the corresponding execution result. For example, the user says "Open the window", and the system replies "The window has been opened for you". Through this single-turn dialogue data, the large language model can learn to recognize the user's intent and associate it with the corresponding API and parameters.
[0106] The multi-turn dialogue data is used to learn multi-turn semantic understanding. The multi-turn dialogue data includes multi-turn conversations between the user and the system. For example, the user first says "Open the window", and then says "Close it a little more". Through this data, the large language model can learn to understand the semantics of multi-turn conversations and infer the user's true intent based on context information.
[0107] First, the first prompt and the single-turn dialogue data are input into the large language model, and the training model understands the user's intent according to the user query and outputs the corresponding API and parameters.
[0108] Next, the first prompt and the multi-turn dialogue data are input into the large language model. The training model not only understands the user's intent according to the current round of user query, but also needs to consider context information, such as the content of the previous round of conversation, in order to more accurately understand the user's intent.
[0109] Thus, the second training data includes single-round conversation data and multi-round conversation data. Based on the first prompt, the large language model is trained for intent understanding according to the single-round conversation data in the second training data. Then, based on the first prompt, the server trains the large language model for semantic understanding according to the multi-round conversation data in the second training data. In this way, through the training with single-round conversation data, the large language model can learn the correspondence between user intents and application interface calls, thereby improving the accuracy of intent understanding. Through the training with multi-round conversation data, the large language model can learn the influence of context information on user intents, thereby enhancing the semantic understanding ability and accurately understanding the true intents of users.
[0110] Please refer to Figure 5 , an embodiment of the present application provides a voice interaction method. The voice interaction method is based on the large language model trained by the above training method. The method includes:
[0111] 021: Obtain the current voice request;
[0112] 022: Based on the large language model, determine a vehicle control instruction corresponding to the current voice request according to the current voice request;
[0113] 023: Send the vehicle control instruction to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction.
[0114] An embodiment of the present application also provides a server, including a memory and a processor. The server can be deployed in a vehicle or in the cloud. The voice interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain the current voice request. And based on the large language model, determine a vehicle control instruction corresponding to the current voice request according to the current voice request to complete the voice interaction. And send the vehicle control instruction to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction.
[0115] An embodiment of the present application also provides a voice interaction device. The voice interaction method of the embodiment of the present application can be implemented by the voice interaction device of the embodiment of the present application. Specifically, the voice interaction device includes an acquisition module, a determination module, and a sending module. The acquisition module is used to obtain the current voice request. The determination module is used to determine a vehicle control instruction corresponding to the current voice request based on the large language model according to the current voice request. The sending module is used to send the vehicle control instruction to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction.
[0116] Specifically, the current voice request refers to the voice command issued by the user in the current conversation turn. Suppose the user previously said "Open the window" and currently says "Close it a little more", then "Close it a little more" is the current voice request. It should be noted that the current voice request, together with the content of the previous conversation turns, is used as the input to the large language model.
[0117] The vehicle control instruction refers to the instruction obtained by the large language model based on the current voice request and used to control the vehicle to perform specific operations. For example, when the user says "Open the window", the system needs to understand the user's intention to open the window and perform corresponding operations according to the current state of the vehicle (such as whether the window is already open).
[0118] The system first obtains the current voice request issued by the user.
[0119] Next, using the previously trained large language model, based on the current voice request, analyze the user's intention and determine the corresponding vehicle control instruction (the corresponding API and parameters). That is, combining the user's current voice request and context information (historical voice requests), such as the previous conversation content, the current state of the vehicle, etc., to determine the corresponding vehicle control instruction.
[0120] Finally, send the determined vehicle control instruction to the vehicle, and the vehicle performs corresponding operations according to the instruction, such as opening the window, adjusting the temperature, etc.
[0121] In this way, obtain the current voice request. Next, based on the large language model, the server determines the vehicle control instruction corresponding to the current voice request according to the current voice request. Finally, send the vehicle control instruction to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction. In this way, the large language model can understand the multi-turn conversation content of the user and, based on the context information, achieve the intention understanding and semantic understanding of the user's voice request, thereby accurately understanding the user's intention and converting the user's intention into a vehicle control instruction, thus realizing the function of voice control of the vehicle and natural and smooth multi-turn interaction, enhancing the user experience.
[0122] Please refer to Figure 6 , in some embodiments, the large language model is configured with a second prompt word, and step 022 (based on the large language model, according to the current voice request, determine the vehicle control instruction corresponding to the current voice request) includes:
[0123] 0221: Based on the second prompt word, guide the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request.
[0124] In some embodiments, the determination module is further configured to guide the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request based on the second prompt word.
[0125] In some embodiments, the processor is further configured to guide the large language model to determine a vehicle control instruction based on the current voice request and the historical voice request according to the second prompt.
[0126] Specifically, the second prompt is a prompt for guiding the large language model to perform reasoning, including a task objective, an input format, and an output format. It guides the model to complete the context information through Chain-of-Thought (COT) reasoning, helping the large language model to better understand the user's intention and generate a vehicle control instruction based on the current voice request and the historical voice request. The second prompt may be as follows:
[0127] "Please first combine the content of the historical conversation round and the current user request to determine whether the user's intention is related to the content of the historical conversation round or depends on the content of the historical conversation round to complete the information for the current request, and output the complete user request after completion. Then, perform an atomic operation splitting of the user intention based on the complete user request. The final result is output in JSON format."
[0128] The second prompt is configured in the large language model. The obtained current voice request and historical voice request are used as the input of the large language model. The large language model will perform reasoning based on the knowledge learned during the training process and under the guidance of the pre-configured second prompt, and output the corresponding APIs and parameters that the vehicle can execute, that is, the vehicle control instruction.
[0129] In this way, based on the second prompt, the server guides the large language model to determine a vehicle control instruction according to the current voice request and the historical voice request. In this way, the second prompt can guide the large language model to perform reasoning according to the pre-designed thinking process, so as to accurately understand the user's intention and output the corresponding vehicle control instruction. Moreover, the second prompt can help the large language model accurately understand the context information, thereby improving the semantic understanding ability and realizing natural and smooth multi-round interaction.
[0130] Please refer to Figure 7 , in some embodiments, step 0221 (guiding the large language model to determine a vehicle control instruction based on the second prompt according to the current voice request and the historical voice request) includes:
[0131] 02211: Based on the second prompt, guide the large language model to determine a target voice request according to the current voice request and the historical voice request;
[0132] 02212: Determine a target intention according to the target voice request;
[0133] 02213: Determine a natural language processing result according to the target intention;
[0134] 02214: Determine a vehicle control instruction based on the natural language processing result.
[0135] In some embodiments, the determination module is further configured to, based on a second prompt word, guide a large language model to determine a target voice request according to the current voice request and the historical voice request. And determine a target intent according to the target voice request. The determination module is further configured to determine a natural language processing result according to the target intent. And determine a vehicle control instruction according to the natural language processing result.
[0136] In some embodiments, the processor is further configured to, based on a second prompt word, guide a large language model to determine a target voice request according to the current voice request and the historical voice request. And determine a target intent according to the target voice request. The processor is further configured to determine a natural language processing result according to the target intent. And determine a vehicle control instruction according to the natural language processing result.
[0137] Specifically, the target voice request refers to the content that the model actually wants to express after analyzing and understanding the user's intention based on the user's current voice request and historical voice request. It is the basis for understanding the user's intention. It may be a supplement or correction to the current voice request, or it may be the current voice request itself.
[0138] The target intent refers to the function or purpose that the large language model identifies that the user wants to achieve according to the target voice request, such as "open the window" and "adjust the temperature".
[0139] The natural language processing result refers to the result of the model converting the target intent into a natural language expression. The form of the natural language processing result can be text, instruction, code, etc., for subsequent vehicle control or information query operations.
[0140] Please refer to Figure 8 , Figure 8 for the schematic diagram of the processing flow of the large language model for the current voice request. Take the current voice request and the historical voice request obtained from the historical conversation database as the input of the large language model, and based on the trained model and the second prompt word, according to the knowledge learned during the training process, combined with the guidance of the second prompt word, reason about the current voice request and the historical voice request, and output the content that the user actually wants to express, that is, the target voice request. For example, if the user previously said "open the window" and currently says "close it a little more", then the target voice request is "close the window a little".
[0141] Next, the large language model analyzes the target voice request, identifies the operation that the user wants to perform, and determines the user's target intent. And determine a natural language processing result according to the target intent, such as API and parameters. Finally, determine a vehicle control instruction according to the natural language processing result.
[0142] For example, the user says "Turn on the air conditioner", and then says "Make it a little higher". The large language model analyzes the current voice request and the historical voice requests, and determines that the target voice request is "Turn up the air conditioner a little more". Then, the large language model identifies the target intention as "Increase the temperature of the air conditioner". Subsequently, based on the user's intention of "Increase the temperature of the air conditioner", the large language model determines the natural language processing result: "API: AcOpen; ARGUMENTS: air conditioner". Finally, based on the natural language processing result of "API: AcOpen; ARGUMENTS: air conditioner", the vehicle control instruction that the vehicle can directly execute is determined.
[0143] In this way, based on the second prompt word, the server guides the large language model to determine the target voice request according to the current voice request and the historical voice requests. Then, the server determines the target intention based on the target voice request. Next, the server determines the natural language processing result based on the target intention. Finally, the server determines the vehicle control instruction based on the natural language processing result. In this way, the second prompt word can guide the large language model to reason according to the pre-designed thinking process, so as to accurately understand the user's intention and output the corresponding vehicle control instruction. Moreover, the second prompt word can help the large language model accurately understand the context information, thereby improving the semantic understanding ability and realizing natural and fluent multi-round interaction.
[0144] Please refer to Figure 9 , in some embodiments, step 02211 (based on the second prompt word, guiding the large language model to determine the target voice request according to the current voice request and the historical voice requests) includes:
[0145] 022111: When the current voice request is associated with the historical voice requests, perform a complement processing on the current voice request according to the historical voice requests to determine the target voice request;
[0146] 022112: When the current voice request is not associated with the historical voice requests, determine the current voice request as the target voice request.
[0147] In some embodiments, the determination module is used to perform a complement processing on the current voice request according to the historical voice requests to determine the target voice request when the current voice request is associated with the historical voice requests. And when the current voice request is not associated with the historical voice requests, determine the current voice request as the target voice.
[0148] In some embodiments, the processor is further used to perform a complement processing on the current voice request according to the historical voice requests to determine the target voice request when the current voice request is associated with the historical voice requests. And when the current voice request is not associated with the historical voice requests, determine the current voice request as the target voice request.
[0149] Specifically, if the current voice request is associated with a historical voice request, the large language model will, based on the historical voice request, understand the information that the user may have omitted in the current voice request. Through the completed full request, the large language model can more accurately identify the user's intention and convert it into executable APIs and parameters. For example, if the user first says "Open the window" and then says "Close it a little more", the large language model will understand that the user means "Close the window a little more".
[0150] If the current voice request is not associated with the historical voice request, that is, the current voice request has nothing to do with the historical voice request, the large language model will directly perform intention recognition and instruction generation based on the current request. For example, when the user says "Play music", the model will directly use the current voice request as the target voice request.
[0151] In this way, when the current voice request is associated with the historical voice request, the server completes the current voice request based on the historical voice request to determine the target voice request. Or when the current voice request is not associated with the historical voice request, the server determines the current voice request as the target voice request. In this way, by completing the current voice request, the user's intention can be accurately understood, avoiding ambiguity and misunderstanding. Moreover, by associating with the historical voice request, multi-round interaction can be achieved, making the conversation smooth and natural, thus enhancing the user experience.
[0152] Please refer to Figure 10 , in some embodiments, step 02213 (determine the natural language processing result according to the target intention) includes:
[0153] 022131: Perform slot recognition on the target voice request to determine the slot recognition result;
[0154] 022132: Perform application programming interface prediction according to the slot recognition result and the target intention to determine the predicted application interface;
[0155] 022133: According to the slot recognition result and the predicted application interface, select the predicted application interface to perform application programming interface parameter filling to determine the natural language processing result.
[0156] In some embodiments, the determination module is used to perform slot recognition on the target voice request to determine the slot recognition result. And perform application programming interface prediction according to the slot recognition result and the target intention to determine the predicted application interface. And according to the slot recognition result and the predicted application interface, select the predicted application interface to perform application programming interface parameter filling to determine the natural language processing result.
[0157] In some embodiments, the processor is further configured to perform slot recognition on the target voice request to determine the slot recognition result, perform application programming interface (API) prediction based on the slot recognition result and the target intent to determine the predicted API, and perform API parameter filling by selecting the predicted API according to the slot recognition result and the predicted API to determine the natural language processing result.
[0158] Specifically, slot recognition refers to extracting specific information fragments from the user's input. These information fragments are usually referred to as "slots". Slots are usually the key information necessary to complete a certain task or request, such as time, location, object, etc. Taking the user voice request "What's the temperature tomorrow?" as an example, the slot information obtained through slot recognition may include ["tomorrow" - Date], that is, the slot information includes the slot value and the slot type. Here, "tomorrow" is the slot value and Date is the slot type. Taking the user voice request "Navigate to address Q" as an example, the slot information obtained through slot recognition is ["address Q" - Place], where "Zhongguancun" is the slot value and Place is the slot type. Slot recognition is an important task in natural language processing. Its goal is to identify specific entities and attributes in the text and map them to predefined slots. Slot recognition can help the large language model module understand the user's intent and generate more accurate and natural responses.
[0159] The slot recognition result refers to the named entities obtained by performing slot recognition on the voice request. Such as the above slot information ["tomorrow" - Date] and slot information ["address Q" - Place], etc.
[0160] API prediction refers to predicting the operation type corresponding to the input text based on the semantics of the input text and generating the corresponding API call instruction. API prediction can help the large language model module understand the user's intent and perform the corresponding operations.
[0161] API parameter filling is a task in natural language processing. Its goal is to specify specific values for the parameters in the API call instruction according to the semantics of the input text and the API call instruction.
[0162] The large language model module first performs slot recognition on the target voice request, decomposing the user's instruction into different slots, such as song name, singer name, and play mode. Then, based on the slot recognition result and the target intent, the large language model module predicts the application programming interface that the user needs to execute, such as "play song" and "query weather". Finally, according to the slot recognition result and the predicted application interface, the large language model module selects the corresponding predicted application interface to perform the filling of the application programming interface parameters. For example, the song name and singer name are filled as parameters into the "play song" interface.
[0163] Finally, the large language model module outputs the natural language processing result. For example, it recognizes that the user's instruction is "play song" and calls the music player to play the song specified by the user.
[0164] In this way, the server performs slot recognition on the target voice request to determine the slot recognition result. Then, based on the slot recognition result and the target intent, the server predicts the application programming interface to determine the predicted application interface. Finally, according to the slot recognition result and the predicted application interface, the server selects the predicted application interface to perform the filling of the application programming interface parameters to determine the natural language processing result. In this way, based on the result of slot recognition, the predicted application programming interface is selected to perform the filling of the application programming interface parameters, and the server can understand and respond to the target voice request, directly output the execution result and send it to the vehicle to complete the voice interaction, thereby reducing the latency of the in-vehicle system, improving the response speed to the user's instruction, and providing a convenient and natural interaction experience.
[0165] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the vehicle control method as described above are implemented.
[0166] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution media, etc.
[0167] In the description of this specification, the descriptions referring to terms such as "specifically", "furthermore", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not necessarily intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0168] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment or part of executable instructions including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, not in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0169] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A method for training a large language model for vehicle voice interaction, characterized in that: The method comprises: Acquiring vehicle knowledge information, the vehicle knowledge information including function point information of vehicle function points and association information between the vehicle function points; Based on a preset mask language model, generating first training data according to the vehicle knowledge information; The base model is trained according to the first training data to obtain the large language model.
2. The training method according to claim 1, characterized in that: The generating the first training data based on the preset mask language model and according to the vehicle knowledge information includes: Based on the preset mask language model, the vehicle knowledge information is randomly masked to generate the first training data, wherein the first training data includes the vehicle knowledge information and tokenized knowledge information, wherein the tokenized knowledge information is the vehicle knowledge information with some words masked.
3. The training method according to claim 1, characterized in that: The method further comprises: Configure the first prompt word; The large language model is subjected to supervised fine-tuning training according to the first prompt word and pre-configured second training data.
4. The training method according to claim 3, characterized in that: The second training data includes single-round dialogue data and multi-round dialogue data, and the supervised fine-tuning training of the large language model according to the first prompt word and the pre-configured second training data includes: Based on the first prompt word, performing intent understanding training on the large language model according to the single-round dialogue data; Based on the first prompt word, semantic understanding training is performed on the large language model according to the multiple rounds of dialogue data.
5. A voice interaction method, characterized in that: The voice interaction method is based on a large language model trained by the training method according to any one of claims 1 to 4, and the method comprises: Get the current voice request; Based on the large language model, and according to the current voice request, determining a vehicle control instruction corresponding to the current voice request; The vehicle control instruction is sent to the vehicle so that the vehicle completes the voice interaction according to the vehicle control instruction.
6. The voice interaction method according to claim 5, characterized in that: The large language model is configured with a second prompt word, and determining a vehicle control instruction corresponding to the current voice request based on the large language model and according to the current voice request includes: Based on the second prompt word, the large language model is guided to determine the vehicle control instruction according to the current voice request and the historical voice requests.
7. The voice interaction method according to claim 6, characterized in that: The step of guiding the large language model to determine the vehicle control instruction according to the current voice request and the historical voice request based on the second prompt word includes: Based on the second prompt word, guiding the large language model to determine a target voice request according to the current voice request and the historical voice requests; Determining a target intention according to the target voice request; Determining a natural language processing result according to the target intention; The vehicle control instruction is determined according to the natural language processing result.
8. The voice interaction method according to claim 7, characterized in that: The step of guiding the large language model to determine a target voice request according to the current voice request and the historical voice requests based on the second prompt word includes: In a case where the current voice request is associated with the historical voice request, completing the current voice request according to the historical voice request to determine the target voice request; In a case where the current voice request is not associated with the historical voice request, the current voice request is determined as the target voice request.
9. The voice interaction method according to claim 7, characterized in that: Determining the natural language processing result according to the target intention includes: Performing slot recognition on the target voice request and determining a slot recognition result; Perform application program interface prediction based on the slot identification result and the target intention, and determine the predicted application interface; According to the slot identification result and the predicted application interface, the predicted application interface is selected to perform application program interface parameter filling to determine the natural language processing result.
10. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 9 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Interaction method, device, equipment, vehicle, readable storage medium and program product
CN120745826A
Intention recognition method, intention recognition device and intelligent glasses
CN120783736A
Voice interaction method, server and computer readable storage medium
CN120853560A
Voice interaction methods, servers, and computer-readable storage media
CN120853560B
Voice interaction method and device and electronic equipment
CN121075331A