Large model-based in-vehicle intelligent scene generation method and apparatus, device, and storage medium
By leveraging large language models and knowledge graph technologies, user intent is identified and personalized vehicle scene cards are generated. This addresses the issue of insufficient interactive experience in traditional in-vehicle intelligent systems, enabling more intelligent and personalized scene generation and improving user satisfaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUIZHOU DESAY SV AUTOMOTIVE
- Filing Date
- 2025-06-27
- Publication Date
- 2026-06-04
Smart Images

Figure CN2025104562_04062026_PF_FP_ABST
Abstract
Description
Method, apparatus, equipment, and storage medium for generating intelligent in-vehicle scenes based on large models
[0001] This application claims priority to Chinese Patent Application No. 202411721582.2, filed with the Chinese Patent Office on November 28, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of vehicle control technology, and for example to a method, apparatus, device and storage medium for generating in-vehicle intelligent scenes based on a large model. Background Technology
[0003] With the development of the automotive industry, especially the rise of intelligent connected vehicle technology, the intelligentization of in-vehicle systems has become a key focus of the industry. Modern cars are no longer just means of transportation; they are transforming into intelligent mobile spaces that integrate multimedia entertainment, navigation, communication, and safety assistance.
[0004] Currently, traditional in-vehicle intelligent systems generally suffer from poor interactive experiences. For example, the scene modes provided by these systems are often relatively fixed, lacking the ability to dynamically adjust based on user preferences and real-time needs. Furthermore, during human-machine interaction, users typically need to communicate with the system using fixed command formats. All of these limitations restrict the user experience and lead to decreased user satisfaction. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for generating in-vehicle intelligent scenes based on a large model, in order to solve the problem of poor interaction experience between vehicles and users.
[0006] Firstly, this application provides a method for generating in-vehicle intelligent scenes based on a large model, including:
[0007] Receive user input information and output user intent information based on the user input information using a preset large language model;
[0008] A new vehicle scene card is generated based on the user intent information, or a target vehicle scene card that matches the user intent information is determined using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards.
[0009] Control the current vehicle according to the actions executed in the new vehicle scene card or the target vehicle scene card.
[0010] Secondly, this application provides an in-vehicle intelligent scene generation device based on a large model, comprising:
[0011] The user intent recognition module is configured to receive user expression information input by the user and output user intent information based on the user expression information using a preset large language model;
[0012] The vehicle scene card determination module is configured to generate a new vehicle scene card based on the user intent information, or to determine a target vehicle scene card that matches the user intent information using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards.
[0013] The vehicle control module is configured to control the current vehicle based on the actions executed in the new vehicle scene card or the target vehicle scene card.
[0014] Thirdly, this application provides an electronic device comprising:
[0015] At least one processor;
[0016] and memory that is communicatively connected to at least one processor;
[0017] The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to execute the large-model-based vehicle intelligent scene generation method of the first aspect described above.
[0018] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a processor to execute the above-described method for generating intelligent vehicle scenes based on a large model.
[0019] This application provides a vehicle-mounted intelligent scene generation solution based on a large-scale model. It receives user input, uses a preset large-scale language model to output user intent information, and determines whether to generate a new vehicle scene card based on the intent information, or uses a knowledge graph to determine a target vehicle scene card matching the intent information. The knowledge graph includes historical vehicle scene cards. The system controls the current vehicle based on the actions executed in the new or target vehicle scene card. By employing this technical solution, the large-scale language model enhances the natural language processing capabilities of the vehicle system, accurately identifying user intent. Furthermore, by utilizing this user intent and knowledge graph technology, the accuracy of scene recognition is improved, allowing the system to better serve the user. By generating new scene cards or finding matching historical scene cards, the solution addresses the lack of flexibility and personalization in scene generation, providing a more intelligent and personalized scene generation solution and enhancing the user's interactive experience.
[0020] It should be understood that the description in this section is not intended to identify key or important features of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a flowchart of a vehicle intelligent scene generation method based on a large model according to Embodiment 1 of this application;
[0023] Figure 2 is a flowchart of a vehicle intelligent scene generation method based on a large model according to Embodiment 2 of this application;
[0024] Figure 3 is a schematic diagram of the structure of an in-vehicle intelligent scene generation device based on a large model according to Embodiment 3 of this application;
[0025] Figure 4 is a schematic diagram of the structure of an electronic device according to Embodiment 4 of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. In the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0028] Example 1
[0029] Figure 1 is a flowchart of a vehicle intelligent scene generation method based on a large model provided in Embodiment 1 of this application. This embodiment can be applied to the situation of controlling a vehicle using a vehicle intelligent scene. The method can be executed by a vehicle intelligent scene generation device based on a large model. The vehicle intelligent scene generation device based on a large model can be implemented in hardware and / or software. The vehicle intelligent scene generation device based on a large model can be configured in an electronic device in the vehicle. The electronic device can be composed of two or more physical entities, or it can be composed of a single physical entity.
[0030] As shown in Figure 1, the in-vehicle intelligent scene generation method based on a large model provided in Embodiment 1 of this application specifically includes the following steps:
[0031] S101. Receive user expression information input by the user, and output user intent information based on the user expression information using a preset large language model.
[0032] In this embodiment, user input, such as voice and / or text information, can be received first. Then, a pre-trained large language model is used to process the user input to obtain user intent information.
[0033] S102. Generate a new vehicle scene card based on the user intent information, or determine a target vehicle scene card that matches the user intent information using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards.
[0034] In this embodiment, a knowledge graph can be used to first determine whether a target vehicle scene card matching the user's intent information exists in the historical vehicle scene cards. If not, a new vehicle scene card can be generated based on the user's intent information. The knowledge graph can contain associations between multiple historical vehicle scene cards and user intent information. The vehicle scene card can include actions to be performed, such as opening a car window.
[0035] S103. Control the current vehicle according to the execution action in the new vehicle scene card or the target vehicle scene card.
[0036] The technical solution of this application embodiment enhances the natural language processing capability of the vehicle system by utilizing a large language model, accurately identifying user intent. Then, by utilizing the user intent and knowledge graph technology, the accuracy of scene recognition is improved, allowing the system to better serve users. By generating new scene cards or finding matching historical scene cards, the problem of lack of flexibility and personalization in scene generation is solved, providing a more intelligent and personalized scene generation solution and enhancing the user's interactive experience.
[0037] Optionally, before receiving user input information, the method further includes: generating historical vehicle scene cards using the components of preset vehicle scene cards, wherein the components include scene name, scene description, scene category, execution action, and triggering conditions.
[0038] Specifically, a scene card template library can be pre-established, and various vehicle scene cards can be pre-defined. The components of a vehicle scene card include: scene name, scene description, scene category, execution action, and triggering conditions.
[0039] For example, scenario names could be "Work Mode" and "Lunch Break Mode," etc. Scenario descriptions are brief descriptions of each scenario to help users understand its function. Examples include: "Creating an efficient and focused environment to help users maintain clear thinking and a good state of mind," and "Creating a comfortable and quiet environment to help users quickly relax, enjoy a short break, and effectively recover their physical and mental well-being." Scenario categories are based on the scenario's function and application area. Examples include: "Car Usage Tips," "Emotional Experience," "Comfortable Space," and "Efficient Travel." Execution actions are the specific actions corresponding to each scenario. Examples include: "Open the window," "Turn off the air conditioner," "Adjust the lights and turn on the fragrance diffuser," "Adjust the seat's undulation effect," and "Turn on the air circulation system and play QQ Music, and play today's recommended playlist," etc. Trigger conditions are the conditions under which the scenario is triggered. For example, the trigger condition for "Work Mode" is: "Start Time: 07:00, End Time: 10:00, Repeat Days: Legal Working Days."
[0040] Optionally, before receiving the user's input expression information, the method further includes: for each historical vehicle scene card, determining the association relationship between the constituent elements of the current historical vehicle scene card, and generating triples based on the association relationship; and generating a knowledge graph using the triples.
[0041] Specifically, entities (i.e., constituent elements) and their relationships can be extracted from historical vehicle scene cards to generate triples. These triples can then be used to create a knowledge graph for the scene cards. A triple is a common knowledge graph representation, formatted as (node 1, edge, node 2). In scene cards, a triple is represented as (main entity, relation, object entity). Each triple consists of two entities and their corresponding relation. Both entities and relations possess descriptive concepts and attributes. Concepts refer to the type and category of the entity or relation, while attributes are its inherent parameters, characteristics, and properties. A knowledge graph is formed by interconnecting multiple triples.
[0042] For example, the relationships between entities include the following four types:
[0043] 1) BELONGS_TO (belongs to) indicates the category to which the entity (such as scene name) belongs;
[0044] 2) DESCRIBED_AS() represents a detailed description of an entity (such as a scene name);
[0045] 3) HAS_ACTION() represents the action to be performed contained in an entity (such as a scene name);
[0046] 4) HAS_TRIGGER indicates the triggering conditions contained in an entity (such as a scene name).
[0047] Triples include:
[0048] (Scene name, belongs to, category), (Scene name, described as, description), (Description, belongs to, category), (Scene name, includes execution action, execution action), (Scene name, includes trigger condition, trigger condition), (Execution action, described as, execution action), (Trigger condition, described as, trigger condition), etc. Among these, if the trigger condition is a complex condition or contains multiple sub-conditions, or the execution action is a complex action or contains multiple sub-actions, further refinement is needed. For example, if the user input is "get in the car," then there could be two corresponding sub-conditions: "Driver's seat: occupied" and "Ignition status: on."
[0049] For example, a specific vehicle scene card can be represented as:
[0050] 1) Scene Name: Lunch Break Mode;
[0051] 2) Scene Description: Create a comfortable and tranquil environment to help users relax quickly, enjoy a short break, and effectively restore their physical and mental well-being;
[0052] 3) Scene category: Comfortable space;
[0053] 4) Actions: Close the car windows, turn on the air freshener, adjust the seat's swaying effect, and turn on the air circulation system;
[0054] 5) Triggering condition: Time is from 12:00 to 13:00.
[0055] The corresponding list of triples is as follows:
[0056] 1)(Lunch Break Mode, BELONGS_TO, Comfort Space);
[0057] 2)(Lunch break mode, DESCRIBED_AS, creates a comfortable and quiet environment to help users relax quickly, enjoy a short break, and allow their body and mind to recover effectively.
[0058] 3)(Creating a comfortable and tranquil environment helps users relax quickly, enjoy a short break, and allow their body and mind to recover effectively. (BELONGS_TO, Comfortable Space)
[0059] 4)(Lunch break mode, HAS_ACTION, close the windows);
[0060] 5)(Lunch break mode, HAS_ACTION, turn on the fragrance sprayer);
[0061] 6)(Lunch break mode, HAS_ACTION, adjusts the seat undulation effect);
[0062] 7)(Lunch break mode, HAS_ACTION, activates the air circulation system).
[0063] One approach is to use the graph database Neo4j to build a relational model and store these entities and relationships in the form of a graph.
[0064] Optionally, the step of outputting user intent information based on the user's expression information using a preset large language model includes: mining user intent information from the user's expression information using the preset large language model based on a first preset prompt, wherein the first preset prompt includes user intent generation guidance information, and the user intent information includes vehicle scene card creation requirement information or emotional expression information; wherein, the step of generating a new vehicle scene card based on the user intent information, or determining a target vehicle scene card matching the user intent information using a knowledge graph includes: if the user intent information includes vehicle scene card creation requirement information, then generating a new vehicle scene card based on the vehicle scene card creation requirement information; if the user intent information includes emotional expression information, then determining a target vehicle scene card matching the user intent information using a knowledge graph.
[0065] Specifically, large models (such as Qwen or GLM) can be used to identify the user's input intent, thus recognizing the type of user's need (i.e., user intent information). Before inputting into the large model, the user's input can be preprocessed. This preprocessing includes text cleaning, removing irrelevant characters (such as redundant spaces and special symbols), standardization (unifying letter case and converting common abbreviations), and word segmentation. For example, if a user inputs "The air outside is terrible," the preprocessed output would be: "outside" "air" "good" "bad".
[0066] Then, the preprocessed user expression information is input into the large model, and a pre-defined first preset prompt is used to assist the large model in determining the type of user need. This setting can increase the accuracy of the large model's output. Types include scene card creation (i.e., vehicle scene card creation requirement information) and emotion expression (i.e., emotion expression information). The first preset prompt can include user intent generation guidance information such as input / output examples of the large model. Specifically, vehicle scene card creation requirement information includes direct vehicle control commands output by the user, while emotion expression information includes implicit vehicle control needs expressed by the user (indirect vehicle control commands), such as information expressing emotions or feelings.
[0067] Next, after identifying the user intent information, if the user intent information includes vehicle scene card creation requirement information, a new vehicle scene card is generated based on the vehicle scene card creation requirement information. If the user intent information includes emotional expression information, a knowledge graph is used to determine the target vehicle scene card that matches the user intent information.
[0068] For example, the first preset prompt can be represented as:
[0069] 1. You are an expert in in-vehicle intelligent scene recognition;
[0070] 2. Define the following user requirement types:
[0071] - Create scene cards (e.g., when I get in the car, the air conditioning automatically turns on and music plays).
[0072] - Expressing emotions (e.g., I want to sleep for a bit)
[0073] 3. Define the following scene categories:
[0074] - Car usage tips
[0075] -Emotional experience
[0076] -Comfortable space
[0077] -Efficient travel
[0078] 4. The returned results will only include the requirement type and scenario category that have been defined above;
[0079] 5. Based on user input, determine the type of user's needs and the category of the scenario, rewrite the user input into a form that better matches the scenario description, and extract keywords and phrases from the user input.
[0080] 6. The results are returned in JSON format.
[0081] 7. Example:
[0082] - Enter: "When I get in the car, the air conditioning automatically turns on and music plays."
[0083] Output: {"Requirement Type":"Create Scene Card","Scenario Category":"Car Usage Tips","Keywords and Phrases":["Get in the Car","Turn on Air Conditioning","Play Music"],"Rewritten Query":"Automatically turn on the air conditioning and play music when the user gets in the car"}
[0084] Type: "I want to take a nap"
[0085] - Output: {"Requirement Type":"Emotional Expression","Scenario Category":"Comfortable Space","Keywords and Phrases":["Want to Sleep"],"Rewritten Query":"I feel a little sleepy and hope to have a quiet and comfortable environment to take a nap"}
[0086] question:
[0087] This user query belongs to which of the following two categories: 1. Creating scene cards, 2. Emotional expression; determine the scene category, rewrite the user input into a form that better matches the scene description, and extract keywords and phrases.
[0088] Example 2
[0089] Figure 2 is a flowchart of a vehicle intelligent scene generation method based on a large model provided in Embodiment 2 of this application. The technical solution of this embodiment is further optimized based on the above optional technical solutions, and gives a specific way of controlling the vehicle using vehicle intelligent scenes.
[0090] Optionally, the step of determining the generation of a new vehicle scene card based on the user intent information includes: if the user intent information includes vehicle scene card creation requirement information, then extracting first target information from the user expression information using a natural language processing model, wherein the first target information is related to the constituent elements of the historical vehicle scene cards; generating an initial vehicle scene card based on the first target information, keywords, and phrases using the preset large language model and the second preset prompt, wherein the user intent information also includes the keywords and the phrases, and the second preset prompt includes vehicle scene card generation guidance information; and generating a new vehicle scene card by calculating the similarity between the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph. The advantage of this setup is that, based on the RAG (Retrieval-Augmented Generation) concept, when the user intent information includes vehicle scene card creation requirement information, an initial vehicle scene card is generated using a large model, and a new vehicle scene card is accurately generated through similarity calculation. This achieves automatic generation of personalized scene cards according to the user's actual needs, ensuring that the generated scene cards not only meet user needs but also have high accuracy and personalization.
[0091] Optionally, determining the target vehicle scene card matching the user intent information using a knowledge graph includes: if the user intent information includes emotional expression information, then extracting second target information from the user expression information and the rewritten statement using a natural language processing model, wherein the second target information is related to the constituent elements of the historical vehicle scene card, and the user intent information also includes the rewritten statement; determining the target vehicle scene card matching the user intent information by performing similarity calculations on the second target information, keywords, and phrases, and the historical vehicle scene cards in the knowledge graph, wherein the user intent information also includes the keywords and phrases. The advantage of this setup is that when the user intent information includes emotional expression information, the use of a large model and similarity calculation enhances natural language processing capabilities, improves the accuracy of scene recognition, promotes the high degree of personalization and intelligence of the in-vehicle intelligent system, and enhances the convenience and comfort of user-vehicle interaction.
[0092] As shown in Figure 2, the in-vehicle intelligent scene generation method based on a large model provided in Embodiment 2 of this application specifically includes the following steps:
[0093] S201. Generate historical vehicle scene cards by using the constituent elements of preset vehicle scene cards; for each historical vehicle scene card, determine the relationship between the constituent elements of the current historical vehicle scene card, and generate triples based on the relationship; use the triples to generate a knowledge graph; wherein, the constituent elements include scene name, scene description, scene category, execution action, and triggering condition.
[0094] S202. Receive user expression information input by the user, and use a preset large language model to extract user intent information from the user expression information based on a first preset prompt, wherein the first preset prompt includes user intent generation guidance information, and the user intent information includes vehicle scene card creation requirement information or emotional expression information.
[0095] S203. Determine whether the user intent information includes vehicle scene card creation requirement information. If yes, proceed to step 204; otherwise, proceed to step 207.
[0096] S204. Extract first target information from the user's expressed information using a natural language processing model, wherein the first target information is related to the constituent elements of the historical vehicle scene card.
[0097] Specifically, if the user intent information includes vehicle scene card creation requirement information, a natural language processing model (such as the SpaCy model) can be used to extract entity words or phrases related to the constituent elements of historical vehicle scene cards from the user's expression information, i.e., the first target information.
[0098] S205. Using the preset large language model and the second preset prompt, generate an initial vehicle scene card based on the first target information, keywords and phrases, wherein the user intent information further includes the keywords and phrases, and the second preset prompt includes vehicle scene card generation guidance information.
[0099] Specifically, a second preset prompt can be set in advance to assist the preset large language model in generating initial vehicle scene cards based on keywords and phrases in the first target information and user intent information. The second preset prompt includes vehicle scene card generation examples and other guidance information for vehicle scene card generation.
[0100] For example, the second preset prompt can be represented as:
[0101] 1. You are an expert in in-vehicle intelligent scene recognition;
[0102] 3. Define the following scene categories:
[0103] - Car usage tips
[0104] -Emotional experience
[0105] -Comfortable space
[0106] -Efficient travel
[0107] 4. The scene categories returned in the results should only be those already defined above;
[0108] 6. The results are returned in JSON format.
[0109] 7. Example of a scene card:
[0110] Example 1:
[0111] Scene Name: Lunch Break Mode
[0112] Scenario Description: To create a comfortable and tranquil environment that helps users relax quickly, enjoy a short break, and effectively restore their physical and mental well-being.
[0113] Scene Category: Comfort Space
[0114] Triggering conditions:
[0115] Actions performed: Close the car windows, turn on the air freshener, adjust the seat's sway function, and turn on the air circulation system.
[0116] Example 2:
[0117] Scene Name: Poor Air Quality Mode
[0118] Scenario Description: When the outside air quality is poor, close all car windows and the sunroof, and activate the air purification system to protect the air quality inside the car.
[0119] Scenario Category: Car Usage Tips
[0120] Triggering conditions:
[0121] Actions to be taken: Close all windows and sunroof; activate the air purification system.
[0122] ##question
[0123] Generate a scene card based on user input and suggested keywords and phrases.
[0124] ##enter
[0125] {user_input}
[0126] ## Output
[0127] {
[0128] Scene Name:"",
[0129] "Scene Description":"",
[0130] Scene Category:"",
[0131] "Execute action":[],
[0132] Triggering condition: []
[0133] }
[0134] For example, when the first target information of the preset large language model is: "When I get in the car, the air conditioning will automatically turn on and music will play," the keywords and phrases are: "get in the car," "turn on the air conditioning," and "play music." The initial vehicle scene card output by the preset large language model is: {"Scenario Name":"Welcome Mode","Scenario Description":"When the user gets in the car, the vehicle automatically adjusts to a comfortable initial state, including turning on the air conditioning to adjust the temperature and playing music to provide the user with a pleasant riding experience.","Scenario Category":"Efficient Travel","Action Execution":["Turn on the air conditioning","Play music"],"Trigger Condition":["Get in the car"]}.
[0135] S206. Generate a new vehicle scene card by calculating the similarity between the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph. Execute step 211.
[0136] Specifically, similarity calculations can be performed on the content of the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph. The target content in the historical vehicle scene card with the highest similarity is written into the initial vehicle scene card, such as the content corresponding to the action, to obtain the new vehicle scene card.
[0137] Furthermore, the step of generating a new vehicle scene card by performing similarity calculations between the initial vehicle scene card and historical vehicle scene cards in the knowledge graph includes: performing similarity calculations between the trigger conditions and execution actions in the constituent elements of the initial vehicle scene card and the trigger conditions and execution actions in the constituent elements of historical vehicle scene cards in the knowledge graph to obtain a first calculation result; and replacing the trigger conditions and execution actions in the initial vehicle scene card with the trigger conditions and execution actions corresponding to the highest first calculation result to generate a new vehicle scene card.
[0138] Optionally, the similarity calculation includes TF-IDF similarity calculation and BERT similarity calculation.
[0139] Specifically, the trigger conditions and actions in the initial vehicle scene card can be compared with the trigger conditions and actions in the knowledge graph for similarity calculation. The trigger conditions and actions of the historical vehicle scene card with the highest similarity in the knowledge graph are determined as the best match. The trigger conditions and actions in the initial vehicle scene card are then replaced with the trigger conditions and actions corresponding to the highest similarity to generate a new vehicle scene card. TF-IDF cosine similarity is typically used for statistical text representation, while BERT is a pre-trained deep learning model that provides richer and more semantic text representations. Combining these two methods to calculate text similarity comprehensively utilizes statistical and semantic information, improving the accuracy and robustness of similarity calculation.
[0140] For example, the methods for determining the first calculation result include:
[0141] The first calculation result = α * TF-IDF similarity calculation result + (1-α) * BERT similarity calculation result;
[0142] Where α is a preset weight parameter, such as 0.4.
[0143] S207. Determine whether the user intent information includes emotional expression information. If yes, proceed to step 208; otherwise, proceed to step 202.
[0144] Specifically, if the user intent information does not include emotional expression information and vehicle scene card creation requirement information, step 202 can be executed to re-receive the user expression information input by the user.
[0145] S208. Extract second target information from the user expression information and the rewritten statement using a natural language processing model, wherein the second target information is related to the constituent elements of the historical vehicle scene card, and the user intent information also includes the rewritten statement.
[0146] Specifically, if the user intent information includes sentiment expression, a natural language processing model (such as the SpaCy model) can be used to extract entity words or phrases related to the components of the historical vehicle scene card from the user's expressed information and the rewritten statement, i.e., the second target information. The rewritten statement is a modified version of the user's expressed information, which is more standardized, fluent, and suitable for processing by large models.
[0147] S209. By performing similarity calculations on the second target information, keywords and phrases, and historical vehicle scene cards in the knowledge graph, a target vehicle scene card matching the user intent information is determined, wherein the user intent information also includes the keywords and phrases.
[0148] Specifically, the similarity between the second target information, keywords and phrases and the content in the historical vehicle scene cards in the knowledge graph can be calculated, and the historical vehicle scene card with the highest similarity can be identified as the target vehicle scene card that matches the user's intent information.
[0149] Furthermore, the step of determining the target vehicle scene card matching the user intent information by performing similarity calculations between the second target information, keywords, and phrases and the historical vehicle scene cards in the knowledge graph includes: performing similarity calculations between the scene descriptions in the historical vehicle scene cards in the knowledge graph and the second target information, keywords, and phrases respectively to obtain a second calculation result; and determining the historical vehicle scene card corresponding to the highest second calculation result as the target vehicle scene card matching the user intent information.
[0150] Specifically, the historical vehicle scene cards in the knowledge graph can be traversed, and the similarity between each scene description under each scene category and the second target information, keywords, and phrases can be calculated. The historical vehicle scene card corresponding to the scene description with the highest similarity is determined as the target vehicle scene card that matches the user intent information.
[0151] S210. Control the current vehicle according to the execution action in the target vehicle scene card.
[0152] Specifically, before executing step 210, a target vehicle scene card can be displayed to allow the user to confirm and receive feedback. After user confirmation, the current vehicle can be automatically controlled according to the actions specified in the target vehicle scene card.
[0153] S211. Control the current vehicle according to the execution action in the new vehicle scene card.
[0154] Specifically, before executing step 211, a new vehicle scene card can be displayed to allow the user to confirm and provide feedback. After user confirmation, the new vehicle scene card can be saved and updated in the knowledge graph to obtain an updated knowledge graph. This allows for continuous optimization and expansion of the knowledge graph, enabling the provision of more intelligent and personalized services based on user descriptions.
[0155] The in-vehicle intelligent scene generation method based on a large model provided in this application is based on the RAG (Retrieval-Augmented Generation) concept. When the user's intent information includes vehicle scene card creation requirements, an initial vehicle scene card is generated using a large model, and a new vehicle scene card is accurately generated through similarity calculation. This realizes the automatic generation of personalized scene cards according to the user's actual needs, ensuring that the generated scene cards not only meet the user's needs but also have high accuracy and personalization. When the user's intent information includes emotional expression information, the large model and similarity calculation enhance natural language processing capabilities, improve the accuracy of scene recognition, promote the high degree of personalization and intelligence of the in-vehicle intelligent system, and improve the convenience and comfort of user interaction with the vehicle.
[0156] Example 3
[0157] Figure 3 is a schematic diagram of a vehicle-mounted intelligent scene generation device based on a large model provided in Embodiment 3 of this application. As shown in Figure 3, the device includes: a user intent recognition module 301, a vehicle scene card determination module 302, and a vehicle control module 303, wherein:
[0158] The user intent recognition module is configured to receive user expression information input by the user and output user intent information based on the user expression information using a preset large language model;
[0159] The vehicle scene card determination module is configured to generate a new vehicle scene card based on the user intent information, or to determine a target vehicle scene card that matches the user intent information using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards.
[0160] The vehicle control module is configured to control the current vehicle based on the actions executed in the new vehicle scene card or the target vehicle scene card.
[0161] The in-vehicle intelligent scene generation device based on a large model provided in this application enhances the natural language processing capabilities of the in-vehicle system by utilizing a large language model, accurately identifying user intent. By leveraging this user intent and knowledge graph technology, the accuracy of scene recognition is improved, allowing the system to better serve users. By generating new scene cards or finding matching historical scene cards, the device solves the problems of lack of flexibility and personalization in scene generation, providing a more intelligent and personalized scene generation solution and enhancing the user's interactive experience.
[0162] Optionally, the device may also include:
[0163] The historical scene card generation module is configured to generate historical vehicle scene cards by using preset vehicle scene card components before receiving user input information. The components include scene name, scene description, scene category, execution action, and triggering conditions.
[0164] Optionally, the device may also include:
[0165] The triple generation module is configured to determine the association relationship between the constituent elements of the current historical vehicle scene card for each historical vehicle scene card before receiving the user expression information input by the user, and generate triples according to the association relationship.
[0166] The knowledge graph generation module is configured to generate a knowledge graph using the triples.
[0167] Optionally, the user intent recognition module includes:
[0168] The user intent recognition unit is configured to use a preset large language model to extract user intent information from the user's expression information based on a first preset prompt. The first preset prompt includes user intent generation guidance information, and the user intent information includes vehicle scene card creation requirement information or emotional expression information.
[0169] Optionally, the vehicle scene card determination module includes:
[0170] The new scene card determination unit is configured to generate a new vehicle scene card based on the vehicle scene card creation requirement information if the user intent information includes vehicle scene card creation requirement information.
[0171] The first target scene card generation unit is configured to determine the target vehicle scene card that matches the user intent information by using a knowledge graph if the user intent information includes emotional expression information.
[0172] Optionally, the vehicle scene card determination module includes:
[0173] The first target information extraction unit is configured to extract first target information from the user expression information using a natural language processing model if the user intent information includes vehicle scene card creation requirement information, wherein the first target information is related to the constituent elements of the historical vehicle scene card.
[0174] The initial scene card generation unit is configured to generate an initial vehicle scene card based on the first target information, keywords, and phrases using the preset large language model and the second preset prompt. The user intent information further includes the keywords and the phrases, and the second preset prompt includes vehicle scene card generation guidance information.
[0175] The new scene card generation unit is configured to generate a new vehicle scene card by calculating the similarity between the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph.
[0176] Furthermore, the step of generating a new vehicle scene card by performing similarity calculations between the initial vehicle scene card and historical vehicle scene cards in the knowledge graph includes: performing similarity calculations between the trigger conditions and execution actions in the constituent elements of the initial vehicle scene card and the trigger conditions and execution actions in the constituent elements of historical vehicle scene cards in the knowledge graph to obtain a first calculation result; and replacing the trigger conditions and execution actions in the initial vehicle scene card with the trigger conditions and execution actions corresponding to the highest first calculation result to generate a new vehicle scene card.
[0177] Optionally, the vehicle scene card determination module includes:
[0178] The second target information extraction unit is configured to extract second target information from the user expression information and the rewritten statement using a natural language processing model if the user intent information includes emotional expression information. The second target information is related to the constituent elements of the historical vehicle scene card, and the user intent information also includes the rewritten statement.
[0179] The second target scene card generation unit is configured to determine the target vehicle scene card that matches the user intent information by performing similarity calculations between the second target information, keywords and phrases and historical vehicle scene cards in the knowledge graph, wherein the user intent information also includes the keywords and phrases.
[0180] Furthermore, the step of determining the target vehicle scene card matching the user intent information by performing similarity calculations between the second target information, keywords, and phrases and the historical vehicle scene cards in the knowledge graph includes: performing similarity calculations between the scene descriptions in the historical vehicle scene cards in the knowledge graph and the second target information, keywords, and phrases respectively to obtain a second calculation result; and determining the historical vehicle scene card corresponding to the highest second calculation result as the target vehicle scene card matching the user intent information.
[0181] The vehicle-mounted intelligent scene generation device based on a large model provided in this application can execute the vehicle-mounted intelligent scene generation method based on a large model provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.
[0182] Example 4
[0183] Figure 4 illustrates a schematic diagram of an electronic device 40 that can be used to implement embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0184] As shown in Figure 4, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded into the RAM 43 from storage unit 48. The RAM 43 can also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0185] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0186] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as a large-model-based in-vehicle intelligent scene generation method.
[0187] In some embodiments, the large-model-based in-vehicle intelligent scene generation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the large-model-based in-vehicle intelligent scene generation method described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to execute the large-model-based in-vehicle intelligent scene generation method by any other suitable means (e.g., by means of firmware).
[0188] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), hybrid programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0189] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0190] The computer equipment provided above can be used to execute the vehicle-mounted intelligent scene generation method based on a large model provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0191] Example 5
[0192] In the context of this application, a computer-readable storage medium may be a tangible medium, and the computer-executable instructions, when executed by a computer processor, are used to perform a method for generating in-vehicle intelligent scenes based on a large model, the method comprising:
[0193] Receive user input information and output user intent information based on the user input information using a preset large language model;
[0194] A new vehicle scene card is generated based on the user intent information, or a target vehicle scene card that matches the user intent information is determined using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards.
[0195] Control the current vehicle according to the actions executed in the new vehicle scene card or the target vehicle scene card.
[0196] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by, or in conjunction with, an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0197] The computer equipment provided above can be used to execute the vehicle-mounted intelligent scene generation method based on a large model provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0198] It is worth noting that in the above embodiments of the vehicle-mounted intelligent scene generation device based on large models, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this application.
Claims
1. A method for generating in-vehicle intelligent scenes based on a large model, comprising: Receive user input of user expression information, and output user intent information based on the user expression information using a preset large language model; A new vehicle scene card is generated based on the user intent information, or a target vehicle scene card that matches the user intent information is determined using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards. Control the current vehicle according to the actions executed in the new vehicle scene card or the target vehicle scene card.
2. The method according to claim 1, further comprising, before receiving user-inputted user expression information: Historical vehicle scene cards are generated by using preset vehicle scene card components, wherein the components include scene name, scene description, scene category, execution action, and triggering conditions.
3. The method according to claim 2, further comprising, before receiving user-inputted user expression information: For each historical vehicle scene card, determine the association relationship between the constituent elements of the current historical vehicle scene card, and generate triples based on the association relationship; A knowledge graph is generated using the triples.
4. The method according to claim 1, wherein, The step of using a preset large language model to output user intent information based on the user's expressed information includes: Using a preset large language model, user intent information is extracted from the user's expression information based on a first preset prompt. The first preset prompt includes user intent generation guidance information, and the user intent information includes vehicle scene card creation requirement information or emotional expression information. The step of determining whether to generate a new vehicle scene card based on the user intent information or to determine a target vehicle scene card that matches the user intent information using a knowledge graph includes: If the user intent information includes vehicle scene card creation requirement information, then a new vehicle scene card is generated based on the vehicle scene card creation requirement information; If the user intent information includes emotional expression information, then a target vehicle scene card matching the user intent information is determined using a knowledge graph.
5. The method according to claim 1, wherein, The step of determining and generating a new vehicle scene card based on the user intent information includes: If the user intent information includes vehicle scene card creation requirement information, then a first target information is extracted from the user expression information using a natural language processing model, wherein the first target information is related to the constituent elements of the historical vehicle scene card; Using the preset large language model and the second preset prompt, an initial vehicle scene card is generated based on the first target information, keywords, and phrases. The user intent information also includes the keywords and phrases, and the second preset prompt includes vehicle scene card generation guidance information. A new vehicle scene card is generated by calculating the similarity between the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph.
6. The method according to claim 5, wherein, The step of generating a new vehicle scene card by calculating the similarity between the initial vehicle scene card and the historical vehicle scene cards in the knowledge graph includes: The trigger conditions and execution actions in the constituent elements of the initial vehicle scene card are compared with the trigger conditions and execution actions in the constituent elements of the historical vehicle scene cards in the knowledge graph to obtain the first calculation result. The trigger conditions and execution actions corresponding to the highest first calculation result are used to replace the trigger conditions and execution actions in the initial vehicle scene card to generate a new vehicle scene card.
7. The method according to claim 1, wherein, The step of using a knowledge graph to determine the target vehicle scene card that matches the user intent information includes: If the user intent information includes emotional expression information, then a second target information is extracted from the user expression information and the rewritten statement using a natural language processing model. The second target information is related to the constituent elements of the historical vehicle scene card, and the user intent information also includes the rewritten statement. By performing similarity calculations on the second target information, keywords, and phrases, and the historical vehicle scene cards in the knowledge graph, a target vehicle scene card matching the user intent information is determined, wherein the user intent information also includes the keywords and phrases.
8. The method according to claim 7, wherein, The step of determining the target vehicle scene card that matches the user intent information by calculating the similarity between the second target information, keywords and phrases, and historical vehicle scene cards in the knowledge graph includes: The scene descriptions in the historical vehicle scene cards in the knowledge graph are compared with the second target information, keywords, and phrases to calculate the similarity and obtain the second calculation result. The historical vehicle scene card corresponding to the highest second calculation result is determined as the target vehicle scene card that matches the user intent information.
9. The method according to any one of claims 5-8, wherein, The similarity calculation includes TF-IDF similarity calculation and BERT similarity calculation.
10. A vehicle-mounted intelligent scene generation device based on a large model, comprising: The user intent recognition module is configured to receive user expression information input by the user and output user intent information based on the user expression information using a preset large language model; The vehicle scene card determination module is configured to generate a new vehicle scene card based on the user intent information, or to determine a target vehicle scene card that matches the user intent information using a knowledge graph, wherein the knowledge graph includes historical vehicle scene cards. The vehicle control module is configured to control the current vehicle based on the actions executed in the new vehicle scene card or the target vehicle scene card.
11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the vehicle-mounted intelligent scene generation method based on a large model as described in any one of claims 1-9.
12. A computer-readable storage medium storing computer instructions for causing a processor to execute the in-vehicle intelligent scene generation method based on a large model as described in any one of claims 1-9.