A light application generation method and device, vehicle and electronic device
By collecting multimodal data of in-vehicle occupants and using lightweight application-based generative large models to generate DSL scripts, the problem of low operating efficiency in traditional in-vehicle interactive systems has been solved. This enables real-time and personalized lightweight application interface generation, improving user experience and system resource utilization efficiency.
Patent Information
- Application Number
- CN202511509087.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Traditional in-vehicle interaction systems suffer from low operational efficiency and poor scene adaptability, failing to meet users' flexible operational needs and immersive human-vehicle interaction experience in real time. Furthermore, existing lightweight application generation methods rely on pre-stored information and cannot generate interfaces that match user intent in real time as needed.
By collecting multimodal data from vehicle occupants, a lightweight application generative big data model is used to identify user intent, generate DSL scripts, and dynamically generate lightweight application interfaces that are adapted to user intent, including dialogue data, facial expressions, motion data, and physiological data. Combined with vehicle configuration information, it supports real-time generation and configuration of lightweight application interfaces.
It enables the rapid generation of lightweight application interfaces that are adapted to user intent, saves system resources, provides an immersive in-vehicle infotainment experience, reduces development costs, and improves the flexibility and personalization of lightweight application generation.
Smart Images

Figure CN121008801B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicles, in particular to the technical field of Internet of Vehicles, and specifically to a light application generation method and device, a vehicle and an electronic device. BACKGROUND
[0002] With the rapid iteration of artificial intelligence, sensor technology and intelligent cockpit, vehicle-mounted multi-modal interaction and generative artificial intelligence are key technologies on the road of intelligentization of vehicle-mounted systems. Traditional vehicle-mounted interaction systems mainly use touch and voice, which can cover most use scenarios, but this interaction mode still has problems such as low operation efficiency and poor scene adaptability. The traditional vehicle-machine interaction interface is an application that has been developed or a defined arrangement scene that consumes time for the user, and when facing new functions and scene requirements, it needs to be redeveloped and deployed, consuming a lot of manpower and resources, and cannot meet the flexible operation requirements and immersive human-vehicle interaction experience in real time. In this case, how to provide a real-time dynamic interaction page for users and the cockpit in a multi-modal scene and support the ability to execute ecological services is particularly important.
[0003] In a related technology, a light application running method is disclosed, which proposes to optimize the running efficiency of vehicle-mounted light applications through an intermediate application layer. However, this technology relies on pre-stored application development information and does not support the function of generating light applications in real time as needed by users. SUMMARY
[0004] The present application provides a light application generation method and device, a vehicle and an electronic device, which can temporarily generate a DSL script of a light application through real-time collected multi-modal data, and then generate an interface of the light application based on the DSL script, the interface including the configuration of at least one vehicle-mounted service covered by the light application. Without the need for pre-deployment of the light application, the light application interface containing multiple vehicle-mounted services that are adapted to the user's use intention is quickly generated and the related configuration options are completed, which greatly saves system resources and provides an immersive vehicle-machine experience that meets the user's intention in real time.
[0005] According to a first aspect of the present application, a light application generation method is provided, the method comprising:
[0006] Collecting multi-modal data of an occupant in a vehicle; the multi-modal data includes at least two of the following: dialogue data, facial expression data, action data, physiological data and body posture data;
[0007] Inputting the multi-modal data and vehicle configuration information into a light application generative large model, identifying the use intention of the occupant for a vehicle-mounted service in a light application through the light application generative large model, and generating a DSL script of the light application; the vehicle configuration information includes at least one of the following: vehicle type information, vehicle condition information, component configuration information and ecological service configuration information;
[0008] display, on a car machine interface, an interface of the light application based on the DSL script; the interface includes a configuration of at least one in-vehicle service in a set of in-vehicle services covered by the light application; the configuration is adapted to the use intention.
[0009] It can be seen that, by applying the embodiment scheme, multi-modal data of a passenger in a vehicle, including at least two of dialogue data, facial expression data, action data, physiological data, and body posture data, can be actively collected, the multi-modal data is input into a light application generative large model, the light application generative large model identifies a use intention of the passenger to an in-vehicle service of a light application, that is, an in-vehicle service that the passenger may want to use, generates a DSL script of the light application, and displays, on a car machine interface, an interface of the in-vehicle service that the passenger may want to use based on the DSL script, to recommend the in-vehicle service interface to the user for in-vehicle service selection; the interface includes a configuration of at least one in-vehicle service in a set of in-vehicle services covered by the light application; the configuration is adapted to the use intention. The DSL script of the light application can be temporarily generated based on the multi-modal data collected in real time, and the light application that meets the use demand of the user can be automatically generated based on the DSL script. The light application does not need to be deployed in advance, which greatly saves system resources, and only needs to call related pre-stored components for reuse when the light application is generated in real time. Moreover, the light application interface that is adapted to the use intention of the user and includes multiple in-vehicle services can be quickly generated and the related configuration options can be completed, which greatly saves system resources and provides an immersive car machine experience that meets the user's intention in real time.
[0010] In a possible manner, the DSL script includes a first DSL script for describing a set of in-vehicle services and / or a second DSL script for describing a single in-vehicle service; wherein the set of in-vehicle services includes a plurality of in-vehicle services pre-configured.
[0011] In a possible manner, the light application generative large model generates the DSL script of the light application by:
[0012] The light application generative large model performs semantic analysis on the multi-modal data to obtain a semantic analysis result;
[0013] The light application generative large model identifies the use intention of the passenger to the light application based on the vehicle configuration information and the semantic analysis result;
[0014] The light application generative large model generates the DSL script of the light application based on the use intention.
[0015] It can be seen that the generative light application provided in the embodiments of the present application has scalability and sustainability. In a traditional manner, custom development is required, and a large amount of human and material resources needs to be invested. The present system only needs to access corresponding ecological services (voice, navigation, vehicle control, etc.) to have the ability of most vehicle-mounted applications, greatly reducing the development cost.
[0016] In a possible manner, the light application generative large model generates a DSL script of the light application based on the use intention, including:
[0017] The light application generative large model generates a DSL script of the light application based on the use intention and the personalized configuration information of the occupant; wherein the personalized configuration information includes at least one of the following: preference information of the interface layout, preference information of the vehicle-mounted service.
[0018] It can be seen that, by applying the embodiments of the present application, when generating the DSL script, the use habits and preferences of the user are combined, and the interface layout and the interaction mode are dynamically adjusted, providing a more personalized experience.
[0019] In a possible manner, before inputting the multi-modal data into the light application generative large model, the method further includes:
[0020] obtaining historical data in a second time window before a first time window; the historical data includes historical action information and / or historical vehicle condition information; wherein the historical action information represents an action performed by the user on the vehicle; the first time window is a time window for obtaining the multi-modal data; the length of the second time window is a first preset time length;
[0021] inputting the multi-modal data into the light application generative large model, identifying the use intention of the occupant for the light application through the light application generative large model, and generating a DSL script of the light application, including:
[0022] inputting the multi-modal data and the historical data into the light application generative large model, identifying the use intention of the occupant for the light application through the light application generative large model, and generating a DSL script of the light application.
[0023] It can be seen that, in the embodiments of the present application, the action or vehicle condition information performed by the user in the previous time window is input into the light application generative large model as additional information. Since the action performed by the user in a relatively short time before also reflects the current intention to a certain extent, combining the above additional information can help the large model to more accurately identify the user intention, and further ensure that the generated light application meets the user demand.
[0024] In a possible manner, the method includes:
[0025] obtain a training sample set; a training sample in the training sample set includes a sample DSL script, sample multi-modal data corresponding to the sample DSL script, and a score;
[0026] train an initial light application generative large model based on the training sample set to obtain the light application generative large model.
[0027] It can be seen that in the embodiments of the present application, a training sample set containing sample DSL scripts, corresponding sample multi-modal data, and scores is obtained, and an initial light application generative large model is trained based on this to obtain a light application generative large model. Since the correspondence between the sample multi-modal data and the sample DSL script enables the model to learn the mapping rule from data to script, and the score can guide the model to optimize the generation quality, the trained model can more accurately generate light application DSL scripts that meet the needs and have high quality according to the input multi-modal data, effectively improving the accuracy and practicality of light application generation.
[0028] In a possible manner, the training sample further includes: sample historical data in a fourth time window before the third time window; the sample historical data includes sample historical action information and / or sample historical vehicle condition information; wherein the third time window is a time window for obtaining the sample multi-modal data; the length of the fourth time window is a second preset time length.
[0029] It can be seen that in the embodiments of the present application, sample historical data in a fourth time window before the third time window is additionally included in the training sample, which includes sample historical action information and / or sample historical vehicle condition information. Since the third time window is the time node for obtaining the sample multi-modal data, and the sample historical data in the fourth time window includes user historical actions and vehicle condition information, it can provide background basis for the user's intention when obtaining the sample multi-modal data in the time dimension. The user's behavior in the recent historical period is often related to the current intention. By integrating these sample historical data into the training, the initial light application generative large model can more comprehensively understand the complex relationship between multi-modal data and user intention in different situations in the learning process, and thus train a light application generative large model with better performance, which can finally more accurately meet the actual needs of the user when generating light application DSL scripts.
[0030] In a possible manner, the training of the initial light application generative large model based on the training sample set to obtain the light application generative large model includes:
[0031] The sample multi-modal data and the sample historical data are taken as training input, the sample DSL script is taken as training output, and the score is taken as satisfaction degree of the training output, so as to train the initial light application generative large model to obtain the light application generative large model.
[0032] It can be seen that, in the embodiment of the application, the initial model is trained to obtain the light application generative large model by taking the sample multi-modal data and the sample historical data (containing historical action / vehicle condition information) as training input, the sample DSL script as training output, and the score as an output satisfaction degree index. Thus, deep fusion learning of multi-dimensional information is realized, so that the model can capture the association rule between "current multi-modal data + recent historical behavior" and user intent; the score mechanism guides the model to iterate towards a high satisfaction degree through feedback optimization, thereby improving the accuracy and practicality of the generated script. Finally, the trained model can not only understand user demand scenarios more comprehensively, but also generate a light application DSL script that highly matches the actual intent of the user, thereby effectively enhancing the quality assurance and user demand matching capability of light application generation. In a possible manner, the interface of the light application is displayed on the car machine interface based on the DSL script, including:
[0033] The script code in the DSL script is parsed, and the script code includes at least one of service element description code, data source definition code, and UI layout description code;
[0034] Based on the parsing result, the interface of the light application is visually rendered on the car machine interface.
[0035] In a possible manner, the visual rendering on the car machine interface based on the parsing result includes:
[0036] A vehicle-mounted service component associated with the parsing result is determined;
[0037] A rendering order of the vehicle-mounted service component is determined;
[0038] Based on the rendering order, vehicle-mounted service interfaces of the vehicle-mounted service component are called in sequence, and based on data returned by the vehicle-mounted service interfaces, the interface of the light application is rendered on the car machine interface; the interface of the light application includes a display interface for displaying service content of the vehicle-mounted service component.
[0039] It can be seen that in the embodiment of the application, only access to the vehicle-mounted ecological service is needed to have the ability of most vehicle-mounted applications. Compared with the traditional custom development application mode, the required human and material resources are greatly reduced, and the scalability and sustainability of the generated light application are improved. In addition, the description code of the service card can be defined in advance, and the service card represents a set of vehicle-mounted services. When generating the DSL script for describing the light application, the description code of the service card can be directly called, so that one description code covers multiple vehicle-mounted services, making the DSL script more efficient in describing the light application, and improving the efficiency of generating the DSL script.
[0040] In a possible manner, the method further includes:
[0041] In the case that the light application described by the DSL script contains third-party ecological services, the API service interface for interfacing the third-party ecological services is accessed by calling the intermediate server associated with the third-party ecological services;
[0042] The visualization rendering is performed on the vehicle machine interface based on the parsing result, and the interface of the light application is displayed, including:
[0043] Based on the parsing result of the script code in the DSL script and the data returned by the API service interface, the interface of the light application is rendered on the vehicle machine interface; the interface of the light application contains a display interface for displaying the service content of the third-party ecological services.
[0044] As can be seen, through the above ecological fusion mode, it is not necessary to pre-deploy third-party ecological software on the vehicle side. In the case that the generated light application involves third-party ecological service capabilities, the corresponding services can be obtained by accessing the API interface provided by the third-party system through the intermediate server, which improves the flexibility of generating the light application and reduces the pressure of deploying software on the vehicle side.
[0045] In a possible manner, the method further includes:
[0046] In response to the denial instruction of the user to the light application, new multi-modal data of the occupant is obtained;
[0047] The new multi-modal data is input into the light application generative large model, the deviation between the generated light application and the use demand of the occupant is identified by the light application generative large model, and the adjusted DSL script of the light application is generated;
[0048] Based on the adjusted DSL script of the light application, the adjusted interface of the light application is displayed on the vehicle machine interface.
[0049] It can be seen that the embodiment of the application provides an intuitive direct interaction interface, provides a real-time generated light application according to user demand, and the real-time generated light application is not fixed and unchangeable. Before the light application takes effect, the user can receive an adjustment operation on the light application, and the related configuration information of the light application is adjusted. That is, when there is a system understanding deviation, the user can optimize through clarification multiple times, or customize the light application on the interface to meet the demand. The flexibility of generating the light application is further improved, and the user experience is improved.
[0050] In a possible manner, the method further includes:
[0051] In response to receiving the adjustment operation of the user on the light application, inputting adjustment information corresponding to the adjustment operation into the light application generative large model, generating a DSL script of the light application after adjustment by the light application generative large model;
[0052] Displaying the interface of the light application after adjustment on the car machine interface based on the DSL script of the light application after adjustment.
[0053] In a possible manner, the method further includes:
[0054] Based on the DSL script corresponding to the light application before adjustment and the DSL script corresponding to the light application after adjustment, optimizing the light application generative large model.
[0055] In a possible manner, the inputting the multi-modal data into the light application generative large model includes:
[0056] Time aligning the multi-modal data;
[0057] Preprocessing and feature extraction are performed on the time-aligned multi-modal data to obtain multi-modal feature information;
[0058] Inputting the multi-modal feature information into the light application generative large model.
[0059] In a possible manner, the inputting the multi-modal data into the light application generative large model includes: in the case of meeting a trigger condition, inputting the multi-modal data into the light application generative large model; the trigger condition includes: determining that the light application generation function is in an enabled state based on a user instruction.
[0060] According to the second aspect provided by the application, a light application generation device is provided, and the device includes:
[0061] The acquisition module is configured to collect multi-modal data of an occupant in a vehicle.
[0062] The generating module is configured to input the multi-modal data and vehicle configuration information into a light application generating large model, identify the use intention through the light application generating large model, and generate a DSL script of the light application; the vehicle configuration information includes at least one of the following: vehicle model information, vehicle condition information, component configuration information, and ecological service configuration information;
[0063] The display module is configured to display an interface of the light application on a vehicle machine interface based on the DSL script; the interface includes a configuration situation of at least one vehicle service in a vehicle service set covered by the light application; and the configuration situation is adapted to the use intention.
[0064] In a possible manner, the DSL script includes a first DSL script for describing a vehicle service group and / or a second DSL script for describing a single vehicle service; and the vehicle service group includes a plurality of vehicle services pre-configured.
[0065] In a possible manner, the light application generating large model generates the DSL script of the light application by:
[0066] The light application generating large model performs semantic analysis on the multi-modal data to obtain a semantic analysis result;
[0067] The light application generating large model identifies the use intention based on the vehicle configuration information and the semantic analysis result;
[0068] The light application generating large model generates the DSL script of the light application based on the use intention.
[0069] In a possible manner, the light application generating large model generates the DSL script of the light application based on the use intention, including:
[0070] The light application generating large model generates the DSL script of the light application based on the use intention and personalized configuration information of the occupant; and the personalized configuration information includes at least one of the following: preference information of an interface layout, and preference information of a vehicle service.
[0071] In a possible manner, the obtaining module is further configured to:
[0072] Obtain historical data in a second time window before a first time window; the historical data includes historical action information and / or historical vehicle condition information; the historical action information represents an action performed by a user on the vehicle; the first time window is a time window in which the multi-modal data is obtained; and the second time window has a first preset length;
[0073] inputting the multi-modal data into a light application generative large model, identifying, by the light application generative large model, a use intention of the occupant for a light application, and generating a DSL script of the light application, including:
[0074] inputting the multi-modal data and the historical data into the light application generative large model, identifying, by the light application generative large model, a use intention of the occupant for a light application, and generating a DSL script of the light application.
[0075] In a possible manner, the apparatus further includes a training module, and the training module is configured to:
[0076] obtain a training sample set, wherein each training sample in the training sample set includes a sample DSL script, sample multi-modal data corresponding to the sample DSL script, and a score;
[0077] train an initial light application generative large model based on the training sample set to obtain the light application generative large model.
[0078] In a possible manner, the training module is specifically configured to:
[0079] train the initial light application generative large model by taking the sample multi-modal data and the sample historical data as training input, the sample DSL script as training output, and the score as satisfaction degree of the training output, to obtain the light application generative large model.
[0080] In a possible manner, the display module is specifically configured to:
[0081] parse script code in the DSL script, wherein the script code includes at least one of service element description code, data source definition code, and UI layout description code;
[0082] visually render the car machine interface based on the parsing result to display an interface of the light application.
[0083] In a possible manner, the display module is specifically configured to:
[0084] determine a vehicle-mounted service component associated with the parsing result;
[0085] determine a rendering order of the vehicle-mounted service component;
[0086] based on the rendering order, sequentially call a vehicle-mounted service interface that interfaces the vehicle-mounted service component, and render the interface of the light application on the car machine interface based on data returned by the vehicle-mounted service interface; the interface of the light application includes a display interface for displaying service content of the vehicle-mounted service component.
[0087] In a possible implementation manner, the apparatus further includes a calling module, configured to, in a case where the light application described by the DSL script contains a third-party ecological service, call an intermediate server associated with the third-party ecological service to access an API service interface for interfacing with the third-party ecological service.
[0088] The display module is specifically configured to:
[0089] render an interface of the light application on the vehicle machine interface based on a parsing result of script code in the DSL script and data returned by the API service interface; the interface of the light application contains a display interface for displaying service content of the third-party ecological service.
[0090] In a possible implementation manner, the obtaining module is further configured to:
[0091] obtain new multi-modal data of the occupant in response to a denial instruction of the light application by the user;
[0092] The generation module is further configured to: input the new multi-modal data into the light application generative large model, identify a deviation between a generated light application and a use requirement of the occupant by using the light application generative large model, and generate an adjusted DSL script of the light application; and display an adjusted interface of the light application on the vehicle machine interface based on the adjusted DSL script of the light application.
[0093] In a possible implementation manner, the apparatus further includes an adjustment module, configured to:
[0094] in response to receiving an adjustment operation of the light application by the user,
[0095] input adjustment information corresponding to the adjustment operation into the light application generative large model, and generate an adjusted DSL script of the light application by using the light application generative large model;
[0096] display an adjusted interface of the light application on the vehicle machine interface based on the adjusted DSL script of the light application.
[0097] In a possible implementation manner, the multi-modal data includes at least two of the following: dialogue data, facial expression data, gesture data, physiological data, and body posture data.
[0098] In a possible implementation manner, the generation module is specifically configured to:
[0099] perform time alignment on the multi-modal data;
[0100] perform preprocessing and feature extraction on the time-aligned multi-modal data to obtain multi-modal feature information;
[0101] input the multi-modal feature information into the light application generation large model.
[0102] In one possible manner, the generation module is specifically configured to:
[0103] input the multi-modal data into the light application generation large model in the case of meeting a trigger condition; the trigger condition includes determining that the light application generation function is in an open state based on a user instruction.
[0104] According to a third aspect provided in the present application, a vehicle is provided, including the device of the above second aspect and any possible implementation manner thereof.
[0105] According to a fourth aspect provided in the present application, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; and wherein the processor is configured to execute the instructions to implement the method of the above first aspect and any possible implementation manner thereof.
[0106] According to a fifth aspect provided in the present application, a computer-readable storage medium is provided, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method of the above first aspect and any possible implementation manner thereof.
[0107] According to a sixth aspect provided in the present application, a computer program product is provided, the computer program product includes computer instructions, when the computer instructions are run on an electronic device, the electronic device executes the method of the above first aspect and any possible implementation manner thereof.
[0108] It should be noted that the technical effects brought by any implementation manner of the second aspect to the sixth aspect can refer to the technical effects brought by the corresponding implementation manner of the first aspect, which will not be repeated here.
[0109] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0110] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application, but are not to be construed as setting forth any limitation of the application.
[0111] Figure 1 is a flowchart of a light application generation method according to an exemplary embodiment;
[0112] Figure 2 is a flowchart of multi-modal preprocessing according to an exemplary embodiment;
[0113] Figure 3 FIG. 1 is a flow diagram of a light application DSL script generation method according to an example embodiment;
[0114] Figure 4 FIG. 2 is a flow diagram of another light application generation method according to an example embodiment;
[0115] Figure 5 FIG. 3 is a block diagram of a light application generation apparatus according to an example embodiment;
[0116] Figure 6 FIG. 4 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION
[0117] In order for those skilled in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings.
[0118] It should be noted that the terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in other than the order illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.
[0119] In the embodiments of the present application, the words "exemplary", "such as", or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design described as "exemplary", "such as", or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of "exemplary", "such as", or "for example" is intended to present concepts in a concrete manner.
[0120] In order to solve the problems of low operation efficiency and poor scene adaptability in the traditional car-machine interaction mode, the embodiments of the present application provide a light application generation method, device and vehicle.
[0121] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all.
[0122] Embodiments of the present application relate to a vehicle, which can also be referred to as a vehicle, a mobile carrier, an electric vehicle (EV), a hybrid electric vehicle (HEV), a plug-in hybrid electric vehicle (PHEV), a fuel cell vehicle (FCV), an autonomous vehicle, an intelligent and connected vehicle (ICV), a driverless vehicle, etc.
[0123] In embodiments of the present application, the vehicle can be a car, a sport utility vehicle (SUV), a truck, an electric vehicle, a motorcycle, a tricycle, a special vehicle (such as an ambulance, a fire truck, a police car, etc.), a driverless taxi, an intelligent and connected bus, an autonomous logistics vehicle, an electric truck, etc. In addition, the method is also applicable to various special vehicles, such as agricultural vehicles, mining vehicles, forestry vehicles, airport vehicles, port vehicles, etc. The present application does not make specific limitations in this regard.
[0124] In embodiments of the present application, the system involved mainly includes: a multi-modal interaction collection system, a large model training system, a large model carrying system, a light application generator, a smart adaptation system, and a configuration system. First, a brief introduction to the system involved is given as follows:
[0125] (1) Multi-modal interaction collection system: This system collects various perception heterogeneous data such as voice, gesture, and expression through multiple channels. The data is time-synchronized, attention-weighted fused, and scored by a modal confidence scoring mechanism to generate a unified semantic expression.
[0126] (2) Large model training system: This system is responsible for constructing a model database including component library metadata, domain-specific language (DSL) adaptation templates, and interaction history records. It uses knowledge enhancement and supervised fine-tuning (SFT) technology to provide models with the ability to accurately identify user intent and generate DSL scripts for light applications.
[0127] (3) Large model carrying system: This system carries a large model that has been specifically trained to preprocess, contextually analyze, and identify the semantic expression generated in the above steps, with the goal of understanding the user's intent and generating a DSL script based on it.
[0128] (4) Light application generator: The light application generator supports user customization of light applications, can import DSL scripts and convert them into interface elements that users can intuitively understand and use, dynamically display them to users, and allow users to directly modify and interact with them and execute related ecological services.
[0129] (5) Intelligent adaptation system: This system is used to capture user usage habits and preferences, build a personal profile, and further personalize services. Based on historical usage data, the system is continuously adjusted and optimized to better capture the user's personalized needs and provide more accurate services.
[0130] (6) Configuration system: This system realizes dynamic assembly of components, configuration of ecological services, and deployment of user individualization configuration files through cloud collaboration, supporting on-demand adaptation and real-time update of vehicle functions.
[0131] The light application generation method provided by the present application will be described in detail below with reference to the accompanying drawings. Referring to Figure 1 , the method can include the following steps:
[0132] S101: Collecting multi-modal data of passengers in the vehicle.
[0133] In the embodiments of the present application, the passengers include the driver and the passengers.
[0134] In some embodiments of the present application, the multi-modal data includes at least two of the following: dialogue data, facial expression data, action data, physiological data, and body posture data.
[0135] For example, the multi-modal data can include a driver, a driver's body temperature of 37.2 degrees, a driver's heart rate of 88 beats per minute, a breathing rate of 19 times per minute, a hand gesture of fanning, a line of sight area of the front, a pointing area of the window, an emotion of irritability, and voice information of "what to eat".
[0136] For example, the voice can be collected by the microphone in the vehicle, and the gesture, facial expression, and body posture data can be collected by the camera in the vehicle. For example, the microphone in the vehicle can collect the dialogue data between the driver and the passenger "the air temperature is high" and "it's too hot", and the camera in the vehicle can collect the action of the passenger fanning with his hand.
[0137] S102: Inputting the multi-modal data and vehicle configuration information into the light application generation large model, identifying the use intention of the passengers for the vehicle service in the light application through the light application generation large model, and generating the DSL script of the light application.
[0138] In the embodiments of the present application, a pre-trained light application generation large model can be integrated in a vehicle. As an example, the supported component library data, pre-edited scene templates, attribute / event specification documents, and interaction data are taken as a data set, the attributes, styles, and bound events of the components and templates are combined and expanded, and more than 500,000 data are provided through data enhancement and negative sample generation to feed the large model for training. Knowledge enhancement and SFT fine-tuning techniques are used to enable the pre-trained large model to accurately identify user intent and generate DSL scripts for light applications. After training is completed, the model is deployed to the vehicle through manual detection and automated testing.
[0139] In one embodiment of the present application, referring to Figure 2 The multi-modal data is input into the light application generation large model, including:
[0140] S201: Time aligning the multi-modal data.
[0141] As an example, the collected multi-modal data is time-stamped and aligned. For example, for data from different devices (cameras, microphones, etc.), the timestamps carried by the data are unified to a consistent time axis. Subsequently, the synchronized processing of the sorted time series data can be performed through a synchronization algorithm.
[0142] S202: Preprocessing and feature extraction of the time-aligned multi-modal data to obtain multi-modal feature information.
[0143] As an example, preprocessing can include data cleaning. For example, for audio data, noise elimination, noise reduction, and other processing can be performed.
[0144] For preprocessed data, feature extraction can be performed to obtain multi-modal feature information.
[0145] As an example, the dialog data collected by the microphone is subjected to feature extraction and text conversion. For gesture data collected by the camera, dynamic feature extraction and gesture judgment are performed. For facial expression data collected by the camera, key point tracking, feature point extraction, and micro-expression detection are performed.
[0146] S203: Inputting the multi-modal feature information into the light application generation large model.
[0147] The above embodiments are an example of the present application, and the model structure related to feature extraction can also be integrated in the light application generation large model. This is not limited.
[0148] In one embodiment of the present application, the light application generation large model receives multi-modal data or multi-modal feature information, identifies the use intent of the occupant for the light application based on the multi-modal data or multi-modal feature information, and generates a DSL script for the light application.
[0149] A DSL script refers to a domain-specific language focusing on a certain application field, which is different from a general cross-field general-purpose computer language. The domain-specific language is only used in certain specific fields. It has the characteristics of lightweight programming, simple structure, and easy system execution.
[0150] For example, for a car machine system, voice control of air conditioning, navigation, and other modules are usually supported, but the interaction process of different vehicle models may differ greatly. For example, the mapping relationship from voice instructions to device control differs, and a DSL script code can be used to define the car machine interaction process. In addition, the calling relationship between vehicle networking services and the processing rules of vehicle data can also be defined based on the DSL script code.
[0151] S103: Display the interface of the light application on the car machine interface based on the DSL script; the interface includes the configuration of at least one vehicle service in the vehicle service set covered by the light application; the configuration is adapted to the use intent.
[0152] In the embodiments of the present application, the light application on the vehicle can be understood as a lightweight and scenario-based application program specially designed for the automobile driving scene, which aims to provide users with safe, convenient, and personalized in-vehicle service experience by simplifying functions, optimizing interactions, and adapting to the vehicle environment.
[0153] In the embodiments of the present application, the light application on the vehicle can involve multiple vehicle services. Correspondingly, the interface of the generated light application can include the configuration of at least one vehicle service in the vehicle service set covered by the light application. The configuration is adapted to the use intent.
[0154] For example, a light application can involve multiple vehicle services such as "seat temperature adjustment", "air conditioner temperature adjustment", "window adjustment", etc. The light application interface can display the display content of each vehicle service. For example, for "seat temperature adjustment", an icon or picture related to the seat can be displayed in the light application interface.
[0155] Further, the light application interface displays the specific configuration of at least one vehicle service in the vehicle service set, which is related to the user's intent. For example, the light application generation model identifies the multi-modal data of the occupant and determines that the occupant is in a tired state and the current intent is to rest. The generated light application can involve multiple vehicle services such as "seat angle adjustment", "air conditioner temperature adjustment", and "window adjustment", and for "seat angle adjustment", the configuration is suitable for a reclining angle; for "air conditioner temperature adjustment", the configuration is suitable for a temperature suitable for rest; and for "window adjustment", the window is closed. As can be seen, the configuration of the above multiple vehicle services is associated with the user's use intent.
[0156] In one embodiment of the present application, after displaying the interface of the light application on the car machine interface, the light application is recommended to the user according to the specific configuration of the vehicle-mounted service covered by the light application, and then each vehicle-mounted service is run according to the user selection. For example, starting to adjust the window, seat, etc.
[0157] For example, according to the collected conversation data "the temperature is high" and "it is too hot" between the driver and the passenger, and the action of the occupant in the vehicle fanning with hands collected by the in-vehicle camera, the interface of the light application of starting the air conditioner and seat ventilation is generated and recommended to the user.
[0158] In another embodiment of the present application, displaying the interface of the light application on the car machine interface can be regarded as a preview interface of the vehicle-mounted service, and after receiving the confirmation of the user, each vehicle-mounted service is run according to the specific configuration of the vehicle-mounted service covered by the light application. For example, when the user has no objection to the vehicle-mounted service and the configuration of each vehicle-mounted service displayed in the light application, the user can issue a confirmation instruction through voice, a control button or other means, and then start to run each vehicle-mounted service according to the specific configuration of the vehicle-mounted service covered by the light application.
[0159] In the embodiment of the present application, the car machine system can integrate the light application generator to realize the parsing of various instructions in the DSL code.
[0160] In some embodiments of the present application, the light application generator parses the script code in the DSL script, and the script code includes at least one of service element description code, data source definition code and UI layout description code; and based on the parsing result, visual rendering is performed on the car machine interface to display the interface of the light application.
[0161] For example, the service element description code is used to describe the service content contained in the light application, for example, the light application to be generated contains the service of adjusting the seat, and the service element description code can contain structured information for describing the amplitude of seat adjustment, temperature, whether the massage function is turned on, etc. Generally, the light application to be generated sets multiple service contents, and each service content can be described by the service element description code.
[0162] For example, the data source definition code is used to describe the data source and modification method of various data in the light application. For example, the light application to be generated involves the acquisition of weather state information, and then the data source definition code can be used to define the acquisition method of the data, and in the process of generating the light application, the data is acquired through the corresponding acquisition method. The acquisition method can include the interface between the car machine system and the backend API.
[0163] For example, the UI layout description code is used to describe the interface layout of the light application to be generated, such as the arrangement and combination of elements in the light application interface, the jump logic and spatial relationship between the main page and the subpage, and the like.
[0164] After the DSL script is parsed, visual rendering is performed on the user interface, and the rendering result can include various graphical elements.
[0165] In the embodiments of the present application, the interface of the generated light application is interactive and can be adjusted in response to user operations. For example, the user can achieve service calls such as orchestration and vehicle control through simple clicking. For the orchestration and vehicle control services bound to the components, the user performs clicking, sliding and other operations, and the generated rendering engine identifies the service information bound to the components and performs corresponding operations. For example, the orchestration service sends an execution script to the scene orchestration, which parses, decides and executes the script.
[0166] In an embodiment of the present application, the vehicle service components and third-party ecological service capabilities can be ecologically integrated to achieve overall calling and orchestration.
[0167] For example, the vehicle service component can be a single vehicle service or a set of vehicle services. For example, if multiple vehicle services are frequently used together, they can be defined as a vehicle service component, which can also be referred to as a service card. For example, a service card is pre-configured to include three vehicle services: seat heating, seat massage and window adjustment.
[0168] In an embodiment of the present application, in the training stage of the light application generated large model, the training data includes DSL scripts for describing vehicle service components and DSL scripts for describing single vehicle services. Each vehicle service component includes multiple vehicle services bound together.
[0169] Correspondingly, the DSL script generated by the light application generated large model based on the multi-modal data can include a first DSL script for describing vehicle service components and / or a second DSL script for describing single vehicle services; wherein the vehicle service component includes a plurality of pre-configured vehicle services.
[0170] For example, in the case of binding the three vehicle services: seat heating, seat massage and window adjustment into a vehicle service component, the DSL script of the light application generated by the light application generated large model can include a DSL script for describing the vehicle service component, and also include a DSL script for describing a single vehicle service, such as "ambience light adjustment". Thus, the interface of the light application displayed on the car machine interface includes the display content of each vehicle service in the vehicle service component and the display content of the single vehicle service.
[0171] It can be seen that in the scheme, the description script of the vehicle-mounted service component can be defined in advance, and the DSL script for describing the light application can be directly generated to generate the DSL script corresponding to the vehicle-mounted service component, so that one script code covers multiple vehicle-mounted services, the DSL script can more efficiently describe the light application, and the efficiency of generating the DSL script is improved.
[0172] In an embodiment of the present application, the vehicle-mounted interface is visually rendered based on the analysis result to display the interface of the light application, which can specifically include: determining the vehicle-mounted service component associated with the analysis result; determining the rendering order of the vehicle-mounted service component; based on the rendering order, sequentially calling the vehicle-mounted service interface connected to the vehicle-mounted service component, rendering the interface of the light application on the vehicle-mounted interface based on the data returned by the vehicle-mounted service interface; the interface of the light application includes a display interface for displaying the service content of the vehicle-mounted service component.
[0173] For example, the vehicle-mounted service interface corresponding to each vehicle-mounted service component can be set in advance. The DSL script is analyzed to determine which vehicle-mounted service components are involved in the to-be-generated light application, and the rendering order of the vehicle-mounted service components is determined, so that the corresponding vehicle-mounted service interface is sequentially called based on the rendering order. In the process of rendering the interface of the light application, the interface is rendered based on the data returned by the vehicle-mounted service interface.
[0174] For example, the vehicle-mounted service component includes a navigation service and various types of vehicle control services. If the to-be-generated light application involves a certain vehicle-mounted service component, the corresponding vehicle-mounted service interface is called to obtain the required data for interface rendering. The required data can be determined based on the analysis result of the DSL script.
[0175] As can be seen, in the embodiments of the present application, only the vehicle-mounted ecological service needs to be accessed to have the ability of most vehicle-mounted applications. Compared with the traditional custom development application mode, the required human and material resources are greatly reduced, and the scalability and sustainability of the generated light application are improved.
[0176] In an embodiment of the present application, a rendering queue can be constructed based on the rendering order of the vehicle-mounted service component. If the vehicle-mounted service interface connected to the vehicle-mounted service component fails to render the interface of the light application on the vehicle-mounted interface based on the data returned by the vehicle-mounted service interface, an exception information is reported, and the current vehicle-mounted service interface is skipped to call the vehicle-mounted service interface connected to the next vehicle-mounted service component in the rendering queue. For example, in the case of using the intent for recommending nearby food, the data returned by the vehicle-mounted service interface can be "acquired nearby food 1, its characteristics are XXX", "picture of nearby food 1", "rating of nearby food 1", "hot comments of nearby food 1", and "address of nearby food 1".
[0177] In an embodiment of the present application, in the case that the light application described in the DSL script contains a third-party ecological service, an intermediate server associated with the third-party ecological service can be called to access an API service interface for interfacing the third-party ecological service.
[0178] For example, if the generated light application involves a third-party ecological server, such as an xx music website or an xxx video website, an API interface provided by the third-party system can be accessed through an intermediate server to obtain corresponding services. The intermediate server can act as a proxy server to avoid direct connection between the third-party system and the vehicle core bus or control system, thereby improving the security of data access.
[0179] Specifically, in the process of calling the third-party ecological service, a request conforming to a standard can be sent to the intermediate server, which performs parsing and verification processing and calls the third-party service or the vehicle-mounted service. If there is a corresponding service, the request result is returned to meet the user's demand.
[0180] Correspondingly, in an embodiment of the present application, the interface of the light application is visualized and rendered on the vehicle machine interface based on the parsing result, including rendering the interface of the light application on the vehicle machine interface based on the parsing result of the script code in the DSL script and the data returned by the API service interface. The interface of the light application contains a display interface for displaying the service content of the third-party ecological service.
[0181] Since the third-party ecological service is called, in the process of rendering the interface of the light application, the script parsing result and the data returned by the API service interface are combined to jointly render the interface. For example, if the light application to be generated contains playing a certain song in xx music, the API service interface interfacing xx music is called to obtain data related to the song for rendering the interface of the light application, and finally a display interface containing the service content of xx music is generated.
[0182] As can be seen, through the above ecological fusion method, it is not necessary to pre-deploy third-party ecological software on the vehicle side. In the case that the generated light application involves third-party ecological service capability, the API interface provided by the third-party system can be accessed through the intermediate server to obtain the corresponding service, thereby improving the flexibility of generating the light application and reducing the pressure of deploying software on the vehicle side.
[0183] By applying the embodiment scheme, multi-modal data of the occupant in the vehicle including at least two of dialogue data, facial expression data, action data, physiological data and body posture data can be actively collected, the multi-modal data is input into the light application generative large model, the use intention of the occupant to the in-vehicle service of the light application is recognized by the light application generative large model, that is, the in-vehicle service that the occupant may want to use, and the DSL script of the light application is generated; based on the DSL script, the interface of the in-vehicle service that the occupant may want to use is displayed on the vehicle machine interface, so as to recommend the in-vehicle service interface to the user for in-vehicle service selection; the interface includes the configuration situation of at least one in-vehicle service in the in-vehicle service set covered by the light application; the configuration situation is adapted to the use intention. The DSL script of the light application can be temporarily generated through the real-time collected multi-modal data, and the light application conforming to the use demand of the user can be automatically generated based on the DSL script. The light application does not need to be deployed in advance, and the system resources are greatly saved, and only the related pre-stored components need to be called for reuse when the light application is generated in real time.
[0184] In an implementation manner, referring to Figure 3 The light application generative large model can generate the DSL script of the light application by the following operations:
[0185] S301: The light application generative large model performs semantic analysis on the multi-modal data to obtain a semantic analysis result.
[0186] In the embodiment, the data of different modalities is converted to obtain a structured and abstract semantic representation that can be understood and operated by a machine. The semantic representation can be a vector or a structured text (for example, in json format).
[0187] For example, the core architecture of the light application generative large model can adopt a transformer architecture to realize multi-modal semantic analysis through self-attention and cross-attention mechanisms.
[0188] S302: The light application generative large model identifies the use intention of the occupant to the light application based on vehicle configuration information and the semantic analysis result.
[0189] As an example, the vehicle configuration information can include vehicle model information, vehicle condition information, component configuration information and ecological service configuration information. The vehicle model information can be used to associate the hardware configuration of the vehicle model, for example, whether it is provided with seat massage and seat electric adjustment. The vehicle condition information can be associated with the state of the vehicle, for example, whether the current state of the vehicle is suitable for starting a specific function. The component configuration information represents the components that can be called by the vehicle configuration, and the ecological service configuration information represents the software ecology that can be called by the vehicle configuration, for example, third-party music software and the like.
[0190] In the embodiment of the application, the light application generation type large model identifies the use intention of the light application of the occupant according to the semantic analysis result, in combination with the context information and the vehicle configuration information. The context information can include in-vehicle state information, user identity information, etc. The in-vehicle state information can include the current air conditioner temperature, seat state, etc.
[0191] S303: The light application generation type large model generates a DSL script of the light application based on the use intention.
[0192] The trained light application generation type large model has the ability to generate a DSL script. After determining the use intention of the light application, the DSL script corresponding to the light application is generated.
[0193] For example, after the user gets into the vehicle, the voice input "I am a little tired" is accompanied by a yawn action. The vehicle-mounted system first collects and processes the voice information and facial information in real time, performs data analysis and weighting to obtain multi-modal semantics. Then, the multi-modal semantics are input into the light application generation type large model. The light application generation type large model identifies the user intention through semantic understanding, in combination with the context and configuration information, obtains the use intention of the light application of the occupant, and further generates a DSL script of the light application. In this scenario, the target light application can be a "rest mode".
[0194] For example, the generated light application interface includes an interactive control element of a graphical element. For example, for the light application of the rest mode, the generated interface can display the adjustment angle of the seat in the vehicle, the temperature of the air conditioner, the wind direction, etc. If the user thinks that the light application meets his own needs, he can confirm it by simply clicking. After receiving the confirmation, the car machine system can adjust the seat and the air conditioner according to the information displayed in the light application interface.
[0195]
[0196] It can be seen that the generated light application provided in the embodiment of the application has extensibility and sustainability. The traditional method needs to be customized and developed, and a large amount of human and material resources need to be invested. The system only needs to access corresponding ecological services (voice, navigation, vehicle control, etc.) to have the ability of most vehicle-mounted applications, so that the development cost is greatly reduced.
[0197] In an embodiment of the application, before the multi-modal data is input into the light application generation type large model, the method further includes:
[0198] The historical data in a second time window before a first time window is obtained, and the historical data includes historical action information and / or historical vehicle condition information. The historical action information represents an action of the user acting on the vehicle; the first time window is a time window for obtaining the multi-modal data; and the length of the second time window is a first preset time length.
[0199] Optionally, the first preset time length can be set according to actual needs. For example, the first preset time length can be the same as the time length corresponding to the first time window. The present application does not make specific limitations on this.
[0200] For example, after the multi-modal data is collected in the first time window, the intention recognition can be performed based on the multi-modal data in combination with the historical action information and the historical vehicle condition information in the second time window before the first time window.
[0201] Optionally, the second time window can be multiple, that is, the historical action information and the historical vehicle condition information in multiple second time windows can be obtained for intention recognition.
[0202] For example, the user opens the window by himself or lowers the temperature of the air conditioner by himself.
[0203] Correspondingly, the multi-modal data is input into the light application generation large model, specifically including: inputting the multi-modal data and the historical data into the light application generation large model, recognizing the use intention of the occupant to the light application through the light application generation large model, and generating the DSL script of the light application.
[0204] In an implementation manner, the above historical data can be mixed in a specific prompt word template and input into the light application. For example, it is detected that the user has just opened the window by himself, and the intention recognition is performed in combination with this information.
[0205] It can be seen that in the embodiment of the present application, the action or vehicle condition information performed by the user in the previous time window is input into the light application generation large model as additional information. Since the action performed by the user in a short time before can also reflect the current intention to a certain extent, in combination with the above additional information, the large model can more accurately recognize the user intention, and further ensure that the generated light application meets the user demand.
[0206] In an embodiment of the present application, the light application generation large model generates the DSL script of the light application based on the use intention, including: the light application generation large model generates the DSL script of the light application based on the use intention and the personalized configuration information of the occupant; wherein the personalized configuration information includes at least one of the following: preference information of interface layout, preference information of vehicle-mounted service.
[0207] For example, the personalized configuration information of the user can be continuously optimized, and the vehicle-mounted system can collect the key interaction information of the user for optimizing the user preference, which can include the preference information of the interface layout and the preference information of the vehicle-mounted service.
[0208] As an example, the personalized configuration information can be optimized in an explicit feedback manner. For example, a feedback entry is provided in the light application display interface, and in response to a user operation, the personalized configuration information input by the user can be received.
[0209] As an example, the personalized configuration information can be optimized in an implicit feedback manner. For example, the behavior data of the user is analyzed to indirectly infer the preferences of the user.
[0210] Therefore, when generating the script of the DSL of the light application, the personalized configuration information of the occupant can be further combined. For example, in the process of generating the DSL script, the related preference settings are queried in the user preference library, and preference conflicts are handled. For example, if the user preference is "light color mode", and the current scene is night, the system can automatically generate the interface in the night mode according to the original rules, but since there is a conflict with the user preference, the mode can be optimized and adjusted.
[0211] It can be seen that by applying the embodiment scheme of the present application, when generating the DSL script, the use habits and preferences of the user are combined, and the interface layout and interaction mode are dynamically adjusted to provide a more personalized experience.
[0212] In an embodiment of the present application, after displaying the interface of the light application on the car machine interface, the method can further include: in response to a denial instruction of the user to the light application, obtaining new multi-modal data of the occupant; inputting the new multi-modal data into the light application generative large model, identifying the deviation between the generated light application and the use demand of the occupant through the light application generative large model, and generating the DSL script of the adjusted light application.
[0213] For example, if the user thinks that the generated light application does not meet the demand of the user, the denial instruction can be input. For example, the denial is performed in a voice dialogue manner or by clicking an interactive control on the car machine interface.
[0214] After receiving the denial instruction of the user to the light application, the new multi-modal data of the occupant can be reacquired. The new multi-modal data is processed, the deviation between the generated light application and the use demand of the occupant is identified by the capability of the large model, and the DSL script is further adjusted.
[0215] In an embodiment of the present application, after displaying the interface of the light application on the car machine interface, the method can further include: in response to receiving an adjustment operation of the user to the light application, inputting adjustment information corresponding to the adjustment operation into the light application generative large model, generating the DSL script of the adjusted light application through the light application generative large model; and displaying the adjusted interface of the light application on the car machine interface based on the DSL script of the adjusted light application.
[0216] For example, if the user considers that the generated light application deviates from the user's needs, the user can adjust the information to be adjusted displayed in the interface through the editing interface of the light application. For example, for the light application of the nap mode, if the user is not satisfied with the seat angle and the air conditioner blowing angle displayed in the interface of the light application, the user can change the seat angle displayed in the interface by manual operation. For example, the user slides the seat in the interface with a finger to change the angle of the seat.
[0217] Subsequently, the car machine system receives the adjustment operation of the user, and displays the adjusted interface of the light application on the car machine interface. If the user is satisfied with the information to be adjusted displayed in the adjusted interface, the user can confirm.
[0218] It can be seen that by using the embodiment of the present application, an intuitive and direct interaction interface is provided, and a real-time generated light application is provided for the user on demand. The real-time generated light application is not fixed and unchangeable, and before it takes effect, the adjustment operation of the user on the light application can be received to adjust the related configuration information of the light application. That is, when there is a system understanding deviation, the user can clarify and optimize multiple times, or customize the light application on the interface to meet the needs. The flexibility of generating the light application is further improved, and the user experience is improved.
[0219] In some embodiments of the present application, the multi-modal data is input into the light application generation large model, including: in the case of meeting the trigger condition, inputting the multi-modal data into the light application generation large model.
[0220] As an example, the trigger condition is set in advance, and the multi-modal data is input into the light application generation large model only in the case of meeting the trigger condition, thereby triggering the generation of the light application.
[0221] For example, the trigger condition includes determining that the light application generation function is in an open state based on the user instruction. For example, the user issues an open instruction of the light application generation function to the car machine system through the car machine operation or the voice instruction, and then the car machine system controls the light application generation function to be in an open state. In this state, the obtained multi-modal data of the occupant is input into the light application generation large model for intent recognition and subsequent light application generation.
[0222] It can be seen that if the multi-modal data of the occupant in the vehicle is obtained each time, and the intent recognition and the generation of the light application are triggered based on the multi-modal data, the light application may be generated frequently, which affects the normal use of the car machine by the user. By setting the trigger condition, the light application generation function is limited, and the car machine system is prevented from frequently generating the light application, which affects the execution of the direct instruction made by the user by the car machine system.
[0223] For ease of understanding, the light application generation method provided by the embodiments of the present application is further introduced below in combination with the drawings. Referring to Figure 4 , Figure 4 The model training process and the model use process are shown.
[0224] For the model training process, a data set for training is determined, including component library metadata, DSL adaptation templates, attribute description documents, and historical records. The data set is preprocessed, including structured conversion, data enhancement, and negative sample generation, and then iterative training is performed based on the data set.
[0225] For the model use process, multi-modal data acquisition and data processing are performed, and the processed data is input into the light application generative large model while carrying personalized configurations. The light application generative large model performs preprocessing, context association, intent recognition, and content generation to obtain a DSL script. Subsequently, the DSL script is input into a generative rendering engine, which includes a DSL parsing engine and a rendering engine. The DSL parsing engine calls a scene engine service through a vehicle-mounted service calling module. The scene engine is configured with a script parsing module, a decision module, an execution module, and a script management module. The rendering engine calls an intermediate server through an ecological service calling module, and the intermediate server is configured with a calling parsing module, a service authentication module, an ecological service execution module, and a database.
[0226] After the scene engine and the intermediate server are processed, the service calling result is returned to the generative rendering engine, and the generative rendering engine generates an interface of the light application that can be interacted according to the service calling result.
[0227] Subsequently, it is determined whether the user confirms, if yes, the light application is executed. And data acquisition is performed for the generation process of the light application, and the collected data is written into the data set for large model training as key interaction records.
[0228] If the user does not confirm, for example, the user clarifies the generated light application interface, the light application is not executed, and multi-modal data acquisition is returned again.
[0229] In yet another embodiment, for the model training process, the light application generation apparatus can obtain a training sample set, that is, a data set for training. Further, the light application generation apparatus can train an initial light application generative large model based on the training sample set to obtain the light application generative large model.
[0230] The training sample in the training sample set includes a sample DSL script, sample multi-modal data corresponding to the sample DSL script, and a score. The training sample further includes sample historical data in a fourth time window before a third time window, wherein the sample historical data includes sample historical action information and / or sample historical vehicle condition information. The third time window is a time window in which the sample multi-modal data is obtained, and the fourth time window has a length of a second preset time length.
[0231] Optionally, the second preset time length can be set according to actual needs. For example, the second preset time length can be the same as the length of the third time window. The present application does not make specific limitations on this.
[0232] Optionally, the fourth time window can be multiple, that is, multiple historical action information and historical vehicle condition information in the fourth time window can be obtained for model training.
[0233] It can be understood that the specific content of the sample multi-modal data can refer to the foregoing description of the multi-modal data, and the sample historical data can refer to the foregoing description of the multi-modal data. Details are not repeated here.
[0234] In a possible implementation manner, the score is an evaluation score for the effect of the sample DSL script. For example, the score range can be set between 1 and 5, wherein the higher the score value, the more the sample DSL script fits the human aesthetic standard. The score can be obtained by manually evaluating the sample DSL script by using a MOS (Mean Opinion Score) evaluation method, which can more intuitively measure the actual experience of the user on the DSL.
[0235] Specifically, the light application generation apparatus can train the initial light application generative large model by taking the sample multi-modal data and the sample historical data as training input, the sample DSL script as training output, and the score as the satisfaction degree of the training output, to obtain the light application generative large model.
[0236] It should be noted that the data fed to the initial light application generation model during the training process is only for helping the model learn the syntax rules of the DSL and master the appropriate layout in different scenarios, and is not for limiting the model to select and adapt within the range of the fed data. The above mainly introduces the scheme provided by the embodiments of the present application from the method aspect. In order to realize the above functions, the light application generation device or the electronic device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present application, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0237] The embodiments of the present application can divide the functional modules of the exemplary light application generation device or electronic device according to the above method. For example, the light application generation device or electronic device can include various functional modules corresponding to each functional division, or two or more functions can be integrated into one processing module. The integrated module can be realized in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. Actual implementation can have another division method.
[0238] Figure 5 is a block diagram of a light application generation device according to an exemplary embodiment. Referring to Figure 5 , the light application generation device includes:
[0239] The acquisition module 501 is configured to collect multi-modal data of an occupant in a vehicle.
[0240] The generation module 502 is configured to input the multi-modal data and vehicle configuration information into a light application generation model, identify, through the light application generation model, a use intention of the occupant for a vehicle service in a light application, and generate a DSL script of the light application; the vehicle configuration information includes at least one of the following: vehicle type information, vehicle condition information, component configuration information, and ecological service configuration information.
[0241] The display module 503 is configured to display an interface of the light application on a vehicle machine interface based on the DSL script; the interface includes a configuration situation of at least one vehicle service in a set of vehicle services covered by the light application; and the configuration situation is adapted to the use intention.
[0242] It can be seen that by applying the embodiment scheme of the present application, multi-modal data of a vehicle occupant including at least two of dialogue data, facial expression data, action data, physiological data, and body posture data can be actively collected, the multi-modal data is input into a light application generative large model, the use intention of the occupant for the vehicle service of the light application is recognized by the light application generative large model, that is, the vehicle service that the occupant may want to use, a DSL script of the light application is generated; based on the DSL script, an interface of the vehicle service that the occupant may want to use is displayed on the vehicle machine interface, to recommend the vehicle service interface to the user for vehicle service selection; the interface includes configuration of at least one vehicle service in the vehicle service set covered by the light application; the configuration is adapted to the use intention. The DSL script of the light application can be temporarily generated through real-time collection of multi-modal data, and the light application that meets the use demand of the user is automatically generated based on the DSL script. The light application does not need to be deployed in advance, which greatly saves system resources, and only needs to call related pre-stored components for reuse when generating the light application in real time. The light application interface containing multiple vehicle services that is adapted to the use intention of the user can be quickly generated and the related configuration options can be completed, which greatly saves system resources and provides immersive vehicle machine experience that meets the user's intention in real time.
[0243] In a possible manner, the DSL script includes a first DSL script for describing a vehicle service component and / or a second DSL script for describing a single vehicle service; wherein the vehicle service component includes a plurality of vehicle services pre-configured.
[0244] In a possible manner, the light application generative large model generates the DSL script of the light application by:
[0245] The light application generative large model performs semantic analysis on the multi-modal data to obtain a semantic analysis result;
[0246] The light application generative large model identifies the use intention of the occupant for the light application based on the vehicle configuration information and the semantic analysis result;
[0247] The light application generative large model generates the DSL script of the light application based on the use intention.
[0248] In a possible manner, the vehicle configuration information includes at least one of the following: vehicle model information, vehicle condition information, component configuration information, and ecological service configuration information.
[0249] In a possible manner, the light application generative large model generates the DSL script of the light application based on the use intention, including:
[0250] The light application generation type large model generates a DSL script of the light application based on the use intention and personalized configuration information of the occupant; wherein the personalized configuration information includes at least one of the following: preference information of interface layout, preference information of vehicle-mounted service.
[0251] In a possible manner, the obtaining module is further configured to:
[0252] obtain historical data in a second time window before the first time window; the historical data includes historical action information and / or historical vehicle condition information; wherein the historical action information represents an action performed by the user on the vehicle; the first time window is a time window in which the multi-modal data is obtained; the length of the second time window is a first preset time period;
[0253] input the multi-modal data into a light application generation type large model, identify the use intention of the occupant for the light application through the light application generation type large model, and generate a DSL script of the light application, including:
[0254] input the multi-modal data and the historical data into the light application generation type large model, identify the use intention of the occupant for the light application through the light application generation type large model, and generate a DSL script of the light application.
[0255] In a possible manner, the device further includes a training module; the training module is configured to:
[0256] obtain a training sample set; a training sample in the training sample set includes a sample DSL script, sample multi-modal data corresponding to the sample DSL script, and a score;
[0257] train an initial light application generation type large model based on the training sample set to obtain the light application generation type large model.
[0258] In a possible manner, the training module is specifically configured to:
[0259] train the initial light application generation type large model by taking the sample multi-modal data and the sample historical data as training input, taking the sample DSL script as training output, and taking the score as the satisfaction degree of the training output, to obtain the light application generation type large model.
[0260] In a possible manner, the display module is specifically configured to:
[0261] parse script code in the DSL script; the script code includes at least one of service element description code, data source definition code, and UI layout description code;
[0262] Based on the analysis result, visual rendering is performed on the vehicle machine interface to display the interface of the light application.
[0263] In one possible manner, the display module is specifically configured to:
[0264] determine the vehicle-mounted service component associated with the analysis result;
[0265] determine the rendering order of the vehicle-mounted service component;
[0266] based on the rendering order, sequentially call the vehicle-mounted service interface for interfacing the vehicle-mounted service component, and based on the data returned by the vehicle-mounted service interface, render the interface of the light application on the vehicle machine interface; the interface of the light application includes a display interface for displaying service content of the vehicle-mounted service component.
[0267] In one possible manner, the apparatus further includes a calling module configured to, in a case where the light application described in the DSL script includes a third-party ecological service, call an intermediate server associated with the third-party ecological service to access an API service interface for interfacing the third-party ecological service;
[0268] The display module is specifically configured to:
[0269] based on the analysis result of the script code in the DSL script and the data returned by the API service interface, render the interface of the light application on the vehicle machine interface; the interface of the light application includes a display interface for displaying service content of the third-party ecological service.
[0270] In one possible manner, the obtaining module is further configured to:
[0271] in response to a denial instruction of the light application from the user, obtain new multi-modal data of the occupant;
[0272] The generation module is further configured to: input the new multi-modal data into the light application generative large model, identify a deviation between a generated light application and a use demand of the occupant through the light application generative large model, generate an adjusted DSL script of the light application, and display an adjusted interface of the light application on the vehicle machine interface based on the adjusted DSL script of the light application.
[0273] In one possible manner, the apparatus further includes an adjustment module configured to:
[0274] in response to receiving an adjustment operation of the light application from the user,
[0275] input adjustment information corresponding to the adjustment operation into the light application generative large model, and generate an adjusted DSL script of the light application through the light application generative large model.
[0276] display, on the car-machine interface, an adjusted interface of the light application based on the adjusted DSL script of the light application.
[0277] In a possible implementation, the apparatus further includes an optimization module configured to:
[0278] optimize the light application generative large model based on the DSL script corresponding to the light application before adjustment and the DSL script corresponding to the light application after adjustment.
[0279] In a possible implementation, the multi-modal data includes at least two of the following: dialogue data, facial expression data, gesture data, physiological data, and body posture data.
[0280] In a possible implementation, the generation module is specifically configured to:
[0281] time-align the multi-modal data;
[0282] preprocess and extract features from the time-aligned multi-modal data to obtain multi-modal feature information;
[0283] input the multi-modal feature information into the light application generative large model.
[0284] In a possible implementation, the generation module is specifically configured to:
[0285] input the multi-modal data into the light application generative large model when a trigger condition is met; the trigger condition includes determining, based on a user instruction, that the light application generation function is in an enabled state.
[0286] Figure 6 is a block diagram of an electronic device according to an example embodiment. As shown in Figure 6 , the electronic device includes but is not limited to a processor 601 and a memory 602.
[0287] The memory 602 is configured to store executable instructions of the processor 601. It can be understood that the processor 601 is configured to execute the instructions to implement the light application generation method in the above embodiments.
[0288] It should be noted that those skilled in the art can understand that the structure of the electronic device shown in Figure 6 does not constitute a limitation on the electronic device. The electronic device can include more or fewer components than those shown in Figure 6 , or combine certain components, or arrange the components differently.
[0289] The processor 601 is the control center of the electronic device, connects all parts of the electronic device by various interfaces and lines, executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 602 and calling data stored in the memory 602, thereby overall monitoring the electronic device. The processor 601 can include one or more processing units. Alternatively, the processor 601 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 601.
[0290] The memory 602 can be used to store software programs and various data. The memory 602 can mainly include a program storage area and a data storage area, wherein the program storage area can store the operating system, the application programs (such as determination units, processing units, etc.) required by at least one function module, etc. In addition, the memory 602 can include a high-speed random access memory, and can also include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.
[0291] In the exemplary embodiments, a computer readable storage medium including instructions is also provided, for example, the memory 602 including instructions, which can be executed by the processor 601 of the electronic device to implement the methods in the above embodiments.
[0292] Alternatively, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, the non-transitory computer readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0293] In the exemplary embodiments, the embodiments of the present application also provide a computer program product including one or more instructions, which can be executed by the processor of the electronic device to complete the methods in the above embodiments.
[0294] It should be noted that the instructions in the above computer readable storage medium or the one or more instructions in the computer program product are executed by the processor of the electronic device to realize each process of the above method embodiments, and can achieve the same technical effects as the above methods. To avoid repetition, it will not be repeated here.
[0295] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0296] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0297] The units described as separate components can or can not be physically separate, and the components shown as units can be one physical unit or multiple physical units, that is, can be located in one place or can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0298] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0299] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions to make a device (which can be a single chip, a chip, etc.) or a processor execute all or part of the steps of the method of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk and various program code storage media.
[0300] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any change or replacement within the technical scope disclosed by the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A light application generation method, characterized by, The method comprises: Collecting multi-modal data of an occupant in a vehicle; the multi-modal data comprises at least two of the following: dialogue data, facial expression data, action data, physiological data, and body posture data; Inputting the multi-modal data and vehicle configuration information into a light application generative large model, identifying, by the light application generative large model, a use intention of the occupant for a vehicle-mounted service in a light application, and generating a DSL script of the light application; the vehicle configuration information comprises at least one of the following: vehicle model information, vehicle condition information, component configuration information, and ecological service configuration information; Displaying, based on the DSL script, an interface of the light application on a vehicle machine interface; the interface comprises a configuration condition of at least one vehicle-mounted service in a set of vehicle-mounted services covered by the light application; the configuration condition is adapted to the use intention.
2. The method of claim 1, wherein: The DSL script comprises a first DSL script for describing a vehicle-mounted service component and / or a second DSL script for describing a single vehicle-mounted service; wherein the vehicle-mounted service component comprises a plurality of vehicle-mounted services pre-configured.
3. The method of claim 1, wherein, The light application generative large model generates the DSL script of the light application by: The light application generative large model performs semantic analysis on the multi-modal data to obtain a semantic analysis result; The light application generative large model identifies the use intention based on the vehicle configuration information and the semantic analysis result; The light application generative large model generates the DSL script of the light application based on the use intention.
4. The method of claim 3, wherein, The light application generative large model generates the DSL script of the light application based on the use intention, comprising: The light application generative large model generates the DSL script of the light application based on the use intention and personalized configuration information of the occupant; wherein the personalized configuration information comprises at least one of the following: preference information of interface layout, preference information of vehicle-mounted service.
5. The method of claim 1, wherein, Before inputting the multi-modal data into the light application generative large model, the method further comprises: Obtaining historical data in a second time window before a first time window; the historical data comprises historical action information and / or historical vehicle condition information; wherein the historical action information represents an action performed by a user on the vehicle; the first time window is a time window in which the multi-modal data is obtained; the length of the second time window is a first preset time period; Inputting the multi-modal data into the light application generative large model, identifying, by the light application generative large model, a use intention of the occupant for a vehicle-mounted service in a light application, and generating a DSL script of the light application, comprising: Inputting the multi-modal data and the historical data into the light application generative large model, identifying, by the light application generative large model, the use intention, and generating the DSL script of the light application.
6. The method of claim 5, wherein, The method comprises: Obtaining a training sample set; a training sample in the training sample set comprises a sample DSL script, sample multi-modal data corresponding to the sample DSL script, and a score; The initial light application generative large model is trained based on the training sample set to obtain the light application generative large model.
7. The method of claim 6, wherein, The training sample further comprises sample historical data in a fourth time window before a third time window; the sample historical data comprises sample historical action information and / or sample historical vehicle condition information; the third time window is a time window for obtaining the sample multi-modal data; and the fourth time window has a second preset length.
8. The method of claim 7, wherein, The initial light application generative large model is trained based on the training sample set to obtain the light application generative large model, comprising: The sample multi-modal data and the sample historical data are taken as training inputs, the sample DSL script is taken as training output, and the score is taken as the satisfaction degree of the training output, so as to train the initial light application generative large model to obtain the light application generative large model.
9. The method of claim 1, wherein, The interface of the light application is displayed on the in-vehicle machine interface based on the DSL script, comprising: The script code in the DSL script is parsed, and the script code comprises at least one of service element description code, data source definition code, and UI layout description code; Based on the parsing result, the light application interface is visually rendered on the in-vehicle machine interface.
10. The method of claim 9, wherein, The parsing result is associated with a vehicle-mounted service component; The rendering order of the vehicle-mounted service component is determined; Based on the rendering order, the vehicle-mounted service interface connected to the vehicle-mounted service component is called in sequence, and the light application interface is rendered on the in-vehicle machine interface based on the data returned by the vehicle-mounted service interface; the light application interface comprises a display interface for displaying service content of the vehicle-mounted service component. The method further comprises:
11. The method of claim 9, wherein, In the case that the light application described in the DSL script comprises a third-party ecological service, an API service interface for connecting to the third-party ecological service is accessed by calling an intermediate server associated with the third-party ecological service; Based on the parsing result of the script code in the DSL script and the data returned by the API service interface, the light application interface is rendered on the in-vehicle machine interface; the light application interface comprises a display interface for displaying service content of the third-party ecological service. The method further comprises: In response to a denial instruction of the light application by the user, new multi-modal data of the occupant is obtained; 12. The method of claim 1, wherein, The new multi-modal data is input into the light application generative large model, the deviation between the generated light application and the use demand of the occupant is identified by the light application generative large model, and the DSL script of the adjusted light application is generated; Based on the DSL script of the adjusted light application, the interface of the adjusted light application is displayed on the in-vehicle machine interface. The method further comprises: 13. The method of claim 1, wherein, In response to receiving the adjustment operation of the user on the light application, input adjustment information corresponding to the adjustment operation into the light application generative large model, and generate a DSL script of the light application after adjustment through the light application generative large model; Display the interface of the light application after adjustment on the vehicle machine interface based on the DSL script of the light application after adjustment.
14. The method according to claim 12 or 13, characterized in that, The method further comprises: Based on the DSL script corresponding to the light application before adjustment and the DSL script corresponding to the light application after adjustment, optimize the light application generative large model.
15. The method of claim 1, wherein, The inputting of the multi-modal data into the light application generative large model comprises: Time aligning the multi-modal data; Preprocessing and feature extraction are performed on the time-aligned multi-modal data to obtain multi-modal feature information; Input the multi-modal feature information into the light application generative large model.
16. The method of claim 1, wherein, The inputting of the multi-modal data into the light application generative large model comprises: In the case of meeting the triggering condition, input the multi-modal data into the light application generative large model; The triggering condition comprises: determining that the light application generation function is in an enabled state based on a user instruction.
17. A light application generating apparatus characterized by comprising: The device comprises: An acquisition module for collecting multi-modal data of an occupant in a vehicle; the multi-modal data comprises at least two of the following: dialogue data, facial expression data, action data, physiological data, and body posture data; A generation module for inputting the multi-modal data and vehicle configuration information into a light application generative large model, identifying a use intention of the occupant for a vehicle service in a light application through the light application generative large model, and generating a DSL script of the light application; the vehicle configuration information comprises at least one of the following: vehicle model information, vehicle condition information, component configuration information, and ecological service configuration information; A display module for displaying an interface of the light application on a vehicle machine interface based on the DSL script; the interface comprises a configuration of at least one vehicle service in a set of vehicle services covered by the light application; the configuration is adapted to the use intention.
18. A vehicle characterized by comprising: The vehicle comprises the device of claim 17.
19. An electronic device, comprising: Comprise: A processor; A memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method of any one of claims 1-16.
Citation Information
Patent Citations
Low-code application development method and system based on AI auxiliary generation model
CN118760427A
Low-code development method and system based on hybrid orchestration of intelligent agent and intelligent service
CN120491977A