Vehicle interaction interface generation method and device based on multi-agent cooperation, equipment and medium
By employing a multi-agent collaborative method for generating vehicle user interfaces, and utilizing intent recognition and streaming rendering technologies, the problem of low development efficiency and weak personalization capabilities in intelligent vehicle HMIs has been solved, enabling rapid and personalized interface generation and improved user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUIZHOU DESAY SV AUTOMOTIVE
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
The development efficiency of existing human-machine interfaces (HMIs) for intelligent vehicles is low, their personalization capabilities are weak, and their iteration cycles are long. They cannot dynamically adjust core elements such as interface layout and font size according to user preferences or driving scenarios, resulting in modification cycles that can last for weeks or even months.
A vehicle interaction interface generation method based on multi-agent collaboration is adopted. By acquiring user interaction information and vehicle multimodal data, an intent recognition model is used to identify the set of intent commands. The interface generation model and streaming rendering technology are combined to dynamically adjust or generate personalized interaction interfaces.
It improves the efficiency of generating interactive interfaces, enables real-time adjustment and personalized adaptation of user interfaces, shortens the interface response time to within 10 seconds, and enhances user experience and security.
Smart Images

Figure CN121900863A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and medium for generating a vehicle interactive interface based on multi-agent collaboration. Background Technology
[0002] With the rapid development of intelligent connected vehicle technology, the in-vehicle human-machine interface (HMI) has become a core carrier for improving the user's driving experience. Its ease of interaction and scenario adaptability directly affect the user's operational safety and user experience.
[0003] Currently, the human-machine interface (HMI) of intelligent vehicles is mainly developed using traditional manual coding methods. The HMI interfaces in in-vehicle systems generally employ a fixed, pre-built template approach. This means that during the system development phase, developers write the corresponding HMI interfaces (such as music playback, navigation, and vehicle control interfaces) into static templates and store them permanently in the in-vehicle system. Users can only access these pre-built, fixed interfaces while driving, unable to dynamically adjust the layout, font size, or other core elements based on user preferences or driving scenarios. Even minor changes to the interface (such as button positions or font sizes) require rewriting code and conducting integration testing, resulting in modification cycles that can last for weeks or even months. Therefore, this fixed HMI interface technology has significant limitations and is no longer able to meet the diverse needs of users and the complex requirements of driving scenarios. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for generating vehicle interactive interfaces based on multi-agent collaboration, which solves the problems of low development efficiency, weak personalization capabilities, and long iteration cycles of vehicle interactive interfaces. It can dynamically adjust or generate personalized interactive interfaces that meet user needs based on user interaction information, thereby improving interface generation efficiency and enhancing the user's interactive experience.
[0005] According to one aspect of the present invention, a method for generating a vehicle interaction interface based on multi-agent cooperation is provided, comprising:
[0006] Acquire user interaction information and vehicle multimodal data;
[0007] Intent recognition is performed on the interaction information and the multimodal data to obtain a set of intent commands;
[0008] The initial interactive interface of the vehicle is determined based on the set of intent commands and the interface generation model; wherein the interface generation model is composed of multiple intelligent agents.
[0009] The initial interactive interface is rendered using streaming rendering technology to obtain the target interactive interface of the vehicle.
[0010] According to another aspect of the present invention, a vehicle interaction interface generation device based on multi-agent cooperation is provided, comprising:
[0011] The data acquisition module is used to acquire user interaction information and vehicle multimodal data;
[0012] The intent recognition module is used to perform intent recognition on the interaction information and the multimodal data to obtain a set of intent commands;
[0013] The interface determination module is used to determine the initial interactive interface of the vehicle based on the intent instruction set and the interface generation model; wherein the interface generation model is composed of multiple intelligent agents.
[0014] The interface rendering module is used to render the initial interactive interface using streaming rendering technology to obtain the target interactive interface of the vehicle.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the vehicle interaction interface generation method based on multi-agent cooperation as described in any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the vehicle interaction interface generation method based on multi-agent cooperation as described in any embodiment of the present invention.
[0020] The technical solution of this invention involves acquiring user interaction information and vehicle multimodal data; performing intent recognition on the interaction information and multimodal data to obtain an intent command set; determining the vehicle's initial interactive interface based on the intent command set and an interface generation model; wherein the interface generation model is composed of multiple intelligent agents; and rendering the initial interactive interface using streaming rendering technology to obtain the vehicle's target interactive interface. This technical solution addresses the problems of low development efficiency, weak personalization capabilities, and long iteration cycles in vehicle interactive interfaces. It can dynamically adjust or generate personalized interactive interfaces that meet user needs based on user interaction information, improving interface generation efficiency and enhancing the user's interactive experience.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a method for generating a vehicle interaction interface based on multi-agent collaboration according to Embodiment 1 of the present invention;
[0024] Figure 2 This is a schematic diagram illustrating the process from original code output to online release according to Embodiment 1 of the present invention;
[0025] Figure 3 This is a schematic diagram illustrating the process from code generation to online release of the technical solution provided in Embodiment 1 of the present invention;
[0026] Figure 4 This is a flowchart of a method for generating a vehicle interaction interface based on multi-agent collaboration according to Embodiment 2 of the present invention;
[0027] Figure 5 This is a schematic diagram of a three-level matching strategy between interface requests and various interface templates provided in Embodiment 2 of the present invention;
[0028] Figure 6 This is a schematic diagram of a vehicle interaction interface generation device based on multi-agent collaboration according to Embodiment 3 of the present invention;
[0029] Figure 7This is a schematic diagram of the structure of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," "initial," and "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1 This is a flowchart of a vehicle interaction interface generation method based on multi-agent collaboration according to Embodiment 1 of the present invention. This embodiment is applicable to the real-time generation of human-machine interaction interfaces in vehicles. The method can be executed by a vehicle interaction interface generation device based on multi-agent collaboration. This device can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:
[0034] S110: Acquire user interaction information and vehicle multimodal data.
[0035] In this embodiment, the user can be any user currently in the vehicle, specifically a user interacting with the vehicle. Interaction information can be user actions or requests actively or passively transmitted to the system via the in-vehicle interactive terminal. In this embodiment, the interaction information can include user voice interaction data and user profile data. Specifically, the voice interaction data can be collected by the vehicle's voice acquisition device and processed through speech recognition to obtain the user's natural language commands. For example, voice interaction data could include commands such as "increase the lyrics," "maximum font size," "change to a tech theme," and "move the list to the right." User profile data can refer to data such as the user's age, driving habits, and personal preferences. In this embodiment, the user profile data can be obtained through comprehensive analysis of the user's personal profile and historical interaction logs with the vehicle. Furthermore, the interaction information in this embodiment can also include touch interaction operations, gesture interaction information, and facial expression interaction information, all of which can be collected by the vehicle's acquisition device. Multimodal data can refer to the vehicle's real-time status data, driving environment data, and device status data. The vehicle's real-time status can include data such as vehicle speed, gear position, and remaining battery power. Driving environment data can refer to environmental data during vehicle operation. For example, driving environment data may include road conditions, time, location, and weather data. Device status data can refer to the status data of in-vehicle devices. For example, device status data may include data such as the current screen brightness, volume level, and seat position of the vehicle's user interface.
[0036] In this embodiment, the vehicle terminal can control various data acquisition devices in the vehicle to obtain voice interaction data transmitted by the user to the in-vehicle interactive terminal, as well as real-time vehicle status data, driving environment data, and device status data.
[0037] S120. Perform intent recognition on the interactive information and multimodal data to obtain a set of intent commands.
[0038] Intent recognition can be an operation that involves deep comprehensive analysis and identification of interactive information and multimodal data to obtain the corresponding intent. The intent instruction set can be a structured instruction set. In this embodiment, the intent instruction set can include various intent instructions corresponding to user goals, scene constraints, and preference information. In this embodiment, a large language model or intent recognition model in the vehicle domain can be used to perform intent recognition operations on interactive information and multimodal data, thereby obtaining multiple intent instruction information, which can then be integrated to obtain the intent instruction set.
[0039] In this embodiment, optionally, intention recognition is performed on the interaction information and multimodal data to obtain a set of intention instructions, including: inputting the interaction information and multimodal data into an intention recognition model to obtain multiple explicit intention instructions and multiple implicit intention instructions; and constructing an intention instruction set based on the multiple explicit intention instructions and multiple implicit intention instructions.
[0040] The intent recognition model can be a pre-trained vehicle-domain related model. In this embodiment, the intent recognition model can be trained based on relevant historical vehicle data. Explicit intent commands refer to operational requirements that the user directly expresses through interaction and needs to be explicitly executed. For example, explicit intent commands could be commands such as "enlarge the lyrics" or "play a children's song." Implicit intent commands can be commands derived from contextual data reasoning based on user interaction information and vehicle multimodal data, including potential needs, scenario constraints, and user preference information. For example, the scenario constraint in an implicit intent command could correspond to the command "When driving, more convenient operation is needed," and the user preference information could correspond to the command "Default large font display, adapted to children's visual needs."
[0041] In this embodiment, interactive information and multimodal data can be input into the intent recognition model. The intent recognition model performs deep fusion analysis on the input interactive information and multimodal data to identify multiple explicit intent commands of the user. It can also infer multiple implicit intent commands by combining context. Finally, the multiple explicit intent commands and multiple implicit intent commands are integrated and processed to obtain an intent command set consisting of commands such as user operation goals, scene constraints and preference information.
[0042] In this embodiment, the intent recognition model in the vehicle domain can be used to perform contextual intent recognition on the input data, which can comprehensively and in real time capture and understand user intent and contextual environment, thereby improving the accuracy and reliability of intent recognition.
[0043] S130. Determine the initial interactive interface of the vehicle based on the set of intent commands and the interface generation model.
[0044] The interface generation model is composed of multiple agents. In this embodiment, the interface generation model may include layout agents, functional agents, interaction agents, and component agents. The initial interactive interface refers to the interface scheme of the vehicle human-machine interface output by the interface generation model. In this embodiment, a complete interactive interface scheme can be composed of interface layout information, interface function information, interface interaction information, and target components. The interface scheme in this embodiment can be described by a structured description language, such as DSL or JSON. In this embodiment, the intent instruction set can be decomposed according to the multiple agents included in the interface generation model to obtain the interface scheme of the vehicle human-machine interface.
[0045] Specifically, the interface generation model in this embodiment adopts a multi-agent collaborative architecture. The set of intent commands is input into the interface generation model, and the interface generation model decomposes this complex HMI generation task into multiple specialized agents for parallel or serial processing. That is, the layout agent, function agent, interaction agent and component agent can be processed in parallel or serially respectively to obtain the interface scheme of the vehicle human-machine interaction interface.
[0046] In this embodiment, the complex task of generating a vehicle interaction interface (HMI) can be broken down into sub-tasks corresponding to multiple intelligent agents, such as layout, function, interaction, and components, for separate processing. Furthermore, a model pool of different sizes and capabilities can be configured for each intelligent agent. This embodiment can dynamically schedule the most suitable model for execution based on the complexity of the sub-task. For example, a lightweight 7B small model can be used for simple data splitting tasks to ensure speed; while a more powerful large model is used for complex intent understanding and layout planning tasks to ensure quality, thus guaranteeing model performance and reducing model costs.
[0047] Furthermore, in this embodiment, after obtaining the initial interactive interface, a test agent can be used to perform joint debugging tests on the initial interactive interface. Specifically, in this embodiment, the interface scheme of the vehicle human-machine interaction interface output by the interface generation model can be used, and then the interface scheme can be input into the test agent, so that the test agent can automatically perform joint debugging tests to obtain a successful interface scheme, and then continue the subsequent rendering processing operations on the successful interface scheme. If a failed interface scheme is obtained, it needs to be regenerated and the reason for the failure needs to be reported so that the interface generation model can make targeted optimizations and adjustments.
[0048] S140. The initial interactive interface is rendered using streaming rendering technology to obtain the target interactive interface of the vehicle.
[0049] Streaming rendering technology is a real-time graphics rendering technique that transmits, decodes, and renders simultaneously. Rendering processing can be the process of transforming abstract interface logic, layout rules, component styles, and interaction logic into a visually perceptible graphical interface for the user. The target interactive interface can be the visual interface obtained after rendering the initial interactive interface. In this embodiment, the target interactive interface can be the human-machine interface that is ultimately presented on the in-vehicle terminal screen.
[0050] The solution in this embodiment can be executed by the vehicle interface generation system. Users can adjust the interface in real time through natural language, and the interface response time can be shortened to within 10 seconds, achieving a "what you say is what you get" technical effect. This adapts to the personalized needs of different users and scenarios, and also reduces the risks of manual operation while driving. Moreover, users no longer need to passively accept unified updates, but can enjoy an instant and safe personalized interactive experience, greatly improving user satisfaction.
[0051] This embodiment employs advanced streaming rendering technology to render the vehicle's initial interactive interface, thereby obtaining the final human-machine interface. Specifically, in this embodiment, when multiple agents in the interface generation model generate corresponding initial interactive interface (HMI) schemes, the streaming rendering technology in the rendering engine can immediately begin parsing and drawing the basic framework (such as layout and placeholders), without waiting for all content to be generated. Subsequently, the styles and data of the interface components are dynamically injected and updated onto the screen in a "stream" format, thus rendering all the interface scheme content generated by multiple agents and obtaining the final human-machine interface presented on the in-vehicle terminal screen. This embodiment, through this setup, provides "second-level" initial visual feedback and a "smooth" final presentation effect, greatly enhancing the immediacy of the interaction.
[0052] Furthermore, in this embodiment, the interface generation system ultimately outputs not only the visual interface but also front-end code that can run directly on the vehicle operating system, achieving a seamless transition from "design" to "implementation." For example, a schematic diagram illustrating the process from initial code output to online deployment in this embodiment is shown below. Figure 2 As shown. Figure 2As shown, producers (e.g., developers) generate development requirements for corresponding pages (Page 1 - Customer 1 to Page N - Customer N) based on input from different customers regarding their needs for the vehicle's user interface. Then, for different functions and different customer requirements, they need to write code N times repeatedly, a process that requires manual intervention. After completing the code, they then perform N rounds of integration, testing, and debugging, again requiring manual repetition. After integration and testing, the system proceeds to final release, and user feedback is then fed back to the producers. It is evident that the current system, from initial code production to deployment, relies entirely on repetitive manual operations, resulting in extremely low efficiency and a high risk of resource waste due to "manual overload" caused by repetitive work.
[0053] The technical solution in this embodiment illustrates the process from code generation to online deployment as shown in the diagram below. Figure 3 As shown, firstly, the producer determines the user's intent based on the interaction information of different customers, thus identifying the user's task for the page requirement (from page 1 to customer 1 to page N to customer N). Then, a generative AI Agent (interface generation model) automates the process. This agent comprises four agents—layout generation, UI generation, interaction strategy, and function aggregation—which process in parallel or sequentially to automatically generate the layout, UI, and interaction strategy, and automatically complete function aggregation, replacing the original manual repetitive coding operations. Next, a "testing agent" automatically completes the integration and testing stages, replacing manual repetitive testing. Finally, after testing, the application is directly released. Simultaneously, by combining user interaction logs and through "private model training," the capabilities of the generative agent are continuously optimized, forming a "feedback-optimization" closed loop.
[0054] Understandable. Figure 2 This shows the current process from code production to deployment, while Figure 3 This diagram illustrates the process from code generation to online deployment using the technical solution described in this embodiment. From Figures 2 to 3 This demonstrates an optimization and upgrade from an "inefficient manual code deployment process" to an "AI-assisted, highly automated process." The technical solution in this embodiment uses AI tools to replace N repetitive manual operations, significantly improving the efficiency from code production to deployment and reducing resource waste.
[0055] In this embodiment, streaming rendering technology can efficiently and in real-time present the HMI solution output by the interface generation model on the in-vehicle terminal screen. Furthermore, during the rendering process, this embodiment can dynamically transmit the interface content generated by the corresponding Agent according to user needs or scene constraints, and simultaneously complete decoding and rendering during data transmission, ultimately achieving low latency and highly smooth visualization effects.
[0056] The technical solution of this invention involves acquiring user interaction information and vehicle multimodal data; performing intent recognition on the interaction information and multimodal data to obtain an intent command set; determining the vehicle's initial interactive interface based on the intent command set and an interface generation model; wherein the interface generation model is composed of multiple agents; and rendering the initial interactive interface using streaming rendering technology to obtain the vehicle's target interactive interface. This technical solution addresses the problems of low development efficiency, weak personalization capabilities, and long iteration cycles in vehicle interactive interfaces. It can dynamically adjust or generate personalized interactive interfaces that meet user needs based on user interaction information, improving interface generation efficiency and enhancing the user's interactive experience.
[0057] Example 2
[0058] Figure 4 This is a flowchart of a vehicle interaction interface generation method based on multi-agent collaboration according to Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, the optimization includes: the interface generation model includes a layout agent, a functional agent, an interaction agent, and a component agent; the initial interaction interface is composed of interface layout information, interface function information, interface interaction information, and target components; the initial interaction interface of the vehicle is determined according to the intent command set and the interface generation model, including: determining interface layout information and interface function information based on the intent command set, the layout agent, and the functional agent respectively; determining interface interaction information based on the interface layout information, the interface function information, and the interaction agent; and determining target components based on the interface layout information, the interface function information, the interface interaction information, and the component agent. Figure 4 As shown, the method includes:
[0059] S410: Acquire user interaction information and vehicle multimodal data.
[0060] S420. Perform intent recognition on the interactive information and multimodal data to obtain a set of intent commands.
[0061] In this embodiment, the interface generation model consists of multiple intelligent agents; the interface generation model includes a layout intelligent agent, a functional intelligent agent, an interaction intelligent agent, and a component intelligent agent. The layout intelligent agent can determine the layout information of the interactive interface based on a set of intent commands. The functional intelligent agent can aggregate or invoke the functions of the interactive interface based on a set of intent commands. The interaction intelligent agent can design the interaction logic strategy of the interactive interface, and can be determined based on the results generated by the layout intelligent agent and the functional intelligent agent. The component intelligent agent can determine the intelligent selection and binding processing of each UI component of the interactive interface; in this embodiment, the specific UI elements can be determined based on the output results of the layout intelligent agent, the functional intelligent agent, and the interaction intelligent agent.
[0062] S430 determines the interface layout information and interface function information based on the intent instruction set, the layout agent, and the function agent, respectively.
[0063] The interface layout information can include the spatial location, size proportion, hierarchical relationship, and layout rules of all functional modules and UI components within the vehicle's interactive interface. Interface function information can refer to the required functional modules within the vehicle's interactive interface, the logical relationships between modules, and the conditions for triggering functions.
[0064] In this embodiment, the set of intent commands can be input into the layout agent, function agent, and interaction agent respectively for intelligent analysis and processing, thereby enabling intelligent decision-making on the layout information, functional modules and logical relationships between modules, and interaction logic information within the corresponding vehicle interaction interface.
[0065] In this embodiment, optionally, determining the interface layout information and interface function information based on the intent instruction set, the layout agent, and the function agent respectively includes: inputting the intent instruction set into the layout agent, analyzing and processing each intent instruction in the intent instruction set through the layout agent, and generating corresponding interface layout information in combination with set layout constraints; inputting the intent instruction set into the function agent, analyzing and processing each intent instruction in the intent instruction set through the function agent, determining each functional module and the logical relationship between each functional module, and using each functional module and the logical relationship between each functional module as interface function information.
[0066] Setting layout constraints here refers to the pre-defined layout constraints for each vehicle based on the limitations of the in-vehicle display interface. In this embodiment, the pre-defined layout constraints in the layout agent can be obtained, and then the layout agent can analyze and process the intent command set to determine the spatial position, size ratio, hierarchical relationship, and layout rules of all functional modules and UI components within the vehicle's interactive interface. It is understandable that the pre-defined layout constraints in the layout agent can be determined based on different vehicle screen information and safety specifications. Different vehicles have different screen size information and layout specification constraints, which can be set according to the actual needs of the vehicle.
[0067] Specifically, in this embodiment, the set of intent commands can be input into the layout agent. Upon receiving the set of intent commands, the layout agent intelligently analyzes and processes each command to extract information such as functional area requirements, scenario constraints, and style preferences. Combined with screen size and safety regulations (e.g., driving distraction rules), the layout agent intelligently determines the overall layout information architecture of the interface. In this embodiment, the overall layout information architecture clearly defines the division, arrangement, and adaptation rules of functional areas, thus determining the spatial location, size proportion, hierarchical relationship, and layout rules of all functional modules and UI components within the vehicle's interactive interface. For example, the layout agent will determine the appropriate layout information for the "listening to nursery rhymes while driving" scenario. This could be a layout with large buttons and fewer layers, placing the core functional components (play, loop, and next track) in the optimal area; the optimal area can be an area favored by the user.
[0068] In this embodiment, the set of intent commands can be input into a functional agent. The functional agent analyzes each intent command in the set to determine the corresponding functional module. Then, it dynamically aggregates the required functional modules from the system function library and determines the logical relationships and function triggering conditions between the modules. Thus, the functional modules and their logical relationships can be used as interface function information. The logical relationships between functional modules can be determined by a fixed functional logic relationship and the user's intent commands. In this embodiment, the functional agent is responsible for aggregating and invoking interface functions. For example, in this embodiment, the functional agent analyzes and processes the set of intent commands and dynamically aggregates the required functional modules from the system function library. For instance, it dynamically aggregates music player and children's mode functional modules from the system function library and determines the logical relationship between them.
[0069] S440. Determine the interface interaction information based on the interface layout information, interface function information, and interactive agent.
[0070] The interface interaction information refers to information regarding the interaction methods, operation processes, feedback mechanisms, and exception handling between the user and interface elements. In this embodiment, the interactive agent can be processed after the layout agent and the functional agent. Typically, the interactive agent intervenes after the interface layout and functional module information are initially determined. The interactive agent can define detailed interaction rules based on the areas and element types defined in the layout, and simultaneously adapt the interaction methods in detail by combining the interface function information generated by the functional agent. Ultimately, it generates interface interaction information that matches both the layout and function. The interaction rules can refer to the setting of interactive areas and the determination of interaction methods. It is understood that the interaction rules in this embodiment need to comply with driving safety regulations.
[0071] In this embodiment, interface layout information and interface function information can be input into the interactive agent. The interactive agent analyzes and processes the interface layout information and interface function information to determine the interaction methods, operation processes, feedback mechanisms, and exception handling information between the user and interface elements, thereby obtaining the corresponding interface interaction information. Specifically, in this embodiment, the interactive agent can design the interface interaction logic based on the interface layout information and interface function information. This can include defining the interaction methods between the user and interface elements, which may include interaction methods such as clicks, voice, or gestures; and configuring the system's feedback mechanism, which may include feedback methods such as visual animations and voice broadcasts. For example, the interactive agent can set that when the "loop button" is clicked in a certain music listening area, the button should have a significant visual change, accompanied by a voice prompt "Single loop started".
[0072] S450: Determine the target component based on the interface layout information, interface function information, interface interaction information, and component agent.
[0073] A target component can refer to the UI element corresponding to each functional point. In this embodiment, there can be multiple target components, and the specific number of component elements can be determined by the interface layout information, interface function information, and interface interaction information. In this embodiment, the component agent can determine the matching UI element for each functional point based on the interface layout information, interface function information, and interface interaction information obtained from the layout agent, function agent, and interaction agent, thus obtaining each target component.
[0074] Specifically, in this embodiment, interface layout information, interface function information, and interface interaction information can be input into the component agent. The component agent synchronously parses the area size and position constraints contained in the interface layout information, the business function types and priorities contained in the interface function information, and the interaction methods and interaction rules contained in the interface interaction information. Then, it can filter suitable candidate components from the component library according to the function type, and eliminate mismatched components by combining the physical constraints of the layout and the rules of interaction, determine the optimal target component, and finally configure parameters such as size, style, and interaction trigger threshold for the target component to generate the final target components.
[0075] In this embodiment, the component Agent can also be bound to the vehicle's brand Token, which can be set according to the actual needs of each vehicle. This embodiment allows for a custom component library, pre-defined by conforming to brand design specifications. Understandably, the custom component library in this embodiment may contain some basic components, and can also generate new components in real time based on the decision information output by the layout Agent and interaction Agent. Then, it is stored or updated based on user feedback on the new components. Therefore, the components in the custom component library will continuously be added or updated earlier, further improving the visual effect and loading speed of the components.
[0076] In this embodiment, a highly reusable UI component library that encapsulates complex business logic and visual effects can be developed as a custom component library for the component agent. In this embodiment, the component agent is responsible for the intelligent selection and binding of UI components. By accessing a predefined custom component library, it selects the most suitable UI component for each functional point based on the interface layout of the layout agent and the interaction agent, as well as the decision information of the interface interaction. For example, a large circular button with an icon, injected with specific data such as song title or cover image.
[0077] It should be noted that in this embodiment, during the generation of a new vehicle interface, the functional agent can run in parallel with the layout agent. This is because the functional agent parses the business capability requirements involved in the user's intent and maps them to callable functional modules in the vehicle system. Although it does not directly rely on the layout details generated by the layout agent, it ultimately determines the placement area of each function. Therefore, in this embodiment, the interface function information output by the functional agent needs to be aligned with the layout structure and other layout details included in the interface layout information generated by the layout agent before it can be delivered to the component agent. If, after the vehicle interface has been generated, only some interface content in the generated interface needs to be modified according to user requirements, one or more intelligent agents can be determined from the interface generation model based on the modified content for corresponding processing.
[0078] In this embodiment, abstract layout rules can be transformed into a visual interface architecture based on interface layout information, thereby determining information such as module position, size, and hierarchy. Then, the target components are bound to the partitions of the interface architecture one by one. By attaching the functional modules and logical relationships in the interface function information to the bound components, and by embedding the operation methods, feedback mechanisms, and exception handling rules in the interface interaction information into the corresponding components, the interaction logic of how the user operates and how the interface responds is clarified. Finally, the generated interface content is checked for consistency to ensure that there are no conflicts between layout, function, and interaction, and the initial interactive interface of the vehicle is obtained.
[0079] In this embodiment, optionally, after determining the initial interactive interface of the vehicle, the method further includes: using the initial interactive interface as an interface template and storing the interface template in an interface repository.
[0080] The interface template can be a reusable initial interactive interface scheme. The interface repository can be a resource database for centralized storage and management of interface templates. In this embodiment, the interface repository can contain a UI repository of various components in the interface. It is understood that in this embodiment, the custom component library can be stored in the interface repository. In this embodiment, after each initial interactive interface is generated by processing user interaction information and vehicle multimodal information, the initial interactive interface can be used as an interface template and stored in the interface repository.
[0081] Furthermore, in this embodiment, each interface template in the interface repository can also contain customized Token information, thereby producing interface information with a unified or customized style, thus enabling the display of different brand logos. Specifically, this embodiment can also design a Token system, abstracting the brand's design specifications (such as colors, font sizes, rounded corners, spacing, and animation curves) into a set of "Tokens" that can be read and applied by machines, ensuring that any interface generated through the interface generation model naturally conforms to the brand's DNA. This embodiment can also construct a UI repository cold start, which can pre-generate a large number of UI templates covering mainstream scenarios based on Tokens and a custom component library, building an initial UI repository.
[0082] In this embodiment, the generated interface schemes can be stored as interface templates in the interface repository for subsequent matching and reuse, further improving the efficiency of interface generation.
[0083] In this embodiment, optionally, the method further includes: determining the corresponding interface request based on the intent instruction set; performing similarity matching between the interface request and various interface templates in the interface repository to obtain a matching result; and determining the vehicle's initial interactive interface based on the matching result.
[0084] The interface request can be generated based on a set of intent instructions and can be used to match templates in the interface repository. In this embodiment, the interface request may include multi-dimensional core feature data. For example, the interface request may include scene features, functional features, device features, and interaction features. Similarity matching can be an operation that determines the similarity between the multi-dimensional core feature data contained in the interface request and the feature data contained in each interface template in the interface repository, and compares the similarity with a pre-set threshold. The matching result can be a matching level obtained through similarity matching. In this embodiment, the matching result may include a first-level matching result, a second-level matching result, and a third-level matching result.
[0085] In this embodiment, the structured intent command set can be transformed into a standardized request format that is machine-recognizable and can be used for interface template matching. Specifically, the input intent command set can be decomposed and standardized to extract core feature data that is strongly correlated with interface template matching, and redundant data can be removed. Then, the corresponding interface request can be generated according to a pre-defined fixed request protocol structure. In this embodiment, the core feature data contained in the interface request can be matched with the core feature data in each interface template contained in the interface repository, and the similarity can be compared with a set threshold to determine the corresponding matching level result. Based on the determined different matching level results, the corresponding interface generation strategy can be determined, and then the initial interaction interface of the vehicle can be determined according to different interface generation strategies.
[0086] For example, a schematic diagram of the three-level matching strategy between interface requests and various interface templates in this embodiment is shown below. Figure 5 As shown. The three-level matching strategy in this embodiment is applied to the runtime delivery phase, which is the stage where personalized UIs are dynamically delivered based on user requests. The three-level caching matching strategy improves delivery efficiency, such as... Figure 5 As shown, the specific process can be as follows: First, determine the user's corresponding interface request. Then, upon receiving the user request, perform feature extraction processing on the interface request to obtain the corresponding request features. Next, perform similarity matching between the request features and various interface templates in the interface repository according to a three-level matching strategy. When there is a perfect match, first-level matching is triggered, and existing template content can be directly recalled and returned as the initial interactive interface scheme. When there is a similar match, second-level matching is triggered, and through partial recall processing and partial dynamic adjustment processing, the interface scheme of the initial interactive interface is determined by combining existing resources through the "dynamic splicing" capability. When there is no match, third-level matching is triggered, and all Agents included in the current interface generation model are used to generate a completely new interface scheme in real time, which serves as the initial interactive interface. In this embodiment, the templates generated by second-level and third-level matching can also be fed back to the resource library. This embodiment converts the intent instruction set into the corresponding interface request and matches it with various interface templates in the interface repository according to a three-level matching strategy, ultimately outputting a personalized UI interface scheme to complete the response to the user request. This ensures 100% coverage while maximizing the reuse of existing results, balancing efficiency and cost.
[0087] Specifically, in this embodiment, the core feature data contained in the interface request can be compared with the core feature data in any interface template contained in the interface repository to calculate the similarity score. Then, the similarity score is compared with a set threshold. For example, the set threshold may include a first threshold and a second threshold. If the similarity score is greater than or equal to the first threshold, the matching result between the interface request and the corresponding interface template is a first-level matching result; if the similarity score is greater than or equal to the second threshold and less than the first threshold, the matching result between the interface request and the corresponding interface template is a second-level matching result; if the similarity score is less than the second threshold, the matching result between the interface request and the corresponding interface template is a third-level matching result. It can be understood that the first-level matching result is a complete match; the second-level matching result is a highly similar match; and the third-level matching result is a complete mismatch.
[0088] In this embodiment, by converting the intent command set into an interface request and matching it with various interface templates stored in the interface repository, different interface generation strategies are determined based on the matching results. This ensures that the system can maintain a high response speed and controllable computing cost when dealing with complex and personalized needs, achieving a balance between high efficiency and low cost.
[0089] In this embodiment, optionally, the matching results include first-level matching results, second-level matching results, and third-level matching results; correspondingly, determining the vehicle's initial interaction interface based on the matching results includes: if the matching result is a first-level matching result, using the matched first interface template as the vehicle's initial interaction interface; if the matching result is a second-level matching result, determining the target Agent in the interface generation model based on the matched second interface template and the interface request, and determining the vehicle's initial interaction interface based on the target Agent; if the matching result is a third-level matching result, determining the vehicle's initial interaction interface based on all Agents included in the interface generation model.
[0090] In this embodiment, the first interface template can refer to the interface template corresponding to the first-level matching result. The second interface template can refer to the interface template corresponding to the second-level matching result. The target agent can refer to the agents required after matching the second interface template and the interface request. In this embodiment, the target agent can be one, two, or even three agents. All agents can refer to all agents included in the interface generation model. For example, all agents can include layout agents, function agents, interaction agents, and component agents. In addition, in this embodiment, a style management agent or other agents can be set according to actual needs.
[0091] In this embodiment, if the matching result is a Level 1 matching result, it means that the current interface request and the matched first interface template are a perfect match. Therefore, the matched first interface template can be directly used as the vehicle's initial interactive interface. If the matching result is a Level 2 matching result, it means that the current interface request and the matched first interface template are highly similar. Therefore, based on the differences between the matched second interface template and the interface request, one or two corresponding Agents can be called as target Agents to process the differences, thereby obtaining the vehicle's initial interactive interface. If the matching result is a Level 3 matching result, it means that the current interface request and the matched first interface template are completely mismatched. In this case, the vehicle's initial interactive interface can be regenerated based on all Agents included in the interface generation model, that is, the initial interactive interface can be generated in real time through layout Agent, function Agent, interaction Agent, and component Agent.
[0092] In this embodiment, when determining the corresponding interface request based on the intent instruction set, the system first searches the UI repository for an interface template that perfectly matches the current intent and core data features. If a perfect match is found, the template can be directly recalled and rendered, achieving a millisecond-level response. If no perfect match is found, the system searches for a highly similar matching template. For example, if the user interface request differs from template A only in color, the system only needs to call a simple style management agent to dynamically adjust the template's color and other token attributes, quickly adapting and delivering the template, achieving a second-level response. For entirely new interface requirements with no reusable templates, a complete multi-agent collaborative generation process is initiated to generate the initial interactive interface in real time, ensuring 100% requirement coverage.
[0093] In this embodiment, all newly generated or adjusted UI components in the secondary and tertiary matching results will be automatically returned to the UI repository as new interface templates. The more interface scenarios the interface generation system processes, the richer the repository will become, and the higher the matching efficiency will be in the future, forming a self-improving "flywheel effect".
[0094] In this embodiment, the three-level matching results can correspond to different interface generation strategies, thereby ensuring the real-time generation of the interactive interface and adapting to the personalized needs of different users and scenarios.
[0095] S460 uses streaming rendering technology to render the initial interactive interface to obtain the target interactive interface of the vehicle.
[0096] In this embodiment, optionally, it also includes: collecting user feedback information and interface interaction logs; and optimizing the interface generation model based on the feedback information and interface interaction logs.
[0097] Feedback information can refer to user evaluations of the generated target interactive interface. In this embodiment, feedback information can include active and passive feedback. For example, active feedback can be provided through in-interface quick ratings or function satisfaction questionnaires. Passive feedback can refer to feedback associated with the user's cancellation of an operation. The interface interaction log refers to the user's operational behavior and interface running status automatically recorded by the vehicle terminal, reflecting the actual usage of the interactive interface.
[0098] In this embodiment, user interaction logs, active feedback information, or passive feedback information can be automatically collected. The interface generation model can be optimized based on these logs and feedback information to obtain an updated interface generation model.
[0099] Furthermore, this embodiment can automatically label the interaction data to form high-quality training data, which can be used for incremental training and fine-tuning of the privately deployed intent recognition model and interface generation model. This embodiment can also collect Bad Cases that the system cannot process. Automatic or semi-automatic attribution analysis is performed on Bad Cases, and the analysis results (such as an intent recognition error or an unreasonable layout) are stored in the HMI knowledge base. The content of the knowledge base can also dynamically optimize the prompts for each agent, guiding the agent to avoid similar errors in the future.
[0100] Furthermore, in this embodiment, the optimized model can be seamlessly deployed online through a hot update mechanism, enabling the system to continuously evolve without interrupting service and become more intelligent with use.
[0101] This embodiment, through such a setup, constructs the model's self-learning and continuous optimization capabilities, forming a positive closed loop of "generation-feedback-evolution," which enables it to continuously learn and optimize from real-world usage data, thereby improving the accuracy and reliability of the interface-generated model.
[0102] The technical solution of this invention involves acquiring user interaction information and vehicle multimodal data; performing intent recognition on the interaction information and multimodal data to obtain an intent command set; determining interface layout information, interface function information, and interface interaction information based on the intent command set, combined with a layout agent, a function agent, and an interaction agent; determining a target component based on the interface layout information, interface interaction information, and a component agent; determining the vehicle's initial interactive interface based on the interface layout information, interface function information, interface interaction information, and the target component; and rendering the initial interactive interface using streaming rendering technology to obtain the vehicle's target interactive interface. This technical solution addresses the problems of low development efficiency, weak personalization capabilities, and long iteration cycles in vehicle interactive interfaces. It can dynamically adjust or generate personalized interactive interfaces that meet user needs based on user interaction information, improving interface generation efficiency and enhancing the user's interactive experience. In this embodiment, user intent is understood through multimodal data perception. Then, multiple specialized agents work together to complete the layout, component binding, interaction strategies and style injection. Combined with a three-level matching mechanism and streaming output technology, the real-time generation of the interactive interface is guaranteed. It can dynamically generate personalized UI that meets the requirements based on the user's natural language commands and output interface code that can be rendered in real time, realizing the "what you say is what you get" interactive experience.
[0103] Example 3
[0104] Figure 6 This is a schematic diagram of a vehicle interaction interface generation device based on multi-agent collaboration according to Embodiment 3 of the present invention. Figure 6 As shown, the device includes:
[0105] Data acquisition module 610 is used to acquire user interaction information and vehicle multimodal data;
[0106] The intent recognition module 620 is used to recognize the intent from the interaction information and multimodal data to obtain a set of intent commands;
[0107] The interface determination module 630 is used to determine the initial interactive interface of the vehicle based on the set of intent commands and the interface generation model; wherein, the interface generation model is composed of multiple intelligent agents.
[0108] The interface rendering module 640 is used to render the initial interactive interface using streaming rendering technology to obtain the target interactive interface of the vehicle.
[0109] Optionally, the interface generation model includes layout agents, functional agents, interaction agents, and component agents; the initial interactive interface is composed of interface layout information, interface function information, interface interaction information, and target components.
[0110] Interface determination module 630 includes:
[0111] The information determination unit is used to determine interface layout information and interface function information based on the intent instruction set, the layout agent, and the functional agent, respectively.
[0112] The interaction information determination unit is used to determine the interface interaction information based on the interface layout information, interface function information, and interactive intelligent agent.
[0113] The component determination unit is used to determine the target component based on the interface layout information, interface function information, interface interaction information, and component agent.
[0114] Optional, the information determination unit is specifically used for:
[0115] The set of intent commands is input into the layout agent, which analyzes and processes each intent command in the set of intent commands and generates the corresponding interface layout information in combination with the set layout constraints.
[0116] The set of intent commands is input into the functional agent. The functional agent analyzes and processes each intent command in the set of intent commands to determine each functional module and the logical relationship between each functional module. The functional module and the logical relationship between each functional module are used as the interface functional information.
[0117] Optionally, the device further includes a template storage module, used to use the initial interactive interface as an interface template after determining the initial interactive interface of the vehicle, and to store the interface template in the interface repository.
[0118] Optionally, the device may also include:
[0119] The UI request determination module is used to determine the corresponding UI request based on the set of intent instructions.
[0120] The matching module is used to perform similarity matching between the interface request and various interface templates in the interface repository to obtain the matching results.
[0121] The determination module is used to determine the initial interactive interface of the vehicle based on the matching results.
[0122] Optionally, the matching results include first-level matching results, second-level matching results, and third-level matching results;
[0123] Accordingly, the module is defined, specifically for:
[0124] If the matching result is a first-level matching result, the first matching interface template will be used as the vehicle's initial interactive interface.
[0125] If the matching result is a secondary matching result, the target intelligent agent in the interface generation model is determined based on the matched second interface template and interface request, and the initial interactive interface of the vehicle is determined based on the target intelligent agent.
[0126] If the matching result is a level 3 matching result, the initial interaction interface of the vehicle is determined based on all the intelligent agents included in the interface generation model.
[0127] Optional, the intent recognition module 620 is specifically used for:
[0128] The interaction information and multimodal data are input into the intent recognition model to obtain multiple explicit intent commands and multiple implicit intent commands;
[0129] An intent instruction set is constructed based on multiple explicit intent instructions and multiple implicit intent instructions.
[0130] The vehicle interaction interface generation device based on multi-agent collaboration provided in this embodiment of the invention can execute the vehicle interaction interface generation method based on multi-agent collaboration provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0131] Example 4
[0132] Figure 7 This is a schematic diagram of an electronic device according to Embodiment 4 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0133] like Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0134] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a method for generating vehicle interaction interfaces based on multi-agent cooperation.
[0136] In some embodiments, the vehicle interface generation method based on multi-agent cooperation can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the vehicle interface generation method based on multi-agent cooperation described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the vehicle interface generation method based on multi-agent cooperation by any other suitable means (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0142] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for generating a vehicle interaction interface based on multi-agent collaboration, characterized in that, include: Acquire user interaction information and vehicle multimodal data; Intent recognition is performed on the interaction information and the multimodal data to obtain a set of intent commands; The initial interactive interface of the vehicle is determined based on the set of intent commands and the interface generation model; wherein the interface generation model is composed of multiple intelligent agents. The initial interactive interface is rendered using streaming rendering technology to obtain the target interactive interface of the vehicle.
2. The method according to claim 1, characterized in that, The interface generation model includes a layout agent, a functional agent, an interaction agent, and a component agent; the initial interactive interface is composed of interface layout information, interface function information, interface interaction information, and target components. The initial interactive interface of the vehicle is determined based on the set of intent commands and the interface generation model, including: Based on the intent instruction set, the layout agent, and the functional agent, the interface layout information and interface function information are determined respectively. The interface interaction information is determined based on the interface layout information, the interface function information, and the interactive agent. The target component is determined based on the interface layout information, the interface function information, the interface interaction information, and the component agent.
3. The method according to claim 2, characterized in that, Based on the intent instruction set, the layout agent, and the functional agent, the interface layout information and interface function information are determined respectively, including: The set of intent commands is input into the layout agent, which analyzes and processes each intent command in the set of intent commands and generates corresponding interface layout information in combination with the set layout constraints. The set of intent commands is input to the functional agent, which analyzes and processes each intent command in the set of intent commands to determine each functional module and the logical relationship between each functional module, and uses each functional module and the logical relationship between each functional module as interface function information.
4. The method according to claim 1, characterized in that, After determining the vehicle's initial user interface, the following is also included: The initial interactive interface is used as an interface template, and the interface template is stored in the interface repository.
5. The method according to claim 4, characterized in that, Also includes: The corresponding interface request is determined based on the set of intent instructions; The interface request is matched with each interface template in the interface repository to obtain the matching result; The initial user interface of the vehicle is determined based on the matching results.
6. The method according to claim 5, characterized in that, The matching results include first-level matching results, second-level matching results, and third-level matching results; Accordingly, the initial interactive interface of the vehicle is determined based on the matching result, including: If the matching result is a first-level matching result, the matching first interface template will be used as the vehicle's initial interactive interface. If the matching result is a secondary matching result, the target intelligent agent in the interface generation model is determined based on the matched second interface template and the interface request, and the initial interactive interface of the vehicle is determined based on the target intelligent agent. If the matching result is a level 3 matching result, the initial interaction interface of the vehicle is determined based on all the intelligent agents included in the interface generation model.
7. The method according to claim 1, characterized in that, Intent recognition is performed on the interaction information and the multimodal data to obtain a set of intent commands, including: The interaction information and the multimodal data are input into the intent recognition model to obtain multiple explicit intent commands and multiple implicit intent commands; An intent instruction set is constructed based on the multiple explicit intent instructions and the multiple implicit intent instructions.
8. A vehicle interaction interface generation device based on multi-agent collaboration, characterized in that, include: The data acquisition module is used to acquire user interaction information and vehicle multimodal data; The intent recognition module is used to perform intent recognition on the interaction information and the multimodal data to obtain a set of intent commands; The interface determination module is used to determine the initial interactive interface of the vehicle based on the intent instruction set and the interface generation model; wherein the interface generation model is composed of multiple intelligent agents. The interface rendering module is used to render the initial interactive interface using streaming rendering technology to obtain the target interactive interface of the vehicle.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the vehicle interaction interface generation method based on multi-agent cooperation as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the vehicle interaction interface generation method based on multi-agent cooperation as described in any one of claims 1-7.