User demand identification method

By combining visual language models and large language models, and incorporating Maslow's hierarchy of needs theory, a structured list of user needs is generated. This addresses the shortcomings of in-vehicle intelligent cockpit systems in identifying implicit user needs and understanding dynamic scenarios, enabling accurate identification of user needs and personalized services.

CN121808239APending Publication Date: 2026-04-07DONGFENG MOTOR GRP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing in-vehicle intelligent cockpit systems are inadequate in recognizing users' implicit needs and understanding dynamic scenarios. They struggle to effectively integrate multi-source heterogeneous information and perform deep semantic analysis, fail to fully capture key elements in the scenario, and lack deep memory and utilization of users' long-term preferences and historical behaviors, resulting in an inability to provide a truly personalized and continuous service experience.

Method used

By employing a pre-trained visual language model and a large language model working together, structured scene labels are generated from multimodal perception data. Combined with Maslow's hierarchy of needs theory, a user needs list is generated, enabling the identification and prioritization of users' implicit needs.

Benefits of technology

It improves the accuracy of identifying and prioritizing users' explicit and implicit needs, and can proactively identify needs that users do not express directly, providing personalized and proactive services that fit user habits, thereby enhancing user stickiness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808239A_ABST
    Figure CN121808239A_ABST
Patent Text Reader

Abstract

The invention provides a user demand identification method, and belongs to the technical field of intelligent cabins, and the user demand identification method comprises the steps: obtaining a scene description through a pre-trained visual language model according to multi-modal perception data; the visual language model is used for performing semantic understanding according to the multi-modal perception data and outputting scene description in a natural language form; according to the scene description and user historical data, generating a scene label by using a first large language model; the first large language model is used for mapping the scene description into a structured scene label; and according to the scene label and historical data of the user, generating a demand list of the user by using a second large language model combined with a Maslow demand hierarchy theory so as to generate a task list. According to the method, through cooperative work of the visual language model and the large language model, the original multi-modal sensing data is converted into the structured scene label, and then the accurate user demand list is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart cockpit technology, and in particular to a method for identifying user needs. Background Technology

[0002] As intelligent vehicles evolve towards a "scenario-driven service" paradigm, proactive service capabilities in the cockpit and driving domain have become core competitive advantages. However, current in-vehicle intelligent cockpit systems still face many challenges in understanding user needs, particularly in identifying implicit user needs and understanding dynamic scenarios. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the prior art, and proposes a method for identifying user needs that can identify implicit user needs.

[0004] In a first aspect, embodiments of the present invention provide a method for identifying user needs, comprising: obtaining a scene description using a pre-trained visual language model based on multimodal perception data; the multimodal perception data including at least image data, audio data, vehicle status data, and external environment data; the visual language model being used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form; generating scene tags using a first language model based on the scene description and user historical data; the first language model being used to map the scene description into structured scene tags; and generating a user's need list for generating a task list using a second language model incorporating Maslow's hierarchy of needs theory based on the scene tags and user historical data.

[0005] In an embodiment of the present invention, scene tags are generated using a first language model based on scene description and user history data. This includes: combining user history data and using prompting engineering techniques to guide the first language model to extract and populate predefined structured scene tags from the scene description; wherein, user history data includes at least short-term memory data and long-term memory data; short-term memory data includes at least one of the following: current scene description, historical scene tags, identified demand list, currently executing task status, and recent user interaction history; long-term memory data includes at least one of the following: user's personality preferences, driving habits, historical commuting routes, family member information, and frequently used services; structured scene tags include at least one of the following: trip stage tags, vehicle usage purpose tags, behavior tags, and vehicle status tags; vehicle status tags include at least one of the following: personnel status tags, object status tags, vehicle environment status tags, and vehicle status tags.

[0006] In embodiments of the present invention, the trip stage label includes at least one of pre-boarding, during driving, parking, charging or refueling, maintenance, and emergency; the vehicle use purpose label includes at least one of commuting, travel, shopping, business reception, picking up family members, logistics transportation, and outdoor activities; the behavior label includes at least one of working, entertainment, leisure, socializing, learning, and rest; the personnel status label includes at least one of driver fatigue status, number and distribution of passengers, and status of children and pets; the object status label includes at least one of item placement location and hazardous materials detection; the vehicle environment status label includes at least one of temperature and humidity, air quality, weather, and road congestion; and the vehicle status label includes at least one of vehicle speed, fuel consumption, battery level, and malfunction indicator light status.

[0007] In an embodiment of the present invention, a user's needs list is generated based on scene tags and user history data using a second language model that incorporates Maslow's hierarchy of needs theory. This includes: configuring prompts and constraints of the second language model according to Maslow's hierarchy of needs theory; and generating the user's needs list based on scene tags and user history data using the second language model.

[0008] In an embodiment of the present invention, according to Maslow's hierarchy of needs, the prompts and constraints of the second language model are configured, including: configuring the second language model as follows: for each scene tag, combined with user historical data, and according to Maslow's hierarchy of needs, the user's physiological needs, safety needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs are identified layer by layer; wherein, the priority of user needs from high to low is safety needs, physiological needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs.

[0009] In an embodiment of the present invention, generating a task list includes: generating corresponding physiological tasks, safety tasks, love and belonging tasks, respect tasks, cognitive tasks, and aesthetic tasks based on physiological needs, safety needs, love and belonging needs, respect tasks, cognitive tasks, and aesthetic needs; and executing the tasks in the task list sequentially according to the priority of the user's needs.

[0010] In embodiments of the present invention, image data is acquired through an in-vehicle camera; audio data is acquired through an in-vehicle microphone array; vehicle status data is acquired through a vehicle CAN bus interface and environmental sensors; and external environmental data is acquired through an external data interface, including at least user schedule and health data.

[0011] A second aspect of the present invention provides a user requirement identification system, which can be used to implement the aforementioned user requirement identification method, comprising: a scene description module, used to obtain a scene description based on multimodal perception data using a pre-trained visual language model; the multimodal perception data includes at least image data, audio data, vehicle status data, and external environment data; the visual language model is used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form; a scene tagging module, used to generate scene tags based on the scene description and user historical data using a first language model; the first language model is used to map the scene description into structured scene tags; and a requirement identification module, used to generate a user requirement list for generating a task list based on the scene tags and user historical data using a second language model incorporating Maslow's hierarchy of needs theory.

[0012] A third aspect of the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the aforementioned user requirement identification method.

[0013] A fourth aspect of the present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the aforementioned user requirement identification method.

[0014] The user demand identification method provided by this invention transforms the original multimodal perception data into structured scene labels through the collaborative work of a visual language model and a large language model. Then, it generates an accurate user demand list by combining Maslow's hierarchy of needs. This method at least partially solves the problem of insufficient understanding of dynamic scenes and the inability to fully capture user needs, and achieves the technical effect of identifying implicit user needs based on dynamic scenes. Attached Figure Description

[0015] Figure 1 A flowchart illustrating a user demand identification method provided in an embodiment of the present invention;

[0016] Figure 2 This is a diagram of an in-vehicle intelligent agent architecture for the method provided in an embodiment of the present invention;

[0017] Figure 3 This is a schematic diagram of one embodiment of the method provided by the present invention;

[0018] Figure 4 A structural block diagram of a user demand identification system provided in an embodiment of the present invention;

[0019] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions of the present invention, exemplary embodiments of the present invention are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0021] Where there is no conflict, the various embodiments of the present invention and the features thereof may be combined with each other.

[0022] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0024] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and the invention, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.

[0025] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information all comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example: appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely locating a specific individual.

[0026] Traditional in-vehicle intelligent cockpit systems (such as mainstream voice assistants or navigation systems) primarily rely on explicit voice commands or touch input from users. Their workflow typically involves the user issuing a command, the system recognizing the intent, executing the corresponding function, and finally providing feedback. For example, if a user says "turn on the air conditioning," the system will turn on the air conditioning. Some systems may combine simple contextual information (such as time and location) for limited scene awareness, such as alerting the user to traffic congestion ahead during navigation.

[0027] Therefore, this solution primarily responds passively to explicitly expressed user needs, failing to proactively identify implicit needs not directly expressed by the user. For example, when a user yawns, the system cannot proactively assess the potential risk of fatigued driving and offer rest suggestions; seeing a user carrying shopping bags into the car on a rainy day, the system struggles to proactively associate it with needs such as "helping the user clean up their items" or "playing soothing music to alleviate fatigue." Its limited ability to understand complex and dynamic in-vehicle scenarios hinders the effective fusion and deep semantic analysis of multi-source heterogeneous information, resulting in an inability to comprehensively capture key elements within the scenario.

[0028] Furthermore, the system lacks a deep memory and utilization of users' long-term preferences and historical behaviors, making it difficult to provide a truly personalized and continuous service experience. It also lacks an understanding of demand priorities, making it unable to effectively prioritize identified demands and make optimal decisions among multiple potential demands.

[0029] In recent years, Large Language Models (LLM) and Visual Language Models (VLM) have made breakthroughs in natural language understanding and cross-modal perception. Some studies have attempted to apply LLM to in-vehicle speech recognition and natural language interaction, and VLM to environmental perception and behavior recognition. However, demand insight still relies on user voice input. For raw, unstructured scenarios, LLM may still suffer from unstable inference results due to ambiguity or missing information, especially in complex or edge-of-the-road scenarios. The accuracy and consistency of LLM output may be affected when there is a lack of explicit, structured guidance. Most existing demand prediction based on LLM relies on the generalized knowledge or simple rules of the large model itself, without systematically integrating psychological theories into the demand reasoning framework. This makes it difficult for the system to identify users' deep, implicit needs, and also makes it impossible to perform truly human-centered prioritization that conforms to human behavioral patterns.

[0030] First, the technical terms involved in this invention will be explained as follows:

[0031] VLM (Visual Language Model): A cross-modal artificial intelligence technology that jointly models visual and language modalities.

[0032] LLM (Large Language Model): A natural language processing model based on deep learning, possessing powerful text generation, understanding, reasoning, and interaction capabilities.

[0033] CL3: Advanced Cognitive Intelligent Cockpit Definition Guide, the standard for classifying intelligent cockpit capability levels in the intelligent vehicle industry.

[0034] Multimodal perception: The system's ability to acquire and fuse different types of data through various sensors (such as cameras, microphones, and environmental sensors).

[0035] Scenario Description: Using VLM, raw multimodal perception data is transformed into natural language text that describes the current environment and user state in detail.

[0036] Scene tags: Structured, multi-dimensional scene feature tags generated by LLM based on scene descriptions, such as trip stage, purpose of use, behavior, etc.

[0037] Implicit needs: Needs that users do not explicitly express but actually exist or have potential, which can only be identified through reasoning and insight.

[0038] Maslow's hierarchy of needs theory divides human needs from low to high into physiological needs, safety needs, love and belonging needs, esteem needs, and self-actualization needs. In this invention, based on the actual needs of the cabin, these needs are expanded to include physiological, safety, love and belonging, esteem, cognitive, and aesthetic needs.

[0039] To address the technical problems of existing in-vehicle intelligent systems in terms of user demand insight, such as "insufficient ability to identify implicit user needs", "lack of depth of scene understanding", "unreasonable priority of needs", and "insufficient and inefficient use of LLM / VLM input", this invention provides a method for identifying user needs. Figure 1 A flowchart illustrating a user demand identification method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the process includes: S1, obtaining a scene description using a pre-trained visual language model based on multimodal perception data; the multimodal perception data includes at least image data, audio data, vehicle status data, and external environment data; the visual language model is used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form; S2, generating scene tags using a first-level language model based on the scene description and user history data; the first-level language model is used to map the scene description into structured scene tags; S3, generating a user's needs list based on the scene tags and user history data using a second-level language model that incorporates Maslow's hierarchy of needs theory to generate a task list.

[0040] Figure 2This is a diagram of an in-vehicle intelligent agent architecture provided by an embodiment of the present invention, such as... Figure 2 As shown, the system mainly consists of a perception and labeling layer, a memory layer, and a reasoning and decision-making layer. The perception and labeling layer is responsible for real-time collection of multimodal information such as the internal and external environment, user behavior, and vehicle status. The memory layer continuously stores perception data, demand insight results, task execution context status, user historical behavior, and long-term preferences, providing personalized and contextual information support for reasoning and decision-making. The reasoning and decision-making layer is the core of the system. This layer receives structured scene labels and memory data, uses a Large Language Model (LLM) as the reasoning engine, and integrates prompts from Maslow's hierarchy of needs theory to perform demand insight. This process transforms complex scene information into a structured demand list and adjusts demand priorities according to Maslow's theory.

[0041] In this embodiment, the VLM-driven scene description generation module inputs multimodal data (images, audio, vehicle status, etc.) into a pre-trained Visual Language Model (VLM). After fine-tuning for the in-vehicle scene, this VLM can perform deep semantic understanding of the fused multimodal information, outputting a precise and detailed scene description in natural language (natural language text, referring to computer-generated sentences or paragraphs that humans can easily understand, used to summarize the current scene). The VLM not only identifies objects and behaviors but also understands their contextual relationships. For example, it recognizes "the driver is driving in the rain, holding their forehead, and the interior temperature is high," rather than simply "there are people" or "it's raining." This high-quality scene description forms the basis for subsequent structured label generation.

[0042] Through embodiments of this invention, a Visual Language Model (VLM) is used to deeply understand raw multimodal data and generate detailed scene descriptions. Then, a Large Language Model (LLM) is used in conjunction with user historical behavior and memory data to transform the scene descriptions into structured scene labels. This invention aims to provide an innovative user needs insight method that integrates multimodal perception, user historical data, LLM and VLM, structured scene labels, and Maslow's hierarchy of needs theory, thereby significantly improving the accuracy of identifying explicit and implicit user needs and the rationality of prioritization.

[0043] Based on the above embodiments, scene tags are generated using a first language model according to scene descriptions and user history data. This includes: combining user history data and using prompting engineering techniques to guide the first language model to extract and populate predefined structured scene tags from the scene descriptions; wherein, user history data includes at least short-term memory data and long-term memory data; short-term memory data includes at least one of the following: current scene description, historical scene tags, identified demand list, currently executing task status, and recent user interaction history; long-term memory data includes at least one of the following: user's personality preferences, driving habits, historical commuting routes, family member information, and frequently used services; structured scene tags include at least one of the following: trip stage tags, vehicle usage purpose tags, behavior tags, and vehicle status tags; vehicle status tags include at least one of the following: personnel status tags, object status tags, vehicle environment status tags, and vehicle status tags.

[0044] In this embodiment, prompting engineering is a method for guiding a large language model to generate output in a specific format or content. Through designed input text (i.e., prompt words), the model can be guided on how to complete a task.

[0045] In this embodiment, the memory module is responsible for storing and updating all key information: short-term memory includes the current scene description, scene tags, identified requirement list, ongoing task status, and recent user interaction history, used to maintain contextual coherence; long-term memory includes the user's personalized preferences (such as seating habits, music preferences, and air conditioning temperature preferences), driving habits, historical commuting routes, family member information, and frequently used services. This information is continuously learned and accumulated through user usage and feedback. The memory module is not simply a database; it provides real-time, personalized contextual support, enabling requirement reasoning and task planning to be adjusted according to the user's unique characteristics. For example, for a user who frequently listens to the news during their commute, the system will prioritize pushing news information.

[0046] In this embodiment of the invention, the scene description generated by the VLM (Visual Language Model) is used as the main input, combined with user historical behavior and long-term preference data stored in the memory layer, and then input into the large language model. This LLM, through prompting engineering techniques, guides the model to extract and fill predefined structured scene tags from the scene description. Leveraging its powerful understanding and inductive capabilities, the LLM maps and quantifies the scene description from free text onto these structured tags, forming a highly generalized contextual information that can be efficiently processed by the machine, improving the accuracy and efficiency of subsequent demand reasoning. For example, from "driver yawning" and "driving at night on the highway," "driver fatigue status: yes," "trip stage: driving," and "road type: highway" can be generated.

[0047] Based on the above embodiments, the trip stage labels include at least one of the following: before boarding, during driving, parking, charging or refueling, maintenance, and emergency; the purpose of use labels include at least one of the following: commuting, travel, shopping, business reception, picking up family members, logistics transportation, and outdoor activities; the behavior labels include at least one of the following: working, entertainment, leisure, socializing, studying, and resting; the personnel status labels include at least one of the following: driver fatigue status, number and distribution of passengers, and status of children and pets; the object status labels include at least one of the following: the placement of items and the detection of hazardous materials; the vehicle environment status labels include at least one of the following: temperature and humidity, air quality, weather, and road congestion; and the vehicle status labels include at least one of the following: vehicle speed, fuel consumption, battery level, and malfunction indicator light status.

[0048] While existing solutions may acquire some real-time context, they fall short in effectively integrating and structuring multi-source contextual information (including scene descriptions generated by VLM, real-time vehicle status, and external environmental data). The lack of an efficient, structured information flow mechanism between the VLM-generated scene description and the LLM input, coupled with insufficient integration of other multi-source real-time data, prevents the system from constructing a comprehensive and accurate real-time scene background. This limits the LLM's understanding and grasp of the current scene's complexity during inference. This lack of information integration means the inference results may not accurately reflect the user's actual immediate needs. Therefore, through embodiments of this invention, a predefined classification system is established, as shown in Table 1, where each label represents a specific aspect of the scene (e.g., time, purpose, behavior). It transforms unstructured natural language descriptions into machine-processable categorized data. The structured scene labels are not simply keyword extraction, but rather a hierarchical and semantically related system as shown in Table 1.

[0049] Table 1

[0050] Based on the above embodiments, according to scene tags and user history data, a user's needs list is generated using the second language model that incorporates Maslow's hierarchy of needs theory. This includes: configuring prompts and constraints of the second language model according to Maslow's hierarchy of needs theory; generating a user's needs list using the second language model based on scene tags and user history data; and the needs list, which includes: a description of the needs (e.g., "the user needs rest," "the user wants to keep the air inside the car fresh"), the Maslow's hierarchy level to which the needs belong (e.g., "physiological needs," "safety needs"), and a precise priority.

[0051] Existing solutions generally lack effective mechanisms for accumulating, storing, and utilizing users' long-term preferences and historical behavioral data (such as driving habits, frequently used routes, and music preferences). This limits LLM's ability to fully consider users' unique personalized characteristics when performing demand inference, resulting in insufficiently customized services and difficulty in building user stickiness. This invention uses users' historical behavior and long-term preference data as core inputs, enabling the demand insight process to fully utilize users' historical behavior and long-term preferences to provide more personalized, proactive services that align with user habits, thereby enhancing user stickiness.

[0052] In the embodiments of this invention, the structured scene tags and user history behavior and long-term preference data provided by the memory layer are used as core inputs to a Large Language Model (LLM) as the inference engine. During the inference process, the LLM not only performs generalization based on context, but more importantly, it incorporates prompts and constraints from Maslow's hierarchy of needs theory, thereby generating a needs list that better aligns with human instincts and emotions.

[0053] Based on the above embodiments, according to Maslow's hierarchy of needs, the prompts and constraints of the second language model are configured, including: configuring the second language model as follows: for each scene tag, combined with user historical data, and according to Maslow's hierarchy of needs, the user's physiological needs, safety needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs are identified layer by layer; among which, the priority of user needs from high to low is safety needs, physiological needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs.

[0054] Through embodiments of this invention, needs are proactively identified and categorized into different levels, such as physiological needs, safety needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs, based on the scenario and user status. By introducing the structured thinking framework of Maslow's hierarchy of needs, LLM is guided to consider deeper, unexpressed but actually existing latent needs of users, enhancing the insight into implicit needs. For example, when a user has not spoken to family members for a long time, the implicit need of "missing loved ones (love and belonging)" may be discerned, and a call may be proactively suggested.

[0055] Based on the above embodiments, a task list is generated, including: generating corresponding physiological tasks, safety tasks, love and belonging tasks, respect tasks, cognitive tasks, and aesthetic tasks according to physiological needs, safety needs, love and belonging needs, respect tasks, cognitive tasks, and aesthetic needs; and executing the tasks in the task list in sequence according to the priority of user needs.

[0056] Through embodiments of the present invention, the identified needs are dynamically prioritized according to Maslow's hierarchy of needs theory. Typically, physiological and safety needs have the highest priority. This avoids the prioritization confusion or irrationality that may occur when relying solely on the generalization ability of LLM.

[0057] Based on the above embodiments, image data is collected through an in-vehicle camera; audio data is collected through an in-vehicle microphone array; vehicle status data is collected through the vehicle CAN bus interface and environmental sensors; and external environmental data is collected through an external data interface, including at least user schedule and health data.

[0058] In this embodiment, the multimodal data acquisition module is deployed in the vehicle environment, integrating various sensors, such as: a high-resolution camera for acquiring image / video data of driver facial expressions (fatigue, emotions), gestures, passenger distribution, in-vehicle item status, external road environment, traffic conditions, and weather (rain, snow, sunlight); a high-sensitivity microphone array for acquiring in-vehicle voice commands, dialogues, and environmental noise; environmental sensors for acquiring air quality data such as in-vehicle temperature, humidity, and PM2.5; a vehicle CAN bus interface for acquiring precise vehicle operating status data such as vehicle speed, mileage, fuel consumption / battery charge, fault codes, steering angle, and pedal opening; and an external data interface for acquiring user schedules (calendar), health data (heart rate, sleep), news subscriptions, frequently used applications, and social media information through interconnection with the user's mobile phone.

[0059] In the embodiments of this invention, multimodal perception data refers to a collection of data from different sources and types. In an in-vehicle environment, this includes multiple data sources such as vision, hearing, the vehicle itself (CAN bus sensors), and external sources (network, calendar), which together constitute a comprehensive perception of the current environment.

[0060] Figure 3 This is a schematic diagram of one embodiment of the method provided by the present invention, such as... Figure 3As shown, when a user walks towards the vehicle with a briefcase in hand, the external camera identifies their walking status, vehicle sensors detect that the doors are locked and the air conditioning is off, and the system simultaneously reads the user's phone schedule (such as important meetings in the morning), news preferences, and health data, and combines this with historical commuting records for multimodal perception. Based on this information, the system generates structured scene labels: the trip stage is "before getting in the car," the purpose of use is "commuting," the vehicle status is "not started," and the user status is "not in the car." Subsequently, through demand reasoning, four core needs are identified: "safety needs: driving safety," "physiological needs: comfortable environment," "cognitive needs: schedule reminders," and "cognitive needs: route planning," and tasks are assigned to the corresponding intelligent agents: the driving task intelligent agent immediately plans the optimal route, assesses real-time road conditions, and recommends the autonomous driving mode; the comfort environment intelligent agent turns on the air conditioning and seat heating in advance; the information service intelligent agent pushes meeting reminders and news briefings on the in-car screen; and the health protection intelligent agent activates heart rate and blood pressure monitoring, forming a closed-loop service from perception and reasoning to execution, and finally continuously optimizing the strategy based on user feedback.

[0061] Figure 4 A structural block diagram of a user demand identification system provided in an embodiment of the present invention is shown below. Figure 4 As shown, the present invention also provides a user demand identification system, which can be used to implement the above-mentioned user demand identification method, including: a scene description module, used to obtain a scene description based on multimodal perception data using a pre-trained visual language model; the multimodal perception data includes at least image data, audio data, vehicle status data, and external environment data; the visual language model is used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form; a scene label module, used to generate scene labels based on the scene description and user historical data using a first language model; the first language model is used to map the scene description into structured scene labels; and a demand identification module, used to generate a user demand list based on the scene labels and user historical data using a second language model combined with Maslow's hierarchy of needs theory for generating a task list.

[0062] Based on the same inventive concept, embodiments of the present invention also provide an electronic device. Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Figure 5 As shown, an embodiment of the present invention provides an electronic device including: one or more processors 101, a memory 102, and one or more I / O interfaces 103. The memory 102 stores one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement any of the user requirement identification methods described in the above embodiments; the one or more I / O interfaces 103 are connected between the processor and the memory, configured to enable information interaction between the processor and the memory.

[0063] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102, and can realize information interaction between the processor 101 and the memory 102, including but not limited to a data bus (Bus).

[0064] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.

[0065] In some embodiments, the one or more processors 101 include a field-programmable gate array.

[0066] This invention also provides a computer-readable medium. The computer-readable medium stores a computer program, which, when executed by a processor, implements the steps in any of the user requirement identification methods described in the above embodiments. The computer-readable storage medium may be volatile or non-volatile.

[0067] This invention also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the aforementioned user requirement identification method.

[0068] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0069] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0070] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0071] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0072] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0073] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0074] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0075] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0077] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A method for identifying user needs, characterized in that, include: Scene descriptions are obtained using pre-trained visual language models based on multimodal perception data. The multimodal perception data includes at least image data, audio data, vehicle status data, and external environment data; the visual language model is used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form. Based on the scene description and user history data, scene tags are generated using a first language model; the first language model is used to map the scene description into structured scene tags. Based on the scene tags and user history data, the second language model, which incorporates Maslow's hierarchy of needs theory, is used to generate a user's needs list for generating a task list.

2. The method according to claim 1, wherein, The step of generating scene tags using the first language model based on the scene description and user history data includes: Based on the user's historical data, prompting engineering techniques are used to guide the first language model to extract and populate predefined structured scene tags from the scene description; The user history data includes at least short-term memory data and long-term memory data; the short-term memory data includes at least one of the following: current scene description, historical scene tags, identified demand list, currently executing task status, and recent user interaction history; the long-term memory data includes at least one of the following: user's personality preferences, driving habits, historical commuting routes, family member information, and frequently used services; the structured scene tags include at least one of the following: trip stage tags, vehicle usage purpose tags, behavior tags, and vehicle status tags; the vehicle status tags include at least one of the following: personnel status tags, object status tags, vehicle environment status tags, and vehicle status tags.

3. The method according to claim 2, wherein, The trip stage labels include at least one of the following: before boarding, during driving, parking, charging or refueling, maintenance, and emergency. The vehicle use purpose labels include at least one of the following: commuting, travel, shopping, business reception, picking up family members, logistics transportation, and outdoor activities. The behavior labels include at least one of the following: working, entertainment, leisure, socializing, studying, and resting. The personnel status labels include at least one of the following: driver fatigue status, number and distribution of passengers, and status of children and pets. The object status labels include at least one of the following: the placement of items and the detection of hazardous materials. The vehicle environment status labels include at least one of the following: temperature and humidity, air quality, weather, and road congestion. The vehicle status labels include at least one of the following: vehicle speed, fuel consumption, battery level, and malfunction indicator light status.

4. The method according to claim 1, wherein, The process involves generating a user's needs list based on the scene tags and user history data, using the second language model incorporating Maslow's hierarchy of needs theory. This list includes: Based on Maslow's hierarchy of needs, configure the prompts and constraints of the second language model; Using the second major language model, a user's needs list is generated based on the scene tags and user history data.

5. The method according to claim 4, wherein, The configuration of prompts and constraints for the second language model, based on Maslow's hierarchy of needs, includes: The second major language model is configured as follows: for each scene tag, combined with user historical data, and based on Maslow's hierarchy of needs, the user's physiological needs, safety needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs are identified layer by layer; wherein, the priority of the user's needs from high to low is safety needs, physiological needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs.

6. The method according to claim 5, wherein, The generated task list includes: Based on the physiological needs, safety needs, love and belonging needs, esteem needs, cognitive needs, and aesthetic needs, corresponding physiological tasks, safety tasks, love and belonging tasks, esteem tasks, cognitive tasks, and aesthetic tasks are generated. Based on the priority of the user's needs, the tasks in the task list are executed sequentially.

7. The method according to claim 1, wherein, The image data is collected through an in-vehicle camera; the audio data is collected through an in-vehicle microphone array; the vehicle status data is collected through the vehicle CAN bus interface and environmental sensors; and the external environment data is collected through an external data interface, including at least user schedule and health data.

8. A user demand identification system, characterized in that, Capable of implementing the method as described in any one of claims 1 to 7, comprising: The scene description module is used to obtain a scene description based on multimodal perception data using a pre-trained visual language model; the multimodal perception data includes at least image data, audio data, vehicle status data, and external environment data; the visual language model is used to perform semantic understanding based on the multimodal perception data and output a scene description in natural language form. The scene tagging module is used to generate scene tags using a first language model based on the scene description and user history data; the first language model is used to map the scene description into structured scene tags. The requirement identification module is used to generate a user's requirement list based on the scene tags and user historical data, using the second language model that combines Maslow's hierarchy of needs theory, in order to generate a task list.

9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.