Two-way social intent interaction method based on VLN and EHMI and application of two-way social intent interaction method
By adopting a closed-loop architecture based on VLN and EHMI, two-way intention interaction of the intelligent autonomous mobile platform is realized, which solves the problems of one-way interaction and insufficient environmental understanding of the existing system, and improves the operational efficiency and safety in public places.
Patent Information
- Application Number
- CN202510930980.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-11-11
AI Technical Summary
Existing intelligent autonomous mobile platform (such as wheelchair) interaction systems suffer from limitations such as one-way interaction, insufficient environmental understanding, and ambiguous expression of intent, resulting in low operating efficiency and safety hazards in public places.
It adopts a closed-loop architecture based on Visual Language Model (VLN) and Extended Human-Computer Interface (EHMI), collects data through multimodal sensors, combines scene semantic understanding technology to achieve bidirectional intent understanding and expression, outputs joint decision results using Visual Language Model (VLM), and feeds back external intent through EHMI to form a continuously optimized closed loop.
It achieves multi-directional interaction, can simultaneously identify the intentions of users and social participants, reduce operational conflicts, improve collaboration efficiency in public places, accurately distinguish the priority of intentions in complex scenarios, and achieve continuous optimization by dynamically adjusting decision parameters through reinforcement learning.
Smart Images

Figure CN120928944A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence human-computer interaction technology, and to a two-way social intent interaction method and its application based on VLN (Visual Language Model) and EHMI (Extended Human-Computer Interface) for autonomous mobile platforms. In particular, it relates to a two-way social intent interaction method and its application based on VLN and EHMI for autonomous mobile platforms, which aims to solve the two-way communication problem of people with disabilities actively expressing their needs and passively recognizing external intentions in public spaces. Background Technology
[0002] As my country's population ages, the performance requirements for products with high demand among the elderly, such as wheelchairs, are increasing. Simultaneously, the rapid development of technologies like language models and artificial intelligence image recognition has led to various ideas and inventions related to intelligent autonomous mobility platforms (wheelchairs), such as automatic positioning and path planning. These inventions have, to some extent, optimized the user experience of autonomous mobility platforms (wheelchairs). However, existing inventions are mostly based on the intentions of autonomous mobility platform (wheelchair) users, neglecting the guidance role of other social participants, especially staff and caregivers, who often assist in guiding users in public places. Close interaction between disabled individuals and staff in public places is time-consuming and laborious, often resulting in low operational efficiency in complex environments.
[0003] Current intelligent autonomous mobility platform (wheelchair) interaction systems suffer from three major drawbacks: 1) Limited one-way interaction: Existing systems (such as eye-tracking control) only support user-initiated operations and struggle to recognize the intentions of pedestrians, leading to avoidance conflicts. For example, in a crowded hospital corridor, when a caregiver gestures for the autonomous mobility platform (wheelchair) to stop, traditional eye-tracking control systems may continue moving forward because they cannot interpret the external intention, forcing pedestrians to make emergency avoidance, which not only reduces traffic efficiency but may also cause collision risks; 2) Insufficient environmental understanding: Traditional sensor solutions struggle to interpret complex social scenarios (such as crowd gesture interactions) and are prone to misjudging emergency situations. For example, at subway station turnstiles, when multiple pedestrians point in different directions simultaneously, traditional systems based on infrared or ultrasonic sensors are prone to conflict and confusion, triggering erroneous emergency stops; 3) Ambiguous intention expression: The movement intentions of autonomous mobility platform (wheelchair) users cannot be effectively conveyed to those around them, resulting in low social efficiency and potential traffic safety hazards.
[0004] Therefore, developing a method that enables multi-directional interaction, allows users to understand their environment independently, and effectively conveys user intent to other people in the environment is of great practical significance. Summary of the Invention
[0005] Due to the aforementioned deficiencies in existing technologies, this invention provides a method that enables multi-directional interaction, self-service understanding of the environment, and effective transmission of user intent to other people in the environment. Specifically, it is a two-way social intent interaction method and its application for autonomous mobile platforms based on VLN (Visual Language Model) and EHMI (Extended Human-Computer Interface). Through the closed-loop architecture of VLM+EHMI, it achieves two-way intent understanding and expression, filling a technological gap and overcoming the deficiencies of existing autonomous mobile platform interaction systems, which can only perform one-way interaction, have insufficient environmental understanding capabilities, and have ambiguous intent expression.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A bidirectional social intent interaction method based on VLN and EHMI includes the following steps:
[0008] S1. Collect voice information and environmental image information of the autonomous mobile platform through the multimodal sensors on the autonomous mobile platform, extract dynamic scene context features by combining scene semantic understanding technology, and fuse the above data to construct a unified decision input vector;
[0009] S2. Input the unified decision input vector into the visual language model. The visual language model outputs a joint decision result, which includes the user's accurate intent I. u °, Diagram of other social participants I e °, Advanced Action Command A for the Autonomous Mobile Platform action and reward weight ω;
[0010] S3, Advanced Action Command A for the Autonomous Mobile Platform action The motion control model is input to obtain the timing control command sequence of the autonomous mobile platform, which controls the autonomous mobile platform to act according to the timing control command sequence and updates the pose and environmental status of the autonomous mobile platform in real time.
[0011] S4. Display the user's intent and the current operating purpose of the autonomous mobile platform on the EHMI in the form of short text or simple graphics, and display it in real time for surrounding people to understand, so as to realize two-way intent communication between the user and external social participants; specifically, after the microphone receives the user's voice instruction, it displays the intent on the EHMI. After receiving instructions from other social participants, it reflects the current processing status through the EHMI. After the VLM issues a command, it displays the next operation strategy on the EHMI through the built-in text corresponding to graphics, so as to achieve a two-way interactive effect and improve the interaction efficiency.
[0012] S5. Closed-loop dynamic optimization: Based on environmental feedback signals and reward weight ω, the VLM decision parameters are dynamically adjusted to form a continuous optimization closed loop of "perception-decision-execution-feedback".
[0013] The bidirectional social intent interaction method based on VLN and EHMI of this invention has a reasonable step sequence design. It can not only realize bidirectional intent parsing, but also use EHMI to feed back to the outside world to reduce operation conflicts (i.e., realize bidirectional intent interaction). Moreover, VLM accurately distinguishes the intent priority in complex scenarios by weighting key feature fragments, reducing the occurrence of erroneous action execution. In addition, the use of reinforcement learning for dynamic optimization can dynamically adjust the VLM decision parameters, which can realize the continuous optimization closed loop of interactive control, and has good application prospects.
[0014] As a preferred technical solution:
[0015] The above-described method for bidirectional social intent interaction based on VLN and EHMI includes a multimodal sensor comprising a wide-angle camera (covering a 120° field of view in front) and a directional microphone array (for directional noise reduction).
[0016] The voice information refers to the audio stream from the user and external personnel, and the environmental image information refers to the environmental video stream. The video stream is sampled at 30fps and uses binocular stereoscopic technology to read stereoscopic images; the audio stream uses beamforming technology to separate user commands from environmental noise.
[0017] The extraction process of dynamic scene context features in the bidirectional social intent interaction method based on VLN and EHMI, as described above, is as follows:
[0018] A graph convolutional neural network is used to analyze each image frame in the environmental image information of the autonomous mobile platform, calculate its node characteristics, and construct a scene semantic graph based on the node characteristics. In the scene semantic graph, nodes represent entity information (such as obstacle coordinates, social participant positions, and location type), and edges represent spatial relationships between entities (such as distance and relative orientation), thereby encoding the topological structure of the dynamic environment. The dynamic scene context features are the scene semantic graph, which includes single-node attribute semantics (such as target category, location, and size) and neighboring node association semantics (such as distance and positional relationship between two targets).
[0019] The decision-making process of VLM introduces an intent-driven attention mechanism: using the user's current command as the query vector, weights are assigned to multimodal input features (visual actions, speech text, scene graphs), prioritizing feature segments with high relevance to intent, thereby improving the robustness of intent recognition in complex scenarios.
[0020] The training process of the visual language model in the bidirectional social intent interaction method based on VLN and EHMI, as described above, is as follows:
[0021] Using the unified decision input vector from the training dataset as input, and the precise intent of the user corresponding to the unified decision input vector, the pointer diagrams of other social participants, and the high-level action instructions A for the next step of the autonomous mobile platform, the input vector is used. action The process of continuously adjusting model parameters, using the reward weight ω as the theoretical output, involves the unified decision input vector of the training data in the training dataset corresponding to the user's accurate intent, the pointer diagrams of other social participants, and the advanced action instructions A of the autonomous mobile platform for the next step. action The reward weight ω is known.
[0022] The training dataset for the Visual Language Model (VLM) primarily includes user intent text and other directional gestures and eye contact data from common public locations such as supermarkets, hospitals, and subway stations. During VLM training, the feedback from the current and previous iterations is stored for subsequent debugging and for optimization of social aspects using reinforcement learning through comparison.
[0023] The aforementioned bidirectional social intent interaction method based on VLN and EHMI includes a timing control command sequence comprising steering angle and speed parameters. The motion control model can be SayCan, which will... action Decomposed into timing control instruction sequences :
[0024]
[0025]
[0026] Among them, A action For high-level action commands, φ represents the model's learned weights (fixed after training), and a t S is the underlying action for time step t. t Let t be the full-dimensional state vector of the wheelchair at time step t.
[0027] During instruction execution, it is possible to set up a dangerous behavior recognition interruption based on graph convolutional networks to ensure the safety of the movement; at the same time, parameters such as running time, speed, acceleration, and jerk are recorded to facilitate subsequent reinforcement learning optimization.
[0028] The above-described bidirectional social intent interaction method based on VLN and EHMI includes environmental feedback signals such as pedestrian reactions and obstacle avoidance effects.
[0029] The aforementioned bidirectional social intent interaction method based on VLN and EHMI utilizes a closed-loop dynamic optimization framework. Specifically, it designs a multi-dimensional reward function (covering safety, efficiency, comfort, and social coordination) and automatically adjusts the VLM decision parameters based on environmental feedback, enabling the system to adapt to diverse scenario requirements. The reward function primarily consists of four parts: comfort, safety, efficiency, and social interaction. By rewarding and penalizing performance, the output of the VLM is dynamically adjusted, continuously optimizing the system's performance. The weights of the reward function are set during VLM pre-training, and during invocation, the VLM continuously provides weight vectors based on the current scenario.
[0030] The present invention also provides a computer device, the computer device comprising:
[0031] At least one processor; and,
[0032] A memory communicatively connected to the at least one processor; wherein,
[0033] The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements a two-way social intent interaction method based on VLN and EHMI as described above.
[0034] Furthermore, the present invention also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement a bidirectional social intent interaction method based on VLN and EHMI as described above.
[0035] The above technical solution is only one feasible technical solution of the present invention. The scope of protection of the present invention is not limited thereto. Those skilled in the art can reasonably adjust the specific design according to actual needs.
[0036] The above invention has the following advantages or beneficial effects:
[0037] (1) The bidirectional social intent interaction method based on VLN and EHMI of the present invention uses VLM to parse bidirectional intent, which can simultaneously identify user intent (voice / action) and social participant intent (gesture, eye contact, location), and then output joint decision instructions;
[0038] (2) The bidirectional social intent interaction method based on VLN and EHMI of the present invention uses EHMI to feed back to the outside world, and displays the user's intent and the autonomous mobile platform's response status and target action in the form of short text in real time (such as "received information"), informing people around, reducing misunderstandings, avoiding operational conflicts (such as pedestrians misjudging the autonomous mobile platform's path), improving the efficiency of collaboration in public places, and realizing bidirectional intent interaction.
[0039] (3) The bidirectional social intent interaction method based on VLN and EHMI of the present invention uses VLM to accurately distinguish the intent priority in complex scenarios by weighting key feature fragments, thereby reducing the occurrence of erroneous action execution;
[0040] (4) The bidirectional social intent interaction method based on VLN and EHMI of the present invention can dynamically adjust the VLM decision parameters by using reinforcement learning dynamic optimization, and can realize the continuous optimization closed loop of interactive control, which has good application prospects. Attached Figure Description
[0041] The invention, its features, shape, and advantages will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. Like reference numerals denote like parts throughout the drawings. The drawings are not drawn to scale; their focus is on illustrating the gist of the invention.
[0042] Figure 1 This is a diagram illustrating the overall control architecture of the bidirectional social intent interaction method based on VLN and EHMI of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, beneficial effects, and significant advancements of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention.
[0044] Obviously, all the embodiments described are only some embodiments of the present invention, and not all embodiments; based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Example 1
[0046] A bidirectional social intent interaction method based on VLN and EHMI, the architecture of which is as follows: Figure 1 As shown, it includes the following steps:
[0047] S1: Using multimodal sensors mounted on the autonomous mobile platform, including a wide-angle camera (covering a 120° field of view in front) and a microphone array (for directional noise reduction), the system collects in real time the user's language commands V, scene image frames A, and scene context information C. This forms the input vector X=[V,A,C], which constitutes the prompt for the VLM. Specifically:
[0048] S101: The microphone array is directional to reduce ambient noise and accurately collect the voice commands of users on the autonomous mobile platform, accurately converting the sound information into text content V;
[0049] S102: Wide-angle camera captures video and decomposes the video stream into each frame image A;
[0050] S103: A wide-angle camera captures video, and combined with binocular stereo vision technology, it uses a graph neural network (GCN) to calculate the characteristics of nodes in each image frame of the video stream. Each node represents an entity (such as a pedestrian, autonomous mobile platform, or obstacle), and includes: coordinate position (x, y) and depth information z to form V. node E is composed of the physical distance (Euclidean distance) and relative angles (such as the angle between the pedestrian's orientation and the autonomous mobile platform) between each node. edge ;
[0051]
[0052]
[0053] Among them, W C represents the parameters of the graph neural network (GCN); C represents the scene context information output by the graph neural network.
[0054] S2: In the scene interaction mode, vector X is used to construct a good prompt input to a VLM model pre-trained for human gestures, and the processed output is the accurate intent I of the autonomous mobile platform user. u °, Diagram of other social participants I e °, Advanced Action Command A for the Autonomous Mobile Platform action And the reward weight ω generated based on the current scenario for subsequent optimization;
[0055] D=VLM(X;θ)
[0056] D=[I e °,I u °,A action ,ω]
[0057] Where D represents the precise intent of the autonomous mobile platform user (I). u °, Diagram of other social participants I e °, Advanced Action Command A for the Autonomous Mobile Platform action The vector consists of reward weights ω generated based on the current scene for subsequent optimization; VLM(X;θ) is the output of the visual language model, where θ is the model parameter, which can be continuously adjusted based on feedback results;
[0058] S201: Training the VLM model, which is a deep representation learning process that integrates multimodal data and scene-specific knowledge. Its core lies in building a model foundation capable of accurately understanding visual-language relationships in complex social scenarios, specifically:
[0059] S2011: Dataset construction involves collecting video sequences of interactions between autonomous mobile platform users and surrounding people in real-world public spaces (such as hospitals, subway stations, and shopping mall passages). The data includes gestures from other social participants as well as the intentions of the autonomous mobile platform users themselves. The data must cover diverse social scenarios (asking for directions, giving way, asking for help, and cooperating in passage).
[0060] S2012: Data annotation: Visual annotation: Fine-grained annotation of video frames, including elements such as surrounding pedestrians and key objects (doors, buttons);
[0061] S2013: Scene Analysis: Pixel-level semantic segmentation to identify ground, walls, obstacles, passage areas, etc.;
[0062] S2014: Text annotation: Intent description: Natural language description of the potential intents of the participants (including autonomous mobile platform users and interactive objects) in the current scene (such as "requesting to give way", "asking for directions", "indicating priority", "providing directions");
[0063] S2015: Context Description: Describe the environmental state (crowding, lighting, spatial structure), social relationships (strangers, medical staff), and interaction state (initiation, response, persistence);
[0064] S2016: Command-Action Pair: The correspondence between possible voice commands from autonomous mobile platform users and the actions expected to be performed by the autonomous mobile platform;
[0065] S2017: By labeling the dataset, the weights of the needs of autonomous mobile platform users for comfort, safety, efficiency and social interaction in each case are labeled, which makes it easier to continuously optimize the LLM output through reinforcement learning during subsequent debugging and use;
[0066] S2018: Using the unified decision input vector from the training dataset as input, and taking the accurate intent of the user corresponding to the unified decision input vector, the pointer diagrams of other social participants, and the high-level action instructions A for the next step of the autonomous mobile platform,... action The process of continuously adjusting model parameters, with the reward weight ω as the theoretical output;
[0067] S202: The three feature segments of the input vector are as follows:
[0068] (1) k1: Visual features (such as the embedding vector of other social participants’ hand gestures)
[0069] (2) k2: Audio features (BERT embedding of the speech “Please avoid pedestrians”)
[0070] (3) k3: Scene graph features (such as the "pedestrian-autonomous mobile platform distance" output by GCN)
[0071] S203: Calculate the weight corresponding to each feature segment using the following formula;
[0072]
[0073] Where q is the intent query vector, corresponding to the voice command of the autonomous mobile platform user; x is the vector input to the model, composed of k1, k2, and k3 mentioned above; α i Weights corresponding to each feature segment; The key transformation matrix linearizes the x vector into a key vector for weight calculation.
[0074] S204: Output the accurate intent I of the autonomous mobile platform user using the following formula. u °, Diagram of other social participants I e ° and the next advanced action command A for autonomous mobile platforms action ;
[0075]
[0076] Where D is the result matrix, containing the accurate intent I of the autonomous mobile platform user. u °, Diagram of other social participants I e ° and the next advanced action command A for autonomous mobile platforms action x is the vector of the input model; The value transformation matrix is used to transform x into a value vector to construct the final representation D, α i Weights corresponding to each feature segment;
[0077] S205: Due to the time required for VLM to process information, and taking security into consideration, during the time between the two generation of D by VLM, the autonomous mobile platform is allowed to be interrupted based on GCN to determine emergency situations (such as stopping the autonomous mobile platform if the road ahead is blocked).
[0078] S206: Once the VLM has completed the next step of the run plan, it passes its contents to the EHMI so that the EHMI can display the run objectives to other participants;
[0079] S3: Based on the current situation of the autonomous mobile platform, the decomposed A will be achieved through an operation control model (such as SayCan). action The sequence of underlying control commands for autonomous mobile platforms (Similar to chassis movement);
[0080] S301: Instruction semantic parsing and motion parameterization, parsing Aaction Extract key motion parameters based on semantic type (e.g., steering, obstacle avoidance, emergency stop, etc.):
[0081] (1) Target direction θ target (Based on S) t (Environment map and obstacle locations in the image).
[0082] (2) Velocity constraint v max (Dynamically adjusted based on scene congestion);
[0083] (3) Target position x target ;
[0084] S302: State-action mapping modeling, using a decision function similar to Google's open-source SayCan framework:
[0085]
[0086] S t It includes various states of the current autonomous mobile platform, including yaw angle, speed, acceleration, etc.
[0087] S303: Generation of Kinematic Command Sequences
[0088] Abstract actions are broken down into low-level instructions that occur continuously over time. For example, consider a hospital caregiver's guidance scenario:
[0089]
[0090] S304: Dynamic Status Updates and Feedback
[0091] After execution a t Update system status
[0092]
[0093] The pose (coordinates and heading angle) of the autonomous mobile platform.
[0094] Control period (determined experimentally based on the reaction time of the VLM and the reaction rate of the control model)
[0095] Real-time updates of scene graph node relationships for the next round of VLM.
[0096] S4: Based on EHMI, the intentions (V) of the autonomous mobile platform user and the current operational purpose (a) of the autonomous mobile platform are displayed on the EHMI in the form of short text or simple graphics to achieve the purpose of two-way interaction.
[0097]
[0098] S401: Upon receiving instructions from the autonomous mobile platform, display them on the EHMI for other social participants to read, facilitating further guidance and improving social efficiency;
[0099] S402: After receiving the action content sent by VLM, it displays the corresponding graphic on EHMI through its built-in simple text-to-graphic correspondence, so that other social participants can read it and improve the efficiency of interaction.
[0100] S5: By utilizing environmental feedback signals (FF), the decision weight parameters of the visual language model (VLM) are dynamically adjusted through reinforcement learning to achieve continuous optimization of the system's decision-making capabilities, forming a closed loop of "decision-execution-feedback-learning".
[0101] S501: The parameter optimization feedback mechanism primarily focuses on the comfort, safety, efficiency, and social interaction of the intelligent autonomous mobile platform. It mainly utilizes sensors to measure the current acceleration and jerk of the autonomous mobile platform; for safety, it is primarily determined by the environmental hazard assessment based on the graph neural network between two VLM outputs; for efficiency, it is primarily determined by the speed at which the intended action is completed; for social interaction, it is primarily determined by the characteristic differences between the VLM recognition intentions of other social participants in the previous moment and the VLM recognition intentions of other social participants in the current moment.
[0102] S502: Based on a weight vector of comfort, safety, efficiency, and sociality. To optimize comfort, a reward function is set up:
[0103]
[0104] Among them, R safety Indicates a safety reward; R efficiency Rewards for efficiency; R comfort For comfort rewards; R social Indicates social reward; ω i The weight vector is generated by VLM based on the scenario (such as a hospital or subway station). The weight allocation is dynamically adjusted according to the scenario to ensure a dynamic balance of multidimensional needs under different situations. The final reward value is then passed to the reinforcement learning algorithm after being weighted and summarized to guide the decision network to optimize the strategy, ensuring the system's efficiency, robustness, and adaptability in complex scenarios.
[0105] Example 2
[0106] A computer device includes: at least one processor and a memory communicatively connected to the at least one processor;
[0107] The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the bidirectional social intent interaction method based on VLN and EHMI as described in Example 1.
[0108] Example 3
[0109] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the bidirectional social intent interaction method based on VLN and EHMI as described in Embodiment 1.
[0110] Verification has shown that the bidirectional social intent interaction method based on VLN and EHMI of this invention utilizes VLM to parse bidirectional intents, simultaneously recognizing user intents (voice / action) and the intents of social participants (gestures, eye contact, location), and then outputting joint decision commands. By using EHMI to feed back to the outside world, the user intent, the autonomous mobile platform's response status, and the target action are displayed in real-time in short text format (e.g., "Message received"), informing those around, reducing misunderstandings, avoiding operational conflicts (e.g., pedestrians misjudging the autonomous mobile platform's path), improving collaboration efficiency in public places, and achieving bidirectional intent interaction. Furthermore, by using VLM to weighted key feature fragments, the priority of intents in complex scenarios is accurately distinguished, reducing the occurrence of erroneous action execution. The use of reinforcement learning for dynamic optimization can dynamically adjust VLM decision parameters, achieving a continuous optimization closed loop for interactive control, demonstrating promising application prospects.
[0111] Those skilled in the art should understand that variations can be implemented by combining existing technology with the above embodiments, which will not be elaborated here. Such variations do not affect the essence of the present invention, and will not be elaborated here either.
[0112] The preferred embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and the devices and structures not described in detail should be understood as being implemented in a conventional manner in the art. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the present invention. This does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the present invention's technical solutions still fall within the protection scope of the present invention.
Claims
1. A bidirectional social intent interaction method based on VLN and EHMI, characterized in that: Includes the following steps: S1. Collect voice information and environmental image information of the autonomous mobile platform through the multimodal sensors on the autonomous mobile platform, extract dynamic scene context features by combining scene semantic understanding technology, and fuse the above data to construct a unified decision input vector; S2. Input the unified decision input vector into the visual language model. The visual language model outputs a joint decision result, which includes the user's accurate intent I. u °, Diagram of other social participants I e °, Advanced Action Command A for the Autonomous Mobile Platform action and reward weight ω; S3, Advanced Action Command A for the Autonomous Mobile Platform action The motion control model is input to obtain the timing control command sequence of the autonomous mobile platform, which controls the autonomous mobile platform to act according to the timing control command sequence and updates the pose and environmental status of the autonomous mobile platform in real time. S4. Display the user's intent and the current operational purpose of the autonomous mobile platform on the EHMI in the form of short text or simple graphics; S5. Closed-loop dynamic optimization: Based on environmental feedback signals and reward weight ω, the VLM decision parameters are dynamically adjusted to form a continuous optimization closed loop of perception-decision-execution-feedback.
2. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The multimodal sensor includes a wide-angle camera and a directional microphone array; The voice information refers to the audio stream of the user and external personnel, and the environmental image information refers to the environmental video stream.
3. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The extraction process of the dynamic scene context features is as follows: A graph convolutional neural network is used to analyze each image frame in the environmental image information of the autonomous mobile platform, calculate its node characteristics, and construct a scene semantic graph based on the node characteristics. The dynamic scene context features are the scene semantic graph, which includes single-node attribute semantics and neighbor node association semantics.
4. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The training process of the visual language model is as follows: Using the unified decision input vector from the training dataset as input, and the precise intent of the user corresponding to the unified decision input vector, the pointer diagrams of other social participants, and the high-level action instructions A for the next step of the autonomous mobile platform, the input is further processed. action The process of continuously adjusting model parameters, using the reward weight ω as the theoretical output, involves the unified decision input vector of the training data in the training dataset corresponding to the user's accurate intent, the pointer diagrams of other social participants, and the advanced action instructions A of the autonomous mobile platform for the next step. action The reward weight ω is known.
5. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The timing control command sequence includes steering angle and speed parameters.
6. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The environmental feedback signals include pedestrian reactions and obstacle avoidance performance.
7. The bidirectional social intent interaction method based on VLN and EHMI according to claim 1, characterized in that, The closed-loop dynamic optimization is implemented based on a reinforcement learning framework. Specifically, it involves designing a multi-dimensional reward function and automatically adjusting the VLM decision parameters of the VLM based on environmental feedback.
8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements a two-way social intent interaction method based on VLN and EHMI as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement a bidirectional social intent interaction method based on VLN and EHMI as described in any one of claims 1 to 7.