Method and system for human-machine trust calibration based on large model-driven digital human
The digital human system driven by a large model enables multimodal data fusion and dynamic interaction, solves the problem of insufficient trust calibration in autonomous driving systems, improves driver trust and safety, and enhances user experience.
Patent Information
- Application Number
- CN202510613385.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing autonomous driving systems lack an effective trust calibration mechanism in Level 3 human-machine co-driving mode, resulting in insufficient or excessive driver trust, which affects driving safety and user experience.
A large-model-driven digital human system is adopted. Through multimodal data fusion and dynamic interaction design, driver feature information is collected in real time to generate personalized interaction strategies, including voice, facial expressions and body movements. Combined with the vehicle cockpit system, multi-channel control is carried out to achieve precise adjustment of trust level.
It improves the accuracy and adaptability of trust calibration in different driving scenarios for autonomous driving systems, enhances driver trust and safety, and improves user experience.
Smart Images

Figure CN120534374B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, specifically to a human-machine trust calibration method and system based on a large model-driven digital human. Background Technology
[0002] As autonomous driving technology gradually evolves towards higher levels, such as Level 3, human-machine co-driving has become an inevitable trend in the development of intelligent transportation systems. According to statistics from the Society of Automotive Engineers (SAE), in Level 3 autonomous driving scenarios, human-machine collaboration efficiency is optimal when the driver's trust level in the system is between 0.4 and 0.6. However, existing systems generally suffer from trust bias due to a lack of trust calibration. Insufficient trust (trust level < 0.4) can lead to multiple unnecessary driver takeovers, significantly increasing the driver's workload; while excessive trust (trust level > 0.6) can increase driver reaction delays, potentially causing significant safety hazards in emergency scenarios.
[0003] Against this backdrop, current interactive systems exhibit the following significant shortcomings in dynamic trust adjustment: 1. Traditional interfaces can only provide one-way information output—traditional agent methods may only offer fixed voice prompts or simple interface information, lacking more natural and nuanced interaction methods. This results in monotonous dialogue and rigid interaction behaviors, making it difficult to achieve long-term effective trust calibration. 2. Virtual agents lack an emotional interaction dimension—while existing virtual agents possess a static, anthropomorphic appearance, their actions lack human-like characteristics, making it difficult to enhance user trust in autonomous driving systems. 3. Dialogue systems are limited by insufficient semantic understanding capabilities and lack adaptive adjustment capabilities—while traditional conversational agents can achieve trust regulation to some extent, their cross-individual effectiveness and timeliness are not high. They are often based on preset rules or keyword matching, unable to accurately understand the driver's complex speech, ambiguous expressions, or emotional language. This makes the interaction process unable to reflect the randomness and variability of human interaction, and the effectiveness is difficult to guarantee.
[0004] These shortcomings make it difficult for existing technologies to construct trust calibration mechanisms that conform to human cognitive patterns. Currently, digital humans, as an innovative interaction method, have made some progress in the field of trust building. They can enhance drivers' trust in autonomous driving systems to a certain extent and provide emotional support in certain scenarios. However, their ability to cope with various scenarios and regulate complex emotions remains limited. Digital humans mainly rely on preset rules and simple emotional feedback strategies. In more complex driving scenarios or when drivers experience significant emotional fluctuations, their performance may not fully meet the driver's needs and they lack the ability to adapt to individual driver differences and driving behavior variations.
[0005] In summary, there is still an urgent need to develop a more adaptable human-computer trust calibration method and system with multimodal interaction capabilities. Summary of the Invention
[0006] The present invention aims to provide a human-machine trust calibration method and system based on a large model-driven digital human, which has strong adaptability and adjustment capabilities, can accurately adapt interaction strategies in diverse driving scenarios, ensure that the driver maintains an appropriate sense of trust in different driving scenarios, thereby improving the safety and user experience of autonomous driving.
[0007] To achieve the above objectives, the present invention provides the following basic solution.
[0008] Option 1
[0009] A human-machine trust calibration system based on a large model-driven digital human includes an in-vehicle computing unit and an in-vehicle HMI system, as well as a driver monitoring system; the driver monitoring system includes an information acquisition module and a trust detection module.
[0010] The information acquisition module is used to collect the driver's feature information, including facial features, gaze features, voice features, and physiological activity features. The trust detection module is used to perform trust detection based on the feature information and a deep learning model, and output the driver's current trust level.
[0011] The in-vehicle computing unit includes a prompt word generation module, an inference module, and a processing module; the in-vehicle HMI system includes a digital human interaction module.
[0012] The prompt generation module generates a target prompt based on the current driving scenario and the driver's current level of trust. The reasoning module outputs a digital human motion adjustment scheme and natural language dialogue content based on the target prompt and the current driving scenario. The processing module converts the digital human motion adjustment scheme and natural language dialogue content into interactive instructions and outputs them to the digital human interaction module. The digital human interaction module displays interactive actions based on the interactive instructions.
[0013] Furthermore, the information acquisition module includes a camera for acquiring facial and gaze features, a voice acquisition device for acquiring voice features, and a physiological sensor for acquiring physiological activity features.
[0014] Furthermore, the prompt generation module includes the following steps when generating the target prompt:
[0015] It receives feature information, driving behavior feature vectors, vehicle environment perception data and trust level output by the trust detection module in real time, and constructs a structured JSON data packet containing four data dimensions: driver status, driving scenario, vehicle status and trust level.
[0016] The JSON data packet is parsed and reconstructed using a preset dynamic template engine to generate an enhanced prompt that incorporates the target elements, which serves as the target prompt.
[0017] Furthermore, the target elements include the digital human's role setting parameters, the semantic description of the current driving scenario, the trust level quantification value, the human-machine collaborative control target constraints, the generation requirements, and the digital human's action-expression coupling descriptor.
[0018] Furthermore, when the inference module outputs the digital human motion adjustment scheme and natural language dialogue content, it includes the following steps:
[0019] Large models adapted to driving scenarios are called through predefined API interfaces;
[0020] By combining a pre-set knowledge base, the large model is driven to perform joint reasoning to generate an interactive response that includes digital human motion adjustment schemes and natural language dialogue content;
[0021] The generated digital human motion adjustment scheme and natural language dialogue content are processed by the processing module, then encapsulated by the TCP / IP protocol stack, and transmitted to the vehicle HMI system via the vehicle Ethernet bus.
[0022] The knowledge base includes semantic understanding rules for driving scenarios, a mapping table of digital human motion feature parameters, and a multimodal interaction strategy database. The digital human motion adjustment scheme includes at least facial expression parameters, limb movement commands, and speech synthesis feature vectors.
[0023] Furthermore, the digital human interaction module includes a pre-set Unity-based digital human.
[0024] Furthermore, the digital human establishes an interactive connection with the vehicle cockpit system; when the interactive command includes cockpit interactive content, after the digital human finishes broadcasting, a control command is sent to the vehicle cockpit system via CAN bus or vehicle communication protocol to trigger response operations of seat vibration, instrument panel display, air conditioning system or windows.
[0025] Furthermore, the digital human interaction module displays interactive actions based on interaction commands, including the following steps:
[0026] Receive interactive commands. When the interactive commands contain voice broadcast content, perform speech synthesis and lip-sync processing through a locally deployed end-to-end deep learning model.
[0027] Among them, based on interactive commands, the Wav2Lip or RTFace model is used to predict the mouth shape trajectory in real time on the audio stream, and generate lip movement data synchronized with the speech waveform for mouth shape alignment.
[0028] When the interactive command contains action content, the HumanTOMATO or EMAGE deep learning model is used to generate a digital human action sequence that matches the interactive command in real time.
[0029] The motion sequence is dynamically segmented using a streaming engine, and continuous motion is decomposed into data packets containing timestamps, joint rotation angles, and motion trajectory vectors according to a preset time window.
[0030] Data packets are streamed to the Unity rendering engine via the UDP protocol, driving the digital human to achieve frame-synchronized rendering of body movements.
[0031] Furthermore, the trust levels include excessive trust, appropriate trust, and insufficient trust;
[0032] When the trust level is excessive trust, the inference module uses trust suppression as the control target when outputting digital human action adjustment schemes and natural language dialogue content; when the trust level is insufficient trust, the inference module uses trust enhancement as the control target when outputting digital human action adjustment schemes and natural language dialogue content.
[0033] Option 2
[0034] The human-machine trust calibration method based on a large model-driven digital human applies the human-machine trust calibration system based on a large model-driven digital human as described in Scheme 1 to perform human-machine trust calibration, including the following steps:
[0035] The driver monitoring system acquires the driver's characteristic information in real time through its information acquisition module. The characteristic information includes facial features, gaze features, voice features, and physiological activity features.
[0036] The trust detection module uses a deep learning model to analyze the feature information to obtain the driver's current trust level.
[0037] The prompt generation module generates a target prompt based on the current driving scenario and the trust level; the target prompt is input into the inference module, which outputs a digital human motion adjustment scheme and natural language dialogue content adapted to the current scenario.
[0038] The processing module converts the digital human motion adjustment scheme and natural language dialogue content into interactive commands that can drive the digital human; the digital human interaction module of the vehicle HMI system executes the interactive commands and displays the digital human's interactive actions and dialogue content in real time.
[0039] The working principle and advantages of this invention are as follows:
[0040] This invention presents a human-machine trust calibration method and system based on a large-model-driven digital human. It possesses strong adaptive adjustment capabilities, enabling precise adaptation of interaction strategies across diverse driving scenarios. This ensures that the driver maintains an appropriate level of trust in different driving situations, thereby improving the safety and user experience of autonomous driving. The key points are:
[0041] First, this solution significantly improves the effectiveness, accuracy, and adaptability of trust calibration through multimodal data fusion and dynamic interactive design. Firstly, this solution employs a multi-channel fusion approach. The system first integrates multi-source information such as facial features, voice emotion, and physiological signals for trust detection, achieving higher accuracy compared to traditional single-dimensional trust detection methods. Then, based on the detection results, multi-channel adjustment is performed in conjunction with a digital human, which can more effectively regulate the driver's trust level, thereby improving user acceptance and optimizing the intelligent system's ability to calibrate driver trust.
[0042] Secondly, this solution employs a highly human-like digital human design, combined with the theory of human empathy, which can both enhance driver trust and suppress trust under appropriate circumstances. Specifically, in digital human control, this solution uses lip-syncing technology and high-precision motion generation models (such as Wav2Lip and HumanTOMATO) to ensure high coordination between the digital human's facial expressions, movements, and voice, enhancing the naturalness and credibility of the interaction. Furthermore, it is linked with the vehicle's cockpit system, and the collaboratively constructed multi-channel control mechanism (visual, auditory, and tactile coordination) can dynamically adjust the interaction strategy according to the driver's state, and the interaction method has a stronger communicative effect. It can enhance trust through empathetic interaction when trust is insufficient, and suppress trust through warning feedback when trust is excessive, forming a closed-loop calibration capability. In addition, the deep integration of the digital human with the vehicle's cockpit system (such as linking seat vibration and instrument warnings via CAN bus) extends virtual interaction to the physical feedback level, enhancing the driver's perception of system interaction.
[0043] Secondly, this solution, through structured data integration and dynamic prompt generation technology, enables in-depth analysis and personalized adaptation of driver characteristics. Compared to traditional template-based digital human interaction technology, this solution dynamically generates targeted prompts in real time by combining multi-dimensional data such as driving scenarios, vehicle status, and driver status. This drives the inference module to call a large model to generate differentiated interaction strategies, thereby constructing a highly flexible digital human that can provide more refined and personalized trust control. This breaks through the limitations of traditional fixed templates, allowing the interaction content to dynamically adapt to the driver's real-time psychological state and scenario needs, no longer limited to simple driving scenarios and simple driving states.
[0044] Third, this solution has strong adaptability to real-world scenarios. Specifically, by inferring from a prompt dynamically generated based on the driving environment and driver trust, the system can dynamically generate interaction strategies deeply bound to the current scenario, making the digital human's trust calibration more consistent with real-world driving scenarios. Furthermore, the combined application of streaming and inference modules ensures real-time interaction while facilitating rapid response to sudden changes in the scenario, guaranteeing the system's real-time interactivity.
[0045] Furthermore, compared to existing human-computer interaction trust processing solutions, such as the system and method for improving driving trust based on multimodal intelligent interaction disclosed in patent publication number CN118918558A, this solution possesses stronger interaction naturalness, higher calibration accuracy, and deeper system integration. While existing patent solutions can dynamically adjust interaction strategies, they focus on physical feedback, with trust levels calculated statically through weighted scoring. Their technical implementation path is based on a finely tuned pre-trained large model (Trust-Baichuan2) and updated memory data, and interaction recommendations rely on fixed templates. Moreover, they lack highly collaborative anthropomorphic interaction capabilities, resulting in insufficient expressiveness and immersion. Although virtual avatars can be set, these avatars mostly provide feedback through preset voice commands, lacking high freedom and dynamic adaptation capabilities.
[0046] This solution focuses on the dynamic generation and emotional interaction of digital humans. It centers on a highly human-like digital human, combining a large model to dynamically generate facial expressions, actions, and voice interactions. This is then integrated with vehicle cockpit systems (such as seat vibration and dashboard) to achieve multi-sensory collaborative feedback and cross-system collaboration. A complete, highly flexible virtual avatar feedback implementation scheme is constructed, ensuring the virtual avatar possesses highly coordinated voice-expression-action linkage capabilities. This not only allows the large model to infer control decisions but also enables it to generate voice dialogues that conform to current emotional semantics. By combining with other deep learning models, it achieves precise synchronization of lip movements and speech, natural matching of the current scene and semantics, and consistent control of action feedback and dialogue content. This significantly improves the human-machine trust control capabilities and user experience of autonomous driving systems (existing patent documents only use virtual avatars for information projection, lacking dynamic generation and emotional interaction capabilities).
[0047] Secondly, this solution sets up a dynamic prompt generation strategy, which combines facial expressions, voice emotions and other data to generate structured JSON packages, and builds an enhanced prompt through a dynamic template engine. This allows the interaction strategy to accurately adapt to the driver's real-time psychological state, making full use of multimodal data and achieving a high degree of interaction adaptability (existing patent documents rely on predefined label data, such as GSR weighted scores, which cannot dynamically parse multidimensional data).
[0048] Furthermore, this solution sets up a differentiated trust adjustment strategy, dividing trust levels into "excessive trust" and "insufficient trust", and designs corresponding control targets (such as a serious expression on the digital human and triggering seat vibration when trust is suppressed), which can achieve two-way calibration (enhancing or suppressing trust), making the calibration more targeted (existing patent documents are one-way trust enhancement). Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the system architecture of the human-machine trust calibration method and system based on a large model-driven digital human according to the present invention;
[0050] Figure 2 This is a schematic diagram of the trust calibration process in an embodiment of the human-machine trust calibration method and system based on a large model-driven digital human of the present invention.
[0051] Figure 3 This is a schematic diagram of a JSON data packet representing an embodiment of the human-machine trust calibration method and system for digital humans based on a large model, as described in this invention.
[0052] Figure 4 This is a prompt example diagram of the human-machine trust calibration method and system based on a large model-driven digital human according to the present invention;
[0053] Figure 5This is a schematic diagram of the large model control process in the inference module of the human-machine trust calibration method and system embodiment of the present invention based on a large model-driven digital human.
[0054] Figure 6 This is a schematic diagram of the digital human driving and control process in an embodiment of the human-machine trust calibration method and system based on a large model driven digital human of the present invention. Detailed Implementation
[0055] The following detailed explanation illustrates the specific implementation methods:
[0056] The basic implementation examples are as follows: Figure 1 As shown: A human-machine trust calibration system based on a large model-driven digital human includes an in-vehicle computing unit, an in-vehicle HMI system, and a driver monitoring system. HMI, or Human-Machine Interface, refers to the interface or interaction platform between humans and machine systems, typically found in common automotive center console screens or in-vehicle touchscreens.
[0057] The driver monitoring system includes an information collection module and a trust detection module.
[0058] The information acquisition module is used to collect the driver's characteristic information, including facial features, gaze features, voice features, and physiological activity features.
[0059] The information acquisition module includes a camera for acquiring facial and gaze features, a voice acquisition device for acquiring voice features, and a physiological sensor for acquiring physiological activity features. In this embodiment, the physiological sensor includes a heart rate sensor, a skin resistance sensor, etc. The physiological activity features include heart rate and its variation parameters, and skin resistance and its variation parameters.
[0060] The trust detection module is used to perform trust detection based on feature information and a deep learning model, and output the driver's current trust level. The trust level includes over-trust, appropriate trust, and under-trust.
[0061] The in-vehicle computing unit includes a prompt word generation module, an inference module, and a processing module; the in-vehicle HMI system includes a digital human interaction module.
[0062] The prompt generation module generates a target prompt based on the current driving scenario and the driver's current level of trust. Here, "prompt" refers to a prompt word; in natural language processing or large model generation tasks, a prompt is a text input or instruction used to guide the model to generate a specific response or perform a specific task. For example... Figure 4 As shown, an example of a target prompt is presented:
[0063] "[Character Setting] You are an autonomous driving trust control assistant, responsible for engaging in natural language conversations with drivers to help them adapt to the autonomous driving system."
[0064] [Driving Environment]
[0065] Road conditions: {road_condition};
[0066] Traffic conditions: {traffic};
[0067] Weather: {weather};
[0068] Vehicle speed: {vehicle_speed};
[0069] [Trust Level]
[0070] Driver trust level: {trust_level} (insufficient trust / appropriate trust / excessive trust);
[0071] [Regulation Target]
[0072] The goal is: {adjustment_goal} (to increase / decrease trust);
[0073] [Generation Requirements]
[0074] Based on the above information, generate a natural and fluent dialogue: If there is insufficient trust, please use a reassurance and explanation strategy to enhance the driver's confidence in autonomous driving; if there is excessive trust, please appropriately remind the driver to stay attentive; the dialogue should be concise and natural (within 50 words).
[0075] [Digital Human Design]
[0076] Based on the above information and the driving scenario, describe the digital human's body movements, facial expressions, and whether it interacts with other cabin devices (windows, ambient lighting, aromatherapy).
[0077] Specifically, the prompt generation module includes the following steps when generating the target prompt:
[0078] It receives feature information, driving behavior feature vectors, in-vehicle environment perception data, and trust levels output by the trust detection module in real time, and constructs a structured JSON data packet containing four data dimensions: driver status, driving scenario, vehicle status, and trust level; such as Figure 3As shown, JSON, or JavaScript Object Notation, is a lightweight data-interchange format file that typically consists of key-value pairs; the key is a string (usually a descriptive name), and the value is the value associated with that key, which can be data of different types.
[0079] The JSON data packet is parsed and reconstructed using a preset dynamic template engine to generate an enhanced prompt that incorporates the target elements, which serves as the target prompt.
[0080] The target elements include the digital human's role setting parameters, the semantic description of the current driving scenario, the trust level quantification value, the human-machine collaborative control target constraints, the generation requirements, and the digital human's action-expression coupling descriptor.
[0081] like Figure 5 As shown, the reasoning module is used to output a digital human motion adjustment scheme and natural language dialogue content based on the target prompt and the current driving scenario.
[0082] When the inference module outputs the digital human motion adjustment scheme and natural language dialogue content, it includes the following steps:
[0083] The system invokes predefined API interfaces to create large models adapted to driving scenarios. Here, API stands for Application Programming Interface, which is an interface that allows different software components to interact, typically providing a set of functions for accessing external services or data. In this embodiment, the large models include allama3, gpt-4o, and deepseek.
[0084] By combining a pre-set knowledge base, the large model is driven to perform joint reasoning to generate interactive responses that include digital human motion adjustment schemes and natural language dialogue content.
[0085] The generated digital human motion adjustment scheme and natural language dialogue content are processed by the processing module, then encapsulated by the TCP / IP protocol stack, and transmitted to the vehicle HMI system via the vehicle Ethernet bus.
[0086] The knowledge base includes semantic understanding rules for driving scenarios, a mapping table of digital human motion feature parameters, and a multimodal interaction strategy database. The digital human motion adjustment scheme includes at least facial expression parameters, limb movement commands, and speech synthesis feature vectors.
[0087] like Figure 2As shown, when the trust level is excessive trust, the inference module uses trust suppression as the control target when outputting the digital human action adjustment scheme and natural language dialogue content; when the trust level is insufficient trust, the inference module uses trust enhancement as the control target when outputting the digital human action adjustment scheme and natural language dialogue content.
[0088] The processing module is used to convert the digital human motion adjustment scheme and natural language dialogue content into interactive instructions and output them to the digital human interaction module; the digital human interaction module displays interactive actions according to the interactive instructions.
[0089] The digital human interaction module includes a pre-installed Unity-based digital human. A digital human, or Digital Human, refers to a virtual character created using artificial intelligence technology that interacts with the user through human-like appearance, behavior, and voice. Unity refers to an existing cross-platform game engine that supports 2D, 3D, VR (virtual reality), and AR (augmented reality) application development, enabling high-quality dynamic interaction and immersive experiences for virtual humans.
[0090] The digital human establishes an interactive connection with the vehicle's cockpit system. When the interactive command includes cockpit interaction content, after the digital human finishes broadcasting, it sends a control command to the vehicle's cockpit system via the CAN bus or vehicle communication protocol, triggering response operations such as seat vibration, instrument panel display, air conditioning system, or windows. For example, during the trust adjustment process, if the driver lacks trust, the digital human can trigger seat vibration to soothe the driver; if the driver has excessive trust, it can open the windows or turn on the warning lights to remind the driver to be ready to take over at any time.
[0091] The digital human interaction module displays interactive actions based on interaction commands, including the following steps:
[0092] Upon receiving an interactive command, if the interactive command includes voice broadcast content, the system performs speech synthesis and lip-syncing processing using a locally deployed end-to-end deep learning model. In this embodiment, models such as GPT-SoVITS can be used to achieve high-quality speech generation and personalized voice synthesis.
[0093] Among them, based on interactive commands, the Wav2Lip or RTFace model is used to predict the mouth shape trajectory in real time on the audio stream, and generate lip movement data synchronized with the speech waveform for mouth shape alignment.
[0094] When the interaction command includes action content, the HumanTOMATO or EMAGE deep learning model is used to generate a sequence of digital human actions that match the interaction command in real time. This setting is not limited to pre-made animation libraries and enhances the naturalness, emotional expression, and immersive interaction of the digital human during voice interaction.
[0095] A streaming engine is used to dynamically segment the action sequence, decomposing continuous actions into data packets containing timestamps, joint rotation angles, and motion trajectory vectors according to a preset time window. This setup enables real-time inference and improves the real-time performance of the interaction.
[0096] Data packets are streamed to the Unity rendering engine via the UDP protocol, driving the digital human to achieve frame-synchronized rendering of body movements.
[0097] like Figure 6 As shown, after the aforementioned multimodal data is synchronously loaded into the Unity rendering engine, it can drive the digital human to perform real-time voice broadcasts, facial expressions, and body movements, achieving a believable virtual interactive avatar. Furthermore, in conjunction with the vehicle's cockpit system, it enables voice-driven intelligent in-vehicle interactive feedback, thus completing a full closed-loop control process from driver emotion recognition and language generation to digital human expression and cockpit response.
[0098] In this embodiment, based on the above operations, the Unity-based digital human can generate corresponding body movements and facial animations, as well as lip-syncing, and can dynamically adjust its interaction strategy according to the driver's level of trust, ensuring that the driver's emotions are properly guided.
[0099] For example, when a driver is in a state of distrust and needs to have their trust restored, the virtual human will use a gentle tone, such as "Rest assured, the system is monitoring the entire process to ensure your safety," and will reassure the driver with a gentle smile, friendly eye contact, and slow body language (such as gently raising a hand) to enhance their sense of security. When a driver is in a state of over-trust, the digital human's expression will be more serious, avoiding being overly friendly or relaxed, and emphasizing the need to remain alert. Its interactive commands may include frowning or displaying a tense expression to convey a stronger demand for attention; and gestures such as gripping the steering wheel, pointing to the steering wheel, or making a gesture of preparing to take over to remind the driver to remain alert; and its interactive commands may include a voice warning: "Please note that although the system is handling the situation, please keep your hands on the steering wheel so that you can take over immediately in an emergency."
[0100] This embodiment also provides a human-machine trust calibration method based on a large model-driven digital human. The method utilizes the aforementioned human-machine trust calibration system based on a large model-driven digital human to perform human-machine trust calibration, and includes the following steps:
[0101] The driver monitoring system acquires the driver's characteristic information in real time through its information acquisition module. The characteristic information includes facial features, gaze features, voice features, and physiological activity features.
[0102] The trust detection module uses a deep learning model to analyze the feature information to obtain the driver's current trust level.
[0103] The prompt generation module generates a target prompt based on the current driving scenario and the trust level; the target prompt is input into the inference module, which outputs a digital human motion adjustment scheme and natural language dialogue content adapted to the current scenario.
[0104] The processing module converts the digital human motion adjustment scheme and natural language dialogue content into interactive commands that can drive the digital human; the digital human interaction module of the vehicle HMI system executes the interactive commands and displays the digital human's interactive actions and dialogue content in real time.
[0105] This embodiment provides a human-machine trust calibration method and system based on a large model-driven digital human, which has strong adaptability and adjustment capabilities. It can accurately adapt interaction strategies in diverse driving scenarios, ensuring that the driver maintains an appropriate sense of trust in different driving scenarios, thereby improving the safety and user experience of autonomous driving.
[0106] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.
Claims
1. A human-machine trust calibration system based on a large model-driven digital human, comprising an in-vehicle computing unit and an in-vehicle HMI system, characterized in that, It also includes a driver monitoring system; the driver monitoring system includes an information collection module and a trust detection module; The information acquisition module is used to collect the driver's characteristic information; The feature information includes facial features, gaze features, voice features, and physiological activity features; The trust detection module is used to perform trust detection based on feature information and a deep learning model, and output the driver's current trust level. The in-vehicle computing unit includes a prompt word generation module, an inference module, and a processing module; the in-vehicle HMI system includes a digital human interaction module. The prompt generation module is used to generate a target prompt based on the current driving scenario and the driver's current level of trust. The prompt generation module includes the following steps when generating the target prompt: It receives feature information, driving behavior feature vectors, vehicle environment perception data and trust level output by the trust detection module in real time, and constructs a structured JSON data packet containing four data dimensions: driver status, driving scenario, vehicle status and trust level. The JSON data packet is parsed and reconstructed using a preset dynamic template engine to generate an enhanced prompt that incorporates target elements, which serves as the target prompt. The target elements include the role setting parameters of the digital human, the semantic description of the current driving scenario, the trust level quantification value, the human-machine collaborative control target constraints, the generation requirements, and the digital human action-expression coupling descriptor. The reasoning module is used to output a digital human motion adjustment scheme and natural language dialogue content based on the target prompt and the current driving scenario; the trust level includes over-trust, appropriate trust and insufficient trust; When the trust level is excessive trust, the inference module uses trust suppression as the control target when outputting digital human action adjustment schemes and natural language dialogue content; when the trust level is insufficient trust, the inference module uses trust enhancement as the control target when outputting digital human action adjustment schemes and natural language dialogue content. The processing module is used to convert the digital human motion adjustment scheme and natural language dialogue content into interactive instructions and output them to the digital human interaction module; the digital human interaction module displays interactive actions according to the interactive instructions; In the digital human interaction module, a Unity-based digital human is pre-set; the digital human interacts with the user through anthropomorphic appearance, behavior and voice; the digital human establishes an interactive connection with the vehicle cockpit system; when the interaction command includes cockpit interaction content, after the digital human finishes broadcasting, a control command is sent to the vehicle cockpit system via CAN bus or vehicle communication protocol to trigger the response operation of seat vibration, instrument panel display, air conditioning system or windows; During the trust adjustment process, if the driver lacks trust, the seat vibration is triggered to soothe the driver; if the driver has excessive trust, the windows are opened or the warning lights are turned on to remind the driver to be ready to take over at any time.
2. The human-machine trust calibration system based on a large model-driven digital human according to claim 1, characterized in that, The information acquisition module includes a camera for acquiring facial and gaze features, a voice acquisition device for acquiring voice features, and a physiological sensor for acquiring physiological activity features.
3. The human-machine trust calibration system based on a large model-driven digital human according to claim 1, characterized in that, When the inference module outputs the digital human motion adjustment scheme and natural language dialogue content, it includes the following steps: Large models adapted to driving scenarios are called through predefined API interfaces; By combining a pre-set knowledge base, the large model is driven to perform joint reasoning to generate an interactive response that includes digital human motion adjustment schemes and natural language dialogue content; The generated digital human motion adjustment scheme and natural language dialogue content are processed by the processing module, then encapsulated by the TCP / IP protocol stack, and transmitted to the vehicle HMI system via the vehicle Ethernet bus. The knowledge base includes semantic understanding rules for driving scenarios, a mapping table of digital human motion feature parameters, and a multimodal interaction strategy database. The digital human motion adjustment scheme includes at least facial expression parameters, limb movement commands, and speech synthesis feature vectors.
4. The human-machine trust calibration system based on a large model-driven digital human according to claim 1, characterized in that, The digital human interaction module displays interactive actions based on interaction commands, including the following steps: Receive interactive commands. When the interactive commands contain voice broadcast content, perform speech synthesis and lip-sync processing through a locally deployed end-to-end deep learning model. Among them, based on interactive commands, the Wav2Lip or RTFace model is used to predict the mouth shape trajectory in real time on the audio stream, and generate lip movement data synchronized with the speech waveform for mouth shape alignment. When the interactive command contains action content, the HumanTOMATO or EMAGE deep learning model is used to generate a digital human action sequence that matches the interactive command in real time. The motion sequence is dynamically segmented using a streaming engine, and continuous motion is decomposed into data packets containing timestamps, joint rotation angles, and motion trajectory vectors according to a preset time window. Data packets are streamed to the Unity rendering engine via the UDP protocol, driving the digital human to achieve frame-synchronized rendering of body movements.
5. A human-machine trust calibration method based on a large model-driven digital human, characterized in that, The human-machine trust calibration system based on a large model-driven digital human as described in any one of claims 1-4 is used to perform human-machine trust calibration, comprising the following steps: The driver monitoring system acquires the driver's characteristic information in real time through its information acquisition module. The characteristic information includes facial features, gaze features, voice features, and physiological activity features. The trust detection module uses a deep learning model to analyze the feature information to obtain the driver's current trust level. The prompt generation module generates a target prompt based on the current driving scenario and the trust level; the target prompt is input into the inference module, which outputs a digital human motion adjustment scheme and natural language dialogue content adapted to the current scenario. The processing module converts the digital human motion adjustment scheme and natural language dialogue content into interactive commands that can drive the digital human; the digital human interaction module of the vehicle HMI system executes the interactive commands and displays the digital human's interactive actions and dialogue content in real time.
Citation Information
Patent Citations
System and method for improving driving credibility based on multi-modal intelligent interaction
CN118918558A
Adaptive trust calibration
CN115465283A