Human-computer interaction method and device based on passenger emotion recognition and storage medium

By identifying the emotional state of the occupants and predicting operational needs of historical data, and generating personalized dialogue content, the problem of lack of active perception and personalized interaction in the on-board system is solved, and intelligent and personalized on-board human-computer interaction is realized, improving user experience and security.

CN120353343APending Publication Date: 2025-07-22AISPEECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510560157.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing vehicle-mounted human-computer interaction system lacks active perception and personalized interactive services, resulting in the user experience being mechanized and passive, and the inability to effectively understand the user's emotional changes and needs.

Method used

By acquiring occupant sensing parameters to identify emotional states, combining occupant historical data to predict operational needs, generating personalized dialogue content and actively triggering human-computer interaction operations, using multimodal data fusion and deep learning models for sentiment analysis and demand prediction.

Benefits of technology

It realizes the active interaction of the on-board system, improves the level of intelligence and personalization, provides personalized services that are closer to user needs, reduces the operating burden of passengers, and enhances driving safety and comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353343A_ABST
    Figure CN120353343A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction method and device based on passenger emotion recognition and a storage medium, and the method comprises the steps: obtaining a passenger sensing parameter, and recognizing a passenger emotion state corresponding to the passenger sensing parameter; passenger historical data are obtained, the passenger operation requirement is predicted according to the passenger historical data and the passenger emotion state, and the passenger historical data comprise at least one of passenger historical interaction information, passenger attribute information, vehicle-mounted environment context information and the passenger emotion change trend; and generating personalized dialogue content based on the predicted passenger operation demand, and triggering man-machine interaction operation according to the personalized dialogue content. Therefore, by introducing passenger emotion recognition and historical data analysis and actively adjusting the vehicle-mounted interaction service content, active interaction of the vehicle-mounted man-machine interaction system is achieved, and the intelligent and personalized level of the vehicle-mounted man-machine interaction system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of in - vehicle human - machine interaction, and in particular, to a human - machine interaction method, device, storage medium, and program product based on occupant emotion recognition. Background Art

[0002] With the rapid development of intelligent and automated technologies, the in - vehicle human - machine interface (HMI) system has become an indispensable part of modern vehicles. Traditional in - vehicle HMI systems mainly rely on methods such as speech recognition, touchscreens, and buttons to interact with drivers and occupants, aiming to improve the convenience and safety of operations. In addition, some in - vehicle HMI systems can also respond to the operations of occupants based on preset conditions, such as voice assistants providing standard instructions or vehicle status feedback. However, these in - vehicle human - machine interaction methods all require interactive operations initiated by the occupants, lacking the ability to actively perceive the environment and user context, making the user experience seem rather mechanical and passive. Summary of the Invention

[0003] This application provides a human - machine interaction method, device, storage medium, and program product based on occupant emotion recognition to at least solve the problems of the lack of active perception ability and personalized interaction service methods in the current in - vehicle systems in related technologies.

[0004] In a first aspect, an embodiment of this application provides a human - machine interaction method based on occupant emotion recognition, including: obtaining occupant sensing parameters and identifying the occupant emotional state corresponding to the occupant sensing parameters; obtaining occupant historical data, and predicting the occupant operation requirements according to the occupant historical data and the occupant emotional state; the occupant historical data includes at least one of the following: occupant historical interaction information, occupant attribute information, in - vehicle environment context information, and occupant emotion change trend; generating personalized dialogue content based on the predicted occupant operation requirements, and triggering a human - machine interaction operation according to the personalized dialogue content.

[0005] In a second aspect, an embodiment of this application provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the human - machine interaction method based on occupant emotion recognition in any embodiment of this application.

[0006] In a third aspect, an embodiment of this application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the human - machine interaction method based on occupant emotion recognition in any embodiment of this application are implemented.

[0007] Fourthly, an embodiment of the present application provides a computer program product, including a computer program / instructions, which when executed by a processor implement the steps of the human-computer interaction method based on occupant emotion recognition according to any embodiment of the present application.

[0008] The beneficial effects of the embodiments of the present application are as follows: By actively sensing the emotional state of the occupant through the occupant sensing parameters, combining with the historical data of the occupant to predict and analyze the operation requirements of the occupant, generating personalized dialogue content accordingly, actively triggering the human-computer interaction operation by the system, and being able to actively adjust the service content according to the emotional state of the occupant and environmental changes, the initiative interaction of the in-vehicle system is realized, and the intelligence and personalization level of the in-vehicle human-computer interaction system are significantly improved. In addition, by introducing occupant emotion recognition and historical data analysis, the interaction operations actively triggered by the system can be closer to the user's needs and meet the needs of the occupant users for personalized interaction services. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 Shows an operation flowchart of an example of the human-computer interaction method based on occupant emotion recognition according to an embodiment of the present application; Figure 2 Shows an operation flowchart of an example of triggering a human-computer interaction operation according to personalized dialogue content according to an embodiment of the present application; Figure 3 Shows an operation flowchart of an example of recognizing the emotional state of the occupant corresponding to the occupant sensing parameters according to an embodiment of the present application; Figure 4 Shows an operation flowchart of an example of generating personalized dialogue content based on the predicted operation requirements of the occupant according to an embodiment of the present application; Figure 5 Shows a schematic diagram of an example of the system working process of an in-vehicle voice digital human system based on vision and emotion analysis according to an embodiment of the present application; Figure 6 Shows a schematic structural diagram of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0012] It should be noted that the demand of passengers for in-vehicle HMI systems is gradually shifting from the traditional "tool type" to the "experience type". Modern consumers expect cars to be not only a means of transportation, but also an intelligent device that provides personalization, context awareness, and emotional resonance. Therefore, in-vehicle HMI systems should also transform from a single operation interface to an intelligent interaction platform with emotional perception capabilities, which can provide more personalized and intelligent services according to the emotional states and needs of drivers or passengers.

[0013] In current related technologies, in-vehicle voice interaction systems mainly rely on speech recognition technology and natural language processing technology. However, these systems usually ignore the emotional states and behavior predictions of users, resulting in insufficiently intelligent and personalized user experiences. Among them, some in-vehicle voice assistants rely on speech recognition and natural language processing for single-round interactions, while some Advanced Driver Assistance Systems (ADAS) use vision recognition technology to monitor the facial expressions and attention of drivers. However, these systems often rely only on speech and visual data for emotion recognition and cannot effectively understand the emotional changes of users, resulting in insufficiently personalized interaction experiences.

[0014] In addition, most current related technologies rely on single-modal data (such as only speech or only vision), resulting in insufficient accuracy and comprehensiveness of emotion recognition. For example, traditional voice assistants can only analyze emotions based on speech content and do not consider the facial expressions of users. In addition, after recognizing emotions, most of these systems can only provide fixed feedback and cannot dynamically adjust according to changes in emotional states. Such a design leads to insufficient interaction experiences during the use of users and cannot meet personalized needs. It should also be noted that although some technologies attempt multi-modal data processing, they often stay at the level of information superposition and fail to deeply analyze and integrate different-modal data, resulting in insufficient information integration and intelligence levels. The problems caused by these defects have existed in the field of in-vehicle voice interaction for a long time. The demand of users for intelligent and personalized interactions is increasing, but this demand has not been effectively met at present. Therefore, it is of great significance to develop more advanced systems to solve these problems.

[0015] It should be understood that the purpose of the above description of the current related technology is only to facilitate the public's better understanding of the inventive spirit and motivation of the present application, and is not regarded as a limitation of the present application. In addition, the technical solutions described in the above current related technology are not prior art and may also be unpublished technical solutions, such as those under research or in the laboratory stage.

[0016] In view of the deficiencies in the above-mentioned current related technology under study, in the embodiments of the present application, by comprehensively considering the emotional state of the occupant, historical interaction information, and in-vehicle environment context, the system can flexibly adjust its interaction method and content, so that the interaction method can be adjusted according to emotional changes, realizing flexible and adaptable human-machine interaction. The in-vehicle HMI system can better adapt to the needs of different occupants and situational changes, improving the intelligence level of the system.

[0017] Figure 1 The operation flowchart of an example of a human-machine interaction method based on occupant emotion recognition according to an embodiment of the present application is shown.

[0018] As Figure 1 shown, in step S110, occupant sensing parameters are obtained, and the occupant emotional state corresponding to the occupant sensing parameters is identified.

[0019] In some embodiments, a variety of sensors are equipped in the vehicle to obtain the physiological and behavioral data of the occupant in real time, such as seat sensors, cameras, microphones, etc., to collect the physiological and behavioral data of the occupant in real time. In addition, the parameter dimensions of the occupant sensing parameters can also be diversified, including but not limited to facial expressions (analyzing facial movements through cameras), voice tones (capturing audio through microphones), body postures and gestures (capturing movements through seat sensors or in-vehicle cameras).

[0020] Furthermore, by using machine learning and deep learning models (such as convolutional neural networks, long short-term memory networks, etc.), the emotional state of the occupant is classified according to real-time sensing data (such as heart rate, facial expressions, voice features, etc.). Here, the types of emotional states can also be diversified, such as pleasure, anger, anxiety, fatigue, relaxation, etc. Thus, by analyzing the occupant's sensing parameters, the corresponding emotional changes are accurately identified, realizing the automatic perception of the occupant's emotions.

[0021] In step S120, occupant historical data is obtained, and the occupant operation requirements are predicted based on the occupant historical data and the occupant emotional state.

[0022] Here, the occupant historical data includes at least one of the following: occupant historical interaction information, occupant attribute information, in-vehicle environment context information, and occupant emotional change trend. Exemplarily, the occupant historical interaction information can be the occupant's past voice commands, touch screen operations, vehicle setting adjustments (such as temperature, seat position, in-vehicle entertainment system settings, etc.). The occupant attribute information can record the occupant's age, gender, driving habits, etc., which is used to create a user profile of the occupant. The in-vehicle environment context information can record the current driving mode (urban driving, highway driving, etc.), weather conditions, traffic flow, vehicle status, etc. The occupant emotional change trend can record the temporal change of the predicted occupant emotional state within a neighboring preset time period.

[0023] Furthermore, by combining the occupant's emotional state with the historical data, methods such as collaborative filtering and deep learning models are used to predict the operations that the occupant may need in the current emotional state. For example, if it is recognized that the occupant is in a low mood, the system may predict that they need to listen to some pleasant music or receive soothing language; if the occupant appears fatigued, the system can actively adjust the in-vehicle environment and provide a rest mode. Thus, by analyzing the occupant's user profile, historical behavior, emotional trend, in-vehicle environment context information, and current emotional state, the system can accurately predict the occupant's operation needs, thereby providing more considerate personalized services.

[0024] In step S130, personalized dialogue content is generated based on the predicted occupant operation needs, and a human-machine interaction operation is triggered according to the personalized dialogue content.

[0025] In some embodiments, according to the predicted occupant operation needs, the system creates personalized dialogue content through a Natural Language Generation (NLG) module, which can generate corresponding personalized dialogue content according to the occupant operation needs to actively ask the user whether to perform the corresponding human-machine interaction operation. In addition, the personalized dialogue content can also be generated by combining the current occupant emotional state and the occupant operation needs. Exemplarily, the intelligent generation of personalized dialogue content is achieved by leveraging large language model technology. Specifically, by filling information such as the predicted occupant operation needs and the occupant emotional state into the corresponding semantic slots of the prompt words, the automated generation of personalized dialogue content is realized.

[0026] In some examples, synthetic speech corresponding to the personalized conversation content can also be generated, and human-machine interaction operations can be triggered based on the synthetic speech. Exemplarily, speech with appropriate intonation, speech rate, and tone is generated according to the occupant's emotional state, and the synthetic audio can also be ensured to conform to the occupant's preferences based on the occupant attributes, thereby enhancing the personalization of the interaction. For example, when it is detected that the occupant is in a low mood, interaction prompts are provided in a gentler and soothing tone, such as "Do you need a break?" or "I can help you adjust the seat to relax."

[0027] Here, the operation modes for the system to trigger human-machine interaction operations based on the personalized conversation content can be diverse, such as automatic voice feedback or in-vehicle terminal interface display. In some cases, the system also needs to detect the occupant feedback information for the triggered human-machine interaction operations. When the occupant feedback information is confirmation, the corresponding human-machine interaction operations are controlled to be executed. Thus, by generating personalized conversation content and automatically triggering operations, the system is no longer just passively relying on the occupant's active input, and can actively initiate interactions based on the emotional state, greatly enhancing the intelligent experience and fluency of the in-vehicle interaction.

[0028] Through the embodiments of the present application, by actively perceiving the emotional state and needs of the occupant, automated and intelligent personalized services can be realized, the user experience can be improved, and driving safety and comfort can be enhanced. In addition, through emotion prediction and automatic response, the system can significantly reduce the operation burden of the occupant, make the interaction smoother and more natural, thereby greatly enhancing the intelligence, personalization, and responsiveness of the in-vehicle HMI system.

[0029] Figure 2 The operation flowchart of an example of triggering human-machine interaction operations according to the personalized conversation content according to the embodiments of the present application is shown.

[0030] As Figure 2 shown, in step S210, a digital human emotional state matching the occupant's emotional state is determined.

[0031] In some embodiments, the real-time emotional state of the occupant can be collected through an emotion analysis module, which can be an occupant emotion vector including multiple emotional dimensions (such as happy, sad, anxious, relaxed, etc.), and can be processed through smoothing filtering to reduce the noise in the real-time emotional fluctuations, making the emotion judgment more stable and coherent.

[0032] Then, based on the psychological scale (such as Prouvé's emotional scale, emotional circle model, etc.), the emotional vector is input into the predefined emotion mapping engine. Here, the emotion mapping engine is responsible for discretizing the continuous emotional space into 12 basic emotions, such as happiness, sadness, anger, surprise, calmness, etc., and calculating the most suitable composite emotion through a dynamic weight allocation algorithm. For example, if the system detects that the user's anxiety index exceeds the set threshold, the "active care" emotion of the digital human will be triggered first.

[0033] In addition, if the system identifies a high level of anxiety during the emotional analysis process (such as anxiety or tension during driving), it will not only trigger the "active care" emotion, but also link other in-vehicle systems (such as navigation system, speed monitoring, etc.) to determine whether it is necessary to superimpose "route guidance" auxiliary emotions. For example, when anxious, the system may automatically trigger gentle voice prompts and combine with the navigation system to give clear route guidance to relieve the user's anxiety.

[0034] In step S220, a digital human image with a digital human emotional state is rendered, and a human-computer interaction operation is triggered according to the personalized dialogue content.

[0035] Specifically, a hybrid mode of pre-recording and parameter-driven can be adopted. For example, by pre-recording core action animation streams, based on the Unity engine, the system uses 50 pre-made core action animation streams (such as smiling, frowning, stroking hands, relaxing, etc.). These animations cover a variety of emotional expressions and can be flexibly switched under different emotions. In addition, due to the different screen sizes and display viewing angles of different models, the system can also use skeletal animation redirection technology to adapt these pre-recorded animations to display devices of different models to achieve universality and consistency across models.

[0036] In real-time interaction, the system calls the corresponding basic animation flow according to the matching results of the emotion tags, and adjusts the digital human's facial expressions, body movements, eyes, etc. For example, when the system recognizes that the occupant is anxious, the digital human will show a soothing and caring expression (such as a smile, soft eyes, and gentle tone). If the emotion is happy or calm, the digital human will appear more active and relaxed.

[0037] In addition, based on the digital human's emotional rendering results, the system will trigger corresponding human-computer interaction operations according to the personalized dialogue content. For example, when the digital human shows a specific emotion (such as "active care" emotion), the system will give a voice prompt through synthesized voice, and the voice is synchronized with the digital human's expression to enhance the naturalness of the interaction. For example: "You seem a little anxious, take a deep breath, I have adjusted the temperature in the car for you, you will be more comfortable." At the same time, the system automatically adjusts the temperature in the car.

[0038] It should be understood that the interaction between the digital human image and the emotional state is not static. During the interaction process, the system dynamically adjusts the emotional expression of the digital human according to the feedback of the occupant (such as voice commands, expressions, or emotional changes). If the system finds that the occupant is still feeling uneasy or dissatisfied, the digital human can further comfort and encourage by updating the intonation, adjusting body movements, or even adding voice. Thus, the digital human emotion rendering technology combined with personalized dialogue enables the system to provide more natural emotional feedback at the visual and language levels. The interaction felt by the occupant is not just a mechanical response, but a "communication" with emotional resonance and interaction.

[0039] Through the embodiments of the present application, by combining emotion analysis and digital human emotion rendering technology, it is not only possible to accurately identify the emotional state of the occupant, but also to provide a personalized and emotional interaction experience through the image and voice feedback of the digital human. Based on the hybrid mode of the emotion mapping engine and the pre-recorded animation stream, the system can achieve consistent display on different vehicle models, and through real-time emotion-driven dynamic rendering and feedback adjustment, each interaction becomes more natural, smooth, and full of emotional depth.

[0040] In some examples of the embodiments of the present application, a multi-modal data fusion processing mechanism is selected for the recognition of the occupant's emotional state, that is, the occupant sensing parameters adopt multi-modal data, which can include the occupant monitoring image and the occupant voice data.

[0041] Figure 3 An operation flowchart showing an example of recognizing the occupant's emotional state corresponding to the occupant sensing parameters according to the embodiments of the present application is shown.

[0042] As Figure 3 shown, in step S310, the occupant visual features corresponding to the occupant monitoring image are extracted, and the occupant acoustic features corresponding to the occupant voice data are extracted. The occupant visual features include the occupant expression features and / or the occupant posture and movement features.

[0043] Regarding the description of the occupant visual features, the facial images of the occupant are collected in real time through an in-vehicle camera or an external sensor, and facial expression recognition technology (such as a facial feature detection algorithm based on the convolutional neural network CNN) is applied to extract the facial expression features of the occupant from them, which can include the raising / lowering of the eyebrows, the rising / falling of the corners of the mouth, the degree of narrowing of the eyes, etc., so as to recognize the corresponding facial expressions. The posture and movement features can be based on a deep learning model, such as OpenPose, to perform real-time analysis on the posture and movements of the occupant, for example, to identify the positioning and movement of human joints, and then to judge whether the occupant is in a tense or relaxed state (such as yawning, supporting the head, etc.).

[0044] Description of occupant acoustic features, which can collect the voice data of the occupant through a microphone array in the vehicle, extract features including pitch, volume, speech rate, tone, speech intensity, etc., and can reflect the emotional fluctuations of the occupant. For example, fast and high-pitched speech may indicate anxiety or excitement, while slow and low-pitched speech may indicate fatigue or frustration. In addition, the MFCC (Mel-Frequency Cepstral Coefficients) algorithm can also be applied to extract acoustic features such as frequency features and energy distribution in the speech signal, so as to identify the emotional state of the occupant (such as anxiety, pleasure, calm, etc.).

[0045] In step S320, the occupant visual features and the occupant acoustic features are fused to generate corresponding multi-modal occupant perception features.

[0046] Here, the methods of multi-modal feature fusion can be diverse, such as feature concatenation or multi-modal neural networks. In feature concatenation, the feature vectors extracted from vision and speech are directly concatenated together to form a higher-dimensional feature vector. In multi-modal neural networks, the visual features and acoustic features are fed into the multi-modal neural network, and the correlation between the two modalities is learned through shared network layers to generate the fused emotion features.

[0047] In step S330, the multi-modal occupant perception features are processed based on the sentiment analysis model to determine the corresponding occupant emotional state.

[0048] Here, the model types of the sentiment analysis model can be diverse, such as support vector machine (SVM), convolutional neural network (CNN), or deep neural network (DNN), etc. Further, the corresponding emotion label probability distribution (such as anxiety, pleasure, fatigue, anger, etc.) is output through the sentiment analysis model, which expresses the possibility of each emotion through probability values. For example, "anxiety: 80%", "pleasure: 10%", "calm: 5%", etc., and also provides a basis for the analysis of compound emotion states, such as the matching of the digital human emotion state. Thus, the system can quickly and efficiently identify the emotional state of the occupant, ensuring the timeliness and accuracy of the emotional feedback.

[0049] Through the embodiments of the present application, by combining visual and acoustic feature extraction, feature fusion, and the sentiment analysis model, accurate emotion recognition based on multi-modal data is realized. The system can efficiently extract and process information from different sensors, realize the real-time perception of the occupant's emotional state during driving, and provide more comprehensive and accurate emotional feedback.

[0050] Regarding the extraction details of the occupant acoustic features in step S310, in some examples of the embodiments of the present application, a bidirectional LSTM is used to model the temporal variation of the occupant voice data to extract the corresponding Mel cepstral coefficients, spectrogram features, and intonation features.

[0051] Specifically, the collected voice signal is segmented (such as segmented by frame, and the duration of each frame is about 20 ms), and feature extraction is performed on each frame of the signal. The voice signal is converted into a frequency-domain signal, and a Mel frequency filter bank is used to extract the Mel spectrogram, and the cepstral coefficients are calculated from the Mel spectrogram. In addition, by performing a short-time Fourier transform (STFT) or a continuous wavelet transform (CWT) on the voice signal, a corresponding spectrogram is generated to capture the high-frequency and low-frequency information in the voice. Intonation (Pitch) is an important emotional feature in speech, which can be achieved by capturing the pitch (i.e., the fundamental frequency of the speech) of each frame.

[0052] Here, a bidirectional LSTM temporal modeling is used to model the temporal dependence in the sequence data. Exemplarily, the extracted Mel cepstral coefficients, spectrogram, and intonation features are used as inputs and fed into a bidirectional LSTM model for modeling. The LSTM captures the long-term dependence information in the voice through memory units, and the bidirectional LSTM can capture the temporal information from both the forward and backward directions, enhancing the modeling ability of emotional changes. Especially in complex emotional states (such as anxiety, tension, etc.), it can better capture the fluctuations in the voice and accurately capture the temporal information of emotional changes in the voice.

[0053] In addition, based on the BERT model, the voice semantic emotion features corresponding to the occupant voice data are extracted.

[0054] Specifically, first, the voice of the occupant is converted into text by using speech recognition (ASR) technology. For example, a speech recognition model based on deep learning is used for high-precision speech transcription, ensuring that the semantic information in the voice is correctly extracted and converted into text. Then, the BERT model is used for semantic emotion analysis. The BERT model can be a model fine-tuned with emotion label data (such as the BERT model - Emotion), which can understand and analyze the semantic and context information in the text. Through its bidirectional attention mechanism, it understands the relationship between each word in the sentence and the surrounding words, thereby identifying the emotional tendency in the text (such as pleasure, anxiety, anger, etc.).

[0055] Furthermore, based on the Mel cepstral coefficients, spectrogram features, intonation features, and voice semantic emotion features, the occupant acoustic features are determined.

[0056] Specifically, the Mel cepstral coefficients, spectrogram features, intonation features, and speech semantic emotion features are fused, such as feature concatenation, to obtain the occupant acoustic features. Thus, the occupant acoustic features comprehensively reflect various emotional information in the speech signal, including multiple levels of information from the speech quality, pitch to semantics, ensuring high accuracy in recognizing the occupant's emotional state.

[0057] Through multi-level feature extraction and fusion based on the bidirectional LSTM and BERT models, the system can accurately capture the emotional changes in the occupant's speech. The combination of Mel cepstral coefficients, spectrograms, intonation features, and semantic emotion features extracted by BERT greatly improves the accuracy and meticulousness of emotional analysis.

[0058] In some embodiments, based on a fusion module adopting the Transformer architecture, the occupant visual features and the occupant acoustic features are jointly modeled in sequence to generate corresponding multi-modal occupant perception features.

[0059] Here, the Transformer architecture can effectively capture the global dependencies between different modalities through the self-attention mechanism. Specifically, in multi-modal emotion recognition, the Transformer architecture can process the feature vectors obtained from visual and speech feature extraction, learn the correlations between them, and thus generate unified multi-modal occupant perception features.

[0060] Exemplarily, after aligning the visual feature sequence and the acoustic feature sequence along the time axis, they are input into the Transformer model. The encoder part of the Transformer models the time series relationships in these multi-modal data through the self-attention mechanism, enabling the model to learn the correlations between different modalities (for example, the pleasure in speech may correspond to a smile on the face and squinting eyes).

[0061] Thus, the multi-modal fusion module based on the Transformer architecture can efficiently integrate the occupant's visual and acoustic features, effectively capture the complex dependencies between visual and acoustic features, enabling the system to not only recognize single-modal emotional information but also fuse multi-modal data to provide more comprehensive and accurate emotional judgments. In addition, based on sequence joint modeling and capturing the temporal correlations and interactions between them through the self-attention mechanism, the Transformer can better understand and model the emotional changes of the occupant at different time points, enhancing the system's perception ability in complex emotional states.

[0062] Regarding the details of predicting the occupant operation requirements in step S120, in some examples of the embodiments of the present application, the occupant historical data and the occupant emotional state are processed based on a demand prediction model to predict the corresponding occupant operation requirements. The demand prediction model adopts a hybrid model combining Transformer and RNN.

[0063] Specifically, first, the input data (including occupant historical interaction information, emotional state, environmental context information, etc.) are respectively input into Transformer and RNN. For historical interaction data and emotional state data, the self-attention mechanism of Transformer converts them into globally dependent feature representations; while for vehicle environment information and short-term emotional fluctuations, RNN can handle the temporal changes of these data and capture short-term dependencies.

[0064] Then, the features processed by Transformer and RNN are fused to form a comprehensive multi-dimensional feature representation, which contains an understanding of the occupant's needs from both global and local perspectives. After fusing the features, operation requirement prediction is performed through a fully connected layer (or multiple fully connected layers), such as may include adjusting the in-vehicle environment (such as air conditioning, seats), operating the in-vehicle entertainment system, adjusting the navigation system, etc. The final output is the possible operation requirements of the occupant, such as whether to adjust the in-vehicle temperature, whether to request the voice assistant for help, or whether to trigger certain automation systems (such as autonomous driving assistance, lane keeping, etc.). In addition, the system can also rank the possibilities of different requirements through probability output to help determine the most likely operation requirements.

[0065] Thus, by combining the global modeling ability of Transformer and the short-term dependence processing ability of RNN, the demand prediction model can comprehensively understand the changes in the occupant historical data and emotional state, so as to accurately predict the occupant's operation requirements under various environmental conditions, improving the accuracy of the demand prediction results.

[0066] Figure 4 The operation flow chart shows an example of generating personalized dialogue content based on the predicted occupant operation requirements according to the embodiments of the present application.

[0067] As Figure 4 shown, in step S410, the predicted occupant operation requirements, the occupant emotional state, the occupant attribute information, and the vehicle-mounted knowledge graph are input into the large language model to determine the corresponding at least one candidate dialogue text.

[0068] Here, the in-vehicle knowledge graph can contain various types of information about the vehicle (such as in-vehicle equipment, functions, status, etc.) and usage instructions for the in-vehicle system. It not only provides background information based on the in-vehicle environment but also helps generate conversation content related to the in-vehicle system. For example, certain operations may need to guide the occupant on how to use the in-vehicle navigation or provide real-time feedback on the vehicle's status.

[0069] Furthermore, by providing these input information to large language models (such as GPT series or Deepseek series), these input information will help the model understand the occupant's personalized needs and emotional state. The large language model will generate multiple candidate conversation texts based on the context and can select the most appropriate conversation content according to the specific situation and requirements.

[0070] In addition, the large model can also adopt Retrieval-augmented Generation (RAG) technology to enhance the generation accuracy of the conversation content. For example, by retrieving relevant knowledge in the database (such as the occupant's historical behavior, common statements, vehicle status, etc.), the system can introduce more accurate and relevant information, enhancing the reliability and relevance of the generated content. For example, if the occupant often requests a break during long drives, relevant services (such as rest suggestions, music selection, etc.) can be integrated into the current conversation by retrieving similar interactions in the historical data.

[0071] In step S420, the in-vehicle environment context information, the occupant's emotional state, and each candidate conversation text are input into the multi-strategy decision maker to determine the personalized conversation content from each candidate conversation text.

[0072] Here, the multi-strategy decision maker will screen out the most suitable one from the candidate conversation texts by comprehensively considering the in-vehicle environment context information, the occupant's emotional state, and the suitability of the candidate conversation text, and determine the final personalized conversation content.

[0073] Exemplarily, the multi-strategy decision maker can adopt multiple strategies (such as emotional strategy, safety strategy, entertainment strategy, etc.) to evaluate the candidate conversation texts, weigh various factors, and determine which candidate text best meets the multiple needs in the current situation. Exemplarily, for the adaptation of emotional needs, the decision maker will preferentially select those conversation texts that can effectively adjust the occupant's mood. For example, if the occupant shows anxiety, the decision maker may select a milder and more soothing text, while if the occupant shows pleasure, the system may select a more exciting or encouraging text.

[0074] Through the multi-strategy decision-making device, it is possible to comprehensively consider the vehicle environment, the emotions of the occupants, and personalized needs, and quickly screen out the most suitable dialogue content from multiple candidate texts. The system screens out the most suitable one from multiple candidate texts to generate the final personalized dialogue content, which not only meets the needs of the occupants but also can be intelligently adjusted according to the emotional state and vehicle environment, making the interaction with the vehicle system more natural and cordial.

[0075] Figure 5 FIG. shows a schematic diagram of the system working process of an example of a vehicle-mounted voice digital human system based on visual and emotional analysis according to an embodiment of the present application.

[0076] In the embodiment of the present application, by combining vision, voice, and emotional analysis, comprehensive understanding and prediction of the user's emotions and behaviors are achieved through multi-modal data fusion. The biggest difficulty lies in the complexity of real-time collection and analysis of emotional data, as well as how to effectively fuse multiple modal data to provide accurate feedback.

[0077] The vehicle-mounted voice digital human system based on visual and emotional analysis provided by the embodiment of the present application aims to improve the intelligence and personalization level of vehicle-mounted voice interaction. The system combines the user's facial expressions, body movements, and emotional states to provide dynamically adjusted dialogue and services.

[0078] It should be noted that currently, traditional vehicle-mounted voice assistants and advanced driver assistance systems have obvious defects in user interaction: they cannot perceive and adapt to the user's emotions in real time, resulting in insufficient user experience. To address this problem, we decided to adopt a multi-modal data fusion method to comprehensively understand the needs and emotions of users from multiple dimensions.

[0079] As Figure 5 shown, first, the user activates the system through a voice command or a touch screen. When the user activates the system through a voice command or a touch screen, the system will perform an environmental self-check at startup, including camera calibration, microphone sensitivity adjustment, and in-vehicle environmental noise assessment, to ensure that the sensors are in the best working condition.

[0080] Then, the system captures the user's expressions and actions through the camera, and at the same time collects the user's voice data. The user's voice data is collected using a multi-microphone array, and beamforming technology is used to eliminate environmental noise interference to ensure the accuracy of data collection.

[0081] Then, the visual recognition module processes the facial expression and motion data. The visual recognition module preprocesses the collected video frames (such as denoising, normalization, and illumination correction). After face detection, an emotion classifier is used to recognize emotions in the cropped facial region. At the same time, body key points are extracted, and then the temporal motion data is analyzed. Finally, a multi-modal fusion of facial expression and motion features is achieved using the Transformer architecture. The emotion analysis module processes the speech data. By extracting the Mel-Frequency Cepstral Coefficients (MFCC) and spectrogram features of the speech data, local emotion features are extracted using a CNN, and a bidirectional LSTM is used to model the temporal changes in speech. Combining the text information converted by Automatic Speech Recognition (ASR), the semantic emotion is enhanced by analyzing the semantic emotion using the BERT-Emotion model. Finally, the features are fused using the Transformer to identify the user's emotional state.

[0082] Thus, the system can comprehensively analyze the user's facial expressions and speech emotional states, improving the ability to accurately recognize the user's emotions. This enables the system to more naturally respond to the user's emotional changes during interaction, thereby enhancing the personalized experience of the interaction.

[0083] Then, the behavior prediction module combines the user's historical data and current emotional state to predict the user's needs. The behavior prediction module integrates the user's historical data (including previous voice commands, in-vehicle environment information, user daily habits, and emotional change trends) and the current emotional state, and stores and extracts the user's long-term behavior patterns through hierarchical memory units. A sequence modeling method combining Transformer and RNN is adopted, and the prediction decision is continuously optimized through reinforcement learning, so as to achieve accurate prediction of the user's needs. For example, when the user is detected to be tired, soothing music or rest reminders are actively recommended; when an urgent emotion is recognized, the in-vehicle temperature is automatically adjusted or relaxation suggestions are provided; or when the user frequently queries navigation information, real-time traffic conditions updates and route optimizations are provided.

[0084] Thus, by combining historical data and the current emotional state, the system can predict the user's needs in advance. This function not only improves the intelligence level of the interaction but also makes the user feel more considerate services. For example, when the user shows fatigue, the system can actively provide rest suggestions or adjust the in-vehicle environment to improve comfort.

[0085] Then, the voice interaction module generates personalized dialogue content based on the prediction results. This module uses the AISpeech DFM-2 large language model and combines the vehicle knowledge graph and user portrait vector to generate personalized dialogue content through a multi-strategy decision maker. The user's current emotional state and historical data are added to the input, and the retrieval enhancement generation technology is used to ensure that the generated content is close to user needs. TTS technology is added to convert text into fluent and emotional voice output, and finally the digital human based on 3D modeling and facial animation drives actively initiates interaction with the user.

[0086] More specifically, the voice interaction module contains two functional units: one for generating personalized conversation content, and the other for generating digital human emotional expressions that match the user's emotional state. The part that generates conversation content uses a large language model (e.g., AISpeech DFM-2) combined with human feedback reinforcement learning to customize responses by inputting the user's current emotional state and historical data, while the part that generates digital human emotions is based on the real-time emotional judgment results provided by the emotion analysis module, and maps the results to the emotional parameters of the digital human after weighted fusion and smoothing filtering, thereby determining the specific emotional state of the digital human, such as happiness, surprise, calmness, etc.

[0087] In terms of emotion determination, the system inputs the emotion vector output by the emotion analysis module into the predefined emotion mapping engine, which discretizes the continuous emotion space into 12 basic emotions based on the psychological scale, and matches the most suitable composite emotion through a dynamic weight allocation algorithm. In the specific processing, when it is detected that the user's anxiety index exceeds the threshold, the system will first trigger the "active care" emotion, and link the navigation system data to determine whether it is necessary to superimpose the "route guidance" auxiliary emotion.

[0088] Emotion rendering uses a hybrid mode of pre-recording and parameter-driven: 50 core action animation streams pre-made based on the Unity engine are adapted to the screen size and viewing angle of different models through skeletal animation redirection technology. During real-time interaction, the system calls the corresponding basic animation stream according to the matched emotion tag.

[0089] It should be noted that the in-vehicle interactive systems in the current related technologies are mostly one-way feedback and lack emotional understanding. However, the system provided by the embodiment of the present application can dynamically adapt to the user's emotional changes through the real-time fusion of multimodal data and provide personalized and intelligent services. This interactive method makes the user experience more smooth and natural, greatly improving the practicality and intelligence level of the in-vehicle voice system.

[0090] Through the in-vehicle voice digital human system based on vision and emotion analysis provided by the embodiments of the present application, a comprehensive understanding of the user's emotions and behaviors is achieved, the system can quickly adapt to changes in user needs, improve the fluency and intelligence of interaction, reduce the user's waiting time, and enhance user satisfaction. In addition, the digital human not only responds passively to user instructions, but also can initiate conversations actively and provide personalized suggestions. This function improves the user's sense of participation, makes the interaction more vivid and user-friendly, and enables vehicle occupants to establish an emotional connection with the in-vehicle interaction system better.

[0091] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0092] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above-mentioned human-computer interaction methods based on occupant emotion recognition.

[0093] In some embodiments, the embodiments of the present application further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute any one of the above-mentioned human-computer interaction methods based on occupant emotion recognition.

[0094] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the human-computer interaction method based on occupant emotion recognition.

[0095] Figure 6 is a schematic hardware structure diagram of an electronic device for executing the human-computer interaction method based on occupant emotion recognition provided by another embodiment of the present application. As Figure 6 shown, the device includes: One or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example.

[0096] The device for executing the human-machine interaction method based on occupant emotion recognition may further include: an input device 630 and an output device 640.

[0097] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected through a bus or other means, Figure 6 Taking connection through a bus as an example.

[0098] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the human-machine interaction method based on occupant emotion recognition in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the human-machine interaction method based on occupant emotion recognition in the above method embodiments.

[0099] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0100] The input device 630 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.

[0101] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, execute the human-machine interaction method based on occupant emotion recognition in any of the above method embodiments.

[0102] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0103] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0104] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.

[0105] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices.

[0106] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.

[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0109] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A human-machine interaction method based on occupant emotion recognition, comprising: Obtaining occupant sensing parameters and recognizing the occupant emotion state corresponding to the occupant sensing parameters; Obtaining occupant historical data and predicting the occupant operation requirements according to the occupant historical data and the occupant emotion state; the occupant historical data includes at least one of the following: occupant historical interaction information, occupant attribute information, vehicle-mounted environment context information, and occupant emotion change trend; Generating personalized dialogue content based on the predicted occupant operation requirements and triggering a human-machine interaction operation according to the personalized dialogue content.

2. The method according to claim 1, wherein The triggering of the human-machine interaction operation according to the personalized dialogue content includes: Generating a synthesized voice corresponding to the personalized dialogue content and triggering a human-machine interaction operation according to the synthesized voice.

3. The method according to claim 1 or 2, wherein The triggering of the human-machine interaction operation according to the personalized dialogue content includes: Determining a digital human emotion state matching the occupant emotion state; Rendering a digital human image with the digital human emotion state and triggering a human-machine interaction operation according to the personalized dialogue content.

4. The method according to claim 1, wherein, The occupant sensing parameters include occupant monitoring images and occupant voice data, The recognizing of the occupant emotion state corresponding to the occupant sensing parameters includes: Extracting the occupant visual features corresponding to the occupant monitoring images and extracting the occupant acoustic features corresponding to the occupant voice data; the occupant visual features include occupant expression features and / or occupant posture movement features; Fusing the occupant visual features and the occupant acoustic features to generate corresponding multi-modal occupant perception features; Processing the multi-modal occupant perception features based on an emotion analysis model to determine the corresponding occupant emotion state.

5. The method according to claim 4, wherein The extracting of the occupant acoustic features corresponding to the occupant voice data includes: Modeling the speech time series change of the occupant voice data based on a bidirectional LSTM to extract corresponding Mel cepstral coefficients, spectrogram features, and intonation features; Extracting the speech semantic emotion features corresponding to the occupant voice data based on a BERT model; Determining the occupant acoustic features according to the Mel cepstral coefficients, spectrogram features, intonation features, and the speech semantic emotion features.

6. The method according to claim 4, wherein, The fusing of the occupant visual features and the occupant acoustic features to generate corresponding multi-modal occupant perception features includes: Based on a fusion module adopting a Transformer architecture, jointly modeling the occupant visual features and the occupant acoustic features in a sequence to generate corresponding multi-modal occupant perception features.

7. The method according to claim 1, wherein, The predicting of the occupant operation requirements according to the occupant historical data and the occupant emotion state includes: Processing the occupant historical data and the occupant emotion state based on a demand prediction model to predict the corresponding occupant operation requirements; the demand prediction model adopts a hybrid model combining Transformer and RNN.

8. The method according to claim 1, wherein The generating of the personalized dialogue content based on the predicted occupant operation requirements includes: Inputting the predicted occupant operation requirements, the occupant emotion state, occupant attribute information, and a vehicle-mounted knowledge graph into a large language model to determine at least one corresponding candidate dialogue text; Input the in-vehicle environment context information, the occupant emotional state, and each of the candidate dialogue texts into a multi-strategy decision maker to determine personalized dialogue content from each of the candidate dialogue texts.

9. A storage medium having a computer program stored thereon, wherein, When executed by a processor, the program implements the steps of the method according to any one of claims 1-8.

10. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method according to any one of claims 1-8.

Citation Information

Cited By

  • Real-time emotion perception and voice interaction system for intelligent cockpit

    CN121009400A

  • Intelligent cockpit-oriented real-time emotion perception and voice interaction system

    CN121009400B

  • Active intelligent dialogue pushing method and system based on user state perception

    CN121924168A

  • A user state perception-based active intelligent conversation pushing method and system

    CN121924168B