Intelligent conversation interaction system and method based on facial recognition and electroencephalogram

By acquiring user information in real time through facial recognition and EEG/EEG technology, identifying user status and needs, and generating personalized interactive responses, this technology solves the problem of the inability to effectively identify user status and needs in existing technologies, thereby improving the intelligence of the interactive system and the user experience.

CN119828888BActive Publication Date: 2025-11-11HUAXIA DIGITAL INTELLIGENCE HOLDINGS (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411892731.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-11-11
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing technologies fail to effectively identify user states and determine user needs by acquiring user modal information, and cannot generate interactive responses that combine user states and needs to optimize interaction.

Method used

An intelligent conversation interaction system based on facial recognition and EEG-cortical conductance is adopted. The system acquires user information in real time through a modality acquisition module, identifies user status through a state recognition module, judges user needs through a speech recognition module, and generates interactive responses through a conversation interaction module, dynamically adjusting coefficients to optimize the interaction strategy.

Benefits of technology

It achieves accurate identification of user status and needs, generates personalized interactive responses, improves user engagement and experience, and optimizes the intelligence level of the interactive system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828888B_ABST
    Figure CN119828888B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent conversation interaction technology, and discloses an intelligent conversation interaction system and method based on facial recognition and EEG / TEV. The system utilizes a modality acquisition module to acquire user modal information in real time via a sensor network, corrects this information based on environmental changes, and synchronizes the corrected modal information to comprehensively reflect changes in user state. A state recognition module extracts features from the modal information and identifies the user state using a state recognition strategy. This strategy includes calculating the user's response value and comparing it with a user response threshold to accurately identify the user's current state. A speech recognition module recognizes the user's speech and determines their needs, providing a deep understanding of the user's needs in different states. A conversation interaction module generates interactive responses based on the user's state and needs, and collects user feedback on these responses to optimize the interactive responses and the state recognition strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent conversation interaction technology, and in particular to an intelligent conversation interaction system and method based on facial recognition and electroencephalography (EEG) and electrodermal transmission. Background Technology

[0002] Currently, researchers are dedicated to fusing facial recognition with various physiological signals such as EEG and EEG to improve the accuracy of physiological recognition. With the help of advanced algorithms and hardware, they aim to achieve real-time monitoring and recognition of physiological signals, providing users with timely feedback and interaction. However, most studies have not yet solved the problems of how to identify user states by acquiring user modal information, how to recognize user speech to determine user needs, and how to combine user states and needs to generate and continuously optimize interactive responses.

[0003] For example, Chinese patent CN113488057B discloses a dialogue implementation method and system for elderly care, addressing the technical problem of how to simulate a unique personal speaking style and tone of voice as much as possible through voice interaction technology, thereby improving the quality of life of widowed elderly people and alleviating the pain suffered by their children who have lost a loved one. The technical solution is as follows: S1, record the conversation between the two parties using a recording device; S2, convert the recorded voice into text and proofread it; S3, input the prepared corpus into a dialogue model for training, and output a personalized dialogue model; S4, use the existing voice corpus to create a speech synthesis model with personal intonation characteristics. The system includes a dialogue model generation unit and a personalized speech synthesis unit; the dialogue model generation unit includes a dialogue recording acquisition module, a speech-to-text module, a text processing and proofreading module, and a model training module; the personalized speech synthesis unit includes a voiceprint encoder, a speech synthesizer, and a voice generator.

[0004] For example, Chinese patent CN108021703B discloses a conversational intelligent teaching system, including interconnected knowledge base units and functional units. The functional units include an input preprocessing module, an answer reasoning module, an evaluation module, and a dialogue management module, all interconnected. The knowledge base unit includes a domain ontology, interaction templates, a semantic dictionary, and a student model. This conversational teaching system utilizes a domain ontology with well-defined semantic relationships and hierarchical structure to model the system's domain knowledge. Through the hierarchical structure and semantic relationships of concepts, it provides the system with reasoning-based knowledge. Furthermore, it proposes an ontology-driven dialogue management mechanism that can quickly generate dialogue content and sequences.

[0005] The aforementioned patents suffer from the problems described in this background: The dialogue implementation methods described above record the conversations of both parties using a data acquisition device, convert the recorded audio into text, proofread it, and then input the processed corpus into a dialogue model for training, outputting a personalized dialogue model. They also utilize existing audio corpus to create a speech synthesis model with individual intonation characteristics, thereby simulating a unique personal speaking style and tone of voice. The aforementioned conversational intelligent teaching system models the system's domain knowledge through a domain ontology with good semantic relationships and hierarchical structure, providing reasoning-based knowledge to the system through the hierarchical structure of concepts and speech relationships. However, the two aforementioned patents do not address how to identify user states by acquiring user modal information, recognize user speech to determine user needs, and generate and continuously optimize interactive responses based on user states and needs. To solve this problem, this invention proposes an intelligent conversational interaction system and method based on facial recognition and EEG / TEG. Summary of the Invention

[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0007] In view of the problems existing in the above-mentioned intelligent conversation interaction systems and methods based on facial recognition and EEG-cortex, the present invention is proposed.

[0008] Therefore, the purpose of this invention is to provide an intelligent conversational interaction system and method based on facial recognition and electroencephalography (EEG) and electrodermal conductance.

[0009] To address the aforementioned technical problems, this invention provides an intelligent conversation interaction system based on facial recognition and EEG / TEG, comprising: a modality acquisition module, a state recognition module, a speech recognition module, and a conversation interaction module;

[0010] The modality acquisition module is used to acquire the user's modality information in real time through a sensor network, correct the user's modality information according to environmental changes, and synchronize the corrected modality information.

[0011] The state recognition module is used to extract features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state.

[0012] The speech recognition module is used to recognize the user's speech and determine the user's needs;

[0013] The conversation interaction module is used to generate interactive responses based on user status and user needs, and collect user feedback on the interactive responses in order to optimize the interactive responses and status recognition strategies. The logic for optimizing the status recognition strategy includes comparing the head posture angle with the angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients according to the user's motion state.

[0014] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEV as described in this invention, the following steps are taken: Features of the modal information are extracted to generate user features, including detection point distance, head posture angle, frequency band ratio, and EEG / TEV amplitude. The user state is then identified using a state recognition strategy, which includes:

[0015] The user response value v is calculated based on user characteristics. The user response threshold is configured. The user response threshold includes a response base value v1 and a response extreme value v2. The user state is obtained by comparing the response base value and the response extreme value with the user response value. If v < v1, the user state is identified as the state base value.

[0016] If v1≤v≤v2, then the user state is identified as a high state value;

[0017] If v > v2, then the user state is identified as an extreme state.

[0018] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEG of the present invention, the calculation formula for the user response value is as follows:

[0019] v=w1×d+w2×θ+w3×r+w4×p;

[0020] In the formula, v represents the user's response value, w1 represents the distance coefficient, d represents the distance to the detection point, w2 represents the angle coefficient, θ represents the head posture angle, w3 represents the ratio coefficient, r represents the frequency band ratio, w4 represents the amplitude coefficient, and p represents the skin conductance amplitude.

[0021] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEG of the present invention, the logic of the optimized state recognition strategy includes:

[0022] Configure an angle threshold, which includes a base angle value θ1 and an extreme angle value θ2. Compare the base angle value and the extreme angle value with the head posture angle to obtain the user's motion state. If |θ|≤θ1, the user's motion state is no motion.

[0023] If θ1 < |θ| < θ2, then the user's motion state is the motion baseline value;

[0024] If |θ|≥θ2, then the user's motion state is the extreme value of motion;

[0025] The coefficients are dynamically adjusted based on the user's motion status. These coefficients include distance coefficient, angle coefficient, ratio coefficient, and amplitude coefficient.

[0026] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEG of the present invention, the logic of the dynamic adjustment coefficient includes:

[0027] The coefficients before adjustment are obtained. Based on the user's motion state, a smoothing factor and a target coefficient are determined. The difference between the target coefficient and the coefficients before adjustment is calculated to obtain the coefficient difference value. The smoothing factor and the coefficient difference value are then processed and combined with the coefficients before adjustment to obtain the adjusted coefficients, thus completing the dynamic adjustment of the coefficients.

[0028] As a preferred embodiment of the intelligent conversational interaction system based on facial recognition and EEG-cortex electrophysiology described in this invention, the modal information includes facial information, EEG information, and cortex electrophysiology information;

[0029] Strategies for acquiring user modal information in real time through sensor networks include:

[0030] The system corrects the user's modal information based on environmental changes, including adaptive image enhancement of facial information, comprehensive artifact control of EEG information, and comprehensive environmental compensation of TE information, and synchronizes the corrected modal information with data.

[0031] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEG of the present invention, the logic for comprehensive environmental compensation of the TEG information includes:

[0032] The raw skin conductance information is obtained, the current skin temperature is measured, a standard skin temperature is configured, and the raw skin conductance information is processed with the current skin temperature and the standard skin temperature to obtain the temperature-compensated skin conductance information.

[0033] Measure the current ambient humidity, configure the standard ambient humidity, and process the temperature-compensated skin conductance information with the current ambient humidity and the standard ambient humidity to obtain temperature and humidity-compensated skin conductance information.

[0034] The average value of the skin conductance information after temperature and humidity compensation is calculated to perform baseline correction on the skin conductance information after temperature and humidity compensation.

[0035] As a preferred embodiment of the intelligent conversation interaction system based on facial recognition and EEG / TEG of the present invention, the step of synchronizing the corrected modal information includes:

[0036] Determine the target synchronization time point, and find two time points close to the target synchronization time point, as well as the values ​​of facial information, EEG information, and TEG information at these two time points, in order to interpolate the target synchronization time point.

[0037] Intelligent conversation interaction methods based on facial recognition and EEG / TEG include:

[0038] S1. Acquire the user's modal information in real time through a sensor network, correct the user's modal information according to environmental changes, and synchronize the corrected modal information.

[0039] S2. Extract the features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state.

[0040] S3. Recognize the user's voice and determine the user's needs;

[0041] S4. Generate interactive responses based on user status and user needs, and collect user feedback on the interactive responses to optimize the interactive responses and status recognition strategies. The logic for optimizing the status recognition strategy includes comparing the head posture angle with the angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients based on the user's motion state.

[0042] A computer device includes: a memory for storing instructions; and a processor for executing the instructions, causing the computer device to perform an intelligent conversational interaction method based on facial recognition and electroencephalography (EEG) and electrocorticography (ECG).

[0043] A computer-readable storage medium having a computer program stored thereon, which, when executed, implements an intelligent conversational interaction method based on facial recognition and EEG / TEG.

[0044] The beneficial effects of this invention are as follows: The modal acquisition module acquires user modal information in real time through a sensor network and corrects this information based on environmental changes. Simultaneously, it synchronizes the corrected modal information, comprehensively reflecting changes in the user's state. The state recognition module extracts features from the modal information and identifies the user's state through a state recognition strategy. This strategy includes calculating the user's response value and comparing it with a user response threshold to obtain the user's state, accurately identifying the user's current state. The speech recognition module recognizes the user's speech and determines their needs, enabling a deeper understanding of the user's true needs and intentions in a specific state. The conversation interaction module generates interactive responses based on the user's state and needs, and collects user feedback on these responses to optimize the interaction response and state recognition strategy. The logic for optimizing the state recognition strategy includes comparing the user's head posture angle with an angle threshold to obtain their movement state and dynamically adjusting coefficients based on the user's movement state to generate multiple forms of interactive responses. This improves user engagement and experience, continuously enhancing the intelligence level and user experience of the conversation interaction system. The intelligent conversation interaction system can be applied in multiple scenarios. Attached Figure Description

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0046] Figure 1 This is a system structure diagram of the intelligent conversation interaction system based on facial recognition and electroencephalography (EEG) and electrodermal conductance (EDC) of the present invention.

[0047] Figure 2 This is a flowchart illustrating the state recognition strategy of the intelligent conversation interaction system based on facial recognition and EEG / EEG / EEG of the present invention.

[0048] Figure 3 This is a flowchart illustrating the optimized state recognition strategy of the intelligent conversation interaction system based on facial recognition and EEG-cortical conductance of the present invention.

[0049] Figure 4 This is a flowchart of the intelligent conversation interaction method based on facial recognition and EEG-cortical conductance of the present invention. Detailed Implementation

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0051] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention can also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0052] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0053] Example 1

[0054] This embodiment provides a system architecture diagram of an intelligent conversation interaction system based on facial recognition and EEG / TEG, as follows: Figure 1 As shown, the intelligent conversation interaction system based on facial recognition and EEG-cortical conductance includes a modality acquisition module, a state recognition module, a speech recognition module, and a conversation interaction module.

[0055] The modal acquisition module is used to acquire the user's modal information in real time through the sensor network, correct the user's modal information according to environmental changes, and synchronize the corrected modal information.

[0056] Modal information includes facial information, electroencephalogram (EEG) information, and electrodermal conductance (EDC) information.

[0057] Strategies for acquiring user modal information in real time through sensor networks include:

[0058] The system corrects the user's modal information based on environmental changes, including adaptive image enhancement of facial information, comprehensive artifact control of EEG information, and comprehensive environmental compensation of TE information, and synchronizes the corrected modal information with data.

[0059] It should be understood that adaptive image enhancement of facial information is used to adjust the contrast of facial images, integrated artifact control of EEG information is used to remove artifacts in EEG information through ICA technology and control the intensity of motion artifact suppression, and integrated environmental compensation of skin conductance information is used to correct skin conductance information and comprehensively consider the effects of skin temperature and ambient humidity.

[0060] For example, the functional expression for adaptive image enhancement of facial information is shown below:

[0061] I(a,b)=λ×log(1+I0(a,b));

[0062] In the formula, I(a,b) represents the pixel value of the enhanced facial image at coordinates (a,b), λ represents the enhancement factor, and I0(a,b) represents the pixel value of the original facial image at coordinates (a,b).

[0063] It needs to be explained that: the enhancement factor λ is used to control the intensity of the logarithmic transformation to adjust the scaling factor of the facial image contrast. The enhancement factor λ is usually a positive real number, ranging from 0 to 5. A larger enhancement factor λ will result in a more obvious enhancement effect on the facial image, and the dark details of the facial image will be stretched, while a smaller enhancement factor λ will result in a weaker enhancement effect on the facial image. If λ = 1, it means that the facial image only undergoes a standard logarithmic transformation; the pixel value I0(a,b) at coordinates (a,b) of the original facial image refers to the pixel intensity value at coordinates (a,b) of the facial image, ranging from 0 to 255; log(·) is a logarithmic function used to adjust the contrast of the facial image to enhance the visibility of low-brightness pixels and reduce the difference in high-brightness areas, thereby enhancing the contrast. The above formula log(1+I0(a,b)) ensures that the input value is at least 1, avoiding the undefined case of the logarithmic function when the input is 0.

[0064] For example, the functional expression for the integrated artifact control of EEG information is shown below:

[0065] b(t) = β × ICA(b0(t));

[0066] In the formula, b(t) represents the observed value of EEG information at time t after ICA artifact removal and motion artifact suppression, β represents the motion suppression coefficient, ICA(·) represents the artifact removal of the original EEG information by ICA technology, and b0(t) represents the observed value of the original EEG information at time t.

[0067] It should be explained that: the observed value b0(t) of the original EEG information at time t typically ranges from -100μV to 100μV; the motion inhibition coefficient β is a weighted coefficient used to control the intensity of motion artifact suppression. The motion inhibition coefficient β is usually a positive number, ranging from 0 to 1. The larger the value of the motion inhibition coefficient β, the stronger the motion artifact suppression, and the smaller the value of the motion inhibition coefficient β, the weaker the motion artifact suppression; the goal of ICA is to retain the useful information of EEG information to the maximum extent while removing artifacts.

[0068] The logic of comprehensive environmental compensation for skin charge information includes:

[0069] The raw skin conductance information is obtained, the current skin temperature is measured, a standard skin temperature is configured, and the raw skin conductance information is processed with the current skin temperature and the standard skin temperature to obtain the temperature-compensated skin conductance information.

[0070] Measure the current ambient humidity, configure the standard ambient humidity, and process the temperature-compensated skin conductance information with the current ambient humidity and the standard ambient humidity to obtain temperature and humidity-compensated skin conductance information.

[0071] The average value of the skin conductance information after temperature and humidity compensation is calculated to perform baseline correction on the skin conductance information after temperature and humidity compensation.

[0072] For example, the functional expression for the comprehensive environmental compensation of skin charge information is as follows:

[0073]

[0074] In the formula, s represents the skin conductance information after comprehensive environmental compensation, s0 represents the original skin conductance information, γ represents the temperature coefficient, T represents the current skin temperature, T0 represents the standard skin temperature, h0 represents the standard ambient humidity, h represents the current ambient humidity, and mean(·) represents the mean value processing.

[0075] It needs to be explained that: the skin conductance information s after comprehensive environmental compensation refers to the information that has removed the effects caused by skin temperature, ambient humidity, and baseline drift during long-term measurement, so that the skin conductance information can reflect more actual physiological activities, and is usually a positive number; the original skin conductance information s0 reflects the change in the conductivity of the skin surface, and is usually a positive number; the current skin temperature T refers to the temperature of the skin surface at the time of measurement, which usually fluctuates between 28℃ and 34℃; the standard skin temperature T0 refers to the skin temperature under stable ambient temperature conditions; the temperature coefficient γ refers to the degree of influence of skin temperature changes on skin conductance information, and its value ranges between 0.01 and 0.05; the standard ambient humidity h0 is usually the humidity at the time of experimental setup, and its value is usually between 30% and 60%; the mean processing mean(·) is used for baseline drift correction, which eliminates the influence of baseline drift during long-term measurement by subtracting the average value in parentheses; the calculation of the above formulas requires dimensionless processing.

[0076] The steps for synchronizing the corrected modal information include:

[0077] Determine the target synchronization time point, and find two time points close to the target synchronization time point, as well as the values ​​of facial information, EEG information, and TEG information at these two time points, in order to interpolate the target synchronization time point.

[0078] For example, the function expression for synchronizing the corrected modal information is shown below:

[0079]

[0080] In the formula, This indicates facial information, EEG information, and ductal nerve conductance information at time point t. i The value after data synchronization, X i-1 X represents the values ​​of facial information, electroencephalogram (EEG) information, and electrodermal conductance (EDC) information at time point i-1. i+1 t represents the values ​​of facial information, EEG information, and ductal skin information at time point i+1. i+1 t represents the timestamp of the (i+1)th time point. i-1 t represents the timestamp of the (i-1)th time point. * Indicates the target synchronization time point.

[0081] What needs to be explained is: These refer to facial information at time point t. i The value of the EEG information at time point t i The values ​​and skin conductance information at time point t i The value at point X needs to be dimensionless; while X i-1 and X i+1 The value of refers to the time t is close to the target synchronization point. * The values ​​corresponding to the two time points; target synchronization time point t * The value of is usually within the range of all timestamps, and time points with smaller intervals are usually selected for alignment.

[0082] Sensor networks consist of RGB cameras, EEG sensors (such as head-mounted electroencephalograms), and electrodermal sensors (such as GSR sensors). RGB cameras capture facial images, and changes in these images are often closely related to the user's state. EEG sensors acquire electroencephalogram (EEG) signals, which provide rich information about brain activity and can be used to analyze the user's state. GSR sensors acquire electrodermal signals, and changes in these signals can be used to detect changes in the user's state. In real-world applications, environmental factors can affect the data acquisition performance of sensor networks; therefore, sensor networks... During data acquisition, the network needs to consider environmental interference correction, including adaptive image enhancement of facial images, comprehensive artifact control of EEG signals, and comprehensive environmental compensation of TE signals, in order to improve the quality of facial images under low light conditions, remove artifacts caused by muscle activity or eye movement, and compensate for changes in skin temperature and ambient humidity after measuring the user's TE response. Since the sampling rates of different sensors may be different, it is necessary to synchronize the modal information of all users to the same point in time for subsequent analysis and processing. Through the above strategies, the modal information of users can be acquired and processed in real time and accurately, providing support for subsequent interactive responses.

[0083] The state recognition module is used to extract features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state.

[0084] Modal information features are extracted to generate user features, including detection point distance, head pose angle, frequency band ratio, and skin conductance amplitude. User state is then identified using a state recognition strategy, such as... Figure 2 As shown, it specifically includes:

[0085] The user response value v is calculated based on user characteristics. The user response threshold is configured. The user response threshold includes a response base value v1 and a response extreme value v2. The user state is obtained by comparing the response base value and the response extreme value with the user response value. If v < v1, the user state is identified as the state base value.

[0086] If v1≤v≤v2, then the user state is identified as a high state value;

[0087] If v > v2, then the user state is identified as an extreme state.

[0088] For example, the function expression for the distance between detection points is shown below:

[0089]

[0090] In the formula, d represents the distance between the detection points, and (a1,b1) and (a2,b2) are the coordinates of the two detection points.

[0091] It needs to be explained that the detection point distance d refers to the distance between facial key points. If you want to measure the distance between the eyes, then the two detection points are the inner corner and outer corner of one eye, respectively. If you want to measure the relative distance between the lips and the nose, then the two detection points are the tip of the nose and the center point of the lower lip, respectively. The range of values ​​for the detection point distance d depends on the position of the detection points in the facial image. A detection point distance of 0 means that the two detection points coincide. Generally speaking, the larger the pixel value, the larger the detection point distance.

[0092] For example, the functional expression for the head pose angle is shown below:

[0093]

[0094] In the formula, θ represents the head posture angle, and (a1,b1) and (a2,b2) are the coordinates of the two detection points.

[0095] It needs to be explained that: the head posture angle θ describes the angle of the head orientation, which is the directional angle from the first detection point (a1, b1) to the second detection point (a2, b2), usually the angle relative to the horizontal line (x-axis). The value range of the head posture angle θ is (-π / 2, π / 2); arctan(·) is the arctangent function, used to calculate the head posture angle from the slope, converting the ratio of the vertical and horizontal distances between two detection points into an angle value; when calculating the pitch angle (that is, the up-and-down rotation of the head), the two detection points are the tip of the nose and the chin; when calculating the yaw angle (that is, the left-and-right rotation of the head), the two detection points are the tip of the nose and the center point of the eyes; when calculating the roll angle (that is, the rotation of the head around the central axis, i.e., the tilt of the head), the two detection points are the center point of the eyes and the center point of the lips.

[0096] The band ratio refers to the power ratio of Theta waves and Beta waves calculated from the acquired EEG information, which can indicate the user's state.

[0097] For example, the functional expression for the magnetic field amplitude is shown below:

[0098] p = max(s) - min(s);

[0099] In the formula, p represents the skin conductance amplitude, and s represents the skin conductance information after comprehensive environmental compensation.

[0100] It should be explained that the skin conductance amplitude p refers to the difference between the maximum and minimum values ​​of skin conductance information after stimulation. It is used to quantify the intensity of the skin conductance response and can reflect the degree of the user's response. The higher the skin conductance amplitude, the stronger the user's response.

[0101] The formula for calculating user response values ​​is shown below:

[0102] v=w1×d+w2×θ+w3×r+w4×p;

[0103] In the formula, v represents the user's response value, w1 represents the distance coefficient, d represents the distance to the detection point, w2 represents the angle coefficient, θ represents the head posture angle, w3 represents the ratio coefficient, r represents the frequency band ratio, w4 represents the amplitude coefficient, and p represents the skin conductance amplitude.

[0104] It needs to be explained that: the distance coefficient w1 refers to the influence of the detection point distance on the user's reaction value, and the value of the distance coefficient w1 is usually a non-negative number; the angle coefficient w2 refers to the influence of the head posture angle on the user's reaction value, and the value of the angle coefficient w2 is usually a non-negative number; the ratio coefficient w3 refers to the influence of the frequency band ratio on the user's reaction value, and the value of the ratio coefficient w3 is usually a non-negative number; the value of the frequency band ratio r ranges from 0 to 1 or is greater than 1, depending on the different states of the brain. In a focused state, Beta waves are dominant, so the frequency band ratio is lower; in a relaxed state, Theta waves are dominant, so the frequency band ratio is higher; the amplitude coefficient w4 refers to the influence of the skin conductance amplitude on the user's reaction value, and the value of the amplitude coefficient w4 is usually a non-negative number; the specific values ​​of the above-mentioned distance coefficient w1, angle coefficient w2, ratio coefficient w3, and amplitude coefficient w4 need to be obtained through experimental optimization, and must satisfy w1+w2+w3+w4=1.

[0105] In practical applications, facial images are analyzed using a deep learning model to extract facial features, including the distance between detection points and head pose angles. The distance between detection points refers to the distance between key facial points (including the distance between the eyes, and the relative distances between the lips and nose). Head pose angles are calculated based on these key facial points to determine the user's facial expressions and movements. Features of the extracted EEG information, including frequency band ratios, are used to assess the user's state. Features of the extracted TE (electrical conductance) information, including TE amplitude, are used to determine... The system analyzes the intensity of user responses. Based on the comparison between user response values ​​and baseline and extreme response values, the user state is identified. When the user state is at the baseline level, facial expressions are typically relaxed, skin conductance is low, and Theta waves dominate. When the user state is at a high level, facial expressions are typically intense, skin conductance is strong, and the frequency band ratio is high. When the user state is at an extreme level, facial expressions are typically tense, head posture changes significantly, and skin conductance is strong. These steps enable real-time monitoring and response to changes in user state, providing a more personalized and adaptive interactive experience.

[0106] The speech recognition module is used to recognize the user's voice and determine the user's needs.

[0107] For example, when a user makes a voice input, the system recognizes the user's voice, converts the user's voice into text information, cleans and segments the converted text to remove noise and irrelevant information, analyzes the text through a semantic understanding model to identify the user's intent, and further refines the intent by combining facial images to generate user needs, and determines the priority of user needs and refines the response.

[0108] The conversation interaction module is used to generate interactive responses based on user status and user needs, and to collect user feedback on the interactive responses in order to optimize the interactive responses and status recognition strategies. The optimization of the status recognition strategy includes comparing the head posture angle with the angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients according to the user's motion state.

[0109] Based on user status and needs, the system selects a suitable response template from the response template library to generate an interactive response. For example, when the system detects that a user is querying weather information, it automatically generates a weather-related interaction with the response template: "Today's weather is 25℃, and tomorrow's forecast is 20℃. Please dress warmly." When the system detects that a user is in an extreme state, it automatically provides corresponding acceleration and care with the response template: "I will check for you as soon as possible. Please wait a moment." After each interaction with the user, the system collects user feedback on the response, including whether the user is satisfied or dissatisfied. It configures a frequency threshold and compares the frequency of user dissatisfaction with the frequency threshold to obtain an optimization strategy. If the frequency of user dissatisfaction is less than or equal to the frequency threshold, the optimization strategy is to optimize the interactive response.

[0110] If the frequency of user dissatisfaction exceeds the frequency threshold, the optimization strategy becomes the optimization of the state recognition strategy.

[0111] For example, optimizing interactive responses includes simplifying the language in response templates, speeding up the response of conversational interaction systems, and enhancing accuracy. For instance, for some complex or ambiguous queries, multi-turn dialogues should be used to accurately determine the user's true needs, and the accuracy of the response should be enhanced by using the user's historical behavior. At the same time, the response should be completed within a reasonable time. For example, the response time for simple queries should not exceed 3 seconds, and for more complex queries, the response time should be appropriately extended while informing the user of the progress.

[0112] The logic for optimizing the state recognition strategy is as follows: Figure 3 As shown, it specifically includes:

[0113] Configure an angle threshold, which includes a base angle value θ1 and an extreme angle value θ2. Compare the base angle value and the extreme angle value with the head posture angle to obtain the user's motion state. If |θ|≤θ1, the user's motion state is no motion, that is, the user is not performing any movement and may be in a static state.

[0114] If θ1 < |θ| < θ2, then the user's motion state is the motion baseline, which means the user is making slight movements;

[0115] If |θ|≥θ2, then the user's motion state is at its extreme value, meaning the user is engaged in vigorous movement.

[0116] The coefficients are dynamically adjusted based on the user's motion status. These coefficients include distance coefficient, angle coefficient, ratio coefficient, and amplitude coefficient. The logic for dynamically adjusting the coefficients includes:

[0117] The coefficients before adjustment are obtained. Based on the user's motion state, a smoothing factor and a target coefficient are determined. The difference between the target coefficient and the coefficients before adjustment is calculated to obtain the coefficient difference value. The smoothing factor and the coefficient difference value are then processed and combined with the coefficients before adjustment to obtain the adjusted coefficients, thus completing the dynamic adjustment of the coefficients.

[0118] An example, the function expression for dynamically adjusting the coefficients is shown below:

[0119]

[0120] In the formula, w j (m) represents the coefficient adjusted under the user's motion state m, w j The coefficients before adjustment are represented by κ(m), and the smoothing factor is represented by κ(m) under motion state m. This represents the target coefficient.

[0121] For example, the logic for determining the smoothing factor includes: if the user's motion state is no motion, then κ(m) = 0;

[0122] If the user's motion state is the motion baseline, then κ(m) = 0.5;

[0123] If the user's motion state is at its extreme value, then κ(m) = 1.

[0124] It needs to be explained that the coefficient w is adjusted under the user's motion state m. j (m) where j represents the index. When j=1, it refers to the distance coefficient adjusted under the user's motion state m; when j=2, it refers to the angle coefficient adjusted under the user's motion state m; when j=3, it refers to the ratio coefficient adjusted under the user's motion state m; and when j=4, it refers to the amplitude coefficient adjusted under the user's motion state m. The smoothing factor κ(m) is used to adjust the coefficient according to the current user's motion state, controlling the magnitude of the coefficient adjustment. The value of the smoothing factor κ(m) ranges from 0 to 1. Target coefficient This refers to the target coefficient that a user wants to adjust to when they are in a certain state of motion. The value range is between 0 and 1.

[0125] In practical applications, when the user is in a static state, the distance and amplitude coefficients are typically higher because the detection point distance and skin conductance amplitude more accurately reflect the user's state in a static condition. Conversely, the angle and ratio coefficients are typically lower because head posture and brain activity change relatively little in a static state. When the user's motion state is at a baseline, the angle and amplitude coefficients should be appropriately increased, especially since motion has a certain impact on the user's state, while the distance and ratio coefficients may remain relatively stable. When the user's motion state is at an extreme motion state, the angle and amplitude coefficients are typically higher because head posture... Changes in facial expression and skin conductance are closely related to the user's state, while the distance coefficient and ratio coefficient are usually set relatively low because facial expressions are more affected by movement, while changes in brain activity are less noticeable than head posture. Based on the adjusted coefficients and the user's state, the response speed and response template of the conversation interaction system are adjusted. For example, if the user is in an extreme state, the conversation interaction system should slow down the response speed, lower the tone, and provide encouraging interactive responses. If the user is in a normal state, the conversation interaction system should speed up the response speed, raise the tone, and provide more interactive response information to provide the user with a more personalized and adaptive interactive experience.

[0126] Example 2

[0127] This embodiment provides a flowchart of a smart conversation interaction method based on facial recognition and EEG / TEG, such as... Figure 4 As shown, intelligent conversation interaction methods based on facial recognition and EEG / TEG include:

[0128] S1. Acquire the user's modal information in real time through a sensor network, correct the user's modal information according to environmental changes, and synchronize the corrected modal information.

[0129] S2. Extract the features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state.

[0130] S3. Recognize the user's voice and determine the user's needs;

[0131] S4. Generate interactive responses based on user status and user needs, and collect user feedback on the interactive responses to optimize the interactive responses and status recognition strategies. The logic for optimizing the status recognition strategy includes comparing the head posture angle with the angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients based on the user's motion state.

[0132] For details on the intelligent conversation interaction method based on facial recognition and EEG-cortex, please refer to the Intelligent Conversation Interaction System Based on Facial Recognition and EEG-cortex, which will not be elaborated here.

[0133] Example 3

[0134] In this embodiment, a computer device is provided, including a memory and a processor. The memory is used to store instructions, and the processor is used to execute the instructions, causing the computer device to perform the steps of implementing the above-described intelligent conversation interaction method based on facial recognition and EEG-cortex conductance.

[0135] Example 4

[0136] In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, it implements the steps of the above-described intelligent conversation interaction method based on facial recognition and EEG-cortex conductance.

[0137] The computer-readable storage medium includes various media that store program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0138] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An intelligent conversation interaction system based on facial recognition and EEG / TEG, characterized in that: include: Modality acquisition module, status recognition module, speech recognition module, and conversation interaction module; The modality acquisition module is used to acquire the user's modality information in real time through a sensor network, correct the user's modality information according to environmental changes, and synchronize the corrected modality information. The state recognition module is used to extract features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state. Features of modal information are extracted to generate user features, including detection point distance, head pose angle, frequency band ratio, and skin conductance amplitude. The state recognition strategy includes: Calculate user response value based on user characteristics Configure user response thresholds, which include response base values. and reaction extreme values The user state is obtained by comparing the reaction base value and reaction extreme value with the user's reaction value. If so, the user's state is identified as the state base value; if If so, the user's state is identified as a high value; if If so, the user's state is identified as an extreme state. The formula for calculating the user's response value is as follows: ; In the formula, This represents the user's response value. Represents the distance coefficient. Indicates the distance between the detection points. Indicates the angle coefficient. Indicates the head posture angle. Represents the ratio coefficient. Indicates the bandwidth ratio. Indicates the amplitude coefficient. Indicates the amplitude of the skin charge; The speech recognition module is used to recognize the user's speech and determine the user's needs; The conversation interaction module is used to generate interactive responses based on user status and needs, and collect user feedback on the interactive responses to optimize the interactive responses and status recognition strategy. The logic for optimizing the status recognition strategy includes comparing the head posture angle with an angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients based on the user's motion state. The logic for optimizing the status recognition strategy includes: Configure angle thresholds, which include angle base values. and angular extrema The user's motion state is obtained by comparing the base angle and extreme angle with the head posture angle. If the user's motion state is no motion, then the user's motion state is no motion. Then the user's motion state is the motion baseline value; if If the user's motion state is extreme, then the user's motion state is considered extreme. The coefficients are dynamically adjusted based on the user's motion state, and these coefficients include distance coefficient, angle coefficient, ratio coefficient, and amplitude coefficient. The logic for the dynamic adjustment coefficient includes: The coefficients before adjustment are obtained. Based on the user's motion state, a smoothing factor and a target coefficient are determined. The difference between the target coefficient and the coefficients before adjustment is calculated to obtain the coefficient difference value. The smoothing factor and the coefficient difference value are then processed and combined with the coefficients before adjustment to obtain the adjusted coefficients, thus completing the dynamic adjustment of the coefficients.

2. The intelligent conversation interaction system based on facial recognition and EEG / TEG as described in claim 1, characterized in that: The modal information includes facial information, electroencephalogram (EEG) information, and electrical activity of the skin. Strategies for acquiring user modal information in real time through sensor networks include: The system corrects the user's modal information based on environmental changes, including adaptive image enhancement of facial information, comprehensive artifact control of EEG information, and comprehensive environmental compensation of TE information, and synchronizes the corrected modal information with data.

3. The intelligent conversation interaction system based on facial recognition and EEG / TEG as described in claim 2, characterized in that: The logic for comprehensive environmental compensation based on the skin contact information includes: The raw skin conductance information is obtained, the current skin temperature is measured, a standard skin temperature is configured, and the raw skin conductance information is processed with the current skin temperature and the standard skin temperature to obtain the temperature-compensated skin conductance information. Measure the current ambient humidity, configure the standard ambient humidity, and process the temperature-compensated skin conductance information with the current ambient humidity and the standard ambient humidity to obtain temperature and humidity-compensated skin conductance information. The average value of the skin conductance information after temperature and humidity compensation is calculated to perform baseline correction on the skin conductance information after temperature and humidity compensation.

4. The intelligent conversation interaction system based on facial recognition and EEG / TEG as described in claim 3, characterized in that: The steps for synchronizing the corrected modal information include: Determine the target synchronization time point, and find two time points close to the target synchronization time point, as well as the values ​​of facial information, EEG information, and TEG information at these two time points, in order to interpolate the target synchronization time point.

5. A method for intelligent conversation interaction based on facial recognition and EEG / TEG, implemented based on any one of claims 1-4, characterized in that: include: S1. Acquire the user's modal information in real time through a sensor network, correct the user's modal information according to environmental changes, and synchronize the corrected modal information. S2. Extract the features of modal information and identify the user state through a state recognition strategy. The state recognition strategy includes calculating the user response value and comparing the user response value with the user response threshold to obtain the user state. S3. Recognize the user's voice and determine the user's needs; S4. Generate interactive responses based on user status and user needs, and collect user feedback on the interactive responses to optimize the interactive responses and status recognition strategies. The logic for optimizing the status recognition strategy includes comparing the head posture angle with the angle threshold to obtain the user's motion state, and dynamically adjusting the coefficients based on the user's motion state.

6. A computer device, characterized in that: include: Memory, used to store instructions; A processor is configured to execute the instructions, causing the computer device to perform the intelligent conversation interaction method based on facial recognition and EEG-cortex as described in claim 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, it implements the intelligent conversation interaction method based on facial recognition and EEG / TEG as described in claim 5.

Citation Information

Patent Citations

  • A conversational intelligent teaching system

    CN108021703B

  • Dialogue Implementation Methods and Systems for Health and Wellness

    CN113488057B

  • Method, device and equipment for generating head posture of responder and storage medium

    CN116402927A

  • Method for dynamically adjusting precision of guide rail of lathe bed of large numerical control planer type milling machine

    CN118768620A