Emotion prediction method and device based on multiple biosignals, equipment and medium
Patent Information
- Application Number
- CN202610610349.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-09-15
AI Technical Summary
但是,在操作者佩戴护具(头盔、口罩、眼镜)或光照变化、头部偏转、操作者有意识地控制表情来掩饰内心状态等条件下,情绪识别的准确率将会降低;并且,这种方式属于事后识别,容易产生干预不及时的问题
[0015] This application embodiment acquires tactile contour subsequences collected by a tactile sensor at each time step, and determines micro-gesture events at each time step based on the tactile contour subsequences; determines the dynamic feature subsequences of the corresponding micro-gesture events based on the tactile contour subsequences collected at each time step, and constructs a micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequences of each time step; for each time step, calculates the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence, and selects the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area; in each time step, physiological signals are acquired from the physiological signal acquisition area by a physiological signal sensor, and constructs a physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step; acquires grip force signals collected by a pressure sensor at each time step, and constructs a grip force sequence corresponding to the target time window by combining the grip force signals corresponding to each time step; and predicts the emotional change trend from the target time window to the next target time window by using a large language model based on the micro-gesture dynamic time sequence, physiological signal sequence, and grip force sequence corresponding to the target time window, thus obtaining the emotional change prediction result. In this way, we can take micro-gestures, a behavioral feature that is more difficult to conceal, in multimodal signals as the main factor, combine dynamic optimization of physiological signal acquisition and grip force to assist in discrimination, and use the model to directly output the emotional trend of future time windows, so as to achieve a more essential and robust characterization of emotional state and forward-looking prediction. Specifically, compared to solutions that rely solely on facial expressions and are reactive in their recognition, this application captures subtle, subconscious hand movements of the operator using tactile sensory micro-gesture dynamic sequences. This fundamentally avoids recognition biases caused by facial occlusion, lighting changes, and intentional facial expression concealment. By dynamically selecting the optimal physiological signal acquisition area in real-time through tactile contour amplitude guidance, it ensures the continuity and signal-to-noise ratio of physiological signals during complex maneuvers such as hand switching and large-angle turns. Furthermore, by combining grip force signals, it effectively distinguishes between normal maneuvers and tense exertion, significantly suppressing false alarms caused by physical manipulation. Finally, a large language model performs cross-time window temporal modeling of multiple biosignal sequences, directly outputting a prediction of emotional changes for the next target time window, thus enabling early intervention before emotional deterioration (such as pre-anger or pre-fatigue) occurs. In summary, this application improves the accuracy of emotion recognition and the timeliness of emotion intervention.
Smart Images

Figure CN122744792A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and medium for emotion prediction based on multiple biological signals. Background Technology
[0002] In various scenarios requiring continuous control (such as driving vehicles, operating aircraft, operating construction machinery, or engaging in virtual reality interaction), the operator's emotional state directly impacts their behavioral decisions and operational safety. Real-time monitoring of emotional states can provide crucial information for behavioral intervention, risk warning, and human-machine collaboration, helping to reduce the risk of accidents caused by negative emotions such as anger, tension, and fatigue, and improving the control experience and safety.
[0003] In related technologies, emotion detection is generally performed through facial expression recognition. Specifically, a camera captures images of the operator's face, and facial features such as eye movements, mouth shapes, and eyebrows are identified to determine the emotional state. However, the accuracy of emotion recognition decreases when the operator wears protective gear (helmet, mask, glasses), or when there are changes in lighting, head tilting, or when the operator consciously controls their expressions to conceal their inner state. Furthermore, this method is a reactive approach, which can easily lead to problems with untimely intervention. Summary of the Invention
[0004] This application proposes a method, device, equipment, and medium for emotion prediction based on multiple biological signals, which can improve the accuracy of emotion recognition and the timeliness of emotion intervention.
[0005] To achieve the above objectives, a first aspect of this application proposes an emotion prediction method based on multiple biological signals, applied to a target device, wherein the target device includes at least a tactile sensor, a physiological signal sensor, and a pressure sensor, and the method includes: The tactile contour subsequences collected by the tactile sensor at each time step are acquired, and the micro-gesture events at each time step are determined based on the tactile contour subsequences. Based on the tactile contour subsequence acquired at each time step, the dynamic feature subsequence of the corresponding micro-gesture event is determined, and combined with the dynamic feature subsequence of each time step, the micro-gesture dynamic time sequence corresponding to the current target time window is constructed. For each time step, the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area is calculated based on the corresponding tactile contour subsequence, and the skin contact area with the largest tactile contour amplitude is selected as the physiological signal acquisition area. In each time step, physiological signals are acquired from the physiological signal acquisition area by the physiological signal sensor, and the physiological signal sequence corresponding to the target time window is constructed by combining the physiological signals corresponding to each time step. The grip force signal collected by the pressure sensor at each time step is acquired, and the grip force signal corresponding to each time step is combined to construct the grip force sequence corresponding to the target time window; Using a large language model, based on the micro-gesture dynamics time series, the physiological signal series, and the grip force series corresponding to the target time window, the emotional change trend from the target time window to the next target time window is predicted, and the emotional change prediction result is obtained.
[0006] Accordingly, a second aspect of this application proposes an emotion prediction device based on multiple biological signals, applied to a target device, wherein the target device includes at least a tactile sensor, a physiological signal sensor, and a pressure sensor, and the device includes: The acquisition module is used to acquire the tactile contour subsequence collected by the tactile sensor at each time step, and determine the micro-gesture event at each time step based on the tactile contour subsequence; The determination module is used to determine the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence collected at each time step, and to construct the micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequence of each time step. The calculation module is used to calculate the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence for each time step, and select the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area. The acquisition module is used to acquire physiological signals from the physiological signal acquisition area through the physiological signal sensor at each time step, and to construct the physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step; A construction module is used to acquire the grip force signal collected by the pressure sensor at each time step, and combine the grip force signal corresponding to each time step to construct the grip force sequence corresponding to the target time window; The prediction module is used to predict the emotional change trend from the target time window to the next target time window based on the micro-gesture dynamics time series, the physiological signal series and the grip force series corresponding to the target time window using a large language model, and to obtain the emotional change prediction result.
[0007] In some implementations, the prediction module is further configured to: Obtain the emotional baseline corresponding to the target object being detected; Based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, input prompt information is generated; Using a large language model, based on the input prompt information, the trend of emotional change from the target time window to the next target time window is predicted, and the emotional change prediction result is obtained.
[0008] In some implementations, the prediction module is further configured to: For each reference time window, the initial micro-gesture dynamics time series, initial physiological signal series, and initial grip force series of the target object in a resting state are collected; Based on multiple initial micro-gesture dynamics time series, multiple initial physiological signal sequences, and multiple initial grip force sequences corresponding to multiple reference time windows, the statistical parameters of micro-gesture dynamics, physiological signal, and grip force sequence are calculated. Based on the micro-gesture dynamics statistical parameters, the physiological signal statistical parameters, and the grip force sequence statistical parameters, the emotional baseline corresponding to the target object is obtained.
[0009] In some embodiments, the emotion prediction device based on multiple biosignals further includes a generation module for: Obtain the first trust parameter corresponding to the micro-gesture dynamics time series, and obtain the second trust parameter corresponding to the physiological signal sequence and the grip force sequence; A first descriptive information for the micro-gesture dynamics time series is generated based on the first trust parameter, and a second descriptive information for the physiological signal sequence and the grip force sequence is generated based on the second trust parameter; The step of generating input prompt information based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence includes: Based on the emotional baseline, the first descriptive information, the second descriptive information, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, input prompt information is generated.
[0010] In some implementations, the generation module is further configured to: Obtain the device context parameters for the target object to control the target device, wherein the device context parameters include at least the steering angular velocity and the speed; A first threshold is calculated based on the speed, wherein the first threshold is positively correlated with the speed; When the steering angular velocity exceeds the first threshold, a first trust parameter corresponding to the micro-gesture dynamics time series is calculated, and a second trust parameter corresponding to the physiological signal sequence and the grip force sequence is calculated, wherein the first trust parameter is less than the second trust parameter.
[0011] In some implementations, the prediction module is further configured to: Using a large language model, based on the input prompt information, emotion prediction based on multiple biological signals is performed on the target object to obtain the emotion vector trajectory corresponding to the target time window; Based on the input prompt information and the emotion vector trajectory, the emotion of the target object is predicted in the next target time window to obtain the predicted emotion vector trajectory corresponding to the next target time window. Based on the emotion vector trajectory and the predicted emotion vector trajectory, a risk level summary for the target object is obtained; Based on the emotion vector trajectory, the predicted emotion vector trajectory, and the risk level summary, the emotion change prediction result is obtained.
[0012] In some implementations, the determining module is further configured to: For each time step, determine the corresponding sliding sub-time window with each time step as the endpoint; Based on the tactile contour subsequence of at least one time point contained in the sliding sub-time window, the micro-gesture frequency corresponding to the micro-gesture event, the average rubbing intensity of the detected target object on the skin contact area corresponding to the tactile sensor, the rhythmic index, the cumulative hand displacement, and the pressure trend slope are determined. The rhythmicity index is used to characterize the degree of aggregation of the micro-gesture events corresponding to the sliding sub-time window on the time axis. The cumulative hand displacement at each time step is obtained by accumulating the tactile contour changes between adjacent time steps included in the sliding sub-time window. The tactile contour changes at each time step are determined based on the differences between the tactile contour subsequences between adjacent time steps included in the sliding sub-time window. The pressure trend slope at each time step is obtained by linearly fitting the micro-gesture event intensity values of multiple time steps included in the sliding sub-time window. Based on the micro-gesture frequency, average rubbing intensity, rhythmic index, cumulative hand displacement, and pressure trend slope corresponding to the sliding sub-time window at each time step, a dynamic feature sub-sequence corresponding to each time step is determined.
[0013] Accordingly, a third aspect of the embodiments of this application proposes a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the emotion prediction method based on multiple biosignals according to any one of the embodiments of the first aspect of this application.
[0014] Accordingly, a fourth aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the emotion prediction method based on multiple biosignals according to any one of the embodiments of the first aspect of this application.
[0015] This application embodiment acquires tactile contour subsequences collected by a tactile sensor at each time step, and determines micro-gesture events at each time step based on the tactile contour subsequences; determines the dynamic feature subsequences of the corresponding micro-gesture events based on the tactile contour subsequences collected at each time step, and constructs a micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequences of each time step; for each time step, calculates the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence, and selects the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area; in each time step, physiological signals are acquired from the physiological signal acquisition area by a physiological signal sensor, and constructs a physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step; acquires grip force signals collected by a pressure sensor at each time step, and constructs a grip force sequence corresponding to the target time window by combining the grip force signals corresponding to each time step; and predicts the emotional change trend from the target time window to the next target time window by using a large language model based on the micro-gesture dynamic time sequence, physiological signal sequence, and grip force sequence corresponding to the target time window, thus obtaining the emotional change prediction result. In this way, we can take micro-gestures, a behavioral feature that is more difficult to conceal, in multimodal signals as the main factor, combine dynamic optimization of physiological signal acquisition and grip force to assist in discrimination, and use the model to directly output the emotional trend of future time windows, so as to achieve a more essential and robust characterization of emotional state and forward-looking prediction. Specifically, compared to solutions that rely solely on facial expressions and are reactive in their recognition, this application captures subtle, subconscious hand movements of the operator using tactile sensory micro-gesture dynamic sequences. This fundamentally avoids recognition biases caused by facial occlusion, lighting changes, and intentional facial expression concealment. By dynamically selecting the optimal physiological signal acquisition area in real-time through tactile contour amplitude guidance, it ensures the continuity and signal-to-noise ratio of physiological signals during complex maneuvers such as hand switching and large-angle turns. Furthermore, by combining grip force signals, it effectively distinguishes between normal maneuvers and tense exertion, significantly suppressing false alarms caused by physical manipulation. Finally, a large language model performs cross-time window temporal modeling of multiple biosignal sequences, directly outputting a prediction of emotional changes for the next target time window, thus enabling early intervention before emotional deterioration (such as pre-anger or pre-fatigue) occurs. In summary, this application improves the accuracy of emotion recognition and the timeliness of emotion intervention. Attached Figure Description
[0016] Figure 1This is a schematic diagram of the architecture of the emotion prediction system based on multiple biological signals provided in the embodiments of this application; Figure 2 This is a flowchart of the emotion prediction method based on multiple biological signals provided in the embodiments of this application; Figure 3 This is a schematic diagram of the functional modules of the emotion prediction device based on multiple biological signals provided in the embodiments of this application; Figure 4 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] In various scenarios requiring continuous control (such as driving vehicles, operating aircraft, operating construction machinery, or engaging in virtual reality interaction), the operator's emotional state directly impacts their behavioral decisions and operational safety. Real-time monitoring of emotional states can provide crucial information for behavioral intervention, risk warning, and human-machine collaboration, helping to reduce the risk of accidents caused by negative emotions such as anger, tension, and fatigue, and improving the control experience and safety.
[0021] In related technologies, emotion detection is generally performed through facial expression recognition. Specifically, a camera captures images of the operator's face, and facial features such as eye movements, mouth shapes, and eyebrows are identified to determine the emotional state. However, the accuracy of emotion recognition decreases when the operator wears protective gear (helmet, mask, glasses), or when there are changes in lighting, head tilting, or when the operator consciously controls their expressions to conceal their inner state. Furthermore, this method is a reactive approach, which can easily lead to problems with untimely intervention.
[0022] Based on this, embodiments of this application provide a method, apparatus, device, and medium for emotion prediction based on multiple biological signals, which can improve the accuracy of emotion recognition and the timeliness of emotion intervention.
[0023] The emotion prediction method, apparatus, device, and medium based on multiple biological signals provided in this application are specifically described through the following embodiments. First, the emotion prediction system based on multiple biological signals in the embodiments of this application are described.
[0024] Please refer to Figure 1 In some embodiments, this application provides an emotion prediction system based on multiple biological signals, including a terminal 11 and a server 12.
[0025] In some implementations, terminal 11 can be a complete edge intelligent hardware device. Specifically, the target device can be a directional control device (such as a smart steering wheel, aviation side stick, game joystick, etc.) integrating tactile sensors, physiological signal sensors, and pressure sensors, with a microcontroller (MCU) or system-on-a-chip (SoC) embedded inside or on its surface, forming a complete edge intelligent hardware device. Terminal 11 can independently complete signal acquisition, feature extraction, emotion trend prediction, and soft intervention triggering at the control site, without relying on real-time cloud communication, to ensure low latency and privacy security.
[0026] In some implementations, the internal architecture of terminal 11 may include a hardware layer, a signal acquisition and processing layer, an algorithm inference layer, and an interaction output layer. Specifically, the hardware layer may include a swept frequency capacitive sensing (SFCS) membrane, an electrocardiogram (ECG) and Galvanic Skin Response (GSR) electrode array, a pressure / grip sensor, and an embedded MCU / SoC; the signal acquisition and processing layer can perform multi-channel synchronous sampling, filtering, and tactile contour extraction; the algorithm inference layer can realize micro-gesture detection, multimodal feature fusion, and short-term trend prediction; the interaction output layer can output emotion prediction summaries via CAN / Ethernet / SDK and drive soft intervention feedback such as lighting and vibration.
[0027] In some implementations, the sensor material layer structure used in the terminal 11 (e.g., the composite film layer in the steering wheel grip area) can be, from top to bottom, as follows: a conductive layer (TPU conductive film, i.e., the outermost layer that the target object's hand directly contacts), a copper layer (ECG electrode) for heart rate detection, a first fiber layer (silver fiber conductive fabric), an EmoSense tactile sensing layer (SFCS double-sided electrode), a second fiber layer (insulating layer), and the main structure of the steering wheel. This multi-layer design achieves the integrated integration of tactile sensing, physiological signal acquisition, and structural support.
[0028] In some implementations, the steering control device can be a car steering wheel. When the driver holds the steering wheel with both hands, the SFCS membrane layer (EmoSense layer) outputs a 160-dimensional tactile contour subsequence at a sampling rate of 20Hz; the MCU locates the hand position based on the tactile contour amplitude, dynamically selects the ECG / GSR electrode with the lowest impedance for physiological signal acquisition, and simultaneously reads the grip force data from the pressure sensor. In the algorithm inference layer, the micro-gesture detection model identifies the normal grip state without micro-gestures. When the driver begins to exhibit agitated rubbing movements, the system detects the micro-gesture event in real time and calculates the micro-gesture dynamics time series. Subsequently, the lightweight large language model built into the terminal 11, combined with the individual's emotional baseline, outputs the predicted emotional changes for the next 30 seconds (increased arousal, decreased valence). Based on this, the interaction output layer changes the steering wheel LED light strip from green to yellow and triggers a gentle vibration. The entire closed-loop process proceeds in an orderly manner between the layers, with the hardware layer and signal acquisition layer responsible for data acquisition, the algorithm layer responsible for intelligent judgment, and the interaction layer responsible for feedback execution. During this process, terminal 11 only intermittently uploads the desensitized risk level summary and event count to server 12, without uploading the original signal in real time.
[0029] Furthermore, the conductive layer can be made of a 0.1mm thick thermoplastic polyurethane (TPU) conductive film, providing a wear-resistant and sweat-proof surface; the copper layer is an etched flexible printed circuit board (FPC) electrode used to collect electrocardiogram signals; the first fiber layer uses silver fiber conductive fabric (0.2mm thick), combined with a cotton thread layer (0.4mm thick) to provide flexibility and conductive stability; the EmoSense layer is a double-sided electrode swept-frequency capacitive sensing unit, whose side facing the hand forms a contact sensing interface with the conductive layer and copper layer, while the side facing the steering wheel body is insulated from the metal frame by a second fiber layer (non-conductive felt). The layers are fixed together by thermoforming or bonding. This multi-layer structure is easily packaged into a detachable steering wheel kit, adaptable to steering wheels of different diameters and shapes. Experiments have shown that this material layer structure improves the signal-to-noise ratio by approximately 40% compared to a single-layer electrode under dynamic control, and is adaptive to sweat and temperature changes, thereby improving the accuracy and robustness of emotion prediction.
[0030] In some implementations, server 12 can be used for aggregate analysis of de-identified statistical features from multiple terminals 11, global model updates, cross-scenario transfer optimization of individualized sentiment baselines, and long-term trend mining. Server 12 can be a cloud server, edge computing node, or local workstation, equipped with high-performance computing units (GPU, TPU) and large-capacity storage; the backend software system deployed on it realizes group baseline statistics, periodically retrains the general large language model to improve prediction accuracy, generates personalized transfer learning alignment parameters for different operators or different device forms, and processes risk assessment reports and scheduling suggestions in fleet or aircraft management scenarios. In addition, server 12 is also responsible for distributing updated model weights, individual baseline calibration coefficients, and intervention strategy parameters to each terminal 11. During the power-on initialization phase, terminal 11 obtains the general pre-trained model and initial baseline parameters from server 12. During operation, it only uploads the desensitized statistical features (such as the trend of pressure level change, micro-gesture frequency summary, and risk event count) to server 12 at a low frequency (such as per minute or per stroke), without uploading the original tactile contour, physiological waveform, and grip force data.
[0031] Furthermore, the server 12 can receive anonymized statistical features uploaded by multiple terminals 11, perform cross-individual and cross-scenario model optimization and baseline drift compensation analysis, and send the updated model parameters and individualized calibration coefficients back to the corresponding terminal 11. This ensures both the timeliness and privacy compliance of real-time emotion intervention, and also enables the continuous evolution of model performance and the transfer of group knowledge.
[0032] The emotion prediction method based on multiple biological signals in this application can be illustrated through the following embodiments.
[0033] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.
[0034] In this application embodiment, the description will focus on a multiple biosignal-based emotion prediction device, which can be integrated into a computer device. See [link to relevant documentation]. Figure 2 , Figure 2The flowchart illustrates the steps of the emotion prediction method based on multiple biosignals provided in this application embodiment. The method is applied to a target device, which includes at least a tactile sensor, a physiological signal sensor, and a pressure sensor. This application embodiment takes the example of an emotion prediction device based on multiple biosignals specifically integrated into a terminal or server. When the processor on the terminal or server executes the program instructions corresponding to the emotion prediction method based on multiple biosignals, the specific process is as follows: Step 101: Obtain the tactile contour subsequence collected by the tactile sensor at each time step, and determine the micro-gesture event at each time step based on the tactile contour subsequence.
[0035] In some implementations, in order to accurately separate and semantically represent the subtle hand movements that reflect subconscious emotional states generated by the operator (target object) during the natural grip direction manipulation of the device, which are generated from the continuous, high-dimensional and noisy raw tactile perception data stream, the high-resolution contact contour characteristics of swept-frequency capacitive tactile sensing can be utilized. The capacitance-frequency response vector obtained at each discrete sampling moment can be used as the basic analysis unit, and a machine learning model can be combined to identify the specific spatial morphology and temporal pattern corresponding to behaviors such as rubbing, friction, and tapping. This will construct the smallest semantic unit (i.e., micro-gesture event) that can be directly processed by the subsequent emotion prediction model, thereby achieving early, accurate, and anchored subconscious emotion-driven behavior.
[0036] Among them, the tactile sensor can be a sensor unit array built based on the principle of swept-frequency capacitive tactile perception. For example, it can be a flexible conductive fabric layer with double-sided electrodes arranged in the grip area of the steering wheel. By applying AC excitation signals at multiple frequency points to the electrodes, the resistance and capacitance responses at the skin-electrode interface that change with frequency are obtained, so as to generate a high-dimensional capacitance-frequency profile vector characterizing subtle tactile interactions such as contact area, pressure, sliding, and rubbing.
[0037] The time step can be the smallest discrete time unit for the system to synchronously sample and process tactile and physiological signals, and can be used to construct the basic analytical scale of the continuous signal flow on the time axis. For example, it can correspond to the period of each complete frequency sweep of the tactile sensor at a fixed frequency of about 20 Hz and the output of a set of tactile contour vectors.
[0038] The tactile profile subsequence can be a set of high-dimensional vector data output by the tactile sensor after completing a full frequency sweep within a single time step. For example, it can be a capacitance-frequency profile vector composed of approximately 160 values representing the amplitude or phase of the capacitance response at different frequency points. This vector quantifies the complex skin-electrode interface state formed between the palm and the sensor at the current moment and can be used to characterize the instantaneous tactile spatial fingerprint generated by actions such as rubbing and friction.
[0039] Among them, micro-gesture events can be short-term, subtle hand movements driven by the operator's subconscious emotions, determined by a pre-defined classification model based on one or more continuous time step tactile contour subsequences. For example, they can be unconscious rubbing, finger tapping, palm rubbing, or local squeezing movements identified using a random forest model. Compared to macroscopic manipulation behaviors (such as turning), these movements have a smaller spatial amplitude and specific tactile contour evolution patterns, and can be used as a reliable and difficult-to-conceal behavioral representation of the operator's true emotions and stress state.
[0040] In some implementations, the tactile sensor can be a swept frequency capacitive sensing (SFCS) unit, which is arranged in the grip area of the directional control device. In order to obtain high-dimensional tactile information that can finely characterize the hand contact state, in each discrete time step, the system can control the MCU to apply AC excitation signals at multiple frequency points to the SFCS electrodes, and collect the amplitude or phase of the capacitive response signal at each frequency point, thereby forming a high-dimensional vector characterizing the electrical characteristics of the skin-electrode interface at the current moment, i.e., the tactile contour subsequence.
[0041] Specifically, taking a car steering wheel as an example, the SFCS double-sided electrode flexible film layer is embedded in the grip sleeve at the 3 o'clock and 9 o'clock positions of the steering wheel. The system performs a frequency sweep operation at a sampling rate of about 20Hz (i.e., the interval between adjacent time steps is about 50 milliseconds). Each frequency sweep traverses 160 frequency points, and records a capacitance response amplitude at each frequency point. Thus, each time step outputs a tactile contour vector containing 160 values.
[0042] For example, to improve signal stability, the system can normalize the vector after acquisition and use sliding median filtering to remove high-frequency noise caused by sweat or momentary contact jitter. To cope with single-handed operation or hand-switching scenarios, the MCU can also dynamically adjust the electrode scanning order based on the tactile heatmap of the previous moment, prioritizing the scanning of the expected palm area to ensure that the tactile contour subsequence can still be continuously output when the control posture changes.
[0043] For example, in a real test, when the driver's right palm is stably resting on the 9 o'clock position of the steering wheel without any micro-gestures, the tactile contour subsequence output at a certain time step can be represented as a vector [0.12, 0.35, 0.78, ..., 0.22] (160 dimensions in total), where the value of the k-th dimension corresponds to the frequency. The normalized capacitive response amplitude was measured. When the driver began unconsciously rubbing their thumb, the amplitudes of multiple frequency points (typically in the mid-to-high frequency range) related to the thumb contact area exhibited regular fluctuations. For example, the values in the 50th to 80th dimensions changed from a range of 0.5–0.7 to 0.8–1.2, showing a periodic rise and fall. The tactile contour subsequences at each time step were stored in a circular buffer according to their time index, forming a tactile contour subsequence of length N (e.g., corresponding to a 2-second time window), providing the raw data foundation for subsequent micro-gesture detection.
[0044] Furthermore, after obtaining the tactile contour subsequence for each time step, the system can further identify whether that time step contains micro-gesture events and the specific category of the micro-gestures. Specifically, the system can perform differential feature extraction on the tactile contour subsequence of the current time step, such as calculating the gradient between adjacent dimensions, the number of contour peaks, and the change in Euclidean distance relative to several previous time steps. Subsequently, these features are input into a pre-trained classification model.
[0045] In some implementations, the classification model can employ a random forest algorithm. Its input consists of approximately 20 dimensions of hand-crafted statistical features (including contour mean, variance, peak frequency, low-frequency component energy, etc.), and its output is a binary label ("with micro-gesture" or "without micro-gesture") or a category label (such as "rubbing," "tapping," "friction," "squeezing," etc.). Training data comes from tactile contour sequences collected from multiple subjects performing standard micro-gesture actions in a simulated cockpit, and these sequences are manually labeled. Taking the "rubbing" micro-gesture as an example, its typical spatial pattern is characterized by synchronized amplitude pulses in multiple dimensions corresponding to the contact areas of the thumb and index finger in the tactile contour. The temporal pattern is characterized by sinusoidal oscillations in amplitude changes over 3-5 consecutive time steps. The random forest model determines whether the current time step belongs to the rubbing category by analyzing the voting results of multiple decision trees. To reduce latency, the inference process is entirely completed on the MCU inside the direction control device (e.g., using lightweight C language to implement decision tree inference), with the processing time for each time step controlled within 10 milliseconds to ensure real-time performance. After identifying a micro-gesture event, the system simultaneously records the event's timestamp, category confidence level, and the original contour index that triggered the event for subsequent dynamic feature extraction.
[0046] Through the above methods, a precise mapping from chaotic physical contact signals to structured emotional and behavioral characteristics can be achieved, thus providing a high-fidelity, low-latency, and privacy-friendly data foundation for the subsequent construction of micro-gesture dynamics time series describing the dynamic evolution of emotions. This significantly improves the robustness and early warning capability of the emotion prediction system in complex manipulation environments.
[0047] Step 102: Determine the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence collected at each time step, and construct the micro-gesture dynamic time series corresponding to the current target time window by combining the dynamic feature subsequence of each time step.
[0048] In some implementations, in order to elevate discrete, isolated micro-gesture events into multidimensional dynamic features that can characterize their frequency of occurrence, rhythm intensity, and evolution acceleration, parameters can be extracted from the tactile contour subsequence of each time step to construct a continuous feature sequence that reflects the spatiotemporal evolution of unconscious hand behavior driven by emotions. This provides a physically meaningful and dominant dynamic input for subsequent prediction of emotional change trends.
[0049] Among them, the dynamic feature subsequence can be a set of quantized parameter vectors extracted for micro-gesture events occurring at a single time step, which can be used to map microscopic tactile morphological changes into dynamic features with physical meaning and temporal evolution attributes.
[0050] The target time window can be a time interval used to aggregate dynamic feature subsequences of multiple consecutive time steps to form a complete analysis unit. For example, it can be a sliding time window that can be set to 30 seconds to 120 seconds and can be used to limit the time domain range of micro-gesture dynamics modeling.
[0051] Among them, the micro-gesture dynamics time series can be a high-dimensional vector sequence formed by arranging the dynamic feature subsequences corresponding to each time step within the current target time window in chronological order. For example, it can be a matrix formed by stacking the micro-gesture frequency, rhythmicity index, cumulative hand displacement and pressure trend slope of multiple consecutive time steps (corresponding to about 2.5 seconds to several seconds, depending on the sampling rate and window length) according to the time index. It can be used to characterize the frequency change, rhythmic fluctuation and pressure accumulation trend of micro-gesture behavior in the time dimension.
[0052] Specifically, the system can count the number of micro-gesture events occurring within a preset, sliding target time window (e.g., 2 seconds long, containing approximately 40 consecutive time steps) with the current time step as the endpoint, and divide by the window duration to obtain the micro-gesture frequency (MGFreq, unit: times / minute); at the same time, it calculates the average tactile contour amplitude corresponding to all time steps in the window where micro-gesture events are detected, as the average rubbing intensity.
[0053] Furthermore, to characterize the clustering of micro-gestures over time, the system can calculate a rhythmicity index, the value of which can be determined as follows: First, obtain the time interval sequence of adjacent micro-gesture events within the window, calculate the coefficient of variation (the ratio of standard deviation to mean) of this sequence, and take the reciprocal of this coefficient of variation as the rhythmicity index; or use an autocorrelation function to extract the intensity of the main peak. A higher rhythmicity index value indicates that the micro-gestures exhibit a sudden, clustered distribution or a regular rhythm, while a lower value indicates random dispersion.
[0054] For example, the cumulative hand displacement can be obtained by accumulating the change in tactile contour between the current time step and the adjacent time steps, where the change at each time step is calculated based on the Euclidean distance between two adjacent tactile contour subsequences.
[0055] Furthermore, the pressure trend slope can be obtained by linearly fitting the micro-gesture event intensity values (which can be tactile contour amplitude or rubbing intensity) of the current time step and several preceding time steps (e.g., the first 5 time steps), and the slope of the fitted line is the pressure trend slope.
[0056] In some implementations, the five parameters mentioned above—micro-gesture frequency, average rubbing intensity, rhythmic index, cumulative hand displacement, and pressure trend slope—can be organized into a five-dimensional vector in a predetermined order, thus forming the dynamic feature subsequence corresponding to the current time step.
[0057] For example, in a driving test in a congested area, the system continuously collected tactile contours at a sampling rate of 20Hz. When the driver began to intermittently tap the steering wheel due to frustration, the system detected a "tapping" micro-gesture event at time step t. Using this time step as the endpoint, statistics were performed on the preceding 2-second time window (out of a total of 40 time steps): A total of 6 micro-gesture events were detected within this window, resulting in a micro-gesture frequency of 180 times / minute; the average rubbing intensity (using the root mean square of the tactile contour amplitude) was 0.72 (normalized units); the time intervals between adjacent micro-gestures were 0.35 seconds, 0.32 seconds, 0.40 seconds, 0.38 seconds, and 0.42 seconds, respectively, with a coefficient of variation of approximately 0.09 and a rhythmicity index of 11.1, indicating a highly regular tapping motion; the cumulative hand displacement was obtained by summing the Euclidean distances between every two adjacent contours in the 40 time steps, with a value of 120 (relative units); linear fitting was performed on the intensity values of the first 5 micro-gesture events (0.65, 0.68, 0.70, 0.73, and 0.72, respectively), yielding a pressure trend slope of +0.015, showing a slight upward trend. Therefore, the dynamic feature subsequence of the current time step can be represented as [180,0.72,11.1,120,0.015].
[0058] In some implementations, after obtaining the dynamic feature subsequence for each time step, the system organizes the dynamic feature subsequences of multiple consecutive time steps into a high-dimensional time series, i.e., a micro-gesture dynamics time series, according to the chronological order of the time steps. The construction of this series depends on a configurable target time window (e.g., 30 to 120 seconds). Specifically, at the start of the current target time window, the system clears the buffer; subsequently, the dynamic feature subsequences generated at each time step are appended to the end of the buffer until the window duration is reached. Each row of this series corresponds to a time step, and each column corresponds to a dynamic feature dimension (frequency, intensity, rhythm, displacement, slope).
[0059] In some implementations, to improve the efficiency of subsequent prediction models, the system can also downsample the micro-gesture dynamics time series in the time dimension (e.g., taking the mean every 5 time steps) or normalize the feature dimension (subtracting the individual baseline mean and dividing by the standard deviation). For example, with a target time window of 30 seconds and a sampling rate of 20Hz, if downsampling is not performed, the micro-gesture dynamics time series contains 600 time steps, each time step corresponding to a 5-dimensional feature vector, forming a 600×5 matrix. This matrix records the entire process of the driver's micro-gesture frequency gradually increasing from 120 times / minute to 210 times / minute, the average rubbing intensity increasing from 0.5 to 0.9, the rhythmic index stabilizing at around 10, the cumulative hand displacement continuously increasing, and the pressure trend slope changing from positive to negative within 30 seconds, providing a dominant input for subsequent multimodal trend prediction.
[0060] In some implementations, the dynamic feature subsequence can further incorporate the category transition entropy of micro-gestures. Specifically, in addition to statistical frequency and intensity, the system records the category transition probability between adjacent micro-gesture events, such as the frequency of transition from rubbing to tapping, and calculates its Shannon entropy (category transition entropy) as an additional feature dimension, incorporating it into the dynamic feature subsequence. This category transition entropy reflects the instability of the operator's emotional state: when emotions worsen, micro-gesture categories often switch randomly between multiple actions, leading to an increase in transition entropy; while repetitive single-pattern movements (such as only rubbing) result in lower entropy values. This can improve the accuracy of predicting pre-anger states.
[0061] Specifically, the system can preset a set of micro-gesture categories. For example, k=4 corresponds to four categories: "rubbing," "tapping," "friction," and "squeezing." Within a target time window (e.g., 30 seconds), the system records the category of micro-gesture events detected at each time step in chronological order, forming a category sequence. Then, it counts the number of category transitions between all adjacent events (i.e., the i-th event and the (i+1)-th event) to construct a [class sequence]. The transition counting matrix M, where Indicates from category Transfer to Category The number of occurrences. Then the counting matrix is converted into a probability matrix. : ; in, Indicates from category The total number of all transfers from the starting point. If this sum is zero (i.e., category...). If it does not appear in the current window, then let =0. Subsequently, the category transition entropy H can be calculated using the following formula: ; in, It can be a very small positive number (e.g.) This is used to avoid infinity when taking the logarithm of zero probability. H reaches its maximum value when the event categories are completely unordered and all transitions occur with equal probability. When category transitions are highly regular (e.g., always rubbing → tapping → rubbing), H is relatively small. The system adds this H value as an additional dimension to the dynamic feature subsequence of each time step within the current target time window (or as a statistical feature of the window as a whole). By introducing category transition entropy, the system can quantify the degree of disorder in the operator's micro-gesture behavior at the category level, thereby more sensitively capturing the instability before emotional outbursts and improving the accuracy of predicting pre-anger and pre-fatigue.
[0062] Through the above methods, continuous temporal modeling can be achieved, from discrete behavioral labels to reflecting the spatiotemporal evolution of unconscious behaviors driven by emotions. This provides a dynamic feature foundation that is dominant, has high temporal resolution, and is difficult to be subjectively faked, for subsequent multimodal emotion trend prediction by combining physiological signal sequences and grip force sequences.
[0063] In some implementations, to quantify and extract the multidimensional dynamic attributes inherent in the original micro-gesture events that characterize the evolution of emotion-driven behavior, the frequency of micro-gestures can be obtained by counting the number of occurrences within a sliding time window, the average rubbing intensity and rhythmicity index can be obtained by analyzing the amplitude and rhythmic clustering of tactile contours, the cumulative hand displacement can be obtained by accumulating the tactile contour changes between adjacent time steps, and the pressure trend slope can be obtained by linearly fitting the micro-gesture event intensity sequence. This constructs multiple dynamic feature parameters including frequency, intensity, rhythm, cumulative displacement, and acceleration, thereby assigning a dynamic feature subsequence with temporal evolution semantics to the micro-gesture event at each time step. For example, step 102, "determining the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence collected at each time step," may include: (102.1) For each time step, determine the corresponding sliding sub-time window with each time step as the endpoint; (102.2) Based on the tactile contour subsequence of at least one time point contained in the sliding sub-time window, determine the micro-gesture frequency, average rubbing intensity of the detected target object on the skin contact area corresponding to the tactile sensor, rhythmic index, cumulative hand displacement and pressure trend slope corresponding to the micro-gesture event; Among them, the rhythmicity index is used to characterize the degree of aggregation of corresponding micro-gesture events on the time axis within the sliding sub-time window. The cumulative hand displacement at each time step is obtained by accumulating the tactile contour changes between adjacent time steps contained within the sliding sub-time window. The tactile contour changes at each time step are determined based on the differences between the tactile contour subsequences between adjacent time steps contained within the sliding sub-time window. The pressure trend slope at each time step is obtained by linearly fitting the intensity values of micro-gesture events at multiple time steps contained within the sliding sub-time window. (102.3) Based on the micro-gesture frequency, average kneading intensity, rhythmic index, cumulative hand displacement and pressure trend slope corresponding to the sliding sub-time window of each time step, determine the dynamic feature sub-sequence corresponding to each time step.
[0064] Among them, the sliding sub-time window can be a time interval with a preset fixed duration (such as 2 seconds). This interval includes the tactile contour subsequence of the current time step and several consecutive time steps before it, which is used to statistically analyze the local tactile change features near the current time step. The micro-gesture frequency can be calculated by counting the total number of micro-gesture events within a sliding sub-time window, dividing the total number of events by the window duration (in minutes), and obtaining the number of events per minute. For example, if 6 events are detected within a 2-second window, the frequency is 180 times / minute, which is used to quantify the level of activity of the operator's emotion-driven hand movements per unit of time.
[0065] Among them, the skin contact area can be the continuous sensing range in which the operator's hand and the tactile sensor form effective physical contact, such as the electrode array area covered by the palm and fingers at the 3 o'clock and 9 o'clock positions of the steering wheel. The position and area of this area can be dynamically located by analyzing the electrode unit cluster with the largest amplitude in the swept-frequency capacitive tactile contour vector. This can be used to guide the physiological signal sensor to select the sampling channel with the lowest impedance, and to calculate the average rubbing intensity on this area.
[0066] The average rubbing intensity can be calculated by extracting the amplitude (e.g., root mean square of capacitance response) of the tactile contour subsequence at each time step of all detected micro-gesture events within the sliding sub-time window, and taking the arithmetic mean to reflect the average excitation intensity of rubbing, friction and other actions, thus characterizing the level of emotional arousal.
[0067] The rhythmicity index can be a sequence of time intervals between adjacent micro-gesture events within a sliding sub-time window. The degree of clustering of events on the time axis is quantified by the reciprocal of its coefficient of variation or the peak of its autocorrelation. The more uniform the intervals (higher the rhythmicity), the larger the value; the more disordered the intervals, the smaller the value. This index can be used to distinguish between regular micro-gestures under sustained stress and disordered movements caused by sudden emotional fluctuations.
[0068] The change in tactile contour can be the difference between the tactile contour subsequences of two adjacent time steps within a sliding sub-time window, for example, measured by Euclidean distance or cosine distance. The cumulative hand displacement can be calculated by summing the changes in tactile contours between all adjacent time steps within a sliding sub-time window, resulting in the net cumulative sliding or rotational amount of the hand on the sensor surface. The cumulative hand displacement can be used to distinguish between macroscopic manipulation actions (such as continuous displacement caused by steering) and localized small-amplitude displacements caused by micro-gesture events.
[0069] The pressure trend slope can be calculated by performing a univariate linear fit on the intensity values of micro-gesture events (such as tactile contour amplitude or rubbing intensity) across multiple time steps (the current time step and several preceding time steps, e.g., the previous 5 to 10 time steps) within a sliding sub-time window. The slope of the fitted line is the pressure trend slope. A positive pressure trend slope indicates an upward trend in intensity, suggesting a possible deterioration in mood; a negative value indicates a downward trend.
[0070] The time axis can be a one-dimensional time coordinate formed by arranging discrete time steps in the order of the first time step of the sliding sub-time window. For example, the sequential index of the first sample, the second sample, ... the Nth sample can be used as the position mark on the time axis. Each time step corresponds to a fixed time interval (such as 0.05 seconds, corresponding to a 20Hz sampling rate).
[0071] Adjacent time steps can be two consecutive and directly connected sampling moments on the time axis of the sliding sub-time window, that is, the current time step and at least the previous time step that is immediately adjacent to it.
[0072] In some implementations, the system can continuously acquire tactile contour subsequences at a fixed sampling rate (e.g., 20Hz, i.e., 50 milliseconds between adjacent time steps). To assign statistical characteristics within the local time domain to each time step, the system determines a corresponding sliding sub-time window by pushing back a preset duration (e.g., 2 seconds) from the current time step t for each time step where dynamic feature subsequences need to be calculated. The starting point of this sub-time window is... (in The total number of time steps contained within the window, for example =2 seconds × 20Hz = 40 time steps), with the endpoint being the current time step t. The sliding sub-time window slides forward with each time step, meaning each time step corresponds to a window with that time step as its endpoint.
[0073] For example, the sliding sub-time window can be configured according to the application scenario (e.g., from 1 second to 5 seconds). This sliding sub-time window can be used to aggregate tactile evolution information around the current time step, providing a data foundation for subsequent calculations of statistics such as micro-gesture frequency and rubbing intensity.
[0074] Furthermore, within each sliding sub-time window, the system can use the tactile contour sub-sequences of all time steps contained within the sliding sub-time window and the detected micro-gesture events to calculate the following five parameters: Micro-gesture frequency: The total number of micro-gesture events occurring within a sliding sub-time window. Divide by the window duration (converted to minutes), that is ,in The duration of a single time step (e.g., 0.05 seconds) is given, with the result in times per minute. For example, if 6 micro-gestures are detected within a 2-second window, the frequency is 180 times per minute.
[0075] Average rubbing intensity: For all time steps within the window where micro-gesture events are detected, extract the amplitude of the tactile contour subsequence at each time step (e.g., the root mean square of the capacitive response at 160 frequency points) and calculate its arithmetic mean. If there are no micro-gesture events within the window, then take the tactile contour background amplitude at the current time step.
[0076] Rhythmicity metrics: Obtain the time interval sequence between adjacent micro-gesture events within a window. (where m is the number of events), calculate the mean of the sequence. with standard deviation Rhythmic indicators ( (To prevent division by zero for extremely small positive numbers). The larger the R value, the more uniform the event intervals and the stronger the rhythmicity.
[0077] Cumulative hand displacement: The changes in tactile contour between all adjacent time steps within the window are summed. The change between each adjacent time step is expressed as Euclidean distance. ,in Let be the tactile contour vector at time step i. Cumulative displacement .
[0078] Pressure trend slope: From all time steps within the window, select the intensity values (which can be haptic contour amplitudes) of the current time step and the previous L time steps (e.g., L=5) that have micro-gesture events, to form a sequence. ,in The time step number. Let k be the intensity value. Performing a univariate linear regression on the above sequence, the slope k of the fitted line represents the pressure trend slope. A positive slope indicates increasing intensity, while a negative slope indicates decreasing intensity.
[0079] For example, in a test on a congested road section, at the end of the current time step t=100 seconds, a sliding sub-time window of 2 seconds (40 time steps) was taken. Six "tapping" events were detected within the window, with a micro-gesture frequency of 180 times / minute; the tactile contour amplitude (normalized) of each event at the corresponding time step was 0.65, 0.68, 0.70, 0.73, 0.72, and 0.71, respectively, with an average rubbing intensity of 0.698; the intervals between adjacent events were 0.35, 0.32, 0.40, 0.38, and 0.42 seconds, with a rhythmicity index R≈9.8; the cumulative Euclidean distance between adjacent time steps yielded a cumulative hand displacement of 120 (relative units); linear fitting was performed on the current time step and the previous 5 intensity values, resulting in a pressure trend slope of approximately +0.015.
[0080] Furthermore, the five parameters calculated above can be combined in a fixed order (micro-gesture frequency, average kneading intensity, rhythmic index, cumulative hand displacement, and pressure trend slope) into a five-dimensional vector, which is the dynamic feature subsequence corresponding to each time step. The system independently executes the above steps for each time step, thereby generating a dynamic feature subsequence for each discrete time step. These subsequences, arranged by time index, constitute the micro-gesture dynamics time series required for subsequent steps.
[0081] In some implementations, when no micro-gesture events are detected within the sliding sub-time window, the micro-gesture frequency and rhythmicity indices are both set to 0, the average kneading intensity is taken as the tactile contour background amplitude of the current time step, the cumulative hand displacement is calculated normally (reflecting the displacement of the manipulation action), and the pressure trend slope is set to 0 to maintain the continuity of the time series data. In this way, each time step obtains a feature vector that can characterize the local tactile dynamic evolution.
[0082] For example, in addition to using a fixed window length, an adaptive sliding window length mechanism can be introduced: the average interval of micro-gesture events within the window is detected in real time. When the average interval is less than a threshold (e.g., 0.3 seconds), the window length is automatically shortened (e.g., from 2 seconds to 0.5 seconds) to improve the response speed to sudden emotional fluctuations; when the average interval is larger (e.g., greater than 1 second), the window length is appropriately extended (up to 5 seconds) to smooth background noise. This mechanism can maintain the stability of feature extraction under different operating conditions such as intense driving or smooth cruising.
[0083] Through the above methods, it is possible to extract continuous features from single event detection to reflect the spatiotemporal evolution of emotion-driven behavior, thereby providing a dominant, high-information-density, and difficult-to-subjectively-falsify dynamic feature foundation for subsequent construction of micro-gesture dynamic time series and short-term emotion trend prediction combined with physiological signals.
[0084] Step 103: For each time step, calculate the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence, and select the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area.
[0085] In some implementations, in order to ensure the continuity and signal-to-noise ratio of physiological signals when poor contact of the predetermined physiological electrodes occurs due to single-handed operation, hand switching, or changes in grip posture during dynamic control, the area with the largest amplitude (characterizing the lowest contact impedance and the clearest signal) can be dynamically set as the current physiological signal acquisition area. This allows for the use of high spatial resolution tactile perception to guide the automatic and optimal switching of physiological electrode channels in real time, thereby solving the problem of signal interruption or quality degradation of traditional fixed electrodes in dynamic grip scenarios.
[0086] The tactile contour amplitude can be a quantified value of the intensity of the capacitive response signal output by the tactile sensor at one or more swept frequency points on a specific skin contact area. For example, it can be the Euclidean norm or peak value of the phase or amplitude response vector measured by the swept frequency capacitive tactile sensing unit. The magnitude of this amplitude is positively correlated with the contact pressure, contact area, and the tightness of the skin-electrode interface. It can be used to evaluate the signal quality of different gripping areas at the current moment, serving as a basis for dynamically and optimally switching physiological electrodes.
[0087] The physiological signal acquisition area can be a local spatial range in the sensor array of the steering control device, dynamically selected based on the tactile contour amplitude distribution of the current time step, where one or more optimal electrode units are located for acquiring electrocardiogram or skin conductance signals through physiological signal sensors. For example, in the multi-unit ECG / GSR electrode array at the 3 o'clock position of the steering wheel, the MCU determines the pair of units with the largest amplitude (usually corresponding to the center of the palm or the thenar eminence where the muscles are fully in contact) based on the tactile thermogram and closes them into the acquisition circuit. This can be used to replace the preset fixed electrodes to achieve adaptive following of the physiological sampling channel in scenarios such as hand switching, sliding, or loosening the grip, ensuring continuous and effective acquisition of heart rate variability and skin conductance signals.
[0088] In some implementations, to evaluate the adhesion quality between different skin contact areas and the tactile sensor, the tactile contour amplitude corresponding to each skin contact area can be calculated for each tactile contour subsequence acquired at each time step. The tactile sensor can consist of multiple independent sensing units (or electrode pairs), each covering different positions of the grip area of the steering control device (e.g., the 3 o'clock, 9 o'clock, and back areas of the steering wheel). Within each time step, each sensing unit outputs a set of swept-frequency capacitive response vectors, the amplitude of which at different frequency points reflects the contact area, pressure, and adhesion tightness between the unit and the skin.
[0089] Furthermore, to obtain the tactile contour amplitude of this unit, energy synthesis can be performed on the response amplitudes at each frequency point in the vector, for example, by calculating the root mean square or the vector magnitude. Taking a sensing unit in the 3 o'clock region of a steering wheel as an example, if its output capacitance amplitudes at 160 frequency points are respectively Then the tactile contour amplitude of the unit Alternatively, a simpler peak method can be used, taking the maximum value in the vector as the amplitude. In a practical test, when the driver's right palm was stably pressed against the 3-point area, the tactile contour amplitude of that area was calculated to be 0.82 (normalized units); while at the same time step, the amplitude of the 9-point area was only 0.12 due to the hand being removed. The system repeats the above calculation for each time step to obtain the tactile contour amplitude of each skin contact area at the current moment.
[0090] Understandably, during dynamic manipulation, operators may use one hand, switch hands, slide their grip, or adjust their posture, leading to poor contact or decreased signal quality of the preset physiological signal electrodes (ECG / GSR). To address this issue, this application utilizes tactile contour amplitude as a proxy indicator of contact quality. A larger amplitude indicates a tighter skin-electrode interface and lower contact impedance, making it most suitable for acquiring physiological signals.
[0091] Specifically, after calculating the tactile contour amplitude of each region at each time step, the amplitude of all regions (e.g., the six candidate regions on the steering wheel) can be compared, and the region with the largest amplitude can be selected as the physiological signal acquisition region for the current time step. Subsequently, the MCU controls the analog switch of the physiological signal sensor (ECG / GSR array) to dynamically close the electrode unit pair corresponding to that region and acquire ECG or ductal skin signals from that region. For example, when the driver changes from holding the steering wheel with both hands to gripping it with one hand at the 3 o'clock position during a turn, the tactile contour amplitude of the 3 o'clock region is significantly higher than that of other regions. The system immediately switches the physiological acquisition channel from the default 9 o'clock region to the 3 o'clock region, thereby ensuring continuous and effective acquisition of heart rate variability and ductal skin signals. Thus, in scenarios such as hand-switching operations and aggressive driving, the loss of multimodal data due to brief interruptions in electrode contact can be effectively avoided.
[0092] Through the above methods, physiological signals can be maintained without interruption and with a high signal-to-noise ratio in dynamic control scenarios such as single-handed operation, hand switching, or changes in grip posture. This provides a seamless and adaptive hardware-algorithm collaboration guarantee for the subsequent construction of continuous and reliable physiological signal sequences, significantly improving the robustness of the multimodal emotion prediction system in actual control environments.
[0093] Step 104: In each time step, physiological signals are collected from the physiological signal acquisition area by a physiological signal sensor, and combined with the physiological signals corresponding to each time step, a physiological signal sequence corresponding to the target time window is constructed.
[0094] In some implementations, to ensure the continuity and temporal consistency of physiological signals acquired from the optimal acquisition area after dynamic selection and switching, raw physiological data (such as ECG waveforms and skin conductance values) can be synchronously sampled from the determined physiological signal acquisition area by a physiological signal sensor in each time step. The data is then organized into a sequence structure aligned with the micro-gesture dynamics time window according to the order of the time steps. This constructs a synchronous physiological signal sequence that can be used for emotional state correction, high-interference scene takeover, and individual baseline update, thereby providing high-quality and strictly aligned auxiliary modal input for multimodal temporal prediction.
[0095] Among them, the physiological signal sensor can be a data acquisition module consisting of multiple electrode units and their driving circuits arranged on the grip area of the directional control device. For example, it can be a multi-channel sensor array that integrates electrocardiogram electrodes and skin conduction electrodes at the same time.
[0096] Physiological signals can be raw electrical quantities or derived quantities that have undergone preliminary processing and can be quantified by physiological signal sensors and collected from the skin surface. For example, the voltage waveform of an electrocardiogram (ECG) signal is used to calculate heart rate and heart rate variability (HRV), and the admittance value of a skin conductance (GSR) signal is used to obtain skin conductance level and skin conductance response peak.
[0097] The physiological signal sequence can be an ordered sequence of physiological signals collected at each discrete time step within a target time window, arranged in the order of time index. For example, the ECG RR interval sequence or skin conductance sampling values of several consecutive time steps (corresponding to 30 seconds to 120 seconds) can be organized into a time series vector, which shares the same time axis scale with the micro-gesture dynamics time series.
[0098] In some implementations, the physiological signal sensor may include an array of electrocardiogram electrodes and skin conductance electrodes distributed at multiple candidate locations within the grip area of the steering control device (e.g., the 3 o'clock and 9 o'clock positions on the steering wheel, and the back grip area). At each time step, one or more physiological signals can be acquired by controlling an analog switch to close the electrode pair corresponding to a dynamically selected physiological signal acquisition area (i.e., the skin contact area with the largest current tactile contour amplitude) via the MCU.
[0099] Specifically, electrocardiogram (ECG) acquisition uses a three-electrode or two-electrode differential method to obtain the operator's ECG waveform (e.g., at a sampling rate of 250 Hz or higher), and calculates heart rate and heart rate variability (HRV) characteristics in real time, such as the root mean square (RMSSD) of the difference between adjacent RR intervals. Skin conductance acquisition uses AC excitation or DC constant voltage method to obtain skin conductance level (SCL) and the amplitude and recovery time of phased skin conductance response (SCR) induced by emotional arousal.
[0100] Taking a congested driving scenario as an example, after the system selects the 9 o'clock area on the steering wheel as the acquisition area, the MCU closes the two ECG electrodes and one GSR electrode pair in that area to continuously acquire ECG voltage signals (typical amplitude 0.5~2mV) and skin conductance admittance values (typically 2~20 microSiemens). During the acquisition process, the system simultaneously performs power frequency notch filtering (50 / 60Hz) and low-pass filtering, and uses sliding median filtering to remove electromyographic noise and motion artifacts, ensuring the quality of the original signal output at each time step.
[0101] Furthermore, to achieve multimodal alignment with the micro-gesture dynamics time series, the physiological signals acquired and preprocessed at each time step can be organized into a physiological signal sequence according to the chronological order of the time steps. This sequence covers a preset target time window (e.g., 30 to 120 seconds), with each time step within the window corresponding to one or more physiological quantitative indicators. For example, for electrocardiogram (ECG) signals, each time step can output a heart rate value (beats / minute) and an HRV indicator (e.g., RMSSD); for skin conductance (SC) signals, each time step can output a skin conductance level value and an indicator of whether a skin conductance response was detected. The system combines these indicators into a low-dimensional vector (e.g., four-dimensional: [heart rate, RMSSD, skin conductance level, skin conductance response event flag]) and stacks them in chronological order.
[0102] Taking a target time window of 30 seconds and an output of one comprehensive physiological vector per second as an example, the physiological signal sequence is presented as a 30-row × 4-column matrix. It should be noted that since ECG R-wave detection and HRV calculation usually require several heartbeat cycles, the system can use a sliding window update: the physiological indicators output at each time step can be calculated in real time based on the original waveforms within the current and previous second-level windows, thereby ensuring the continuity of the sequence and low latency.
[0103] For example, in addition to outputting the aforementioned conventional physiological indicators, respiratory rate can be estimated by synchronously acquiring and utilizing the respiratory sinus arrhythmia (RSA) component of the electrocardiogram at each time step. Specifically, power spectrum analysis or peak detection is performed on the RR interval sequence to extract the power in the high-frequency band (0.15~0.4Hz), which is correlated with respiratory depth and frequency. Adding respiratory rate as an additional feature dimension to the physiological signal sequence can improve the ability to distinguish between the operator's relaxed / tense state, especially when combined with micro-gesture rhythmic indicators, enabling a more accurate differentiation between tension caused by focus and tension caused by anxiety.
[0104] Through the above methods, seamless and continuous collection and temporal alignment of physiological data can be achieved in dynamic control environments such as single-hand switching and changes in grip posture. This provides a high-quality, low-interruption auxiliary data branch with independent correction capabilities for subsequent multimodal fusion prediction, significantly enhancing the robustness and reliability of emotion trend prediction in real-world scenarios.
[0105] Step 105: Acquire the grip force signal collected by the pressure sensor at each time step, and combine the grip force signal corresponding to each time step to construct the grip force sequence corresponding to the target time window.
[0106] In some implementations, in order to distinguish between micro-gestures driven by emotional tension and force changes generated by physical manipulation actions (such as large-angle turns or pushing and pulling joysticks), and to suppress false alarms caused by manipulation load, the grip force values collected by pressure sensors at each time step can be acquired, and the grip force of consecutive time steps can be organized into a grip force sequence aligned with the micro-gesture dynamics sequence according to time windows. This provides auxiliary behavioral evidence for distinguishing manipulation actions from tense force, thereby achieving confidence correction of the tactile dominant mode and suppression of false alarms in interference scenarios in multimodal fusion.
[0107] The pressure sensor can be one or more sensing elements embedded in the grip area of the steering control device that can convert contact pressure into an electrical signal, such as a thin-film pressure sensor, strain gauge, or piezoresistive sensor array. These sensors are distributed at the 3 o'clock and 9 o'clock positions of the steering wheel or the palm pad area of the control stick grip, and can be used to continuously detect the radial grip force, clamping force and their dynamic changes applied by the operator to the grip surface.
[0108] The grip force signal can be the raw force value output by the pressure sensor at each time step or a mechanical quantity after preliminary filtering. For example, it can be the instantaneous grip force value in Newtons obtained by converting the resistance change output by the thin-film pressure sensor.
[0109] The grip force sequence can be an ordered sequence of grip force signals collected at each discrete time step within a target time window, arranged in chronological order. For example, 600 grip force value points are continuously recorded at a sampling rate of 20Hz within a 30-second time window. This sequence shares the same time axis as the micro-gesture dynamics time series and physiological signal sequence.
[0110] In some implementations, the pressure sensor may be a thin-film pressure sensor or a strain gauge sensor, embedded at multiple locations in the grip area of the steering control device (e.g., the 3 o'clock and 9 o'clock positions on the steering wheel, and the palm contact area on the back). At each time step, the pressure sensor may output an electrical signal (e.g., a resistance value or a voltage value) proportional to the grip force, which is then sampled by the analog-to-digital converter (ADC) built into the MCU and converted into a digital force value.
[0111] Taking the steering wheel as an example, a three-layer thin-film pressure sensor is arranged at the 3 o'clock position with a range of 0~50N. When the driver holds the steering wheel normally with both hands, a set of force values can be output at each time step (sampling rate 20Hz), such as 12N for the left hand grip and 15N for the right hand grip.
[0112] Furthermore, to eliminate instantaneous noise caused by fine-tuning of grip posture or road vibration, the system performs sliding median filtering (window length of 5 time steps) on the continuously sampled signal, and the smoothed value is used as the grip force signal for that time step. In some implementations, the system simultaneously collects force values from multiple sensors and takes the maximum value as the representative grip force for the current time step, or records the left-hand force and right-hand force separately to retain more information. When the operator removes one hand from the tray, the force value on the side away from the tray approaches zero, and the system still records this zero value normally to reflect the change in grip state.
[0113] In some implementations, in order to perform temporal alignment and multimodal fusion with other modalities (micro-gesture dynamics time series, physiological signal series), the system organizes the grip force signals acquired at each time step into a grip force sequence in chronological order, which covers a preset target time window (e.g., 30 to 120 seconds).
[0114] In actual testing, for example, during a 30-second period in a congested area, the grip force sequence exhibited a dynamic change, gradually increasing from an initial 15N to 28N and then decreasing. This change showed a time correlation with the increase in micro-gesture frequency and the decrease in HRV. To facilitate subsequent feature extraction, the system can also perform differential calculations on the grip force sequence to obtain a sequence of grip force change rates (reflecting rapid, tense gripping), or calculate statistical features within a window (mean, variance, peak value, etc.), and use these derived features as supplementary representations of the grip force sequence.
[0115] For example, in addition to using a single global grip force value, the grip force can be decomposed into two components: radial grip force and tangential torque. Specifically, by arranging multiple pressure sensing units around the circumference of the steering wheel and combining them with steering angle velocity information, the pushing / pulling force (radial) and the frictional torque in the rotational direction applied by the operator's hands to the steering wheel can be calculated. Radial grip force mainly reflects muscle tension, while tangential torque is more related to steering control actions. In subsequent multimodal fusion, only the radial grip force component is used as an auxiliary feature, while ignoring the tangential torque, which can further suppress tactile misjudgments caused by normal steering. In this way, by employing force decomposition, the false alarm rate can be effectively reduced in large-amplitude steering scenarios.
[0116] The above methods can provide key auxiliary evidence for multimodal fusion models to distinguish between emotion-driven micro-gestures and physical manipulation actions. This can automatically reduce the weight of tactile modalities and suppress false alarms in scenarios with large-scale turning or high-load manipulation, significantly improving the discrimination accuracy and robustness of the emotion trend prediction system in complex dynamic manipulation environments.
[0117] Step 106: Using a large language model, based on the micro-gesture dynamics time series, physiological signal series and grip force series corresponding to the target time window, predict the emotional change trend from the target time window to the next target time window, and obtain the emotional change prediction result.
[0118] In some implementations, in order to provide sufficient advance intervention before the onset of emotional deterioration (such as pre-anger or pre-fatigue) and overcome the limitation of only recognizing the current state, the time series of micro-gesture dynamics, physiological signal sequences, and grip force sequences aligned within the same target time window can be input into a time-series prediction model (such as a sequence-to-sequence architecture constructed by a large language model). The temporal evolution patterns contained therein can be encoded and decoded to predict the trajectory of emotional changes and risk trends within a predetermined future window from the end of the current target time window to the end of the next target time window, thereby outputting an emotional change prediction result with advance warning.
[0119] Among them, the large language model can be a sequence generation model based on massive text data pre-trained, with a multi-head self-attention mechanism and a deep Transformer architecture, such as the GPT series or its lightweight variants.
[0120] In some implementations, large language models can be given the ability to understand the context and extrapolate trends of multimodal time series (micro-gesture dynamics time series, physiological signal sequences, grip force sequences) through fine-tuning or cue learning. They can be used to replace or enhance traditional ConvLSTM / TCN time series prediction networks to achieve generative prediction of the future state of non-verbal behavior sequences.
[0121] The next target time window can be a future time interval that follows the current target time window on the time axis and has the same length, or it can be a configurable time interval. For example, if the current target time window covers the period from 0 to 30 seconds, then the next target time window covers the period from 30 to 60 seconds.
[0122] Among them, the emotion change prediction results can be structured information output by a large language model that describes the trend of the target object's emotional state evolution within a predetermined future time window. For example, it can include a continuous vector sequence containing the three-dimensional trajectory of valence-arousal-dominance, the peak risk level within the future window, the direction and confidence of the upward / downward trend, and semantic risk summaries such as pre-anger and pre-fatigue. This information can be used to drive parameterized soft intervention strategies such as gradual changes in light color, adjustment of rhythm and vibration intensity, and cabin-linked load reduction to achieve proactive adjustment before the emotion deteriorates.
[0123] In some implementations, a finely tuned large language model (such as a lightweight GPT or LLaMA variant) can be used as the core engine for multimodal time series prediction. This model first serializes and encodes the micro-gesture dynamics time series, physiological signal series, and grip force series within a target time window (e.g., 30 seconds). Specifically, the micro-gesture dynamics features (micro-gesture frequency, average rubbing intensity, rhythmicity indicators, cumulative hand displacement, pressure trend slope), physiological signal features (heart rate, HRV, skin conductance level, etc.), and grip force values at each time step can be concatenated into a multimodal feature vector, which is then arranged in the order of the time steps to form an input sequence.
[0124] Furthermore, to adapt to the text input paradigm of large language models, the system can also map the input sequence to the word vector space through linear projection or quantized embedding, or directly generate a natural language description (e.g., "In the past 30 seconds, the frequency of micro-gestures increased from 120 to 180 times / minute, the rubbing intensity increased by 20%, HRV decreased, and grip strength increased from 15N to 28N"). After receiving the above input, the model can use its internal self-attention mechanism to capture long-range dependencies and cross-modal associations between multimodal features, and predict the trend of emotional state changes in the next target time window (e.g., the next 30 seconds) in an autoregressive manner. For example, the model outputs the valence, arousal, and dominance three-dimensional emotional vector trajectory for each time step in the next 30 seconds, as well as the comprehensive risk level (level 1-3). To ensure real-time operation at the edge (e.g., the MCU inside the steering wheel), model quantization (INT8) and operator fusion techniques are used to keep the inference latency within 100 milliseconds.
[0125] For example, in a traffic jam driving simulation test, if data is collected within a 30-second target time window: the micro-gesture dynamics time series shows a continuous increase in micro-gesture frequency (150→220 times / minute), increased rubbing intensity, a rhythmicity index decreasing from 9 to 5 (tending towards disorder), and a positive stress trend slope; the physiological signal sequence shows a heart rate increasing from 78 to 92 beats / minute, an RMSSD (HRV index) decreasing from 35ms to 20ms, and an increase in skin conductance; the grip strength sequence increases from 12N to 25N. Based on these inputs, the large language model outputs a predicted emotional change for the next target time window (30~60 seconds): a predicted increase in arousal of 0.3 (normalized scale), a decrease in valence of 0.2, and a decrease in dominance of 0.4; the overall risk level is level 3 (severe), and a pre-anger state can be described in words, with intervention recommended. This predicted result highly matches the observed frequent tapping and tense facial expressions of the driver in the subsequent 30 seconds, verifying the effectiveness of the prediction.
[0126] In some implementations, once the predicted mood change is obtained, a parameterized soft intervention strategy corresponding to the predicted risk level can be triggered immediately. For example, for Level 2 (moderate) risk, the cabin HMI can be directly driven via the CAN bus or by gradually changing the LED light strip in the steering wheel grip area from green to yellow, while simultaneously initiating gentle rhythmic vibrations (frequency approximately 30Hz, duty cycle 50%), and outputting a guidance prompt via the in-vehicle voice assistant: "We have detected that you are slightly fatigued. We suggest you take deep breaths or activate driver assistance." For Level 3 (severe) risk, the intervention can be further enhanced: the vibration intensity is increased, the light strip turns red and flashes, and it is linked with the vehicle / cabin / engine cabin systems, such as automatically reducing the multimedia volume, lowering the air conditioning temperature, suggesting a rest stop (within the limits allowed by the autonomous driving level), or sending a desensitized risk summary to the fleet management platform. After the intervention is triggered, the system continuously monitors the signal for the next time window. If the operator's condition improves (e.g., micro-gesture frequency decreases, HRV recovers), the intervention intensity is gradually reduced or the system returns to normal; if the condition continues to deteriorate, the intervention is escalated and high-risk events are recorded. All predictive summaries, intervention records, and desensitized statistical features can be stored locally and uploaded to the cloud for long-term trend analysis with user authorization.
[0127] For example, besides directly encoding multimodal sequences into numerical vectors and inputting them into a large language model, a thought chain prompt strategy can be used to guide the model to analyze the reasons for emotional evolution frame by frame. For instance, the input prompt can explicitly require the model to reason step by step: "First, analyze whether changes in the frequency and rhythm of micro-gestures indicate increased stress; second, combine this with whether a decrease in HRV corroborates autonomic nerve activation; then, determine whether the increase in grip strength stems from tension rather than turning movements; finally, synthesize the results to derive future trends." In this way, the model's predictions are not only numerically accurate but also include interpretable reasoning paths, facilitating system debugging and increasing user trust.
[0128] This application embodiment acquires tactile contour subsequences collected by a tactile sensor at each time step, and determines micro-gesture events at each time step based on the tactile contour subsequences; determines the dynamic feature subsequences of the corresponding micro-gesture events based on the tactile contour subsequences collected at each time step, and constructs a micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequences of each time step; for each time step, calculates the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence, and selects the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area; in each time step, physiological signals are acquired from the physiological signal acquisition area by a physiological signal sensor, and constructs a physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step; acquires grip force signals collected by a pressure sensor at each time step, and constructs a grip force sequence corresponding to the target time window by combining the grip force signals corresponding to each time step; and predicts the emotional change trend from the target time window to the next target time window by using a large language model based on the micro-gesture dynamic time sequence, physiological signal sequence, and grip force sequence corresponding to the target time window, thus obtaining the emotional change prediction result. In this way, we can take micro-gestures, a behavioral feature that is more difficult to conceal, in multimodal signals as the main factor, combine dynamic optimization of physiological signal acquisition and grip force to assist in discrimination, and use the model to directly output the emotional trend of future time windows, so as to achieve a more essential and robust characterization of emotional state and forward-looking prediction. Specifically, compared to solutions that rely solely on facial expressions and are reactive in their recognition, this application captures subtle, subconscious hand movements of the operator using tactile sensory micro-gesture dynamic sequences. This fundamentally avoids recognition biases caused by facial occlusion, lighting changes, and intentional facial expression concealment. By dynamically selecting the optimal physiological signal acquisition area in real-time through tactile contour amplitude guidance, it ensures the continuity and signal-to-noise ratio of physiological signals during complex maneuvers such as hand switching and large-angle turns. Furthermore, by combining grip force signals, it effectively distinguishes between normal maneuvers and tense exertion, significantly suppressing false alarms caused by physical manipulation. Finally, a large language model performs cross-time window temporal modeling of multiple biosignal sequences, directly outputting a prediction of emotional changes for the next target time window, thus enabling early intervention before emotional deterioration (such as pre-anger or pre-fatigue) occurs. In summary, this application improves the accuracy of emotion recognition and the timeliness of emotion intervention.
[0129] In some implementations, the predictive framework can be injected into the model in a contextual format that the large language model can understand, thereby guiding the model to make accurate sentiment trend predictions in an individual coordinate system. For example, step 106 may include: (106.1) Obtain the emotional baseline corresponding to the target object being detected; (106.2) Based on the emotional baseline, the micro-gesture dynamics time series, physiological signal series and grip force series corresponding to the target time window, input prompt information is generated; (106.3) Using a large language model, based on input prompts, the trend of emotion change from the target time window to the next target time window is predicted, and the emotion change prediction results are obtained.
[0130] Among them, the emotional baseline can be the individualized reference range or distribution characteristics of micro-gesture dynamic parameters, physiological signal parameters and grip force parameters obtained by statistical calculation of the initial data of 1 to 3 minutes collected from the target object in the quasi-resting state at the beginning of power-on (such as constant speed driving or non-control load stage). For example, the mean and standard deviation of micro-gesture frequency, the median RMSSD of heart rate variability, and the resting baseline value of skin conductance level can be used to convert the currently detected feature values to the deviation score or Z score in the individual coordinate system, thereby compensating for physiological fatigue drift caused by individual differences and long-term operation.
[0131] The input prompt information can be structured text or serialized tensors constructed according to the input format that the large language model is accustomed to during pre-training. This information semantically organizes the emotion baseline parameters (as the individual's context), the micro-gesture dynamics time series and its feature description within the current target time window, the physiological signal sequence and its feature description, and the grip force sequence and its feature description. For example, a natural language template such as "the frequency range of micro-gestures in the individual's normal state is between X and Y, the current window frequency is Z, and it shows an upward trend" can be used to guide the large language model to align to the individual's specific coordinate scale when performing emotion trend prediction.
[0132] In some implementations, the emotional baseline can be established upon system initial power-on or the first use by the operator. Specifically, multiple consecutive reference time windows (e.g., three 30-second windows) can be selected, during which the target subject is in a resting state—i.e., the vehicle is traveling at a constant speed in a straight line, without significant turns or obvious emotional stimuli. Within each reference window, the system collects the initial micro-gesture dynamics time series, initial physiological signal series, and initial grip force series in the exact same manner as described above. Subsequently, statistical analysis is performed on each series within these windows: the mean and standard deviation of micro-gesture frequency, the typical range of average rubbing intensity, the baseline value of rhythmicity indicators, the resting level of cumulative hand displacement, and the zero-value neighborhood of the pressure trend slope are calculated; simultaneously, the median RMSSD of heart rate variability, the resting mean of skin conductance, and the long-term average of grip force are calculated. These statistical parameters collectively constitute the operator's emotional baseline.
[0133] Furthermore, to compensate for physiological fatigue drift caused by long-term driving, the system can automatically collect new short-term resting segments every 30 minutes of driving (e.g., during high-speed cruising) and slowly update the baseline parameters using an exponentially weighted moving average, so that the baseline can adaptively adjust to the operator's state changes.
[0134] Furthermore, the system can fuse emotional baseline parameters with multimodal sequences within the current target time window to generate structured input prompts. These prompts can be in a natural language format that large language models can understand, and include three parts: an individual background description (based on the baseline's normal range), a summary of the current observation data (key statistical features extracted from micro-gesture dynamics time series, physiological signal sequences, and grip force sequences), and task instructions.
[0135] For example, the statistical parameters of the baseline can be converted into descriptive statements, such as "The driver's resting micro-gesture frequency is 30±8 times / minute, resting heart rate is 72±5 beats / minute, and resting grip strength is 12±3N"; then the real-time features in the current window are compared with the baseline to generate a deviation description, such as "The current micro-gesture frequency is 180 times / minute, exceeding the baseline by 5 standard deviations; the RMSSD of HRV is 40% lower than the baseline"; finally, a prediction instruction is added: "Based on the above information, please predict the changing trends of the driver's emotional valence, arousal, and dominance in the current 30 seconds and the next 30 seconds, and output the risk level (1 mild / 2 moderate / 3 severe)".
[0136] In some implementations, descriptive information marked with trust parameters (such as the current tactile signal being affected by steering interference and having low confidence) can also be embedded as a prompt to guide the model to allocate modal weights appropriately.
[0137] For example, in a test on a congested road section, the system generated the following input prompt information: "
Individual Baseline
[0138] [Current 30-second window data] Micro-gesture frequency 180 times / minute (5 times higher than baseline), kneading intensity 0.72, rhythmicity index 3.1 (significantly reduced), pressure trend slope +0.02; heart rate 92 beats / minute, RMSSD 20ms, skin conductance level 9μS; grip strength 25N (left) and 20N (right), grip strength asymmetry coefficient 0.2.
[0139] [Context] When the steering angular velocity is below the threshold, the tactile trust level is high, and the physiological trust level is high.
[0140] [Instruction] Please predict the continuous trend of the driver's emotional arousal (0~1, 0 represents relaxation), valence (-1~1, negative values represent negativity), and sense of dominance (0~1) over the next 30 seconds, as well as the overall risk level (1 / 2 / 3). Output format: "Predicted trend: Arousal increases from 0.7 to 0.9, valence decreases from -0.2 to -0.6, sense of dominance decreases from 0.5 to 0.2; Risk level: 3; Confidence level: 0.92".
[0141] Furthermore, the aforementioned input prompts can be fed into a finely tuned, lightweight large language model (such as a Phi-3 Mini deployed at the edge or a quantized LLaMA-3B). The large language model can then utilize its self-attention mechanism to parse the numerical contrast relationships and temporal semantics within the prompts and output prediction results.
[0142] Specifically, this process can employ an autoregressive generation method. The model first outputs the three-dimensional sentiment vector trajectory for each time step within the next target time window (e.g., one sampling point every 2 seconds), and then summarizes the data to obtain the risk level and confidence score. To ensure real-time performance, model inference is performed on the neural network acceleration unit built into the steering wheel, with a typical latency of 80-150 milliseconds.
[0143] Furthermore, the prediction results can then be routed to the soft intervention module to drive feedback such as lighting, vibration, and voice, and can also be sent to the cockpit domain controller via the CAN bus. In some implementations, the model also generates brief decision-making text, such as frequent micro-gestures combined with a sharp decrease in HRV and an increase in grip strength, indicating high tension and a high probability of emotional deterioration within the next 30 seconds, to improve the interpretability of the system.
[0144] In some implementations, the sentiment baseline and multimodal sequences can be directly embedded into the input layer of a large language model via learnable continuous vectors, without explicitly generating text.
[0145] By using the above methods, the model can perform normalized evaluation and trend extrapolation of the current signal in the individual coordinate system, thereby providing individualized and drift-compensated contextual guidance for predicting the emotional changes from the current time window to the next time window, significantly improving the generalization ability of cross-individual emotional prediction and the prediction accuracy in long-term manipulation scenarios.
[0146] In some implementations, to eliminate the impact of inherent physiological differences between different operators and signal drift caused by prolonged operation (such as baseline shift due to fatigue) on the accuracy of emotion prediction, an emotion baseline representing the operator's normal state can be constructed under no-operational-load or low-load resting conditions. This provides a quantitative reference benchmark for subsequent individualized normalization, drift compensation, and anomaly detection of real-time signals. For example, (106.1) may include: (106.1.1) For each reference time window, collect the initial micro-gesture dynamics time series, initial physiological signal series and initial grip force series of the target object in a resting state; (106.1.2) Based on multiple initial micro-gesture dynamics time series, multiple initial physiological signal sequences and multiple initial grip force sequences corresponding to multiple reference time windows, calculate the statistical parameters of micro-gesture dynamics, physiological signal sequences and grip force sequences; (106.1.3) Based on the statistical parameters of micro-gesture dynamics, the statistical parameters of physiological signals and the statistical parameters of grip force sequence, the emotional baseline corresponding to the target object is obtained.
[0147] The initial micro-gesture dynamics time series can be an ordered sequence of dynamic characteristic parameters such as micro-gesture frequency, rubbing intensity, rhythmicity index, cumulative hand displacement, and pressure trend slope continuously collected by the system according to time steps within a reference time window when the target object is in a resting state (e.g., the vehicle is traveling in a straight line at a constant speed, without significant turning, and the operator is in a stable mood).
[0148] The initial physiological signal sequence can be the original or pre-processed physiological data sequence obtained from the selected acquisition area by the physiological signal sensor within the same resting state reference time window and synchronized with time.
[0149] The initial grip force sequence can be a sequence of grip force values continuously collected and organized by pressure sensors at time steps within a reference time window under the same resting state.
[0150] Among them, the statistical parameters of micro-gesture dynamics can be a summary quantitative index obtained by statistical analysis of multiple initial micro-gesture dynamic time series obtained from multiple reference time windows. For example, the overall mean and standard deviation of micro-gesture frequency, the peak distribution range of rubbing intensity, the typical range of rhythmic index, and the zero value neighborhood of pressure trend slope can be used to define the statistical boundary of the normal tactile behavior pattern of the object, and serve as the basis for subsequent real-time data Z-score normalization and anomaly judgment.
[0151] Among them, the physiological signal statistical parameters can be a summary of individualized physiological characteristics obtained by statistically calculating multiple initial physiological signal sequences obtained from multiple reference time windows. For example, they can be the mean and coefficient of variation of heart rate variability indicators (such as RMSSD), the typical range of low-frequency to high-frequency power ratio (LF / HF), the median of skin conductance level, and the conventional amplitude of response peak. These parameters can be used to construct an individualized baseline of resting tension and emotional arousal of the subject's autonomic nervous system.
[0152] Among them, the grip force sequence statistical parameters can be a summary of mechanical characteristics obtained by statistically analyzing multiple initial grip force sequences obtained from multiple reference time windows, such as the long-term mean, standard deviation, maximum and minimum range, and typical distribution of the rate of change of grip force. These parameters can be used to distinguish between natural grip force fluctuations under normal operation of the object and abnormal force driven by emotional tension.
[0153] In some implementations, resting state data for multiple reference time windows can be collected when the target object uses the system for the first time or when the user actively triggers the calibration process. Taking a vehicle steering wheel application as an example, the system's criteria for determining the resting state include: vehicle speed stable at 30-60 km / h with no significant steering (steering angular velocity below a preset threshold), no sudden acceleration / braking, and no high-intensity voice interaction or media playback in the cabin. Once these conditions are met, the system automatically opens three consecutive reference time windows (each window lasts for example, 30 seconds, and adjacent windows may or may not overlap). Within each reference window, the system collects an initial micro-gesture dynamics time series (including micro-gesture frequency, rubbing intensity, rhythmicity indicators, cumulative hand displacement, and pressure trend slope at each time step), an initial physiological signal series (heart rate, HRV indicators, skin conductance level, and skin conductance response events at each time step), and an initial grip force series (grip force value at each time step). These three initial sequences together constitute the resting state baseline data for that reference window. To ensure data reliability, the system can check for abnormal events within the window (such as an operator sneezing or answering a phone call). If any such events occur, the window is discarded and the data collection time is extended until three valid reference windows are obtained.
[0154] Furthermore, after obtaining the initial sequence of multiple reference windows, a set of statistical parameters can be calculated for each mode. For micro-gesture dynamics, the micro-gesture frequency values of all time steps within all windows can be aggregated, and their mean and standard deviation can be calculated; similarly, the mean and standard deviation of kneading intensity, the mean and standard deviation of rhythmicity index, the mean and standard deviation of cumulative hand displacement, and the mean and standard deviation of pressure trend slope (the pressure trend slope should be close to zero in the resting state) can be calculated.
[0155] Specifically, for physiological signals, the mean and standard deviation of heart rate, the mean and standard deviation of heart rate variability (RMSSD), the mean and standard deviation of skin conductance level, and the average frequency of skin conductance response events can be calculated.
[0156] Specifically, for grip strength, the mean and standard deviation of grip strength can be calculated, as well as the mean and standard deviation of the difference between left and right grip strength.
[0157] These statistical parameters are referred to as micro-gesture dynamics statistical parameters, physiological signal statistical parameters, and grip force sequence statistical parameters, respectively.
[0158] For example, in a real-world calibration, the statistical results of a driver's three reference window data can be as follows: The mean micro-gesture frequency was 25 times / minute, with a standard deviation of 8 times / minute; the mean rubbing intensity was 0.30 (normalized), with a standard deviation of 0.05; the mean rhythmicity index was 10.2, with a standard deviation of 1.5; the mean cumulative hand displacement was 35 (relative units), with a standard deviation of 10; and the mean pressure trend slope was 0.001, with a standard deviation of 0.008. Physiological signals: the mean heart rate was 72 beats / minute, with a standard deviation of 5 beats / minute; the mean RMSSD was 40 ms, with a standard deviation of 6 ms; the mean skin conductance was 3.2 μS, with a standard deviation of 0.8 μS; and the mean skin conductance response events were 0.2 times / minute. The mean grip strength was 13 N, with a standard deviation of 3 N. These values constitute the driver's specific statistical baseline.
[0159] Furthermore, all the aforementioned statistical parameters can be packaged into a sentiment baseline for the target object. The sentiment baseline can be represented as a set of numerical vectors (e.g., arranging the means and standard deviations of each modality into a long vector in a fixed order), or as a probability distribution model (e.g., assuming each feature follows a Gaussian distribution, described by mean and variance). In subsequent real-time predictions, the system can compare the real-time feature values of the current target time window with the sentiment baseline to calculate the degree of deviation (e.g., Z-score) or Mahalanobis distance. The formula for calculating Mahalanobis distance is: ; in, It can be a real-time multimodal feature vector. It can be the baseline mean vector. This can be a covariance matrix (composed of the standard deviations of each statistical parameter and the correlation between features). Mahalanobis distance can eliminate the influence of dimensions and take into account the correlation between features, more accurately determining whether the current state significantly deviates from the individual's normal range. When the Mahalanobis distance exceeds a preset threshold (e.g., 2.5), the system determines that the operator is in an abnormal emotional state. The emotional baseline is also slowly updated over long driving periods (e.g., every 30 minutes using a new resting segment with an exponentially weighted moving average to update the mean and variance) to compensate for fatigue or physiological adaptation changes.
[0160] By using the above methods, a unique emotional baseline that reflects the normal physiological and behavioral characteristics can be constructed, thereby providing a unified individual coordinate system and drift compensation reference for subsequent real-time signals, significantly improving the generalization ability of cross-individual emotion prediction and the discrimination stability under long-term manipulation scenarios.
[0161] In some implementations, to suppress misjudgments of tactile modalities caused by physical motion interference in complex control scenarios (such as large-angle steering, continuous curves, or bumpy road sections), a first trust parameter and a second trust parameter dynamically calculated based on the device context (e.g., steering angular velocity and vehicle speed) can be obtained. Semantic descriptions describing the credibility of the micro-gesture dynamics modality and the physiological / grip force modality can be generated accordingly. These descriptions, along with the multimodal sequence and the emotion baseline, are injected into the input prompt to guide the large language model to adaptively adjust the dependence weights on each modality's evidence during fusion prediction, thereby reducing the false alarm rate caused by environmental disturbances. For example, before (106.2), that is, before "generating input prompt information based on the emotion baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence," the following can be included: (A.1) Obtain the first trust parameter corresponding to the micro-gesture dynamics time series, and obtain the second trust parameter corresponding to the physiological signal sequence and grip force sequence; (A.2) Generate first descriptive information for the micro-gesture dynamics time series based on the first trust parameter, and generate second descriptive information for the physiological signal sequence and grip force sequence based on the second trust parameter; Based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, input prompt information is generated, including: Based on the emotional baseline, first descriptive information, second descriptive information, micro-gesture dynamics time series corresponding to the target time window, physiological signal sequence, and grip force sequence, input prompt information is generated.
[0162] The first trust parameter can be a numerical indicator used to quantify the reliability of the micro-gesture dynamics time series within the current target time window. For example, when the steering angular velocity read through the CAN bus exceeds the first threshold positively correlated with the vehicle speed, the system determines that the physical control action (such as rapid steering) produces a large number of tactile contour changes, and thus automatically sets the parameter to a low value (such as 0.3). This can be used to reduce the confidence of the tactile dominant mode in subsequent multimodal fusion and suppress the risk of being misjudged as anxiety-related micro-gestures due to physical friction.
[0163] The second trust parameter can be a numerical indicator used to quantify the reliability of the physiological signal sequence and grip force sequence within the current target time window. For example, in a scenario with a large turning radius, due to severe interference with tactile sensation, the system automatically sets the second trust parameter to a higher value (such as 0.8). This can be used to increase the priority of ECG, skin conductance and grip force data in this scenario, so that they can play a greater role in emotion prediction or even take over the dominant judgment.
[0164] The first descriptive information can be a text fragment that converts the specific value, level, or state of the first trust parameter into a semantic phrase that the large language model can directly understand. For example, if the current tactile signal is interfered with by a large turning action, the confidence level is low, or the confidence level of the micro-gesture dynamics time series is 30%, it can be used to clearly inform the model of the degree of trust in the tactile modality evidence in the input prompt, and guide the model to reduce its dependence on micro-gesture dynamics features during this time period.
[0165] The second descriptive information can be the conversion of the value or level of the second trust parameter into semantic descriptive text that is instructive for the large language model. For example, the current physiological signals (ECG, skin conductance) and grip strength signals are less affected by physical interference, have a high confidence level, or have a physiological and mechanical modality confidence level of 80%. This information can be used to guide the model in the input prompt to prioritize the physiological and grip strength order in this scenario.
[0166] In some implementations, the system can acquire the steering angular velocity and speed of a target device (e.g., a car) in real time. The steering angular velocity (unit: degrees / second) is obtained by reading the steering wheel angle sensor signal via the CAN bus and calculating the rate of change of angle per unit time; simultaneously, the speed (km / h) is obtained by reading the vehicle speed sensor signal. The system dynamically calculates a first threshold based on the speed, which is positively correlated with the speed, for example, using a piecewise linear function: the threshold is 180° / s when the vehicle speed is 10 km / h, and 60° / s when the vehicle speed is 60 km / h.
[0167] Furthermore, the current steering angular velocity can be compared with this threshold: if the steering angular velocity exceeds the first threshold, it indicates that the operator is making a sharp turn (such as a U-turn or a sharp bend). In this case, the physical friction and sliding between the palm and the steering wheel will severely contaminate the tactile signal, making micro-gesture detection unreliable. Therefore, the system can set the first trust parameter corresponding to the micro-gesture dynamics time series to a lower value (e.g., 0.3), while setting the second trust parameter corresponding to the physiological signal sequence and grip force sequence to a higher value (e.g., 0.8). Conversely, if the steering angular velocity does not exceed the threshold, both can be set to a moderately high value by default (e.g., first trust parameter 0.9, second trust parameter 0.7, micro-gesture as the dominant modality). The value range of the trust parameter is usually from 0 to 1, with higher values indicating higher reliability of the evidence for that modality in the current scenario.
[0168] In some implementations, the trust parameter can also be fine-tuned based on other contextual factors, such as the degree of road bumps (detected by an accelerometer) and cabin humidity (which affects the quality of the skin conductance signal).
[0169] Furthermore, after obtaining the trust parameters, the system can convert them into semantic descriptive text that a large language model can directly understand. The first descriptive information targets the micro-gesture dynamics time series, and the second descriptive information targets the physiological signal sequence and grip force sequence. The conversion rules can be based on preset templates and threshold ranges: for example, when the first trust parameter is greater than 0.8, it generates "The current tactile signal is clear, and the micro-gesture dynamics time series has high confidence"; when the trust parameter is between 0.5 and 0.8, it generates "The tactile signal has slight interference, and the micro-gesture dynamics time series has moderate confidence"; when the trust parameter is less than 0.5, it generates "The current turning action is violent, the tactile signal is severely interfered with by physical friction, the micro-gesture dynamics time series has low confidence, and it is recommended to reduce its weight."
[0170] Similarly, for the second trust parameter, descriptions such as "physiological signals and grip force signals are minimally affected, and data quality is good" or "the skin conductance signal drifts slightly due to sweating, but remains generally reliable" can be generated. These descriptions can be concatenated into a concise contextual hint paragraph. In some implementations, the specific value of the trust parameter (e.g., "0.3") can also be included in the description to provide a quantitative reference.
[0171] Furthermore, the system can organize the emotional baseline (which already includes individual statistical parameters), the aforementioned first and second descriptive information, and the micro-gesture dynamics time series, physiological signal series, and grip force series within the current target time window into a complete input prompt message according to a predetermined format. This prompt message can be in a hybrid form of natural language and structured data: first, the individual baseline background is stated, then a modal confidence description is added, followed by a list of key statistical characteristics of the current time window (e.g., "micro-gesture frequency 180 times / minute, exceeding the baseline by 5 times; heart rate 92 beats / minute, 20% higher than the baseline"), and finally, a clear prediction instruction is given. For example: "[Individual Baseline]...[Modal Confidence] Current turning angular velocity exceeds the threshold, tactile signal confidence is low (0.3), physiological and grip signal confidence is high (0.8). [Real-time Data]...[Instruction] Based on the above information, please predict the emotional change trend within the next 30 seconds." This prompt message is fed into the large language model as input for generating the emotion prediction result.
[0172] By using the above methods, the reliability of each evidence source in the current control scenario can be explicitly marked in the multimodal input prompts. This guides the large language model to adaptively adjust the dependence weights on different modalities during fusion prediction, significantly reducing emotional misjudgments caused by physical disturbances such as turning and bumps, and improving the robustness and prediction accuracy of the system in real dynamic control environments.
[0173] In some implementations, to effectively suppress the misinterpretation of tactile signals such as palm friction and sliding caused by rapid steering wheel rotation as anxiety-related microgestures in scenarios with strong physical control, such as sharp turns and continuous curves, the steering angular velocity and driving speed of the target device can be acquired in real time. A positively correlated first threshold can be dynamically calculated based on the speed. When the steering angular velocity exceeds this threshold, it is determined to be a physically controlled scenario. In this scenario, the first trust parameter of the microgesture dynamics time series is actively set to a second trust parameter lower than that of the physiological signal series and grip force signal series. This reduces the acceptance weight of tactile evidence in multimodal fusion, thereby avoiding the misreporting of normal steering actions as emotional distress events. For example, (A.1) may include: (A.1.1) Obtain the device context parameters for the target object to manipulate the target device, wherein the device context parameters include at least the steering angular velocity and the velocity; (A.1.2) Calculate the first threshold based on speed, where the first threshold is positively correlated with speed; (A.1.3) When the steering angular velocity exceeds the first threshold, calculate the first trust parameter corresponding to the micro-gesture dynamics time series, and calculate the second trust parameter corresponding to the physiological signal sequence and grip force sequence, wherein the first trust parameter is less than the second trust parameter.
[0174] The target device can be a directional control device or its vehicle that has a multimodal sensing and emotion prediction system installed or integrated, such as a car steering wheel and CAN bus network on the vehicle, a flight simulator joystick, a ship's rudder, a construction machinery's control handle, or a game / VR controller. The target device can provide the system with real-time control data (such as turning angle and speed) and provide power and interactive interfaces to the system.
[0175] Among them, the device context parameters can be a set of dynamic variables that are output in real time by the target device’s Controller Area Network (CAN) bus, Ethernet or dedicated interface, describing the current control state and environmental conditions. Examples include the vehicle’s instantaneous speed, steering wheel angular velocity, steering torque, auxiliary system activation status, and task phase indicators (such as takeoff / landing, competitive game intensity). These parameters can be used to quantify the degree of interference of physical control actions on tactile signals in the current scenario.
[0176] Among them, the steering angular velocity can be the rate of change of the angle of rotation of the target device's directional control interface (such as steering wheel, joystick, steering wheel) around its steering axis per unit time. It can be obtained by reading the steering wheel angle sensor signal through the CAN bus and calculating the time derivative. The magnitude of this parameter directly reflects whether the operator is performing violent operations such as rapid U-turns, emergency avoidance, or continuous curves. It can be used as a key feature to determine whether the tactile channel is dominated by physical friction.
[0177] Speed can be the rate of movement of the target equipment (vehicle) relative to the ground or a reference frame, such as the speed of a vehicle (kilometers per hour) or the airspeed of an aircraft (knots).
[0178] The first threshold can be an angular velocity threshold used to determine whether the current steering operation belongs to a strong physical interference scenario.
[0179] In some implementations, control-related device context parameters can be obtained in real time via the target device's communication bus (e.g., vehicle control area network CAN bus, Ethernet, or dedicated interface). Taking a car steering wheel as an example, the steering wheel angle sensor signal and vehicle speed sensor signal can be subscribed to on the bus. At each time step (e.g., a sampling period of 50ms), the current steering angle value (in degrees) and speed value (in kilometers per hour) can be read. By performing time difference on the steering angle value (the current steering angle minus the steering angle of the previous time step, divided by the sampling interval), the steering angular velocity (degrees per second) can be calculated.
[0180] Optionally, other context parameters, such as steering torque, auxiliary system activation status, and road surface roughness (via accelerometer), can be read for fine-tuning of subsequent trust parameters.
[0181] Furthermore, an angular velocity threshold can be dynamically calculated based on the current speed to obtain a first threshold. The first threshold can be positively correlated with speed: at low speeds, a larger turning angular velocity is permissible (e.g., making a U-turn in a parking lot), while at high speeds, even a small angular velocity may pose a danger or cause strong interference. In one specific implementation, the first threshold... It can be calculated using the following piecewise linear function: ; in The threshold is set to speed (km / h) and degrees per second. For example, at a speed of 10km / h, the threshold is 180° / s; at 30km / h, the threshold is 140° / s; and at speeds of 60km / h and above, the threshold stabilizes at 60° / s. By setting a first threshold, it is ensured that trust parameter adjustments are not mistakenly triggered during sharp U-turns at low speeds, while even slight steering in high-speed scenarios is considered strong interference, thus protecting the reliability of sentiment prediction.
[0182] Specifically, after calculating the current steering angular velocity at each time step... and the first threshold Then, the system compares: if > If the current situation involves strong physical manipulation interference (such as rapid steering or sharp turns), the tactile signal is dominated by friction and sliding between the palm and the steering wheel, resulting in low confidence in micro-gesture detection. Conversely, physiological signals (ECG, skin conductance) and grip force signals are less affected by such interference. Therefore, the first confidence parameter corresponding to the micro-gesture dynamics time series can be used... Set to a low value (e.g., 0.3) to assign the second trust parameter corresponding to the physiological signal sequence and grip strength sequence. Set it to a high value (e.g., 0.8). If Then the default value (e.g.) will be used. =0.9, =0.7), meaning the default haptic mode is dominant. To smooth transient fluctuations, the system can also perform an exponentially weighted moving average of the trust parameter and its value from the previous moment. For example, during a series of curves at a speed of 50 km / h, the calculated first threshold is 180. 2 (50 10) = 100° / s, the actual peak steering angular velocity is 120° / s, exceeding the threshold, the system immediately outputs... =0.3, =0.8. This trust parameter is then fed into the descriptive information generation module, which causes the large language model to reduce the weight of tactile evidence during prediction, thus avoiding misjudging normal turning as anxious microgestures.
[0183] By employing the above methods, it is possible to automatically reduce the weight of tactile modal evidence and increase the weight of physiological / grip force modal evidence in scenarios dominated by physical control. This significantly suppresses false alarms caused by misjudging normal physical friction as emotionally driven micro-gestures during conditions such as large-angle steering, U-turns, or continuous curves, thereby improving the system's robustness in real dynamic driving environments and enhancing the user experience.
[0184] In some implementations, to transform the encoding results of multimodal time series by a large language model into interpretable and interventionizable emotional state representations, and to provide early warning of emotional deterioration within a short future time window, the large language model can generate a continuous emotional vector trajectory for the current target time window based on input prompts. Then, using the current trajectory and prompts as conditions, an autoregressive prediction of the emotional vector trajectory for the next target time window is made. Finally, peak values, slopes, and confidence levels are extracted from the temporal evolution of both and quantified into a risk level summary. The current trajectory, the predicted trajectory, and the risk summary are then used as the output. For example, (106.3) may include: (106.3.1) Using a large language model, based on input prompt information, the emotion prediction of the target object based on multiple biological signals is performed to obtain the emotion vector trajectory corresponding to the target time window; (106.3.2) Based on the input prompt information and the emotion vector trajectory, the emotion of the target object in the next target time window is predicted to obtain the predicted emotion vector trajectory corresponding to the next target time window; (106.3.3) Based on the emotion vector trajectory and the predicted emotion vector trajectory, a risk level summary for the target object is obtained; (106.3.4) Based on the emotion vector trajectory, the predicted emotion vector trajectory and the risk level summary, the emotion change prediction results are obtained.
[0185] The emotion vector trajectory can be a continuous path formed by connecting the multidimensional emotion representation vectors output by the large language model at each time step or each inference frame in chronological order within the current target time window. For example, it can be a set of time points composed of the three dimensions of valence, arousal and dominance output at the second-level granularity within a 30-second window. This trajectory reflects the real-time state of emotion and also contains short-term evolution patterns such as rise, fall and fluctuation. It can be used to determine whether the current emotion is in the early signs of deterioration or in an abnormal range.
[0186] Among them, the predicted emotion vector trajectory can be the emotion vector trajectory of the large language model based on the current target time window and input prompt information (including emotion baseline, trust description, etc.), and output through sequence generation or extrapolation. It is a continuous path of expected emotional state covering the next target time window (e.g., the next 30 to 120 seconds). For example, it can predict the valence-arousal-dominance three-dimensional coordinates of each time step in the future window. This trajectory can reveal trends such as pre-anger, pre-fatigue, or a continuous increase in arousal in advance, and can be used to trigger soft intervention strategies before the emotion actually deteriorates.
[0187] The risk level summary can be a hierarchical semantic label and brief description output after comprehensively evaluating the key features (such as trajectory slope, peak exceedance, rise duration, and confidence interval) in the current emotion vector trajectory and the predicted emotion vector trajectory. For example, "Level 1 (mild): Arousal is on the rise but still within the baseline range" or "Level 3 (severe): Arousal is predicted to exceed the individual threshold by 2 standard deviations within the next 30 seconds, accompanied by a sharp drop in dominance." It can be used to drive the cockpit / cabin / game engine to perform parameterized soft interventions (such as color change, vibration, and difficulty reduction) while meeting privacy compliance requirements.
[0188] In some implementations, after receiving input prompts (including sentiment baseline, modal trust description, and multimodal sequence statistical features of the current window), the large language model can use its internal Transformer self-attention mechanism to encode the numerical and semantic associations in the input prompts and generate the sentiment vector trajectory within the current target time window using an autoregressive approach.
[0189] For example, an emotion vector trajectory can consist of a series of three-dimensional emotion vectors at discrete time points (e.g., one point every 2 seconds). Each vector's three dimensions correspond to valence (ranging from -1 to 1, with negative values representing negativity), arousal (ranging from 0 to 1, with higher values indicating excitement / tension), and dominance (ranging from 0 to 1, with higher values indicating a stronger sense of control). Using a 30-second window as an example, the large language model can output 15 emotion vectors (one every 2 seconds), forming a continuous path from the beginning to the end of the window. The model also outputs the confidence score for each vector during generation. To ensure real-time performance, this generation can employ greedy decoding or constrained bundle search.
[0190] Furthermore, the emotion vector trajectory within the current target time window can be concatenated with the original input prompt to create an expanded prompt (e.g., "Current trajectory: ..., please predict the emotion trajectory for the next 30 seconds based on this"). This is then fed back into the large language model, which uses the state of the last few time points of the current trajectory as conditions to predict the emotion vector sequence with the same temporal resolution within the next target time window (e.g., seconds 30-60). This process is similar to temporal extrapolation: the model learns that under patterns of increasing micro-gesture frequency, decreasing HRV, and increasing grip strength, future arousal usually continues to rise, while valence and dominance may decrease. For example, if the current trajectory shows arousal rising from 0.5 to 0.7, the model predicts that in the future window, arousal will further rise to 0.9, valence will decrease from -0.2 to -0.6, and dominance will decrease from 0.6 to 0.3. The prediction results are also output in the form of vector trajectories, accompanied by point-by-point confidence intervals.
[0191] In some implementations, the system can calculate multiple risk indicators based on the current and predicted emotion vector trajectories: the maximum value and upward slope of arousal within the prediction window; the minimum value and rate of decline of valence; the minimum value and rate of decline of dominance; and whether these indicators exceed preset absolute thresholds or deviate from the individual baseline (e.g., Mahalanobis distance). Combining these indicators, the risk can be classified into 1-3 levels: Level 1 (mild, arousal < 0.6 and valence > -0.3, deviation from baseline < 1 standard deviation); Level 2 (moderate, arousal 0.6-0.8 or valence -0.6 to -0.3, deviation 1-2 standard deviations); Level 3 (severe, arousal > 0.8 or valence < -0.6, deviation > 2 standard deviations, or a consistently positive prediction slope). In some implementations, the system also outputs key factors corresponding to the risk level, such as "Risk Level 3, Main Driving Factors: Rapid Increase in Predicted Arousal + Sharp Decrease in Heart Rate Variability." This information is presented in a structured summary format.
[0192] Furthermore, the three pieces of information mentioned above—the current emotion vector trajectory, the predicted emotion vector trajectory, and the risk level summary—can be encapsulated to form the final emotion change prediction result. This result can be either a comprehensive text description (e.g., "Current valence -0.2, arousal 0.7, dominance 0.5; predicted valence for the next 30 seconds -0.6, arousal 0.9, dominance 0.2; overall risk level 3 (severe). Soft intervention is recommended immediately.") or a structured data object (containing fields such as trajectory array, risk level, confidence level, and timestamp). This result is then sent to the soft intervention module and the HMI interface.
[0193] For example, in a car scenario, when the system predicts a risk level of 3, it will immediately turn the LED light strip to flash red, increase the vibration intensity, and prompt the driver via voice to rest or activate driver assistance. Simultaneously, this prediction result can also be uploaded to the fleet management platform (after anonymization) for driver fatigue management.
[0194] In some implementations, to improve the reliability of the prediction results, the system can require the large language model to simultaneously output explanatory text, such as "Model prediction basis: The frequency of micro-gestures has been continuously increasing and the rhythm is disordered over the past 30 seconds. Combined with the decrease in HRV and the increase in grip strength, this all points to a state of high stress; based on the pattern in the training data, the arousal level is highly likely to exceed 0.9 within the next 30 seconds." This explanatory text not only facilitates debugging for developers but can also be displayed in the user interface (e.g., "The system determines that you are currently under a lot of stress. Please take a break"), increasing user acceptance of automatic intervention.
[0195] Through the above methods, abstract temporal physiological-behavioral signals can be transformed into three-dimensional emotional pathways and hierarchical early warning information with physical semantics and forward-looking capabilities. This provides a quantitative basis for the perception-prediction-soft intervention closed loop, which can directly drive the linkage of lights, vibration, voice and HMI, and realizes a key leap from state recognition to trend prediction, and from post-event alarm to early intervention.
[0196] Please see Figure 3 This application also provides an emotion prediction device based on multiple biosignals, applied to a target device. The target device includes at least a tactile sensor, a physiological signal sensor, and a pressure sensor. The emotion prediction device based on multiple biosignals can implement the above-mentioned emotion prediction method based on multiple biosignals. The emotion prediction device based on multiple biosignals may include: The acquisition module 31 is used to acquire the tactile contour subsequence collected by the tactile sensor at each time step, and to determine the micro-gesture event at each time step based on the tactile contour subsequence; The determination module 32 is used to determine the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence collected at each time step, and to construct the micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequence of each time step. The calculation module 33 is used to calculate the tactile contour amplitude of the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence for each time step, and select the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area. The acquisition module 34 is used to acquire physiological signals from the physiological signal acquisition area through the physiological signal sensor at each time step, and to construct the physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step. Module 35 is used to acquire the grip force signal collected by the pressure sensor at each time step, and combine the grip force signal corresponding to each time step to construct the grip force sequence corresponding to the target time window; The prediction module 36 is used to predict the emotional change trend from the target time window to the next target time window by using a large language model based on the micro-gesture dynamics time series, physiological signal series and grip force series corresponding to the target time window, and obtain the emotional change prediction result.
[0197] The specific implementation of this emotion prediction device based on multiple biosignals is basically the same as the specific embodiment of the emotion prediction method based on multiple biosignals described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the emotion prediction device based on multiple biosignals may also be equipped with other functional modules to implement the emotion prediction method based on multiple biosignals in the above embodiments.
[0198] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned emotion prediction method based on multiple biosignals. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0199] Please see Figure 4 , Figure 4 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 41 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 42 can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 42 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 42 and called and executed by the processor 41 to execute the emotion prediction method based on multiple biosignals of the embodiments of this application. Input / output interface 43 is used to implement information input and output; The communication interface 44 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 45 transmits information between various components of the device (e.g., processor 41, memory 42, input / output interface 43, and communication interface 44); The processor 41, memory 42, input / output interface 43 and communication interface 44 are connected to each other within the device via bus 45.
[0200] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described emotion prediction method based on multiple biosignals.
[0201] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0202] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0203] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0205] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0206] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0207] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0208] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0209] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0210] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0211] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0212] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A sentiment prediction method based on multiple biological signals, characterized in that, Applied to a target device, the target device comprising at least a tactile sensor, a physiological signal sensor, and a pressure sensor, the method includes: The tactile contour subsequences collected by the tactile sensor at each time step are acquired, and the micro-gesture events at each time step are determined based on the tactile contour subsequences. Based on the tactile contour subsequence acquired at each time step, the dynamic feature subsequence of the corresponding micro-gesture event is determined, and combined with the dynamic feature subsequence of each time step, the micro-gesture dynamic time sequence corresponding to the current target time window is constructed. For each time step, the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area is calculated based on the corresponding tactile contour subsequence, and the skin contact area with the largest tactile contour amplitude is selected as the physiological signal acquisition area. In each time step, physiological signals are acquired from the physiological signal acquisition area by the physiological signal sensor, and the physiological signal sequence corresponding to the target time window is constructed by combining the physiological signals corresponding to each time step. The grip force signal collected by the pressure sensor at each time step is acquired, and the grip force signal corresponding to each time step is combined to construct the grip force sequence corresponding to the target time window; Using a large language model, based on the micro-gesture dynamics time series, the physiological signal series, and the grip force series corresponding to the target time window, the emotional change trend from the target time window to the next target time window is predicted, and the emotional change prediction result is obtained.
2. The emotion prediction method based on multiple biological signals according to claim 1, characterized in that, The method uses a large language model to predict the emotional change trend from the target time window to the next target time window based on the micro-gesture dynamics time series, the physiological signal series, and the grip force series corresponding to the target time window, obtaining the emotional change prediction result, including: Obtain the emotional baseline corresponding to the target object being detected; Based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, input prompt information is generated; Using a large language model, based on the input prompt information, the trend of emotional change from the target time window to the next target time window is predicted, and the emotional change prediction result is obtained.
3. The emotion prediction method based on multiple biological signals according to claim 2, characterized in that, The acquisition of the emotional baseline corresponding to the target object being detected includes: For each reference time window, the initial micro-gesture dynamics time series, initial physiological signal series, and initial grip force series of the target object in a resting state are collected; Based on multiple initial micro-gesture dynamics time series, multiple initial physiological signal sequences, and multiple initial grip force sequences corresponding to multiple reference time windows, calculate the statistical parameters of micro-gesture dynamics, physiological signal statistical parameters, and grip force sequence statistical parameters. Based on the micro-gesture dynamics statistical parameters, the physiological signal statistical parameters, and the grip force sequence statistical parameters, the emotional baseline corresponding to the target object is obtained.
4. The emotion prediction method based on multiple biological signals according to claim 2, characterized in that, Before generating input prompt information based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, the process includes: Obtain the first trust parameter corresponding to the micro-gesture dynamics time series, and obtain the second trust parameter corresponding to the physiological signal sequence and the grip force sequence; A first descriptive information for the micro-gesture dynamics time series is generated based on the first trust parameter, and a second descriptive information for the physiological signal sequence and the grip force sequence is generated based on the second trust parameter; The step of generating input prompt information based on the emotional baseline, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence includes: Based on the emotional baseline, the first descriptive information, the second descriptive information, the micro-gesture dynamics time series corresponding to the target time window, the physiological signal sequence, and the grip force sequence, input prompt information is generated.
5. The emotion prediction method based on multiple biological signals according to claim 4, characterized in that, The step of obtaining the first trust parameter corresponding to the micro-gesture dynamics time series and the second trust parameter corresponding to the physiological signal sequence and the grip force sequence includes: Obtain the device context parameters for the target object to control the target device, wherein the device context parameters include at least the steering angular velocity and the speed; A first threshold is calculated based on the speed, wherein the first threshold is positively correlated with the speed; When the steering angular velocity exceeds the first threshold, a first trust parameter corresponding to the micro-gesture dynamics time series is calculated, and a second trust parameter corresponding to the physiological signal sequence and the grip force sequence is calculated, wherein the first trust parameter is less than the second trust parameter.
6. The emotion prediction method based on multiple biological signals according to claim 2, characterized in that, The process involves using a large language model to predict the emotional change trend from the target time window to the next target time window based on the input prompt information, thereby obtaining the emotional change prediction result, including: Using a large language model, based on the input prompt information, emotion prediction based on multiple biological signals is performed on the target object to obtain the emotion vector trajectory corresponding to the target time window; Based on the input prompt information and the emotion vector trajectory, the emotion of the target object is predicted in the next target time window to obtain the predicted emotion vector trajectory corresponding to the next target time window. Based on the emotion vector trajectory and the predicted emotion vector trajectory, a risk level summary for the target object is obtained; Based on the emotion vector trajectory, the predicted emotion vector trajectory, and the risk level summary, the emotion change prediction result is obtained.
7. The emotion prediction method based on multiple biological signals according to claim 1, characterized in that, The determination of the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence acquired at each time step includes: For each time step, determine the corresponding sliding sub-time window with each time step as the endpoint; Based on the tactile contour subsequence of at least one time point contained in the sliding sub-time window, the micro-gesture frequency corresponding to the micro-gesture event, the average rubbing intensity of the detected target object on the skin contact area corresponding to the tactile sensor, the rhythmic index, the cumulative hand displacement, and the pressure trend slope are determined. The rhythmicity index is used to characterize the degree of aggregation of the micro-gesture events corresponding to the sliding sub-time window on the time axis. The cumulative hand displacement at each time step is obtained by accumulating the tactile contour changes between adjacent time steps included in the sliding sub-time window. The tactile contour changes at each time step are determined based on the differences between the tactile contour subsequences between adjacent time steps included in the sliding sub-time window. The pressure trend slope at each time step is obtained by linearly fitting the micro-gesture event intensity values of multiple time steps included in the sliding sub-time window. Based on the micro-gesture frequency, average rubbing intensity, rhythmic index, cumulative hand displacement, and pressure trend slope corresponding to the sliding sub-time window for each time step, a dynamic feature sub-sequence corresponding to each time step is determined.
8. A mood prediction device based on multiple biological signals, characterized in that, Applied to a target device, the target device comprising at least a tactile sensor, a physiological signal sensor, and a pressure sensor, the device comprising: The acquisition module is used to acquire the tactile contour subsequence collected by the tactile sensor at each time step, and determine the micro-gesture event at each time step based on the tactile contour subsequence; The determination module is used to determine the dynamic feature subsequence of the corresponding micro-gesture event based on the tactile contour subsequence collected at each time step, and to construct the micro-gesture dynamic time sequence corresponding to the current target time window by combining the dynamic feature subsequence of each time step. The calculation module is used to calculate the tactile contour amplitude output by the tactile sensor for the corresponding skin contact area based on the corresponding tactile contour subsequence for each time step, and select the skin contact area with the largest tactile contour amplitude as the physiological signal acquisition area. The acquisition module is used to acquire physiological signals from the physiological signal acquisition area through the physiological signal sensor at each time step, and to construct the physiological signal sequence corresponding to the target time window by combining the physiological signals corresponding to each time step; A construction module is used to acquire the grip force signal collected by the pressure sensor at each time step, and combine the grip force signal corresponding to each time step to construct the grip force sequence corresponding to the target time window; The prediction module is used to predict the emotional change trend from the target time window to the next target time window by using a large language model based on the micro-gesture dynamics time series, the physiological signal series and the grip force series corresponding to the target time window, and obtain the emotional change prediction result.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the emotion prediction method based on multiple biological signals as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the emotion prediction method based on multiple biological signals as described in any one of claims 1 to 7.