AI robot emotional state image recognition system based on micro expression analysis

By measuring the mechanical background vibration frequency, extracting local phase information, and constructing a quantized residual map, the problem of misjudgment of mechanical vibration in robot emotion recognition was solved, achieving efficient emotion state recognition and system self-healing effect.

CN121962780APending Publication Date: 2026-05-01深圳市永迦电子科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳市永迦电子科技有限公司
Filing Date
2026-03-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to distinguish between mechanical vibrations and emotional expressions in robot emotion recognition. They lack decoupling mechanisms for micron-level phase changes and quantified motion trajectories, and cannot effectively identify the unique non-rigid texture flickering patterns of robots, resulting in poor anti-interference capabilities and low specificity in emotion state determination.

Method used

The mechanical background vibration frequency is measured by a visual data acquisition unit, local phase information is extracted by an Euler phase enhancement unit, a quantization residual decoupling unit is constructed to generate a quantization residual map, and the state classification and feedback unit is combined to determine the conflict state of emotion computing and generate an emotion state recognition signal or system optimization instruction.

Benefits of technology

It can accurately filter mechanical vibrations in environments with extremely low signal-to-noise ratios, significantly improve the anti-interference ability of emotion recognition on non-biological carriers, accurately capture the subtle stiffness and tremors of robots caused by computing power allocation delays or logical conflicts, achieve deep analysis of non-rigid texture flickering patterns, and provide a complete closed-loop effect from emotion recognition to system self-healing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962780A_ABST
    Figure CN121962780A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and robots, in particular to a robot emotional state image recognition system based on micro-expression analysis, which comprises a data acquisition and frequency domain locking step: acquiring a facial time sequence image, measuring a mechanical background vibration frequency and delimiting a target enhancement frequency band; a phase amplification and enhancement step: extracting a local phase in a complex pyramid domain, and performing Euler video amplification on a component in a frequency band; a residual error decoupling and modeling step: constructing an actual quantization and ideal smooth motion model, and calculating a difference value to generate a quantization residual error map; a state judgment and feedback step: extracting space-time pulse characteristics, judging an emotion calculation conflict state and generating a tuning instruction; according to the method, micron-sized tremor is effectively captured, and a bottom layer driving state is converted into visual emotion features.
Need to check novelty before this filing date? Find Prior Art

Description

An AI robot emotion state image recognition system based on micro-expression analysis Technical Field

[0001] This invention relates to the fields of computer vision and robotics, specifically to a method based on micro-expression analysis. Robot emotional state image recognition system. Background Technology

[0002] In the field of human-computer interaction, vision-based micro-expression analysis technology has gradually become the mainstream means of endowing artificial intelligence robots with human-like emotional understanding and expression capabilities because it can capture subtle emotional changes. At present, the emotion recognition of robots mainly uses feature extraction algorithms for biological faces. If the system needs to accurately determine the emotional state of the robot, it usually relies on traditional temporal motion analysis or spatial feature matching of robot facial image sequences to obtain emotional feature data for classification.

[0003] However, when constructing an emotion recognition system based on a non-biological carrier using the above methods, directly applying biometric algorithms makes it difficult to eliminate the robot's inherent mechanical background vibration frequencies, such as the regular vibrations generated by cooling fans or motors. This leads to the misjudgment of normal mechanical physical vibrations as emotional expressions, resulting in poor anti-interference capabilities. At the same time, this processing method mainly focuses on macroscopic pixel displacement or texture changes, lacking a decoupling mechanism for micrometer-level phase changes and quantized motion trajectories. It cannot effectively distinguish between subtle stiffness and tremors caused by computational power allocation delays or logical conflicts, nor can it quantify mechanical damping characteristics through the energy attenuation rate of spatiotemporal pulse features. This results in difficulty in capturing physical mapping characteristics caused by deep logical conflicts. Furthermore, it ignores the quantization residual between the actuator's executed action and the ideal smooth command, making it difficult to identify the robot's unique non-rigid texture flickering patterns, resulting in low specificity for emotion state determination. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a method based on micro-expression analysis. The robot emotion state image recognition system, specifically, includes the following technical solution: a visual data acquisition unit for acquiring time-series image data of the robot's face, performing spectral analysis on the static background region in the time-series image data, determining the mechanical background vibration frequency, and defining a target enhancement frequency band based on the mechanical background vibration frequency; an Euler phase enhancement unit for transforming the time-series image data to the complex pyramid domain, extracting local phase information, and performing Euler video amplification processing on the phase change components within the target enhancement frequency band to generate a phase-enhanced image sequence; a quantization residual decoupling unit for constructing an actual quantized motion model and an ideal smooth motion model based on the phase-enhanced image sequence, and calculating the difference between the actual quantized motion model and the ideal smooth motion model to generate a quantization residual map; and a state classification and feedback unit for extracting spatiotemporal pulse features from the quantization residual map, determining the emotion calculation conflict state based on the spatiotemporal pulse features, and generating an emotion state recognition signal or system optimization instruction based on the determination result.

[0005] Preferably, the process of the visual data acquisition unit delineating the target enhancement frequency band is as follows: identifying the higher harmonic components of the mechanical background vibration frequency, setting the frequency range where the higher harmonic components are located as the target enhancement frequency band, or setting the complement outside the mechanical background vibration frequency and its neighboring frequencies as the target enhancement frequency band.

[0006] Preferably, the process of generating a phase-enhanced image sequence using the Euler phase enhancement unit is as follows: the image is decomposed into sub-bands of different scales and directions using the complex Smith-Pyramid transform, the phase of the sub-band corresponding to the target enhancement frequency band is subjected to temporal filtering and amplitude amplification in the complex domain, and the phase-enhanced image sequence is obtained by inverse transformation.

[0007] Preferably, the process of generating the quantization residual map by the quantization residual decoupling unit is as follows: performing pixel-by-pixel difference calculation on the actual quantized motion trajectory data and the ideal smooth motion trajectory data in the spatiotemporal domain, obtaining the absolute value of the motion vector difference, mapping the absolute value of the motion vector difference to grayscale values ​​or heatmaps, and obtaining the quantization residual map.

[0008] Preferably, the process of generating the quantization residual map by the quantization residual decoupling unit is as follows: performing pixel-by-pixel difference calculations on the actual quantized motion model and the ideal smooth motion model in the spatiotemporal domain, obtaining the absolute value of the motion vector difference, mapping the absolute value of the motion vector difference to grayscale values ​​or heatmaps, and obtaining the quantization residual map.

[0009] Preferably, the process by which the state classification and feedback unit determines the state of emotional computing conflict is as follows: calculate the energy density integral of the quantized residual map within the time window, and use the energy density integral as the residual strength index; monitor the fluctuation pattern of the residual strength index in the time domain, and when the fluctuation pattern shows high-frequency oscillation characteristics, it is determined that there is an emotional computing conflict state.

[0010] Preferably, the process by which the state classification and feedback unit generates an emotional state recognition signal is as follows: the residual intensity index is compared with a preset silence threshold and a conflict threshold; when the residual intensity index is less than or equal to the silence threshold, a fake smooth state signal is generated; when the residual intensity index is greater than the silence threshold and less than or equal to the conflict threshold, a potential computational anxiety signal is generated; when the residual intensity index is greater than the conflict threshold, a real emotional leakage signal is generated.

[0011] Preferably, the process of the state classification and feedback unit generating system tuning instructions is as follows: in response to the judgment result being an emotion computing conflict state, the position coordinates of the actuator that caused the conflict are extracted, and a low-level drive smoothing parameter correction instruction is generated for the position coordinates, or a computing resource reallocation instruction is generated for the upper-level logic.

[0012] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention uses a visual data acquisition unit to perform spectral analysis on the static background area to determine the mechanical background vibration frequency, and further locks the higher harmonic components or complementary frequency bands as the target enhancement frequency band. Combined with the complex pyramid domain transformation of the Euler phase enhancement unit, it achieves the effect of accurately filtering the inherent vibration of the cooling fan or motor in an environment with extremely low signal-to-noise ratio. Compared with the traditional algorithm that directly processes the original image, it can explicitly amplify the micron-level nonlinear vibration caused by abnormal current pulses or torque counteraction, solve the problem of misjudging normal mechanical physical vibration as emotional expression, and significantly improve the anti-interference ability of non-biological carrier emotion recognition.

[0013] 2. This invention constructs an actual quantized motion model and an ideal smooth motion model through a quantized residual decoupling unit, and introduces an expectation-reality difference mechanism to generate a quantized residual map, thereby achieving the effect of transforming the underlying driving state into a visual emotional feature. By identifying the step effect in the phase-enhanced image sequence and calculating its deviation from the polynomial fitting trajectory, it can accurately capture the subtle stiffness and vibration of the robot caused by the stepper motor step angle limitation or computing power allocation delay. This solves the problem that traditional spatial domain amplification algorithms cannot identify the robot's unique quantized step features and cannot effectively distinguish the physical mapping features caused by logical conflicts.

[0014] 3. This invention extracts spatiotemporal pulse features from the quantized residual spectrum through a state classification and feedback unit, and uses the logarithmic least squares method to fit the energy decay rate to quantify the mechanical damping characteristics, achieving the effect of deep analysis of non-rigid texture flickering modes; combined with the comprehensive judgment of zero-crossing rate and residual intensity index, it can effectively distinguish between undamped random electronic noise and mechanical structure resonance with specific damping characteristics, avoiding misjudging sensor dark current noise as micro-expressions, and solving the problem that existing technologies cannot verify conflicting states of emotion calculation through physical attributes on non-biological faces.

[0015] 4. This invention establishes a hierarchical early warning mechanism and closed-loop optimization strategy based on residual strength indicators. It utilizes homography matrix transformation to accurately locate the physical coordinates of actuators causing conflicts and generates low-level drive smoothing parameter correction instructions or upper-level computing resource reallocation instructions based on the judgment results. This achieves a complete closed-loop effect from emotion recognition to system self-healing. It can not only identify fake smoothing or genuine emotion leakage states but also dynamically adjust... By controlling gain or process priority, the problem that simple emotion recognition systems cannot provide intuitive basis and substantial intervention means for robot maintenance and performance optimization is solved. Attached Figure Description

[0016] The present invention will be further explained below with reference to the accompanying drawings and embodiments: Figure 1 is a structural diagram of the system of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0018] Example 1: Please refer to Figure 1, a method based on micro-expression analysis. A robot emotion state image recognition system includes: a visual data acquisition unit for acquiring time-series image data of the robot's face, performing spectral analysis on the static background region in the time-series image data, determining the mechanical background vibration frequency, and defining a target enhancement frequency band based on the mechanical background vibration frequency; an Euler phase enhancement unit for transforming the time-series image data to the complex pyramid domain, extracting local phase information, and performing Euler video amplification processing on the phase change components within the target enhancement frequency band to generate a phase-enhanced image sequence; a quantization residual decoupling unit for constructing an actual quantized motion model and an ideal smooth motion model based on the phase-enhanced image sequence, and calculating the difference between the actual quantized motion model and the ideal smooth motion model to generate a quantization residual map; and a state classification and feedback unit for extracting spatiotemporal impulse features from the quantization residual map, determining the emotion calculation conflict state based on the spatiotemporal impulse features, and generating an emotion state recognition signal or system optimization instruction based on the determination result.

[0019] This embodiment details an image decoupling architecture based on first principles of physics, aiming to address the adaptability issues of traditional biometric recognition algorithms on non-biological carriers. The system uses a visual data acquisition unit to acquire raw time-series image data of the robot's face using a high-frame-rate industrial camera. To ensure the accuracy of subsequent analysis, this unit employs a temporal variance thresholding method to identify static background regions, specifically calculating the gray-level variance of each pixel in the image sequence. The calculation method involves selecting a length of... It is recommended to use a sliding window of 20-50 frames to calculate the pixels within the window. The variance; to eliminate calculation ambiguity, this embodiment explicitly uses a divisor of . The overall variance formula quantifies the absolute dispersion of the current window's data, and the calculation formula is as follows:

[0020] in, : Length of the time sliding window, sourced from a preset value, in frames; Pixels within this time window The grayscale mean is calculated using the following formula: The unit is grayscale value; :time The pixel grayscale value, in grayscale values; areas with a variance value lower than a preset macroscopic motion threshold are marked as static background areas, such as the edge of a robot's shell or a fixed support. Specifically, this threshold... The method for determining this is as follows: data is collected under conditions of no light. For example, in a 50-frame dark field image, the mean of the grayscale variance of all pixels is calculated as the sensor's dark current noise floor. ;set up ;in, Safety factor, derived from a preset constant, dimensionless, with a value range of 3.0 to 5.0. In this embodiment, it is preferably 3.0 to ensure that it can cover 99.7% of random electronic noise; The sensor has low dark current noise, derived from statistical calculations, with units of grayscale value squared. Simultaneously, this calculation result must be less than a preset minimum grayscale variance threshold for active facial expressions, for example, 2.0, corresponding to grayscale changes caused by minute muscle twitches. The generated binarized mask is then processed. Morphological opening operations are performed on the structuring elements to eliminate isolated salt-and-pepper noise points, thereby eliminating interference from active facial expression movements. The specific algorithm flow for the spectral analysis performed by this unit is as follows: spatial averaging is performed on pixels marked as background static regions to generate a background average brightness time series.

[0021] in, The set of pixels in the static background region, derived from the variance threshold determination mentioned above; :gather The base number, which is the total number of pixels in the background area, is 1. :time The average background brightness, expressed in grayscale values; A Hanning window is applied to suppress spectral leakage, and a fast Fourier transform is performed. Obtain the spectrum ; within the preset effective mechanical frequency band For example, 5 100 The frequency point with the largest internal search amplitude, i.e. This value is used as the measured mechanical background vibration frequency. Based on this, the Euler phase enhancement unit focuses on local phase changes rather than simple pixel displacement, explicitly amplifying micro-vibrations invisible to the naked eye through complex pyramid domain transformation. The quantization residual decoupling unit introduces an expectation-reality difference mechanism, subtracting the fitted smooth trajectory from the actually observed step-like or oscillating motion trajectory, thereby constructing a quantization residual map reflecting the physical defects of the actuator. The state classification and feedback unit identifies non-rigid texture flicker patterns. Specifically, the extraction process of spatiotemporal pulse features is as follows: Second-order temporal difference analysis is performed on each pixel in the quantization residual map, and the calculation formula is:

[0022] in, The temporal second-order difference value of a pixel is used to characterize abrupt changes in acceleration, and the unit is grayscale value; : Quantization residual plot at time The pixel grayscale values ​​originate from the quantization residual decoupling unit; the second-order difference values... Points exceeding three times the mean of their local neighborhood are marked as pulse points, where the local neighborhood is defined as centered on the current pixel. Spatial window, mean The average level of the absolute differences within this window is used to aggregate features into a feature vector that includes spatial location, pulse duration, and energy decay rate; where the energy decay rate... The calculation formula is as follows: Select the falling edge of the pulse, that is, from the peak value... The time period up to the regression threshold For discrete sampling points, if the number of sampling points is less than 3, i.e. the pulse decreases extremely rapidly, the pulse half-width time is directly used as a substitute feature; a residual intensity sequence is constructed. and relative time series The peak time is defined as Subsequent sampling point times are offsets relative to the peak time, in seconds; to avoid logarithmic calculation anomalies, ... Perform non-zero truncation, i.e., take , For a preset minimum positive number, such as Using log-least squares to fit the exponential model Transform it into a linear regression form passing through the origin. Its analytical solution is:

[0023] in, Energy decay rate, in physical terms, is a quantitative mechanical damping characteristic, and its unit is... ; Time offset relative to the peak time, in units of ; The pulse peak value is expressed as a grayscale value. This value is used to quantify the mechanical damping characteristics. Based on these characteristics, the unit determines the emotional computing conflict state and generates corresponding control commands in response to the determination result. In scenarios where the robot is under high-load computing or logical deadlock, this embodiment, through the collaborative work of the above units, can accurately capture the micron-level stiffness or tremor caused by the robot's computing power allocation delay or logical conflict in an environment with extremely low signal-to-noise ratio. This effectively transforms the underlying driving state into visualized emotional features and avoids the defect of misjudging the vibration of the cooling fan as emotional expression.

[0024] Example 2: The process of the visual data acquisition unit delineating the target enhancement frequency band is as follows: identify the higher harmonic components of the mechanical background vibration frequency, set the frequency range where the higher harmonic components are located as the target enhancement frequency band, or set the complement of the mechanical background vibration frequency and its neighboring frequencies as the target enhancement frequency band.

[0025] This embodiment further defines the frequency band allocation logic of the visual data acquisition unit; considering that conflicts in robot emotion calculations often lead to nonlinear clipping or distortion of motor drive signals, this nonlinearity manifests as higher harmonics of the fundamental frequency in the frequency domain; therefore, the system executes the following frequency band locking strategy:

[0026] in, : Harmonic order, derived from a preset integer, physically represents the number of harmonics of the nonlinear distortion, with a unit of 1; : Upper limit of harmonic order, which is derived from a preset constant. The specific value range is recommended to be 3~7. In this embodiment, it is preferably 5, which aims to cover the main mechanical nonlinear characteristics while avoiding high-frequency noise interference. The physical meaning is the maximum harmonic order that participates in band-locking, and the unit is 1. The mechanical background vibration frequency, derived from FFT analysis of the background static region, physically represents the robot's fundamental vibration frequency during standby, and is measured in units of... ; The bandwidth (half-width at half maximum) is derived from the camera sampling rate calculation, and the calculation formula is:

[0027] in, For camera sampling rate, The length of the spectrum analysis window, i.e., the number of sampling points, is 128 in this embodiment. Its physical meaning is the frequency resolution determinant, with a unit of 1. Alternatively, the system adopts a complement filtering strategy, treating the inherent frequency and its neighborhood as whitelisted noise and removing them, while retaining the remaining frequency bands as the target enhancement band. In complex industrial or service scenarios, this embodiment can accurately filter out normal physical vibrations of the robot, such as regular noise caused by fan rotation, by locking high-order harmonics or complements, while specifically amplifying high-frequency nonlinear vibration signals caused by abnormal current pulses or torque counteraction, thereby greatly improving the specificity of emotional state recognition.

[0028] Example 3: The process of generating a phase-enhanced image sequence using the Euler phase enhancement unit is as follows: The image is decomposed into sub-bands of different scales and directions using the complex Smith-Pyramid Transform. In the complex domain, the phase of the sub-band corresponding to the target enhancement frequency band is subjected to temporal filtering and amplitude amplification, and the phase-enhanced image sequence is obtained by inverse transformation.

[0029] This embodiment further defines the processing flow of the Euler phase enhancement unit. To amplify minute motions without introducing optical flow artifacts, the unit utilizes the complex Smith-Pyramid transform, an improved multi-scale complex domain decomposition method. The specific implementation of this transform involves employing a frequency-domain recursive filtering architecture to ensure the completeness of phase information. A complex Smith filter bank is designed, defining frequency-domain polar coordinates. Construct radial basis functions The domain is Used to capture energy in specific frequency bands, and angular basis functions:

[0030] in, This represents the total number of directions; in this embodiment, 4 corresponds to 0°, 45°, 90°, and 135°. Normalization constant ; Perform the pyramid construction steps: for the current layer image Obtain the spectrum by performing Fast Fourier Transform (FFT). ; the spectrum Multiply by the bandpass filter respectively And perform the inverse FFT to obtain Complex subband coefficients in each direction, including local amplitude and phase; the spectrum Multiplied by low-pass filter It also performs interval downsampling, i.e., frequency domain clipping, as the input spectrum for the next layer of the pyramid. This process replaces the traditional real-field Laplace pyramid, thus supporting precise operations on sub-pixel phase; specifically, the number of pyramid layers... Based on the image resolution setting, the calculation formula is satisfied:

[0031] To ensure the top subband contains sufficient spatial information, the filter's radial bandwidth is set to 1 octave and its directional bandwidth to 30 degrees to achieve tight frequency domain coverage. The subband phase corresponding to the target frequency band is subjected to time-domain filtering, and an amplification factor is applied, calculated using the following formula:

[0032] in, The original local phase, derived from complex pyramid transform, physically represents the spatial phase information of an image in a specific sub-band, measured in units of... ; Euler amplification factor, which is preset based on the image noise level, physically represents the gain factor for phase changes, and is in units of 1. The transfer function, derived from the target enhancement frequency band setting, specifically employs a Gaussian smoothing bandpass filter to suppress time-domain ringing effects. Its mathematical expression is:

[0033] in, For the first The center frequency of each target frequency band. , is the bandwidth of the k-th frequency band, which physically represents the frequency response characteristic of the bandpass filter, and its unit is 1; The transform operator, derived from the algorithm definition, physically represents a one-dimensional fast Fourier transform along the time axis. This is an inverse transformation, with a unit of 1; it should be noted that this transformation operator It performs frequency domain filtering on the phase change of each pixel over time, rather than spatial domain transformation on a single frame of image; during the reconstruction phase, the system explicitly executes amplitude and phase reconstruction logic: preserving the original amplitude of each sub-band. Remain unchanged, and compare it with the corrected phase. Combined to generate new complex subband signals , Using the imaginary unit, a phase-enhanced image sequence is obtained by reconstructing the modified complex subband signal through inverse transform; this is to address potential issues with the robot's face. Screen refresh rate jitter or The high-frequency flickering of the array, in this embodiment, utilizes the phase translation characteristics of the complex domain to replace the traditional pixel displacement analysis, which can handle sub-pixel level motion, capture fine features that cannot be identified by traditional spatial domain magnification algorithms, and realize in-depth analysis of electronic facial micro-expressions.

[0034] Example 4: The process of constructing the model using the quantization residual decoupling unit is as follows: perform polynomial fitting or low-pass filtering to smooth the motion trajectory in the phase-enhanced image sequence to generate an ideal smooth motion model; retain the original motion trajectory containing the step effect in the phase-enhanced image sequence as the actual quantization motion model.

[0035] This embodiment further defines the model construction process of the quantization residual decoupling unit; this step aims to establish a benchmark for comparing the expected and the actual. The system performs polynomial fitting on the motion trajectory in the phase-enhanced image sequence to construct an ideal smooth motion model, the formula of which is expressed as:

[0036] in, Ideal position value, derived from calculation results, physically represents the macroscopic command intent issued by the upper-level control logic, and is measured in units of... or ; Fitting coefficients, derived from least squares calculations, physically represent the shape parameters of the trajectory curve, and are measured in units of... ,when When it is a pixel, or ,when When the value is in radians, this ensures that the physical dimensions of each term in the polynomial are consistent with those on the left side. Time variable, sourced from the system clock, physically represents the sampling time, and is measured in seconds (s). The order of the fitted polynomial is derived from preset parameters and its physical meaning is the upper limit of the complexity of the smoothing model, with a unit of 1; its parameter selection is based on: setting This corresponds to a cubic spline curve because the macroscopic motion control of a robot usually follows the principle of minimum jerk programming. A cubic polynomial is sufficient to fit the smooth rigid body motion trajectory of this type. Any high-frequency fluctuations that deviate from this order are considered as micro-expressions or quantization noise that need to be separated. At the same time, the system directly retains the original motion trajectory containing the step effect in the phase-enhanced image sequence and uses it as the actual quantization motion model to truly reflect the physical execution of the actuator. When the robot attempts to execute a smooth expression but is limited by the stepper motor step angle or drive conflict, this embodiment can identify the robot's unique quantization step characteristics by distinguishing between the ideal model and the actual model. This high-frequency pause is the key physical evidence for judging whether the robot is experiencing computational anxiety.

[0037] Example 5: The process of generating a quantization residual map by the quantization residual decoupling unit is as follows: Perform pixel-by-pixel difference calculation on the actual quantized motion trajectory data and the ideal smooth motion trajectory data in the spatiotemporal domain to obtain the absolute value of the motion vector difference. Map the absolute value of the motion vector difference to grayscale values ​​or heatmaps to obtain the quantization residual map.

[0038] This embodiment further defines the generation process of the quantized residual map; after obtaining the ideal and actual models, the system calculates their differences in the spatiotemporal domain; it is worth noting that, in order to unify physical dimensions, the system performs first-order time derivative processing on the constructed actual quantized motion model and ideal smooth motion model based on position trajectory, respectively, to obtain the corresponding velocity vector data; in particular, if the ideal smooth motion model... Based on joint angles, units If we construct the model, we need to multiply by the projection transformation coefficients before differentiating. The unit is Symbols are used here. To distinguish it from the energy decay rate in Example 1 This is done by mapping the values ​​to the image pixel domain to eliminate computational errors caused by inconsistencies in units; then, for each pixel in the image, its quantization residual value is calculated:

[0039] in, Quantitative residual value, derived from differential calculation, physically represents the degree of deviation between the actuator's executed action and the ideal command, and is measured in units of... ; The actual motion vector is obtained by differentiating the actual quantized motion model. Its physical meaning is the actual rate of change of the current pixel, and its unit is... ; The ideal motion vector, derived from an ideal smooth motion model after differentiation and dimensional normalization, represents the expected rate of change of the current pixel, with units of . The system maps the residual value to a grayscale value, specifically using the following saturated linear mapping formula:

[0040] in, : pixel grayscale value at the location, range ; : Saturation velocity threshold, derived from a preset empirical value, such as 5 The physical meaning is the upper limit of the quantization error; any error exceeding this rate will be displayed as pure white highlight. The truncation function ensures that the values ​​are within the effective grayscale range, resulting in a quantized residual map. The map generated in this embodiment visually displays the stress distribution on the robot's face. When the robot is in a logical conflict or resource contention, specific areas on the face, such as the servo motor connection, will appear as bright spots on the map, realizing the physical visualization of invisible logical conflicts and enabling maintenance personnel to quickly locate the source of the fault.

[0041] Example 6: The process of the state classification and feedback unit to determine the sentiment computing conflict state is as follows: Calculate the energy density integral of the quantized residual spectrum within the time window, and use the energy density integral as the residual strength index; monitor the fluctuation pattern of the residual strength index in the time domain, and when the fluctuation pattern shows high-frequency oscillation characteristics, it is determined that there is a sentiment computing conflict state.

[0042] This embodiment further defines the process for determining the conflict state of emotion computing. This process is a specific algorithm implementation of Embodiment 1 for determining the conflict state of emotion computing based on spatiotemporal pulse features. That is, discrete spatiotemporal pulse features, such as spatial location and pulse duration, are mapped to global energy indicators for state monitoring. In order to quantify the degree of conflict, the system introduces a residual intensity index, which is a global energy aggregation representation of spatiotemporal pulse features. Its calculation method follows the following integral logic:

[0043] in, The residual strength index, derived from integral calculation, physically represents the total error energy density within a time window, and its unit is... Equivalent value; The equivalent energy conversion coefficient is derived from system calibration. Its specific value must satisfy dimensional balance, and its unit is set to [unit missing]. This is because of the integral term. The physical dimensions are Therefore, this coefficient needs to be eliminated. And introduce This ensures the final target It has a clear meaning of physical energy density; the specific implementation steps of system calibration are as follows: in a controlled environment, a standard grayscale target is driven by an exciter to achieve a frequency of... , such as 10 Amplitude is , such as 0.5 The simple harmonic motion, its theoretical energy density is calculated according to the physical dynamics formula. The unit is Simultaneously acquire the image sequence of this process and calculate the original pixel integral value. The unit is Ultimately, through the formula Determine the specific value of this coefficient; : Time window length, sourced from a preset value, physically represents the observation period of micro-expression duration, corresponding to the upper limit of pulse duration statistics in spatiotemporal pulse characteristics, and is measured in units of ; Spatial integration region, sourced from the image plane coordinate system, physically represents the set of facial regions involved in energy calculation, corresponding to the spatial location distribution in spatiotemporal pulse features, with units of [unit missing]. ; Quantitative residual value, derived from the calculation results of the previous level, physically represents the intensity of single-point error, and is measured in units of... Based on this, the system calculates the zero-crossing rate. The high-frequency oscillation characteristics are quantified by combining physical damping properties. The specific algorithm steps are as follows: For the residual strength index sequence... By performing high-pass filtering or subtracting the moving average, the fluctuation component is obtained:

[0044] in, Define the width of the sliding window; use the sign function to calculate the number of sign changes per unit time:

[0045] in, To monitor the length of the time window, the unit is... , This represents the number of sampling points within the window. For symbolic functions; during the system startup initialization phase, data is collected. Frame, such as Residual strength index sequence under completely static state Calculate its variance as the silent noise variance:

[0046] This establishes the background noise level of the environment; the calculated noise level is then used to determine the background noise level. With the preset oscillation frequency threshold Comparison was performed; among them, The source is set as the inherent modal frequency of the robot's mechanical structure, obtained through factory frequency sweep testing, for example, 20. 0.5 times; simultaneously, in order to verify the physical properties of the oscillation, namely the energy attenuation rate in the spatiotemporal pulse characteristics of Example 1, the system calls the energy attenuation rate calculated in Example 1. Since a single discrete pulse point cannot be independently fitted exponentially, feature aggregation logic is performed here: a set of spatiotemporally continuous pulse points is defined as an independent pulse event. The specific criterion for determining spatiotemporal continuity is: for any two pulse points... and If the time interval is satisfied In this embodiment, one frame is used, meaning frames that are temporally adjacent and spatially Euclidean distance apart.

[0047] This embodiment takes If a pixel is spatially 8-neighborhood connected, then two points are determined to belong to the same event; and a unique fitted decay rate value is assigned to this event, thereby calculating the average decay rate of all identified pulse events within the current window. The following conditions must be met: ; within the current window variance 3 represents the preset signal-to-noise ratio coefficient; ,For example The system characterizes the typical mechanical damping range and eliminates undamped random electronic noise; the system determines that the fluctuation pattern exhibits high-frequency oscillation characteristics, thus confirming the existence of an emotional computing conflict state.

[0048] Example 7: The process of the state classification and feedback unit generating emotional state recognition signals is as follows: The residual intensity index is compared with the preset silence threshold and conflict threshold; when the residual intensity index is less than or equal to the silence threshold, a fake smooth state signal is generated; when the residual intensity index is greater than the silence threshold and less than or equal to the conflict threshold, a potential computational anxiety signal is generated; when the residual intensity index is greater than the conflict threshold, a real emotional leakage signal is generated.

[0049] This embodiment further defines the generation process of emotion state recognition signals; the system sets a silence threshold and a conflict threshold based on the statistical distribution of historical operating data; the specific threshold calculation logic is as follows: data is collected during the self-check phase of robot startup. Calculate the mean of the residual intensity data in the idle state of the frame. with standard deviation Set the silence threshold to To cover the sensor's background noise, the collision threshold is set to... To define significant abnormal oscillations, the system executes the following hierarchical judgment logic: When the residual strength index is less than or equal to the silent threshold, the system generates a fake smooth state signal, indicating extremely smooth robot motion that conforms to the ideal model; when the index is between the silent threshold and the conflict threshold, the system generates a potential computational anxiety signal, indicating slight quantization error oscillations; when the index exceeds the conflict threshold, the system generates a genuine emotion leakage signal, indicating that the actual motion deviates significantly from the ideal trajectory. This embodiment constructs a hierarchical early warning mechanism for robot emotions, particularly capable of identifying fake smooth states. This is of great significance in social robot interaction scenarios, as it can alert users that the robot may be executing a preset deceptive expression program, or conversely, indicate a shortage of computing resources when genuine emotion leakage is detected, providing an intuitive basis for maintaining system stability.

[0050] Example 8: The process of the state classification and feedback unit generating system tuning instructions is as follows: In response to the judgment result being a conflict state in emotion computing, the position coordinates of the actuator that caused the conflict are extracted, and a low-level drive smoothing parameter correction instruction is generated for the position coordinates, or a computing resource reallocation instruction is generated for the upper-level logic.

[0051] This embodiment further defines the generation process of system tuning instructions; when the judgment result is a conflict state in emotion computing, the system executes the following closed-loop control logic: extract the actuator position coordinates that caused the conflict; to avoid black-boxing of the mapping process, the system adopts a combination of weighted centroid method and homography matrix transformation: in the quantized residual map, calculate the weighted centroid coordinates of the high residual region. :

[0052] in, To determine the region of interest (ROI), a pre-calibrated camera-mechanical structure homography matrix was used. The source is the factory calibration, which converts image coordinates into physical structure coordinates; the specific calculation process is as follows: matrix multiplication is performed. Perform homogeneous coordinate normalization to obtain the actual physical coordinates. ,Right now This step eliminates the nonlinear distortion caused by perspective projection, ensuring the geometric accuracy of coordinate mapping; in the robot actuator layout database, searches are performed to... The nearest motor to Euler As the target actuator, it generates low-level drive smoothing parameter correction instructions for the position coordinates; the system adopts a gain scheduling strategy based on residual strength, and the specific correction formula is as follows:

[0053] in, : The differential gain parameter of the controller; increasing this value can enhance damping to suppress high-frequency oscillations. : Adjustment rate coefficient, which is determined empirically, such as 0.1; : Conflict threshold, derived from the definition in Example 7; The theoretical limit of saturated integrals is derived from the calculation formula:

[0054] in, The maximum optical flow velocity that the sensor can measure, in units of , The image domain integral area of ​​the region of interest, in units Its value is equal to the total number of pixels within the region of interest; it needs to be clarified here that the physical area in the original statement refers to the physical region range corresponding to this parameter, but the pixel dimension must be used in the formula calculation to be consistent with... In The denominators are dimensionally canceled out to ensure have The dimension of energy density, This is the length of the time window; this parameter is used to convert the integral dimension to... Normalize to the 0,1 interval to avoid [the following] The value is too large, relative to the scalar sensor threshold, causing... The logical flaw of gain explosion; if the oscillation originates from the main control unit scheduling delay, i.e., the extracted actuator location is in the head main control board area, the system generates a computing resource reallocation instruction, specifically adjusting the Nice value of the emotion computing process, i.e., its priority, to:

[0055] This formula uses the normalized residual ratio to dynamically adjust the priority, ensuring that when the conflict intensity is close to the saturation limit, the priority is increased to the maximum extent and the Nice value is reduced significantly, thereby achieving a reasonable redistribution of computing resources.

[0056] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method based on micro-expression analysis A robot emotion state image recognition system, characterized in that, include: The visual data acquisition unit is configured to acquire time-series image data of the robot's face, perform spectral analysis on the background static area in the time-series image data, determine the mechanical background vibration frequency, and delineate the target enhancement frequency band based on the mechanical background vibration frequency. The Euler phase enhancement unit is configured to transform the time-series image data to the complex pyramid transform domain, extract local phase information, and perform Euler video amplification processing on the phase change components within the target enhancement frequency band to generate a phase-enhanced image sequence. The quantization residual decoupling unit is configured to construct actual quantized motion trajectory data and ideal smooth motion trajectory data based on the phase-enhanced image sequence, and calculate the difference between the actual quantized motion trajectory data and the ideal smooth motion trajectory data to generate a quantization residual map. The state classification and feedback unit is configured to extract spatiotemporal pulse features from the quantization residual map, determine the emotion calculation conflict state based on the spatiotemporal pulse features, and generate an emotion state recognition signal or system optimization instruction based on the determination result.

2. A method based on micro-expression analysis according to claim 1 A robot emotion state image recognition system, characterized in that, The process of defining the target enhancement frequency band by the visual data acquisition unit is configured as follows: identifying the higher harmonic components of the mechanical background vibration frequency, setting the frequency range where the higher harmonic components are located as the target enhancement frequency band, or setting the set of frequencies outside the mechanical background vibration frequency and its neighboring frequencies as the target enhancement frequency band.

3. A method based on micro-expression analysis according to claim 1 A robot emotion state image recognition system, characterized in that, The process of generating a phase-enhanced image sequence by the Euler phase enhancement unit is configured as follows: the image is decomposed into sub-bands of different scales and directions using complex pyramid transformation, the phase of the sub-band corresponding to the target enhancement frequency band is subjected to temporal filtering and amplitude amplification in the complex domain, and the phase-enhanced image sequence is obtained by inverse transformation.

4. A method based on micro-expression analysis according to claim 1 A robot emotion state image recognition system, characterized in that, The process of constructing data by the quantization residual decoupling unit is configured as follows: performing polynomial fitting or low-pass filtering smoothing on the motion trajectory in the phase-enhanced image sequence to generate the ideal smooth motion trajectory data; retaining the original motion trajectory containing the step effect in the phase-enhanced image sequence as the actual quantized motion trajectory data.

5. The AI ​​robot emotion state image recognition system based on micro-expression analysis according to claim 1, characterized in that, The process of generating the quantization residual map by the quantization residual decoupling unit is as follows: performing pixel-by-pixel difference calculation on the actual quantized motion trajectory data and the ideal smooth motion trajectory data in the spatiotemporal domain, obtaining the absolute value of the motion vector difference, mapping the absolute value of the motion vector difference to grayscale values ​​or heatmaps, and obtaining the quantization residual map.

6. A method based on micro-expression analysis according to claim 1 A robot emotion state image recognition system, characterized in that, The process of the state classification and feedback unit in determining the sentiment computing conflict state is configured as follows: calculate the energy density integral of the pixel values ​​of the quantized residual map within a preset time window, and use the energy density integral as the residual strength index; monitor the fluctuation pattern of the residual strength index in the time domain, and when the oscillation frequency of the fluctuation pattern exceeds the preset oscillation frequency threshold, it is determined that there is a sentiment computing conflict state.

7. A method based on micro-expression analysis according to claim 6 A robot emotion state image recognition system, characterized in that, The process of generating the emotional state recognition signal by the state classification and feedback unit is configured as follows: comparing the residual intensity index with a preset silence threshold and a conflict threshold, wherein the silence threshold is less than the conflict threshold. If the residual strength index is less than or equal to the silence threshold, a fake smooth state signal is generated; if the residual strength index is greater than the silence threshold and less than or equal to the conflict threshold, a potential computational anxiety signal is generated; if the residual strength index is greater than the conflict threshold, a real emotional leakage signal is generated.

8. A method based on micro-expression analysis according to claim 6 A robot emotion state image recognition system, characterized in that, The process of the state classification and feedback unit generating system tuning instructions is as follows: In response to the judgment result being an emotion computing conflict state, the position coordinates of the actuator that caused the conflict are extracted, and a low-level drive smoothing parameter correction instruction is generated for the position coordinates, or a computing resource reallocation instruction is generated for the upper-level logic.