A method, system, device and medium for analyzing interpersonal emotional synchronization

CN122799474APending Publication Date: 2026-09-22刘和钰
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610891577.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0010]上述的一种人际互动情感同步分析方法、系统、设备及介质,从双方的面部动作单元强度序列出发,通过双流卷积网络逐层提取个体表情动态特征,并在多个深度层级上依据信息压缩状态自适应生成门控向量对两流特征进行渐进式跨流融合,从而在保护单一个体微表情独立动态的同时充分捕获双方互动耦合的时序信息,获得互动感知特征序列;在此基础上,对互动感知特征序列进行互相关计算获得互相关函数,利用连续小波变换将互相关函数分解为多个尺度层级的同步分量,并通过跨尺度注意力机制使不同尺度的同步信息进行交互融合,构建出层级化同步表征向量,实现对微尺度情感传染、中尺度互动规范协调和宏尺度关系融合的分离与协同建模;与此同时,从互动感知特征序列中计算双方动作单元激活的协方差矩阵,基于协方差矩阵迹的统计量生成离散度指数并编码为离散度特征向量,用于量化面部表情激活模式的多样性;将层级化同步表征向量与离散度特征向量拼接后输入采用径向基函数激活层的非单调映射网络,借助径向基函数天然具备的峰值响应特性,使网络能够学习同步度与心理默契度之间的倒U型映射关系,自动区分高离散度中等同步的真实情感共鸣与低离散度高同步的表演性伪同步,输出心理默契度评估值;最终通过将互动过程划分为多个重叠窗口并获取心理默契度评估值时间序列及其一阶差分,生成抵触风险预判结果。由此,本方案通过渐进式融合、多尺度时域分解与非单调映射三个技术环节的有机协同,突破了单一时间尺度同步分析的信息混淆缺陷,从根本上克服了同步度越高默契度越高的错误假设,实现了对人际互动中真实情感共鸣的准确量化和对伪同步的有效甄别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799474A_ABST
    Figure CN122799474A_ABST
Patent Text Reader

Abstract

The application relates to a kind of interpersonal interaction emotion synchronization analysis method, system, equipment and medium, the facial feature of both sides is extracted by double-flow convolution network, and gate vector is adaptively generated according to information compression state in multiple levels to carry out progressive fusion, obtain interactive perception feature sequence;The cross-correlation calculation is carried out to interactive perception feature sequence and is decomposed in multiple scales by continuous wavelet transform, and hierarchical synchronization representation vector is constructed using cross-scale attention to separate mirror response of different time scales;The dispersion index of action unit covariance matrix trace is calculated from interactive perception feature sequence, and the dispersion feature vector is used to quantify expression diversity;After synchronization representation vector and dispersion feature vector are spliced, input radial basis function non-monotonic mapping network, output psychological tacit degree evaluation value by learning inverse U-shaped mapping relationship, distinguish real emotional resonance and pseudo synchronization;Finally, by multiple window psychological tacit degree evaluation value sequence, generate resistance risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of computer vision and artificial intelligence, and in particular relates to a method, system, device and medium for synchronous analysis of interpersonal interaction emotions. Background Technology

[0002] In the field of interpersonal interaction analysis, quantifying the degree of emotional synchronization between two parties through facial expressions has significant application value for assessing communication quality and predicting relationship trends. Existing emotional synchronization analysis methods typically begin by identifying action units in the facial videos of both parties to obtain their respective action unit intensity time series. Then, they use statistical measures such as Pearson correlation coefficient, mutual information, or dynamic time warping within a fixed window to calculate the degree of follow-up of facial expression changes. Alternatively, they feed the two-person sequences separately into a two-stream neural network for feature extraction, and then concatenate or weightedly fuse the two stream features to predict the level of synchronization or interaction quality.

[0003] However, the aforementioned techniques generally have significant shortcomings. First, existing methods mostly rely on a single time scale when analyzing synchronicity, such as using the same-sized analysis window for the entire interaction process. This leads to the mixing of mirror response signals at different time levels. Psychological research shows that mirror responses reflect instinctive emotional contagion, interaction norm coordination, and relational emotional fusion at the micro, meso, and macro scales, respectively. Single-scale analysis cannot separate these levels, causing high-frequency synchronization at the micro scale to mask emotional divergence at the macro level, and vice versa, severely damaging the hierarchical integrity and explanatory power of synchronization analysis. Second, existing methods almost invariably assume a monotonically increasing mapping relationship between synchronicity and psychological tacit understanding, meaning that higher synchronicity equates to higher tacit understanding. However, numerous social psychology experiments show that excessively high facial expression synchronization is often associated with performative imitation, conformist behavior, or socially normative actions, constituting "pseudo-synchronicity." Genuine emotional resonance, on the other hand, occurs in situations with moderate synchronicity and rich, varied facial expressions. This monotonous assumption makes existing technologies unable to identify pseudo-synchronicity, easily leading to serious misjudgments in scenarios such as business negotiations and psychological counseling. Furthermore, the fusion design of the two-stream network in processing the features of the two interacting parties is too crude. Whether shallow early fusion or high-level decision fusion is used, there is a problem of information loss: early fusion destroys the integrity of the individual's facial expression dynamics, while late fusion makes it difficult to capture the subtle temporal coupling information between the two parties. At the same time, there is a lack of effective mechanisms to suppress the extraction of redundant features by the two streams, which makes it impossible for the network to efficiently learn the interaction-specific complementary representations.

[0004] Due to the aforementioned shortcomings, existing technologies struggle to accurately and robustly quantify emotional synchronization in interpersonal interactions, limiting their effectiveness in real-world social interaction scenarios. Therefore, there is an urgent need for an emotional synchronization analysis method capable of separating multi-timescale mirror responses, overcoming the monotonic mapping assumption, and achieving efficient fusion of interactive features. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, system, device, and medium for synchronous analysis of interpersonal interaction emotions to address the aforementioned technical problems.

[0006] Firstly, this application provides a method for synchronous analysis of interpersonal interaction emotions, including: S1. Perform motion unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously to obtain the first facial motion unit intensity time series of the first participant and the second facial motion unit intensity time series of the second participant. S2. Input the first facial motion unit intensity time series and the second facial motion unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. Extract the first feature map and the second feature map of multiple levels layer by layer through the first processing stream and the second processing stream. Adaptively generate a gating vector according to the information compression state of each level. Use the gating vector to perform progressive cross-stream fusion of the first feature map and the second feature map to obtain the first interactive perception feature sequence and the second interactive perception feature sequence. S3. By performing cross-correlation calculation on the first interactive sensing feature sequence and the second interactive sensing feature sequence, a cross-correlation function is obtained; continuous wavelet transform is used to perform multi-scale time-domain decomposition on the cross-correlation function to obtain synchronization components at multiple scale levels; cross-scale attention mechanism is used to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector. S4. Based on the first interactive perception feature sequence and the second interactive perception feature sequence, calculate the covariance matrix of the activation of the action units of both parties; generate the dispersion index according to the statistics of the trace of the covariance matrix, and encode the sequence composed of the dispersion index to obtain the dispersion feature vector. S5. Input the concatenation result of the hierarchical synchronization representation vector and the discrete feature vector into a non-monotonic mapping network with radial basis function activation layer, and output the psychological compatibility evaluation value. S6. Divide the interaction process into multiple overlapping time windows, and execute S2 to S5 in each overlapping time window to obtain the time series of psychological tacit understanding evaluation values; generate the conflict risk prediction result based on the time series of psychological tacit understanding evaluation values ​​and its first difference.

[0007] Secondly, this application also provides an interpersonal interaction emotion synchronization analysis system for implementing the method described in the first aspect, the system comprising: The multimodal action unit recognition module is used to perform action unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously, and to obtain the first facial action unit intensity time series of the first participant and the second facial action unit intensity time series of the second participant. The dual-stream progressive fusion feature extraction module is used to input the first facial action unit intensity time series and the second facial action unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. The first and second feature maps of multiple levels are extracted layer by layer through the first and second processing streams. The gating vector is adaptively generated according to the information compression state of each level. The first feature map and the second feature map are progressively fused across streams using the gating vector to obtain the first interactive perception feature sequence and the second interactive perception feature sequence. The multi-scale synchronization feature parsing module is used to obtain a cross-correlation function by cross-correlation calculation of the first interactive sensing feature sequence and the second interactive sensing feature sequence; to perform multi-scale time-domain decomposition of the cross-correlation function by continuous wavelet transform to obtain synchronization components at multiple scale levels; and to use a cross-scale attention mechanism to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector. The synchronous discreteness encoding module is used to calculate the covariance matrix of the activation of the action units of both parties based on the first interactive perception feature sequence and the second interactive perception feature sequence; generate a discreteness index according to the statistics of the trace of the covariance matrix; and encode the sequence composed of the discreteness index to obtain the discreteness feature vector. The non-monotonic psychological mapping module is used to input the concatenation result of the hierarchical synchronous representation vector and the discrete feature vector into the non-monotonic mapping network with the radial basis function activation layer, and output the psychological compatibility evaluation value. The time-series dynamic assessment and risk prediction module is used to divide the interaction process into multiple overlapping time windows. Within each overlapping time window, the dual-stream progressive fusion feature extraction module, multi-scale synchronous feature parsing module, synchronous discreteness encoding module, and non-monotonic psychological mapping module are executed to obtain the time series of psychological tacit understanding evaluation values. Based on the time series of psychological tacit understanding evaluation values ​​and its first-order difference, the conflict risk prediction result is generated.

[0008] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a method for synchronous analysis of interpersonal interaction emotions as described in the first aspect.

[0009] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for synchronous analysis of interpersonal interaction emotions as described in the first aspect.

[0010] The aforementioned interpersonal interaction emotion synchronization analysis method, system, device, and medium start from the facial action unit intensity sequence of both parties. It extracts individual facial expression dynamic features layer by layer through a two-stream convolutional network, and progressively fuses the two-stream features by adaptively generating gating vectors based on information compression states at multiple depth levels. This process preserves the independent dynamics of individual micro-expressions while fully capturing the temporal information of the interaction coupling between the two parties, obtaining an interaction perception feature sequence. Based on this, cross-correlation calculations are performed on the interaction perception feature sequence to obtain a cross-correlation function. Continuous wavelet transform is used to decompose the cross-correlation function into synchronization components at multiple scale levels. A cross-scale attention mechanism is then used to interactively fuse synchronization information at different scales, constructing a hierarchical synchronization representation vector. This enables the analysis of micro-scale emotion contagion, meso-scale interaction coordination, and macro-scale interaction. The system integrates separation and collaborative modeling; simultaneously, it calculates the covariance matrix of the activation of action units of both parties from the interactive perception feature sequence, generates a dispersion index based on the statistics of the trace of the covariance matrix and encodes it as a dispersion feature vector to quantify the diversity of facial expression activation patterns; the hierarchical synchronization representation vector and the dispersion feature vector are concatenated and input into a non-monotonic mapping network with radial basis function activation layer. With the help of the peak response characteristics inherent in radial basis function, the network can learn the inverted U-shaped mapping relationship between synchronization degree and psychological tacit understanding, automatically distinguish between real emotional resonance with high dispersion and medium synchronization and performative pseudo-synchronization with low dispersion and high synchronization, and output psychological tacit understanding evaluation value; finally, by dividing the interaction process into multiple overlapping windows and obtaining the time series of psychological tacit understanding evaluation value and its first difference, a conflict risk prediction result is generated. Therefore, this solution overcomes the information confusion defect of single-time-scale synchronous analysis by organically coordinating three technical links: progressive fusion, multi-scale time-domain decomposition, and non-monotonic mapping. It fundamentally overcomes the erroneous assumption that the higher the degree of synchronization, the higher the degree of tacit understanding, and achieves accurate quantification of genuine emotional resonance in interpersonal interactions and effective identification of pseudo-synchronization. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a method for synchronizing interpersonal interaction emotions provided by the present invention; Figure 2 This is a schematic diagram of the process for calculating the discrete feature vector in one optional embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an interpersonal interaction emotion synchronization analysis system provided by the present invention. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0014] This invention provides a method for synchronizing interpersonal interaction emotions, which can be widely applied in scenarios such as psychological counseling, business negotiations, and team collaboration where an objective assessment of the emotional synchronization status of both parties is required. This method is typically deployed on one or more electronic devices with data acquisition, processing, and computing capabilities. In a typical application environment, the interpersonal interaction emotion synchronization analysis system includes two synchronously operating digital camera devices, a data preprocessing terminal, and a backend analysis server. The two camera devices are respectively aimed at the first and second participants, acquiring real-time facial video streams of both parties and transmitting the video data to the data preprocessing terminal or directly to the backend analysis server via wired or wireless networks. The data preprocessing terminal is responsible for performing face detection and action unit recognition on each frame of the image, generating a time series of facial action unit intensity for both parties. The backend analysis server is equipped with a dual-stream convolutional neural network model, a multi-scale decomposition module, a discreteness extraction module, a non-monotonic mapping network, and a conflict risk prediction model. It performs feature extraction, synchronization analysis, tacit understanding assessment, and risk prediction on the received action unit intensity time series, ultimately outputting a psychological tacit understanding assessment value and a conflict risk prediction result. The analysis results can be displayed on the client's graphical user interface for reference by consultants, negotiation experts, or researchers.

[0015] Example 1: refer to Figure 1 The document presents a flowchart illustrating a method for synchronizing interpersonal interaction emotions, as provided in this application. This method includes the following steps: S1. Perform motion unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously to obtain the time series of the first facial motion unit intensity of the first participant and the time series of the second facial motion unit intensity of the second participant.

[0016] Specifically, two synchronously triggered color digital cameras are used to capture facial video streams of the first and second participants, respectively. The frame rates of the two cameras are kept consistent, and precise timestamp alignment is achieved through hardware synchronization signals or network time protocols. For each captured frame, a face detection algorithm based on a deep convolutional neural network is first used to locate the face bounding box. This detection algorithm outputs the bounding box coordinates and confidence score for each face. If multiple faces are detected in a frame, the bounding box with the largest area is selected as the target face. Subsequently, a regression-based facial landmark detection algorithm is used to locate the coordinates of dozens of feature points such as eyebrows, eyes, nose, and mouth. Affine transformations are performed based on the landmark positions to achieve face alignment, and the face region is cropped into image patches of fixed resolution. The cropped face image patches are then input into a pre-trained action unit intensity estimation model. This action unit intensity estimation model is based on a deep convolutional neural network architecture and is trained in a supervised manner using a large number of facial images with labeled AU (Facial Action Unit Intensity). It can output the activation intensity value of each AU in a preset action unit set for each input face image patch, with the value typically normalized to between 0 and 1. The selected set of action units covers core AUs closely related to emotional expression, such as AU1 (inner eyebrow raised), AU2 (outer eyebrow raised), AU4 (eyebrow lowered), AU5 (upper eyelid raised), AU6 (cheek lifted), AU7 (eyelid tightened), AU9 (nose wrinkled), AU10 (upper lip raised), AU12 (corner of mouth pulled), AU14 (dimples appear), AU15 (corner of mouth lowered), AU17 (chin lifted), AU20 (lip stretched), AU23 (lip tightened), AU25 (lips parted), AU26 (chin lowered), etc., the total number of which is denoted as [missing information]. .

[0017] For the first participant, the AU intensity vectors of all frames are arranged in chronological order to form the time series of the first facial motion unit intensity, denoted as a matrix. ,in Let be the total number of video frames, and be the th element of the matrix. Line 1 Column elements Indicates the first participant in the The first frame The intensity value of each action unit, For frame index, This is the AU index. Similarly, for the second participant, the intensity time series of the second facial action unit is obtained, denoted as a matrix. Because facial muscle movement amplitude and expression baseline vary among individuals, Z-score normalization needs to be performed individually on the time series of each AU channel. For the first participant... Calculate the mean of the time series of each AU. and standard deviation Then perform standardization: ;in, The standardized strength value. The original strength value before standardization. For the first participant The average intensity of each AU across all frames. This represents the corresponding standard deviation. The same operation is performed on the corresponding AU channel for the second participant. For missing frames where the action unit detection model fails to output valid values ​​due to drastic changes in head posture or occlusion, linear interpolation results from adjacent valid frames before and after the missing frame can be used to fill in the gaps, ensuring the continuity and integrity of the time series.

[0018] S2. Input the first facial motion unit intensity time series and the second facial motion unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. Extract the first feature map and the second feature map of multiple levels layer by layer through the first processing stream and the second processing stream. Adaptively generate the gating vector according to the information compression state of each level. Use the gating vector to perform progressive cross-stream fusion of the first feature map and the second feature map to obtain the first interactive perception feature sequence and the second interactive perception feature sequence.

[0019] Specifically, the two-stream convolutional neural network comprises two structurally symmetrical one-dimensional convolutional processing streams, used to process the time series of intensity of the first facial action unit and the second facial action unit, respectively. Taking the first processing stream as an example, its internal structure is described; the second processing stream is completely symmetrical to it. The input to the first processing stream is a normalized matrix. First, an initial one-dimensional convolutional layer is used, with a kernel size of 7 and a number of kernels of [value missing]. With a step size of 1, for the input sequence The AU channels are blended and feature mapped to obtain a channel number of . The initial feature map is then generated. Subsequently, the network passes through three residual blocks, denoted as the first, second, and third residual blocks. Each residual block consists of two concatenated one-dimensional convolutional layers. The first convolutional layer has a kernel size of 5, and the second convolutional layer has a kernel size of 3. Both convolutional layers have the same number of output channels. subscript Indicates the sequence number of the residual block. This corresponds to three residual blocks. Within each residual block, the input features first pass through a first convolutional layer, a batch normalization layer, and a ReLU activation function, then through a second convolutional layer and a batch normalization layer. Afterward, the output is residually concatenated with the input (i.e., element-wise added), and the output of that residual block is obtained by passing the ReLU activation function. The expression for the ReLU activation function is: That is, when Output when greater than 0 Otherwise, output 0. After each residual block, apply a one-dimensional max pooling operation with a stride of 2, downsampling along the time dimension to halve the time length. The number of output channels for the three residual blocks are respectively , , The network is incremented block by block to accommodate more abstract, high-level semantic information. The second processing stream processes the intensity time series of the second facial action unit using the same network structure. The network parameters of the two streams can be shared or independent. In this embodiment, the parameter-independent approach is adopted to capture individual differences more flexibly.

[0020] Progressive cross-stream fusion is performed between two processing streams at selected multiple levels. In this embodiment, the three fusion levels—after the first residual block, after the second residual block, and after the third residual block—are selected as level 1, level 2, and level 3, respectively. At each fusion level, the three-dimensional tensor currently output by the first processing stream is denoted as the first feature map. The three-dimensional tensor currently output by the second processing stream is denoted as the second feature map. ,in The time dimension length of this level. For hierarchical indexes, This refers to the number of channels at this level. The specific implementation of the fusion operation is as follows: First, on the channel dimension, [the following is done]... and By splicing, a product with Joint feature tensor of each channel The joint feature tensor is input into a small gated generation network, which consists of two fully connected layers. The first fully connected layer reduces the number of channels from... Mapped to The second fully connected layer maintains the number of channels at a certain value, followed by the ReLU activation function. This is followed by a Sigmoid activation function, which outputs the first gating vector for the first processing stream. Each element has a value between 0 and 1. The formula for this generation process is:

[0021] In the formula, This is the weight matrix of the first fully connected layer, used to combine the joint feature tensor. The number of channels from Transform to an intermediate dimension; This is the bias vector of the first fully connected layer; It is the ReLU activation function, which performs activation on each element of the input tensor. operate; This is the weight matrix of the second fully connected layer, used to transform the intermediate features into the final gate vector; This is the bias vector for the second fully connected layer; The Sigmoid activation function has the following mathematical form: Map any real number to interval; This is the first gating vector generated, where each element represents the proportion of information retained by the first processing stream at the corresponding position. Simultaneously, another set of completely independent gating generation networks is used to process the same joint feature tensor. Processing is performed to generate a second gating vector for the second processing stream. Calculation method and Completely symmetric, meaning it is obtained through two fully connected layers and corresponding activation functions, but the weights and bias parameters are independent.

[0022] To achieve controlled mixing of cross-stream information, it is also necessary to map the feature maps of the other stream to the feature space of the current stream. This involves introducing a learnable first linear transformation matrix. Perform linear projection on the second feature map: In the formula, This is the second feature map after mapping, and its time dimension is still [missing information]. The channel dimension is Matrix multiplication is performed along the channel dimension; that is, for each time step, the second feature map is multiplied in that step. Row vectors multiplied by the transformation matrix Subsequently, using the first gating vector The first fused feature map is obtained by performing an element-wise weighted summation of the first feature map and the mapped second feature map: In the formula, This represents element-wise multiplication, which means multiplying two tensors with the same dimension at corresponding positions; To and All dimensions are exactly the same tensor; This is the first feature map after fusion.

[0023] Symmetrically, a second linear transformation matrix is ​​introduced. The first feature map after mapping is obtained. and using the second gating vector The second fused feature map is obtained: Feature map fusion and The original first and second feature maps are replaced respectively, and then fed into the subsequent network layers of their respective processing streams. This operation is performed at all three fusion levels, allowing cross-stream information to gradually permeate and interact across multiple abstraction levels. After traversing all network layers, the first and second processing streams output the first interactive perception feature sequence, respectively. Second interactive sensory feature sequence ,in For the final time dimension at the network end, This represents the final number of output feature channels.

[0024] The gating mechanism enables the network to adaptively adjust the degree of cross-stream fusion: by generating learnable parameters in the network through gating, the network automatically learns during task-driven training what proportion of information from the other stream should be incorporated at each level and feature dimension. When the features at a certain level are already sufficiently rich, the gating value tends to be smaller to protect the information from being diluted; when a certain level needs information from the other stream to enhance interactive perception, the gating value tends to be larger to promote information exchange. The gating capability in this embodiment relies entirely on end-to-end backpropagation learning.

[0025] S3. By performing cross-correlation calculation on the first interactive sensing feature sequence and the second interactive sensing feature sequence, a cross-correlation function is obtained; continuous wavelet transform is used to perform multi-scale time-domain decomposition on the cross-correlation function to obtain synchronization components at multiple scale levels; cross-scale attention mechanism is used to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector.

[0026] Specifically, firstly, the first interactive perception feature sequence... Second interactive sensory feature sequence Cross-correlation calculations were performed to quantify the following and being followed relationship in the facial expressions of both parties. To reduce computational complexity and focus on the main directions of variation, principal component analysis was performed on both feature sequences along their respective feature dimensions. The eigenvalues ​​and eigenvectors of the covariance matrix are selected, and the top eigenvalues ​​and eigenvectors whose cumulative variance contribution rate reaches a preset proportion (e.g., 95%) are chosen. The principal components and their corresponding eigenvectors form the projection matrix. The first feature sequence after dimensionality reduction is: The second feature sequence is The dimensionality reduction operation uniformly uses the projection matrix calculated from the first feature sequence to ensure that the features of both sides are compared in the same low-dimensional space. This indicates the number of principal components retained.

[0027] For each principal component dimension ( The cross-correlation function is defined as the cross-correlation function at different time delays. Below, the mean of the pointwise product of the features of the first feature sequence and the delayed features of the second feature sequence. Delay Take integer values; positive values ​​indicate that the second participant lags behind the first participant, and negative values ​​indicate that the first participant lags behind. This cross-correlation function... The mathematical expression is:

[0028] in, express In the The first time step Principal component values, express At the corresponding time step Principal component values, The total time length of the sequence after dimensionality reduction. For integer delays, the summation range is limited. To ensure that the index does not go out of bounds, and Representing the given delay The lower and upper bounds of the effective summation. Finally, for all The average of the cross-correlation functions of each principal component dimension is used to obtain the comprehensive cross-correlation function: . It is a one-dimensional discrete function, reflecting the overall synchronization strength of the facial expression features of both parties under different time delays. When exist When a high value is obtained nearby, it indicates that the changes in the expressions of both parties tend to be synchronized; when the peak appears at a non-zero delay, it indicates that there is a stable leader-follower relationship.

[0029] Then, continuous wavelet transform is used to analyze the comprehensive cross-correlation function. Multi-scale time-domain decomposition is performed to separate the mirror response signals at different time scale levels. A complex Morlet wavelet is selected as the mother wavelet function, which consists of the product of a complex sine wave and a Gaussian envelope, and its expression is: .in, For continuous time variables, The imaginary unit satisfies , The center angular frequency, The normalization constant is used to make the wavelet energy 1. The commonly used value for is 6, which achieves a good balance between the time-domain and frequency-domain localization characteristics of the wavelet. (Scale factor) The range of values ​​is usually within the smallest scale. To the maximum scale Sampling is performed at logarithmically uniform intervals, and the total number of samples is denoted as 1. Each scale is denoted as... Small-scale factors correspond to high-frequency, short-term windows, while large-scale factors correspond to low-frequency, long-term windows. Translation factor. All integer delay positions can be taken.

[0030] At every scale and each translation position Above, calculate the wavelet transform coefficients: In the formula, for The complex conjugate function, This indicates that the time delay variable According to translation and scale Perform scaling and translation. The energy normalization factor is used to sum over all possible discrete time delays. Wavelet transform coefficients It is a complex number whose modulus squared Constructing a time-scale energy spectrum, representing the energy at the translational position. and scale Synchronous energy density.

[0031] Based on psychological research on the time constants of emotional contagion, interaction norm coordination, and relational emotional integration, two scales are set to divide the threshold, including all... The scale factor is divided into three continuous scale bands: microscale band. Includes all scales with scale factors smaller than the first dividing threshold, corresponding to fast, transient mimicry behavior on the order of seconds; mesoscale band. Includes scales with scaling factors between the first and second segmentation thresholds, corresponding to facial expression matching and interaction coordination lasting from several seconds to tens of seconds; macro-scale frequency band. This includes all scales with scale factors greater than the second threshold, corresponding to emotional trend coordination over periods exceeding ten seconds, and even the entire interaction process. For each scale band, three statistical features are extracted to characterize the synchronization properties of that band: the average energy, energy standard deviation, and energy skewness of the energy spectrum across all scales within that band along the translation axis. Taking the microscale band as an example, let the translation factor be... The total number of possible values ​​is The calculation is as follows: , , In the formula, Indicates the number of scales within the microscale frequency band; and All One translation position; For scale The average energy below; The average energy at the microscale measures the overall intensity of synchronization at that scale. The standard deviation of energy at the microscale measures the degree of fluctuation in synchrotron energy over time. The microscale energy skewness measures the asymmetry in energy distribution. These three statistics together constitute the microscale synchronization component vector. Similarly, performing the same calculations on the mesoscale and macroscale frequency bands yields the mesoscale synchronization component. and macro-scale synchronization components .

[0032] Subsequently, a cross-scale attention mechanism is used to interactively fuse the three synchronization components to construct a hierarchical synchronization representation vector. This embodiment employs a concise self-attention implementation scheme. First, the three synchronization components are concatenated along the last dimension into a joint vector. The joint vector is mapped to a higher dimension through a fully connected layer, and then the query vector is generated through three independent linear projection matrices. Key vector Sum value vector Then calculate the scaled dot product attention: calculate and The dot product of the values ​​is scaled by the square root of the key vector dimension, normalized using the softmax function to obtain the attention weight matrix, and then multiplied by the value vector to obtain the weighted output. This weighted output is then mapped back through a fully connected layer. The dimensional space is finally divided into three updated components according to the dimension. , and The vectors are then concatenated again to form a hierarchical synchronous representation vector. This simplified self-attention implementation allows the components at different scales to reference and calibrate each other: if one scale exhibits abnormally high synchronization while another scale deviates, the attention mechanism can suppress the contribution of the abnormal scale in the vector space, thereby enhancing the robustness and semantic consistency of the representation.

[0033] S4. Based on the first interactive perception feature sequence and the second interactive perception feature sequence, calculate the covariance matrix of the activation of the action units of both parties; generate the dispersion index according to the statistics of the trace of the covariance matrix, and encode the sequence composed of the dispersion index to obtain the dispersion feature vector.

[0034] Specifically, this step utilizes action unit information contained in the interactive perception feature sequence to quantify the diversity of facial expression activation. The feature map after the first residual block and before the first fusion in the two-stream convolutional neural network is selected as the base feature because this layer's features have not undergone extensive cross-stream fusion, thus better preserving the independent response information of their respective original AU activations. Let the feature map corresponding to the first participant in the first interactive perception feature sequence of this layer be... The feature map corresponding to the second participant is ,in For the time dimension of this layer, Let be the number of feature channels. A learnable linear mapping matrix is ​​introduced. Convert the feature channel dimension into the total number of action units. This allows the vector at each time step after mapping to be interpreted as the reconstructed activation intensity values ​​of each AU: , .in, This is the mapping sequence for the first action unit. The second action unit mapping sequence is then mapped. Subsequently, along the time dimension with a fixed window length... and fixed sliding step size ( The time frame is divided into a series of overlapping sliding time windows. Each sliding time window is indexed by the starting frame. The icon, contained within the window A continuous time step.

[0035] For windows Collect the first action unit mapping sequence within this window indivual The first sample set is composed of dimensional vectors. Each of them The first participant's action unit activation covariance matrix within this window. The matrix is ​​calculated from the sample set using an unbiased estimator. The Line 1 The column elements are:

[0036] in, For the first The th sample vector AU components For all within this window The sample at the th The mean over each AU, i.e. , and For AU indexes, the value range is 1 to The diagonal elements of this matrix Indicates the first The activation variance of each AU within this window, with off-diagonal elements representing the activation covariance between different AUs. Similarly, the activation covariance matrix of the action units of the second participant is calculated based on the second sample set. :

[0037] in, For the second participant The th sample vector AU components This corresponds to the mean.

[0038] After obtaining the covariance matrix of both parties, calculate the window. Action unit activation discreteness index :

[0039] In the formula, The trace operation represents the sum of all diagonal elements of a matrix; for traces ; This represents the element-wise addition of two matrices; the denominator contains... is the normalization factor, where The total number of action units is such that the dispersion index eliminates the influence of the total number of AUs. The higher the value, the more diverse the AU combinations used by both parties within this window and the greater the variation in activation magnitude; The lower the value, the more monotonous the expressions of both parties become, relying solely on the mechanical repetition of a few AUs.

[0040] Arrange the dispersion indices of all sliding time windows in chronological order of their center times to form a dispersion index sequence. ,in This represents the total number of windows. To obtain a fixed-length feature representation for subsequent network processing, a one-dimensional convolutional encoder is used to encode the sequence. This encoder consists of two one-dimensional convolutional layers. The first convolutional layer maps the discrete sequence of a single channel to... There are feature channels, and the kernel size is . The first convolution has a stride of 1, followed by a ReLU activation function; the second convolution maintains the same number of channels, and the kernel size is also the same. This is followed by a global average pooling layer, which compresses the entire time dimension into a single statistical value, outputting a discrete feature vector. ,in This represents the output dimension of the encoder.

[0041] S5. Input the concatenation result of the hierarchical synchronization representation vector and the discrete feature vector into a non-monotonic mapping network with radial basis function activation layer, and output the psychological compatibility evaluation value.

[0042] Specifically, the hierarchical synchronization representation vector obtained in step S3 is... The discrete eigenvector obtained in step S4 By concatenating along the dimensional direction, a comprehensive feature vector is formed. The non-monotonic mapping network consists of three parts: a fully connected hidden layer, a radial basis function activation layer, and a linear output layer.

[0043] Comprehensive feature vector First, dimensionality blending and non-linear feature extraction are performed using fully connected hidden layers. The fully connected layers contain... There are n hidden units, and the weight matrix is ​​as follows: The bias vector is The calculation is as follows: .in, To hide the representation vector, Weight matrix transpose; It is a linear rectified function, taking element-wise... ; This represents the number of hidden units.

[0044] Following the hidden layer is a radial basis function activation layer, which contains... The nth radial basis function unit. In this embodiment, each radial basis function unit adopts an isotropic Gaussian kernel form. For the nth... Units ( Let its learnable center vector be... The learnable scaling parameter is The activation response of this unit is represented by the hidden representation vector. With the center vector The square of the Euclidean distance, multiplied by the scaling parameter and then taken as an exponential function, yields: .in, The square of the Euclidean distance. and They are respectively and The One portion, For dimension indexing; For the first The scaling parameter of each unit controls the size of the unit's response range. and When they completely overlap, The greater the distance between the two, the closer the response value becomes. .

[0045] all The activation responses of each radial basis function unit are linearly combined using learnable weights, and after adding a bias term, are input into the sigmoid activation function to obtain the final psychological compatibility assessment value. :

[0046] In the formula, For the first The weight coefficient of each unit can be positive or negative; For bias terms; The sigmoid function has the following mathematical form: Map any real number to interval; This represents the total number of radial basis function units. All learnable parameters of a non-monotonic mapping network include the weight matrices of the fully connected layers. and bias The center vector of each radial basis element and scaling parameters and the weights of the output layer and bias .

[0047] The non-monotonic mapping capability of this network structure stems from the local response characteristics of the radial basis function units. Each unit produces a high response only in a local region near its center vector. Through data-driven training, different radial basis units move to different typical regions of the synchronous-discrete joint input space. For example, the center of some units learns to locate in a region of moderate synchronousity and high dispersion, and its output weights... Some units are learned to have larger positive values, causing the network to output high synchronization values ​​in that region; the centers of other units will move to pseudo-synchronization regions with high synchronization and low dispersion, and their weights will be... The negative value is learned, causing the network's output in that region to be suppressed to a low value. Thus, the entire network automatically learns an inverted U-shaped mapping surface that conforms to psychological principles without any explicit functional constraints, fundamentally overcoming the technical bias in traditional methods that require a monotonically increasing relationship between synchronization and tacit understanding.

[0048] S6. Divide the interaction process into multiple overlapping time windows, and execute S2 to S5 in each overlapping time window to obtain the time series of psychological tacit understanding evaluation values; generate the conflict risk prediction result based on the time series of psychological tacit understanding evaluation values ​​and its first difference.

[0049] Specifically, with a fixed window length Frames and fixed overlap step frame( The entire interactive video is sequentially segmented along the timeline, starting from the first frame, and divided into sections. A number of overlapping time windows. The frame index range corresponding to each window is: , For each time window, extract the data segment corresponding to the frame range from the complete AU intensity time series obtained in step S1 to form the first facial motion unit intensity time series segment and the second facial motion unit intensity time series segment for that window. Using this segment as input, repeat all operations from steps S2 to S5 sequentially to obtain the psychological compatibility assessment value corresponding to that window. .

[0050] All The evaluation values ​​of each window are arranged in chronological order to form a time series of psychological compatibility evaluation values. This sequence reflects the dynamic evolution of the emotional synchronization quality between the two parties over time during the interaction. To capture the direction and rate of change in the level of understanding, the first-order difference sequence of this sequence is calculated. ,in: , .in, Positive values ​​indicate an increase in rapport compared to the previous window, while negative values ​​indicate a decrease. The psychological rapport assessment value for each time window is... Its first difference (The first window only uses) The features are concatenated (with zero-padding) to form a two-dimensional feature vector, which is then input into a pre-trained binary classification conflict risk prediction model. This model can employ logistic regression or a shallow fully connected neural network, outputting the conflict risk probability corresponding to that window. The overall risk of conflict in the interaction can be obtained by taking the maximum or average of the risk probabilities of all windows. When the risk probability exceeds a preset warning level, the system can issue a prompt to the user.

[0051] In addition, all learnable parameters of the aforementioned dual-stream convolutional neural network, progressive fusion gating network, cross-scale attention module, discrete encoder, non-monotonic mapping network, and conflict risk prediction model need to be obtained through end-to-end supervised training and joint optimization.

[0052] When constructing the training dataset, each sample contains a video of two people interacting and its corresponding real psychological compatibility rating label. This label is generated by multiple trained psychology professionals independently rating the video, and then the average value is taken, with the values ​​normalized to a certain level. The training samples can also be labeled with pseudo-synchronization category labels and conflict risk labels to assist supervision. To improve the model's generalization ability and robustness, random data augmentation operations are applied to the original AU time series of each training sample, including random time pruning (cutting off a portion of the original sequence), random time dilation (stretching or compressing the time axis by a small scale), and adding small Gaussian noise to the AU intensity values, generating two different augmented views. These two views should have similar high-level semantic representations in the feature space.

[0053] The total loss function for end-to-end training is composed of the weighted sum of the following four components.

[0054] 1) First, psychological understanding predicts loss. This loss is used to measure the psychological compatibility assessment value of the output of a non-monotonic mapping network. With real labels The difference between them. The mean squared error loss function is used: .in, This represents the number of samples in each training iteration. This is the index of the samples in the batch. The first in the batch The predicted value for each sample, It corresponds to the real label.

[0055] 2) Secondly, there is the loss of contrast between scales. The purpose of this loss term is to maintain semantic consistency in the embedding space for the synchronization representations of the same interaction pair at different time scales, while distinguishing the representations of different interaction pairs. During training, the two augmented views of each training sample are processed separately to complete the multi-scale decomposition in step S3, obtaining their respective micro-scale synchronization components. Mesoscale synchronization components and macro-scale synchronization components A separate fully connected projection head is set up for each scale, mapping the synchronization components to low-dimensional projection vectors with the unit norm. For example, a microscale projection head... Map the microscale synchronization components as ,satisfy ,in Let be the dimension of the projection vector. The same logic applies to mesoscale and macroscale, yielding the following results respectively. and .

[0056] Calculate the inter-scale contrast loss for any two distinct scales. Consider the contrast loss between the micro-scale and meso-scale. For example, In the formula, The first in the batch The microscale projection vector of the first augmented view of each sample; The medium-scale projection vector of the second augmented view of the same sample is used as the positive sample; The first in the batch The midscale projection vector of the second augmented view of each sample. , of which only When a sample is positive, the rest are negative samples; The temperature coefficient controls the sharpness of the similarity distribution. The numerator brings together the representations of the same interaction pair across scales, while the denominator pushes away the representations of different interaction pairs. Similarly, the contrast loss between micro-macro and meso-macro pairs is calculated, and the sum of the losses for all scale pair combinations constitutes the total inter-scale contrast loss. .

[0057] 3) Secondly, there is the loss in suppressing cross-stream information redundancy. This loss term forces the two processing streams to extract complementary rather than redundant interaction information. Before the fusion operation at each fusion level is performed, the first feature map from the current level is used... Second feature map Feature vector pairs are sampled at corresponding positions along the time dimension. Let there be n feature vector pairs obtained from sampling. These samples come from the joint distribution Where is the sampling index. By randomly shuffling... The pairing relationships are used to construct samples of the marginal distribution product. ,in for A random permutation. Introduce a discriminator network. It takes a pair of feature vectors as input and outputs a scalar value. These are the learnable parameters of the discriminator. The discriminator consists of two fully connected layers, with ReLU activation in the middle, and the last layer outputs a scalar. Mutual information is estimated using the Donsker-Varadhan lower bound. In the formula, the first term is the discriminant output mean of the joint distribution samples, and the second term is the logarithmic mean of the discriminant output of the marginal distribution samples after taking the exponent.

[0058] During training, the discriminator parameters This estimate is maximized through gradient ascent to approximate the true mutual information, while the backbone parameters of the two-stream network are reduced for redundancy by minimizing this estimate. The cross-stream information redundancy suppression loss is obtained by summing the mutual information estimates from all fusion layers: In the formula, and The first The first and second feature maps before fusion at each fusion level It traverses all fusion layers. This adversarial training forces each processing stream to focus on unique interactive information that the other cannot provide.

[0059] 4) Finally, there is the loss due to information bottlenecks. At each fusion layer, information bottleneck regularization is applied to guide the network to adaptively adjust the degree of information compression. For the first... There are fusion layers, and their input features are . Introduce an encoder network The output of this layer represents random variables. The posterior distribution of is assumed to be a diagonal Gaussian distribution, i.e. , where the mean Sum of logarithmic variance From the encoder network The prior distribution is calculated from the given information. Set as a standard multivariate Gaussian distribution The information bottleneck loss at this level is the Kullback-Leibler divergence between the posterior and prior distributions: .in, for Dimensions For dimensional indexing, and They are respectively and The One component; This represents the KL divergence operation. The KL divergence term penalizes the deviation from the prior, forcing the intermediate representation to be as compact as possible, while the posterior variance adaptively learns which dimensions need to retain more information. The total information bottleneck loss is obtained by summing the information bottleneck losses at each level. .

[0060] The total loss function is a weighted combination of the four losses mentioned above: .in, , , , The weights for each loss term are positive real numbers. In each training iteration, a batch of samples is randomly sampled from the training set. Two augmented views are generated for each sample. Forward propagation is used to calculate the loss for each term, and then backpropagation is used to calculate the remaining loss. For the gradients of all learnable parameters, a gradient descent optimization algorithm based on an adaptive learning rate is used to update the parameters. For the discriminator network in the cross-stream redundancy suppression loss, several additional discriminator parameter updates are performed in each iteration to more accurately estimate mutual information. Training continues until the total loss no longer decreases on the validation set. After training, the parameters of the two-stream convolutional neural network, progressive fusion gating, cross-scale attention module, discrete encoder, non-monotonic mapping network, and conflict risk prediction model are fixed, making it suitable for sentiment synchronization analysis of new interactive videos.

[0061] This embodiment constructs an end-to-end analysis chain through the aforementioned complete steps and training methods, encompassing facial motion unit extraction, interactive perception feature learning, multi-scale synchronous decomposition and fusion, discreteness evaluation, non-monotonic mapping, and conflict risk prediction. This scheme utilizes cross-scale attention and progressive fusion mechanisms to separate mirror response signals at different time levels, avoiding information confusion caused by single-scale analysis. By introducing a discreteness index and radial basis function network, it overcomes the erroneous assumption that synchronization and tacit understanding must monotonically increase, fundamentally achieving an effective distinction between performative pseudo-synchronization and genuine emotional resonance. This has significant application value in high-risk social scenarios such as psychological counseling and business negotiations.

[0062] The aforementioned method for analyzing interpersonal interaction emotion synchronization starts with the intensity sequence of facial action units from both parties. It extracts individual facial expression dynamic features layer by layer using a two-stream convolutional network, and progressively fuses the two-stream features across multiple depth levels by adaptively generating gating vectors based on information compression states. This approach preserves the independent dynamics of individual micro-expressions while fully capturing the temporal information of the interaction coupling between the two parties, obtaining an interaction-perceived feature sequence. Based on this, cross-correlation calculations are performed on the interaction-perceived feature sequence to obtain a cross-correlation function. Continuous wavelet transform is used to decompose the cross-correlation function into synchronization components at multiple scale levels. A cross-scale attention mechanism is then used to interactively fuse synchronization information at different scales, constructing a hierarchical synchronization representation vector. This achieves the fusion of micro-scale emotion contagion, meso-scale interaction normative coordination, and macro-scale relationship fusion. Separation and collaborative modeling are employed. Simultaneously, the covariance matrix of the activation of action units of both parties is calculated from the interactive perception feature sequence. Based on the statistics of the trace of the covariance matrix, a dispersion index is generated and encoded as a dispersion feature vector to quantify the diversity of facial expression activation patterns. The hierarchical synchronization representation vector and the dispersion feature vector are concatenated and input into a non-monotonic mapping network with a radial basis function activation layer. By leveraging the peak response characteristics inherent in the radial basis function, the network can learn the inverted U-shaped mapping relationship between synchronization and psychological tacit understanding, automatically distinguishing between genuine emotional resonance with high dispersion and moderate synchronization and performative pseudo-synchronization with low dispersion and high synchronization, and outputting a psychological tacit understanding evaluation value. Finally, by dividing the interaction process into multiple overlapping windows and obtaining the time series of psychological tacit understanding evaluation values ​​and their first-order differences, a conflict risk prediction result is generated. Therefore, this solution overcomes the information confusion defect of single-time-scale synchronous analysis by organically coordinating three technical links: progressive fusion, multi-scale time-domain decomposition, and non-monotonic mapping. It fundamentally overcomes the erroneous assumption that the higher the degree of synchronization, the higher the degree of tacit understanding, and achieves accurate quantification of genuine emotional resonance in interpersonal interactions and effective identification of pseudo-synchronization.

[0063] Example 2: Based on the information compression state of each level, an adaptive gating vector is generated. The gating vector is then used to progressively fuse the first and second feature maps across streams to obtain a first interactive perception feature sequence and a second interactive perception feature sequence. This process includes the following steps: S11. At the current level, the first feature map currently output by the first processing stream and the second feature map currently output by the second processing stream are concatenated along the channel dimension to obtain a concatenated feature map.

[0064] Specifically, let the current position be the [number]th [unit]. The fusion layer, the first processing stream, outputs the first feature map as follows: The second feature map output by this layer of the second processing stream is ,in This represents the length of the time dimension at this level. This represents the number of feature channels at this level. This is a hierarchical index. The two are concatenated along the channel dimension to obtain a concatenated feature map. , before Each channel originates from the first feature map, and then... Each channel originates from the second feature map. The splicing operation enables the gating network to simultaneously observe the feature states of both sides at the current level, thereby making a more reasonable decision on the fusion ratio.

[0065] S12. Perform one-dimensional convolution, ReLU activation, one-dimensional convolution and Sigmoid activation on the spliced ​​feature map in sequence to generate the first gating vector for the first processing stream.

[0066] Specifically, a gated generative network consisting of two layers of one-dimensional convolutions is constructed. The first layer of one-dimensional convolutions uses... There are n convolutional kernels, each with a size of 0 in the time dimension. ( Can be taken as ), step size is The convolutional layer uses boundary padding to maintain the time dimension invariant. This convolutional layer then stitches together the feature maps. The number of channels from Compress to The output is denoted as Subsequently, on Applying the ReLU activation function element by element, we obtain ,in The second layer of one-dimensional convolution also uses... There are 1 convolutional kernel, and the kernel size is also 1. Keep the number of channels as The output remains unchanged, and is then passed through the Sigmoid activation function to generate the first gating vector. Each element of the first gating vector has a value between 0 and 1. The entire process can be represented as: In the formula, and These represent the first and second layer one-dimensional convolution operations for the first processing stream gated network, respectively. The ReLU activation function is applied to each element of the feature map. ; The Sigmoid activation function is expressed as follows: ; Each element Indicates the first The time step, the first The proportion of information retained by the first processing stream on each feature channel.

[0067] S13. Perform matrix multiplication between the second feature map and the first linear transformation matrix to obtain the mapped second feature map.

[0068] Specifically, the feature space of the second processing stream may differ in distribution from that of the first processing stream, and direct weighted summation would introduce bias. Therefore, a first linear transformation matrix is ​​introduced. , the second feature map Mapped to the feature space of the first processing stream: In the formula, matrix multiplication is performed along the channel dimension, that is, for each time step, The time step vector (dimension: Right multiplication This yields the mapped vector. This is the second feature map after mapping. The linear transformation matrix... These are learnable parameters that automatically find the optimal cross-stream alignment method through training.

[0069] S14. Using the first gating vector, perform element-wise weighted summation on the first feature map and the mapped second feature map to obtain the first fused feature map.

[0070] Specifically, the first fused feature map It is obtained by weighted summation of the self-retained portion of the first feature map and the cross-flow supplement portion of the mapped second feature map: In the formula, This represents element-wise multiplication; To and All dimensions are exactly the same tensor; For the first The first feature map after layer fusion. For any location... If the gate value near If the first fused feature map retains almost completely the value of the original first feature map at that position; if the gate value is close to The value of the second feature map after mapping is used almost entirely. This soft gating mechanism enables smooth, differentiable adaptive information fusion.

[0071] S15. Perform one-dimensional convolution, ReLU activation, one-dimensional convolution and Sigmoid activation on the spliced ​​feature map in sequence to generate a second gating vector for the second processing stream.

[0072] Specifically, symmetrical to step S12, but using a completely independent set of convolutional layers. and Generate the second gating vector: In the formula, and These represent the first and second layer one-dimensional convolution operations of the gating generative network for the second processing stream, respectively, with parameters independent of the gating network for the first processing stream. The independence of the parameters of the two gating networks allows the two processing streams to flexibly determine the fusion ratio according to their respective needs, enabling the capture of asymmetric leader-follower relationships in interactions.

[0073] S16. Perform matrix multiplication between the first feature map and the second linear transformation matrix to obtain the mapped first feature map.

[0074] Specifically, a second linear transformation matrix is ​​introduced. Map the first feature map to the feature space of the second processing stream: .in, This is the first feature map after mapping.

[0075] S17. Using the second gating vector, perform element-wise weighted summation on the second feature map and the mapped first feature map to obtain the second fused feature map.

[0076] Specifically, the fusion formula for the second fusion feature map is: In the formula, This is the second fusion feature map.

[0077] S18. Input the first fused feature map and the second fused feature map into the next network layer respectively. After traversing multiple layers, the first interactive perception feature sequence and the second interactive perception feature sequence are obtained.

[0078] Specifically, the operations described in S11 to S17 are performed at all preset fusion layers (after the first, second, and third residual blocks). The fused feature maps continue to be passed forward along their respective processing flows, through subsequent residual blocks and downsampling layers, until the end of the network, ultimately generating the first interactive sensing feature sequence and the second interactive sensing feature sequence, respectively.

[0079] Compared to Embodiment 1, this embodiment replaces the fully connected gated network with a one-dimensional convolutional gated network, enabling the generation of gated vectors to utilize contextual information from adjacent time steps. This allows for a more accurate determination of the appropriate degree of cross-stream fusion at the current moment. The sliding of the convolutional kernel along the temporal dimension provides the gated decision with temporal smoothness and local awareness, avoiding the drastic fluctuations in the fusion ratio that might result from independent decisions at each time step. Furthermore, by setting independent gated network parameters and linear transformation matrices for each processing stream, the fusion process is made asymmetric, naturally modeling the asymmetric coupling pattern of one party leading and the other following in the interaction—something traditional symmetric fusion methods cannot achieve. This technique allows the fused interactive perception feature sequence to more precisely and robustly preserve both the internal dynamics of individuals and the cross-individual coupling information, providing a higher-quality input foundation for subsequent cross-correlation calculations and synchronous analysis.

[0080] Example 3: The cross-correlation function is decomposed into multiple scales in the time domain using continuous wavelet transform to obtain synchronization components at multiple scale levels. The steps include: S21. Using the complex Morlet wavelet as the mother wavelet function, perform continuous wavelet transform on the cross-correlation function under multiple scale factors to obtain the time-scale energy spectrum.

[0081] Specifically, the complete mathematical expression for the complex Morlet wavelet mother function is: .in, For continuous time variables, The imaginary unit satisfies , The center angular frequency is the value of which determines the trade-off between the time-domain oscillation frequency and the frequency-domain bandwidth of the wavelet. The normalization factor ensures the energy of the wavelet. The real part of the mother wavelet is a frequency of The cosine wave is modulated by a Gaussian window, and the imaginary part is modulated by the same Gaussian window. The combination of the two allows wavelet transform to capture both the amplitude and phase information of the signal simultaneously.

[0082] Let the comprehensive cross-correlation function obtained in step S3 be... ,in For integer time delay variables. Select a set of scaling factors. Uniform sampling along the logarithmic scale axis covers multiple orders of magnitude from high-frequency microscale to low-frequency macroscale. This represents the total number of scales. For each scale factor... and each translation factor ( Taking all time delay positions, the formula for calculating the discrete continuous wavelet transform is: In the formula, express The complex conjugate function, time delay variable According to translation and scale Perform scaling and translation. This is the energy normalization factor, ensuring the comparability of wavelet coefficients at different scales. Summation iterates through all possible discrete time delays. Wavelet transform coefficients It is a complex number, its modulus Indicates the translation position ,scale Synchronous energy amplitude at the point, modulus squared This constitutes a two-dimensional time-scale energy spectrum. The rows of this energy spectrum correspond to different scales, and the columns correspond to different translation positions. The magnitude of each element value reflects the strength of the cross-correlation energy between the expressions of the two parties at a specific time delay position and a specific time scale.

[0083] S22. Divide the sequence composed of scale factors into multiple scale bands according to the preset scale division threshold.

[0084] Specifically, based on empirical psychological research regarding the functional significance of mirror responses at different time scales in interpersonal interactions, two scale thresholds are set to divide all... The scale factor is divided into three continuous scale bands. The first threshold corresponds to a point on the scale factor axis, separating the high-frequency microscale from the mid-frequency mesoscale; the second threshold separates the mid-frequency mesoscale from the low-frequency macroscale. Specifically, the microscale band... This frequency band includes all scales with scale factors less than the first dividing threshold. The corresponding wavelets in this band have high center frequencies and narrow time-domain windows, primarily capturing instantaneous facial expression imitations and instinctive emotional contagion signals in the millisecond to second range. (Mesoscale band) This includes scales with scaling factors between the first and second segmentation thresholds, corresponding to expression duration matching and interaction norm coordination from several seconds to tens of seconds. Macro-scale frequency band. It includes all scales with scale factors greater than the second dividing threshold, corresponding to the gradual change in emotional tendency and the fusion of relational levels from ten seconds to the entire dialogue process. These three frequency bands do not overlap and cover the entire scale range.

[0085] S23. Calculate the time-averaged energy, energy standard deviation, and energy skewness of the energy spectrum within each scale band to form the synchronization component of the corresponding scale level.

[0086] Specifically, taking the microscale frequency band as an example, the calculation of three statistics will be explained in detail. Let the translation factor be... The total number of possible values ​​is First, calculate the average energy of all scales within the microscale band on the time-shift axis: ,in This represents the number of scales contained in the frequency band, with the scale value in parentheses. The average energy along the lower translation axis is summed in the outer layers and then divided by the number of scales to obtain the overall average energy of the frequency band. This value reflects the average intensity of instantaneous imitation between the two parties at the microscale.

[0087] Secondly, the energy standard deviation of the microscale frequency band is calculated to measure the degree of fluctuation of the microscale synchronization energy along the time shift axis: ,in For scale The average energy under the given conditions. The larger the standard deviation, the more drastic the fluctuations in microscale synchronization over time during the interaction, with periods of strength and weakness.

[0088] Third, calculate the energy skewness of the microscale frequency band to measure the degree of asymmetry in the microscale synchronization energy distribution: skewness A positive skewness indicates that the energy distribution is skewed above the mean, suggesting intermittent high-intensity synchronous bursts; a negative skewness indicates that the energy distribution is skewed below the mean, suggesting overall weaker synchronization with occasional troughs. These three statistics jointly describe the total energy, stability, and distribution pattern of microscale synchronization, together forming the microscale synchronization component vector. Similarly, performing the exact same calculations on the mesoscale and macroscale frequency bands yields the following results. and .

[0089] Example 4: A hierarchical synchronization representation vector is constructed by interactively fusing synchronization components at multiple scale levels using a cross-scale attention mechanism, including the following steps: S31. Based on each component in the synchronization components of multiple scale levels, generate the corresponding query vector, key vector, and value vector through the corresponding learnable linear projection matrix.

[0090] Specifically, let's assume that the synchronization components at three scale levels are obtained from Example 3: microscale synchronization components. Mesoscale synchronization components and macro-scale synchronization components ,in The vector dimension for each scale component (in Example 3) ). For each scale level Set up three independent learnable linear projection matrices: query projection matrix Key projection matrix Sum projection matrix ,in The dimensions of the query and key vectors. Query, key, and value vectors at each scale are generated through matrix multiplication: , , In the formula, For the first Scale synchronization component row vector ( ), For the first A query vector for a scale, used to retrieve information from other scales; For the first The key vector of the scale is used for matching with other scales; For the first The value vector of the scale carries the actual information content that the scale transmits to other scales.

[0091] S32. For any two scale levels, calculate the attention weight of the second scale on the first scale based on the query vector of the first scale and the key vector of the second scale.

[0092] Specifically, regarding the target scale Source Scale Attention weight Query vector of target scale Key vectors at source scale The dot product similarity is obtained by softmax normalization:

[0093] In the formula, for transpose, Let be the dot product of the two (scalar), representing the first... Scale to the first The degree of need for scale information; divided by In order to prevent when Large time-mapping values ​​can cause softmax gradient vanishing; the denominator varies across the three source scales. Summation, such that satisfy That is, all source scales relative to the target scale The sum of attention weights is . Quantified the update target scale When representing, source scale What percentage of information should be included?

[0094] S33. Using the attention weights of all scale levels for each scale level, perform a weighted summation of the value vectors of all scale levels to obtain the updated components of the corresponding scale level.

[0095] Specifically, for each target scale Its updated components It is obtained by weighting the value vectors of all three source scales according to the attention weights: In the formula, For the first The value vector of the scale, The attention weights are calculated in step S32. This operation ensures that the updated component at each scale not only contains its own original information but also incorporates information selectively passed from other scales based on the attention weights.

[0096] S34. The updated components of each scale level are concatenated to obtain a hierarchical synchronous representation vector.

[0097] Specifically, the updated microscale components Mesoscale components and macroscale components Concatenate them sequentially to form a hierarchical synchronous representation vector: .

[0098] Compared to the simplified self-attention in Example 1, this embodiment sets an independent projection matrix for each scale, allowing each scale to learn its own query, key, and value representation space. This design allows for the development of differentiated information retrieval and delivery strategies at different scales, and the explicit computation of cross-scale attention weights also provides a window into the model's interpretability. This multi-scale information interaction mechanism significantly improves the semantic richness of hierarchical synchronous representations and the ability to model complex interaction patterns.

[0099] Example 5: refer to Figure 2 Based on the first and second interactive perception feature sequences, the covariance matrix of the activation of the action units of both parties is calculated; a dispersion index is generated according to the statistics of the trace of the covariance matrix, and the sequence composed of the dispersion index is encoded to obtain the dispersion feature vector, including the following steps: S41. Perform linear mapping on the first interactive perception feature sequence and the second interactive perception feature sequence respectively to obtain the first action unit mapping sequence and the second action unit mapping sequence; wherein, the vector dimension of each time step of the first action unit mapping sequence and the second action unit mapping sequence is equal to the total number of action units.

[0100] Specifically, the feature map of the two-stream convolutional neural network after the first residual block and before fusion at the first fusion layer is selected as the basis, because the features at this position have not yet undergone cross-stream mixing, and each processing stream retains a relatively pure response to its own facial action unit. Let the output feature map of the first processing stream at this position be denoted as... The feature map corresponding to the second processing stream is ,in This is the effective time dimension for this early level. Let be the number of channels. A learnable linear mapping matrix is ​​introduced. ,Will The feature channel linear transformation of the dimension is as follows The action unit intensity reconstruction vector in the dimensional dimension. Matrix multiplication is performed in the channel dimension: , .in, This is the mapping sequence for the first action unit. This is the mapping sequence for the second action unit. The first action unit of this mapping sequence... Line 1 Column elements can be interpreted as the first participant in the first column. The first time step Reconstruction intensity of each action unit. Mapping matrix. End-to-end learning is achieved through backpropagation of subsequent task losses, enabling the network to automatically extract the AU activation information most useful for discreteness evaluation from early features.

[0101] S42. Divide the first action unit mapping sequence and the second action unit mapping sequence into multiple overlapping sliding time windows along the time dimension.

[0102] Specifically, let the length of the sliding time window be... The interval between the starting frames of adjacent windows is [frame number]. frame( This results in overlap. Starting from the first frame of the sequence, the frames are sequentially incremented by a step size... Sliding cut, a total of The first window. The frame index range covered by each window is ,in , For window indexing.

[0103] S43. For each sliding time window, calculate the first action unit activation covariance matrix of the first participant based on the first sample set composed of the vectors of the first action unit mapping sequence within the window, and calculate the second action unit activation covariance matrix of the second participant based on the second sample set composed of the vectors of the second action unit mapping sequence within the window.

[0104] Specifically, for windows ,from Extract the window coverage area The row vectors corresponding to the frames form the first sample set. , where each sample vector Include Reconstruction intensity value of each AU, This is the index of the frame within the window. Calculate the sample mean vector of the first sample set. , its first The components are: .

[0105] The first participant at the window Action unit activation covariance matrix The Line 1 The column elements are given by the following formula: .in, For the first The th sample vector AU components Let be the mean of AU. and For AU indices, the value range is 1 to 1. The diagonal elements of this matrix Indicates the first An AU in the window The variance of the reconstruction intensity within the value reflects the magnitude of the change in activation of the AU. Off-diagonal elements represent covariance relationships between different AUs. Completely symmetrically, from... Extracting windows The second sample set is used to calculate the action unit activation covariance matrix of the second participant. Its elements The calculation method is the same as above, only the corresponding AU component is replaced with the data of the second participant.

[0106] S44. Based on the traces of the activation covariance matrix of the first action unit and the traces of the activation covariance matrix of the second action unit, calculate the dispersion index of the corresponding sliding time window; wherein, the formula for calculating the dispersion index is:

[0107] in, For the index of the sliding time window, For the first participant in the window The activation covariance matrix of the first action unit within the matrix, For the second participant in the window The second action unit within the matrix activates the covariance matrix. Represents the trace of a matrix. This represents the total number of action units. For window The divergence index.

[0108] Specifically, It is the sum of all diagonal elements of the first participant's covariance matrix, which is the sum of all AU variances; Similarly, this is the sum of all AU variances for the second participant. Add the two together and divide by... The average activation variance of each active user (AU) is obtained, eliminating the influence of the total number of AUs and making the index comparable across different AU set sizes. This index comprehensively measures the richness and variability of facial expressions between the two parties within the window, serving as a key quantitative basis for distinguishing genuine emotional resonance from mechanical pseudo-synchronization.

[0109] S45. Arrange the dispersion indices of all sliding time windows in the entire time period in chronological order to obtain a dispersion index sequence; use a one-dimensional convolutional encoder to encode the dispersion index sequence to obtain a dispersion feature vector of fixed length.

[0110] Specifically, indexed by window Arrange the dispersion indices into a one-dimensional sequence from smallest to largest. This sequence describes the temporal evolution trajectory of facial expression diversity throughout the interaction. Features are extracted from this sequence using a one-dimensional convolutional encoder. The first layer of this encoder is a one-dimensional convolutional layer with... There are n convolutional kernels, and the kernel length is 1. With a step size of 1, the single-channel input is mapped to... Each feature channel has several channels, followed by ReLU activation. The second layer is also a one-dimensional convolution, maintaining the same number of channels and kernel length, followed by a global average pooling layer, which aggregates all time steps within each feature channel into a single value, ultimately outputting a fixed-length discrete feature vector. .

[0111] Example 6: The concatenation result of the hierarchical synchronization representation vector and the discrete feature vector is input into a non-monotonic mapping network with radial basis function activation layers, and the output psychological compatibility evaluation value is obtained by the following steps: S51. The concatenation result of the hierarchical synchronous representation vector and the discrete feature vector is input into the fully connected layer and passed through the ReLU activation function to obtain the hidden representation vector.

[0112] Specifically, hierarchical synchronization representation vectors (In Example 4) In Example 1 ) and discrete eigenvectors By concatenating the features, a comprehensive feature vector is obtained. ,in To hierarchically synchronize the dimension of the representation vector, Let be the dimension of the discrete feature vectors. Introduce the weight matrix of the fully connected hidden layer. and bias vector Calculate the hidden representation vector: .in, This is the element-wise operation of the ReLU activation function. To hide the dimension of the representation vector, this fully connected layer performs an initial interactive mixing of synchronization and discrete information, extracting their joint features, which lays the foundation for the nonlinear mapping of subsequent radial basis function layers.

[0113] S52. Input the hidden representation vector into a non-monotonic mapping layer composed of multiple radial basis function units. For each radial basis function unit, calculate the activation response based on the weighted Euclidean distance between the hidden representation vector and the learnable center vector of the radial basis function unit; wherein, the formula for calculating the activation response is:

[0114] in, To hide the representation vector, For the first The learnable center vector of each radial basis function unit. For the first Learnable scaling matrix of radial basis function units To hide the difference vector between the representation vector and the learnable center vector, For the first The radial basis function units with respect to the hidden representation vector Activation response, This indicates the transpose operation.

[0115] Specifically, the non-monotonic mapping layer is the core of the network, containing... Radial basis function unit. Radial basis function units ( It has two learnable parameters: the center vector. and scaling matrix Scaling matrix Constrained to be a symmetric positive definite matrix, in practice it can be parameterized as a diagonal matrix to reduce the number of parameters and ensure positive definiteness, i.e. , where all diagonal elements , For dimensional indexing. (The first...) The activation response of each unit is represented by the hidden representation vector. With the center vector The squared Mahalanobis distance (in) (The metric matrix) is obtained through exponential decay: In the formula, It is a difference vector; ,when When the matrix is ​​diagonal, the expression simplifies to That is, the weighted squared Euclidean distance across all dimensions. Coefficients Mapping distance to negative values, exponential function Transform it into arrive The activation value between. When With the center When they completely overlap, the distance is , Reaching the maximum value; when keep away hour, Rapid decay approaches Scaling matrix Controlling the response range of this unit in each direction within the input space, The larger, in the first The narrower the response on a dimension, the stronger the selectivity.

[0116] S53. Based on the activation responses of all radial basis function units, calculate the psychological compatibility assessment value using the Sigmoid function; wherein, the formula for calculating the psychological compatibility assessment value is:

[0117] in, For the first The activation response of a radial basis function unit; For the first Learnable weight coefficients for each radial basis function unit; For learnable bias terms; This represents the total number of radial basis function elements. For the Sigmoid function, This is an assessment value for psychological compatibility.

[0118] Specifically, the non-monotonic mapping capability and interpretability of this network structure stem from the following mechanism. During training, the center of each radial basis unit... Gradient updates move the cells through the joint synchronous-discrete input space, eventually distributing them to regions with varying degrees of representativeness. For example, the centers of some cells might shift to regions with moderate synchronousity and high dispersion (representing genuine and natural emotional resonance), and their output weights would change accordingly. Some units are learned to have larger positive values, contributing positively to the final output. Other units, however, may shift their centers to regions of high synchronization but low dispersion (representing pseudo-synchronization in the machine), and their weights... The values ​​are learned to be negative or extremely positive, thus suppressing the output in that region. Due to the exponential decay of the radial basis function, the network naturally forms a ridge-like response surface across the entire input space—forming a high-value "peak" in the region of moderate synchronization and high dispersion, and falling into a "valley" in the region of high synchronization and low dispersion, perfectly realizing an inverted U-shaped non-monotonic mapping. This mapping relationship is entirely learned through data-driven learning, without requiring any explicit function form to be specified in advance, fundamentally breaking through the erroneous assumption in existing technologies that synchronization and tacit understanding must be monotonically increasing. Compared to the simplified scheme using isotropic Gaussian kernels in Embodiment 1, this embodiment introduces a learnable diagonal scaling matrix. This allows each radial basis unit to acquire different sensitivities in different dimensions of the input space, significantly enhancing the network's ability to fit complex distributed data and its generalization performance.

[0119] The aforementioned method for analyzing interpersonal interaction emotion synchronization starts with the intensity sequence of facial action units from both parties. It extracts individual facial expression dynamic features layer by layer using a two-stream convolutional network, and progressively fuses the two-stream features across multiple depth levels by adaptively generating gating vectors based on information compression states. This approach preserves the independent dynamics of individual micro-expressions while fully capturing the temporal information of the interaction coupling between the two parties, obtaining an interaction-perceived feature sequence. Based on this, cross-correlation calculations are performed on the interaction-perceived feature sequence to obtain a cross-correlation function. Continuous wavelet transform is used to decompose the cross-correlation function into synchronization components at multiple scale levels. A cross-scale attention mechanism is then used to interactively fuse synchronization information at different scales, constructing a hierarchical synchronization representation vector. This achieves the fusion of micro-scale emotion contagion, meso-scale interaction normative coordination, and macro-scale relationship fusion. Separation and collaborative modeling are employed. Simultaneously, the covariance matrix of the activation of action units of both parties is calculated from the interactive perception feature sequence. Based on the statistics of the trace of the covariance matrix, a dispersion index is generated and encoded as a dispersion feature vector to quantify the diversity of facial expression activation patterns. The hierarchical synchronization representation vector and the dispersion feature vector are concatenated and input into a non-monotonic mapping network with a radial basis function activation layer. By leveraging the peak response characteristics inherent in the radial basis function, the network can learn the inverted U-shaped mapping relationship between synchronization and psychological tacit understanding, automatically distinguishing between genuine emotional resonance with high dispersion and moderate synchronization and performative pseudo-synchronization with low dispersion and high synchronization, and outputting a psychological tacit understanding evaluation value. Finally, by dividing the interaction process into multiple overlapping windows and obtaining the time series of psychological tacit understanding evaluation values ​​and their first-order differences, a conflict risk prediction result is generated. Therefore, this solution overcomes the information confusion defect of single-time-scale synchronous analysis by organically coordinating three technical links: progressive fusion, multi-scale time-domain decomposition, and non-monotonic mapping. It fundamentally overcomes the erroneous assumption that the higher the degree of synchronization, the higher the degree of tacit understanding, and achieves accurate quantification of genuine emotional resonance in interpersonal interactions and effective identification of pseudo-synchronization.

[0120] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0121] Based on the same inventive concept, this application also provides a system for implementing the interpersonal interaction emotion synchronization analysis method described above. The solution provided by this system is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more embodiments of the interpersonal interaction emotion synchronization analysis system provided below can be found in the limitations of the interpersonal interaction emotion synchronization analysis method described above, and will not be repeated here.

[0122] In one exemplary embodiment, such as Figure 3 As shown, an interpersonal interaction emotion synchronization analysis system 30 is provided to implement the methods in the above-described method embodiments. The system includes: The multimodal action unit recognition module 31 is used to perform action unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously, and to obtain the first facial action unit intensity time series of the first participant and the second facial action unit intensity time series of the second participant.

[0123] The dual-stream progressive fusion feature extraction module 32 is used to input the first facial motion unit intensity time series and the second facial motion unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. The first and second feature maps of multiple levels are extracted layer by layer through the first and second processing streams. The gating vector is adaptively generated according to the information compression state of each level. The first feature map and the second feature map are progressively fused across streams using the gating vector to obtain the first interactive perception feature sequence and the second interactive perception feature sequence.

[0124] The multi-scale synchronization feature parsing module 33 is used to obtain a cross-correlation function by performing cross-correlation calculation on the first interactive sensing feature sequence and the second interactive sensing feature sequence; to perform multi-scale time-domain decomposition on the cross-correlation function using continuous wavelet transform to obtain synchronization components at multiple scale levels; and to use a cross-scale attention mechanism to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector.

[0125] The synchronous discreteness encoding module 34 is used to calculate the covariance matrix of the activation of the action units of both parties based on the first interactive perception feature sequence and the second interactive perception feature sequence; generate a discreteness index according to the statistics of the trace of the covariance matrix; and encode the sequence composed of the discreteness index to obtain the discreteness feature vector.

[0126] The non-monotonic psychological mapping module 35 is used to input the concatenation result of the hierarchical synchronous representation vector and the discrete feature vector into the non-monotonic mapping network with the radial basis function activation layer, and output the psychological compatibility evaluation value.

[0127] The time-series dynamic assessment and risk prediction module 36 is used to divide the interaction process into multiple overlapping time windows. Within each overlapping time window, the operations of the dual-stream progressive fusion feature extraction module 32, the multi-scale synchronous feature parsing module 33, the synchronous discreteness encoding module 34, and the non-monotonic psychological mapping module 35 are executed respectively to obtain the time series of psychological tacit understanding evaluation values. Based on the time series of psychological tacit understanding evaluation values ​​and its first-order difference, the conflict risk prediction result is generated.

[0128] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.

[0129] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0130] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0131] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for synchronous analysis of interpersonal interaction emotions, characterized in that, The method includes: S1. Perform motion unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously to obtain the first facial motion unit intensity time series of the first participant and the second facial motion unit intensity time series of the second participant. S2. Input the first facial motion unit intensity time series and the second facial motion unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. Extract the first feature map and the second feature map of multiple levels layer by layer through the first processing stream and the second processing stream. Adaptively generate a gating vector according to the information compression state of each level. Use the gating vector to progressively fuse the first feature map and the second feature map across streams to obtain the first interactive perception feature sequence and the second interactive perception feature sequence. S3. By performing cross-correlation calculation on the first interactive sensing feature sequence and the second interactive sensing feature sequence, a cross-correlation function is obtained; the cross-correlation function is decomposed into multiple scale time domains using continuous wavelet transform to obtain synchronization components at multiple scale levels; the synchronization components at multiple scale levels are interactively fused using a cross-scale attention mechanism to construct a hierarchical synchronization representation vector. S4. Based on the first interactive perception feature sequence and the second interactive perception feature sequence, calculate the covariance matrix of the activation of the action units of both parties; generate a dispersion index according to the statistics of the trace of the covariance matrix, and encode the sequence composed of the dispersion index to obtain a dispersion feature vector. S5. Input the concatenation result of the hierarchical synchronization representation vector and the discrete feature vector into a non-monotonic mapping network with radial basis function activation layer, and output the psychological tacit understanding evaluation value. S6. Divide the interaction process into multiple overlapping time windows, and execute S2 to S5 in each overlapping time window to obtain a time series of psychological compatibility assessment values; generate a conflict risk prediction result based on the time series of psychological compatibility assessment values ​​and its first difference.

2. The method according to claim 1, characterized in that, The step of adaptively generating a gating vector based on the information compression state of each level, and using the gating vector to progressively fuse the first feature map and the second feature map across streams, yields a first interactive perception feature sequence and a second interactive perception feature sequence, including: S11. At the current level, the first feature map currently output by the first processing stream and the second feature map currently output by the second processing stream are concatenated in the channel dimension to obtain a concatenated feature map. S12. Perform one-dimensional convolution, ReLU activation, one-dimensional convolution and Sigmoid activation sequentially on the spliced ​​feature map to generate a first gate vector for the first processing stream. S13. Perform matrix multiplication between the second feature map and the first linear transformation matrix to obtain the mapped second feature map; S14. Using the first gating vector, perform element-wise weighted summation on the first feature map and the mapped second feature map to obtain the first fused feature map; S15. Perform one-dimensional convolution, ReLU activation, one-dimensional convolution and Sigmoid activation sequentially on the spliced ​​feature map to generate a second gate vector for the second processing stream. S16. Perform matrix multiplication between the first feature map and the second linear transformation matrix to obtain the mapped first feature map; S17. Using the second gate vector, perform element-wise weighted summation on the second feature map and the mapped first feature map to obtain the second fused feature map; S18. Input the first fused feature map and the second fused feature map into the next network layer respectively. After traversing multiple layers, the first interactive perception feature sequence and the second interactive perception feature sequence are obtained.

3. The method according to claim 1, characterized in that, The method employs continuous wavelet transform to perform multi-scale time-domain decomposition of the cross-correlation function, obtaining synchronization components at multiple scale levels, including: S21. Using the complex Morlet wavelet as the mother wavelet function, perform continuous wavelet transform on the cross-correlation function under multiple scale factors to obtain the time-scale energy spectrum. S22. Divide the sequence composed of the scale factors into multiple scale bands according to the preset scale division threshold; S23. Calculate the time-averaged energy, energy standard deviation, and energy skewness of the energy spectrum within each scale frequency band to form the synchronization component of the corresponding scale level.

4. The method according to claim 1, characterized in that, The method of using a cross-scale attention mechanism to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector includes: S31. Based on each component in the synchronization components at multiple scale levels, generate corresponding query vectors, key vectors, and value vectors through the corresponding learnable linear projection matrices. S32. For any two scale levels, calculate the attention weight of the second scale on the first scale based on the query vector of the first scale and the key vector of the second scale. S33. Using the attention weights of all scale levels for each scale level, perform a weighted summation of the value vectors of all scale levels to obtain the updated components of the corresponding scale level. S34. The updated components at each scale level are concatenated to obtain the hierarchical synchronization representation vector.

5. The method according to claim 1, characterized in that, The covariance matrix of the activation of the action units of both parties is calculated based on the first interactive perception feature sequence and the second interactive perception feature sequence. A dispersion index is generated based on the statistics of the trace of the covariance matrix, and the sequence composed of the dispersion index is encoded to obtain a dispersion feature vector, including: S41. Perform linear mapping on the first interactive perception feature sequence and the second interactive perception feature sequence to obtain the first action unit mapping sequence and the second action unit mapping sequence; wherein, the vector dimension of each time step of the first action unit mapping sequence and the second action unit mapping sequence is equal to the total number of action units. S42. Divide the first action unit mapping sequence and the second action unit mapping sequence into multiple overlapping sliding time windows along the time dimension; S43. For each sliding time window, calculate the first action unit activation covariance matrix of the first participant based on the first sample set composed of the vectors of the first action unit mapping sequence within the window, and calculate the second action unit activation covariance matrix of the second participant based on the second sample set composed of the vectors of the second action unit mapping sequence within the window. S44. Based on the trace of the activation covariance matrix of the first action unit and the trace of the activation covariance matrix of the second action unit, calculate the dispersion index of the corresponding sliding time window; wherein, the formula for calculating the dispersion index is: ; in, The index of the sliding time window. For the first participant in the window The first action unit within the matrix activates the covariance matrix. For the second participant in the window The second action unit within the matrix activates the covariance matrix. Represents the trace of a matrix. This represents the total number of action units. For window The divergence index; S45. Arrange the dispersion indices of all the sliding time windows in the entire time period in chronological order to obtain a dispersion index sequence; use a one-dimensional convolutional encoder to encode the dispersion index sequence to obtain a dispersion feature vector of fixed length.

6. The method according to claim 1, characterized in that, The concatenation result of the hierarchical synchronization representation vector and the discrete feature vector is input into a non-monotonic mapping network with a radial basis function activation layer, and the output psychological compatibility evaluation value includes: S51. Input the concatenation result of the hierarchical synchronous representation vector and the discrete feature vector into the fully connected layer, and obtain the hidden representation vector after passing through the ReLU activation function; S52. Input the hidden representation vector into a non-monotonic mapping layer composed of multiple radial basis function units. For each radial basis function unit, calculate the activation response based on the weighted Euclidean distance between the hidden representation vector and the learnable center vector of the radial basis function unit; wherein, the formula for calculating the activation response is: ; in, The hidden representation vector, For the first The learnable center vector of each radial basis function unit, For the first Learnable scaling matrix of radial basis function units Let be the difference vector between the hidden representation vector and the learnable center vector. For the first The radial basis function units with respect to the hidden representation vector Activation response, Indicates the transpose operation; S53. Based on the activation responses of all radial basis function units, calculate the psychological compatibility assessment value using the Sigmoid function; wherein, the formula for calculating the psychological compatibility assessment value is: ; in, For the first The activation response of each radial basis function unit; For the first Learnable weight coefficients for each radial basis function unit; For learnable bias terms; This represents the total number of radial basis function elements. For the Sigmoid function, The psychological compatibility assessment value is given.

7. A system for synchronous analysis of interpersonal interaction emotions, used to implement the method according to any one of claims 1 to 6, characterized in that, The system includes: The multimodal action unit recognition module is used to perform action unit recognition on each frame of the facial video of the first participant and the second participant acquired synchronously, so as to obtain the first facial action unit intensity time series of the first participant and the second facial action unit intensity time series of the second participant. The dual-stream progressive fusion feature extraction module is used to input the first facial action unit intensity time series and the second facial action unit intensity time series into the first processing stream and the second processing stream of the dual-stream convolutional neural network, respectively. The module extracts first and second feature maps at multiple levels through the first and second processing streams, adaptively generates gate vectors based on the information compression state of each level, and uses the gate vectors to progressively fuse the first and second feature maps across streams to obtain the first interactive perception feature sequence and the second interactive perception feature sequence. The multi-scale synchronization feature parsing module is used to obtain a cross-correlation function by performing cross-correlation calculation on the first interactive sensing feature sequence and the second interactive sensing feature sequence; to perform multi-scale time-domain decomposition on the cross-correlation function using continuous wavelet transform to obtain synchronization components at multiple scale levels; and to use a cross-scale attention mechanism to interactively fuse the synchronization components at multiple scale levels to construct a hierarchical synchronization representation vector. The synchronous discreteness encoding module is used to calculate the covariance matrix of the activation of the action units of both parties based on the first interactive perception feature sequence and the second interactive perception feature sequence; generate a discreteness index according to the statistics of the trace of the covariance matrix; and encode the sequence composed of the discreteness index to obtain a discreteness feature vector. The non-monotonic psychological mapping module is used to input the concatenation result of the hierarchical synchronous representation vector and the discrete feature vector into a non-monotonic mapping network with a radial basis function activation layer, and output a psychological compatibility evaluation value. The time-series dynamic assessment and risk prediction module is used to divide the interaction process into multiple overlapping time windows. Within each overlapping time window, the operations of the dual-stream progressive fusion feature extraction module, the multi-scale synchronous feature parsing module, the synchronous discreteness encoding module, and the non-monotonic psychological mapping module are executed to obtain a time series of psychological compatibility assessment values. Based on the time series of psychological compatibility assessment values ​​and its first-order difference, a conflict risk prediction result is generated.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.