A depression detection system and method based on multi-modal emotion conflict quantification

CN122531728APending Publication Date: 2026-08-07THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV
Filing Date
2026-05-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

因此,现有系统无法有效捕获并度量“面部愉悦极性”与“语音/语义悲伤极性”之间的剧烈冲突,难以从底层算法逻辑上区分正常的生理性情绪波动与病理性的情感掩饰

Benefits of technology

[0039]1.突破了传统多模态网络“模态一致性假设”的原理性局限。现有技术致力于模态间的特征融合,而本发明首创了基于“特征极性互斥性”度量的检测范式。通过物理隔离的双路提取网络,强行解耦受躯体神经控制的外显特征与受自主神经控制的内隐特征,彻底解决了传统注意力机制易被高置信度社交伪装信号“劫持”和误导的算法缺陷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531728A_ABST
    Figure CN122531728A_ABST
Patent Text Reader

Abstract

The application discloses a depression detection system and method based on multi-modal emotion conflict quantification, and belongs to the field of artificial intelligence and biological information processing. In view of the defect that the prior art is easily misled by subjective social camouflage such as "smile depression", the application discloses a double neural drive decoupling network: through an explicit branch, facial action unit features controlled by somatic nerves are extracted, and through an implicit branch, speech perturbation and negative semantic features controlled by autonomic nerves are extracted. Then, in the emotion representation space, Mahalanobis space deviation and fractional order time series dynamics gradient are introduced to accurately calculate the cross-modal emotion conflict index (CI). Finally, combined with a time convolution network and an adaptive dynamic confidence boundary, a depression risk probability is output. The application can effectively quantify the polarity deviation phenomenon of "smile", and greatly reduce the missed diagnosis rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and bioinformatics processing technology, specifically relating to a system and method for detecting depression in patients with highly defensive psychological characteristics such as "smiling depression" through multimodal emotional conflict quantification. Background Technology

[0002] In the interdisciplinary field of intelligent health monitoring and psychiatry, early screening and objective assessment of depression are evolving from traditional self-reporting scales to multimodal biometric recognition based on artificial intelligence. With the development of deep learning technology, the integrated use of multimodal signals such as facial expressions, speech acoustic features, and text semantics for diagnosis has become a key technological approach to improve the accuracy of automated screening for depression.

[0003] In existing technologies, deep learning-based multimodal depression detection systems generally adhere to the "Modality Consistency Assumption." Their mainstream architectures (such as multi-scale feature pyramids and cross-modal attention mechanisms) aim to uncover commonalities between visual and auditory features, achieving "convergence" at the decision or feature levels through joint representation learning. However, our research has revealed a serious mechanistic error in the aforementioned "Modality Consistency Assumption" in depression detection scenarios involving highly psychological defense mechanisms, such as "smiling depression." Depressed patients often forcibly drive facial action units (e.g., activating AU12 to fake a smile) during interactions, while genuine negative emotions are physiologically leaked through autonomic neural channels beyond conscious control (e.g., high-frequency micro-vibrations in speech). The existing "convergence" architectures mathematically seek covariance between modalities. This mechanism not only fails to recognize the aforementioned polarity reversal phenomenon but also treats high-confidence facial faking features as valid signals and filters out weak speech leakage features as noise. This blind suppression of "feature exclusivity" in the underlying algorithm logic is the root cause of the high rate of missed diagnoses when existing systems face psychological deception.

[0004] First, existing systems are extremely vulnerable to interference from "subjective pretense." Patients with smiling depression often activate psychological defense mechanisms in social or diagnostic situations, forcibly controlling facial muscles controlled by the somatic nervous system (such as pulling the zygomaticus major muscle to fake a smile) to project a positive social facade. Existing multimodal consistency fusion models, lacking decoupling from biological features, are easily hijacked by these high-confidence "obvious positive features," leading to the model being misled by "social illusions" during feature fusion and resulting in serious missed diagnoses.

[0005] Secondly, existing algorithms lack a quantitative modeling mechanism for the "modal polarity divergence phenomenon." The true negative emotions of depressed patients often leak through "implicit physiological channels" regulated by the autonomic nervous system and the hypothalamic-pituitary-adrenal (HPA) axis (such as micro-vibrations in the fundamental frequency of speech, formant shifts caused by vocal cord tension, or negative biases in subconscious text). Existing attention fusion mechanisms tend to seek "covariance" and commonalities between modalities; their mathematical nature dictates that they suppress or even filter out conflicting outliers. Therefore, existing systems cannot effectively capture and measure the intense conflict between "facial pleasure polarity" and "vocal / semantic sadness polarity," making it difficult to distinguish between normal physiological emotional fluctuations and pathological emotional masking from the underlying algorithmic logic.

[0006] Finally, existing evaluation models lack the ability to analyze multi-dimensional spatiotemporal context. Emotional concealment is often accompanied by a temporal delay between micro-expression flashes and speech pauses. Traditional methods based on simple time window splicing cannot accurately capture the dynamic micro-delay between overt concealment behavior and implicit physiological disclosure, resulting in significant deficiencies in the system's real-time performance and evaluation robustness.

[0007] Therefore, there is an urgent need in this field for a system that can break through the limitations of the traditional "modal consistency" framework, comprehensively decouple explicit social masking features from implicit physiological disclosure features from the underlying logic, and achieve a high-precision, low-underreporting depression risk assessment system by accurately quantifying the "conflict degree" between the two in the emotional space, so as to solve the problem of the lack of accuracy and reliability of existing technologies when faced with complex psychological disguises. Summary of the Invention

[0008] To address the fundamental flaw in existing multimodal consistency fusion models, which are prone to misdiagnosis due to the subjectively controlled social deception of subjects with highly developed psychological defense mechanisms, such as those exhibiting "smiling depression," this invention aims to provide a depression detection system and method based on multimodal emotional conflict quantification. This invention overcomes the reliance of traditional algorithms on "modal commonality" by separating explicit and implicit emotional pathways at the underlying architecture, accurately quantifying cross-modal polarity divergence, thereby achieving objective identification of emotional concealment behavior.

[0009] To achieve the above objectives, the technical solution adopted by the present invention to solve its technical problems is as follows:

[0010] This invention provides a depression detection system based on multimodal affective conflict quantification. The system achieves detection by constructing a closed-loop architecture of "data alignment—physiological / physiological decoupling—spatial mapping—conflict quantification," specifically including the following core modules:

[0011] 1. Spatiotemporal alignment module for heterogeneous multimodal stream signals based on phase space synchronization

[0012] This module is used to acquire high-definition facial video streams (spatially distributed high-dimensional matrix), physiological-grade audio streams (one-dimensional temporal high-frequency waveforms), and semantic text streams (discrete semantic encoding) of subjects during interaction in parallel. To address the non-stationary phase delays in heterogeneous sensor signals across sampling rates, dimensions, and transmission nodes, this module constructs a sub-millisecond-level joint calibration mechanism of timestamp hard alignment and dynamic time warping (DTW). Through multi-dimensional temporal interpolation and resampling filters, the system eliminates frequency mismatches in multi-source sensor data and maps the original multimodal data streams into a time-scale strictly isomorphic reference tensor sequence, providing a zero-phase-error data basis for subsequent instantaneous dynamic gradient calculations of cross-modal features.

[0013] 2. Heterogeneous Feature Decoupling and Representation Network Driven by Both Somatic and Autonomic Nervous Systems

[0014] The system abandons the traditional early feature cascade blind fusion mechanism and constructs physically isolated dual-path orthogonal feature extraction branches based on the essential differences in neurophysiological drivers:

[0015] Overt social masking branch (somatic neural representation)

[0016] Targeting facial expression muscles strongly controlled by the somatic nervous system and subjective consciousness, a lightweight 3D convolutional neural network incorporating attention mechanisms is deployed. By constructing a refined spatiotemporal receptive field for the face, dynamic evolution maps of facial action units (AUs) are extracted. The transient activation intensities of the zygomaticus major muscle (AU12), corresponding to the upward movement of the corners of the mouth, and the orbicularis oculi muscle (AU6), corresponding to the contraction of the corners of the eyes, which are strongly correlated with the "social mask," are quantified. The explicit emotional valence is then mapped and output via a nonlinear fully connected layer. ) and explicit arousal ( ).

[0017] Implicit physiological leakage branches (autonomic nervous system / HPA axis representation)

[0018] To address the physiological stress response regulated by the hypothalamic-pituitary-adrenal (HPA) axis and extremely difficult to fake, a high-frequency acoustic encoder is deployed to extract physiological non-stationary features such as fundamental frequency jitter, amplitude shimmer, and formant drift from the speech waveform. Simultaneously, a bidirectional long short-term memory (Bi-LSTM) network with self-attention mechanism is connected in parallel to extract the subconscious negative semantic topology of discrete text sequences in the latent semantic space. After orthogonal concatenation of the acoustic and semantic latent features, the implicit emotional valence is mapped and output. ) and implicit arousal ( ).

[0019] 3. Nonlinear dynamic conflict quantification engine for cross-modal emotion representation space

[0020] The core innovation of this invention lies in abandoning the traditional linear fusion mechanism and instead constructing a nonlinear conflict metric model based on adaptive Mahalanobis distance and temporal dynamics gradient. This engine can accurately capture and quantify polarity mutual exclusion and dynamic divergence in the feature space.

[0021] Step S130-A: Construct cross-modal joint representation and Markov space deviation.

[0022] Since explicit features (facial AU activation) and implicit features (voice perturbation frequency domain values) are heterogeneous data, directly calculating the Euclidean distance is susceptible to interference from dimensions and variance. This system defines the features within the time window t as a two-dimensional column vector: explicit vector... Implicit vector .

[0023] First, the system calculates the inverse of the joint covariance matrix. Eliminate linear correlations between features and calculate instantaneous spatial deviation. :

[0024]

[0025] in, The covariance inverse matrix, obtained through pre-training with large-scale baseline data, can adaptively apply a nonlinear penalty to the deviation between valence and arousal.

[0026] Step S130-B: Introduce the emotional evolution gradient difference in cognitive dynamics

[0027] Patients with smiling depression expend a significant cognitive load maintaining a "social mask," resulting in a mechanical (low rate of change) maintenance of their overt facial expressions, while implicit physiological indicators fluctuate dramatically (high rate of change) due to psychological stress. Since facial motor unit activation and acoustic perturbations are typical high-frequency, high-noise signals with physiological hysteresis memory effects, this system abandons the traditional first-order partial derivatives and introduces the Grünwald-Letnikov fractional differential operator. Calculate the time-series dynamic gradient difference :

[0028]

[0029] In the formula, Fractional order of differential order (preferred) This fractional-order operator constitutes an intrinsic low-pass attenuation and long-range memory filter, precisely quantifying the "evolutionary speed conflict" between the two modalities in the time dimension, that is, when "facial expression stagnation" and "abrupt change in speech emotion" occur simultaneously. This will result in a surge.

[0030] Step S130-C: Generate an emotional conflict index with time-decaying memory.

[0031] Integrating spatial divergence and dynamic gradient, the system defines the instantaneous integrated conflict energy. To eliminate transient noise and capture the cumulative effect of concealment behavior over time, the system calculates the final Conflict Index (CI) using an integral model with an exponential decay factor within a sliding window T.

[0032]

[0033] in, As a time-forgetting decay factor, conflicting data that is closer to the current moment is given higher weight. The generalized differentiable directional polarity triggering function is constructed using the Softplus smooth approximation algorithm:

[0034]

[0035] In the formula, Soft For smooth approximate activation function; the coefficient of the first term Used to apply asymmetric, specific, and strong gain to the typical masking pattern of "explicit valence tending to be positive and implicit valence tending to be negative"; the second term coefficient By utilizing the cosine distance in the feature column vector space, a generalized penalty basis is provided for cross-modal sentiment polarity divergence. This mechanism achieves exponential nonlinear gain amplification of sentiment concealment behavior while ensuring the global continuity of the backpropagation gradient.

[0036] 4. Risk decision-making module based on causal time series tracking and Bayesian adaptive confidence boundary

[0037] The system abandons the traditional static threshold determination mechanism, which is prone to false alarms, and deploys a long short-term memory sequence tracking architecture that integrates a temporal convolutional network (TCN) and a self-attention mechanism. This module uses the sentiment conflict index (CI) sequence within a specific time period and the baseline hidden valence moving average. As feature input, causal convolution is used to ensure that future information is not leaked into the prediction of the current state. At the same time, the system incorporates an adaptive dynamic confidence boundary model based on Bayesian state estimation. The network outputs the smiling depression risk assessment result through the fully connected layer only when the continuous fluctuation trajectory of the CI sequence breaks through the dynamic confidence upper bound and the global baseline valence of the implicit branch drifts to the severely negative nonlinear trap interval.

[0038] Compared with the prior art, the present invention has the following significant advantages:

[0039] 1. This invention overcomes the fundamental limitations of the traditional multimodal network's "modal consistency assumption." Existing technologies focus on feature fusion between modalities, while this invention pioneers a detection paradigm based on the "feature polarity mutual exclusion" metric. Through physically isolated dual-path extraction networks, it forcibly decouples explicit features controlled by somatic nerves from implicit features controlled by autonomic nerves, completely resolving the algorithmic flaw of traditional attention mechanisms being easily "hijacked" and misled by high-confidence social spoofing signals.

[0040] 2. An anti-interference spatiotemporal dynamic model for masking high-noise features was constructed. Addressing the technical bias of existing systems being susceptible to high-frequency acquisition noise interference when extracting weak physiological leakage features, this invention innovatively introduces a first-order dynamic evolution gradient difference with an embedded Gaussian smoothing kernel. This model accurately quantifies the temporal divergence between "facial micro-expression stiffness (low evolution rate) and physiological involuntary micro-tremors (high evolution rate)," greatly improving the temporal sensitivity for capturing complex psychological camouflage features compared to traditional static distance metrics.

[0041] 3. An end-to-end differentiable emotional polarity penalty mechanism is implemented, demonstrating excellent engineering feasibility. Unlike existing multimodal systems that rely on manually set static thresholds and are therefore vulnerable, this invention introduces a directional polarity triggering function based on the Softplus smoothing approximation algorithm into the integral model. This mechanism applies an exponential nonlinear gain only to the typical masking pattern of "explicit polarity tending to be positive and implicit polarity tending to be negative" while ensuring the global continuity of the backpropagation gradient. This effectively avoids neuron death during deep network training and gives the system extremely high industrial applicability as an end-to-end clinical auxiliary diagnosis and treatment workstation. Attached Figure Description

[0042] Figure 1 This is an overall module architecture diagram of a depression detection system based on multimodal emotional conflict quantification provided in an embodiment of the present invention.

[0043] Figure 2A flowchart illustrating a depression detection method based on multimodal affective conflict quantification provided in this embodiment of the invention (corresponding to steps S110-S140 in the claims).

[0044] Figure 3 A schematic diagram of a dual-path feature decoupling extraction network architecture driven by somatic and autonomic nervous systems provided in an embodiment of the present invention (showing the neural network flow of video stream and audio / text stream).

[0045] Figure 4 A schematic diagram of the cross-modal emotional representation spatial conflict quantification calculation model provided in the embodiments of the present invention (showing the valence-arousal VA two-dimensional coordinate system and the divergence vector).

[0046] Figure 5 The system physical hardware deployment and data anonymization flow topology diagram based on the edge-cloud collaborative architecture provided in this embodiment of the invention.

[0047] Figure 6 This is a schematic diagram of the front-end polarity radar chart rendering logic and visualization dashboard provided in an embodiment of the present invention. Detailed Implementation

[0048] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0049] The specific embodiments described herein are merely illustrative of the invention and are not intended to limit the scope of protection of the invention. Those skilled in the art will recognize that various modifications and improvements can be made to the invention without departing from its principles, and these modifications and improvements also fall within the scope of protection of the claims.

[0050] 1. System overall architecture and spatiotemporal reference isomorphism processing

[0051] Reference Figure 1 and Figure 2 This invention provides a depression detection system based on multimodal emotional conflict quantification. At the physical level, this system can operate on a computing device equipped with a high-performance graphics processing unit (GPU, with a computing power of at least 10 TFLOPS).

[0052] Step S110: Perform high-frequency synchronous sensing and benchmark alignment of multi-source heterogeneous data.

[0053] In a specific implementation, the multi-source heterogeneous data high-frequency synchronous sensing module (101) is connected to both a visual sensor and an acoustic sensor. The visual sensor is preferably a high-definition wide-angle camera with a resolution of at least 1080P and a frame rate set between 60fps and 120fps; the acoustic sensor is preferably an omnidirectional array microphone with a sampling rate set between 16kHz and 48kHz.

[0054] To eliminate the non-stationary phase delay caused by heterogeneous sensors during data acquisition and bus transmission, this module introduces a Dynamic Time Warping (DTW) algorithm and a multidimensional timing interpolator. Specifically, the system sets a global time sliding window T (preferably ranging from 1.0s to 3.0s, 1.5s in this embodiment), with a step size of... (Preferred time: 0.5s). Using the timestamp of the acoustic signal as the absolute reference, cubic spline interpolation or keyframe resampling is performed on the visual frame sequence to map the discrete multi-source signals into a reference tensor sequence that is strictly isomorphic in time scale. This ensures that the phase space error during subsequent differentiation calculations is close to zero.

[0055] 2. Decoupling and Representation of Orthogonal Features in Dual Neural Drives

[0056] Reference Figure 3 In step S120, the somatic-autonomic neural-driven feature decoupling extraction module (102) abandons the traditional early feature cascade (Early Fusion) method and constructs a physically isolated dual-path orthogonal feature extraction network:

[0057] First branch (Explicit social disguise branch):

[0058] This branch is used to extract facial spatiotemporal features strongly controlled by somatic nerves and subjective consciousness. The input is a visual baseline tensor, and the main body of the network uses a lightweight 3D convolutional neural network (such as 3D-ResNet18). Multiple binary classification support vector machines (SVMs) or fully connected layers are connected in parallel at the network ends to quantify the transient activation intensity of specific facial action units (AUs). This embodiment focuses on extracting feature maps of the zygomaticus major muscle (AU12) corresponding to the upward movement of the corners of the mouth and the orbicularis oculi muscle (AU6) corresponding to the contraction of the corners of the eyes. After the extracted features are nonlinearly mapped by the ReLU activation function and the Softmax layer, the output is a two-dimensional explicit feature column vector of the subject:

[0059]

[0060] In the formula, Indicates the valence of the substance (1 represents extreme pleasure). This indicates explicit wakefulness.

[0061] The second branch (the implicit physiological leakage branch):

[0062] This branch is used to extract non-stationary physiological stress responses that are controlled by the hypothalamic-pituitary-adrenal (HPA) axis and are extremely difficult to artificially simulate. The inputs are an acoustic reference tensor and a semantic tensor.

[0063] First, the acoustic encoder performs a short-time Fourier transform (STFT, with a Hamming window length of 25ms and a frame shift of 10ms) on the acoustic signal to extract the fundamental frequency jitter, amplitude jitter, and Mel frequency cepstral coefficients (MFCC).

[0064] Secondly, the semantic tensor is input into a bidirectional long short-term memory network (Bi-LSTM) configured with a self-attention mechanism. The system maps the discrete text vectors output from the hidden layer of the network to the distribution space of a pre-constructed clinical psychology emotion lexicon (such as the LIWC lexicon), calculates the projection distance of the text sequence on the negative emotion dimension (such as anxiety, sadness, cognitive rigidity), and thereby extracts the subconscious level of negative semantic topological weights.

[0065] After orthogonally concatenating the acoustic and semantic latent features, they are mapped to a two-dimensional implicit feature column vector through a multilayer perceptron (MLP):

[0066]

[0067] In the formula, and The range of values ​​is [-1, 1].

[0068] 3. Cross-modal nonlinear dynamic conflict quantization calculation model

[0069] Reference Figure 4 In step S130, the cross-modal emotion representation space conflict quantification engine (103) calculates the conflict parameters according to the following formula:

[0070] (1) Adaptive Mahalanobis spatial deviation calculate:

[0071] To eliminate the dimensional differences and linear correlation between valence and arousal in heterogeneous feature spaces, an adaptive Mahalanobis distance model is used to construct the divergence degree model:

[0072]

[0073] In the formula, It is the joint covariance inverse matrix.

[0074] In this embodiment, initial matrix It is not set arbitrarily, but is calculated based on a pre-constructed clinical baseline dataset for depression. The baseline dataset is constructed by collecting synchronous high-definition facial video streams and physiological audio streams of several subjects during structured clinical interviews (such as Hamilton Depression Rating Scale (HAMD) inquiries), and extracting the corresponding explicit and implicit feature column vector distributions.

[0075] To further address the covariance matrix drift problem caused by individual physiological expression differences among subjects, the system does not rely on a fixed [variable] during operation. Instead, it employs the exponential moving average (EMA) algorithm, combined with the newly extracted multimodal feature covariance matrix within the current time sliding window T. Online fine-tuning is performed. The adaptive weight update formula is as follows:

[0076]

[0077] In the formula, The dynamic learning rate momentum coefficient is defined as [0.90, 0.99] (preferred range). This closed-loop mechanism enables the system to adaptively penalize the drift of the masking feature distribution of specific subjects in real time, thereby significantly improving the robustness of polarity deviation detection.

[0078] (2) Gradient difference in nonstationary temporal dynamics evolution based on fractional calculus calculate:

[0079] Since facial AU activation and acoustic jitter perturbations are typical high-frequency, high-noise signals with physiological hysteresis memory effects, traditional integer-order differential calculations would exponentially amplify the high-frequency noise and lose long-range memories of emotional evolution. Therefore, this system abandons the traditional first-order partial derivatives and introduces the Grünwald-Letnikov fractional-order differential operator. Construct a noise-resistant and historically memory-based dynamic gradient model:

[0080]

[0081] In the formula, The fractional derivative order (preferred interval is) This fractional-order operator constitutes an intrinsic low-pass attenuation and long-range memory filter, which can accurately extract the low-frequency evolution rate conflict of heterogeneous modes on the time manifold. It fits the viscoelastic dynamic characteristics of masked micro-expression stiffness (low-order variation) and physiological involuntary tremor (high-frequency variation) from the underlying mathematical mechanism.

[0082] (3) Calculation of the end-to-end differentiable time-decaying memory-emotional conflict index CI:

[0083] Integrating the above spatial and dynamic characteristics, instantaneous comprehensive energy is defined. Within a sliding window T, an exponentially decaying integral model is used to output a continuous conflict exponent:

[0084]

[0085] To ensure the global continuity of gradients during end-to-end backpropagation and to overcome the limitations of a single emotion masking mode, the system constructs a generalized differentiable polar triggering function based on the cosine of the vector angle and asymmetric polar amplification. :

[0086]

[0087] In the formula, For smooth approximate activation function; the coefficient of the first term Used to apply asymmetric, specific, and strong gain to the typical smiling depression pattern of "overt laughter and underlying sadness"; the second term coefficient Using the cosine distance in the eigenvector space ( This provides a generalized penalty basis for emotional orientation deviations (including arousal conflicts) occurring in any quadrant. This logic gate, while maintaining global conflict capture capabilities, avoids the neuron death phenomenon observed during deep network training.

[0088] 4. Causal and temporal risk decision-making and front-end interaction implementation scenarios

[0089] In step S140, the adaptive risk decision-making module (104) based on causal time series receives serialized feature input, including the emotional conflict index CI sequence and implicit feature column vector.

[0090] Specific implementation scenario: Clinical auxiliary diagnosis and full-stack front-end interaction

[0091] Combined with reference Figure 5The diagram shows the physical hardware deployment and data anonymization flow topology of the system based on the edge-cloud collaborative architecture. This invention employs a strict edge-cloud isolation mechanism in its physical deployment. Specifically, the multi-source data perception alignment module (101), the dual neural feature decoupling extraction module (102), and the cross-modal emotional conflict quantification engine (103) are all deployed in the intelligent terminal (such as an edge computing device) at the clinic end. To avoid the risk of leakage of sensitive biometric data, the intelligent terminal extracts one-dimensional features such as the conflict index CI from its local volatile memory (RAM) and immediately physically overwrites and destroys the original video pixels and audio waveforms. Subsequently, the intelligent terminal only sends the encrypted one-dimensional anonymized sequence (CI, Gtime, Vi) to the private cloud cluster through the network transmission channel. The causal time series and adaptive risk decision module (104), deployed in the private cloud cluster, receives the anonymized sequence to calculate the posterior risk probability and ultimately connects the decision result with the enterprise EAP / HIS database and outputs an alarm. This architecture completely cuts off the cloud link of the original data from a physical perspective.

[0092] In this implementation scenario, the system's backend decision network employs a Temporal Convolutional Network (TCN). TCN utilizes dilated causal convolution to extract long-range temporal dependencies of the conflict index (CI) sequence, preventing future information leakage. Simultaneously, the system incorporates an adaptive dynamic confidence boundary model based on a Gaussian distribution. Let the mean conflict index over the past N time windows at time t be... The standard deviation is Then the formula for the confidence upper bound Th(t) calculated dynamically by the system is:

[0093]

[0094] In the formula, k is the confidence level adjustment coefficient (preferred). (corresponding to a 99.7% confidence interval).

[0095] To effectively filter out false positives caused by brief emotional fluctuations, the system introduces an implicit valence moving average. As a joint triggering constraint boundary, its calculation formula is:

[0096]

[0097] In the formula, M is the number of time windows. The system strictly sets the mathematical boundary of the preset severe negative interval to [-1.0, -0.6].

[0098] If and only if the TCN network determines that CI(t) > Th(t) at the current time (i.e., the dynamic confidence upper bound is exceeded), and the calculated moving average value is... The system will only activate the final posterior risk probability output for depression (confidence > 95%) when the value falls into the severely negative range [-1.0, -0.6]. At this point, combined with the reference... Figure 6 The diagram shown illustrates the front-end polarity radar chart rendering logic and visualization dashboard. The system will simultaneously trigger the following full-stack front-end interactive actions:

[0099] (1) Bottom data push: After completing the risk assessment of the current time window T, the system backend uses WebSockets technology to push the warning status signal (Payload) containing key parameters such as conflict index (CI), dynamic threshold (Th), explicit valence (Ve) and implicit valence (Vi) to the front-end interactive interface in milliseconds.

[0100] (2) Multidimensional feature polarity radar chart rendering: After the front-end interface parses the streaming data, it automatically renders and generates the core visual conflict dashboard. In the polygon radar chart grid, the points corresponding to explicit positivity (Ve) are drawn in the outer high-threshold area, and the points corresponding to implicit negativity (Vi) are drawn in the inner low-threshold area. Through extremely asymmetrical polygonal lines, the emotional polarity deviation of the subjects' 'explicit positivity and implicit negativity' is intuitively presented.

[0101] (3) Dual-condition trigger warning pop-up: A pathological depression risk warning pop-up (fatal red line) is rendered on the side of the radar chart. The pop-up clearly indicates the dual triggering condition logic that has been met concurrently. At the same time, the system automatically anchors and latches the multimodal original data slice matrix that triggered the high-risk polarity inversion time period in the background database to form an immutable pathological objective evidence chain to assist clinicians in accurate review.

Claims

1. A depression detection system based on multi-modal affective conflict quantification, characterized in that, include: The multi-source heterogeneous data high-frequency synchronous sensing module is used to collect facial video streams and physiological audio streams of subjects in parallel during the interaction process, and to convert the audio streams into semantic text streams synchronously through an automatic speech recognition (ASR) engine, thereby performing spatiotemporal reference isomorphic alignment of multimodal signals. The feature decoupling extraction module driven by the somatic autonomic nervous system is used to forcibly split the synchronized data stream into two physically isolated network branches, extracting the spatial distribution and dynamic temporal features of muscle action units in the facial region of interest as explicit social masking features, and extracting the high-frequency jitter features of the acoustic spectrum and the negative semantic weights of the text in the audio stream as implicit physiological leakage features. A cross-modal emotion representation space conflict quantification engine is used to map the explicit social masking features and implicit physiological disclosure features to a unified two-dimensional emotion coordinate system, calculate the Mahalanobis spatial divergence degree and temporal dynamic gradient of the two in terms of emotion polarity, and output a continuous emotion conflict index (CI) sequence. The causal time-series-based adaptive risk decision-making module combines the fluctuation trajectory of the emotional conflict index sequence with the baseline emotional state, and outputs the posterior risk probability of depression through a temporal convolutional prediction network and an adaptive dynamic confidence boundary model based on Gaussian distribution.

2. The depression detection system based on multimodal affective conflict quantification according to claim 1, characterized in that: The multi-source heterogeneous data high-frequency synchronous sensing module introduces a dynamic time warping (DTW) algorithm and a timing interpolator to perform hard alignment of the video frames and text word vectors to be processed, using the timestamps of the acoustic signals as the absolute reference; and sets the length to T and the step size to... The global time sliding window maps discrete multi-source signals into a reference tensor sequence that is strictly isomorphic in time scale.

3. The depression detection system based on multimodal affective conflict quantification according to claim 1, characterized in that, The somatic-autonomic neural drive feature decoupling extraction module includes an explicit masking branch and an implicit leakage branch; The explicit masking branch includes a lightweight 3D convolutional neural network for extracting facial regions of interest from the input video stream and extracting the spatiotemporal receptive field features of action units (AUs) based on 3D convolutional kernels. It focuses on quantizing the transient activation intensity corresponding to the amplitude of mouth slant (AU12) and eye corner contraction (AU6), and outputs the explicit valence via a nonlinear mapping. ) and explicit arousal ( ), construct the explicit feature column vector E(t); The implicit leakage branch includes an acoustic perturbation encoder and a deep semantic extraction network. The acoustic perturbation encoder uses short-time Fourier transform to extract the fundamental frequency perturbation (jitter) and amplitude perturbation (shimmer) of the audio stream. The deep semantic extraction network uses a bidirectional long short-term memory network (Bi-LSTM) with self-attention mechanism to extract subconscious negative semantic topological weights by mapping text word vectors to a pre-trained clinical psychology sentiment lexicon space. The acoustic and semantic hidden features are orthogonally concatenated to output the implicit valence. ) and implicit arousal ( ), construct the implicit feature column vector I(t).

4. The depression detection system based on multimodal affective conflict quantification according to claim 3, characterized in that: The conflict quantification engine in the cross-modal emotion representation space calculates conflict parameters in a two-dimensional valence-arousal coordinate system using the following model: Introducing a pre-trained joint covariance inverse matrix During operation, the exponential moving average algorithm combined with newly added multimodal features is used for online fine-tuning to calculate the adaptive Mahalanobis spatial deviation. ; Introducing fractional differential operators The fractional order The temporal dynamic evolution gradient difference is calculated by performing dynamic differentiation on the explicit feature column vector E(t) and the implicit feature column vector I(t) while taking into account long-range memory and high-frequency noise suppression. ; The and The weighted fusion is the instantaneous comprehensive energy C(t).

5. The depression detection system based on multimodal affective conflict quantification according to claim 4, characterized in that: The conflict quantification engine in the cross-modal emotion representation space further generates the emotion conflict index CI sequence through an integral model with an exponential decay factor. In the integral calculation, a time forgetting decay factor is introduced to give higher weight to conflicting data that are close to the current time t. Simultaneously, a generalized differentiable polarity triggering function is introduced. The trigger function is constructed using the Softplus smooth approximation algorithm. Internally, it consists of an asymmetric penalty term for typical masking polarity reversal and a global deviation penalty term based on the cosine of the angle between the feature column vectors. Under the premise of ensuring the continuity of the backpropagation gradient, it achieves generalized capture and exponential amplification of multi-quadrant cross-modal emotional conflicts.

6. The depression detection system based on multimodal affective conflict quantification according to claim 1, characterized in that: The adaptive risk decision-making module based on causal time series employs a temporal convolutional network (TCN). Its input layer contains the sentiment conflict index (CI) sequence of the N consecutive sliding windows preceding the current time, and the implicit valence moving average over the past M time windows. (i.e., baseline emotional state); The temporal convolutional network utilizes dilated causal convolution to extract long-range temporal dependencies and simultaneously calculates the moving mean and standard deviation of the CI sequence to construct a Gaussian dynamic confidence upper bound; when it is determined that the CI sequence continuously exceeds the dynamic confidence upper bound, and the moving mean... When the data falls into the preset severely negative range, a high-risk alert is triggered, and the multimodal raw data slice matrix that caused the polarity reversal is automatically latched.

7. A method for detecting depression based on multimodal affective conflict quantification, characterized in that, Using the system as described in any one of claims 1 to 6, the method includes the following steps: S110 acquires video and audio streams of subjects in parallel through multi-source sensing modules, extracts semantic text by combining with the ASR engine, and performs tensor alignment based on dynamic time warping. S120 utilizes the explicit masking branch to extract transient activation features of facial muscles controlled by somatic nerves and maps them to explicit feature vectors; simultaneously, it utilizes the implicit leakage branch to extract speech perturbations and negative semantic topology controlled by autonomic nerves and maps them to implicit feature vectors. S130, the explicit and implicit feature vectors are projected onto the emotional space, the Mahalanobis space deviation degree and temporal dynamic gradient are calculated, and a dynamic emotional conflict index sequence is generated through a decay integral model with a directional polarity triggering function. S140, the conflict index sequence is input into a temporal convolutional network, combined with the Bayesian dynamic confidence boundary and the baseline implicit valence, to assess the posterior risk probability, and output the smiling depression risk assessment result and interpretable data slices when the threshold is exceeded.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the depression detection method based on multimodal affective conflict quantification as described in claim 7.