Multimodal depression risk fusion screening method and system for adolescent group

CN122531772APending Publication Date: 2026-08-07INST OF HEALTH & MEDICINE HEFEI COMPREHENSIVE NAT SCI CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF HEALTH & MEDICINE HEFEI COMPREHENSIVE NAT SCI CENT
Filing Date
2026-05-14
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0008]本发明要解决现有通用多模态筛查技术在应用于12至18岁青少年时,因忽略该群体特有的生理发育差异性、高情绪掩饰率及积极情绪保护价值而导致筛查假阳性率偏高、准确性不足的技术问题

Benefits of technology

本发明根据受试者年龄动态调整人脸检测窗口尺寸和语音质量评估中频率微扰、振幅微扰与谐噪比的权重组合,从信号采集源头消减了青春期变声与面部发育变化引入的生理噪声,提升了表情与语音特征的检出准确率和抗干扰能力;同时,通过在融合判定阶段将正性表情比例作为保护因子并赋予负权重,纠正了传统方法仅累积病理性指标所带来的负性偏差,有效降低了将心理韧性较强的正常个体误判为高风险的概率;此外,结合对微表情事件的精细化检测与跨模态交叉验证,能够主动识别并标记青少年群体中普遍存在的情绪掩饰行为,进一步保障了筛查结果的精准性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531772A_ABST
    Figure CN122531772A_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal depression risk fusion screening method and system for a youth group, relates to the technical field of psychological measurement, and comprises the following steps: acquiring facial video, speech audio and psychological scale answering data of a youth subject; dividing the subject into age groups according to the age of the subject, and differentially configuring corresponding facial analysis parameters and speech analysis parameters according to the divided age group differences; extracting facial expression features based on the facial analysis parameters corresponding to the age group, extracting speech features based on the speech analysis parameters corresponding to the age group, and calculating scale scores according to the psychological scale answering data; constructing a fusion feature vector containing the facial expression features, the speech features and the scale scores, linearly weighting and fusing the fusion feature vector to obtain a fusion score, wherein the positive expression proportion feature in the fusion feature vector is given a negative weight; and determining the depression risk level of the subject according to the fusion score. The accuracy of the screening result is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychometrics, and in particular to a multimodal depression risk fusion screening method and system for adolescents. Background Technology

[0002] With the development of artificial intelligence technology, multimodal depression risk identification technology based on computer vision and speech signal processing has made a series of advances. This type of technology typically analyzes behavioral signals such as facial expressions and speech rhythm of subjects, combined with machine learning models, to achieve an objective assessment of depressive state, thus alleviating the limitations of traditional scale assessments to some extent.

[0003] However, when existing general multimodal screening technologies are directly applied to adolescents aged 12 to 18, several shortcomings still exist due to the unique physiological development and psychological and behavioral characteristics of this group: 1. Failure to adapt to age-related developmental differences. The period from 12 to 18 years old is a period of dramatic change in the development of facial bones and vocal cords. There are significant differences in facial contours and vocal cord structures among adolescents of different age groups. When dealing with rapidly developing adolescents, the commonly used facial detection window and speech analysis parameters will lead to a decrease in facial feature recognition rate and will confuse the abnormal acoustic indicators such as frequency perturbations and amplitude perturbations caused by physiological voice changes during puberty with the speech signal characteristics of depressive symptoms.

[0004] 2. Insufficient detection of prevalent emotional concealment behavior among adolescents. Research indicates that adolescents exhibit a high degree of social desirability bias during socialization, often tending to conceal their true negative emotions and present a positive image during psychological screening. This emotional concealment behavior causes screening systems based on conventional facial expression analysis to receive distorted emotional signals, resulting in a large number of false negative results.

[0005] 3. Insufficient differentiation between physiological voice changes during puberty and depressive speech abnormalities makes it easy for screening results to misjudge individuals with normal emotional fluctuations or positive psychological resilience as high risk, leading to a high false positive rate.

[0006] 4. The lack of a control mechanism for the quality of existing collected data leads to a high false positive rate in uncontrolled environments.

[0007] In summary, there is an urgent need for a multimodal depression risk screening program specifically designed for adolescents that can effectively overcome differences in their physiological development, emotional masking noise, and data quality interference, in order to achieve large-scale depression risk screening with high accuracy and low false positives. Summary of the Invention

[0008] This invention aims to address the technical problem that existing general multimodal screening technologies, when applied to adolescents aged 12 to 18, suffer from high false positive rates and insufficient accuracy due to neglecting the unique physiological developmental differences, high emotional masking rate, and positive emotional protection value of this group.

[0009] In a first aspect, to address the aforementioned technical problems, this invention provides a multimodal depression risk fusion screening method for adolescents, comprising the following steps: Acquire facial video, audio recordings, and psychological scale responses from adolescent subjects; The subjects were divided into age groups based on their age, and the corresponding facial analysis parameters and voice analysis parameters were configured according to the differences in the age groups. Facial expression features are extracted based on the facial analysis parameters corresponding to the age group, speech features are extracted based on the speech analysis parameters corresponding to the age group, and the scale score is calculated based on the psychological scale response data. A fusion feature vector containing the facial expression features, the voice features, and the scale score is constructed, and the fusion feature vector is linearly weighted and fused to obtain a fusion score; wherein, the positive expression proportion feature in the fusion feature vector is given a negative weight; The subject's depression risk level was determined based on the fusion score.

[0010] Furthermore, the speech features include speech prosody features and speech quality features, wherein: The prosodic features include fundamental frequency, speech rate, and pause ratio; The speech quality characteristics include frequency perturbation, amplitude perturbation, and harmony-to-noise ratio.

[0011] Furthermore, the extraction of the speech quality features includes: Calculate the frequency perturbation, amplitude perturbation, and harmonic-to-noise ratio of the speech audio; Based on the age group, the weights of the frequency perturbation, the amplitude perturbation, and the harmonic-to-noise ratio (HNR) are differentiated in calculating the overall speech quality score, such that the HNR weight corresponding to the age group in puberty is higher than the HNR weight corresponding to other age groups.

[0012] Furthermore, the step of extracting facial expression features based on the facial analysis parameters corresponding to the age group also includes: Micro-expression events with a duration of less than a preset threshold are detected from the facial video to obtain the number and intensity of negative micro-expressions; Based on the number of negative micro-expressions, the intensity of the negative micro-expressions, and the self-reported emotional state derived from the scale scores, it is comprehensively determined whether the subject has emotional concealment behavior, and the determination result is used to construct a feature of the fused feature vector.

[0013] Furthermore, the step of detecting micro-expression events from the facial video with a duration less than a preset threshold also includes: The facial expression sequence frames of the video are traversed according to a sliding time window; Within each window, determine whether there is a frame sequence in three consecutive frames where the intensity of negative facial expressions satisfies the low-high-low peak pattern; If it exists, calculate the duration of the peak pattern; If the duration is less than 200ms, it is recorded as a micro-expression event.

[0014] Furthermore, the comprehensive determination of whether the subject has engaged in emotional concealment behavior is specifically determined when the following three conditions are met simultaneously: The number of detected negative micro-expression events exceeds a preset threshold; The average intensity of the negative micro-expression events exceeds a preset intensity threshold; The proportion of subjective positive emotion reports derived from the scale's response data exceeded a preset threshold.

[0015] Furthermore, the method also includes: The overall confidence level of the current screening task is calculated based on the availability or quality indicators of the facial video, the audio, and the psychological scale response data.

[0016] Furthermore, it also includes adaptively adjusting the confidence level, specifically: Initialize baseline confidence; For each available data mode, a preset basic contribution value corresponding to that mode is accumulated on the baseline confidence level; The accumulated confidence scores are multiplied by a quality factor that reflects the quality of the modality data to implement punitive reduction.

[0017] Furthermore, the step of dividing the subjects into age groups based on their age, and configuring corresponding facial analysis parameters and voice analysis parameters according to the differences in the divided age groups, specifically includes: Subjects aged 12 to 14 were matched to early puberty and assigned the first minimum window size for face detection and the first combination of speech analysis weights. Subjects aged 15 to 16 were matched to mid-adolescence and assigned a second face detection minimum window size larger than the first face detection minimum window size and a second speech analysis weight combination. The sum of the weights of frequency perturbation and amplitude perturbation in the second speech analysis weight combination was lower than the sum of the corresponding weights in the first speech analysis weight combination. Subjects aged 17 to 18 were matched to late adolescence and assigned a third face detection minimum window size and a third speech analysis weight combination that was larger than the second face detection minimum window size.

[0018] Further, the calculation of the scale score based on the responses to the psychological scale includes: Obtain paired scale data before and after the emotional triggering; Calculate the change in a specified dimension of the paired scale data; Using a preset correction coefficient, the change is superimposed on the original scale score to obtain the scale score used to construct the fusion feature vector.

[0019] Furthermore: The facial analysis parameters include the minimum window size for face detection; The speech analysis parameters include the weights of frequency perturbation, amplitude perturbation, and harmony-to-noise ratio in calculating the overall speech quality score.

[0020] Furthermore, the facial expression features include the proportion of negative expressions, the proportion of positive expressions, and the intensity of at least one facial movement unit selected from the group consisting of lowered eyebrows, clenched corners of the mouth, and closed eyes.

[0021] Further, the step of extracting facial expression features based on the facial analysis parameters corresponding to the age group includes: Detect the face region and scale it to a preset pixel size; The scaled-down face area is divided into four regions of interest: eyebrows, eyes, cheeks, and mouth. Based on the region of interest, the intensity of the specified facial motion unit is calculated.

[0022] Further, the calculation of the intensity of the specified facial motion units specifically includes: The intensity reduction between the eyebrows is obtained by calculating the vertical gradient of the eyebrow region; The corner compression strength of the mouth is obtained by calculating the curvature change of the lower edge of the mouth region; The eye closure intensity is obtained by calculating the pixel ratio of the upper and lower halves of the eye region.

[0023] Further, determining the subject's depression risk level based on the fusion score includes: The fusion score is mapped to a depression risk probability using a mapping function; The probability of depression risk is compared with multiple preset risk level thresholds to determine the corresponding risk level.

[0024] A second aspect of the present invention provides a multimodal depression risk fusion screening system for adolescents that implements the method, comprising: The data acquisition module is used to acquire facial video, voice audio, and psychological scale responses from subjects aged 12 to 18. An age group recognition module is used to classify the subject into one of a number of preset age groups based on the subject's age, and to configure facial analysis parameters and voice analysis parameters differently based on the classified age group. The facial analysis module is used to extract facial expression features from the facial video based on facial analysis parameters corresponding to the age group. The speech analysis module is used to extract speech prosodic features and speech quality features from the speech audio based on speech analysis parameters corresponding to the age group. The scale scoring module is used to calculate the scale score based on the response data of the psychological scale; and The feature fusion and risk assessment module is used to construct a fusion feature vector that includes the facial expression features, the speech prosody features, and the scale score; to perform linear weighted fusion on the fusion feature vector to obtain a fusion score; and to determine the subject's depression risk level based on the fusion score.

[0025] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: This invention dynamically adjusts the face detection window size and the weighted combination of frequency perturbation, amplitude perturbation, and harmonic-to-noise ratio in voice quality assessment based on the subject's age. This reduces physiological noise introduced by puberty voice changes and facial development at the signal acquisition source, improving the detection accuracy and anti-interference ability of facial expression and voice features. Simultaneously, by using the proportion of positive expressions as a protective factor and assigning it a negative weight in the fusion judgment stage, it corrects the negative bias caused by traditional methods that only accumulate pathological indicators, effectively reducing the probability of misjudging psychologically resilient, normal individuals as high-risk. Furthermore, by combining refined detection of micro-expression events with cross-modal cross-validation, it can proactively identify and label prevalent emotional masking behaviors among adolescents, further ensuring the accuracy of screening results. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of the overall process for multimodal depression risk screening in adolescents disclosed in this invention; Figure 2 This is a flowchart of the age group recognition and facial parameter configuration disclosed in this invention; Figure 3 This is a flowchart of the age group recognition and voice parameter configuration disclosed in this invention; Figure 4 This is a flowchart of the facial motion unit extraction process disclosed in this invention; Figure 5 This is a flowchart of the micro-expression detection process disclosed in this invention; Figure 6 This is a flowchart of the emotion concealment determination method disclosed in this invention; Figure 7 This is a flowchart of the speech quality index extraction and comprehensive evaluation process disclosed in this invention; Figure 8 This is a flowchart of the scale scoring process disclosed in this invention; Figure 9 This is a flowchart of the multimodal feature fusion and risk probability mapping process disclosed in this invention; Figure 10 This is a flowchart of the confidence-adaptive mechanism disclosed in this invention; Figure 11 This is a diagram of the overall architecture disclosed in this invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] This invention aims to provide a multimodal depression risk fusion screening method for adolescents aged 12 to 18, which solves the technical problems of high false positive rate and insufficient accuracy in the existing technology due to ignoring the unique physiological development differences, high emotional masking rate and positive emotion protection value of this group.

[0030] Please see Figure 1The method mainly includes the following steps: Step S0, Data Acquisition In this scheme, the data collection environment can be an uncontrolled environment such as a school classroom or community activity room. In a preferred embodiment, facial video of the subject is captured by a camera at 720p resolution and 30 frames per second; the subject's voice audio is captured by a microphone at a 16kHz sampling rate; simultaneously, the subject's age information is obtained, and the subject is asked to complete a psychological scale. This step outputs the original video frame sequence, original audio segments, age values, and scale response data.

[0031] When capturing facial videos, the camera should be directly facing the subject's face, and the shooting distance should be controlled within the range of 50-80cm to ensure that the facial area occupies no less than one-third of the frame.

[0032] When collecting voice audio, it is advisable to use a unidirectional microphone to reduce interference from environmental noise.

[0033] Psychological scales may be one or more of the following: Self-Assessment Portrait Map (SAM) scale, Childhood Depression Epidemiology Survey (CES-DC) scale, Patient Health Questionnaire-A (PHQ-A) for Adolescents, or Childhood Depression Rating Scale (CDRS-R) revised version. Scale responses can be obtained through data entry from paper questionnaires or by interactive completion on electronic devices.

[0034] Step S1: Age Group Identification This step receives the age value output from step S0 as input and, based on preset age range division rules, divides the subject into one of three age groups: early puberty, middle puberty, or late puberty. Simultaneously, based on the determined age group, it loads the corresponding facial analysis parameter configuration and voice analysis parameter configuration from a predefined configuration table for use in subsequent steps.

[0035] The following combination Figure 2 and Figure 3 The specific age group division rules and parameter configuration schemes in step S1 are explained in detail.

[0036] Studies have found that the period between 12 and 18 years old is a period of dramatic change in the development of facial bones and vocal cords, with significant differences in facial contour size and vocal cord structure among adolescents of different age groups. Using generic face detection window parameters and speech analysis weights would lead to a decrease in detection rates for younger adolescents due to their smaller facial sizes, and would also confuse the abnormal acoustic indicators caused by physiological voice changes during puberty with the speech signal characteristics of depressive symptoms. Therefore, this approach introduces a preprocessing step for age group identification, dynamically matching appropriate technical parameters based on the specific age of the subject.

[0037] Please see Figure 2 This demonstrates the process of age group recognition and facial parameter configuration, which begins with receiving the subject's age value as input. First, age group classification is performed. This embodiment divides the age range of 12-18 years into three age groups: early puberty (12-14 years), middle puberty (15-16 years), and late puberty (17-18 years). This classification is primarily based on the general patterns of facial bone and vocal cord development in adolescents: adolescents aged 12-14 have not yet fully developed facial bones and their facial size is relatively small; 15-16 years enter the peak period of rapid development and voice change; and 17-18 years have facial bones and vocal cords that are essentially at adult levels.

[0038] The facial detection configuration parameters for different age groups are shown in Table 1. Among them, minSize is the detection window, which is the input parameter of the face detector.

[0039] Table 1. Configuration parameters for face detection at different ages

[0040] Specifically: If the subject's age falls between 12 and 14 years old, they are classified into the early adolescence group, and the corresponding face detection configuration parameters for this group are loaded. The minimum detection window size (minSize) is configured to 30×30 pixels. This smaller detection window better matches the relatively small face size of adolescents in this age group, avoiding missed facial areas due to an excessively large detection window. If the subject's age falls within the 15-16 age range, they are classified into the mid-adolescence group, and the corresponding face detection configuration parameters for this group are loaded. The minimum detection window size (minSize) is configured to 40×40 pixels. This window size is larger than that of the early adolescence group to accommodate the rapid growth of facial bones during this stage.

[0041] If the subject's age falls within the 17-18 age range, they are classified into the late adolescence group, and the corresponding face detection configuration parameters for this group are loaded. The minimum detection window size (minSize) is configured to 50×50 pixels. This size is close to commonly used parameters for adult face detection and can effectively cover the faces of adolescents who are approaching maturity.

[0042] If the entered age value exceeds the applicable range of 12-18 years old, an exception message will be thrown, indicating that the age exceeds the applicable range of this method.

[0043] The inventors verified the effectiveness of the aforementioned differentiated configuration through experiments. In an experiment involving 200 adolescents aged 12-18, the face detection rate for each age group was tested under different minimum detection window sizes. The experimental results are shown in Table 2. Table 2. Experimental results of face detection rate under different detection windows.

[0044] Experimental data shows that using the corresponding optimized parameters for each age group can improve the detection rate by 2-5 percentage points. This verifies the rationality and effectiveness of configuring face detection parameters differently according to age groups.

[0045] Please see Figure 3 This demonstrates the process of age group recognition and voice parameter configuration. This process runs parallel to the facial parameter configuration process, and while it also begins with age group segmentation, the output configuration content is the weight parameters for voice analysis. The weight parameters for voice analysis of different age groups are shown in Table 3.

[0046] Specifically: For subjects classified into the early puberty group (12-14 years old) or the late puberty group (17-18 years old), a standard speech analysis weighting configuration was used. In this standard configuration, the weights corresponding to frequency perturbation jitter are... The value is 0.30, which corresponds to the weight of the amplitude perturbation shimmer. The value is 0.30, which corresponds to the weight of the harmonic noise ratio (HNR). The value is 0.40.

[0047] For subjects classified as belonging to the mid-adolescence group (15-16 years old), a voice change-adjusted weighting was used. In this adjustment, the weight corresponding to the frequency perturbation jitter was... The value is reduced to 0.20, and the weight corresponding to the amplitude perturbation shimmer is... The value is reduced to 0.20, and the weight corresponding to the harmonic noise ratio (HNR) is... The value is increased to 0.60.

[0048] The core design principle of the differentiated weighting configuration in Table 3 above is to actively avoid physiological noise introduced during puberty. For adolescents aged 15-16, who are in the peak of puberty, their vocal cords are undergoing rapid physiological changes. Congestion and incomplete closure of the vocal cords can lead to a physiological increase in the values ​​of frequency perturbation and amplitude perturbation, two perturbation-related indicators. This physiological increase has certain similarities in acoustic manifestation to psychomotor retardation caused by depression, both potentially manifesting as increased jitter and shimmer values. If a standard weighting configuration is used, the system may misinterpret the normal physiological fluctuations during puberty as depression-related speech abnormalities, leading to a rise in false positives. By actively reducing the jitter and shimmer weights for puberty subjects from 0.30 to 0.20, the contribution of these two indicators, which are susceptible to physiological noise interference, can be effectively suppressed. Meanwhile, the harmonic-to-noise ratio (HNR) reflects the overall regularity of vocal cord vibration and the concentration of harmonic energy. Although this indicator is also affected by puberty to some extent, the degree of influence is smaller compared to perturbation-type indicators. Furthermore, it has the highest single-indicator discrimination in differentiating depressive symptoms, with an AUC value of 0.72, higher than shimmer's 0.65 and jitter's 0.68. Therefore, increasing the HNR weight of puberty subjects from 0.40 to 0.60 can enhance the contribution of this stable anchoring indicator, effectively compensating for the weighting loss of perturbation-type indicators, thereby reducing false positives caused by physiological noise during puberty without sacrificing the sensitivity of depressive signal detection.

[0049] Step S2, Facial Analysis This step receives the video frame output from step S0 and the facial analysis parameter configuration output from step S1 as input, and performs sub-tasks such as face detection, facial motion unit intensity extraction, expression classification, and micro-expression detection.

[0050] Specifically: First, face detection is performed on video frames based on the minimum window size parameter for face detection corresponding to the age group, and facial regions are extracted.

[0051] Then, the facial region was divided into multiple regions of interest, and the intensity values ​​of 17 facial action coding system (FACS) action units were calculated. Among them, the intensity of five action units that are strongly associated with depression were calculated: glabellar reduction AU4, cheek lifting AU6, corner of mouth stretching AU12, corner of mouth compression AU15, and eye closure AU43.

[0052] Next, facial expressions are classified based on the intensity of the action units, yielding statistical features such as the proportion of positive and negative expressions. Simultaneously, a sliding window peak detection method is used to identify micro-expression events with a duration of less than 200ms, and a list of micro-expression events is output, including the type, duration, intensity, and timestamp of each micro-expression.

[0053] Finally, the output includes the proportion of negative expressions, the proportion of positive expressions, the intensity values ​​of the five key action units, and a list of micro-expression events.

[0054] In this scheme, the facial analysis steps include three core sub-tasks: facial motor unit intensity extraction, micro-expression detection, and emotion masking determination. These three sub-tasks are sequentially linked, forming a complete analysis chain from basic facial movement perception to high-level psychological state interpretation. The following section combines... Figures 4-6 The specific implementation steps of facial analysis in step S2 will be explained in detail.

[0055] Please see Figure 4 It demonstrates the complete processing flow of the facial motion unit extraction workflow, as detailed below: Step S2-0: The input is the face region image face_roi that has been cropped after face detection.

[0056] Step S2-1: Image standardization. This involves uniformly scaling the input facial region image to a fixed size of 100×100 pixels.

[0057] It should be noted that the selection of 100×100 pixels as the standardized size is based on experiments showing that when the size is smaller than 50×50 pixels, the extraction accuracy of fine motion units such as eyebrow reduction (AU4) and corner-of-the-mouth compression (AU15) drops below 65%; when the size is larger than 200×200 pixels, the computational load increases by approximately four times, but the accuracy only improves by about 2 percentage points. Therefore, 100×100 pixels represents the optimal balance between computational efficiency and extraction accuracy.

[0058] Step S2-2: Divide the standardized facial image into four regions of interest by row.

[0059] The division is based on the relative positions of organs in facial anatomy. Specifically, the eyebrow region corresponds to rows 0 to 30 of the image, occupying approximately 30% of the upper face; activation of the glabella reduction AU4 primarily affects the vertical texture changes in this region. The eye region corresponds to rows 30 to 50 of the image, occupying approximately 20% of the facial height; the expression of eye closure AU43 is concentrated in this region. The cheek region corresponds to rows 30 to 55 of the image, occupying approximately 25% of the facial height; activation of cheek lifting AU6 shows a significant positive gradient change in this region. The mouth region corresponds to rows 60 to 100 of the image, occupying approximately 40% of the lower face; several key action units, such as corner stretching AU12 and corner compression AU15, produce observable changes in this region.

[0060] Step S2-3: Implement specific computational strategies for the five action units that are strongly associated with depression, as shown in steps S2-4 to S2-8.

[0061] Step S2-4: Extract the intensity of the AU4 intensity reduction between the eyebrows. The brow reduction is one of the hallmark features of a depressed facial expression, characterized by eyebrows converging and pressing down towards the center of the brow, creating vertical skin wrinkles in the brow area. This embodiment quantitatively characterizes this texture change by calculating the vertical Sobel gradient in the region of interest of the eyebrows. Specifically, a 3×3 Sobel vertical gradient operator is used to perform convolution operations on the eyebrow region to obtain the vertical gradient magnitude at each pixel location. The average gradient magnitude of the entire region is calculated and then divided by a preset gradient normalization factor of 50.0 to obtain the AU4 intensity value. The gradient normalization factor of 50.0 is an empirical parameter determined based on statistics from 200 labeled facial images. In this batch of statistical samples, the average vertical gradient value when AU4 is activated is distributed in the range of 25-75. Dividing by 50.0 effectively maps the intensity value to the range of 0 to 1.

[0062] Step S2-5: Extract the intensity of cheek lift AU6. AU6 is a key component of the Duchenne smile, manifested as an upward bulge in the cheek area caused by cheek muscle contraction. In this embodiment, the vertical Sobel gradient of the cheek area is calculated, and the proportion of positive gradient pixels (i.e., pixels with gradient values ​​greater than 0) is used as a measure of cheek lift. This proportion is then normalized by dividing it by a baseline value of 0.3, and truncated to 1.0 when the result exceeds 1.0. The baseline value of 0.3 is an empirical value based on the proportion of positive gradient pixels when AU6 is fully activated in the labeled data.

[0063] Step S2-6: Extract the intensity of the AU12 stretch at the corners of the mouth. AU12 is also a key component of the Duchenne smile, manifested as the corners of the mouth stretching to both sides. In this embodiment, Canny edge detection is performed on the mouth region, using 50 and 150 as dual threshold parameters. These two parameters are the standard recommended values ​​for the Canny edge detector in the OpenCV library. After detection, the sum of the number of edge pixels in the left and right 10 columns of the mouth region is counted as a measure of the edge activity at the corners of the mouth. This sum is normalized by dividing by 1000.0, and values ​​exceeding 1.0 are truncated to 1.0.

[0064] Step S2-7: Extract the intensity of AU15, the compression of the corners of the mouth. AU15 is highly correlated with depression, manifested as downward compression of the corners of the mouth, causing the lower edge of the mouth to curve downwards in an arc. In this embodiment, this action is characterized by calculating the curvature change of the lower edge of the mouth region. The first and second derivatives at each point of the lower edge are calculated to obtain the curvature value, and then the curvature change is normalized to the range of 0 to 1.

[0065] Step S2-8: Extract the intensity of eye closure AU43. AU43 is an important facial manifestation of psychomotor retardation, often presenting as ptosis and narrowed palpebral fissures in depressed patients. This embodiment quantifies the degree of eye closure by calculating the ratio of the average pixel value of the upper half to the lower half of the eye region. The smaller the ratio, the higher the degree of eye closure. The intensity value is taken as 1.0 minus this ratio, so that the intensity value increases with the increase of the degree of eye closure.

[0066] Step S2-9: Truncate all calculated intensity values ​​to a closed interval of 0-1 to ensure the validity of the values.

[0067] Step S2-10: Output the final sequence of action unit intensity values.

[0068] It should be noted that, among the 17 facial movement units listed in Table 5 below, apart from the 5 movement units that are highly correlated with depression, the remaining 12 movement units, including inner eyebrow raising AU1, outer eyebrow raising AU2, upper eyelid lifting AU5, eyelid tightening AU7, nose wrinkling AU9, upper lip lifting AU10, dimple tightening AU14, chin lifting AU17, lip stretching AU20, lip tightening AU23, lip parting AU25, and chin drooping AU26, have a lower correlation with depression. In this embodiment, a similar regional feature analysis method is used for intensity extraction as an auxiliary feature to participate in expression classification, but its contribution weight to depression screening is relatively small.

[0069] Table 5 List of Facial Motion Units

[0070] After extracting the intensity of each facial motion unit, basic expression classification is performed based on the intensity combinations of 17 facial motion units. The basic expression categories include seven types: happy, sad, angry, fearful, disgusted, surprised, and neutral. Classification can be achieved using rule-based methods or a pre-trained classifier. The proportions of positive and negative expressions can be further statistically analyzed from the classification results. The positive expression proportion is defined as the percentage of frames classified as happy out of the total number of valid frames. The negative expression proportion is defined as the percentage of the sum of frames classified as sad, angry, fearful, and disgusted out of the total number of valid frames. These two proportions are important input dimensions for the subsequent fusion and determination stage.

[0071] In a further proposed solution, the micro-expression detection subtask corresponds to... Figure 5The process is illustrated. Microexpressions are extremely short-lived facial expressions that are difficult to control voluntarily, typically lasting between 40-200ms. Compared to ordinary facial expressions, which can be subjectively controlled, microexpressions are more likely to reflect an individual's true emotional state. For adolescents, due to social desirability bias, they often tend to display positive or neutral emotions at the level of ordinary facial expressions during psychological screenings, while true negative emotions may only be revealed through fleeting microexpressions. Therefore, microexpression detection is of significant value in identifying emotional concealment behavior in adolescents.

[0072] Please see Figure 5 The document demonstrates the workflow of micro-expression detection. The input to this workflow is a sequence of frame-by-frame expression analysis results (frame_results), which is the sequence of probability values ​​for each expression category obtained after expression classification for each frame. Micro-expression detection focuses on four types of negative expressions: sadness, anger, fear, and disgust.

[0073] Step S2-11: Input the sequence of frame-by-frame facial expression analysis results.

[0074] Step S2-12: Perform a frame count check to determine if the total number of frames in the sequence is greater than or equal to 3 frames. Peak detection of micro-expressions requires at least three consecutive frames to determine the low-high-low three-segment pattern. If the number of frames is less than 3, an empty micro-expression event list is returned directly in step S2-13.

[0075] Step S2-14: Start the sliding window traversal. Take three consecutive frames, namely the (i-1)th frame, the ith frame, and the i+1th frame in the window, as a group, and slide the window frame by frame starting from the second element of the sequence, continuing the traversal until the second to last element.

[0076] Step S2-15 indicates that the window continues to slide until it covers the entire video sequence.

[0077] At each position of the sliding window, step S2-16 performs peak detection for each of the four negative expression types.

[0078] Step S2-17: Determine the peak mode by checking if there is a peak value that satisfies the low-high-low mode within the current three-frame window. The specific conditions for determining the low-high-low mode are: the expression probability value corresponding to prev_val in the previous frame is less than the preset intensity threshold T, curr_val in the middle frame is greater than the intensity threshold T, and next_val in the next frame is less than the intensity threshold T.

[0079] In an optimized embodiment, the intensity threshold T is set to 0.3. This value is based on the fact that in a parametric sensitivity test conducted on a subsample of 100 subjects, when T=0.2, the false positive rate was as high as 18.5%; when T=0.3, the false positive rate decreased to 8.2%, while the false negative rate remained within an acceptable range; when T=0.4, although the false positive rate further decreased to 5.1%, the false negative rate increased significantly to 22.1%. Setting T=0.3 achieves the optimal balance between the false positive and false negative rates.

[0080] If the peak condition is met, the duration is determined in step S2-18. The duration corresponding to the number of frames spanned by the peak pattern is calculated. In this embodiment, the video analysis frame rate is 5 frames per second, that is, 1 frame is taken out of every 6 frames in the original video of 30 frames per second for analysis. The time interval between two frames is 200ms, so the time width corresponding to a three-frame window is 400ms. When the peak only appears in the middle frame, the calculated duration is equal to 2 divided by the frame rate, that is, 2 divided by 5 equals 400ms, which does not meet the constraint that the duration of micro-expressions is less than 200ms. However, since micro-expressions usually only occupy a very narrow interval in the time series, they usually only reach a peak in a single frame under the sampling of 5 frames per second and then fall back quickly in the frames before and after. Therefore, the actual detected duration of micro-expression events is the interval between two frames, and it is necessary to make a judgment based on the comparison result of this interval time and 200ms. If the duration is less than 200 ms, proceed to step S2-20 to record the event as a micro-expression event. The recorded information includes the timestamp of occurrence, expression type, duration, peak intensity, and start and end locations.

[0081] If the duration is ≥ 200 ms, it indicates that the peak value may be a normal facial expression rather than a micro-expression. In step S2-19, the peak value is ignored, and the process returns to step S2-15 to continue the iteration. 200 ms is a classic definition in the field as the time boundary between micro-expressions and normal expressions, and it is widely cited in the research of scholars such as Ekman.

[0082] After each event is recorded or the judgment is ignored, the process returns to step S2-15, and the window continues to slide until the entire video sequence has been traversed.

[0083] Studies have shown that emotional masking is more prevalent among adolescents. This is because adolescents gradually learn behavioral patterns of conforming to social expectations and concealing their true inner emotions during the socialization process. If screening systems rely solely on self-reported scale data or subjectively controllable facial expressions, they are highly likely to capture distorted emotional signals, leading to an underestimation or even missed diagnosis of depression risk.

[0084] Please see Figure 6 This demonstrates the process for determining emotional concealment. It cross-validates the results of micro-expression detection with the subject's self-reported subjective emotional state to determine whether the subject is concealing their emotions. The process takes two inputs: first, the list of micro-expression events (micro_expressions) output from the aforementioned micro-expression detection subtask; and second, the subject's self-reported positive emotion ratio (self_report_positive_ratio). The self-reported positive emotion ratio can be derived from the SAM scale's pleasure rating or responses to positive emotion-related questions on other scales.

[0085] Specifically: Step S2-21: Receive the above input data.

[0086] Step S2-22: Filter out events from the micro-expression event list that belong to the four negative emotions of sadness, anger, fear and disgust, and obtain the negative micro-expression subset negative_micro.

[0087] Step S2-23: Perform an empty check on the screening results. If the negative micro-expression subset is empty, that is, no negative micro-expressions were detected, it means that the subject did not reveal any negative emotions. In step S2-24, return directly to FALSE and determine that there is no emotional concealment.

[0088] If the negative micro-expression subset is not empty, steps S2-25 to S2-27 calculate the three judgment conditions in parallel. All three conditions must be met to finally determine that emotional concealment exists. These three conditions are comprehensively considered from three dimensions: the number and intensity of micro-expressions, and the contradiction between subjective reports and objective signals.

[0089] in: The first condition is the count_condition, which determines whether the number of negative micro-expression events exceeds 3. This threshold of 3 is based on experimental data analysis. Under the condition of analyzing a 60-second video at a frame rate of 5 frames per second, normal adolescents typically exhibit 0-2 occasional negative micro-expressions in a natural state. In a comparative experiment, the average number of negative micro-expressions in the non-depressed group was 0.8, while the average number in the depressed group was 4.2. ROC curve analysis showed that 3 was the optimal cutoff point for the Youden index to reach its maximum value of 0.71.

[0090] The second condition is the intensity_condition, which determines whether the average peak intensity of negative micro-expression events is greater than 0.4. The average intensity is defined as the arithmetic mean of the intensity values ​​of all negative micro-expression events selected in step S2-22. Micro-expressions with an intensity greater than 0.4 are considered to have higher emotional information content; this threshold is supported by Ekman's 2003 study. Experimental sample analysis showed that the average intensity of random facial fluctuations was approximately 0.25, while the average intensity of genuine micro-expressions was approximately 0.55. Therefore, a threshold greater than 0.4 can effectively distinguish between genuine micro-expressions and random fluctuations.

[0091] The third condition is the discrepancy condition, which determines whether the proportion of self-reported positive emotions is greater than 0.5. When a subject's self-reported positive emotions exceed 50%, but negative micro-expressions frequently appear, it indicates a significant inconsistency between the subject's subjectively reported emotional state and the objectively detected micro-emotional signals. This contradiction between self-reported and true emotions is strong evidence for determining emotional concealment.

[0092] Step S2-28: Perform a logical AND operation on the results of the three judgment conditions. The overall judgment result is_concealing is TRUE only if all three conditions—quantity, intensity, and reporting inconsistency—are TRUE, confirming that the subject is concealing emotions. If any condition is not met, it is determined that there is no emotional concealment.

[0093] Step S2-29: Output the judgment result of the depressive mood masking detection.

[0094] Step S3, Speech Analysis This step receives the audio segment output from step S0 and the speech analysis parameter configuration output from step S1 as input, and performs sub-tasks such as audio preprocessing, prosodic feature extraction, and speech quality index extraction.

[0095] Specifically, firstly, the audio undergoes preprocessing operations such as resampling and frame segmentation. Then, prosodic features such as fundamental frequency sequence, speech rate, and pause ratio are extracted, along with three speech quality indicators: frequency perturbation (jitter), amplitude perturbation (shimmer), and harmony-to-noise ratio (HNR). Based on the weighted configuration corresponding to each age group, the three quality indicators are weighted and summed to calculate a comprehensive speech quality score. Step S3 ultimately outputs binary features of low-high pitch, reduced pitch range, slow speech rate, high pause ratio, negative emotion ratio, and the comprehensive speech quality score.

[0096] The following combination Figure 7 The specific implementation of the speech analysis in step S3 will be explained in detail.

[0097] Figure 7 This demonstrates the processing flow for extracting and comprehensively evaluating speech quality metrics. The inputs to this flow are the audio file path (audio_path) and the target sampling rate (sr), which is set to 16000Hz by default.

[0098] Step S3-0: Audio Preprocessing. First, load the audio file from the specified path and resample it to a 16kHz sampling rate to ensure all input audio has a uniform time-frequency resolution. Then, perform frame segmentation on the resampled audio signal. Framing uses a frame length of 2048 sampling points, corresponding to a time window of approximately 128ms. This frame length is sufficient to cover 2-3 complete fundamental frequency cycles at a 16kHz sampling rate, ensuring the stability of the spectrum estimation. Frame shift is 512 sampling points, corresponding to approximately 32ms, achieving an inter-frame overlap rate of 75%. This high overlap rate ensures smooth inter-frame transitions while maintaining time resolution, avoiding artifacts introduced by frame edge effects.

[0099] Steps S3-1a and S3-1b are responsible for calculating the amplitude perturbation shimmer.

[0100] In step S3-1a, the root mean square (RMS) energy value is calculated frame by frame for the segmented audio, resulting in an RMS energy sequence. Step S3-1b then calculates the shimmer value based on this sequence. The shimmer is defined as the standard deviation of the RMS energy sequence divided by its mean, reflecting the periodic variation in glottal excitation amplitude between adjacent periods. To prevent division by zero errors in silent segments, a minimum value, epsilon (1e-6), is added to the denominator.

[0101] Steps S3-2a and S3-2b are responsible for calculating the frequency perturbation jitter.

[0102] Step S3-2a extracts the frame-by-frame fundamental frequency value sequence pitch_values ​​using a fundamental frequency tracking algorithm. Fundamental frequency tracking can be implemented using common algorithms such as autocorrelation or cepstral method. Step S3-2b calculates the jitter value based on this sequence. Jitter is defined as the standard deviation of the fundamental frequency difference sequence between adjacent frames divided by the mean of the fundamental frequency sequence. When the number of effective fundamental frequency frames is less than two, the jitter value is directly set to 0.0.

[0103] Steps S3-3a and S3-3b are responsible for calculating the harmonic-to-noise ratio (HNR).

[0104] In step S3-3a, the spectral centroid sequence `spectral_centroids` is calculated for the original audio signal. The spectral centroid reflects the concentrated frequency band of the signal's spectral energy; the clearer and more regular the harmonic structure, the more stable the value of the spectral centroid. Step S3-3b uses the mean of the spectral centroid sequence divided by 1000.0 as an approximate measure of HNR. The physical basis of this approximation is that speech signals with concentrated harmonic energy typically have a higher and more stable spectral centroid, while noise-dominated signals have a lower and more fluctuating spectral centroid.

[0105] Steps S3-5: Based on the subject's age group, select the corresponding weight combination from the predefined weight configuration table. The weight configuration scheme has been explained in detail in the age group identification section. If the subject's age is 15-16 years old, which is in the middle of puberty and voice change, then apply the voice change adjustment weight, i.e. Take 0.20, Take 0.20, Set the weight to 0.60. If the subject belongs to another age group, apply the standard weight, i.e. Take 0.30, Take 0.30, Take 0.40.

[0106] Step S3-6: Calculate the weighted comprehensive score. The formula for calculating the comprehensive voice quality score (voice_quality_score) is:

[0107] Those skilled in the art will understand that the comprehensive speech quality score (voice_quality_score) is the sum of the normalized values ​​of the three indicators multiplied by their respective weights, resulting in a comprehensive penalty value reflecting the degree of speech quality defects. This penalty value is then subtracted from 1 to obtain the quality score. A higher quality score indicates better audio quality and less susceptibility to noise and physiological interference. Here, jitter and shimmer are themselves normalized perturbation ratios, typically ranging from 0 to 0.1, and are directly used as defect contributors. HNR is normalized to the 0-1 range by dividing min(HNR,20) by 20, where 20 is the upper bound of the HNR for normal adults, cited from the 1982 study by Yumoto, Gould, and Baer. The normal adult HNR range is 7-20 dB. The value obtained by subtracting the normalized HNR from 1 reflects the degree of harmonic component loss and is used as a defect contributor. After multiplying by their respective weights, the sum of the three defect contributions yields a comprehensive defect value, which is then subtracted from 1 to obtain the final quality score.

[0108] In addition to speech quality metrics, the speech analysis step also extracts prosodic features, including the mean and range of the fundamental frequency, speech rate, and pause ratio. These prosodic features are compared with preset thresholds to generate four binary features. See Table 6 below for details: Table 6. Speech Detection Parameter Values ​​and Basis

[0109] Step S4, Scale Scoring This step receives the scale response data output from step S0 as input, calculates the original score according to the scoring rules corresponding to the scale type used, and maps the score result to a uniform numerical range of 0 to 100. If the scale type is SAM and paired administration data before and after emotion induction exists, further emotion induction correction is performed. A preset correction coefficient is used to add the change in emotional state to the original score to obtain the corrected scale score. This step ultimately outputs a value within the range of 0 to 100 as the scale score.

[0110] The following combination Figure 8 The specific implementation of the scale scoring calculation is explained in detail.

[0111] Figure 8 The flowchart illustrates the process of scale rating calculation and emotion evoked correction. The inputs to the process include the scale type (scale_type), response data (scale_responses), and optional pre- and post-emotional evoked SAM data (pre_emotion_data and post_emotion_data).

[0112] Specifically: Step S4-0: Receive the above input data.

[0113] Step S4-1: Calculate the raw score (raw_score) according to the scale type and the corresponding scoring formula, and map the result to a scale range of 0 to 100.

[0114] In a preferred embodiment, the four scales shown in Table 7 below are supported: Table 7 Four scales

[0115] The scoring methods for the above four scales are as follows: The raw score of the SAM scale is calculated as follows: raw_score equals the pleasure score plus 9 minus the arousal score plus the dominance score, the sum of these three scores divided by the maximum score of 27, and then multiplied by 100. The maximum SAM score of 27 is based on a score of 9 for each of the three dimensions.

[0116] The CES-DC scale raw score normalization formula is: raw_score equals the sum of the 20 item scores divided by 60 and then multiplied by 100. The scale contains 20 items, each rated on a four-point scale from 0 to 3, with a maximum score of 60. Higher scores indicate more severe depressive symptoms; the cutoff point in the Chinese version is a score of 15 or higher indicating a risk of depression.

[0117] The normalization formula for the raw score of the PHQ-A scale is: raw_score equals the sum of the scores of the 9 items divided by 27 and then multiplied by 100. The scale contains 9 items, each rated on a four-point scale from 0 to 3, with a maximum score of 27. A score of 10 or higher indicates moderate to severe depression.

[0118] The CDRS-R scale raw score normalization formula is: raw_score equals the sum of the 17 scores divided by 80 and then multiplied by 100. This scale is administered by professionals through semi-structured interviews and contains 17 items, each rated on a five-point scale from 1 to 5, with a maximum score of 80. A score of 40 or higher indicates depression.

[0119] Step S4-2: Determine the conditions for emotion-evoked correction. The basis for this determination is whether complete scoring data of the SAM scale are available both before and after the emotion-evoked task.

[0120] If both the initial and subsequent SAM data are available, proceed to step S4-4, emotion-induced correction, as described below. The correction is calculated as follows: First, calculate the difference between the post-test SAM and the pre-test SAM across three dimensions. Specifically, the change in pleasure_change equals the post-test pleasure minus the pre-test pleasure; the change in arousal_change equals the post-test arousal minus the pre-test arousal; and the change in dominance_change equals the post-test dominance minus the pre-test dominance.

[0121] Then, the sum of the changes in the three dimensions is multiplied by a preset correction coefficient α, and this sum is added to the original score to obtain the adjusted score (adjusted_score). The correction coefficient α is set to 0.2.

[0122] It should be noted that this value was determined based on a grid search experiment conducted on a sample of 150 adolescents, with a search space of 0.1–0.3 and a step size of 0.05. When α=0.2, the corrected score showed the highest correlation with the clinical diagnosis, with a Pearson correlation coefficient r reaching 0.82, a significant improvement compared to r=0.78 when α=0.1 and r=0.79 when α=0.3. As a baseline comparison, without correction, i.e., when α=0, r was only 0.73; after introducing α=0.2, r increased to 0.82, a statistically significant difference (p=0.001). Five-fold cross-validation results showed that the r values ​​for each fold ranged from 0.80 to 0.84, with a mean of 0.82 and a standard deviation of 0.014, indicating that the correction effect of α=0.2 has good stability. Furthermore, when α>0.3, excessive task-induced noise is introduced, and when α<0.1, the correction effect is limited; therefore, the value range is limited to 0.1–0.3.

[0123] If SAM data is unavailable, i.e., CES-DC, PHQ-A, or CDRS-R was only performed once and there is no paired data before and after, then the raw score raw_score is directly output as the final score in step S4-3.

[0124] Step S4-5: Truncate the corrected score to the legal range of 0 to 100.

[0125] Step S4-6: Output the final corrected score.

[0126] Step S5, Feature Fusion This step receives the feature data output from steps S2, S3, and S4 as input. First, a 12-dimensional fusion feature vector is constructed according to a predefined order and dimensions, containing an 11-dimensional AI feature vector and a 1-dimensional scale rating feature. The 11-dimensional AI features include 6-dimensional features from facial analysis and 5-dimensional features from speech analysis. Specifically, the 6-dimensional facial features are: negative expression ratio, positive expression ratio, AU4 intensity reduction between eyebrows, AU15 intensity of corner compression at the corners of the mouth, AU43 intensity of eye closure, and negative micro-expression count; the 5-dimensional speech features are: low fundamental frequency marker, fundamental frequency range narrowing marker, speech rate slowing marker, high pause ratio marker, and negative emotion ratio. The scale rating feature is obtained by dividing the score output from step S4 by 100 and normalizing it to the [0,1] interval. Then, the fusion feature vector is linearly weighted and summed according to 12 predefined weight values ​​to obtain the fusion score. The weight corresponding to the positive expression ratio feature is negative, acting as a protective factor. The specific weight configuration and fusion calculation formula will be explained in detail later.

[0127] The following combination Figure 9 The specific implementation process of step S5 will be explained in detail.

[0128] Step S5-1: According to the above correspondence, concatenate the 11-dimensional AI features and the 1-dimensional scale features into 12-dimensional fusion feature vectors F1 to F12.

[0129] Step S5-2: Perform linear weighted fusion calculation. The fusion score S is calculated as the sum of the products of each dimension's feature value and its corresponding weight, i.e.:

[0130] In an optimized embodiment, the list of 12-dimensional fusion weights is shown in Table 8: Table 8. List of 12-dimensional fusion weights

[0131] Facial channel: |0.15+(-0.10)+0.12+0.10+0.08+0.05| = 0.40.

[0132] Voice channel: 0.12 + 0.08 + 0.10 + 0.07 + 0.03 = 0.40.

[0133] Scale channel: 0.30.

[0134] Based on the weight list in Table 8, the sum of the absolute values ​​of the three channel weights is: 0.40 + 0.40 + 0.30 = 1.10, where the positive expression protection factor has a negative weight (-0.10).

[0135] The fusion result is mapped to the [0,1] interval by the Sigmoid function.

[0136] It should be noted that in this scheme, multi-channel features are first fused using a linear weighting method, without strictly limiting the sum of the weight coefficients to 1. The reason for this is that the fusion result obtained by linear weighting is subsequently processed by the Sigmoid function, which can automatically normalize the output result to the [0,1] value range. The value range constraint can be completed by relying on the normalization property of Sigmoid itself.

[0137] It should be noted that in the above weighting configuration, the weight of the proportion of positive expressions... The value of -0.10 is a key design feature of this scheme. The proportion of positive facial expressions, as a protective factor, is given a negative weight. Physically, this means that the higher the proportion of positive facial expressions displayed by the subject during the calculation of the fusion score, the greater the contribution of this factor to the fusion score. However, because the weight is negative, the accumulation of this factor will shift the overall fusion score towards a reduced risk of depression. This design aligns with the broadening-construction theory in positive psychology. Proposed by Fredrickson in 2001, this theory states that positive emotions can broaden an individual's cognitive scope, enhance psychological resilience, and have a protective effect against depression. Adolescents have high psychological plasticity, and the protective effect of positive emotions against depression is more significant than in adults. By explicitly introducing this protective factor into the fusion model and allowing it to function with a negative weight, the negative bias of traditional methods that only accumulate pathological indicators can be effectively corrected.

[0138] In a further step, to verify the technical effectiveness of the aforementioned negative weighting protection factor, this approach employs an ablation experiment. In a study involving 150 adolescents, the screening performance under different positive expression weighting values ​​was tested. The experimental results are shown in Table 9 below: Table 9. Experimental Results of Different Positive Expression Weights

[0139] As shown in Table 9, when the protective factor is not used, i.e. When the value is 0.00, the accuracy is 81.2%, the sensitivity is 78.5%, and the specificity is 83.8%. When the value is negative 0.05, the three indicators rise to 83.4%, 81.2%, and 85.5%, respectively. When the value is negative 0.10, the three indicators reach their optimal levels, at 85.7%, 83.1%, and 88.0% respectively. When further reduced to -0.15 or -0.20, performance declined, to 84.9% and 83.8% respectively. Five-fold cross-validation results indicate that... The average accuracy was 85.7% when the value was -0.10. The average accuracy rate of 81.4% when the weight is 0.00 represents an improvement of 4.3 percentage points. The paired t-test showed a t-statistic of 8.52 and a p-value less than 0.001, indicating a statistically significant difference. These experimental data strongly demonstrate the practical effectiveness of the negative weighting protection factor in improving screening accuracy and reducing the probability of overdiagnosis.

[0140] Step S6, Risk Probability Mapping This step receives the fusion score output from step S5 as input. Using the mapping function, namely the sigmoid function, the fusion score is mapped from the real number field to the interval [0,1] to obtain the probability of depression risk.

[0141] The following combination Figure 9 The specific implementation process of step S6 will be explained in detail.

[0142] Step S6-1: Perform Sigmoid probability mapping.

[0143] The expression for the mapping function in this step is: This function continuously and monotonically maps the fusion score S in the real number field to an open interval between 0 and 1. The probability approaches 0 when the fusion score S approaches negative infinity; it is 0.5 when S equals 0; and it approaches 1 when S approaches positive infinity. The output of the Sigmoid function has good probabilistic semantics, facilitating the subsequent setting of interpretable risk level thresholds, as shown in Table 10.

[0144] Table 10 Risk Level Threshold Basis

[0145] Step S6-2: Output the final probability of depression, P, with a value between 0 and 1.

[0146] Step S7: Confidence Calculation This step receives the data quality metrics output from steps S2 and S3 as input, including the face detection rate of the face channel and the audio quality score of the voice channel, and also receives the scale data availability information from step S4. Following a pre-defined cumulative-penalty model, starting from the baseline confidence level, the data availability and quality of the face channel, voice channel, and scale channel are evaluated sequentially. The corresponding baseline contribution values ​​are accumulated and multiplied by the corresponding quality factor for penalized reduction, ultimately truncating the result to an upper limit of 1.0. This step outputs a comprehensive confidence value within the range of 0.5 to 1.0.

[0147] Step S8: Risk Level Determination This step receives the depression risk probability output from step S6 and the confidence level output from step S7 as input. The depression risk probability is compared with three preset thresholds of 0.4, 0.6, and 0.8, classifying it into four levels: no depression, mild depression, moderate depression, and severe depression. Simultaneously, the confidence level is labeled according to the range in which the confidence value falls. If the confidence level is lower than the preset low confidence threshold, a suggestion to retake the test is added to the output.

[0148] The following combination Figure 10 The specific implementation of the confidence level adaptive adjustment is explained in detail.

[0149] It should be noted that when conducting large-scale screening in uncontrolled school or community environments, data quality is inevitably affected by various factors. For example, poor ambient lighting or subjects not facing the camera directly will reduce face detection rates; background noise will contaminate audio quality; and some subjects may not have completed the questionnaire completely. If the system outputs high-confidence conclusions when data quality is low, it will seriously affect the reliability of the screening results. The confidence adaptive adjustment mechanism introduced in this solution aims to quantitatively assess the availability and quality of data from each modality and dynamically adjust the confidence level of the final conclusion accordingly.

[0150] Figure 10 The processing flow of the confidence-adaptive mechanism is illustrated. This flow employs an additive-penalty model, starting with a baseline confidence level, sequentially evaluating the data availability and quality of the facial, voice, and scale channels, accumulating the corresponding baseline contribution values, and multiplying them by a quality penalty factor. That is, the confidence level. The calculation formula is:

[0151] In the formula, The baseline uncertainty is set to 0.5. For modal indexing (face=f, voice=v, scale=s), where, The base contribution to the facial channel is set to 0.2. The base contribution for the voice channel is 0.2. The base contribution of the scale channel is set to 0.1. The facial data quality factor is equal to the face detection rate, with a value range of [0,1]. The speech data quality factor is equal to the audio quality score, with a value range of [0,1]. The scale data quality factor (available = 1, unavailable = 0) has a value range of {0, 1}.

[0152] Specifically: Step S8-0: Initialize confidence.

[0153] The initial value of the overall confidence level C is set to 0.5. 0.5 represents the baseline uncertainty when there is no valid evidence; mathematically, it means that under the current information conditions, the probability that the system's conclusion is correct is only at the level of random guessing. When all three modalities of data are unavailable, the system outputs a minimum confidence level of 0.5, indicating that the screening conclusion has extremely low reliability.

[0154] Step S8-1: Assess the availability of facial channel data.

[0155] The criterion is a face detection rate greater than or equal to 50%, meaning that the number of frames that effectively detect faces in step S2 accounts for more than half of the total processed frames. If this condition is met, the facial data is deemed usable, and the process proceeds to step S8-2 to perform quality assessment and contribution accumulation for the facial channel. If the data is unusable, step S8-2 is skipped, and the process directly proceeds to step S8-3 to evaluate the speech channel.

[0156] Step S8-2: Perform confidence update for the facial channel. This update consists of the following two steps: The first step is to add the current confidence level C to the base contribution value of the facial channel, which is 0.2. The facial channel provides 6 of the 12 fusion features, accounting for 55% of the total feature dimensions, so its base contribution value is set to 0.2.

[0157] The second step is to multiply the accumulated confidence score by the quality factor of the face channel, i.e., the face detection rate, as a penalty reduction. For example, if the accumulated confidence score C is 0.7 and the face detection rate is 0.93, then C is updated to 0.7 multiplied by 0.93, which equals 0.651.

[0158] The physical meaning of the detection rate penalty is that a 10% decrease in the face detection rate is equivalent to losing approximately 12% of the temporal information of facial action units, directly affecting the reliability of micro-expression peak detection and expression ratio calculation. Experiments show that when the detection rate is below 50%, the intra-group consistency (ICC) of facial feature extraction drops to 0.42, which is unacceptable. Therefore, at this point, this channel should be skipped instead of continuing to contribute to the basic analysis.

[0159] Step S8-3: Assess the data availability of the voice channel.

[0160] Specifically, the criterion is that the audio quality score is greater than or equal to 0.6, meaning the comprehensive voice quality score output in step S3 is higher than or equal to the preset threshold of 0.6. This threshold is set based on the fact that when the voice quality score is below 0.6, the signal-to-noise ratio of the voice features is insufficient to support reliable prosodic analysis. If this condition is met, the voice data is deemed usable, and the process proceeds to step S8-4 to perform voice channel quality assessment and contribution accumulation. If the data is unusable, step S8-4 is skipped, and the process directly proceeds to step S8-5 to evaluate the scale channel.

[0161] Step S8-4: Perform a confidence update for the voice channel. The update consists of the following two steps: The first step is to add the current confidence level C to the base contribution value of the voice channel, which is 0.2. The voice channel provides 5 of the 12 fusion features, accounting for 45% of the total feature dimensions, and its base contribution value is the same as that of the facial channel.

[0162] The second step involves multiplying the accumulated confidence scores by the quality factor of the speech channel, resulting in the comprehensive speech quality score output in step S3. This score integrates three indicators: shimmer, jitter, and HNR, comprehensively reflecting the acoustic quality of the audio. After this penalized reduction, the screening conclusions generated by low-quality audio will have a lower confidence level.

[0163] Step S8-5: Evaluate the data availability of the scale channel. The judgment criterion is that the scale scores are not empty and are within the valid range, that is, the valid scores were successfully output in step S4. If available, proceed to step S8-6.

[0164] Step S8-6: Execute confidence updates for the scale channels.

[0165] Since the scale data is subjective self-reported data and only provides one of the 12 features, its base contribution value is set to 0.1, lower than the 0.2 for the facial and voice channels. In this embodiment, the quality factor for the scale channel is simplified to binary: 1 when data is available and 0 when unavailable. Therefore, when scale data is available, it is directly incremented by 0.1 without additional penalty reduction.

[0166] Step S8-7: Perform confidence level truncation. Take the minimum value between the current confidence level C and the theoretical upper limit 1.0, i.e., C equals min(C, 1.0). This ensures that in the ideal case, when all three modal data are available and of perfect quality, the confidence level is limited to 1.0.

[0167] Step S8-8: Output the final overall confidence level C, with a value ranging from 0.5 to 1.0.

[0168] This confidence level value will be used in the subsequent risk level determination step S8 to indicate the confidence level of the conclusion or to trigger a prompt to recommend retesting.

[0169] The values ​​of the relevant parameters for the confidence adaptive mechanism are shown in Table 11 below: Table 11 Values ​​of relevant parameters for the confidence-adaptive mechanism

[0170] Step S9, Output Results This step receives the judgment results output from step S8, integrates the risk probability, severity level, confidence level, and intermediate score information of each modality, and forms a structured screening report for output.

[0171] Those skilled in the art will understand that this invention has been adapted for screening depression risk in adolescents: First, the subjects were divided into three different age groups—early, middle, and late puberty—based on their age. The face detection window parameters for facial analysis and the prosody and quality weight parameters for speech analysis were configured differently accordingly. This reduced the physiological noise caused by facial bone development and vocal cord changes at the source of signal acquisition and feature extraction.

[0172] Secondly, in the facial analysis path, a peak detection method based on sliding time windows is introduced to capture negative micro-expressions lasting less than 200ms, and these are cross-validated with self-reported emotional states to achieve automatic judgment of emotional concealment behavior, thereby making up for the shortcomings of existing technologies in detecting adolescents' social desirability bias.

[0173] Furthermore, in the fusion decision-making stage, the proportion of positive facial expressions in the 12-dimensional fusion feature vector is innovatively used as a protective factor and given a negative weight to correct the negative bias caused by the traditional method of only accumulating pathological indicators, effectively reducing the probability of overdiagnosis.

[0174] Finally, by combining an adaptive multimodal confidence assessment mechanism, the credibility of the final screening results is dynamically adjusted based on the quality of signals from each channel, thereby comprehensively improving the accuracy and robustness of large-scale screening in uncontrolled campus or community environments.

[0175] This invention also provides a multimodal depression risk fusion screening system for adolescents, including a data acquisition module, an age group identification module, a facial analysis module, a voice analysis module, a scale scoring module, a confidence assessment module, and a feature fusion and risk determination module. Specifically: The data acquisition module is used to obtain facial videos, audio recordings, and psychological scale responses from subjects aged 12 to 18.

[0176] The age group recognition module is used to classify subjects into one of several preset age groups based on their age, and to configure facial analysis parameters and voice analysis parameters differently based on the classified age group.

[0177] The facial analysis module is used to extract facial expression features from facial videos based on facial analysis parameters corresponding to age groups.

[0178] The speech analysis module is used to extract speech prosodic features and speech quality features from speech audio based on speech analysis parameters corresponding to age groups.

[0179] The scale scoring module is used to calculate scale scores based on the responses to psychological scales.

[0180] The confidence assessment module is used to calculate the overall confidence level of the current screening based on the availability or quality indicators of the data acquired by the data acquisition module.

[0181] The feature fusion and risk assessment module is used to construct a fusion feature vector that includes facial expression features, speech prosody features and scale scores; and to perform linear weighted fusion on the fusion feature vector to obtain a fusion score, wherein the positive expression proportion feature in the fusion feature vector is given a negative weight; and based on the fusion score, the subject's depression risk level is determined.

[0182] Please see Figure 11 This invention also provides a system architecture diagram, which includes a presentation layer, a business layer, a data layer, and an infrastructure layer. The presentation layer provides a user interface, allowing subjects to access the screening system via a browser. The business layer encapsulates the core algorithm modules described in the embodiments of this application, including age group identification services, facial analysis services, voice analysis services, scale scoring services, feature fusion and risk assessment services, and confidence assessment services. These service modules work together to complete the entire process from data collection to screening report generation. The data layer is responsible for persistent data storage and caching. The infrastructure layer provides the system's operating environment support, ensuring environmental consistency and portability for each service module.

[0183] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal depression risk fusion screening method for adolescents, characterized in that, Includes the following steps: Acquire facial video, audio recordings, and psychological scale responses from adolescent subjects; The subjects were divided into age groups based on their age, and the corresponding facial analysis parameters and voice analysis parameters were configured according to the differences in the age groups. Facial expression features are extracted based on the facial analysis parameters corresponding to the age group, speech features are extracted based on the speech analysis parameters corresponding to the age group, and the scale score is calculated based on the psychological scale response data. A fusion feature vector containing the facial expression features, the voice features, and the scale score is constructed, and the fusion feature vector is linearly weighted and fused to obtain a fusion score; wherein, the positive expression proportion feature in the fusion feature vector is given a negative weight; The subject's depression risk level was determined based on the fusion score.

2. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The speech features include speech prosody features and speech quality features, wherein: The prosodic features include fundamental frequency, speech rate, and pause ratio; The speech quality characteristics include frequency perturbation, amplitude perturbation, and harmony-to-noise ratio.

3. The multimodal depression risk fusion screening method for adolescents according to claim 2, characterized in that, The extraction of the speech quality features includes: Calculate the frequency perturbation, amplitude perturbation, and harmonic-to-noise ratio of the speech audio; Based on the age group, the weights of the frequency perturbation, the amplitude perturbation, and the harmonic-to-noise ratio (HNR) are differentiated in calculating the overall speech quality score, such that the HNR weight corresponding to the age group in puberty is higher than the HNR weight corresponding to other age groups.

4. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The step of extracting facial expression features based on the facial analysis parameters corresponding to the age group also includes: Micro-expression events with a duration of less than a preset threshold are detected from the facial video to obtain the number and intensity of negative micro-expressions; Based on the number of negative micro-expressions, the intensity of the negative micro-expressions, and the self-reported emotional state derived from the scale scores, it is comprehensively determined whether the subject has emotional concealment behavior, and the determination result is used to construct a feature of the fused feature vector.

5. The multimodal depression risk fusion screening method for adolescents according to claim 4, characterized in that, The step of detecting micro-expression events from the facial video with a duration less than a preset threshold also includes: The facial expression sequence frames of the video are traversed according to a sliding time window; Within each window, determine whether there is a frame sequence in three consecutive frames where the intensity of negative facial expressions satisfies the low-high-low peak pattern; If it exists, calculate the duration of the peak pattern; If the duration is less than 200ms, it is recorded as a micro-expression event.

6. The multimodal depression risk fusion screening method for adolescents according to claim 4, characterized in that, The comprehensive determination of whether the subject has engaged in emotional concealment behavior is specifically based on the following three conditions being met simultaneously: The number of detected negative micro-expression events exceeds a preset threshold; The average intensity of the negative micro-expression events exceeds a preset intensity threshold; The proportion of subjective positive emotion reports derived from the scale's response data exceeded a preset threshold.

7. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The method further includes: The overall confidence level of the current screening task is calculated based on the availability or quality indicators of the facial video, the audio, and the psychological scale response data.

8. The multimodal depression risk fusion screening method for adolescents according to claim 7, characterized in that, It also includes the aforementioned confidence level adaptive adjustment, specifically: Initialize baseline confidence; For each available data mode, a preset basic contribution value corresponding to that mode is accumulated on the baseline confidence level; The accumulated confidence scores are multiplied by a quality factor that reflects the quality of the modality data to implement punitive reduction.

9. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The step of dividing the subjects into age groups based on their age and configuring corresponding facial analysis parameters and voice analysis parameters according to the differences in the age groups specifically includes: Subjects aged 12 to 14 were matched to early puberty and assigned the first minimum window size for face detection and the first combination of speech analysis weights. Subjects aged 15 to 16 were matched to mid-adolescence and assigned a second face detection minimum window size larger than the first face detection minimum window size and a second speech analysis weight combination. The sum of the weights of frequency perturbation and amplitude perturbation in the second speech analysis weight combination was lower than the sum of the corresponding weights in the first speech analysis weight combination. Subjects aged 17 to 18 were matched to late adolescence and assigned a third face detection minimum window size and a third speech analysis weight combination that was larger than the second face detection minimum window size.

10. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The calculation of the scale score based on the responses to the psychological scale includes: Obtain paired scale data before and after the emotional triggering; Calculate the change in a specified dimension of the paired scale data; Using a preset correction coefficient, the change is superimposed on the original scale score to obtain the scale score used to construct the fusion feature vector.

11. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that: The facial analysis parameters include the minimum window size for face detection; The speech analysis parameters include the weights of frequency perturbation, amplitude perturbation, and harmony-to-noise ratio in calculating the overall speech quality score.

12. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The facial expression features include the proportion of negative expressions, the proportion of positive expressions, and the intensity of at least one facial movement unit selected from the group consisting of lowered eyebrows, pursed corners of the mouth, and closed eyes.

13. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The step of extracting facial expression features based on the facial analysis parameters corresponding to the age group includes: Detect the face region and scale it to a preset pixel size; The scaled-down face area is divided into four regions of interest: eyebrows, eyes, cheeks, and mouth. Based on the region of interest, the intensity of the specified facial motion unit is calculated.

14. The multimodal depression risk fusion screening method for adolescents according to claim 13, characterized in that, The calculation of the intensity of the specified facial motion units specifically includes: The intensity reduction between the eyebrows is obtained by calculating the vertical gradient of the eyebrow region; The corner compression strength of the mouth is obtained by calculating the curvature change of the lower edge of the mouth region; The eye closure intensity is obtained by calculating the pixel ratio of the upper and lower halves of the eye region.

15. The multimodal depression risk fusion screening method for adolescents according to claim 1, characterized in that, The determination of the subject's depression risk level based on the fusion score includes: The fusion score is mapped to a depression risk probability using a mapping function; The probability of depression risk is compared with multiple preset risk level thresholds to determine the corresponding risk level.

16. A multimodal depression risk fusion screening system for adolescents implementing the method as described in any one of claims 1-15, characterized in that, include: The data acquisition module is used to acquire facial video, voice audio, and psychological scale responses from subjects aged 12 to 18. An age group recognition module is used to classify the subject into one of a number of preset age groups based on the subject's age, and to configure facial analysis parameters and voice analysis parameters differently based on the classified age group. The facial analysis module is used to extract facial expression features from the facial video based on facial analysis parameters corresponding to the age group. The speech analysis module is used to extract speech prosodic features and speech quality features from the speech audio based on speech analysis parameters corresponding to the age group. The scale scoring module is used to calculate the scale score based on the response data of the psychological scale; and The feature fusion and risk assessment module is used to construct a fusion feature vector containing the facial expression features, the speech prosody features, and the scale score; to perform linear weighted fusion on the fusion feature vector to obtain a fusion score, wherein the positive expression proportion feature in the fusion feature vector is given a negative weight; and to determine the subject's depression risk level based on the fusion score.