Undisturbed depression recognition method based on gait behavior analysis and face recognition
By using Microsoft Kinect and channel topology refinement graph convolution network combined with facial recognition technology, the cost and interference problems of existing gait analysis equipment are solved, and the interference-free depression recognition is achieved, which improves the recognition accuracy and popularity.
Patent Information
- Application Number
- CN202510411119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing gait analysis methods are expensive in clinical environments and require professional operation, and cannot obtain natural gait behavior characteristics, resulting in inaccurate identification of depression.
Microsoft Kinect was used to shoot the trial walking video, extract the skeleton data, combine the channel-based topological refinement graph convolution for gait behavior analysis, and combine facial recognition technology to mark without interfering with the subject's walking to build a depression recognition model.
Disturbance-free depression recognition is achieved, which improves the accuracy and popularity of recognition, reduces equipment costs, and reduces interference to subjects' walking.
Smart Images

Figure CN120259773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and medical health, and in particular to a non-intrusive depression recognition method based on gait behavior analysis and face recognition. Background Art
[0002] In recent years, a number of studies have shown that there is a close relationship between mental health and physical behavior, and behavioral characteristics can be used to identify mental characteristics. Combining computer technology, digital non-invasive methods have been widely used in mental health research. Digital non-invasive methods refer to using computer technology to extract features, model and classify the behaviors of patients with depression, and then identify depression. The current behavior recognition methods mainly include facial expression, speech and gait recognition. The specificity and sensitivity of the depression recognition method based on gait behavior analysis are quite good, but no related technology for non-intrusive depression recognition based on gait behavior analysis has been published. The equipment required for current quantitative gait analysis is relatively expensive and requires professional operation, so it is not common in the clinical environment; in addition, the current gait analysis methods still have many interferences on the walking process of the subjects, and it is impossible to obtain completely natural gait behavior characteristics. Summary of the Invention
[0003] To solve the above problems, the present invention provides a non-intrusive depression recognition method based on gait behavior analysis and face recognition, which uses Microsoft Kinect to capture the walking video of the subject, extracts the skeleton data during the walking process of the subject for extracting gait behavior characteristics, uses channel-based topological refinement graph convolution to extract representative features in different channels for skeleton-based action recognition; and proposes a key population marking method based on face recognition technology, which marks the subject with face recognition technology without disturbing the walking of the subject, and further collects and identifies data of the key population to help identify individuals at risk of mental disorders.
[0004] To achieve the above object, the present invention provides a non-intrusive depression recognition method based on gait behavior analysis and face recognition, specifically including the following steps:
[0005] S1: Face recognition, input the face information of the subject, extract the face features and compare them with the key collection objects in the database until it is determined as a collection object to start collecting the skeleton data of the subject;
[0006] S2: Capture skeleton data, screen subjects covering a large age range, record natural walking videos, and extract the joint coordinates of the subjects at each time node to form a data set;
[0007] S3: Data preprocessing, which involves data cleaning, coordinate standardization, time alignment, and data augmentation for the joint coordinates of the subjects.
[0008] S4: Model classification, which extracts the features of the joint coordinate data of the subjects and constructs a depression recognition model based on the data features.
[0009] Preferably, in step S1, it specifically includes the following steps:
[0010] S11: Face information entry, using Dlib to detect faces from the camera and store the captured face images locally.
[0011] S12: Face feature extraction. For each subject, based on the multiple captured face images, extract the 128D feature vector of each face, calculate the mean of the 128D feature vectors of the subject, detect 68 feature points of the face image through the ResNet-34 deep learning model, and map the 68 feature points into a 128-dimensional vector space. The vectors in the 128-dimensional vector space are used for face recognition and comparison. The Euclidean distance is used to calculate the distance between two faces:
[0012]
[0013] S13: Face feature comparison. Detect faces from the real-time input camera video stream, extract their 128D feature vectors, and compare them with the feature vectors in the stored face dataset. Calculate the Euclidean distance to determine whether they are the same face. The judgment criterion is as follows:
[0014] If the Euclidean distance between the vector spaces of two faces exceeds the discrimination threshold, it is considered that they are not the same person;
[0015] If the Euclidean distance between the vector spaces of two faces is less than the discrimination threshold, it is considered that they are the same person.
[0016] Preferably, in step S2, filter the natural walking video data of the subjects to ensure that the ages of all subjects are normally distributed. Use the Kinect SDK to extract 25 joint coordinates of the subjects. The specific steps are as follows:
[0017] S21: Use the automatic capture process. When the subject enters the face recognition range, after completing face recognition at the entrance, if it is a subject, data information is collected; if not, no collection is performed.
[0018] S22: Install a face collection module at the exit. After the subject walks out of the collection range and the face collection module recognizes the face information, the collection ends.
[0019] Preferably, in step S3, the data preprocessing specifically includes the following steps:
[0020] S31: Remove invalid data from the data. The invalid data includes unnecessary data, duplicate values, and outliers, and perform data cleaning;
[0021] S32: Eliminate the problem of inconsistent spatial coordinates of the subject relative to the camera and the influence of different body sizes by normalizing the joint coordinates;
[0022] S33: For the time series data collected by different cameras, perform time alignment through frame sampling and time interpolation;
[0023] S34: Perform data augmentation through rotation, shifting, and joint loss simulation to increase the amount of relevant data.
[0024] Preferably, in step S4, it specifically includes the following steps:
[0025] S41: Feature extraction, use a channel topology optimized graph convolutional network to extract the features of the skeleton data;
[0026] The representation of the human skeleton is: G=(V, E, X); where V={v1, v2, …, ν N} is a set composed of N vertices; E is the edge set, represented as an adjacency matrix:
[0027] v i 's neighborhood is represented as N(v i )={v i |a ij ≠0}; the element a ij reflects the correlation strength between v i and v j ;
[0028] X is the feature set of N vertices, represented as a matrix
[0029] where C is the initial feature dimension, and the feature of v i is represented as
[0030] The channel topology optimized graph convolutional network modeling consists of spatial modeling, temporal modeling, and residual connections;
[0031] S42: Load the depression recognition model, pre-train the model with the NTU RGB+D human action recognition dataset, and fine-tune it on the free dataset. During the fine-tuning process, adjust the input layer of the model to match the dimension of the new data, and modify the last layer of the classifier to set the output dimension to 2 to achieve a binary classification task. The binary classification task is depression or normal;
[0032] S43: Model effect evaluation. The model was evaluated using the five-fold cross-validation method. The dataset was randomly divided into five subsets of equal size. Each subset was used as a test set in turn, and the remaining four subsets were combined as training sets. The depression recognition model was trained on the training set and evaluated on the test set. The evaluation results were calculated and averaged five times. The positive and negative sample discriminant features of the last layer of the depression recognition model were extracted. The features of the five test sets were concatenated together, and the difference between the positive sample group and the negative sample group was calculated.
[0033] Preferably, the spatial modeling of step S41 is composed of three parallel channel topology optimization graph convolution modules, specifically including the following steps:
[0034] S411: Linear variation of features:
[0035] Where W represents the learnable weight matrix, and C' represents the output feature dimension;
[0036] S412: Perform channel topology modeling, use an adjacency matrix Α as the basic topology of all channels, capture the association of global joints; use linear transformations φ and ψ to reduce the feature dimension from C' to C low , reducing the amount of calculation; for joints (ν i ,ν j ), through the subtraction form q ij =ξ(σ(ψ(xi)-φ(xj))) and MLP form q ij =ξ(MLP(ψ(xi)||φ(xj))) Calculate channel correlation q ij ;
[0037] Where ξ is a linear transformation of dimension increase, σ is an activation function, and || represents a concatenation operation;
[0038] Combine the basic topology with the channel correlation to generate the channel topology: R = A + α·Q;
[0039] Where α is a learnable scaling factor, is the channel-specific correlation tensor;
[0040] S413: Perform channel feature aggregation. For channel c, use the corresponding topology Perform feature aggregation:
[0041] in, is the feature vector of the cth channel;
[0042] Concatenate the aggregation results of all channels into the final output: Z = [Z1||Z2||…‖Z C' ].
[0043] Preferably, in the temporal modeling of step S41, multi-scale convolution is adopted, which includes four branches. Each branch contains a 1×1 convolution to reduce the channel dimension, and temporal convolutions with different dilation rates (1 / 2 / 3) are used to capture short-term, medium-term, and long-term actions. The results of the four branches are concatenated as the output.
[0044] Preferably, in the residual connection of step S41, the module retains the original input information to alleviate the vanishing gradient: Y = CTR-GC(X) + X;
[0045] The entire channel topology optimized graph convolutional network is cascaded by ten basic blocks, and the number of channels gradually increases (64 → 64 → 64 → 64 → 128 → 128 → 128 → 256 → 256 → 256). The deep network captures high-order spatio-temporal features.
[0046] Preferably, the processing method of the depression scoring task depends on the data situation:
[0047] Sufficient and evenly distributed data: Replace the classifier with a regressor and directly perform regression prediction on the depression score;
[0048] Limited and unevenly distributed data: Manually divide the data into depressed samples and normal samples first, and continue to perform the binary classification task. The probability value of the sample output by the depression recognition model belonging to the depressed category is used as the basis for the depression score.
[0049] Therefore, the present invention adopts the above-mentioned non-intrusive depression recognition method based on gait behavior analysis and face recognition. Based on gait behavior analysis, a gait recognition method for skeletons is proposed. The walking video of the subject is captured using Microsoft Kinect, the skeleton data during the subject's walking process is extracted, the gait behavior features are extracted, and the action recognition based on the skeleton is performed. Furthermore, an analysis model can be established to identify the depression risk population; and a key population marking method based on face recognition technology is proposed, which can mark the subject with face recognition technology without disturbing the subject's walking, and further collect and identify data for the key population.
[0050] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0051] Figure 1 It is a flowchart of a non-intrusive depression recognition method based on gait behavior analysis and face recognition of the present invention;
[0052] Figure 2 It is a flowchart of the face recognition method in the embodiment of the present invention;
[0053] Figure 3It is the basic module of the channel topology optimized graph convolutional network used in the embodiments of the present invention;
[0054] Figure 4 It is the full-automatic acquisition flow chart in the embodiments of the present invention;
[0055] Figure 5 It is the training result of screening part of the data for binary classification in the embodiments of the present invention. Specific embodiments
[0056] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0057] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0058] The terms "including" or "comprising" and the like used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0059] Embodiment
[0060] As Figure 1 shown, a non-intrusive depression recognition method based on gait behavior analysis and face recognition specifically includes the following steps:
[0061] S1: Face recognition, input the face information of the subject, extract the face features and compare them with the key acquisition objects in the database until it is determined as the acquisition object to start collecting the skeletal data of the subject; In step S1, it specifically includes the following steps, as Figure 2 shown:
[0062] S11: Input face information, use Dlib to detect the face from the camera, and input the face picture and store it locally;
[0063] S12: Facial feature extraction. For each subject, based on multiple input face images, extract the 128D feature vector of each face, calculate the mean of the 128D feature vectors of the subject, detect 68 feature points of the face image through the ResNet-34 deep learning model, map the 68 feature points into a 128-dimensional vector space, and the vectors in the 128-dimensional vector space are used for face recognition and comparison. The Euclidean distance is used to calculate the distance between two faces:
[0064]
[0065] S13: Facial feature comparison. Detect the face from the real-time input camera video stream, extract its 128D feature vector, compare it with the feature vectors in the stored face dataset, calculate the Euclidean distance, and determine whether it is the same face. The judgment criterion is:
[0066] If the Euclidean distance between the vector spaces of two faces exceeds the discrimination threshold, it is considered that they are not the same person;
[0067] If the Euclidean distance between the vector spaces of two faces is less than the discrimination threshold, it is considered that they are the same person.
[0068] S2: Capture skeletal data. Select subjects covering a large age range, record natural walking videos, and extract the joint coordinates of the subjects at each time node to form a dataset;
[0069] In step S2, screen the natural walking video data of the subjects to ensure that the ages of all subjects are normally distributed. Use the Kinect SDK to extract 25 joint coordinates of the subjects. Screen the natural walking video data of the subjects to ensure that the ages of all subjects are normally distributed. Use 6 to prevent collecting two batches of data with Kinect devices in different positions. Use the Kinect SDK to extract 25 joint coordinates of the subjects;
[0070] The differences among the subjects in the first batch of data are relatively large. After data cleaning and preprocessing, a total of 30 people are selected from the first batch of data. Among them, those with a scale value greater than or equal to 10 points are positive samples, a total of 15 people, 90 pieces of data, and those with a scale value less than 1 point are negative samples, a total of 14 people, 82 pieces of data. The second batch of data is more balanced. According to the evaluation criteria of the Self-Rating Depression Scale, samples with an SDS score greater than or equal to 42.4 and a PHQ-9 score greater than or equal to 5 are taken as positive samples with depressive tendencies, a total of 39 people; there are 75 people who can obtain the minimum values of SDS and PHQ-9 at the same time. Randomly select 39 people as negative samples to ensure the balance of positive and negative samples. The specific steps are as follows:
[0071] S21: As Figure 4The automatic capture process is used as follows. When the subject enters the face recognition range and completes face recognition at the entrance, if the subject, data information is collected; if not, no collection is performed.
[0072] S22: Install a face collection module at the exit. After the subject walks out of the collection range and the face collection module recognizes the face information, the collection ends.
[0073] The discrimination threshold in this embodiment is set to 0.6.
[0074] S3: Data preprocessing, including data cleaning, coordinate standardization, time alignment, and data augmentation for the joint coordinates of the subject.
[0075] In step S3, the data preprocessing specifically includes the following steps:
[0076] S31: Remove invalid data in the data, including unnecessary data, duplicate values, and outliers, for data cleaning.
[0077] S32: With the center of the human body as the origin, normalize the joint coordinates to the interval [-1, 1] to eliminate the problem of inconsistent spatial coordinates of the subject relative to the camera and the influence of different human body sizes.
[0078] S33: For the time series data collected by different cameras, perform time alignment through frame sampling and time interpolation.
[0079] S34: Enhance the robustness through random rotation (±30°), shifting (±0.1 ratio), and joint loss simulation for data augmentation, increase the data volume of relevant data, and improve the performance of the trained model.
[0080] S4: Model classification, extract the features of the subject's joint coordinate data, and construct a depression recognition model for the data features.
[0081] In step S4, it specifically includes the following steps:
[0082] S41: Feature extraction, use a channel topology optimized graph convolutional network to extract the features of the skeleton data.
[0083] The representation of the human skeleton is: G=(V, E, X); where V={ν1, v2,…, v N} is a set composed of N vertices; E is the edge set, represented as an adjacency matrix:
[0084] v i The neighborhood of is represented as N(v i )={v i |a ij ≠0}; the element aij reflects v i and v j the correlation strength between;
[0085] X is the feature set of N vertices, represented as a matrix
[0086] where C is the initial feature dimension, and the feature representation of v i is
[0087] The channel topology optimization graph convolutional network modeling consists of spatial modeling, temporal modeling, and residual connections, and the basic modules are as Figure 3 shown.
[0088] The spatial modeling of step S41 consists of three parallel channel topology optimization graph convolutional modules, which specifically include the following steps:
[0089] S411: Perform a linear transformation on the features:
[0090] where W represents the learnable weight matrix, and C' represents the output feature dimension;
[0091] S412: Perform channel topology modeling. Use an adjacency matrix Α as the basic topology for all channels to capture the associations of global joints; Use linear transformations φ and ψ to reduce the feature dimension from C' to C low , reducing the computational load; For the joint (v i , v j ), calculate the channel correlation q through the subtraction form q ij = ξ(σ(ψ(xi)-φ(xj))) and the MLP form q ij = ξ(MLR(ψ(xi)||φ(xj))); ij ;
[0092] where ξ is the upsampling linear transformation, σ is the activation function, and || represents the concatenation operation;
[0093] Combine the basic topology with the channel correlation to generate the channel topology: R = A + α·Q;
[0094] where α is the learnable scaling coefficient, is the channel-specific correlation tensor;
[0095] S413: Perform channel feature aggregation. For channel c, use the corresponding topology to perform feature aggregation:
[0096] where, is the feature vector of the c-th channel;
[0097] Concatenate the aggregation results of all channels into the final output: Z = [Z1||Z2||…‖Z C' ].
[0098] In the temporal modeling of step S41, multi-scale convolution is used, which includes four branches. A single branch contains a 1×1 convolution to reduce the channel dimension. Temporal convolutions with different expansion rates (1 / 2 / 3) are used to capture short-term, medium-term and long-term actions. The results of splicing the four branches are output.
[0099] In the residual connection of step S41, the module retains the original input information and alleviates the gradient disappearance: Y = CTR-GC(X)+X;
[0100] The entire channel topology optimization graph convolutional network is composed of ten basic blocks cascaded, and the number of channels gradually increases (64→64→64→64→128→128→128→256→256→256), and the deep network captures high-order spatiotemporal features.
[0101] S42: Load the depression recognition model, pre-train the model with the NTU RGB+D human action recognition dataset, and fine-tune it on the free dataset. During the fine-tuning process, adjust the input layer of the model to match the dimension with the new data, modify the last layer of the classifier, set the output dimension to 2, and implement the binary classification task, which is depression or normal;
[0102] How the depression rating task is handled depends on the data:
[0103] The data is sufficient and evenly distributed: replace the classifier with a regressor and directly perform regression prediction on the depression score;
[0104] The data is limited and unevenly distributed: First, manually divide the data into depression samples and normal samples, and then continue to perform the binary classification task. The probability value of the sample output by the depression recognition model belongs to the depression category as the basis for the depression score.
[0105] S43: Model effect evaluation. The model was evaluated using the five-fold cross-validation method. The dataset was randomly divided into five subsets of equal size. Each subset was used as a test set in turn, and the remaining four subsets were combined as training sets. The depression recognition model was trained on the training set and evaluated on the test set. The evaluation results were calculated and averaged five times. The positive and negative sample discriminant features of the last layer of the depression recognition model were extracted. The features of the five test sets were concatenated together, and the difference between the positive sample group and the negative sample group was calculated.
[0106] In the first batch of collected data, 30 people’s data were selected for binary classification. The training results are as follows Figure 5As shown, as the number of training rounds increases, the loss value continuously decreases and converges to around 0.5. The accuracy of the training set converges to around 80%, and the classification accuracy of the test set converges to around 73%.
[0107] For the second batch of data, the number of positive and negative samples is 39 each. In this example, the five-fold cross-validation method is used to evaluate the model. The data set is randomly divided into five subsets of the same size. Each subset is taken turns as the test set alone, and the remaining four subsets are combined as the training set. The model is trained on the training set and evaluated on the test set. The evaluation results of each time are calculated and the average value of the five times is taken. The evaluation results are shown in Table 1.
[0108] Table 1
[0109]
[0110] The accuracy rate is 64.9%, which represents the proportion of samples correctly classified by the model in the total samples. The AUC is 63.3%, which is the area under the ROC curve and represents the ability of the model to distinguish positive and negative samples. The sensitivity is 60.8%, which represents the ability of the model to correctly identify positive samples. The specificity is 66.8%, which represents the ability of the model to correctly identify negative samples. The F1 score is 62.9%, which is the harmonic mean of the accuracy rate and the sensitivity. The precision is 70.5%, which represents the proportion of positive samples correctly identified by the model among all samples predicted as positive samples. Extract the discriminant features of positive and negative samples in the last layer of the model, splice the features of the five test sets together, and calculate the significant difference between the positive sample group and the negative sample group. The results are shown in Table 2.
[0111] Table 2
[0112]
[0113] The p-value of the Levene's test for equality of variances is 0.047 < 0.05, indicating that there are significant differences in the variances of different sample groups; the t-test results also show that the p-value is 2.595×10 -6 (two-tailed), far less than 0.05, indicating that the mean difference between the positive and negative sample groups is statistically significant. Specifically, the mean difference is 0.481, and its 95% confidence interval is (0.379, 0.583), and both the upper and lower limits are greater than zero, further confirming the effectiveness of the classification effect
[0114] Therefore, the present invention adopts the above-mentioned non-intrusive depression recognition method based on gait behavior analysis and face recognition. By collecting the walking gait data of the subject and combining the channel topology optimized graph convolutional network with the face recognition algorithm, non-invasive and non-intrusive depression recognition is realized. The above results verify the provided bone data feature extraction method and depression classification model, which have strong accuracy and specificity.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An unobtrusive depression recognition method based on gait behavior analysis and facial recognition, characterized in that: Specifically, it includes the following steps: S1: Facial recognition. Enter the face information of the subject, extract the face features and compare them with the key collection objects in the database until it is determined as a collection object and then start collecting the skeletal data of the subject; S2: Capture skeletal data. Screen the subjects covering a large age range, record the natural walking video, and extract the joint coordinates of the subject at each time node to form a data set; S3: Data preprocessing. Clean the data, standardize the coordinates, align the time, and augment the data for the joint coordinates of the subject; S4: Model classification. Extract the features of the subject's joint coordinate data and build a depression recognition model for the data features.
2. The non-intrusive depression recognition method based on gait behavior analysis and face recognition according to claim 1, characterized in that: In step S1, it specifically includes the following steps: S11: Enter face information. Use Dlib to detect the face from the camera and store the face picture locally; S12: Extract face features. For the subject, according to the multiple face pictures entered, extract the 128D feature vector of each face, calculate the mean value of the 128D feature vectors of the subject, detect 68 feature points of the face image through the ResNet-34 deep learning model, map the 68 feature points to a 128-dimensional vector space, and the vectors in the 128-dimensional vector space are used for face recognition and comparison. The Euclidean distance is used to calculate the distance between two faces: S13: Compare face features. Detect the face from the real-time input camera video stream and extract its 128D feature vector, compare it with the feature vectors in the stored face data set, calculate the Euclidean distance, and determine whether it is the same face. The judgment criterion is: If the Euclidean distance between the vector spaces of two faces exceeds the discrimination threshold, it is considered not the same person; If the Euclidean distance between the vector spaces of two faces is less than the discrimination threshold, it is considered the same person.
3. The non-intrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 2, characterized in that: In step S2, screen the natural walking video data of the subjects to ensure that the ages of all subjects are normally distributed. Use the Kinect SDK to extract 25 joint coordinates of the subjects. The specific steps are as follows: S21: Use the automatic capture process. When the subject enters the facial recognition range and completes facial recognition at the entrance, if it is a subject, data information is collected; if not, it is not collected; S22: Install a face collection module at the exit. After the subject walks out of the collection range and the face collection module recognizes the face information, the collection ends.
4. The non-intrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 3, wherein: In step S3, the data preprocessing specifically includes the following steps: S31: Remove the invalid data in the data. The invalid data includes unnecessary data, duplicate values, and outliers, and perform data cleaning; S32: Eliminate the problem of inconsistent spatial coordinates of the subject relative to the camera and the influence of different body sizes by normalizing the joint coordinates; S33: For the time series data collected by different cameras, perform time alignment through frame sampling and time interpolation; S34: Perform data augmentation through rotation, shifting, and joint loss simulation to increase the data volume of relevant data.
5. A non-intrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 4, characterized in that: In step S4, it specifically includes the following steps: S41: Feature extraction. Use the channel topology optimized graph convolutional network to extract the features of the skeleton data; The representation of the human body skeleton is: G = (V, E, X); where V = {v1, v2, …, v N} is a set composed of N vertices; E is the edge set, represented as an adjacency matrix: v i The neighborhood of i )={v i |a ij ≠0}; element a ij Reflects v i and v j The strength of the correlation between X is a feature set of N vertices, represented as a matrix Among them, C is the initial feature dimension, and the feature representation of v i is The channel topology optimization graph convolutional network modeling consists of spatial modeling, temporal modeling and residual connection; S42: Load the depression recognition model, pre-train the model with the NTU RGB+D human action recognition dataset, and fine-tune it on the free dataset. During the fine-tuning process, adjust the input layer of the model to match the dimension with the new data, modify the last layer of the classifier, set the output dimension to 2, and implement the binary classification task, which is depression or normal; S43: Model effect evaluation. The model was evaluated using the five-fold cross-validation method. The dataset was randomly divided into five subsets of equal size. Each subset was used as a test set in turn, and the remaining four subsets were combined as training sets. The depression recognition model was trained on the training set and evaluated on the test set. The evaluation results were calculated and averaged five times. The positive and negative sample discriminant features of the last layer of the depression recognition model were extracted. The features of the five test sets were concatenated together, and the difference between the positive sample group and the negative sample group was calculated.
6. The non-intrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 5, characterized in that: The spatial modeling of step S41 is composed of three parallel channel topology optimization graph convolution modules, which specifically include the following steps: S411: Perform a linear transformation on the feature: Where W represents the learnable weight matrix, and C' represents the output feature dimension; S412: Perform channel topology modeling. Use an adjacency matrix Α as the basic topology for all channels to capture the associations of global joints; use linear transformations φ and ψ to reduce the feature dimension from C' to C, reducing the computational load. For joints (v low , v i , v j ), calculate the channel correlation q through the subtraction form q ij = ξ(σ(ψ(xi) - φ(xj))) and the MLP form q ij = ξ(MLP(ψ(xi) || φ(xj))); ij Where ξ is a dimensional linear transformation, σ is an activation function, and ∥ represents a concatenation operation; Combine the basic topology with the channel correlation to generate the channel topology: R = A + α·Q; where α is a learnable scaling factor, is a channel-specific correlation tensor; S413: Perform channel feature aggregation. For channel c, use the corresponding topology Perform feature aggregation: Among them, is the feature vector of the c-th channel; Concatenate the aggregation results of all channels into the final output: Z = [Z1||Z2||…||Z C' .
7. An unobtrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 5, characterized in that: In the temporal modeling of step S41, multi-scale convolution is used, which includes four branches. A single branch contains a 1×1 convolution to reduce the channel dimension. Temporal convolutions with different expansion rates (1 / 2 / 3) are used to capture short-term, medium-term and long-term actions. The results of splicing the four branches are output.
8. An unobtrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 5, characterized in that: In the residual connection of step S41, the module retains the original input information and alleviates the gradient disappearance: Y = CTR-GC(X)+X; The entire channel topology optimization graph convolutional network is composed of ten basic blocks cascaded, and the number of channels gradually increases (64→64→64→64→128→128→128→256→256→256), and the deep network captures high-order spatiotemporal features.
9. The non-intrusive depression recognition method based on gait behavior analysis and facial recognition according to claim 5, characterized in that: How the depression rating task is handled depends on the data: The data is sufficient and evenly distributed: replace the classifier with a regressor and directly perform regression prediction on the depression score; The data is limited and unevenly distributed: First, manually divide the data into depression samples and normal samples, and then continue to perform the binary classification task. The probability value of the sample output by the depression recognition model belongs to the depression category as the basis for the depression score.