Apparatus and method for depression risk assessment based on fusion and enhanced isomorphic modal representation
By employing a device based on fusion and enhancement of isomorphic modal representations in the diagnosis of depression, utilizing feature extraction and cross-modal fusion modules, combined with a boundary loss function, the problem of lacking global features and intra-class discriminative features in existing technologies is solved, achieving higher accuracy and robustness in depression risk assessment.
Patent Information
- Application Number
- CN202311419878.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-10-30
AI Technical Summary
Existing technologies lack global features, feature fusion between isomorphic modalities, and intraclass distinguishing features in the diagnosis of depression, resulting in a lack of objectivity and accuracy in diagnostic results.
A depression risk assessment device based on fusion and enhancement of isomorphic modal representation is adopted. The feature extraction module extracts multiple isomorphic modal abstract features from gait skeleton data, and the feature is fused through a cross-modal fusion module. The boundary loss function is used to constrain the distance between feature centers in the feature space to enhance the unique features of intra-class modalities.
It improved the accuracy of depression risk assessment, enhanced the feature differentiation of isomorphic modalities and the unique features of intra-class modalities, and improved the assessment accuracy and robustness of the model.
Smart Images

Figure CN117497186B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a device and method for assessing depression risk based on the fusion and enhancement of isomorphic modal representations. Background Technology
[0002] A definitive diagnosis of depression requires a psychiatrist to conduct a systematic consultation, mental examination, and supplementary tests, such as the Hamilton Depression Rating Scale (HAMD) and the Personal Assessment of Depression Screening Scale (PHQ9). Besides high scores on these scales, further assessment is made based on the degree of psychological distress and the patient's performance in daily work and life activities. For example, if the psychological distress is intense and accompanied by significant physical symptoms such as fatigue, sleep disturbances, decreased appetite, and chronic pain, the diagnosis may be more objective. Therefore, artificial intelligence-based depression diagnostic models have emerged.
[0003] Currently, artificial intelligence methods for diagnosing depression include those based on audio and video, physiological signals, and social media data. These methods can provide objective diagnostic results for depression to some extent, but due to the difficulty in collecting patient data, they all face a lack of training and testing data.
[0004] Numerous studies have shown that depression can lead to pathological gait behavior. Currently, spatiotemporal graph convolutional models are widely used as the basic model in fields such as depression risk assessment, action-emotion recognition, and action recognition based on gait skeleton data. However, this model suffers from problems such as a lack of global features, a lack of feature fusion between isomorphic modalities, and a lack of intra-class discriminative features among isomorphic modalities. Summary of the Invention
[0005] This application provides a depression risk assessment device based on the fusion and enhancement of isomorphic modal representations to address the problems in the prior art, such as the lack of global features, the lack of feature fusion between isomorphic modalities, and the lack of intra-class distinguishing features of isomorphic modalities.
[0006] Accordingly, this application also provides a method for assessing depression risk based on the fusion and enhancement of isomorphic modal representations, to ensure the implementation and application of the above method.
[0007] To address the aforementioned technical problems, this application discloses a depression risk assessment device based on the fusion and enhancement of isomorphic modal representations, the device comprising:
[0008] The feature extraction module is used to extract abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer; wherein, there are one or more network layers;
[0009] A cross-modal fusion module, located in one or more network layers, is used to fuse abstract features from multiple isomorphic modalities.
[0010] The loss calculation module is used to constrain the distance between feature centers of multiple isomorphic modalities in the feature space using the boundary loss function;
[0011] The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanism modules employ a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively.
[0012] This application also discloses a method for assessing depression risk based on the fusion and enhancement of isomorphic modal representations, implemented using the apparatus described in any of the above claims, the method comprising:
[0013] The feature extraction module extracts abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer; wherein the network layer has one or more layers.
[0014] By utilizing cross-modal fusion modules located in one or more network layers, abstract features of multiple isomorphic modalities can be fused.
[0015] The boundary loss function in the loss calculation module is used to constrain the distance between the feature centers of multiple isomorphic modalities in the feature space;
[0016] The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanism modules employ a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively.
[0017] In this embodiment, a feature extraction module is used to extract abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer. The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanisms employ multi-head self-attention mechanisms to model global features in the spatial and temporal dimensions respectively, enabling preliminary depression risk assessment. The network layer has one or more layers. A cross-modal fusion module is located in one or more of these layers to fuse abstract features of multiple isomorphic modalities, employing not only a post-fusion strategy but also intermediate feature fusion of isomorphic modalities. The loss calculation module uses a boundary loss function to constrain the distance between feature centers of multiple isomorphic modalities in the feature space, thereby enhancing the unique features of intra-class modalities, reducing similar features within intra-class modalities, and improving the device's assessment accuracy.
[0018] Additional aspects and advantages of the embodiments of this application will be set forth in the following description, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0020] Figure 1 A schematic diagram of a depression risk assessment device based on fusion and enhancement of isomorphic modal representations provided in an embodiment of this application;
[0021] Figure 2 A schematic diagram of a spatiotemporal Transformer encoder provided in an embodiment of this application;
[0022] Figure 3 Human body topology diagram provided for embodiments of this application;
[0023] Figure 4 A schematic diagram of a mask provided for an embodiment of this application;
[0024] Figure 5 A schematic diagram of a cross-modal fusion module provided in an embodiment of this application;
[0025] Figure 6 A schematic diagram of boundary loss provided for an embodiment of this application;
[0026] Figure 7 A schematic diagram of the individual emotional feature generation module provided in an embodiment of this application;
[0027] Figure 8 A flowchart of a depression risk assessment method based on fusion and enhancement of isomorphic modal representations provided in this application embodiment. Detailed Implementation
[0028] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0029] Those skilled in the art will understand that, unless explicitly stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0030] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0031] The solution provided in this application can be executed by any electronic device, such as a terminal device or a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. Regarding the technical problems existing in the prior art, the depression risk assessment device and method based on fusion and enhancement of isomorphic modal representation provided in this application aim to solve at least one of the technical problems of the prior art.
[0032] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0033] This application provides a possible implementation method, such as... Figure 1The diagram shows a schematic of a depression risk assessment device based on the fusion and enhancement of isomorphic modal representations. This scheme can be executed by any electronic device, and optionally, it can be executed on a server or a terminal device.
[0034] like Figure 1 As shown, the device may include the following modules:
[0035] The feature extraction module is used to extract abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer; wherein, there are one or more network layers;
[0036] A cross-modal fusion module, located in one or more network layers, is used to fuse abstract features from multiple isomorphic modalities.
[0037] The loss calculation module is used to constrain the distance between feature centers of multiple isomorphic modalities in the feature space using the boundary loss function;
[0038] The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanism modules employ a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively.
[0039] For gait skeleton data with temporal and spatial features, assuming there are T frames in the temporal sequence, and the corresponding gait skeleton joints in the spatial dimension are represented by a vertex set. This represents the 3D position coordinates of the skeleton joints at frame t, where N represents the number of joints. Optionally, in this embodiment, there are 17 joints, i.e., N=17, and a total of 64 frames, i.e., T=64. For each vertex... The feature dimension is set to C, with an initial value of 3, i.e., 3D coordinates. Therefore, the gait joint coordinates with T frames can be represented as follows: In addition, the orientation of the bones It can be calculated using the following method:
[0040]
[0041] in, This represents the adjacent nodes of the target node in a predefined adjacency matrix. This applies to two isomorphic modes. and The depression risk assessment device in this application embodiment is used to model gait skeletal data, thereby obtaining the prediction results for the test subject. .
[0042] The depression risk assessment device in this embodiment includes a feature extraction module, a cross-modal fusion module, and a loss calculation module. Inspired by the Transformer model in natural language processing tasks, the feature extraction module in this embodiment treats gait joint coordinates as discrete text words and uses the Transformer model to generate the basic device for depression risk assessment. This basic device mainly includes two parts: a spatial self-attention mechanism module and a temporal self-attention mechanism module. The multi-head self-attention mechanism in the spatial self-attention mechanism module and the multi-head self-attention mechanism in the temporal self-attention mechanism module are decoupled from each other.
[0043] Optionally, the multi-head attention mechanism in the spatial self-attention mechanism module can model global features in the spatial dimension of gait skeleton data. For example... Figure 1 The spatiotemporal Transformer encoder and Figure 2 As shown, using gait joint coordinates For example, during training or testing, the data of a certain batch (Batch size, Bs) is set as follows: To decouple the features from the time and space dimensions and facilitate the computation of multi-head self-attention mechanisms, the data format can be modified as follows: This can be considered as the batch size being T times the original. (Refer to...) Figure 2 The spatial multi-head self-attention mechanism is implemented as follows in the spatial self-attention mechanism module:
[0044]
[0045] , , ,
[0046] ,
[0047] ,
[0048] The feature embedding layer has 128 output channels. The number of headers is set to 4; , and The matrix represents the learnable parameters, with 128 input and output channels; the normalized exponential function is... .
[0049] Reference Figure 2The temporal mask multi-self-attention mechanism module (i.e., the temporal self-attention mechanism module) in the temporal self-attention mechanism module can model the global features of gait skeleton data in the temporal dimension through its multi-head attention mechanism. Similar to the feature transformation in the spatial dimension, it uses gait joint coordinates... For example, during training or testing, the data of a certain batch (Batch size, Bs) is set as follows: To decouple the features from the time and space dimensions and facilitate the computation of multi-head self-attention mechanisms, the data format has been modified to... This can be considered as the batch size being N times the original.
[0050] In this embodiment, a feature extraction module is used to extract abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer. The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanisms employ multi-head self-attention mechanisms to model global features in the spatial and temporal dimensions respectively, thus initially realizing the function of depression risk assessment. The network layer has one or more layers. A cross-modal fusion module is located in one or more of these network layers to fuse the abstract features of multiple isomorphic modalities. It not only performs a post-fusion strategy but also achieves intermediate feature fusion of isomorphic modalities. The loss calculation module uses a boundary loss function to constrain the distance between the feature centers of multiple isomorphic modalities in the feature space, thereby enhancing the unique features of intra-class modalities, reducing the similarity of intra-class modalities, and improving the model evaluation accuracy.
[0051] In an optional embodiment, the isomorphic modality in the gait skeleton data includes gait joint coordinates and bone orientation; the network layer includes two feature extraction modules, used to extract joint features corresponding to the joint coordinates as abstract features, and bone features corresponding to the bone orientation as abstract features.
[0052] In an optional embodiment, the spatial self-attention mechanism module uses the embedding features generated by the shortest distance and degree as the location encoding;
[0053] Wherein, the shortest distance is the shortest distance between any joint in the pre-generated topology graph and the other joints; the degree is the number of edges in the topology graph connected to the corresponding joint.
[0054] Because the Transformer model uses a self-attention mechanism to replace the human body topology and locality that graph convolutional networks rely on, the spatial location information of these key points is lost. The model then has no way of knowing the relative and absolute location information of each key point within the human body topology. Therefore, in this embodiment, the shortest distance (SD) and degree embedding features from the topological graph are introduced as location encoding, referring to... Figure 2 The shortest distance coding and degree coding in the model.
[0055] Human body topology diagram as follows Figure 3 There are a total of 17 joints. The shortest distance is the shortest distance from a given joint to all other joints in the topology graph. Table 1 below uses joints 1 and 2 as examples:
[0056] Table 1. Shortest distances between key points 1 and 2
[0057]
[0058] Since the longest distance between nodes in this topology diagram is 8, the input size of the embedding layer is set to 8, and the output channel size is set to 128. After embedding, the data format is as follows: ', replicate T times in the time dimension to facilitate addition with the embedded skeleton features.
[0059] Similarly, the degree features of key points 1 and 2 are shown in Table 2:
[0060] Table 2. Degrees of joints 1 and 2
[0061]
[0062] The degree in the topology graph represents the number of edges connected to that node, with a maximum value of 4. Similarly, after passing through an embedding layer with an input size of 4 and an output channel size of 128, it is added to the embedded skeleton features.
[0063] In an optional embodiment, the temporal self-attention mechanism module uses a masking mechanism to calculate the attention score of any keypoint in the temporal dimension;
[0064] The attention score of any keypoint in the time dimension is calculated based on the features of that keypoint at any given time point and the features prior to that time point.
[0065] Because the Transformer model uses a self-attention mechanism, the temporal information of these key points is lost, and the model has no way of knowing the relative and absolute position information of each key point in a time series. Therefore, this application's embodiment introduces a masking mechanism, implemented as follows:
[0066]
[0067] ,
[0068] in, This is a lower triangular matrix, where green represents 1 to indicate preservation and gray represents 0 to indicate a mask, such as... Figure 4 As shown, the features of a certain frame at a certain key point are used to calculate the attention score only with the features of its previous time points. Represents trainable random position encoding, such as Figure 2 As shown. The above method ensures that the multi-head self-attention mechanism only focuses on information prior to the current time point, guaranteeing temporal precedence.
[0069] Reference Figure 2 The feature extraction module in this embodiment further includes a residual and normalization layer (residual & normalization) and a feedforward layer. Based on the above embodiments, the process of the feature extraction module is described as follows:
[0070] ))
[0071] ))
[0072] in, The encoder is linearly stacked in L layers, with the output of the previous encoder serving as the input of the next encoder. Specifically, it is implemented in 9 layers. go through Figure 1 The spatiotemporal Transformer encoder in the model employs batch normalization (BN). The feedforward layer is constructed from a multilayer perceptron (MLP) with the GELU activation function.
[0073] Similarly, the feature extraction module corresponding to the bone orientation also adopts the exact same structure and has independent learnable parameters.
[0074] In an optional embodiment, the feature extraction module is constructed using a Transformer model; the cross-modal fusion module uses the decoder of the Transformer model to fuse abstract features from multiple isomorphic modalities.
[0075] Specifically, the feature extraction module is built using the encoder of the Transformer model, i.e. Figure 1 The spatiotemporal Transformer encoder in the model. To extract multimodal information, embodiments of this application utilize the decoder of the Transformer model to achieve intermediate feature fusion of isomorphic modalities. For example... Figure 1 The cross-modal fusion module and Figure 5Taking gait joint coordinates as an example, its specific implementation is as follows:
[0076] , , ,
[0077] ,
[0078] ,
[0079] in, (query) is generated from gait joint coordinate features through a linear layer mapping (for...). Figure 5 In ), and Mapped from skeletal features through linear layers (respectively) Figure 5 In If the skeleton orientation feature is used as input, then the subscripts in the above formula are swapped.
[0080] The encoding process of the spatiotemporal Transformer encoder with a cross-modal fusion module is modified as follows compared to the spatiotemporal Transformer encoder without a cross-modal fusion module:
[0081] ))
[0082] ))
[0083] Experiments have shown that the outputs provided by deeper layers of the model have higher-level semantics, which is helpful for modality fusion. Therefore, in the embodiment of this application, the feature extraction modules of the 6th, 7th, and 8th layers of the joint and skeletal model (specifically to...) Figure 1 The cross-modal fusion module is added to the spatiotemporal Transformer encoder to realize the interaction and fusion of multimodal high-level semantic features.
[0084] The feature extraction module in this embodiment extends the spatiotemporal graph convolution model to a Transformer-based model for 3D skeleton feature modeling, successfully applying it to the field of depression risk assessment. It overcomes the locality limitations of graph convolution and human topology, resulting in a depression risk assessment device with higher accuracy compared to graph convolution models.
[0085] In an optional embodiment, the feature centers of the isomorphic modality include the feature centers of the joint features of the healthy sample, the feature centers of the skeletal features of the healthy sample, the feature centers of the joint features of the depressed sample, the feature centers of the skeletal features of the depressed sample, the feature centers of the healthy sample, and the feature centers of the depressed sample.
[0086] The boundary loss function is used to calculate the multimodal boundary loss, which is obtained by calculating the first constraint loss, the second constraint loss, and the third constraint loss. The first constraint loss constrains the distance between the feature centers of the joint features of healthy samples and the feature centers of the skeletal features of healthy samples; the second constraint loss constrains the distance between the feature centers of the joint features of depressed samples and the feature centers of the skeletal features of depressed samples; the third constraint loss constrains the maximum value of the intra-class distances to be less than the inter-class distance. Specifically, the intra-modal distances include the distances between the feature centers of the joint features of healthy samples and the feature centers of the skeletal features of healthy samples, as well as the distances between the feature centers of the joint features of depressed samples and the feature centers of the skeletal features of depressed samples; the inter-class distance is the distance between the feature centers of healthy samples and the feature centers of depressed samples.
[0087] like Figure 6 As shown, the circles represent the encoded features of gait joint coordinates. The triangle represents the coded features of the skeleton. After the classification layer, the feature space is as follows: Figure 6 As shown above, the healthy and depressed samples are separable between classes. The intra-class boundaries of joint and skeletal features are more complex. Some samples have two features that are very close to each other, which means they lack the uniqueness of modal features.
[0088] To learn modality-specific information from gait skeleton data, this application proposes a novel multimodal margin loss function to widen the distance between two modality centers (joint features and skeletal features) in the feature space. , The purpose is to increase the uniqueness and diversity of modal features within a class, while ensuring the distinguishability between classes. That is, to expand... , The distance, and ensure .
[0089] The feature center can be described as follows:
[0090] , ,
[0091] , ,
[0092] ,
[0093] ,
[0094] Among these methods, the feature centers of each class are approximated using batch samples. This represents the number of samples in a particular batch. This represents the healthy samples in a particular batch. This represents samples from a particular batch that are classified as depressed. Feature centers representing joint features in healthy samples within a batch. Feature centers representing the skeletal features of healthy samples in a particular batch. Feature centers representing the joint features of depressed samples in a certain batch. Feature centers representing the skeletal features of depressed samples in a certain batch. The feature centers representing healthy samples in a certain batch, The feature center represents the depression samples in a certain batch.
[0095] In this embodiment, the multimodal boundary loss function can be expressed as follows:
[0096] , , ,
[0097] ,
[0098] ,
[0099] ,
[0100] + + ,
[0101] in, represent First constraint loss in feature space distance, This is a hyperparameter and can be set to 1. Third constraint loss. Constrain the distance between intra-class modes The maximum value in ) should be less than the inter-class distance. First constraint loss Second constraint loss Constrain the intra-class inter-modal distance for healthy and depressed samples respectively. Able to reach the desired distance and By minimizing the multimodal boundary loss This ensures the distinguishability between classes while increasing the uniqueness and diversity of modal features within a class.
[0102] In an optional embodiment, the device further includes an individual emotional feature generation module for embedding personal emotional features;
[0103] Personal emotional characteristics include individual differences and emotional characteristics; individual differences include age and gender;
[0104] Emotional features include angle, distance ratio, and area ratio; where angle refers to the angle between any two joints; distance ratio refers to the distance ratio between any joint and the other joints; and area ratio refers to the area ratio of any two groups of regions consisting of three joints.
[0105] like Figure 7 As shown, individual differences (age and gender) are the main factors affecting gait changes. Therefore, in this embodiment, the samples are divided into children (7-14 years old), young adults (15-35 years old), middle-aged (36-60 years old), and elderly (61 years old and above) according to age groups, and the corresponding characteristics are denoted as follows. Mapped through the embedding layer:
[0106]
[0107] This application also incorporates emotional features into gait-based depression risk assessment, as detailed below:
[0108] Angle: The angle formed by two joints at a third joint, such as between the head and neck (to calculate the degree of head tilt), between the neck and shoulders (to calculate whether there is hunchback), between the hip and thigh (to calculate stride), etc.
[0109] Distance ratio: The ratio of the distances from any joint point to the other joint points, for example, the ratio of the distance from the hand to the neck to the distance from the hand to the hip (for calculating the arm swing distance);
[0110] Area ratio: The ratio of the areas formed by any two groups of three joints. For example, the ratio of the area formed between the elbow and neck to the area formed between the elbow and hip (calculating arm swing). The area ratio can be viewed as a combination of features based on angle and distance ratios, supplementing these features.
[0111] The details are shown in Table 3 below:
[0112] Table 3. Emotional Characteristics
[0113]
[0114] The corresponding features are denoted as The three types of features are embedded as follows:
[0115]
[0116]
[0117]
[0118] like Figure 7 and Figure 1 As shown, supplementary high-level features The following is an explanation of the relationship between this feature and the main model's output features. splicing along the channel dimension:
[0119] )
[0120]
[0121] In an optional embodiment, the apparatus further includes a prediction output module; the prediction output module uses cross-entropy loss and preset labels to supervise the prediction output and obtain the prediction result.
[0122] Reference Figure 1 In this embodiment, the spatial self-attention mechanism module and the temporal self-attention mechanism module are stacked together, and the output after the last layer is processed... Figure 1 The average pooling layer, linear layer, and normalized exponential function in the model utilize cross-entropy loss. and preset tags Supervised prediction output As shown in the following formula:
[0123]
[0124] Among them, the prediction results .
[0125] The embodiments of this application take into account the individual differences of patients and the emotional characteristics of depressive gait, and successfully process and extract relevant features, thereby further improving the accuracy of depression risk assessment and enhancing the robustness of the device.
[0126] In this embodiment, to verify the effectiveness and advancement of the device, extensive experiments were conducted on the proposed depression risk assessment device based on fusion and enhancement of isomorphic modal representations. The dataset collected from Shenzhen People's Hospital included data from 656 subjects (358 healthy subjects and 298 diagnosed depression patients), with gait feature data extracted to obtain 9046 gait segments suitable for training. The results were compared with current advanced depression risk assessment models based on gait skeleton data. Automatic evaluation metrics (F1 Score, Accuracy) were used in the experiments. Experimental results show that the proposed embodiment outperforms the best current gait-based depression risk assessment models in terms of automatic evaluation metrics.
[0127] Furthermore, this application also conducted ablation experiments on a isomorphic modality feature fusion module based on a Transformer decoder, an isomorphic modality feature enhancement module based on metric learning, and an advanced feature embedding module. Under controlled module variables, all models with modules demonstrated better accuracy in depression risk assessment compared to models without modules.
[0128] Based on the same principles as the methods provided in the embodiments of this application, the embodiments of this application also provide a method for assessing depression risk based on the fusion and enhancement of isomorphic modal representations, implemented using the apparatus described in any of the above embodiments, such as... Figure 8 As shown, the method includes:
[0129] Step 801: Use the feature extraction module to extract abstract features of multiple isomorphic modalities in the gait skeleton data in the preset network layers; wherein, there are one or more network layers;
[0130] Step 802: Use the cross-modal fusion module located in one or more of the network layers to fuse the abstract features of multiple isomorphic modalities;
[0131] Step 803: Use the boundary loss function in the loss calculation module to constrain the distance between the feature centers of multiple isomorphic modalities in the feature space;
[0132] The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanism modules employ a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively.
[0133] In this embodiment, a feature extraction module is used to extract abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer. The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module. Both the spatial and temporal self-attention mechanisms employ multi-head self-attention mechanisms to model global features in the spatial and temporal dimensions respectively, thus initially realizing the function of depression risk assessment. The network layer has one or more layers. A cross-modal fusion module is located in one or more of these network layers to fuse the abstract features of multiple isomorphic modalities. It not only performs a post-fusion strategy but also achieves intermediate feature fusion of isomorphic modalities. The loss calculation module uses a boundary loss function to constrain the distance between the feature centers of multiple isomorphic modalities in the feature space, thereby enhancing the unique features of intra-class modalities, reducing the similarity of intra-class modalities, and improving the model evaluation accuracy.
[0134] The depression risk assessment method based on fusion and enhanced isomorphic modality representation provided in this application can achieve... Figures 1 to 7The various processes implemented in the device embodiment will not be described again here to avoid repetition.
[0135] The depression risk assessment method based on fusion and enhancement of isomorphic modal representation in this application can realize the function of the depression risk assessment device based on fusion and enhancement of isomorphic modal representation provided in this application. The implementation principle is similar. The steps in the depression risk assessment method based on fusion and enhancement of isomorphic modal representation in each embodiment of this application correspond to the actions performed by each module and unit in the depression risk assessment device based on fusion and enhancement of isomorphic modal representation in each embodiment of this application. For detailed functional descriptions of each step of the depression risk assessment method based on fusion and enhancement of isomorphic modal representation, please refer to the description of the corresponding depression risk assessment device based on fusion and enhancement of isomorphic modal representation shown above, which will not be repeated here.
[0136] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A depression risk assessment device based on the fusion and enhancement of isomorphic modal representations, characterized in that, The device includes: The feature extraction module is used to extract abstract features of multiple isomorphic modalities in gait skeleton data from a preset network layer; wherein, the network layer has one or more layers. The isomorphic modalities in the gait skeleton data include gait joint coordinates and bone orientation; the network layer includes two feature extraction modules, which are used to extract the joint features corresponding to the gait joint coordinates as the abstract features, and the bone features corresponding to the bone orientation as the abstract features, respectively. The feature centers of the isomorphic modality include the feature centers of joint features of healthy samples, the feature centers of skeletal features of healthy samples, the feature centers of joint features of depressed samples, the feature centers of skeletal features of depressed samples, the feature centers of healthy samples, and the feature centers of depressed samples. The boundary loss function is used to calculate the multimodal boundary loss, which is obtained by calculating the first constraint loss, the second constraint loss, and the third constraint loss. The first constraint loss is used to constrain the distance between the feature centers of the joint features of the healthy sample and the feature centers of the skeletal features of the healthy sample; The second constraint loss is used to constrain the distance between the feature centers of the joint features of the depressed sample and the feature centers of the skeletal features of the depressed sample; The third constraint loss is used to constrain the maximum value of the distance between intra-class modalities to be less than the inter-class distance; Wherein, the distance between intra-class modalities includes the distance between the feature centers of the joint features of the healthy sample and the feature centers of the skeletal features of the healthy sample, and the distance between the feature centers of the joint features of the depressed sample and the feature centers of the skeletal features of the depressed sample; the inter-class distance is the distance between the feature centers of the healthy sample and the feature centers of the depressed sample; A cross-modal fusion module, located in one or more of the network layers, is used to fuse abstract features of multiple isomorphic modalities; The loss calculation module is used to constrain the distance between the feature centers of various isomorphic modes in the feature space using a boundary loss function; The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module; both the spatial self-attention mechanism module and the temporal self-attention mechanism module adopt a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively. The spatial self-attention mechanism module uses embedded features generated by the shortest distance and degree as position encoding; wherein, the shortest distance is the shortest distance between any joint and the other joints in the pre-generated topology graph; the degree is the number of edges connected to the corresponding joint in the topology graph; the temporal self-attention mechanism module uses a masking mechanism to calculate the attention score of any joint in the time dimension; the attention score of any joint in the time dimension is the attention score calculated based on the features of the joint at any time point and the features before that time point.
2. The depression risk assessment device based on fusion and enhancement of isomorphic modal representations according to claim 1, characterized in that, The device also includes an individual emotional feature generation module for embedding personal emotional features; The personal emotional characteristics include individual differences and emotional characteristics; The individual differences include age and gender; The emotional features include angle, distance ratio, and area ratio; wherein, the angle refers to the angle formed by two joints at a third joint; the distance ratio refers to the distance ratio between any joint and the other joints; and the area ratio refers to the area ratio of any two groups of regions composed of three joints.
3. The depression risk assessment device based on fusion and enhancement of isomorphic modal representations according to claim 1, characterized in that, The device also includes a prediction output module; The prediction output module uses cross-entropy loss and preset labels to supervise the prediction output and obtain the prediction result.
4. The depression risk assessment device based on fusion and enhancement of isomorphic modal representations according to claim 1, characterized in that, The feature extraction module is constructed using the Transformer model; The cross-modal fusion module utilizes the decoder of the Transformer model to fuse abstract features of multiple isomorphic modalities.
5. A method for assessing depression risk based on the fusion and enhancement of isomorphic modal representations, implemented using the apparatus described in any one of claims 1-4, characterized in that, The method includes: The feature extraction module extracts abstract features of multiple isomorphic modalities from gait skeleton data in a preset network layer; wherein the network layer has one or more layers. The abstract features of multiple isomorphic modalities are fused using cross-modal fusion modules located in one or more of the network layers; The boundary loss function in the loss calculation module is used to constrain the distance between the feature centers of various isomorphic modes in the feature space; The feature extraction module includes a spatial self-attention mechanism module and a temporal self-attention mechanism module; both the spatial and temporal self-attention mechanism modules employ a multi-head self-attention mechanism to model global features in the spatial and temporal dimensions, respectively.
Citation Information
Patent Citations
Multi-mode depressive emotion recognition method and device
CN115641543A
Action recognition method of three-flow adaptive graph convolution model fusing joint capture
CN116343334A
Emotion recognition method based on individual gait and group feature fusion
CN116503946A
Road vehicle sensing method based on multi-sensor fusion
CN116625383A