Text-based gait depression detection method and device, equipment and medium

By fusing gait silhouette features with general text features, a depression model was constructed and trained, which solved the problem of low accuracy in existing gait depression detection and achieved higher detection accuracy and generalization ability.

CN120823967APending Publication Date: 2025-10-21BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511018620.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing methods for detecting gait depression suffer from low accuracy and fail to account for both general depressive characteristics and individual differences, thus reducing the accuracy and generalization ability of the test results.

Method used

By extracting sample gait silhouette features from multi-frame sample gait silhouette images and fusing them with sample general text features extracted in advance based on sample scales, an initial depression model is constructed. The sample personalized text features are used to guide model training, and gait silhouette images of target detection objects are obtained for depression detection.

Benefits of technology

It improves the accuracy of depression detection and the generalization ability of the model, enabling more precise identification of individual depressive mood changes and enhancing its adaptability to practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823967A_ABST
    Figure CN120823967A_ABST
Patent Text Reader

Abstract

The invention provides a text-based gait depression detection method and device, equipment and a medium, and effectively solves the problem of low accuracy of an existing gait depression detection mode. The method comprises the following steps: extracting sample gait silhouette features from multiple frames of obtained sample gait silhouette images, and fusing the sample gait silhouette features and sample general text features to obtain sample gait text features; collecting sample scales with personalized answers replied by the plurality of sample detection objects to extract sample personalized text features; constructing an initial depression model based on the sample gait text features, and guiding the training of the initial depression model through the sample individual text features to obtain a depression model; and acquiring multiple frames of to-be-detected gait silhouette images of a target detection object during gait walking, and detecting the multiple frames of to-be-detected gait silhouette images based on the depression model to obtain a depression detection result of the target detection object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of depression detection, and in particular to a text-based gait depression detection method, apparatus, device and medium. Background Art

[0002] Depression is a common, insidious, and long-term mental illness, and has become one of the leading causes of disability worldwide. Studies have found that people with depression exhibit distinct characteristic differences in their walking patterns, such as decreased stride length and speed, reduced limb swing, and a slouched posture. This suggests that gait has great potential for analyzing depression.

[0003] However, existing methods for detecting gait depression have several challenges: First, since depression labels are usually determined using the results of self-assessment scales, these labels are often single, static, and coarse-grained quantitative scores that cannot capture the user's multi-dimensional psychological state, resulting in low accuracy in the test results.

[0004] Second, gait data itself exhibits significant individual differences. Existing gait depression detection methods cannot account for both general depression characteristics and individual differences during the modeling process, further reducing the accuracy of the test results and the generalization ability and practical application adaptability of existing gait depression detection methods.

[0005] Therefore, it is of great practical significance to develop a gait recognition depression method that can integrate text information of the scale, enhance feature discrimination and improve the generalization ability of the model. Summary of the Invention

[0006] In view of this, the purpose of this application is to provide a text-based gait depression detection method, device, equipment and medium, which effectively solves the problem of low accuracy of existing gait depression detection methods.

[0007] The present application embodiment provides a text-based gait depression detection method, the method comprising: Extracting sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fusing the sample gait silhouette features with sample general text features pre-extracted based on a sample scale to obtain sample gait text features; the sample scale includes depression detection question-answering questions; Collecting sample scales with personalized answers from multiple sample test subjects to extract sample personalized text features from the sample scales with personalized answers; constructing an initial depression model based on the sample gait text features, and guiding the training of the initial depression model through the sample personality text features to obtain a depression model; A plurality of frames of silhouette images of a target detection object to be tested are obtained when the target detection object is walking, and a depression detection result of the target detection object is obtained by detecting the plurality of frames of silhouette images of the target detection object based on the depression model.

[0008] In combination with the first aspect, the present application provides a first possible implementation of the first aspect, wherein the step of guiding the training of the initial depression model by using the sample personality text features to obtain the depression model includes: Mapping the sample personality text features and the sample gait text features in the initial depression model to the same feature space; The sample personality text features and the sample gait text features are aligned and fused to guide the training of the initial depression model.

[0009] In combination with the first aspect, the embodiment of the present application provides a second possible implementation of the first aspect, wherein the aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model includes: Setting an alignment network, and inputting the sample personality text features and the sample gait text features into the alignment network; The sample personality text features and the sample gait text features are processed based on the alignment network to perform alignment.

[0010] In combination with the first aspect, the embodiment of the present application provides a third possible implementation of the first aspect, wherein the aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model includes: Setting multiple loss functions to constrain corresponding sample features based on each loss function in the multiple loss functions; The multiple loss functions are integrated to constrain the training of the initial depression model to obtain the depression model.

[0011] In combination with the first aspect, an embodiment of the present application provides a fourth possible implementation of the first aspect, wherein the step of detecting the multiple frames of gait silhouette images to be tested based on the depression model to obtain a depression detection result of the target detection subject includes: extracting the silhouette features of the gait to be measured corresponding to the multiple frames of the silhouette images to be measured based on the depression model; The gait silhouette feature to be measured is integrated with the preset general text feature to obtain a gait text feature to obtain the depression detection result based on the gait text feature.

[0012] In combination with the first aspect, the embodiment of the present application provides a fifth possible implementation of the first aspect, wherein the fusing of the sample gait silhouette feature with the sample general text feature pre-extracted based on a general scale to obtain the sample gait text feature includes: Pre-extracting sample general text features from the sample scale to input the sample general text features and the sample gait silhouette features into a preset fusion network; The fusion network is controlled to process the sample general text features and the sample gait silhouette features to obtain the sample gait text features.

[0013] In combination with the first aspect, the embodiment of the present application provides a sixth possible implementation of the first aspect, wherein obtaining multiple frames of silhouette images of the target detection object to be measured when the target detection object is walking includes: Extracting a valid gait video from a gait video of a target detection object, and converting the valid gait video to obtain multiple frames of valid gait images; The multiple frames of valid gait images are sorted to segment multiple frames of gait silhouette images to be measured from the sorted multiple frames of valid gait images.

[0014] In a second aspect, an embodiment of the present application provides a text-based gait depression detection device, the device comprising: A fusion module is used to extract sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fuse the sample gait silhouette features with the sample general text features extracted in advance based on the general scale to obtain sample gait text features; An extraction module, configured to collect sample scales with personalized answers from a plurality of sample test subjects, and extract sample personalized text features from the sample scales with personalized answers; A construction module, configured to construct an initial depression model based on the sample gait text features, and guide the training of the initial depression model through the sample personality text features to obtain a depression model; The detection module is used to obtain multiple frames of gait silhouette images to be tested when the target detection object is walking, and to detect the multiple frames of gait silhouette images to be tested based on the depression model to obtain a depression detection result of the target detection object.

[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of any one of the text-based gait depression detection methods are performed.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, executes the steps of any one of the text-based gait depression detection methods.

[0017] The embodiment of the present application provides a text-based gait depression detection method, which first extracts sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fuses the sample gait silhouette features with sample general text features pre-extracted based on a sample scale to obtain sample gait text features; the sample scale contains depression detection question-and-answer questions; secondly, a scale with personalized answers answered by multiple sample detection subjects is collected to extract sample personalized text features from the scale with personalized answers; then, an initial depression model is constructed based on the sample gait text features, and the sample personalized text features are used to guide the training of the initial depression model to obtain a depression model; finally, multiple frames of the target detection subject are obtained when gaiting. A gait silhouette image to be tested, and the depression detection result of the target detection object is obtained by detecting the multiple frames of gait silhouette images to be tested based on the depression model. The gait depression detection method provided in the present application constructs an initial depression recognition model based on a combination of sample text features and sample gait features, which solves the problem of low accuracy caused by the single method based on depression labels in the existing depression detection method. The initial depression model is trained based on sample individual text features to obtain a depression model, thereby taking into account the difference between general depression features and individual features, further ensuring the accuracy of the depression detection results output by the depression model of the present application, improving the generalization ability and practical application adaptability of the model, and realizing accurate detection of the target detection object. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A flowchart of a text-based gait depression detection method provided in an embodiment of the present application is shown; Figure 2 Shows the structural block diagram and training diagram of the depression model provided in the embodiment of the present application; Figure 3 A schematic diagram of a process for obtaining sample gait text features provided in an embodiment of the present application is shown; Figure 4The following is a structural block diagram of a text-based gait depression detection device provided in an embodiment of the present application; Figure 5 The figure shows a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0021] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0022] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0023] The main drawbacks of current depression detection methods are that they are unable to capture the user's multi-dimensional psychological state, and that gait data itself exhibits significant individual differences. This makes it impossible to account for both general depression characteristics and individual differences, reducing the accuracy of the test results and the generalization ability and practical application adaptability of existing gait depression detection methods.

[0024] Based on this, the embodiments of the present application provide a text-based gait depression detection method, device, equipment and medium, which are described below through examples.

[0025] Example 1 To facilitate understanding of this embodiment, a text-based gait depression detection method disclosed in the embodiment of this application is first described in detail. Figure 1The flowchart of a text-based gait depression detection method shown in FIG. 1 is a flowchart of a text-based gait depression detection method provided in the present application, and the method includes: S101. Extracting sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fusing the sample gait silhouette features with sample general text features pre-extracted based on a sample scale to obtain sample gait text features; the sample scale includes depression detection question-answering questions; S102, collecting sample scales with personalized answers from multiple sample test subjects, and extracting sample personalized text features from the sample scales with personalized answers; S103, constructing an initial depression model based on the sample gait text features, and guiding the training of the initial depression model through the sample personality text features to obtain a depression model; S104: Acquire multiple frames of silhouette images of a target detection subject when the target detection subject is walking, and detect the multiple frames of silhouette images of the target detection subject based on the depression model to obtain a depression detection result of the target detection subject.

[0026] In step S101, the embodiment of the present application pre-acquires a multi-frame sample gait silhouette image, and the multi-frame sample gait silhouette image can be generated by different sample detection objects, such as a diseased detection object that has been determined to have depression and a normal detection object that does not have depression. The multi-frame sample gait silhouette image is two-dimensional, with a total of 30-60 frames. The multi-frame sample gait silhouette image is arranged according to the corresponding time sequence to obtain a coherent gait silhouette, and the sample gait silhouette feature is extracted from the acquired multi-frame sample gait silhouette image. The sample gait silhouette feature is a combination of gait features in the multi-frame sample gait silhouette image. The sample gait silhouette features are all high-dimensional. The specific extraction method is to extract them through the GaitModel model, where Gait The model is based on CNN. The sample gait silhouette features are fused with the sample general text features extracted in advance based on the sample scale to obtain the sample gait text features. The sample general text features are a combination of general text features extracted from the general scale. The sample gait text features are all high-dimensional. The sample scale is pre-designed and contains multiple depression detection questions and answers. Whether there is a tendency to depression is detected in a text-based semantic form. The sample scale is based on the LLM text feature extraction model to extract sample general text features. The sample general text features are high-dimensional.

[0027] In the specific implementation process of step S101, there is an embodiment: Figure 3 As shown, the fusion of the sample gait silhouette feature and the sample general text feature pre-extracted based on the general scale to obtain the sample gait text feature includes: S1011, pre-extracting sample general text features from the sample table, and inputting the sample general text features and the sample gait silhouette features into a preset fusion network; S1012: Control the fusion network to process the sample general text features and the sample gait silhouette features to obtain the sample gait text features.

[0028] In steps S1011-S1012, the embodiment of the present application pre-extracts the sample universal text features from the sample scale based on the LLM text feature extraction model, wherein the text feature extraction model may also be other and is not limited here. After obtaining the sample universal text features, the sample universal text features and the sample gait silhouette features are input into a preset fusion network. The fusion network first reduces the dimension of the sample universal text features so that they are mapped to the same feature space as the sample gait silhouette features. The feature fusion model based on the attention mechanism, after the sample universal text features and the sample gait silhouette features are mapped to the same feature space, adopts an attention-based fusion strategy to integrate the two modalities, thereby improving the model's understanding of depression through interaction. Specifically, the gait features serve as keys (K) and values ​​(V), while universal text that is not related to the individual serves as a query (Q). The fusion process is represented by the following attention mechanism using formula (1): (1); in, represents the general text features of the sample after dimension adjustment, Represents the sample gait silhouette feature, d is used for scaling and is also the attention dimension. This is to scale to prevent the gradient from being too large. WQ, WK, and WV are learnable projection matrices that will input features. (Q) and (K and V) are mapped, represents the sample gait text feature obtained by fusing the sample general text feature with the sample gait silhouette feature. Softmax normalization is used to convert the similarity score into a probability distribution, which represents the degree of attention paid to each part. Formula (1) is used to pay attention to the key information (K) in the input feature according to the input semantic clue (Q) and extract relevant content (V) from it. Formula (1) can also be converted into: .

[0029] In step S102, the embodiment of the present application also collects sample scales with personalized answers answered by multiple sample test subjects, wherein the sample test subjects include diseased test subjects who have been determined to have depression and normal test subjects who do not have depression. Since different sample test subjects generate personalized answers when answering the sample scale, the sample scale with personalized answers and the sample scale without answers are subjected to feature extraction by the same text feature extraction model such as LLM to extract sample personalized text features from the sample scale with personalized answers. The sample personalized text features are a combination of personalized text features extracted from the sample scale with personalized answers. The sample personalized text features are high-dimensional. Based on the sample personalized text features, the sample depression detection results can be directly predicted and verified with the type of the sample test subject. If the sample depression detection results are accurate, then based on L MSE The loss function of the mean square error loss is rewarded, and if the sample depression detection result is inaccurate, it is based on L MSE The loss function is used to punish the sample, thereby ensuring the accuracy of the depression detection results. L MSE The purpose of this method is to establish a mapping between sample personality text features and actual depression scores, which serves as a basic supervision of the model's language understanding ability. The sample personality text features are not only used to predict the sample depression detection results, but also to guide the training process of the constructed initial depression model to obtain a depression model.

[0030] In step S103, the embodiment of the present application constructs an initial depression model based on the sample gait text features, and also guides the training of the initial depression model through the sample personality text features to obtain a depression model, such as Figure 2 Shown are the structural block diagram and training diagram of the depression model. The initial depression model has a feature extraction layer, a fully connected layer, a batch normalization layer, a pooling layer and a 1*1 convolution layer, wherein the batch normalization layer is a batch normalization layer, which is used to standardize the feature distribution. The initial depression model has a dual-branch gait text fusion module, which is used to learn the sample gait text features and the sample personality text features. By introducing scale question information and user answer content to guide the gait representation learning process, more accurate and comprehensive gait feature modeling is achieved. By guiding the initial depression model to align the semantic differences between the sample personality text features and the gait text features, it has a stronger perception of the depressive mood changes of each detection object, thereby obtaining the depression model.

[0031] The depression model includes a general text-gait fusion module, an individual text-gait fusion module and a depression detection module, wherein the general text-gait fusion module includes a fully connected layer and a normalization layer, which are not only used for dimensionality reduction, but also for fusing the sample gait silhouette features with the sample general text features extracted in advance based on a general scale to obtain the sample gait text features. The individual text-gait fusion module includes a batch normalization layer, a pooling layer, a fully connected layer and a 1*1 convolution layer, which are used to align the sample individual text features with the sample gait text features. The depression detection module is used to output the depression detection result.

[0032] In the specific implementation process of step S103, there is an embodiment in which the training of the initial depression model is guided by the sample individual text features to obtain the depression model, including: S1031, mapping the sample personality text features and the sample gait text features in the initial depression model to the same feature space; S1032: Align and fuse the sample personality text features and the sample gait text features to guide the training of the initial depression model.

[0033] In steps S1031-S1032, the embodiment of the present application converts the sample individual text features Fs and the sample gait text features F ig Mapped into a common feature space, the feature space is shared to accommodate multiple features and ensure compatibility of subsequent alignment. Specifically, the sample personality text features Fs It goes through a transformation process that includes batch normalization (BN) to normalize the feature distribution and stabilize training, and average pooling to reduce the dimension. T And obtain a compact feature representation, and a linear layer maps it to the required feature dimension, because the sample personality text features and sample gait text features F ig The feature dimensions of is a 4096-channel dimension, F ig is a 256-channel dimension), so dimension adjustment and scaling operations are required to align the two features. Therefore, after aligning the sample personality text features and the sample gait text features, the sample personality text features and the sample gait text features are fused to guide the training process of the initial depression model to obtain a depression model after the training is completed. By guiding the depression model to align the semantic differences between the sample personality text features and the gait text features, the depression model has a stronger perception of the depressive mood changes of each detection subject.

[0034] The individual text-gait fusion module is based on formula (2) to analyze the individual text features of the sample. Perform dimensionality reduction: (2); in, To map the sample individual text features into the same feature space, AvgPool is AveragePooling, which is used for average pooling. Represents the dimension of the adjusted sample personality text features, N is the batch size, C is the number of channels, Linear It is a linear layer that transforms the sample personality text features The channel dimension is adjusted to 2048. BN is batch normalization, which is used to standardize feature distribution, accelerate convergence and avoid gradient disappearance.

[0035] The individual text-gait fusion module is used to analyze the sample gait text features. F ig , a 1×1 convolutional layer is applied to adjust its feature dimension while preserving spatial relationships, followed by batch normalization (BN) to maintain numerical stability, which is expressed as: (3); in, To map the sample gait text features to the same feature space, Conv is a 1×1 convolution operation that only transforms the channel dimension (without changing the spatial structure), increasing the channel dimension of the sample gait text features to 2048 to match the channel dimension of the sample personality text features after dimensionality reduction.

[0036] In the specific implementation process of step S1032, there is an embodiment in which the aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model includes: A1. Setting an alignment network, and inputting the sample personality text features and the sample gait text features into the alignment network; A 2. Processing the sample personality text features and the sample gait text features based on the alignment network to perform alignment.

[0037] In steps A1-A2, the embodiment of the present application sets the alignment network based on the distribution difference between the sample personality text features and the sample gait text features. The alignment network is implemented based on the loss function of Wasserstein distance and is expressed as L WThe Wasserstein distance provides a smooth and effective optimization gradient even when the two feature distributions do not overlap. It is more stable than cosine or KL divergence, preventing the depression model from being dominated by individual noise. Essentially, it provides a "soft constraint" for the depression model, guiding it during training to extract individual characteristics from the common text features of the samples. Different test subjects exhibit significant individual differences in their textual expressions. For example, some describe "insomnia" while others describe "low mood." The Wasserstein distance loss can map these heterogeneous texts into a unified feature space, mitigating differences in expression and focusing on common indicators of depression, such as "sleep disorders" and "emotional abnormalities." The sample's individual text features and the sample's gait text features are two heterogeneous features. The sample's individual text features tend to be subjective, while the sample's gait text features tend to be objective. By calculating the "optimal transfer cost" of these two feature distributions, the Wasserstein distance helps the depression model learn their underlying relationships. For example, a depressed patient may simultaneously experience "fatigue" and "slow gait," effectively integrating cross-modal features.

[0038] In the specific implementation process of step S1032, there is another embodiment in which the aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model includes: B1. Setting multiple loss functions to constrain corresponding sample features based on each loss function in the multiple loss functions; B2. Integrating the multiple loss functions to constrain the training of the initial depression model to obtain the depression model.

[0039] In steps B1-B2, in order to ensure the accuracy of the depression model, the embodiment of the present application sets a multi-loss function mechanism, based on each loss function in the multi-loss function, respectively constraining the corresponding sample features, such as setting a loss function for the sample personality text feature and the sample gait text feature. L MSE The loss function of the mean square error loss is expressed as: and , used to strengthen the connection between it and the depression detection result. Among them, the sample gait text feature is set L MSE The goal is to evaluate the ability of the depression model to predict depression severity by combining gait information with general scale text information, and to verify whether the depression model can extract comprehensive information from behavior and text that is effective in representing depression. This embodiment of the application also constrains the training of the depression model by adding all loss functions, expressed as formula (4): (4); Where L is the sum of all loss functions.

[0040] In step S104, after the depression model is established and trained, multiple frames of gait silhouette images of the target subject are obtained while walking, with the multiple frames of gait silhouette images being 30-60 frames. These images are then input into the depression model. Based on the depression model, these multiple frames of gait silhouette images are tested to obtain a depression detection result for the target subject. The depression detection result is used to characterize the presence of depression in the target subject. If the depression detection result indicates that the target subject is likely to be depressed, then the target subject requires treatment. In practice, target clients who receive the depression detection result generally exhibit early signs of depression, such as decreased stride length and speed, reduced limb swing, and slouched posture. This application divides the text content of the universal scale into "individual-independent general question stem information" and "individual-specific answer content," introducing a semantic representation obtained by converting the universal scale content and designing a multi-level fusion mechanism to supervise and guide the extraction of gait features, prompting the depression model to learn the correspondence between gait behavior and psychological semantics in the feature space, thereby enhancing the gait representation's ability to understand depressive symptoms.

[0041] In the specific implementation process of step S104, there is an embodiment in which the step of obtaining multiple frames of silhouette images of the target detection object when the target detection object is walking includes: S10411. Extracting a valid gait video from the gait video of the target detection object, and converting the valid gait video to obtain multiple frames of valid gait images; S10412: Sort the multiple frames of valid gait images to segment multiple frames of gait silhouette images to be measured from the sorted multiple frames of valid gait images.

[0042] In steps S10411-S10412, the embodiment of the present application shoots the gait of the target detection object based on a camera or other means to obtain a gait video of the target detection object, and extracts a valid gait video from the gait video of the target detection object, that is, deletes the video in which the target object turns around or there are multiple people, and only retains the video of the target object's personal gait, that is, the video at this time generally contains at least 2 gait cycles, and the video at this time is a valid gait video, and converts the valid gait video to obtain multiple frames of valid gait images, generally 30-60 frames, and sorts the valid gait images according to the time sequence of the valid gait video, so as to segment multiple frames of gait silhouette images to be tested from the sorted multiple frames of valid gait images based on RGB image background segmentation technology.

[0043] In the specific implementation process of step S104, there is another embodiment: the step of detecting the multiple frames of gait silhouette images to be tested based on the depression model to obtain a depression detection result of the target detection object includes: S10421, extracting the gait silhouette features to be measured corresponding to the multiple frames of gait silhouette images to be measured based on the depression model; S10422. Fusing the gait silhouette feature to be measured with preset general text features to obtain gait text features to obtain the depression detection result based on the gait text features.

[0044] In steps S10421-S10422, the gait silhouette features to be tested corresponding to the multiple frames of gait silhouette images to be tested are extracted based on the built-in feature extraction layer of the depression model, and the gait silhouette features to be tested are fused with the preset general text features. The specific fusion method is the attention mechanism described in steps S1011-S1012. The preset general text features are extracted by the feature extraction layer of the depression model from a preset scale for evaluating depression that includes depression detection questions and answers, to obtain gait text features. The depression model can output the depression detection result based on the gait text features, thereby detecting the possibility of depression in the target detection object.

[0045] Example 2 This application also provides a text-based gait depression detection device, such as Figure 4 The block diagram of a multi-granularity depression risk identification device is shown. The functions implemented by this text-based gait depression detection device correspond to the steps of executing a text-based gait depression detection method on a terminal device. This device can be understood as a component of a server including a processor. The text-based gait depression detection device described in this application includes: A fusion module 401 is configured to extract sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fuse the sample gait silhouette features with sample general text features pre-extracted based on a general scale to obtain sample gait text features; Extraction module 402, for collecting sample tables with personalized answers from multiple sample test subjects, and extracting sample personalized text features from the sample tables with personalized answers; A construction module 403 is configured to construct an initial depression model based on the sample gait text features, and guide the training of the initial depression model through the sample personality text features to obtain a depression model; The detection module 404 is configured to obtain a plurality of frames of silhouette images of a target detection subject when the target detection subject is walking, and detect the plurality of frames of silhouette images of the target detection subject based on the depression model to obtain a depression detection result of the target detection subject.

[0046] In a feasible implementation, the building module includes: A mapping module, configured to map the sample personality text features and the sample gait text features in the initial depression model to the same feature space; An alignment module is used to align and fuse the sample personality text features and the sample gait text features to guide the training of the initial depression model.

[0047] In a feasible implementation manner, the building module further includes: An input module, configured to set up an alignment network and input the sample personality text features and the sample gait text features into the alignment network; A processing module is used to process the sample personality text features and the sample gait text features based on the alignment network to perform alignment.

[0048] In a feasible implementation, the building block also includes: A setting module, configured to set multiple loss functions to constrain corresponding sample features based on each loss function in the multiple loss functions; An integration module is used to integrate the multiple loss functions to constrain the training of the initial depression model to obtain the depression model.

[0049] In a feasible implementation, the detection module includes: A first extraction module is configured to extract, based on the depression model, features of the gait silhouettes to be tested corresponding to the multiple frames of gait silhouette images to be tested; The first fusion module is used to fuse the gait silhouette feature to be measured with the preset general text feature to obtain a gait text feature to obtain the depression detection result based on the gait text feature.

[0050] In a feasible implementation, the fusion module includes: A second extraction module is used to pre-extract sample general text features from the sample table, so as to input the sample general text features and the sample gait silhouette features into a preset fusion network; The control module is used to control the fusion network to process the sample general text features and the sample gait silhouette features to obtain the sample gait text features.

[0051] In a feasible implementation manner, the detection module further includes: A conversion module is used to extract a valid gait video from the gait video of the target detection object, and convert the valid gait video to obtain multiple frames of valid gait images; The segmentation module is used to sort the multiple frames of valid gait images to segment multiple frames of gait silhouette images to be measured from the sorted multiple frames of valid gait images.

[0052] Example 3 The present application also provides an electronic device, such as Figure 5 As shown, it includes: a processor 501, a memory 502 and a bus 503, the memory 502 stores machine-readable instructions executable by the processor 501, and when the electronic device is running, the processor 501 and the memory 502 communicate with each other through the bus 503, and when the machine-readable instructions are executed by the processor 501, any one of the steps of the text-based gait depression detection method is performed.

[0053] Example 4 The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of any one of the text-based gait depression detection methods.

[0054] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0055] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0056] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0057] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, platform server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0058] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A text-based gait depression detection method, characterized in that: The method comprises: Extracting sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fusing the sample gait silhouette features with sample general text features pre-extracted based on a sample scale to obtain sample gait text features; the sample scale includes depression detection question-answering questions; Collecting sample scales with personalized answers from multiple sample test subjects to extract sample personalized text features from the sample scales with personalized answers; constructing an initial depression model based on the sample gait text features, and guiding the training of the initial depression model through the sample personality text features to obtain a depression model; A plurality of frames of silhouette images of a target detection object to be tested are obtained when the target detection object is walking, and a depression detection result of the target detection object is obtained by detecting the plurality of frames of silhouette images of the target detection object based on the depression model.

2. The method according to claim 1, characterized in that The step of guiding the training of the initial depression model by using the sample personality text features to obtain the depression model includes: Mapping the sample personality text features and the sample gait text features in the initial depression model to the same feature space; The sample personality text features and the sample gait text features are aligned and fused to guide the training of the initial depression model.

3. The method according to claim 2, characterized in that The aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model includes: Setting an alignment network, and inputting the sample personality text features and the sample gait text features into the alignment network; The sample personality text features and the sample gait text features are processed based on the alignment network to perform alignment.

4. The method according to claim 2, characterized in that The aligning and fusing the sample personality text features and the sample gait text features to guide the training of the initial depression model further includes: Setting multiple loss functions to constrain corresponding sample features based on each loss function in the multiple loss functions; The multiple loss functions are integrated to constrain the training of the initial depression model to obtain the depression model.

5. The method according to claim 1, wherein The step of detecting the multiple frames of gait silhouette images to be tested based on the depression model to obtain a depression detection result of the target detection object includes: extracting the silhouette features of the gait to be measured corresponding to the multiple frames of the silhouette images to be measured based on the depression model; The gait silhouette feature to be measured is integrated with the preset general text feature to obtain a gait text feature to obtain the depression detection result based on the gait text feature.

6. The method according to claim 1, characterized in that The fusing of the sample gait silhouette feature and the sample general text feature extracted in advance based on the general scale to obtain the sample gait text feature includes: Pre-extracting sample general text features from the sample scale to input the sample general text features and the sample gait silhouette features into a preset fusion network; The fusion network is controlled to process the sample general text features and the sample gait silhouette features to obtain the sample gait text features.

7. The method according to claim 1, characterized in that The step of obtaining a plurality of frames of silhouette images of a target detection object when the target detection object is walking comprises: Extracting a valid gait video from a gait video of a target detection object, and converting the valid gait video to obtain multiple frames of valid gait images; The multiple frames of valid gait images are sorted to segment multiple frames of gait silhouette images to be measured from the sorted multiple frames of valid gait images.

8. A text-based gait depression detection device, characterized in that: The device comprises: A fusion module is used to extract sample gait silhouette features from the acquired multi-frame sample gait silhouette images, and fuse the sample gait silhouette features with the sample general text features extracted in advance based on the general scale to obtain sample gait text features; An extraction module, configured to collect a plurality of scales with personalized answers provided by sample test subjects, and extract sample personalized text features from the scales with personalized answers; A construction module, configured to construct an initial depression model based on the sample gait text features, and guide the training of the initial depression model through the sample personality text features to obtain a depression model; The detection module is used to obtain multiple frames of gait silhouette images to be tested when the target detection object is walking, and to detect the multiple frames of gait silhouette images to be tested based on the depression model to obtain a depression detection result of the target detection object.

9. An electronic device, characterized in that: include: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the text-based gait depression detection method according to any one of claims 1 to 7 are performed.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the text-based gait depression detection method according to any one of claims 1 to 7.