Gait Quantification Method and Device Based on Low-Resolution Monocular Camera

The fusion analysis of joint feature points and depth spatiotemporal characterization information obtained by low-resolution monocular cameras solves the problem of high-profile equipment and professionals dependence in the prior art, and realizes low-cost gait quantitative analysis.

CN115965994BActive Publication Date: 2025-07-29INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211643838.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-07-29
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

The existing gait analysis methods rely on high experimental equipment and professional medical personnel, limiting their wide application.

Method used

The gait quantization method based on a low-resolution monocular camera is adopted. By acquiring joint feature points in the target image, combining body size parameters and feature spatiotemporal representation learning modules, deep spatiotemporal representation information is extracted using long and short-term memory networks and one-dimensional convolutional networks, and fusion analysis is performed to achieve gait analysis.

Benefits of technology

It reduces the cost of gait analysis, realizes quantitative analysis of common gait indicators of target objects using only a monocular camera, and improves the usage scenarios of gait analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965994B_ABST
    Figure CN115965994B_ABST
Patent Text Reader

Abstract

The present invention provides a gait quantification method and device based on a low-resolution monocular camera. The method includes: obtaining image feature points corresponding to joints, and determining scale feature points of the image feature points; determining the depth spatio-temporal representations of the image feature points and the scale feature points; fusing the depth spatio-temporal representations of the image feature points and the scale feature points to obtain a fused feature, and inputting the fused feature into a gait analysis model to obtain a gait analysis result of a target object. The gait quantification method and device based on a low-resolution monocular camera provided by the present invention extract the depth spatio-temporal representation information of the image feature points and the scale feature points, fuse the extracted depth spatio-temporal representation information, and input the fused feature into the gait analysis model to implement the gait analysis process of the target object. It realizes the quantification analysis of common indexes of the gait of the target object only by using a monocular camera, and improves the application scenarios of gait analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a gait quantification method and device based on a low-resolution monocular camera. Background Art

[0002] More and more evidence from clinical trials and theories reveals the association between gait characteristics and cognition. A slow walking speed and unstable gait are closely related to inattentiveness in cognition and the weakening of motor functions caused by nervous system damage.

[0003] Most of the existing gait analysis methods rely on expensive experimental equipment and professional medical personnel, and the harsh conditions limit the wide application of gait analysis. Summary of the Invention

[0004] The present invention provides a gait quantification method and device based on a low-resolution monocular camera to solve the technical problem that most of the existing gait analysis methods rely on expensive experimental equipment and professional medical personnel, and the harsh conditions limit the wide application of gait analysis.

[0005] The present invention provides a gait quantification method based on a low-resolution monocular camera, including:

[0006] Obtaining image feature points corresponding to joints of a target object in a target image, where the joints are joints related to the gait of the target object;

[0007] Based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determining scale feature points of the image feature points;

[0008] Based on a feature spatio-temporal representation learning module, determining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0009] Fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and inputting the fused feature into a gait analysis model to obtain a gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0010] According to the gait quantification method based on a low-resolution monocular camera provided by the present invention, fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points includes:

[0011] Based on the attention mechanism algorithm, determine the correlation weight between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points;

[0012] Based on the correlation weight, fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points.

[0013] According to a gait quantification method based on a low-resolution monocular camera provided by the present invention, based on the feature spatio-temporal representation learning module, determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points, including:

[0014] Input the image feature points into a long short-term memory network to obtain the state sequence of the image feature points, and input the state sequence of the image feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the image feature points;

[0015] Input the scale feature points into a long short-term memory network to obtain the state sequence of the scale feature points, and input the state sequence of the scale feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the scale feature points.

[0016] According to a gait quantification method based on a low-resolution monocular camera provided by the present invention, obtain the image feature points corresponding to the joints of the target object in the target image, including:

[0017] Obtain the pixel coordinate points corresponding to the joints of the target object in the target image, and use the pixel coordinate points as the image feature points corresponding to the joints.

[0018] According to a gait quantification method based on a low-resolution monocular camera provided by the present invention, based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points, including:

[0019] Based on the image feature points and the body size parameters of the target object, determine the motion scale feature of the target object;

[0020] Based on the topological relationship between the motion scale feature and the image feature points, determine the scale feature points of the image feature points.

[0021] According to a gait quantification method based on a low-resolution monocular camera provided by the present invention, the joints of the target object related to gait include at least one of the left hip joint, right hip joint, left knee joint, right knee joint, left ankle joint, right ankle joint, left toe joint, and right toe joint.

[0022] The present invention also provides a gait quantification device based on a low-resolution monocular camera, including:

[0023] An acquisition module, configured to acquire image feature points corresponding to joints of a target object in a target image, where the joints are joints of the target object related to gait;

[0024] A topological relationship determination module, configured to determine scale feature points of the image feature points based on body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points;

[0025] A depth spatio-temporal characterization module, configured to determine the depth spatio-temporal characterization of the image feature points and the depth spatio-temporal characterization of the scale feature points based on a feature spatio-temporal characterization learning module, where the feature spatio-temporal characterization learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0026] An analysis module, configured to fuse the depth spatio-temporal characterization of the image feature points and the depth spatio-temporal characterization of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain a gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the method for gait quantification based on a low-resolution monocular camera as described above is implemented.

[0028] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the method for gait quantification based on a low-resolution monocular camera as described above is implemented.

[0029] The present invention also provides a computer program product, including a computer program, where when the computer program is executed by a processor, the method for gait quantification based on a low-resolution monocular camera as described above is implemented.

[0030] The gait quantification method and device based on a low-resolution monocular camera provided by the present invention determine the image feature points corresponding to the joints of the target object and the scale feature points of the image feature points, respectively extract the depth spatio-temporal representation information of the image feature points and the scale feature points, fuse the extracted depth spatio-temporal representation information, and input the fused features into a gait analysis model, thereby realizing the gait analysis process of the target object. The gait analysis method based on images realizes the quantitative analysis of common indexes of the gait of the target object only by using a monocular camera, reduces the cost of gait analysis, and improves the application scenarios of gait analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly describe the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1 is a schematic flowchart of the gait quantification method based on a low-resolution monocular camera provided by the present invention;

[0033] Figure 2 is a schematic structural diagram of the feature spatio-temporal representation learning module provided by the present invention;

[0034] Figure 3 is a schematic structural diagram of the one-dimensional convolutional network provided by the present invention;

[0035] Figure 4 is a schematic structural diagram of the device applying the gait quantification method based on a low-resolution monocular camera provided by the present invention;

[0036] Figure 5 is a schematic structural diagram of the gait quantification device based on a low-resolution monocular camera provided by the present invention;

[0037] Figure 6 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0039] More and more evidence from clinical trials and theories reveals the association between gait characteristics and cognition. Slow walking speed and unstable gait are closely related to inattentiveness in cognition and the weakening of motor functions caused by nervous system damage. For example: in the elderly group with weakened cognitive function, the probability of falling is twice that of normal elderly people; at the same time, nervous system damage also makes patients suffer from chronic pain and muscle atrophy. These situations greatly limit the normal activities of patients and reduce the quality of life of patients. Gait metrics are clinically used to evaluate patients with nervous system damage or cognitive impairment. Gait analysis plays a crucial role in diseases such as stroke, Parkinson's syndrome, cerebral palsy, arthritis, and muscle atrophy.

[0040] However, most of the gait analysis methods in related methods rely on expensive experimental equipment and professional medical personnel, and such demanding requirements limit the wide application of gait analysis. Moreover, related methods cannot meet the dual requirements of low cost and analysis automation in gait analysis tasks.

[0041] In view of the defects in related methods, the present invention proposes a gait quantification method based on a low-resolution monocular camera. Figure 1 It is a schematic flowchart of the gait quantification method based on a low-resolution monocular camera provided by the present invention. Referring to Figure 1 , the gait quantification method based on a low-resolution monocular camera provided by the present invention may include:

[0042] Step 110, obtaining image feature points corresponding to joints of a target object in a target image, where the joints are joints related to the gait of the target object;

[0043] Step 120, determining scale feature points of the image feature points based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points;

[0044] Step 130, determining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on a feature spatio-temporal representation learning module, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0045] Step 140, fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and inputting the fused feature into a gait analysis model to obtain a gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0046] The execution subject of the gait quantification method based on a low-resolution monocular camera provided by the present invention can be an electronic device, a component in the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), or a personal computer (PC), etc. The present invention does not make specific limitations.

[0047] Taking a computer executing the gait quantification method based on a low-resolution monocular camera provided by the present invention as an example, the technical solution of the present invention will be described in detail below.

[0048] In step 110, image feature points corresponding to the joints of the target object in the target image are obtained. The joints of the target object in the target image are the joints of the target object in the target image that are related to the gait.

[0049] The target image can be an image obtained by shooting the movement of the target object using a low-resolution monocular camera during the walking process of the target object. Among them, the resolution of the low-resolution monocular camera can be 640×480.

[0050] The joints are the joints of the target object that are related to the gait. For example, the knee joint, the ankle joint, the hip joint, etc. Based on each joint point of the target object, the movement state of the target object can be analyzed.

[0051] The image feature points of the joints refer to the pixel coordinate points corresponding to each joint of the target object in the target image. The image feature points of the joints can be extracted based on the Openpose algorithm. The image feature points of the joints of the target image in the video image are extracted based on the Openpose algorithm.

[0052] In step 120, based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, the scale feature points of the image feature points are determined.

[0053] The body size parameters of the target object in the target image refer to the body size parameter information such as the height and leg length of the target object in the target image.

[0054] Based on the body size parameters of the target object, the missing scale information of the target image is restored to obtain the scale feature points of the image feature points.

[0055] Optionally, the specific implementation method for determining the scale feature points of the image feature points can be as follows: Assume the height of the target object is h, the leg length of the target object is j, the distance from the head to the eyes is 0.1h, and the hip is used as the coordinate origin. Then the coordinates of the eyes of the target object are:

[0056] (x e ,y e ) = (0, 0.9h - j);

[0057] where x e represents the x-axis coordinate in the scale feature points of the target object's eyes; y e represents the y-axis coordinate of the scale feature points of the target object's eyes. Next, the coordinates of the knees and ankles can be calculated:

[0058]

[0059] where y 1a represents the y-axis coordinate in the scale feature points of the target object's left knee; y l ′ a represents the y-axis coordinate of the target object's left knee in the image feature points; y lk represents the y-axis coordinate of the target object's left ankle in the scale feature points; y l ′ k represents the y-axis coordinate of the target object's left ankle in the image feature points; y ra represents the y-axis coordinate of the target object's right knee in the scale feature points; y r ′ a represents the y-axis coordinate of the target object's right knee in the image feature points; y rk represents the y-axis coordinate of the target object's right ankle in the scale feature points; y r ′ k represents the y-axis coordinate of the target object's right ankle in the image feature points. Next, the distance d hk between the hip and the knee and the distance d ka between the ankle and the knee need to be calculated:

[0060] d hk = max[max(y lk ), max(y rk )];

[0061] d ka = max[max(y la -ylk ), max(y ra -y rk )];

[0062] In the above formula for d hk and d ka , max() in the calculation formula means taking the larger value, where y lk , y rk , y la , y ra are all for the entire video image sequence of the target image. Then, define a symbolic operation function:

[0063]

[0064] Finally, calculate the coordinates of the knee and ankle in the x-axis direction:

[0065]

[0066]

[0067]

[0068]

[0069] where x la represents the x-axis coordinate of the left knee of the target object under the scale feature points; x l ′ a represents the x-axis coordinate of the left knee of the target object under the image feature points; x lk represents the x-axis coordinate of the left ankle of the target object under the scale feature points; x l ′ k represents the x-axis coordinate of the left ankle of the target object in the image feature points; x ra represents the x-axis coordinate of the right knee of the target object in the scale feature points; x r ′ a represents the x-axis coordinate of the right knee of the target object in the image feature points; x rk represents the x-axis coordinate of the right ankle of the target object in the scale feature points; x r ′ k represents the x-axis coordinate of the right ankle of the target object in the image feature points.

[0070] In step 130, based on the feature spatio-temporal representation learning module, determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points. Among them, the feature spatio-temporal representation learning module is determined based on the long short-term memory network and the one-dimensional convolutional network.

[0071] After respectively obtaining the image feature points of the joints and the scale feature points of the image feature points in the target image, the depth representation information of the feature points is respectively extracted using the feature spatio-temporal representation learning module. Among them, the feature spatio-temporal representation learning module is determined based on the long short-term memory network and the one-dimensional convolutional network.

[0072] Optionally, as Figure 2 As shown in the schematic diagram of the feature spatio-temporal representation learning module provided by the present invention, the feature spatio-temporal representation learning module can be composed of a long short-term memory network and three one-dimensional convolutional networks. Among them, the long short-term memory network consists of 12 long short-term memories, and the output is a state sequence. The structure of the one-dimensional convolutional network is as Figure 3 As shown in the schematic diagram of the one-dimensional convolutional network structure provided by the present invention, each one-dimensional convolutional network is composed of 3 consecutive one-dimensional convolutional layers and a max pooling layer. The convolution sum size is 8, and the output dimension is 32. Among them, after each one-dimensional convolutional layer, the relu activation function is used for activation and batch normalization operations. After the max pooling layer, a dropout operation is performed with a ratio of 0.5.

[0073] The image feature points and the scale feature points are respectively input into the feature spatio-temporal representation learning module to obtain the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points.

[0074] In step 140, after obtaining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points in step 130, the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points are fused to obtain the fused feature, and the fused feature is input into the gait analysis model to obtain the gait analysis result of the target object.

[0075] The depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points are fused to obtain the fused feature, and the fused feature is input into the gait analysis model. The gait analysis model is a supervised model for determining gait parameters.

[0076] The training of the gait analysis model can be obtained based on the gait image samples and the corresponding gait parameter labels in the gait image samples. The gait parameter labels can be parameters such as walking speed and step frequency.

[0077] Optionally, the gait analysis model can be a two-layer fully connected neural network. The fused feature is input into the two-layer fully connected neural network to obtain the output gait analysis result. Among them, the gait analysis result can be parameters such as walking speed and step frequency.

[0078] The gait quantization method based on a low-resolution monocular camera provided by an embodiment of the present invention determines image feature points corresponding to joints of a target object and scale feature points of the image feature points, respectively extracts depth spatio-temporal representation information of the image feature points and the scale feature points, fuses the extracted depth spatio-temporal representation information, and inputs the fused features into a gait analysis model, thereby realizing the gait analysis process of the target object. The gait analysis method based on images realizes the quantitative analysis of common indexes of the gait of the target object only by using a monocular camera, reduces the cost of gait analysis, and improves the application scenarios of gait analysis.

[0079] In one embodiment, fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points includes: determining the correlation weight between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on the attention mechanism algorithm; and fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on the correlation weight.

[0080] After obtaining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points, the obtained depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points are fused. Based on the attention mechanism algorithm, first determine the correlation weight between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points, and based on the determined correlation weight, fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points. The specific method for fusing according to the weight relationship can be:

[0081] f = [σ(f i ) * f s , σ(f s ) * f i ;

[0082]

[0083] f i is the depth spatio-temporal representation of the image feature points, and f s represents the depth spatio-temporal representation of the scale feature points, and f is the feature obtained after fusion.

[0084] The gait quantization method based on a low-resolution monocular camera provided by an embodiment of the present invention determines the correlation weight between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points through the attention mechanism algorithm, and realizes the determination of the correlation degree between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points.

[0085] In one embodiment, based on the feature spatio-temporal representation learning module, determining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points includes: inputting the image feature points into a long short-term memory network to obtain the state sequence of the image feature points, and inputting the state sequence of the image feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the image feature points; inputting the scale feature points into a long short-term memory network to obtain the state sequence of the scale feature points, and inputting the state sequence of the scale feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the scale feature points.

[0086] After obtaining the image feature points and the scale feature points, based on the feature spatio-temporal representation learning module, the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points are obtained. Among them, the feature spatio-temporal representation learning module is composed of a long short-term memory network and a one-dimensional convolutional network.

[0087] Specifically, first input the image feature points into a long short-term memory network to obtain the state sequence of the image feature points, and input the state sequence of the image feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the image feature points.

[0088] Input the scale feature points into a long short-term memory network to obtain the state sequence of the scale feature points, and input the state sequence of the scale feature points into a one-dimensional convolutional network to obtain the depth spatio-temporal representation of the scale feature points.

[0089] Optionally, the feature spatio-temporal representation learning module can be composed of a long short-term memory network and three one-dimensional convolutional networks. The long short-term memory network consists of 12 long short-term memories, and the output is a state sequence. Each one-dimensional convolutional network consists of 3 consecutive one-dimensional convolutional layers and a max pooling layer. The convolution sum size is 8, and the output dimension is 32. After each one-dimensional convolutional layer, the relu activation function is used for activation and batch normalization operations. After the max pooling layer, a dropout operation is performed with a ratio of 0.5.

[0090] The gait quantization method based on a low-resolution monocular camera provided by the embodiments of the present invention determines the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points through the feature spatio-temporal representation learning module, providing a basis for subsequent input into the model for gait analysis.

[0091] In one embodiment, obtaining the image feature points corresponding to the joints of the target object in the target image includes: obtaining the pixel coordinate points corresponding to the joints of the target object in the target image, and using the pixel coordinate points as the image feature points corresponding to the joints.

[0092] After obtaining the target image, analyze the target image. Extract the pixel coordinate points corresponding to the joints of the target object in the target image. After obtaining the pixel coordinate points corresponding to the joints of the target object in the target image, use the obtained pixel coordinate points as the image feature points corresponding to the joints of the target object.

[0093] The gait quantization method based on a low-resolution monocular camera provided by an embodiment of the present invention extracts the pixel coordinate points corresponding to the joints of the target object in the target image. After obtaining the pixel coordinate points corresponding to the joints of the target object in the target image, use the obtained pixel coordinate points as the image feature points corresponding to the joints of the target object, thereby realizing the determination of the image feature points in the target image.

[0094] In one embodiment, based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points, including: based on the image feature points and the body size parameters of the target object, determine the motion scale feature of the target object; based on the topological relationship between the motion scale feature and the image feature points, determine the scale feature points of the image feature points.

[0095] The body size parameters of the target object in the target image refer to the body size parameter information such as the height and leg length of the target object in the target image.

[0096] Based on the image feature points and the body size parameters of the target object, determine the motion scale feature of the target object. Among them, the motion scale feature may include: the angles of the left and right knees, the angles of the left and right ankles, the distances between the left and right ankles and the knees in the image, and the distance between the knees. These distances and angles are based on pixel values. After determining the motion scale feature of the target object, based on the determined topological relationship between the motion scale feature and the image feature points, determine the scale feature points of the image feature points.

[0097] Specifically, the specific implementation method for determining the scale feature points of the image feature points can be as follows: Assume that the height of the target object is h, the leg length of the target object is j, the distance from the head to the eyes is 0.1h, and the hip is used as the coordinate origin. Then the coordinates of the eyes of the target object are:

[0098] (x e ,y e ) = (0, 0.9h - j);

[0099] where x e represents the x-axis coordinate in the eye scale feature points of the target object; y e represents the y-axis coordinate of the eye scale feature points of the target object. Next, the coordinates of the knees and ankles can be calculated:

[0100]

[0101] where y la represents the y-axis coordinate among the scale feature points of the left knee of the target object; y l ′ a represents the y-axis coordinate of the left knee of the target object among the image feature points; y lk represents the y-axis coordinate of the left ankle of the target object among the scale feature points; y l ′ k represents the y-axis coordinate of the left ankle of the target object among the image feature points; y ra represents the y-axis coordinate of the right knee of the target object among the scale feature points; y r ′ a represents the y-axis coordinate of the right knee of the target object among the image feature points; y rk represents the y-axis coordinate of the right ankle of the target object among the scale feature points; y r ′ k represents the y-axis coordinate of the right ankle of the target object among the image feature points. Next, it is necessary to calculate the distance d between the hip and the knee hk and the distance d between the ankle and the knee ka :

[0102] d hk = max[max(y lk ), max(y rk )];

[0103] d ka = max[max(y la - y lk ), max(y ra - y rk )];

[0104] In the above calculation formulas of d hk and d ka , max() means taking the larger value, and the y lk , y rk , y la , y ra are all for the entire video image sequence of the target image. Then, define a symbolic operation function:

[0105]

[0106] Finally, calculate the coordinates of the knee and the ankle in the x-axis direction:

[0107]

[0108]

[0109]

[0110]

[0111] where x la represents the x-axis coordinate of the left knee of the target object under the scale feature points; x l ′ a represents the x-axis coordinate of the left knee of the target object under the image feature points; x lk represents the x-axis coordinate of the left ankle of the target object under the scale feature points; x l ′ k represents the x-axis coordinate of the left ankle of the target object among the image feature points; x ra represents the x-axis coordinate of the right knee of the target object among the scale feature points; x r ′ a represents the x-axis coordinate of the right knee of the target object among the image feature points; x rk represents the x-axis coordinate of the right ankle of the target object among the scale feature points; x r ′ k represents the x-axis coordinate of the right ankle of the target object among the image feature points.

[0112] The gait quantization method based on a low-resolution monocular camera provided by the embodiment of the present invention determines the motion scale feature of the target object by based on the image feature points and the body size parameters of the target object. After determining the motion scale feature of the target object, based on the topological relationship between the determined motion scale feature and the image feature points, the scale feature points of the image feature points are determined, realizing the determination of the scale feature points corresponding to the image feature points.

[0113] In one embodiment, the joints of the target object related to gait include at least one of the left hip joint, the right hip joint, the left knee joint, the right knee joint, the left ankle joint, the right ankle joint, the left toe joint, and the right toe joint.

[0114] In the target image, the joints of the target object related to gait include at least one of the left hip joint, the right hip joint, the left knee joint, the right knee joint, the left ankle joint, the right ankle joint, the left toe joint, and the right toe joint.

[0115] It can be understood that based on each joint point of the target object, the moving state of the target object can be analyzed.

[0116] The gait quantization method based on a low-resolution monocular camera provided by an embodiment of the present invention determines the joints of a target object related to gait, providing a basis for subsequent gait analysis of the joints.

[0117] The following is a schematic structural diagram of a device applying the gait quantization method based on a low-resolution monocular camera provided by the present invention Figure 4 as an example to illustrate the technical solution provided by the present invention.

[0118] The device includes: an image feature extraction module 410, a scale feature extraction module 420, a feature spatio-temporal representation learning module 430, a representation fusion module 440, and a representation decoding module 450.

[0119] The image feature extraction module 410 uses the Openpose algorithm to extract the target image in the video and determines the image feature points of the joints of the target object in the target image.

[0120] The scale feature extraction module 420 calculates a set of scale feature points containing scale information based on the height and leg length of the target object and the topological relationship between the image feature points. The relative distances of the coordinates of these scale feature points are similar to the actual body size of the target object.

[0121] The feature spatio-temporal representation learning module 430 uses a hybrid long short-term memory network and a one-dimensional convolutional network to extract the deep spatio-temporal representation information of the image feature points and the deep spatio-temporal representation information of the scale feature points.

[0122] The representation fusion module 440 uses an attention mechanism to determine the correlation weights between the deep spatio-temporal representation of the image feature points and the deep spatio-temporal representation of the scale feature points. Based on the determined correlation weights, the deep spatio-temporal representation of the image feature points and the deep spatio-temporal representation of the scale feature points are fused.

[0123] The representation decoding module 450 analyzes the fused spatio-temporal representation information using a two-layer fully connected neural network to obtain the final gait analysis result of the model.

[0124] Figure 5 is a schematic structural diagram of the gait quantization device based on a low-resolution monocular camera provided by the present invention. As Figure 5 shown, the device includes:

[0125] An acquisition module 510 is configured to acquire the image feature points corresponding to the joints of the target object in the target image, and the joints are the joints of the target object related to gait;

[0126] A topological relationship determination module 520, configured to determine scale feature points of the image feature points based on body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points;

[0127] A depth spatio-temporal representation module 530, configured to determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on a feature spatio-temporal representation learning module, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0128] An analysis module 540, configured to fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain a gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0129] The gait quantization device based on a low-resolution monocular camera provided by an embodiment of the present invention determines image feature points corresponding to joints of a target object and scale feature points of the image feature points, respectively extracts depth spatio-temporal representation information of the image feature points and the scale feature points, fuses the extracted depth spatio-temporal representation information, and inputs the fused feature into a gait analysis model, thereby implementing the gait analysis process of the target object. The gait analysis method based on images realizes the quantitative analysis of common indicators of the gait of the target object using only a monocular camera, reduces the cost of gait analysis, and expands the application scenarios of gait analysis.

[0130] In one embodiment, the analysis module 540 is specifically configured to:

[0131] Fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points includes:

[0132] Determining the correlation weight between the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on an attention mechanism algorithm;

[0133] Fusing the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on the correlation weight.

[0134] In one embodiment, the depth spatio-temporal representation module 530 is specifically configured to:

[0135] Determining the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on a feature spatio-temporal representation learning module includes:

[0136] Input the image feature points into a long short-term memory network to obtain a state sequence of the image feature points, and input the state sequence of the image feature points into a one-dimensional convolutional network to obtain a deep spatio-temporal representation of the image feature points;

[0137] Input the scale feature points into a long short-term memory network to obtain a state sequence of the scale feature points, and input the state sequence of the scale feature points into a one-dimensional convolutional network to obtain a deep spatio-temporal representation of the scale feature points

[0138] In one embodiment, the acquisition module 510 is specifically configured to:

[0139] Obtain the image feature points corresponding to the joints of the target object in the target image, including:

[0140] Obtain the pixel coordinate points corresponding to the joints of the target object in the target image, and use the pixel coordinate points as the image feature points corresponding to the joints.

[0141] In one embodiment, the topological relationship determination module 520 is specifically configured to:

[0142] Based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points, including:

[0143] Based on the image feature points and the body size parameters of the target object, determine the motion scale feature of the target object;

[0144] Based on the topological relationship between the motion scale feature and the image feature points, determine the scale feature points of the image feature points.

[0145] In one embodiment, the acquisition module 510 is further specifically configured to:

[0146] The joints of the target object related to gait include at least one of the left hip joint, right hip joint, left knee joint, right knee joint, left ankle joint, right ankle joint, left toe joint, and right toe joint.

[0147] Figure 6 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute a gait quantization method based on a low-resolution monocular camera. The method includes:

[0148] Obtain the image feature points corresponding to the joints of the target object in the target image, where the joints are the joints of the target object related to gait;

[0149] Based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points;

[0150] Based on the feature spatio-temporal representation learning module, determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0151] Fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain the gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0152] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0153] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the gait quantization method based on a low-resolution monocular camera provided by the above-mentioned various methods. The method includes:

[0154] Obtain the image feature points corresponding to the joints of the target object in the target image, where the joints are the joints of the target object related to gait;

[0155] Based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points;

[0156] Based on the feature spatio-temporal representation learning module, determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points. The feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0157] Fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain the gait analysis result of the target object. The gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0158] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the gait quantization method based on a low-resolution monocular camera provided by the above-mentioned various methods. The method includes:

[0159] Obtain the image feature points corresponding to the joints of the target object in the target image, where the joints are the joints of the target object related to gait;

[0160] Based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points, determine the scale feature points of the image feature points;

[0161] Based on the feature spatio-temporal representation learning module, determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points. The feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network;

[0162] Fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain a gait analysis result of the target object. The gait analysis model is trained based on gait image samples and their corresponding gait parameter labels.

[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A gait quantification method based on a low-resolution monocular camera, characterized in that, The method includes: Obtaining image feature points corresponding to joints of a target object in a target image, where the joints are joints of the target object related to gait; Determining scale feature points of the image feature points based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points; Determining the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points based on a feature spatio-temporal representation learning module, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network; Fusing the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points to obtain a fused feature, and inputting the fused feature into a gait analysis model to obtain a gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels; The feature spatio-temporal representation learning module is composed of a long short-term memory network and three one-dimensional convolutional networks. The long short-term memory network consists of 12 long short-term memories, and the output is a state sequence. Each one-dimensional convolutional network consists of 3 consecutive one-dimensional convolutional layers and a max pooling layer. The convolution sum size is 8, and the output dimension is 32. After each one-dimensional convolutional layer, the relu activation function is used for activation and batch normalization operations are performed. After the max pooling layer, a dropout operation is performed with a ratio of 0.

5.

2. The gait quantification method based on a low-resolution monocular camera according to claim 1, wherein The fusing the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points includes: Determining the correlation weight between the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points based on an attention mechanism algorithm; Fusing the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points based on the correlation weight.

3. The gait quantification method based on a low-resolution monocular camera according to claim 1, characterized in that The determining the depth-time-space representation of the image feature points and the depth-time-space representation of the scale feature points based on the feature spatio-temporal representation learning module includes: Inputting the image feature points into a long short-term memory network to obtain a state sequence of the image feature points, and inputting the state sequence of the image feature points into a one-dimensional convolutional network to obtain the depth-time-space representation of the image feature points; Inputting the scale feature points into a long short-term memory network to obtain a state sequence of the scale feature points, and inputting the state sequence of the scale feature points into a one-dimensional convolutional network to obtain the depth-time-space representation of the scale feature points.

4. The gait quantification method based on a low-resolution monocular camera according to claim 1, wherein The obtaining image feature points corresponding to joints of a target object in a target image includes: Obtaining pixel coordinate points corresponding to joints of a target object in a target image, and using the pixel coordinate points as the image feature points corresponding to the joints.

5. The gait quantification method based on a low-resolution monocular camera according to claim 1, characterized in that The determining scale feature points of the image feature points based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points includes: Determining the motion scale feature of the target object based on the image feature points and the body size parameters of the target object; Determine the scale feature points of the image feature points based on the topological relationship between the motion scale feature and the image feature points.

6. The gait quantization method based on a low-resolution monocular camera according to claim 1, characterized in that The joints of the target object related to gait include at least one of the left hip joint, right hip joint, left knee joint, right knee joint, left ankle joint, right ankle joint, left toe joint, and right toe joint.

7. A gait quantification device based on a low-resolution monocular camera, characterized in that, Comprising: An acquisition module, configured to acquire the image feature points corresponding to the joints of the target object in the target image, where the joints are the joints of the target object related to gait; A topological relationship determination module, configured to determine the scale feature points of the image feature points based on the body size parameters of the target object in the target image and the topological relationship between the body size parameters and the image feature points; A depth spatio-temporal representation module, configured to determine the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points based on a feature spatio-temporal representation learning module, where the feature spatio-temporal representation learning module is determined based on a long short-term memory network and a one-dimensional convolutional network; An analysis module, configured to fuse the depth spatio-temporal representation of the image feature points and the depth spatio-temporal representation of the scale feature points to obtain a fused feature, and input the fused feature into a gait analysis model to obtain the gait analysis result of the target object, where the gait analysis model is trained based on gait image samples and their corresponding gait parameter labels; The feature spatio-temporal representation learning module is composed of a long short-term memory network and three one-dimensional convolutional networks. The long short-term memory network consists of 12 long short-term memories, and the output is a state sequence. Each one-dimensional convolutional network consists of 3 consecutive one-dimensional convolutional layers and a max pooling layer. The convolution kernel size is 8, and the output dimension is 32. After each one-dimensional convolutional layer, the relu activation function is used for activation and batch normalization operations are performed. After the max pooling layer, a dropout operation is performed with a ratio of 0.

5.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the gait quantization method based on a low-resolution monocular camera according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the gait quantization method based on a low-resolution monocular camera according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the gait quantization method based on a low-resolution monocular camera according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal sensor data synthesis method and device and storage medium

    CN113610212A

  • Detection device, detection method, insole, training method and recognition method

    CN113647937A