A method and system for capturing and evaluating motion gait based on visual perception

By combining motion capture systems and visual perception technologies, and employing deep learning methods for label-free 3D gait capture and evaluation, the limitations of existing gait analysis scenarios and processes are overcome, enabling high-precision gait detection and telemedicine in home environments.

CN117011940BActive Publication Date: 2026-01-09SHANGHAI CHILDRENS HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310993837.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2026-01-09
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

Existing 3D gait analysis methods are limited by specific scenarios and professional equipment, making it impossible to perform high-precision gait detection and analysis in a home environment. Furthermore, the detection process is complex and requires the application of markers.

Method used

By combining motion capture systems and visual perception technologies, a mapping model is constructed using deep learning methods. Keypoint matching and coordinate transformation are performed using RGB video and 3D motion capture data to achieve label-free 3D gait capture and evaluation.

Benefits of technology

It enables high-precision gait analysis in a home environment, simplifies the detection process, improves the universality and convenience of gait analysis, and supports telemedicine applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011940B_ABST
    Figure CN117011940B_ABST
Patent Text Reader

Abstract

The application discloses a motion gait capture and evaluation method and system based on visual perception, and relates to the field of visual perception, which comprises the following steps: collecting three-dimensional motion capture data of anatomical key points under world coordinates and corresponding RGB videos under different visual angles when a target is moving, estimating two-dimensional lower limb key points of the target in the image after video segmentation, constructing a deep learning model by minimizing reconstruction error and combining human anatomy constraints to realize mapping from two-dimensional key points to three-dimensional key points through paired two-dimensional key points and three-dimensional key points, obtaining three-dimensional human lower limb key point position estimation under a camera coordinate system, and forming a set of three-dimensional motion trajectories estimated from video information by the motion of the key points in the time dimension; extracting gait information of the target according to the three-dimensional motion trajectories, analyzing and extracting gait characteristic quantities, and comprehensively evaluating the gait characteristic quantities to obtain an evaluation result. The method simplifies the gait detection process and improves the universality of the gait analysis environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of visual perception, and in particular to a motion gait capture and evaluation method and system based on visual perception. BACKGROUND

[0002] The three-dimensional gait analysis currently used in the field of clinical medicine is based on a multi-sensor system or a multi-camera optical capture system to complete corresponding data acquisition. The principle is to capture the body surface marker points pasted on the torso and lower limbs of the detected person through multiple infrared cameras at the same time, and then obtain high-precision three-dimensional motion trajectories of the marker points, and calculate the joint angles and corresponding gait parameters accordingly. Although the gait analysis data obtained by this method has high accuracy, there are the following limitations, but not limited to:

[0003] 1. Scene is limited: the examinee needs to be detected in a specific gait analysis room, not in a general home indoor environment;

[0004] 2. Detection process is limited: the examinee needs to expose the skin of the upper and lower body before gait analysis examination, and the professional gait analysis examiner pastes markers on the body surface at all places, so that the motion trajectory of the marker can be captured by the camera in the subsequent examination process;

[0005] 3. The motion trajectory of the marker needs to be captured by an expensive and professional multi-infrared camera motion capture system.

[0006] The above limitations make the gait analysis examination only available in a specific scene by a professional examiner.

[0007] The existing research mainly aims to extract the biological features contained in the target contour feature, rather than to collect kinematic parameters. At the same time, since the video in the home environment is usually taken from a non-fixed perspective, the estimated three-dimensional human kinematic information is represented in the moving camera coordinate system, which makes it impossible to accurately calculate the key gait parameters, so the existing video analysis method cannot meet the accuracy of the estimation of three-dimensional kinematic parameters required by clinical medicine. SUMMARY

[0008] In view of the defects in the prior art, the motion gait capture and evaluation method and system based on visual perception provided by the embodiments of the present application simplify the detection process and improve the universality of the gait analysis environment, making gait examination and rehabilitation evaluation simple and easy to implement.

[0009] In a first aspect, the motion gait capture and evaluation method based on visual perception provided by the embodiments of the present application comprises: collecting three-dimensional motion capture data of anatomical key points in the world coordinate when the target is moving through a motion capture system;

[0010] acquire motion RGB video of the corresponding target from another perspective by using a camera, and estimate the pixel coordinates of the two-dimensional lower limb key points of the target in the image sequence in the RGB image sequence by using a pose estimation method;

[0011] match and correspond the two-dimensional lower limb key points and the three-dimensional motion capture data in the time sequence;

[0012] construct a mapping model based on deep learning, and learn the mapping function of the two-dimensional lower limb key points and the three-dimensional position estimation in the camera coordinate system;

[0013] take the motion capture system coordinate system as the world coordinate system, obtain a coordinate system conversion matrix in combination with the camera parameters and the 6D pose of the camera in the world coordinate system, calculate the three-dimensional position estimation of the two-dimensional lower limb key points in the world coordinate system according to the three-dimensional position estimation in the camera coordinate system and the coordinate system conversion matrix, perform minimum error reconstruction on the three-dimensional lower limb key points obtained by the motion capture system and the three-dimensional position estimation of the two-dimensional lower limb key points, obtain the three-dimensional human lower limb key point position estimation in the camera coordinate system, and form a set of three-dimensional motion trajectories estimated from video information in the time dimension by the motion of the key points;

[0014] extract gait information of the target according to the three-dimensional motion trajectories, analyze the gait information, extract gait feature quantities, comprehensively evaluate the gait feature quantities by using a gait feature analysis method, and obtain a gait evaluation result.

[0015] In a second aspect, a motion gait capture and evaluation system based on visual perception is provided, which includes a motion capture data acquisition module, a video data acquisition module, a matching module, a data processing module, and a gait analysis module.

[0016] The motion capture data acquisition module is configured to acquire three-dimensional motion capture data of anatomical key points in the world coordinate system when a target moves by using a motion capture system.

[0017] The video data acquisition module is configured to acquire motion RGB video of the corresponding target from another perspective by using a camera, and estimate the pixel coordinates of the two-dimensional lower limb key points of the target in the image sequence in the RGB image sequence by using a pose estimation method;

[0018] The matching module is configured to match and correspond the two-dimensional lower limb key points and the three-dimensional motion capture data in the time sequence.

[0019] The data processing module is configured to construct a mapping model based on deep learning, and learn the mapping function of the two-dimensional lower limb key points and the three-dimensional position estimation in the camera coordinate system.

[0020] The motion capture system coordinate system is taken as a world coordinate system, a coordinate system conversion matrix is obtained by combining camera parameters and a 6D posture of the camera in the world coordinate system, a three-dimensional position estimation of a two-dimensional lower limb key point in the world coordinate system is calculated according to the three-dimensional position estimation of the two-dimensional lower limb key point and the coordinate system conversion matrix, a three-dimensional lower limb key point position estimation in the camera coordinate system is obtained by minimizing error reconstruction of the three-dimensional lower limb key point obtained by the motion capture system and the three-dimensional position estimation of the two-dimensional lower limb key point, and the motion of the key point forms a set of three-dimensional motion trajectories estimated by video information in the time dimension.

[0021] The gait analysis module is used for extracting gait information of the target according to the three-dimensional motion trajectories, analyzing the gait information, extracting gait characteristic quantities, comprehensively evaluating the gait characteristic quantities by using a gait feature analysis method, and obtaining a gait evaluation result.

[0022] The present application has the following beneficial effects:

[0023] The motion gait capture and evaluation method and system based on visual perception provided by the embodiment of the present application simplify the existing clinical gait analysis method, make the gait analysis not limited to the laboratory level, and can realize the universal and simple gait detection in the home environment, provide a new possibility for the universal and convenient gait detection, make the gait examination and rehabilitation evaluation simple and easy, and finally achieve the purpose of remote medical treatment. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the specific embodiments or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual proportion.

[0025] Figure 1 A flow chart of the motion gait capture and evaluation method based on visual perception provided by the first embodiment of the present application is shown;

[0026] Figure 2 A structural block diagram of the motion gait capture and evaluation system based on visual perception provided by the first embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0028] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0029] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0030] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0031] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0032] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0033] Example 1

[0034] like Figure 1 The diagram shows a flowchart of a visual perception-based gait capture and evaluation method provided in the first embodiment of the present invention. The method includes the following steps:

[0035] Three-dimensional motion capture data of anatomical key points in world coordinates is collected by a motion capture system during target movement.

[0036] From another perspective, a camera is used to capture the corresponding motion RGB video of the target, and the pose estimation method is used to estimate the pixel coordinates of the two-dimensional lower limb key points of the target in the RGB image sequence.

[0037] Match the two-dimensional lower limb key points with the three-dimensional motion capture data in the time series;

[0038] A mapping model based on deep learning is constructed to learn a mapping function for three-dimensional position estimation of the two-dimensional lower limb key points and the camera coordinate system;

[0039] The motion capture system coordinate system is taken as the world coordinate system, a coordinate system conversion matrix is obtained by combining the camera parameters and the 6D pose of the camera in the world coordinate system, the three-dimensional position estimation of the two-dimensional lower limb key points in the world coordinate system is calculated according to the three-dimensional position estimation in the camera coordinate system and the coordinate system conversion matrix, the three-dimensional lower limb key points obtained by the motion capture system are reconstructed by minimizing the error of the two-dimensional lower limb key points, and the three-dimensional human lower limb key point position estimation in the camera coordinate system is obtained, and the motion of the key points forms a set of three-dimensional motion trajectories estimated by video information in the time dimension.

[0040] According to the three-dimensional motion trajectory, the gait information of the target is extracted, the gait information is analyzed, the gait feature quantity is extracted, the gait feature analysis method is used to comprehensively evaluate the gait feature quantity, and the gait evaluation result is obtained.

[0041] Specifically, by the motion capture system of clinical gait analysis, high-precision three-dimensional motion capture data of the anatomical key points of the target motion in the world coordinate system {E} is directly collected {E} S m (t m ). At the same time, the corresponding target motion RGB video is collected under another view, and the existing pose estimation method is used to accurately estimate the two-dimensional lower limb key points p j =[x j ,y j ] T in the image, wherein j represents the key point number, and j is an integer. In order to match the two-dimensional and three-dimensional human key point information, when collecting the RGB video, a marker fixed camera or a mobile camera can be used for shooting, and a motion (such as holding hands over the head and then putting down, kicking legs, etc.) can be used to mark the video time node, so as to unify the time axis starting point t0 of the motion capture system and the RGB video collection data. After the time axis starting point is unified, the motion capture system is taken as the time axis reference, the time resolution of the RGB video data and the motion capture system is unified, and the two-dimensional lower limb key points p j (t p ) and the three-dimensional motion capture data {E} S m (t m ) are matched in time sequence. If the data collection frequency of the motion capture system is inconsistent with the video collection system, the following processing is performed: for a certain collection time point data {E} S m (t m,k), where m and k are both integers, find the closest RGB video data sampling time point t. p,n Where p and n are both integers. Using the two-dimensional human lower limb keypoint coordinate sequence [...] extracted from RGB video [...] j (t p,n-1 ), p j (t p,n ), p j (t p,n+1 ), ...] Calculate the pixel coordinates p of the two-dimensional human lower limb key points in the video at the corresponding time point. j (t p,k Interpolation methods (such as bilinear interpolation and trilinear interpolation) are used to supplement the data, thereby achieving temporal matching between two-dimensional and three-dimensional data.

[0042] By combining camera imaging models with human anatomical information, a deep learning-based mapping model is constructed to learn the aforementioned two-dimensional human lower limb key points p. j =[x j y j ] T Perform 3D position estimation in camera coordinate system {C} {c} P c,j =[X j Y j Z j ] T The mapping function f(p) j ) = P c,j .

[0043] Using the motion capture system's coordinate system as the world coordinate system, if a fixed camera is used, the motion capture system is used to acquire the camera's 6D pose in the world coordinate system {E}; if a handheld, mobile camera is used, visual Simultaneous Localization and Mapping (vSLAM) is used to construct an environmental map and estimate the camera's 6D pose in the world coordinate system. Combining the camera parameters and its 6D pose in the world coordinate system, the coordinate transformation matrix can be obtained. {E} T {C} Three-dimensional position estimation in the world coordinate system can be achieved by... {E} P c,j = {E} T {C} {C} P c,j The calculations were performed using the 3D lower limb keypoints obtained from the motion capture system. {E} P m,j Compared with the estimation from two-dimensional lower limb key points {E} P c,j Minimize the error reconstruction e = ||{E} P m,j - {E} p c,j The aforementioned deep learning model is optimized to achieve 3D human lower limb keypoint position estimation under a monocular camera. The algorithm's estimation accuracy can be measured using the Mean Per Joint Position Error (MPJPE), which is the average of the prediction errors of all keypoints. The motion of these keypoints constitutes a set of 3D motion trajectories estimated from video information over time. {E} S c (t), using this trajectory, the spatiotemporal parameters and kinematic parameters (such as stride frequency, stride length, etc.) of the target can be accurately extracted for analysis.

[0044] Visual Localization and Navigation (vSLAM) uses images as input to construct a map of the camera's environment and enables camera relocalization within that environment. vSLAM allows for real-time estimation of the pose of a moving camera in the environment, enabling the transformation of the aforementioned 3D human lower limb keypoint positions from the camera coordinate system {C} to the world coordinate system {E}. Using the previously obtained 2D human keypoint information, a human mask is generated to eliminate the influence of moving human bodies in the RGB video, helping the SLAM algorithm extract key feature points of objects in a static environment. This effectively improves the accuracy of camera pose estimation when moving human bodies are present in the field of view, determines the actual movement distance of the detected target's 3D lower limb keypoints, and thus achieves gait analysis in a regularized coordinate system.

[0045] In practical applications, fixed cameras (such as surveillance cameras) or mobile cameras (such as mobile phone cameras) can be used to collect target gait data. If a fixed camera is used, the aforementioned world coordinate system {E} and camera coordinate system {C} are the same. The aforementioned algorithm combined with the camera parameter model can be used to estimate the position of key points of the lower limbs in 3D and track their trajectory. If a mobile camera is used, the aforementioned vSLAM algorithm should be used to calculate the 6D pose of the camera relative to the environment, thereby transforming the estimated position of the key points of the lower limbs in 3D to the world coordinate system and extracting the motion trajectory of the key points without the need for a motion capture system.

[0046] Using the above method, using fixed or mobile cameras, a large number of diversified (different gender, age, height, weight, etc.) normal gait data are collected, a set of normal gait parameters is determined, and a normal gait database is established. The gait parameters (such as step frequency, step length, step angle, etc.) of the detection target are analyzed, and the gait characteristic quantity is extracted. Then, using mainstream gait feature analysis methods, such as using gait deviation index (GDI, Gait Deviation Index), Gillette gait index (Gillette Gait Index, GGI), etc., the gait is comprehensively evaluated. Taking GDI as an example, GDI extracts 15 gait characteristic quantities according to the gait parameters and takes them as the normal gait standard. The absolute Euclidean distance between the gait of the subject and the normal gait is calculated through logarithmic conversion and Z-score transformation. Using the above analysis method, the deviation degree of the target gait super parameter from the normal gait parameter set is scored, so as to realize the screening, diagnosis and staging of the target disease and the prognosis recovery diagnosis in the home environment.

[0047] The motion gait capture and evaluation method based on visual perception provided by the embodiment of the application realizes the following:

[0048] The motion gait capture and evaluation method based on visual perception provided by the embodiment of the application realizes the following:

[0049] The universality of the gait analysis environment is improved, so that the examinee can complete the motion gait analysis and evaluation in the home environment, the gait examination and rehabilitation evaluation become simple and easy, and the purpose of remote medical treatment is finally achieved;

[0050] The detection process is simplified, that is, the marker-free human posture estimation is realized through visual perception, that is, the markers do not need to be pasted on the body surface.

[0051] In the first embodiment described above, a visual perception-based motion gait capturing and evaluation method is provided, and a visual perception-based motion gait capturing and evaluation system is also provided. Please refer to Figure 2 which is a structural block diagram of a visual perception-based motion gait capturing and evaluation system provided by the second embodiment of the present application. Since the device embodiment is basically similar to the method embodiment, it is described more simply, and the relevant part can be referred to the part of the description of the method embodiment. The device embodiment described below is only illustrative.

[0052] Embodiment 2

[0053] As shown in Figure 2 , a structural block diagram of a visual perception-based motion gait capturing and evaluation system provided by another embodiment of the present application is shown. The visual perception-based motion gait capturing and evaluation system provided by the embodiment includes a motion capture data acquisition module, a video data acquisition module, a matching module, a data processing module, and a gait analysis module. The motion capture data acquisition module is used to acquire three-dimensional motion capture data of anatomical key points in a world coordinate system through a motion capture system when a target moves. The video data acquisition module is used to acquire a motion RGB video of the corresponding target by using a camera from another perspective and estimate the pixel coordinates of two-dimensional lower limb key points of the target in the image by using a pose estimation method. The matching module is used to match and correspond the two-dimensional lower limb key points and the three-dimensional motion capture data in the time sequence. The data processing module is used to construct a mapping model based on deep learning, learn the mapping function of the two-dimensional lower limb key points and the three-dimensional position estimation in the camera coordinate system, obtain a coordinate system conversion matrix by taking the motion capture system coordinate system as the world coordinate system, combining the camera parameters and the 6D pose of the camera in the world coordinate system, calculate the three-dimensional position estimation of the two-dimensional lower limb key points in the world coordinate system according to the three-dimensional position estimation in the camera coordinate system and the coordinate system conversion matrix, and perform minimum error reconstruction on the three-dimensional lower limb key points obtained by the motion capture system and the three-dimensional position estimation of the two-dimensional lower limb key points to obtain the three-dimensional human lower limb key point position estimation in the camera coordinate system. The motion of the key points forms a set of three-dimensional motion trajectories estimated by the video information in the time dimension. The gait analysis module is used to extract gait information of the target according to the three-dimensional motion trajectories, analyze the gait information, extract gait feature quantities, comprehensively evaluate the gait feature quantities by using a gait feature analysis method, and obtain a gait evaluation result. The gait feature analysis method is a gait deviation index.

[0054] The matching module includes a first matching unit. The first matching unit uses an agreed motion to mark the video time nodes, so as to unify the time axis starting points of the motion capture system and the RGB video acquisition data, take the time of the motion capture system as the time axis reference, and unify the time resolution of the RGB video and the motion capture system.

[0055] The matching module further comprises a second matching unit, which is configured to, when the motion capture system and the RGB video data acquisition frequency are inconsistent, find the closest RGB video data sampling time point for one acquisition time point data in the motion capture system, and calculate the two-dimensional human lower limb key point pixel coordinates in the RGB video at the corresponding time point using the two-dimensional human lower limb key point coordinate sequence extracted from the RGB video by interpolation.

[0056] The data processing module comprises a mobile camera processing unit, which uses visual positioning and navigation technology to construct an environment map where the camera is located, estimates the pose of the camera in the environment in real time, and estimates the 6D pose of the camera in the world coordinate system.

[0057] The gait capture and evaluation system based on visual perception simplifies the existing clinical gait analysis method, makes the gait analysis not limited to the laboratory level, and can realize the ubiquitous and simple gait detection in the home environment, provides a new possibility for the ubiquitous and convenient gait detection, makes the gait examination and rehabilitation evaluation simple and easy, and finally achieves the purpose of remote medical treatment.

[0058] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and they should be covered in the scope of the claims and the specification of the present application.

Claims

1. A method for capturing and evaluating the gait of a motion based on visual perception, characterized in that, The method comprises the following steps: acquiring three-dimensional motion capture data of anatomical key points in a world coordinate system through a motion capture system when a target is moving; acquiring a moving RGB video of the corresponding target from another perspective through a camera, and estimating pixel coordinates of two-dimensional lower limb key points of the target in the RGB image sequence in the image through a pose estimation method; matching and corresponding the two-dimensional lower limb key points and the three-dimensional motion capture data in a time sequence; constructing a mapping model based on deep learning, and learning a mapping function of the two-dimensional lower limb key points and three-dimensional position estimation in a camera coordinate system; taking the motion capture system coordinate system as the world coordinate system, combining camera parameters and a 6D pose of the camera in the world coordinate system to obtain a coordinate system conversion matrix, calculating three-dimensional position estimation of the two-dimensional lower limb key points in the world coordinate system according to the three-dimensional position estimation in the camera coordinate system and the coordinate system conversion matrix, and performing minimum error reconstruction on the three-dimensional lower limb key points obtained by the motion capture system and the three-dimensional position estimation of the two-dimensional lower limb key points to obtain three-dimensional position estimation of the lower limb key points of the human body in the camera coordinate system, and the motion of the key points forms a set of three-dimensional motion trajectories estimated from video information in the time dimension; extracting gait information of the target according to the three-dimensional motion trajectories, analyzing the gait information, extracting gait feature quantities, comprehensively evaluating the gait feature quantities through a gait feature analysis method, and obtaining a gait evaluation result; the specific method of matching and corresponding the two-dimensional lower limb key points and the three-dimensional motion capture data in a time sequence comprises: marking video time nodes by using an agreed action, and taking the time axis of the motion capture system and the RGB video acquisition data as a starting point; taking the time of the motion capture system as a time axis reference, and unifying the time resolution of the RGB video and the motion capture system; after the step of unifying the time resolution of the RGB video and the motion capture system, the method further comprises the following steps: when the acquisition frequency of the motion capture system and the RGB video data is inconsistent, finding the closest RGB video data sampling time point for one acquisition time point data in the motion capture system; using interpolation to extract the two-dimensional human lower limb key point coordinate sequence from the RGB video, and calculating the two-dimensional human lower limb key point pixel coordinates in the RGB video at the corresponding time point.

2. The method of claim 1, wherein, The camera is a mobile camera, and the method for obtaining the 6D pose of the camera in the world coordinate system comprises the following steps:

3. The method of claim 1, wherein, the gait feature analysis method is a gait deviation index.

4. A visual perception based motion gait capture and evaluation system, characterized by, The method comprises the following steps: a motion capture data acquisition module, a video data acquisition module, a matching module, a data processing module and a gait analysis module; the motion capture data acquisition module is used to acquire three-dimensional motion capture data of anatomical key points in a world coordinate system through a motion capture system when a target is moving; the video data acquisition module is used to acquire a moving RGB video of the corresponding target from another perspective through a camera, and estimate pixel coordinates of two-dimensional lower limb key points of the target in the RGB image sequence in the image through a pose estimation method; The matching module is configured to match and correspond the two-dimensional lower limb key points and the three-dimensional motion capture data in a time sequence. The data processing module is configured to construct a mapping model based on deep learning, and learn a mapping function of the two-dimensional lower limb key points and three-dimensional position estimation in a camera coordinate system. The action capture system coordinate system is taken as a world coordinate system, a coordinate system conversion matrix is obtained by combining camera parameters and a 6D pose of the camera in the world coordinate system, three-dimensional position estimation of the two-dimensional lower limb key points in the world coordinate system is calculated according to the three-dimensional position estimation in the camera coordinate system and the coordinate system conversion matrix, three-dimensional lower limb key points obtained by the action capture system are reconstructed by minimizing error with the three-dimensional position estimation of the two-dimensional lower limb key points, three-dimensional human lower limb key point position estimation in the camera coordinate system is obtained, and motion of the key points forms a set of three-dimensional motion trajectories estimated by video information in a time dimension. The gait analysis module is configured to extract gait information of the target according to the three-dimensional motion trajectories, analyze the gait information, extract gait characteristic quantities, comprehensively evaluate the gait characteristic quantities by using a gait feature analysis method, and obtain a gait evaluation result. The matching module includes a first matching unit, which uses an agreed action to mark a video time node, takes a time axis of the motion capture system and the RGB video acquisition data as a starting point, and takes time of the motion capture system as a time axis reference to unify time resolution of the RGB video and the motion capture system. The matching module further includes a second matching unit, which is configured to, when acquisition frequencies of the motion capture system and the RGB video data are inconsistent, find a closest RGB video data sampling time point for one acquisition time point data in the motion capture system, use interpolation to extract a two-dimensional human lower limb key point coordinate sequence from the RGB video, and calculate two-dimensional human lower limb key point pixel coordinates in the RGB video at a corresponding time point.

5. The system of claim 4, wherein, The data processing module includes a mobile camera processing unit, which uses visual positioning and navigation technology to construct an environment map where the camera is located, estimates a pose of the camera in the environment in real time, and estimates a 6D pose of the camera in the world coordinate system.

6. The system of claim 4, wherein, The gait feature analysis method is a gait deviation index.

Citation Information

Patent Citations

  • Remote dynamic gait image capturing and evaluating method and system based on machine learning

    CN115731233A

  • Monocular video-based multi-stage human motion capture method and device, and medium

    CN116386141A