Monocular RGB image-based real-time three-dimensional human body sensing system and method
Through a real-time three-dimensional human body perception system based on monocular RGB images, combined with deep learning algorithms, the shortcomings in real-time and accuracy of the existing technology are solved, and efficient three-dimensional human body reconstruction and pose estimation under complex conditions are achieved.
Patent Information
- Application Number
- CN202510118529.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing three-dimensional human perception technology has shortcomings in real-time and accuracy, and is difficult to effectively solve in complex backgrounds and occlusions.
A real-time three-dimensional human body perception system based on monocular RGB images is adopted, and real-time three-dimensional reconstruction and pose estimation of the human body is realized through modules such as data acquisition and preprocessing, feature extraction and matching, three-dimensional reconstruction, size measurement and pose estimation, combined with deep learning algorithms.
It realizes efficient three-dimensional human body reconstruction and pose estimation under monocular RGB image conditions, reduces hardware costs, improves processing accuracy and real-time performance, and is suitable for complex backgrounds and occlusion situations.
Smart Images

Figure CN120047970A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional human perception, and more particularly to a real-time three-dimensional human perception system and method based on a monocular RGB image. Background Art
[0002] With the continuous development of computer vision technology, three-dimensional human perception technology has gradually become a research hotspot. In the fields of virtual reality (VR), augmented reality (AR), human-computer interaction, etc., real-time and accurate three-dimensional human perception technology has broad application prospects.
[0003] However, traditional three-dimensional human perception methods often rely on multi-camera systems or depth cameras, and these systems have certain limitations in deployment and use. For example, traditional three-dimensional human reconstruction methods usually rely on multi-vision systems (such as binocular stereo vision or depth cameras). Although these systems can provide relatively accurate depth information, they have problems such as complex equipment, high cost, and inconvenient installation, which limit their wide use in certain application scenarios. Another common method is to fit based on a predefined human model (such as the SMPL model). These methods usually require a large amount of training data and complex optimization processes, and have poor real-time performance in dynamic scenarios. In recent years, deep learning-based methods have made remarkable progress in three-dimensional human reconstruction. These methods usually directly predict the three-dimensional pose and shape of the human body from a monocular image through a convolutional neural network (CNN). It can be seen that the existing methods still need to be improved in terms of real-time performance and accuracy, especially in complex backgrounds and occlusion situations.
[0004] Therefore, how to provide a real-time three-dimensional human perception system and method based on a monocular RGB image that can solve the above problems is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a real-time three-dimensional human perception system and method based on a monocular RGB image to solve the technical problems existing in the above-mentioned prior art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A real-time three-dimensional human perception system based on a monocular RGB image, comprising:
[0008] A data acquisition and preprocessing module, which acquires a human body image and preprocesses the human body image;
[0009] A feature extraction and matching module, which extracts human feature points in the preprocessed human body image, matches the feature points in the current frame with those in the key frame, and obtains the motion information of the human body;
[0010] A 3D reconstruction module that performs 3D reconstruction based on the motion information to obtain a 3D reconstruction result;
[0011] A dimension measurement module that measures and evaluates the dimensions of key parts of the human body using 3D geometric measurement methods to obtain measurement and evaluation results;
[0012] A pose estimation module that estimates the human pose using a deep learning algorithm to obtain a human pose estimation result;
[0013] A result display module that displays the 3D reconstruction result, the measurement and evaluation result, and the human pose estimation result in a visual form.
[0014] Preferably, the feature extraction and matching module includes:
[0015] A human feature extraction unit that extracts the pre-processed human body image to obtain a set of key points of the human body;
[0016] A feature point direction determination unit that determines the direction information between key points for each key point according to its position in the human body image and its relationship with adjacent feature points;
[0017] A chain-like feature construction unit that constructs the chain-like features of the human body using the extracted key points and the direction information between the key points;
[0018] A feature vector construction unit that constructs a feature vector for each chain-like feature;
[0019] A label vector construction unit that assigns a unique label vector to each chain-like feature;
[0020] An iteration unit that repeats the feature extraction and matching process for the key frames in the pre-processed human body image, extracts and stores the chain-like features and their feature vectors and label vectors, generates the motion information of the human body and outputs it.
[0021] Preferably, the key points of the human body include: shoulders, elbows, hands, wrists, knees.
[0022] Preferably, the feature vector includes the length, direction, and degree of curvature of the chain-like feature.
[0023] Preferably, the 3D reconstruction module includes:
[0024] A monocular RGB camera that captures the human body image to be measured;
[0025] A camera calibration unit that determines the internal parameters and external parameters of the monocular RGB camera and establishes the relationship between the pixel positions in the camera image and the 3D point positions of the human body;
[0026] The 3D point coordinate recovery unit for feature points recovers the 3D point coordinates of feature points in 3D space by using the motion information of the human body and the camera calibration result generated by the feature extraction and matching module;
[0027] The 3D model construction unit constructs a 3D model of the human body according to the 3D point coordinates, including: a bone structure and a muscle structure;
[0028] The 3D pose estimation unit estimates the 3D pose of the human body based on the constructed 3D model, including: joint angle information and body orientation information;
[0029] The 3D motion trajectory tracking unit combines the matching results of human feature points in consecutive frames to track the motion trajectory of the human body in 3D space and outputs a dynamic 3D reconstruction result.
[0030] Preferably, the calculation formula of the camera calibration unit is:
[0031]
[0032] In the formula, (x 1 , y 1 ), (x 2 , y 2 ) are a pair of matching feature points, s and's' are scale factors, X, Y, Z are the point coordinates in the world coordinate system, K is the internal parameter matrix of the camera, t is the translation vector, R is the rotation matrix, R T R = I, R T is the transpose of R, and I represents the identity matrix.
[0033] Preferably, the calculation formula of the 3D point coordinate recovery unit for feature points is:
[0034]
[0035]
[0036] In the formula, P 1 , P 2 are a pair of matching feature points, P 1 = (x 1 , y 1 ), P 2 = (x 2 , y 2 ), a 1 , a 2 are the direction vectors of two feature points, λ 1 , λ 2 are the proportionality factors of the distances from two feature points to the camera center respectively, C 1 , C 2 are the center positions of the camera under different perspectives.
[0037] Preferably, the 3D reconstruction module further includes a structure optimization unit for optimizing the initial 3D point coordinates, and the calculation formula is:
[0038]
[0039] In the formula, N is the number of 3D point coordinates, V i is the view angle set of the i-th 3D point coordinate, p ij is the projection position of the i-th 3D point coordinate under the j-th view angle, ρ is the robust loss function, π is the projection function, K is the internal parameter matrix of the camera, X i represents the coordinate of the i-th 3D point, R j represents the rotation matrix of the j-th view angle, and t j represents the translation vector of the j-th view angle.
[0040] Preferably, the 3D model construction unit further includes: smoothing the 3D model.
[0041] On the other hand, the present invention also provides a real-time 3D human body perception method based on a monocular RGB image, including:
[0042] S100: Collect a human body image and preprocess the human body image;
[0043] S200: Extract the human body feature points in the preprocessed human body image, and match the feature points in the current frame with those in the key frame to obtain the motion information of the human body;
[0044] S300: Perform 3D reconstruction based on the motion information to obtain a 3D reconstruction result;
[0045] S400: Use a 3D geometric measurement method to measure and evaluate the dimensions of the key parts of the human body to obtain a measurement and evaluation result;
[0046] S500: Estimate the human body posture using a deep learning algorithm to obtain a human body posture estimation result;
[0047] S600: Display the 3D reconstruction result, the measurement and evaluation result, and the human body posture estimation result in a visual form.
[0048] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a real-time 3D human body perception system based on a monocular RGB image, which uses a monocular RGB image for 3D reconstruction, dimension measurement, and posture estimation, reducing the hardware cost. At the same time, combining a deep learning algorithm improves the processing accuracy and real-time performance. At the same time, during the 3D reconstruction process, the 3D model is optimized, improving the reconstruction accuracy and stability. Description of the Drawings
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained according to the provided accompanying drawings.
[0050] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0051] Figure 2 It is a schematic diagram of the method flow of the present invention. Detailed implementation manners
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] See Figure 1 , the embodiments of the present invention disclose a real-time three-dimensional human perception system based on a monocular RGB image, including:
[0054] A data acquisition and preprocessing module, which acquires a human body image and preprocesses the human body image;
[0055] A feature extraction and matching module, which extracts human feature points in the preprocessed human body image, matches the feature points in the current frame with those in the key frame, and obtains the motion information of the human body;
[0056] A three-dimensional reconstruction module, which performs three-dimensional reconstruction based on the motion information to obtain a three-dimensional reconstruction result;
[0057] A size measurement module, which measures and evaluates the sizes of key parts of the human body using three-dimensional geometric measurement methods to obtain measurement and evaluation results;
[0058] A pose estimation module, which estimates the human body pose using a deep learning algorithm to obtain a human body pose estimation result;
[0059] A result display module, which displays the three-dimensional reconstruction result, the measurement and evaluation result, and the human body pose estimation result in a visual form.
[0060] In a specific embodiment, in the data acquisition and preprocessing module, the human body images (including key frames and current frames) are preprocessed, including steps such as denoising, grayscale conversion, and binarization, to improve the accuracy of subsequent feature extraction.
[0061] In a specific embodiment, the feature extraction and matching module includes:
[0062] A human body feature extraction unit that extracts the preprocessed human body image to obtain a set of key points of the human body;
[0063] A feature point direction determination unit that, for each key point, determines the direction information between key points (for example, the vector from the shoulder to the elbow) according to the position in the human body image and the relationship with adjacent feature points, which helps the subsequent feature matching process;
[0064] A chain-like feature construction unit that constructs the chain-like features of the human body using the extracted key points and the direction information between the key points. Specifically, a series of line segments (such as the line segments of the arm and leg) can be formed by connecting adjacent key points, and the direction and length of these line segments are considered.
[0065] A feature vector construction unit that constructs a feature vector for each chain-like feature;
[0066] A label vector construction unit that assigns a unique label vector to each chain-like feature for the subsequent matching process;
[0067] An iteration unit that repeats the feature extraction and matching process for the key frames in the preprocessed human body image, extracts and stores the chain-like features and their feature vectors and label vectors, generates the motion information of the human body and outputs it.
[0068] Specifically, the human body feature extraction unit uses a human body key point detection algorithm (such as OpenPose, DeepLabCut, etc.) to process the preprocessed image to obtain a set of key points of the human body.
[0069] Specifically, the key points of the human body include: shoulders, elbows, hands, wrists, knees.
[0070] In a specific embodiment, the feature vector includes the length, direction, and bending degree of the chain-like feature.
[0071] In a specific embodiment, the 3D reconstruction module includes:
[0072] A monocular RGB camera that captures the human body image to be measured;
[0073] A camera calibration unit that determines the internal parameters and external parameters of the monocular RGB camera and establishes the relationship between the pixel positions of the camera image and the 3D point positions of the human body;
[0074] The 3D point coordinate recovery unit of feature points recovers the 3D point coordinates of feature points in the 3D space by using the motion information of the human body and the camera calibration result generated by the feature extraction and matching module;
[0075] The 3D model construction unit constructs a 3D model of the human body according to the 3D point coordinates, including: a bone structure and a muscle structure;
[0076] The 3D pose estimation unit estimates the 3D pose of the human body based on the constructed 3D model, including: joint angle information and body orientation information;
[0077] The 3D motion trajectory tracking unit combines the matching results of human body feature points in consecutive frames to track the motion trajectory of the human body in the 3D space and outputs a dynamic 3D reconstruction result.
[0078] Specifically, through the above technical solution, human body feature points can be extracted from multi-view images, the camera motion parameters can be estimated, triangulation can be performed to obtain an initial 3D point cloud, then the reprojection error can be reduced through optimization, and finally surface reconstruction can be performed to obtain a 3D model of the human body.
[0079] In a specific embodiment, the calculation formula of the camera calibration unit is:
[0080]
[0081]
[0082] In the formula, (x 1 , y 1 ), (x 2 , y 2 ) are a pair of matching feature points, s and's' are scale factors, X, Y, Z are the point coordinates in the world coordinate system, K is the internal parameter matrix of the camera, t is the translation vector, R is the rotation matrix, R T R = I, R T is the transpose of R, and I represents the identity matrix.
[0083] Specifically, the present invention can start from the feature points matched from two perspectives and calculate the corresponding 3D coordinates of the feature points.
[0084] In a specific embodiment, the calculation formula of the 3D point coordinate recovery unit of feature points is:
[0085]
[0086] In the formula, P 1 , P 2 are a pair of matching feature points, P 1 = (x 1, y 1 ), P 2 = (x 2 , y 2 ), a 1 , a 2 are the direction vectors of two feature points, λ 1 , λ 2 are the scale factors of the distances from the two feature points to the camera center respectively, C 1 , C 2 are the center positions of the camera under different perspectives.
[0087] In a specific embodiment, the 3D reconstruction module further includes a structure optimization unit for optimizing the initial 3D point coordinates. The calculation formula is:
[0088]
[0089] In the formula, N is the number of 3D point coordinates, V i is the perspective set of the i-th 3D point coordinate, p ij is the projection position of the i-th 3D point coordinate under the j-th perspective, ρ is the robust loss function, π is the projection function, K is the internal parameter matrix of the camera, X i represents the coordinate of the i-th 3D point, R j represents the rotation matrix of the j-th perspective, and t j represents the translation vector of the j-th perspective.
[0090] Specifically, the 3D point cloud data is optimized to reduce the reprojection error, thereby improving the accuracy of 3D reconstruction.
[0091] In a specific embodiment, the 3D model construction unit further includes: smoothing the 3D model.
[0092] Specifically, in the present invention, through functions such as camera calibration, 3D coordinate recovery of feature points, 3D model construction, 3D pose estimation, and 3D motion trajectory tracking, the recovery and tracking of the human body structure and pose from 2D images to 3D space are realized.
[0093] Specifically, the dimension measurement module uses 3D geometric measurement methods to measure and evaluate the dimensions of key parts of the human body, and obtains the measurement and evaluation results, specifically including:
[0094] On the basis of 3D reconstruction, using 3D geometric measurement algorithms, the dimensions of key parts of the human body are measured, such as height, shoulder width, chest circumference, etc. Through comparison with the standard database, accurate evaluation of human body dimensions is realized.
[0095] Specifically, the pose estimation module uses deep learning algorithms to estimate the human body pose and obtains the human body pose estimation result, specifically including:
[0096] Based on 3D reconstruction, deep learning algorithms (such as convolutional neural networks, recurrent neural networks, etc.) are used to estimate human poses. When training the model, a large number of human pose datasets are used for training to improve the accuracy of pose estimation. The captured human images are processed in real time to output the 3D coordinates of human joint points, thereby obtaining the pose information of the human body.
[0097] See Figure 2 , the embodiment of the present invention also discloses a real-time 3D human perception method based on a monocular RGB image, including:
[0098] S100: Collect human images and preprocess the human images;
[0099] S200: Extract the human feature points in the preprocessed human images, match the feature points in the current frame with those in the key frames, and obtain the motion information of the human body;
[0100] S300: Perform 3D reconstruction based on the motion information to obtain a 3D reconstruction result;
[0101] S400: Use 3D geometric measurement methods to measure and evaluate the dimensions of key parts of the human body to obtain measurement and evaluation results;
[0102] S500: Use deep learning algorithms to estimate human poses to obtain human pose estimation results;
[0103] S600: Display the 3D reconstruction result, the measurement and evaluation result, and the human pose estimation result in a visual form.
[0104] Specifically, the present invention captures human images through a monocular RGB camera and performs 3D reconstruction, dimension measurement, and pose estimation in real time. This system combines advanced computer vision technology and deep learning algorithms to achieve precise perception and understanding of human shape and pose.
[0105] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0106] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A real-time three-dimensional human perception system based on monocular RGB images, characterized in that: include: A data acquisition and preprocessing module, which acquires human body images and preprocesses the human body images; The feature extraction and matching module extracts the human body feature points in the preprocessed human body image, and matches the feature points in the current frame with the key frame to obtain the human body motion information; A three-dimensional reconstruction module performs three-dimensional reconstruction based on the motion information to obtain a three-dimensional reconstruction result; The size measurement module uses three-dimensional geometric measurement methods to measure and evaluate the size of key parts of the human body and obtain measurement and evaluation results; The posture estimation module uses a deep learning algorithm to estimate the human body posture and obtain the human body posture estimation result; The result display module displays the three-dimensional reconstruction result, the measurement and evaluation result, and the human body posture estimation result in a visual form.
2. The real-time three-dimensional human perception system based on monocular RGB images according to claim 1, characterized in that: The feature extraction and matching module includes: A human feature extraction unit extracts the preprocessed human image to obtain a set of key points of the human body; A feature point direction determination unit determines, for each key point, the direction information between the key points according to the position of the key point in the human body image and the relationship with the adjacent feature points; A chain feature construction unit, which constructs a chain feature of a human body by using the extracted key points and the direction information between the key points; A feature vector construction unit, which constructs a feature vector for each chain feature; The label vector construction unit assigns a unique label vector to each chain feature; The iteration unit repeats the feature extraction and matching process for the key frames in the preprocessed human body image, extracts and stores the chain features and their feature vectors and label vectors, generates and outputs the human body's motion information.
3. The real-time three-dimensional human perception system based on monocular RGB images according to claim 2, characterized in that: The key points of the human body include: shoulders, elbows, hands, wrists, and knees.
4. The real-time three-dimensional human perception system based on monocular RGB images according to claim 2, characterized in that: The feature vector includes the length, direction, and curvature of the chain feature.
5. The real-time three-dimensional human perception system based on monocular RGB images according to claim 1, characterized in that: The three-dimensional reconstruction module comprises: A monocular RGB camera to capture the image of the human body to be tested; The camera calibration unit determines the internal and external parameters of the monocular RGB camera and establishes the relationship between the camera image pixel position and the human body 3D point position; The feature point three-dimensional point coordinate recovery unit uses the human body motion information generated by the feature extraction and matching module and the camera calibration result to recover the three-dimensional point coordinates of the feature point in the three-dimensional space; A three-dimensional model building unit, which builds a three-dimensional model of a human body according to the three-dimensional point coordinates, including: a bone structure and a muscle structure; The 3D posture estimation unit estimates the 3D posture of the human body based on the constructed 3D model, including: joint angle information and body orientation information; The 3D motion trajectory tracking unit combines the human feature point matching results of continuous frames to track the motion trajectory of the human body in 3D space and outputs dynamic 3D reconstruction results.
6. The real-time three-dimensional human perception system based on monocular RGB images according to claim 5, characterized in that: The calculation formula of the camera calibration unit is: Where (x1, y1) and (x2, y2) are a pair of matching feature points, s and ′s′ are scale factors, X, Y, Z are the point coordinates in the world coordinate system, K is the camera's internal parameter matrix, t is the translation vector, R is the rotation matrix, and R T R=I,R T is the transpose of R, and I represents the identity matrix.
7. The real-time three-dimensional human perception system based on monocular RGB images according to claim 5, characterized in that: The calculation formula of the feature point three-dimensional point coordinate recovery unit is: Where P1 and P2 are a pair of matching feature points, P1 = (x1, y1), P2 = (x2, y2), a1 and a2 are the direction vectors of the two feature points, λ1 and λ2 are the scale factors of the distances from the two feature points to the camera center, and C1 and C2 are the center positions of the camera at different viewing angles.
8. The real-time three-dimensional human perception system based on monocular RGB images according to claim 7, characterized in that: The 3D reconstruction module also includes a structure optimization unit to optimize the initial 3D point coordinates. The calculation formula is: Where N is the number of three-dimensional point coordinates, V i is the viewing angle set of the i-th 3D point coordinates, p ij is the projection position of the i-th 3D point coordinates under the j-th perspective, ρ is the robust loss function, π is the projection function, K is the internal parameter matrix of the camera, X i represents the coordinates of the i-th 3D point, R j represents the rotation matrix of the jth view, t j Represents the translation vector of the j-th view.
9. The real-time three-dimensional human perception system based on monocular RGB images according to claim 5, characterized in that: The three-dimensional model building unit further includes: performing smoothing processing on the three-dimensional model.
10. A real-time 3D human perception method based on a monocular RGB image using the real-time 3D human perception system based on a monocular RGB image according to any one of claims 1 to 9, characterized in that: include: S100: collecting a human body image and preprocessing the human body image; S200: extracting human body feature points in the preprocessed human body image, and matching the feature points in the current frame with the key frame to obtain human body motion information; S300: Performing three-dimensional reconstruction based on the motion information to obtain a three-dimensional reconstruction result; S400: using a three-dimensional geometric measurement method to measure and evaluate the dimensions of key parts of the human body, and obtaining measurement and evaluation results; S5 00: Use deep learning algorithm to estimate human body posture and obtain human body posture estimation result; S600: Displaying the three-dimensional reconstruction result, the measurement and evaluation result, and the human body posture estimation result in a visual form.
Citation Information
Cited By
Distance measurement system based on artificial intelligence
CN120593631A