A method for multi-person 3D reconstruction in a wide field of view and large scene
Through the end-to-end large-scene single-image multi-person reconstruction framework and ground-guided progressive positioning method, the deep ambiguity problem of multi-person three-dimensional reconstruction in wide-field large scenes is solved, and the global spatially consistent multi-person pose and shape reconstruction is achieved, and the position prediction accuracy is improved.
Patent Information
- Application Number
- CN202210778162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-29
AI Technical Summary
Existing methods cannot achieve multi-person reconstruction of depth sort consistency on wide field of view large scene datasets, especially in scenes with hundreds of people, where people cannot accurately obtain the three-dimensional spatial position and pose of people.
Using an end-to-end large-scene single-image multi-person reconstruction framework, a human-centered scale adaptive hierarchical representation scheme is designed. Combined with a ground-guided progressive positioning method, the scene-level global 3D positioning is converted into local 2D positioning and 3D offset by estimating scene-level camera parameters and public ground, to achieve accurate global spatial positioning of multiple people, and scene-level fine-tuning is performed in the test stage.
The multi-person pose and shape reconstruction with global spatial consistency in large scenarios is realized, which improves the accuracy of people's position prediction in the new scene, overcomes the depth ambiguity problem under single-color camera acquisition, and is suitable for multi-person three-dimensional reconstruction of large scene images.
Smart Images

Figure CN115131504B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional vision, and relates to a method for multi-person three-dimensional reconstruction in a wide field of view and large scene. Background Art
[0002] The three-dimensional reconstruction of the human body refers to recovering the pose and shape of a person from the input pictures or videos. The geometric and motion information provided by it has wide applications in games and movies. With the development of fields such as deep learning, computer vision, and computer graphics, the technology of human body three-dimensional reconstruction based on images has become a research hotspot in computer vision. Due to the difficulty of 3D annotation, most of the datasets used for multi-person three-dimensional reconstruction are obtained through data synthesis methods or in laboratory environments. Existing research is all based on medium and small scene datasets, lacking the analysis of wide field of view and large scene data. Wide field of view and large scene data can provide rich spatial information and are closer to the real-world scene. In a large scene, especially in a scene with hundreds of people, performing monocular multi-person reconstruction is beneficial to scene understanding and crowd analysis. In addition to the three-dimensional pose and shape of people, the accurate three-dimensional spatial position of people is also crucial for analyzing interpersonal relationships and crowd behavior.
[0003] Although wide-field large-scale datasets are closer to real-world scenarios, existing multi-person reconstruction methods have been studied based on medium and small-scale datasets. Hongsuk et al. (Hongsuk C, Gyeongsik M, JoonKyu P, et al. Learning to estimate robust 3d human mesh from in-the-wild crowded scenes[C]. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022) proposed a two-stage method, 3DCrowdNet, which uses 2D poses to distinguish different people and estimates human model parameters using joint-based regressors. This method focuses more on the accuracy of poses and shapes but ignores the 3D spatial positions of people. To obtain consistent multi-person reconstruction results, Jiang et al. (Jiang W, Kolotouros N, Pavlakos G, et al. Coherent Reconstruction of Multiple Humans from a Single Image[C]. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020) proposed the CRMH model based on Faster R-CNN (Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks[J]. Advances in neural information processing systems (NeurIPS), 2015). First, it detects people in the image and then regresses the SMPL parameters of people. During training, it uses penetration loss and depth sorting loss to obtain relatively accurate position relationships. However, this method calculates the depth of people based on the assumption of consistent human heights, which will estimate a larger depth for shorter people. To solve the inherent ambiguity of height and depth, Ugrinovic et al. (Ugrinovic N, Ruiz A, Agudo A, et al.Body Size and Depth Disambiguation in Multi-Person Reconstruction from Single Images [C] proposed a method based on multi-stage optimization to optimize the scale and 3D translation of the body meshes estimated by CRMH; the multi-stage method brings computational redundancy problems. Zhang et al. (Zhang J, Yu D, Liew J H, et al. Body meshes as points [C] In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021) proposed a single-stage method BMP that associates the depth of a person with features at different scales; Sun et al. (Sun Y, Bao Q, Liu W, et al. Monocular, one-stage, regression of multiple 3d people [C]. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), 2021) proposed ROMP, which extracts camera information and SMPL information based on the human center feature map. This method is based on the assumption of weak perspective projection and can only infer the 2D position of a person in the image; to further solve the position problem, Sun et al. (Sun Y, Liu W, Bao Q, et al. Putting people in their place: Monocular regression of 3d people in depth. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022) proposed BEV, which uses a bird's-eye view representation to simultaneously infer the body center and depth in the image; however, all of the above methods can only obtain relative depth and cannot obtain absolute positions, and cannot be directly applied to large scenes.
[0004] In view of the above problems, the present invention proposes an end-to-end large-scene single-image multi-person reconstruction framework. For large-scene images at the billion-pixel level, a human-centered scale-adaptive hierarchical representation scheme is designed, a global and local joint representation model is constructed to overcome the depth ambiguity problem under single-color camera acquisition, and globally spatially consistent multi-person pose and shape reconstruction is achieved. A ground-guided progressive positioning method is proposed. By estimating scene-level camera parameters and the common ground, the global 3D positioning at the scene level is converted into local 2D positioning and 3D offsets, realizing accurate global spatial positioning of multiple people in the scene. Scene-level fine-tuning is performed in the test stage, thereby effectively improving the accuracy of predicting the positions of people in new scenes. Summary of the Invention
[0005] (1) Technical problems to be solved by the present invention:
[0006] The purpose of the present invention is: in view of the fact that existing methods cannot obtain multi-person reconstruction results with consistent depth sorting on wide-field large-scene data sets, an end-to-end large-scene single-image multi-person reconstruction framework is proposed. For large-scene images at the billion-pixel level, a human-centered scale-adaptive hierarchical representation scheme is designed, a global and local joint representation model is constructed to overcome the depth ambiguity problem under single-color camera acquisition, and globally spatially consistent multi-person pose and shape reconstruction is achieved. A ground-guided progressive positioning method is proposed. By estimating scene-level camera parameters and the common ground, the global 3D positioning at the scene level is converted into local 2D positioning and 3D offsets, realizing accurate global spatial positioning of multiple people in the scene. Scene-level fine-tuning is performed in the test stage, thereby effectively improving the accuracy of predicting the positions of people in new scenes.
[0007] (2) To achieve the above object, the present invention adopts the following technical solutions:
[0008] A multi-person three-dimensional reconstruction method for wide-field large scenes, comprising the following steps:
[0009] S1. Preprocess the large-scene image, and obtain cropped images with different resolutions through a human-centered adaptive hierarchical representation, so that people occupy an appropriate proportion in the cropped images. On the basis of maintaining the original aspect ratio of the image, scale the cropped images to a unified size for training the network;
[0010] S2. Estimate the 2D joints of the large-scene image through an existing 2D joint estimation method, correct the estimated incorrect or missing 2D joints through an artificial correction method, and use the 2D joints to estimate the ground equation and camera internal parameters;
[0011] S3. Use the cropped images obtained by preprocessing in S1 to train the network. The network extracts features through a backbone network, and then uses three different branch networks to perform human detection, 2D position estimation, and 3D offset and human parameter model estimation respectively;
[0012] S4. Through a ground-guided progressive positioning method, use the camera internal parameters and ground equation obtained in S2, and the rough 3D position of the human body obtained based on the 2D position obtained in S3, combined with the 3D offset obtained in S3, to obtain the accurate 3D position of the human body;
[0013] S5. Perform scene-level fine-tuning on the model in the test stage, and perform multi-person reconstruction on the new scene images to obtain better 2D projection results;
[0014] S6. By merging the multi-person reconstruction results of all the cropped images and removing the repeatedly estimated people, obtain the globally spatially consistent multi-person reconstruction results in a wide-field large scene.
[0015] Preferably, the preprocessing process described in S1 mainly includes the following steps:
[0016] S101. Define the minimum and maximum heights of people in the large scene image as h min and h max , define the upper and lower bounds of the cropping area as s and e respectively, use a square sliding window to crop the large scene image, and the length of the i-th sliding window in the y direction is c i . To make the height of the person in the cropped image half of the height of the cropped image, c1 = 2×h min . For the last sliding window in the y direction, that is, the n-th sliding window, its length is c n = c1×q n-1 and where q is a proportionality coefficient;
[0017] The human-centered adaptive hierarchical representation described in S1 is as follows:
[0018]
[0019] To ensure that each person can appear completely in the cropped image, add an overlapping sliding window between two adjacent sliding windows in the y direction, and its length is half of the sum of the lengths of the adjacent sliding windows;
[0020] S102. Keep the original aspect ratio of the cropped images with different resolutions, and unify them to (512, 512) through bicubic interpolation, and fill the insufficient parts with 0.
[0021] Preferably, the estimation of the ground equation and camera parameters described in S2 mainly includes the following steps:
[0022] S201. Estimate the 2D joint points of the cropped image through the RMPE method, manually correct the estimated incorrect or missing 2D joint points, merge the obtained results to get the 2D joint point information of the large scene image, and filter the poses according to the prior information, only keeping the standing poses;
[0023] S202. Use the pinhole camera model with a focal length of f (f = f x = f y ), the principal point is the center point of the image, and the ground equation is N T P G + D = 0, where is the ground normal and ||N||2 = 1, D is a constant term reflecting the position of the ground, is a point on the ground;
[0024] S203. Define the midpoint of the left and right ankle points as and its projection point on the image is x b = (u b , v b ), the center point of the left and right shoulders is and its projection point on the image is x t = (u t , v t ), assume X b is a point on the ground, the person stands on the ground with a fixed height h, and the line passing through X b and X t is parallel to the ground normal;
[0025] S204. According to the principle of pinhole imaging, we can get where is the homogeneous coordinate of x b , K is the camera intrinsic matrix, and Z b is the depth of X b ; since X b is a point on the ground and satisfies N T X b + D = 0, we can get:
[0026]
[0027] The projection point of the midpoint of the left and right shoulders can be calculated by the following equation:
[0028]
[0029] where Z t is the depth of X t ;
[0030] S205. Solve the camera parameters and the ground equation through an optimized method. The loss function for the $i$-th person is specifically as follows:
[0031]
[0032] where $L$ 余弦 denotes the cosine distance, and $\lambda$ 角度 , $\lambda$ 模长 are the weights of the corresponding loss terms respectively;
[0033] S206. Translate the obtained ground 0.1 meters along the normal direction to obtain the real ground, rather than the ground where the ankle is located.
[0034] Preferably, the specific implementation process of S3 is as follows:
[0035] S301. Extract features from the input image through the backbone network, and then input the obtained features into three different branch networks. Each branch network consists of two ResNet blocks and batch normalization;
[0036] S302. The first branch network obtains the human center feature map, and uses a Gaussian kernel combined with the body scale to represent the possibility of the center position of the person in the feature map;
[0037] S303. The second branch network obtains the 2D position feature map, estimates the 2D coordinates and 2D offsets of the left and right ankle points, and the sum of the midpoint of the left and right ankle points and the 2D offset is the required 2D position;
[0038] S304. The third branch network obtains the SMPL and offset feature maps, and estimates the pose and shape parameters of SMPL and the 3D offset;
[0039] S305. According to the position obtained from the human center feature map, extract the corresponding parameters from the 2D position feature map and the SMPL and offset feature maps to obtain the 2D position, SMPL parameters, and 3D offset required to estimate the position and pose of the person;
[0040] S306. First, train the human center feature map and the 2D position feature map so that the subsequent learned human mesh has a suitable initial position. After 20 iterations, train the entire network, and the entire network is iterated 70 times.
[0041] Preferably, the ground-guided progressive positioning method described in S4 mainly includes the following steps:
[0042] S401. Define the projection point of the center of the human torso on the ground as the landing point The projections of the landing point on the large-scale scene image and on the cropped image are respectively and $p$局部 That is, the 2D position obtained from the 2D position feature map, p = p 局部 + t p , where is the pixel position of the upper left corner of the cropped image in the large scene image. Since P is a point on the ground, based on the predicted camera parameters and the ground equation, the rough 3D position P can be calculated by the following equation:
[0043]
[0044] where is the homogeneous coordinate of p;
[0045] S402. Obtain the accurate 3D position through the 3D position offset Δ 3d T 3D = -mean(J 脚踝 ) + Δ 3d + P, where J 脚踝 is the 3D coordinates of the left and right ankles. The human body network in the camera coordinate system can be calculated by the following equation:
[0046] M 相机 = M + T 3D
[0047] where M is the human body mesh in the SMPL standard space;
[0048] S403. Prevent the phenomenon of people penetrating the ground in the large scene by performing L1 constraint on the point with the most serious penetration of the ground in the human body mesh in the camera coordinate system.
[0049] Preferably, the penetration loss function corresponding to the penetration loss in S403 is specifically as follows:
[0050]
[0051] where v i ∈ M 相机 , is the homogeneous equation of v i , G = [N T , D] T .
[0052] Preferably, the specific implementation process of the scene-level fine-tuning described in S5 is as follows:
[0053] S501. For a new scene image, obtain the cropped image through the preprocessing process described in S1, and obtain the corresponding ground equation and camera parameters through S2;
[0054] S502. Fix most of the network described in S3, and only optimize the two branch networks for obtaining the 2D position feature map as well as the SMPL and offset feature maps. After the network is iterated for 5 generations, a scene-level fine-tuned model is obtained.
[0055] Preferably, the specific implementation process of the merging process described in S6 is as follows:
[0056] S601. According to the human body center feature map of all the cropped images obtained in S302, scan from left to right and from top to bottom according to the positions of the cropped images in the large scene image.
[0057] S602. Set the threshold to be If the distance between two center points is less than this threshold, it means that these two center points are the center points of the same person. Calculate the distance between the center point and the boundary of the cropped image where it is located. The farther the center point is from the boundary, the lower the possibility that the person it represents is truncated. Retain this center point.
[0058] (3) The beneficial effects of the present invention include the following points:
[0059] (1) The present invention proposes a multi-person 3D reconstruction method under a wide field of view and large scene. Through an end-to-end large scene single-image multi-person reconstruction framework, the above method can achieve globally spatially consistent multi-person pose and shape reconstruction; this method proposes a human-centered scale adaptive hierarchical representation scheme; and proposes a ground-guided progressive positioning method to overcome the depth ambiguity problem under single color camera acquisition.
[0060] (2) The present invention provides a multi-person 3D reconstruction method under a wide field of view and large scene, which solves the problem that the prior art cannot obtain accurate multi-person 3D reconstruction results of spatial positions in large scene images.
[0061] (3) The present invention proposes a human-centered scale adaptive hierarchical representation scheme. By setting sliding windows of different sizes, the human body has a suitable proportion in the cropped image, which is more suitable for the network to learn and improves the robustness of the prediction results; to obtain the accurate 3D positions of people in the large scene, a ground-guided progressive positioning method is proposed. By estimating scene-level camera parameters and the common ground, the global 3D positioning at the scene level is converted into local 2D positioning and 3D offset, realizing accurate global spatial positioning of multiple people in the scene.
[0062] (4) In order to solve the problem of misalignment of 2D projections in new scenes during the testing process, the present invention fine-tunes the model during the testing process. This fine-tuning does not involve new operations and has a low time cost, and can effectively improve the position prediction accuracy of people in new scenes.
[0063] (5) A method for multi-person 3D reconstruction in a wide field of view and large scene is proposed in the present invention, achieving the optimal results on the datasets Panoptic and Crowd-Location. Among them, Crowd-Location is a newly proposed large-scene test dataset, including 20 images of 2 scenes, providing rich annotation information, including bounding boxes, 2D poses, projections of landing points on the images, and homography matrices for expressing the perspective transformation between the real-world ground and the corresponding images. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 FIG. is a block diagram of end-to-end large-scene single-image multi-person reconstruction in a method for multi-person 3D reconstruction in a wide field of view and large scene proposed by the present invention;
[0065] Figure 2 FIG. is a schematic diagram of multi-person reconstruction results on the Crowd-Location test set in a method for multi-person 3D reconstruction in a wide field of view and large scene proposed by the present invention;
[0066] Figure 3 FIG. is a comparison chart of qualitative results between a method for multi-person 3D reconstruction in a wide field of view and large scene proposed by the present invention and the mainstream multi-person reconstruction methods in the prior art. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0068] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0069] Embodiment 1:
[0070] A method for multi-person 3D reconstruction in a wide field of view and large scene includes the following steps:
[0071] S1. Preprocess the large-scene images, obtain cropped images of different resolutions through human-centered adaptive hierarchical representation, so that people occupy an appropriate proportion in the cropped images, and scale the cropped images to a unified size while maintaining the original aspect ratio of the images for training the network. The preprocessing process in S1 mainly includes the following steps:
[0072] S101. Define the minimum and maximum heights of people in the large - scene image as h min and h max . Define the upper and lower bounds of the cropping region as s and e respectively. Use a square sliding window to crop the large - scene image. The length of the i - th sliding window in the y - direction is c i . To make the height of the person in the cropped image be half of the height of the cropped image, c1 = 2×h min . For the last sliding window in the y - direction, that is, the n - th sliding window, its length is c n = c1×q n-1 and where q is the proportionality coefficient. The human - centered adaptive hierarchical representation described in S1 is as follows:
[0073]
[0074] . To ensure that each person can appear completely in the cropped image, add an overlapping sliding window between two adjacent sliding windows in the y - direction, and its length is half of the sum of the lengths of the adjacent sliding windows;
[0075] S102. Keep the original aspect ratio of the cropped images with different resolutions, and unify them to (512, 512) through bicubic interpolation. Fill the insufficient part with 0;
[0076] S2. Estimate the 2D joints of the large - scene image through the existing 2D joint estimation method, and correct the estimated wrong or missing 2D joints manually. Use the 2D joints to estimate the ground equation and camera internal parameters. The ground estimation and camera parameter estimation described in S2 mainly include the following steps:
[0077] S201. Estimate the 2D joints of the cropped image through the RMPE method, manually correct the estimated wrong or missing 2D joints, merge the obtained results to get the 2D joint information of the large - scene image, and filter the poses according to the prior information, only keeping the standing poses;
[0078] S202. Use the pinhole camera model, whose focal length is f (f = f x = f y ), the principal point is the center point of the image, and the ground equation is N T P G + D = 0, where is the ground normal, and ||N||2 = 1, D is the constant term, reflecting the position of the ground, is the point on the ground;
[0079] S203. Define the mid - point of the left and right ankle points as Its projection point on the image is x b =(u b , v b ), and the center point of the left and right shoulders is Its projection point on the image is x t =(u t , v t ). Assume X b is a point on the ground. A person stands on the ground with a fixed height h. The line passing through X b and X t is parallel to the ground normal;
[0080] S204. According to the principle of pinhole imaging, we can get where is the homogeneous coordinate of x b , K is the camera internal parameter matrix, and Z b is the depth of X b ; because X b is a point on the ground and satisfies N T X b +D = 0, we can get:
[0081]
[0082] The projection point of the midpoint of the left and right shoulders can be calculated by the following equation:
[0083]
[0084] where Z t is the depth of X t ;
[0085] S205. Solve the camera parameters and the ground equation through an optimization-based method. The loss function of the i-th person is specifically as follows:
[0086]
[0087] where L 余弦 represents the cosine distance, λ 角度 , λ 模长 are the weights of the corresponding loss terms respectively;
[0088] S206. Translate the obtained ground 0.1 meters along the normal direction to obtain the real ground, rather than the ground where the ankles are located;
[0089] S3. Use the image preprocessed in S1 to train the network. This network realizes feature extraction through the backbone network, and then uses three different branch networks to perform human detection, 2D position estimation, and 3D offset and human parameter model estimation respectively. The specific implementation process of S3 is as follows:
[0090] S301. Extract features from the input image through the backbone network, and then input the obtained features into three different branch networks. Each branch network consists of two ResNet blocks and batch normalization;
[0091] S302. The first branch network obtains the human center feature map, and uses a Gaussian kernel combined with the body scale to represent the possibility of the center position of the person in the feature map;
[0092] S303. The second branch network obtains the 2D position feature map, estimates the 2D coordinates and 2D offsets of the left and right ankle points, and the sum of the midpoint of the left and right ankle points and the 2D offset is the required 2D position;
[0093] S304. The third branch network obtains the SMPL and offset feature maps, and estimates the pose and shape parameters of SMPL and the 3D offset;
[0094] S305. According to the position obtained from the human center feature map, extract the corresponding parameters from the 2D position feature map and the SMPL and offset feature maps, and obtain the 2D position, SMPL parameters and 3D offset required to estimate the position and pose of the person;
[0095] S306. First, train the human center feature map and the 2D position feature map so that the subsequent learned human mesh has a suitable initial position. After 20 iterations, train the entire network, and the entire network iterates 70 times;
[0096] S4. Through the ground-guided progressive positioning method, use the camera parameters and ground equation obtained in S2, and the 2D position obtained in S3 to obtain the rough 3D position of the human body, and use the 3D offset obtained in S3 to obtain the accurate 3D position of the human body; The ground-guided progressive positioning method described in S4 mainly includes the following steps:
[0097] S401. Define the projection point of the human torso center on the ground as the landing point The projections of the landing point on the large-scale scene image and on the cropped image are respectively and p 局部 That is, the 2D position obtained from the 2D position feature map, p = p 局部 +t p where is the pixel position of the upper left corner of the cropped image in the large-scale scene image. Since P is a point on the ground, according to the predicted camera parameters and ground equation, the rough 3D position P can be calculated by the following equation:
[0098]
[0099] where is the homogeneous coordinate of p;
[0100] S402. Obtain the precise 3D position T through the 3D position offset Δ 3d =-mean(J 3D ) + Δ 脚踝 + P, where J 3d is the 3D coordinates of the left and right ankles. The human body network in the camera coordinate system can be calculated by the following equation: 脚踝 M
[0101] M 相机 = M + T 3D
[0102] where M is the human body mesh in the SMPL standard space;
[0103] S403. Prevent the phenomenon of people penetrating the ground in large scenes by performing L1 constraint on the points in the human body mesh in the camera coordinate system that penetrate the ground most severely. The specific penetration loss function corresponding to the penetration loss in S403 is as follows:
[0104]
[0105] where v i ∈ M 相机 , is the homogeneous equation of v i , G = [N T , D] T .
[0106] S5. Perform scene-level fine-tuning on the model during the test phase to obtain better 2D projection results. The scene-level fine-tuning does not involve new operations and has low time cost;
[0107] The specific implementation process of the scene-level fine-tuning described in S5 is as follows:
[0108] S501. For the pictures of a new scene, obtain the cropped images through the preprocessing process described in S1, and obtain the corresponding ground equation and camera parameters through S2;
[0109] S502. Fix most of the network described in S3, and only optimize the two branch networks for obtaining the 2D position feature map and the SMPL and offset feature maps. After the network iterates 5 generations, obtain the model after scene-level fine-tuning;
[0110] S6. Merge the multi-person reconstruction results of all the cropped images, remove the repeatedly estimated people, and obtain the multi-person reconstruction results with consistent depth sorting in the wide-field large scene. The specific implementation process of the merging process described in S6 is as follows:
[0111] S601. According to the human center feature maps of all the cropped images obtained in S302, scan from left to right and top to bottom according to the positions of the cropped images in the large scene image.
[0112] S602. Set the threshold to of the width of the cropped image. If the distance between two center points is less than this threshold, it means that these two center points are the center points of the same person. Calculate the distance between the center point and the boundary of the cropped image where it is located. The farther the center point is from the boundary, the lower the possibility that the person it represents is truncated. Retain this center point.
[0113] Embodiment 2:
[0114] Please refer to Figures 1-3 , and the specific implementation process of a multi-person three-dimensional reconstruction method in a wide field of view large scene described in Embodiment 1 is as follows:
[0115] (1) Data preprocessing:
[0116] In the present invention, the publicly available Human3.6M, MuCo-3DHP, Agora, and PANDA datasets are used. The above datasets include crowd activities in various situations. In order to use the PANDA dataset for training, 2D joint point annotations are performed on four scenes in the PANDA dataset, namely OCT Habour, Basketball Court, University Campus, and Huaqiangbei. First, 2D joint point detection is performed using RMPE, and the estimated incorrect or missing joint points are manually corrected to obtain the dataset PANDA-Pose. The large scene dataset is used to obtain cropped images using a human-centered adaptive hierarchical representation scheme. In order to unify training, on the basis of maintaining the original aspect ratio of the image, the image is fixed to (512, 512) through bicubic interpolation, and the insufficient area is filled with 0.
[0117] (2) Ground equation and camera parameter estimation:
[0118] The datasets Human3.6M, MuCo-3DHP, and Agora provide the camera intrinsics and camera extrinsics. Assuming that in a unified coordinate system, a person stands on the XOZ plane, which is the ground, and the Y direction is the direction in which the person stands. The ground in the unified coordinate system is rotated using the camera extrinsics to obtain the ground equation in the camera coordinate system.
[0119] For the PANDA-Pose dataset, the ground equation and camera parameters are estimated using 2D joint points. First, the 2D poses are filtered using prior information, and only the standing poses are retained. Assuming that a person has a fixed height, an optimization-based method is used to solve the ground equation and camera parameters.
[0120] (3) Human body shape and position information estimation network:
[0121] During the training process, after the image is input into the backbone network to extract features, the features are respectively input into three branch networks; the first branch network predicts the central feature map of the person, the second branch network predicts the 2D position feature map, and the third branch network predicts the SMPL and offset feature maps; first, the backbone network, the first branch network, and the second branch network are trained to give the human body mesh an initial position, and then the entire network is trained to obtain the SMPL parameters, 2D position, and 3D displacement required for human body shape and position estimation.
[0122] Specifically, the HRNet-32 is used as the backbone network to extract features, and each branch network consists of two ResNet blocks and batch normalization.
[0123] (4) Ground-guided progressive positioning method:
[0124] Based on the above steps, the ground-guided progressive positioning method is used to obtain the accurate 3D position of the person; the 2D position estimated by the above network is the projection of the landing point on the picture. According to the principle of pinhole imaging, using the ground equation and camera parameters, the 2D position is converted into the 3D position of the landing point to obtain the rough 3D position of the person; the SMPL estimated by the above network is used to obtain the human body mesh, and the origin of the coordinate system of the human body mesh is converted into the midpoint of the left and right ankle points. The accurate 3D position of the person is obtained by adding the 3D offset estimated by the above network to the rough 3D position.
[0125] (5) Scene-level fine-tuning and merging:
[0126] During the multi-person reconstruction of new scene images in the test phase, the model is fine-tuned at the scene level to obtain better 2D projection results. The scene-level fine-tuning does not involve new operations and has a low time cost. Most of the parameters of the model are fixed, and only the two branch networks for obtaining the 2D position feature map and the SMPL and offset feature maps are optimized. After the network is iterated 5 times, the model after scene-level fine-tuning is obtained; the multi-person reconstruction results of all images are merged, and the repeatedly estimated people are removed to obtain the multi-person reconstruction results with a consistent global spatial distribution.
[0127] As Figure 1 shown, the end-to-end large-scene single-image multi-person reconstruction framework proposed by the present invention is shown. The large-scene image is cropped according to the human-centered hierarchical representation method, and the 2D joints of the large-scene image are obtained by using the existing 2D joint point estimation method and manual correction. The ground equation and camera parameters are estimated, the human body shape and position information is estimated by the network, and the accurate 3D position of the person is obtained through the ground-guided progressive positioning method.
[0128] AsFigure 2 As shown, it presents the multi-person 3D reconstruction results of the present invention on the Crowd-Location test set. It can be fully shown from the results that the large-scale scene multi-person reconstruction method proposed by the present invention can accurately obtain the 3D positions of people, obtain multi-person reconstruction results with consistent global spatial distribution, and can obtain reasonable postures.
[0129] As Figure 3 shown, it presents the comparison of the qualitative results between the present invention and the current mainstream multi-person reconstruction methods. It can be seen that this method can better predict the 3D positions of people. For the reconstructed human body models, this method can obtain reasonable human body shape and posture estimations.
[0130] Table 1 and Table 2 respectively list the comparison of the quantitative results of the present invention and the current mainstream multi-person reconstruction methods in the S1 scenario and S2 scenario of the Crowd-Location dataset; SMAP was proposed by Zhen et al. (Zhen J, Fang Q, Sun J, et al. SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation [C]. In Proceedings of the IEEE / CVF European Conference on Computer Vision (ECCV), 2020) in 2020, CRMH was proposed by Jiang et al. (Jiang W, Kolotouros N, Pavlakos G, et al. Coherent Reconstruction of Multiple Humans from a Single Image [C]. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020) in 2020, ROMP was proposed by Sun et al. (Sun Y, Bao Q, Liu W, et al. Monocular, one-stage, regression of multiple 3d people [C]. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), 2021) in 2021, and BEV was proposed by Sun et al. (Sun Y, Liu W, Bao Q, et al. Putting people in their place: Monocular regression of 3d people in depth. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022) in 2022. These multi-person reconstruction methods cannot be directly applied to large-scale scene images. The multi-person reconstruction results of large-scale scene images are obtained in a way corresponding to these methods. Cropped images are obtained as the input of the multi-person reconstruction method through a human-centered adaptive hierarchical representation method, and the camera parameters used are the camera parameters obtained by the present invention;
[0131] For CRMH, the global position is obtained by using the transformation from the bounding box coordinate system to the global coordinate system. According to the result obtained by CRMH, the camera parameters of the weak perspective projection of the $i$-th person are $\pi$ i $=$ {s i , x i , y i}, and its bounding box $B$ i $=$ [x min , y min , x max , y max . The center of the bounding box is $c$ i $=$ [(x min + x max ) / 2, (y min + y max ) / 2], and the scale is $\alpha$ i $=$ max(x max - x min , y max - y min ). According to these parameters, the depth of the $i$-th person Based on the calculated depth, the global spatial position of the $i$-th person can be calculated by the following equation:
[0132]
[0133] where $w$ and $h$ are the width and height of the large-scale scene image respectively;
[0134] For ROMP, using the result obtained by ROMP and the camera parameters estimated by the present invention, the global spatial position of the person is solved by the PNP algorithm;
[0135] For SMAP and BEV, these two methods can directly obtain the spatial position of the person in the input image, but each image has independent camera coordinate parameters, and its focal length is the width $w$ of the image c , and the principal point is the center point of the image; in the SMAP method, the spatial position of the $j$-th joint point of the $i$-th person is $T$ ij $=$ {X ij , Y ij , Z ij}, and the projection point on the large-scale scene is {x ij , y ij}. To unify it into the large-scale scene camera coordinate system, the depth $Z$ of the joint point ij_全局 $=$ $Z$ ij × $f$ / $w$ c . Using the principle of perspective projection, the 3D position of the $j$-th joint point of the $i$-th person in the large-scale scene camera coordinate system is:
[0136]
[0137] For BEV, only calculate the offset in the independent camera coordinate system of the cropped image result instead of all joint points.
[0138] The evaluation metrics for the quantitative results are Matched, which is used to measure the matching rate between the prediction result and the ground truth; PCOD, which is used to measure the accuracy of the depth ranking between people; PPDError, which is used to measure the accuracy of the distance prediction result between people; OKS, which is used to measure the accuracy between the predicted pose and the ground truth pose. The datasets are Scenario S1 and Scenario S2 in Crowd-Location. It can be seen that the reconstruction result of the present invention achieves the best effect in terms of spatial position in the problem of multi-person 3D reconstruction in a wide field of view and large scene, and at the same time obtains reasonable poses;
[0139] Table 1
[0140]
[0141] Table 2
[0142]
[0143] Table 3 lists the comparison of the quantitative results of the present invention and the current mainstream multi-person reconstruction methods on the Panoptic dataset; SMAP was proposed by Zhen et al. (Zhen J, Fang Q, Sun J, et al. SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation [C]. In Proceedings of the IEEE / CVF European Conference on Computer Vision (ECCV) 2020) in 2020, CRMH was proposed by Jiang et al. (Jiang W, Kolotouros N, Pavlakos G, et al. Coherent Reconstruction of Multiple Humans from a Single Image [C]. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020) in 2020, ROMP was proposed by Sun et al. (Sun Y, Bao Q, Liu W, et al. Monocular, one-stage, regression of multiple 3d people [C]. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), 2021) in 2021, BEV was proposed by Sun et al. (Sun Y, Liu W, Bao Q, et al. Putting people in their place: Monocular regression of 3d people in depth. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022) in 2022, 3DCrowdNet was proposed by Hongsuk et al. (Hongsuk C, Gyeongsik M, JoonKyu P, et al. Learning to estimate robust 3d human mesh from in-the-wild crowded scenes [C].It was proposed in 2022 in the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR). The evaluation metrics for quantitative results include MPJPE, which is used to measure the Euclidean distance between the predicted 3D human model and the real human model; RtError, which is used to measure the Euclidean distance between the predicted root node and the real root node; PCOD, which is used to measure the accuracy of depth ranking between people; PPDError, which is used to measure the accuracy of the distance prediction result between people; the dataset is the four common sub-datasets of Panoptic, namely Haggling, Mafia, Ultim., and Pizza; it can be seen that the reconstruction results of the present invention can also achieve good results in the problem of multi-person 3D reconstruction in small and medium-sized scenes.
[0144] Table 3
[0145]
[0146]
[0147] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for multi-person 3D reconstruction in a wide field of view and large scene, characterized in that: It includes the following steps: S1. Preprocess the large-scale scene image, obtain cropped images of different resolutions through a human-centered adaptive hierarchical representation, so that the person occupies an appropriate proportion in the cropped image, and scale the cropped image to a unified size while maintaining the original aspect ratio of the image for training the network; S2. Estimate the 2D joints of the large-scale scene image through existing 2D joint estimation methods, correct the estimated incorrect or missing 2D joints through manual correction methods, and use the 2D joints to estimate the ground equation and camera internal parameters; S3. Use the cropped images obtained by preprocessing in S1 to train the network. The network realizes feature extraction through a backbone network, and then uses three different branch networks to perform human detection, 2D position estimation, and 3D offset and human parameter model estimation respectively; S4. Through a ground-guided progressive positioning method, use the camera internal parameters and ground equation obtained in S2, and the 2D position obtained based on S3 to obtain the rough 3D position of the human body, and combine the 3D offset obtained in S3 to obtain the accurate 3D position of the human body; The ground-guided progressive positioning method includes the following steps: S401. Define the projection point of the center of the human torso on the ground as the landing point. , and the projections of the landing point on the large - scene image and on the cropped image are respectively and , , which is the 2D position obtained from the 2D position feature map. , where is the pixel position of the upper - left corner of the cropped image in the large - scene image. Because is a point on the ground, according to the predicted camera parameters and the ground equation, the rough 3D position is calculated by the following equation: Among them , is the homogeneous coordinates of; is the camera intrinsic matrix; is the ground normal, and ; is the constant term, reflecting the position of the ground; S402. Obtain an accurate 3D position through a 3D position offset Obtain an accurate 3D position , where are the 3D coordinates of the left and right ankles. The human body network in the camera coordinate system is calculated through the following equation: Among them is the human body mesh in the SMPL standard space; S403. Perform L1 constraint on the points in the human body grid in the camera coordinate system that penetrate the ground most severely to prevent the phenomenon of people penetrating the ground in the large-scale scene; S5. Perform scene-level fine-tuning on the model in the test stage, and perform multi-person reconstruction on the new scene image to obtain better 2D projection results; S6. Merge the multi-person reconstruction results of all cropped images, remove the repeatedly estimated people, and obtain a globally spatially consistent multi-person reconstruction result in the wide-field large-scale scene.
2. The multi-person three-dimensional reconstruction method in a wide field of view and large scene according to claim 1, wherein: The preprocessing process described in S1 includes the following steps: S101. Define the minimum and maximum heights of people in the large-scale scene image as and , define the upper and lower bounds of the cropping region as and . Crop the large-scale scene image using a square sliding window. In the direction, the length of the th sliding window is . To make the height of the person in the cropped image half of the height of the cropped image, . In the direction, the last sliding window, that is, the nth sliding window, has a length of and , where is the proportionality coefficient; The human-centered adaptive hierarchical representation described in S1 is as follows: To ensure that each person appears completely in the cropped image, add an overlapping sliding window between two adjacent sliding windows in the y direction, and its length is half of the sum of the lengths of the adjacent sliding windows; S102. Keep the original aspect ratio of the cropped images of different resolutions, and unify them to (512, 512) through bicubic interpolation method, and fill the insufficient parts with 0.
3. A method for multi-person three-dimensional reconstruction in a wide field of view and large scene according to claim 1, characterized in that: The estimation of the ground equation and camera parameters described in S2 includes the following steps: S201. Estimate the 2D joints of the cropped image through the RMPE method, manually correct the estimated incorrect or missing 2D joints, merge the obtained results to obtain the 2D joint information of the large-scale scene image, and filter the poses according to prior information, only keep the standing pose; S202. Use a pinhole camera model with a focal length of f , f = f x =f y . The principal point is the center point of the image, and the ground equation is , where is a point on the ground; S203. Define the midpoint of the left and right ankle points as , and its projection point on the image is . Define the center point of the left and right shoulders as , and its projection point on the image is . Assume that is a point on the ground. A person stands on the ground with a fixed height . The straight line passing through and is parallel to the ground normal. S204. Obtained according to the principle of pinhole imaging , where is the homogeneous coordinates of , and is the depth of ; since is a point on the ground and satisfies , we get: The projection point of the midpoint of the left and right shoulders in homogeneous coordinates is calculated using the following equation: Among them is depth; S205. Solve the camera parameters and the ground equation by an optimized method. The loss function of the n-th person is specifically as follows: where denotes the cosine distance, , are the weights of the corresponding loss terms respectively; S206. Translate the obtained ground 0.1 meters along the normal direction to obtain the real ground, rather than the ground where the ankles are located.
4. A method for multi-person three-dimensional reconstruction in a wide field of view and large scene according to claim 1, characterized in that: The specific implementation process of S3 is as follows: S301. Extract features from the input image through the backbone network, and then input the obtained features into three different branch networks. Each branch network consists of two ResNet blocks and batch normalization; S302. The first branch network obtains the human body center feature map, and uses a Gaussian kernel combined with the body scale to represent the possibility of the center position of the person in the feature map; S303. The second branch network obtains a 2D position feature map, estimates the 2D coordinates and 2D offsets of the left and right ankle points, and the sum of the midpoint of the left and right ankle points and the 2D offset is the required 2D position; S304. The third branch network obtains the SMPL and offset feature maps, estimates the pose and shape parameters of SMPL and the 3D offset; S305. According to the position obtained from the human body center feature map, extract the corresponding parameters from the 2D position feature map and the SMPL and offset feature maps, and obtain the 2D position, SMPL parameters and 3D offset required to estimate the position and pose of the person; S306. First, train the human body center feature map and the 2D position feature map so that the subsequent learned human body mesh has a suitable initial position. After 20 iterations, train the entire network, and the entire network iterates 70 times.
5. A multi-person three-dimensional reconstruction method in a wide field of view and large scene according to claim 1, characterized in that: The penetration loss function corresponding to the penetration loss in S403 is specifically as follows: wherein , is a homogeneous equation of .
6. A method for multi-person three-dimensional reconstruction in a wide field of view and large scene according to claim 1, characterized in that: The specific implementation process of the scene-level fine-tuning described in S5 is as follows: S501. For a new scene image, obtain a cropped image through the preprocessing process described in S1, and obtain the corresponding ground equation and camera parameters through S2; S502. Fix most of the network described in S3, and only optimize the two branch networks that obtain the 2D position feature map and the SMPL and offset feature maps. After 5 iterations of the network, obtain the model after scene-level fine-tuning.
7. A method for multi-person three-dimensional reconstruction in a wide field of view and large scene according to claim 1, characterized in that: The specific implementation process of the merging process described in S6 is as follows: S601. According to the human body center feature maps of all the cropped images obtained in S302, scan from left to right and from top to bottom according to the positions of the cropped images in the large scene image; S602. Set the threshold to be of the width of the cropped image. If the distance between two center points is less than this threshold, it means that these two center points are the center points of the same person. Calculate the distance between the center point and the boundary of the cropped image where it is located. The farther the center point is from the boundary, the lower the possibility that the person it represents is truncated. Retain this center point.
Citation Information
Patent Citations
Aircraft ground guidance system and method based on controller instruction semantic recognition
CN111667831A
Method and device for generating depth image and camera external parameters by using multiple color pictures
CN113496521A