Object reconstruction method and apparatus, device, and storage medium
Through the adaptive division of image blocks and ground geometric constraints, the accuracy problem of three-dimensional human body reconstruction at a long-distance perspective of large scenes is solved, and high-precision three-dimensional morphology and motion recovery of dense populations is achieved.
Patent Information
- Application Number
- PCT/CN2024/110506
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-31
AI Technical Summary
The prior art is difficult to achieve high-precision three-dimensional human body reconstruction in images from large scenes and long-distance perspectives, especially the three-dimensional morphology and movement recovery of dense populations, and there is a problem of insufficient model receptive field and lack of spatial reference information.
The method of adaptively dividing image blocks is adopted, and the target image is divided into multiple image blocks based on the human body reference size, and the object reconstruction result is generated using the trained object reconstruction model, combined with the ground geometric prior constraints, the spatial consistency of the reconstruction result is improved.
High-precision three-dimensional human body reconstruction in large scenes and long-distance perspectives can be realized, and the three-dimensional shape, posture and spatial position of dense populations can be accurately estimated, improving the accuracy and consistency of reconstruction.
Smart Images

Figure CN2024110506_31072025_PF_FP_ABST
Abstract
Description
Object reconstruction method, device, equipment and storage medium
[0001] This application claims priority to the Chinese invention patent application entitled “Object Reconstruction Method, Apparatus, Device and Storage Medium” filed on December 5, 2023, with application number 202311660086.6, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Example embodiments of the present disclosure relate generally to the field of computer vision, and more particularly, to object reconstruction methods, apparatuses, devices, and computer-readable storage media. Background Art
[0003] In recent years, with the continuous development of the internet and the increasing popularity of wireless communications, human digitization technology has rapidly advanced and is becoming increasingly integrated into people's daily lives. For example, in some scenarios, it is necessary to identify human posture from captured two-dimensional images for subsequent analysis or tracking. Given this background, the use of electronic devices to achieve high-quality human digitization, particularly three-dimensional reconstruction of the human body, has become an increasingly important research direction in computer graphics and vision.
[0004] Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for object reconstruction is provided. The method comprises: determining corresponding reference sizes of multiple objects to be reconstructed in a target image; dividing the target image into multiple first image blocks based on the corresponding reference sizes of the multiple objects, wherein the size of each of the multiple first image blocks is positively correlated with the reference size of at least one object to be reconstructed in the first image block; generating corresponding object reconstruction results for each of the multiple first image blocks using a trained object reconstruction model; and generating a first object reconstruction result of the target image based on the corresponding object reconstruction results of the multiple first image blocks.
[0006] In a second aspect of the present disclosure, an object reconstruction apparatus is provided. The apparatus includes: a determination module configured to determine corresponding reference sizes of multiple objects to be reconstructed in a target image; a division module configured to divide the target image into multiple first image blocks based on the corresponding reference sizes of the multiple objects, wherein the size of each of the multiple first image blocks is positively correlated with the reference size of at least one object to be reconstructed in the first image block; an image block reconstruction module configured to generate corresponding object reconstruction results for each of the multiple first image blocks using a trained object reconstruction model; and a target image reconstruction module configured to generate a first object reconstruction result of the target image based on the corresponding object reconstruction results of the multiple first image blocks.
[0007] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium and can be executed by a processor to perform the method according to the first aspect of the present disclosure.
[0009] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, which, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0010] It should be understood that the contents described in the summary of the present invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0012] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0013] FIG2 shows a schematic block diagram of a process of object reconstruction according to some embodiments of the present disclosure;
[0014] FIG3 is a schematic block diagram illustrating an example process of dividing a first image block according to some embodiments of the present disclosure;
[0015] FIG4 is a schematic block diagram illustrating an example process of determining three-dimensional position information of a reference surface according to some embodiments of the present disclosure;
[0016] FIG5 illustrates a block diagram of an example architecture for object reconstruction according to some embodiments of the present disclosure;
[0017] FIG6 shows a block diagram of a process for obtaining a reference surface based on a sample human body fitting according to some embodiments of the present disclosure;
[0018] FIG7 shows a flowchart of an object reconstruction method according to some embodiments of the present disclosure;
[0019] FIG8 shows a block diagram of an object reconstruction apparatus according to some embodiments of the present disclosure; and
[0020] FIG9 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0022] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit and implicit definitions may be included below.
[0023] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably in this article.
[0024] As briefly mentioned above, 3D reconstruction of the human body has become an increasingly important research area in computer graphics and vision. Currently, some solutions for human geometry reconstruction can be implemented based on a single image. However, for single images taken from large scenes and long distances, recovering the 3D human form and motion of densely packed crowds from such images remains challenging due to the small proportion of the human body in the image, interactions between people, and occlusions. Large scenes here typically refer to ultra-high-resolution images, such as gigapixel images, and long-distance perspectives can include, for example, surveillance camera views or drone aerial photography.
[0025] Currently, a number of methods for 3D reconstruction of the human body have been developed. One type is a multi-step method, which first divides the entire image into blocks and uses a model to perform image processing and 3D human body reconstruction within each image block. However, this type of method only has a single limited receptive field, and the model cannot obtain complete image information of the entire image. This will cause the model to lack 3D reference information of the scene where the image is located, resulting in depth ambiguity in the reconstruction result. The second type is a single-step method, which takes the entire image as input and can show relatively complete relative relationship information of the human body. However, this type of method is usually suitable for small scenes and close-range images, and has problems such as missed detection for large scenes and images from long-distance perspectives.
[0026] In order to at least partially solve the above problems, an embodiment of the present disclosure provides an object reconstruction method. The method includes determining the corresponding reference sizes of multiple objects (e.g., human bodies) to be reconstructed in a target image. Based on the corresponding reference sizes of these objects, the target image is divided into multiple image blocks (hereinafter also described as first image blocks), and the size of each of these image blocks is positively correlated with the reference size of at least one object to be reconstructed in the image block. For example, if the size of the human body contained in some image blocks is larger, then the size of the image block is larger. If the size of the human body contained in some image blocks is smaller, then the size of the image block is smaller. That is, the method of dividing image blocks in the embodiment of the present disclosure is a method of adaptive division based on object size.
[0027] Afterwards, the trained object reconstruction model is used to generate corresponding object reconstruction results (e.g., 3D human body reconstruction results) for each of these image blocks. Based on the corresponding object reconstruction results of these image blocks, an object reconstruction result of the target image (hereinafter also referred to as a first object reconstruction result) is generated.
[0028] When applying the solution of the disclosed embodiments to 3D human body reconstruction from large scenes and long-distance images, by designing image blocks that adapt to human body size and employing a single-step neural network model, the network model's receptive field can be expanded to encompass large scenes. This enables more accurate estimation of the 3D shape, posture, and relative spatial position of densely populated people from a single image in a large scene, leading to more precise 3D human body reconstruction.
[0029] In addition, in some embodiments of the present disclosure, by performing 3D reconstruction on a reference surface (e.g., the ground) in the target image and introducing a geometric prior constraint that the object is located on the reference surface (e.g., a human body standing on the ground), the spatial consistency of the reconstruction result can be further constrained.
[0030] Example environment and basic working principles
[0031] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0032] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. The environment 100 involves a model application device 120, which is configured to reconstruct a target image 110. For example, the target image 110 can be a two-dimensional image of a large scene and a long distance. The target image can be various types of images, such as human images, animal images, vehicle images, etc. Therefore, the object to be reconstructed can be various types of objects in the target image, such as human bodies, animals, vehicles, etc. An object reconstruction model 130 is deployed in the model application device 120. In embodiments of the present disclosure, object reconstruction of the target image 110 is performed with the help of the trained object reconstruction model 130.
[0033] The object reconstruction model 130 can be configured as any appropriate type of model suitable for target detection. As an example, the object reconstruction model 130 can be a single-step network that can directly estimate the spatial position information of each target (which can also be described as an object in the embodiment of the present disclosure) in the input image (e.g., an image block), such as the shape, posture or position of each target. Exemplarily, the object reconstruction model 130 may include a main view position detection network, a bird's-eye view position detection network and a position combination network. The main view position detection network can be used to determine the position information of the object to be reconstructed under the main view perspective (which can be used to indicate the main view perspective of the object to be reconstructed), such as the plane position information of the object to be reconstructed in the target image. The bird's-eye view position detection network can be used to determine the position of the object to be reconstructed in the depth direction under the bird's-eye view perspective (which can be used to indicate the top view perspective of the object to be reconstructed), that is, the depth position information, which can represent the depth information of the object to be reconstructed. The position combination network can be used to combine the plane position information determined by the main view position detection network and the depth position information determined by the bird's-eye view position detection network to obtain the three-dimensional position information of the object to be reconstructed, that is, the spatial position information. The network structures of the primary view position detection network, the bird's-eye view position detection network, and the position combination network can be flexibly constructed or configured based on actual application requirements or application scenarios. For example, the primary view position detection network and the bird's-eye view position detection network can be convolutional neural networks, and the position combination network can be a recurrent neural network. It is understood that this network is an exemplary architecture of the object reconstruction model 130, and the object reconstruction model 130 can also have other network structures. This disclosure is not limited to this.
[0034] The model application device 120 may utilize the deployed object reconstruction model 130 to process the target image 110 to perform object reconstruction on the target image 110. Furthermore, the model application device 120 may generate an object reconstruction result 140 corresponding to the target image 110.
[0035] In some embodiments of the present disclosure, the object reconstruction result 140 may be a three-dimensional object reconstruction result, which may include the relative position relationship of multiple objects, posture information of each object, and three-dimensional shape information, etc.
[0036] In the environment 100, the model training device 125 is configured to train an object reconstruction model 130. The training of the object reconstruction model 130 can be implemented by any appropriate computing system or device, such as a server, a cloud computing device, an edge computing node, etc. In some embodiments, the training of the object reconstruction model 130 can be completed at the model training device 125, and the trained object reconstruction model 130 can be provided to the model application device 120 for use. In some embodiments, the training of the object reconstruction model 130 can be partially or entirely implemented at the model application device 120.
[0037] In environment 100, the model application device 120 or the model training device 125 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device may, for example, include a computing system / server, such as a mainframe, an edge computing node, an electronic device in a cloud environment, and the like. Although shown as separate devices, the model application device 120 and the model training device 125 can be the same device, or included in the same system.
[0038] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0039] Example Embodiment of Object Reconstruction
[0040] 2 , in some embodiments of the present disclosure, a process 200 of performing object reconstruction on a target image 110 using an object reconstruction model 130 includes: partitioning the target image 110 into a plurality of first image blocks 210 using an adaptive partitioning method based on the size of the object to be reconstructed. Then, processing the plurality of first image blocks 210 using the object reconstruction model 130 to obtain object reconstruction results 220 corresponding to the plurality of first image blocks 210. Then, based on the object reconstruction results 220 of the plurality of first image blocks, a first reconstruction result 230 of the target image is obtained.
[0041] For ease of description and distinction, in the embodiments of the present disclosure, the image block obtained by adaptive partitioning is referred to as the first image block; the object reconstruction result obtained based on the first image block is referred to as the first object reconstruction result. The image block obtained by uniform partitioning is referred to as the second image block; the object reconstruction result obtained based on the second image block is referred to as the second object reconstruction result.
[0042] The following first describes the process of dividing the target image 110 into a plurality of first image blocks 210 .
[0043] In the embodiment of the present disclosure, first, corresponding reference sizes of multiple objects to be reconstructed in a target image are determined, and then based on the corresponding reference sizes of these objects, the target image is divided into multiple first image blocks.
[0044] In some embodiments, the target image is a two-dimensional image. With the object to be reconstructed as the target, the contours of multiple objects to be reconstructed can be determined from the target image based on a target detection algorithm for a two-dimensional image, and the size of the corresponding detection frame (which can be regarded as the corresponding reference size of the object) can be determined based on the contours of the multiple objects to be reconstructed. Exemplarily, the size of the detection frame can be the diagonal length of the detection frame. The size of the first image block is further determined based on the size of the detection frame. The size of the first image block can be M times the size of the detection frame, where M is a real number greater than or equal to 1. M can be determined based on empirical values.
[0045] In some embodiments, the target image may be first divided into multiple second image blocks using a uniform partitioning method. A rough object reconstruction result is then obtained by processing the second image blocks using the object reconstruction model. The reference sizes of at least some objects are then determined based on the rough object reconstruction result. The first image blocks are then adaptively partitioned based on the sizes of these objects.
[0046] 3 , the target image 110 is first divided into a plurality of second image blocks 310 having the same size, and then the trained object reconstruction model 130 is used to obtain an object reconstruction result 320 of each second image block.
[0047] The object reconstruction result 320 may include initial three-dimensional information of the object, such as posture information of each human body in the multiple human bodies, three-dimensional shape information of each human body in the multiple human bodies, and relative positional relationships between different human bodies in the multiple human bodies.
[0048] Then, based on the object reconstruction results 320 of the plurality of second image blocks, a second object reconstruction result 330 of the target image is generated. Based on the second object reconstruction result 330, corresponding reference sizes 340 of the plurality of objects to be reconstructed are determined. Furthermore, based on the corresponding reference sizes 340 of the objects to be reconstructed, the target image 110 is adaptively partitioned to obtain a plurality of first image blocks 210.
[0049] For example, the second object reconstruction result 330 is projected onto the two-dimensional target image to obtain contour information of multiple objects, and then the size of the corresponding detection box (which can be regarded as the reference size of the object) is determined based on the contour information of the multiple objects. The size of the first image block is further determined based on the size of the detection box.
[0050] When dividing the first image block based on the second object reconstruction result 330, a "dynamic division" or "dynamic windowing" method can be used. Specifically, for the area to be divided in the target image, the reference object in the area to be divided is determined based on the corresponding reference size of the object to be reconstructed in the area to be divided. Exemplarily, the object to be reconstructed is an object with relatively significant three-dimensional information based on the second object reconstruction result 330, such as the largest object. Then, based on the reference size of the reference object, the image block size is determined; and a sub-area in the area to be divided that has the image block size and includes the reference object is determined as one of the multiple first image blocks.
[0051] In other words, this approach can be understood as not obtaining multiple first image blocks through a single division. Instead, a reference object is determined for the i-th time, and the first image block corresponding to the i-th division is obtained based on the reference object (which can also be described as the i-th sliding window). For the i+1th time, the reference object is determined again for the area where no image blocks have yet been divided, and the image block is further divided again to obtain the first image block corresponding to the i+1th division (which can also be described as the i+1th sliding window). In this way, reference objects are continuously determined, and first image blocks are continuously obtained through division.
[0052] For a three-dimensional object, the position information of the three-dimensional object includes plane position information and depth information. Depth information can be understood as the distance of the three-dimensional object from the shooting lens. The depth values of multiple three-dimensional objects in the same plane may be different, that is, the distances of multiple objects from the lens are different. The plane information of multiple objects with the same depth value may also be different, that is, there may be multiple objects at the same distance from the lens. In the embodiment of the present disclosure, for a two-dimensional target image, its size can be expressed as x*y, then the depth direction can be described as the longitudinal direction, that is, the direction represented by y. The x direction is described as the transverse direction.
[0053] In some embodiments, in the aforementioned "dynamic segmentation" approach, the multiple first image blocks are segmented along the depth direction of the target image (e.g., the y-direction mentioned above). Consequently, the multiple first image blocks obtained have different depth positions. Object reconstruction results obtained based on image blocks with different depth values are more accurate.
[0054] Of course, it is understandable that if the lateral size of the two-dimensional target image is relatively large, the image blocks may be divided along the lateral direction (eg, the x direction mentioned above).
[0055] In some embodiments, multiple first image blocks may be constructed based on the average reference size of multiple objects to be reconstructed. For example, for an area where multiple objects are concentrated, the average reference size of the multiple objects in the area may be used to construct the first image block corresponding to the area. The size of the first image block is proportional to the average reference size of the objects.
[0056] Next, a process of using the object reconstruction model 130 to process the plurality of first image blocks 210 to obtain the object reconstruction results 220 corresponding to the plurality of first image blocks 210 will be described.
[0057] By processing each first image block 210 separately using the object reconstruction model 130, an object reconstruction result 220 corresponding to each first image block 210 can be obtained. The object reconstruction model 130 can be a multi-step network or a single-step network. The specific processing process is not repeated in this disclosure.
[0058] When the object to be reconstructed is located in the middle of the first image block, the resulting object reconstruction is more robust. Therefore, the first image block can be expanded so that the object to be reconstructed is located in the middle of the first image block. This approach is particularly suitable when the first image block is located at the edge of the target image, such as at the bottom of the target image.
[0059] Therefore, in some embodiments, for a given first image block among the plurality of first image blocks, the given first image block is expanded by adding pixels having a preset pixel value in at least one direction, and then the trained object reconstruction model is used to process the expanded given first image block to generate an object reconstruction result for the given first image block.
[0060] The at least one direction here can be the longitudinal axis of the image, i.e., the depth direction. For example, it can be along two opposite depth directions. The preset pixel value can be any suitable pixel value, such as a pixel value representing black or white. For another example, the preset pixel value can be a pixel value that is significantly different from the object to be reconstructed. For example, if the object to be reconstructed is a human body, the pixel values filled in the expanded area can be pixel values that are not of the human body, such as pixel values of the earth.
[0061] Next, a process of generating a first reconstruction result 230 of a target image by using the object reconstruction results 220 corresponding to the plurality of first image blocks 210 will be described.
[0062] In some embodiments, the object reconstruction results 220 corresponding to the first image blocks may be directly spliced together according to the relative relationship between the first image blocks to obtain the first reconstruction result 230 of the target image.
[0063] In some embodiments, to improve the accuracy of the first reconstruction result 230 of the target object, a reference surface (e.g., the ground) in the target image is reconstructed and a geometric prior constraint that the object is located on the reference surface (e.g., a human body standing on the ground) is introduced, thereby further constraining the spatial consistency of the reconstruction result.
[0064] Therefore, under the a priori constraint that a predetermined portion of the object reconstructed in each first image block contacts the reference surface, the three-dimensional position information of the reference surface in the target image can also be determined. Based on the three-dimensional position information of the reference surface, the corresponding object reconstruction results of the multiple first image blocks are updated. Based on the relative positions of the multiple first image blocks in the target image, the updated corresponding object reconstruction results are combined into a first object reconstruction result.
[0065] For example, based on the prior condition that a person's footsteps will contact the ground when standing on it, the 3D position information of the ground as a reference surface can be determined. The depth values in the object reconstruction results of the person in each first image block are then adjusted based on the 3D position information of the ground.
[0066] In other examples, the reference surface can also be a desktop, a water surface, etc. For example, if the target image is an image of some animals resting on the water surface, and the objects to be reconstructed are these animals, the reference surface can be the water surface. The embodiments of this disclosure do not limit the form of the reference surface.
[0067] The process of obtaining the three-dimensional position information of the reference surface is given below.
[0068] Referring to FIG4 , process 400 includes obtaining a sample image 410 that includes a reference surface. The sample image may have the same field of view as the target image with respect to the reference surface. For example, when the camera lens is fixed, a set of image sequences (which may be a video) captured have the same field of view. The sample image can then be determined from this set of image sequences with the same field of view. For example, multiple images of the object to be reconstructed (e.g., a human body) evenly distributed in the depth direction are selected as sample images.
[0069] Then, a plurality of sample objects are selected from the plurality of reconstructed objects indicated by the object reconstruction result of the sample image.
[0070] Specifically, the sample object 430 is determined from the high-confidence reconstruction result 420 obtained by processing the sample image 410 using the object reconstruction model 130 .
[0071] In some embodiments, the object reconstruction results of the sample images include corresponding reconstruction confidence levels for multiple reconstructed objects, indicating the accuracy or reliability of the object reconstruction results. Therefore, reconstructed objects with corresponding reconstruction confidence levels above a preset threshold can be selected from multiple regions of different depths in the sample image as sample objects. In other words, sample objects with relatively accurate reconstruction results are selected, and the selected sample images have different depths.
[0072] Next, based on the object reconstruction results of the sample images, the three-dimensional position information 440 of the corresponding predetermined parts of these sample objects 430 is determined. The three-dimensional position information of the corresponding predetermined parts of these sample objects is then fitted into the three-dimensional position information 450 of the reference surface. For example, if the sample images are human images and the sample objects are human bodies, multiple sample human bodies can be selected, the three-dimensional position information of their footsteps determined, and then the three-dimensional position information of the ground surface can be obtained by fitting.
[0073] After the three-dimensional position information of the reference surface is determined in the above manner, the object reconstruction result can be further optimized and adjusted.
[0074] In the disclosed embodiments, object reconstruction model 130 can be trained using any suitable method. The following describes an example training process. Object reconstruction model 130 generates an object reconstruction result for a training image based on the training image. Based on the object reconstruction result for the training image, spatial information about the object in the training image and the plane information projected onto the training image are determined. Object reconstruction model 130 is trained using this spatial information, plane information, and labeling information for the training image.
[0075] Specifically, during the network training phase, the object reconstruction model is jointly optimized by combining the 3D information from the object reconstruction results of the training image, the 2D projection information projected into the training image, and the label information in the training image, such as the position of each object in the training image and the distance between the training image and the ground. In other words, the model is trained by combining the object's 3D pose and shape, 3D spatial position, the object's 2D pose and 2D position projected into the training image, and the distance between the 2D projection and the ground plane. The loss function, for example, can be the L2-loss between the true value and the predicted value of each metric.
[0076] The above describes some embodiments of the object reconstruction solution according to the present disclosure. Next, for a clearer understanding of the present disclosure, an example implementation of the object reconstruction method of the present disclosure is introduced by taking the reconstruction of a three-dimensional human body model in a large scene high-resolution image as an example.
[0077] First, we introduce an example of a single-step multi-person 3D human reconstruction network used to reconstruct a 3D human model. The algorithm first uses a backbone network to extract image features. These features are then fed into multiple network branches, which are used to obtain a rough front-view human center position map, a bird's-eye view human center position map, a precise front-view human center position offset vector map, a precise bird's-eye view human center position offset vector map, and a human surface feature map. By combining the front-view and bird's-eye view information, we obtain a coarse 3D human center position probability map and a precise 3D human center position offset vector map. Combining these two information sets yields the 3D positions of all persons in the image. By simultaneously learning the front-view and bird's-eye view images, we explicitly model the positions of the human body in the image plane and depth. Then, using a normalized camera representation based on weak-perspective projection, the predicted 3D position is transformed from this 3D latent space into camera-centric 3D space, obtaining the precise 3D position of the human body in camera space. Finally, based on the estimated spatial position, samples are taken at the corresponding positions of the human body surface feature map to obtain the feature vector of each person in the map. This feature vector is combined with the corresponding depth position code to regress each person's 3D body parameters, thereby obtaining each person's 3D shape, posture and relative spatial position.
[0078] An example process 500 for 3D human body reconstruction is described with reference to FIG5 . In step 1 of process 500 , a large-scene, high-resolution image 510 is processed by a uniform partitioning module to produce a plurality of uniformly partitioned image blocks 511 . Image blocks 511 are then processed using a single-step, multi-person 3D human body reconstruction network to produce a coarse reconstruction result 512 .
[0079] This single-step multi-person 3D human reconstruction method cannot directly process large, high-resolution images. Conventional uniform blocking methods often result in distant objects being too small to be recognized, while nearby objects are too large to fully visualize the human body within the window. Therefore, some embodiments of the present disclosure utilize a human-size-aware dynamic windowing module. Unlike the uniform blocking method, this dynamic windowing module adjusts the sliding window size based on the proportion of human body size within a local image block.
[0080] In step 2 of process 500 , the image 510 is divided again based on the dynamic windowing module to obtain human-size-aware dynamic windowing blocks 513 .
[0081] Exemplarily, the reconstructed three-dimensional human body model is projected two-dimensionally onto the image 510, a rectangular human body detection frame is obtained through the human body contour, and the diagonal length of the detection frame is used as the size of the detection target (the human body in this example).
[0082] In addition, when designing the dynamic window module, the human body closest to the lens on the vertical axis of the image is selected as the reference target of the i-th sliding window, and the size of the i-th sliding window is proportional to the size of the reference target.
[0083] In addition, considering that the reconstruction effect is more robust when the target human body is located in the middle of the image, the image blocks on the vertical axis (that is, the depth direction) can be expanded so that the model can accurately reconstruct the human body at the bottom of the image.
[0084] Through the dynamic windowing process in the above steps, the large scene high-resolution image 510 is adaptively divided into image blocks of dynamic sizes that are perceived based on human size, and the depth positions of these image blocks are different.
[0085] In step 3 of process 500, the previously described multi-person 3D human reconstruction method is again used to process each image block, resulting in a precise reconstruction 514 of the 3D human figures within the image blocks at different depths. Compared to the coarse reconstruction 512, a more accurate reconstruction is achieved because the image block size adapts to the size of the human figures in image 510.
[0086] Alternatively or additionally, to obtain a 3D human reconstruction result for the entire crowd image, appropriate constraints need to be added to the precise reconstruction results 514 within each image block obtained in step 3 above to obtain the accurate 3D position of each reconstructed object. It is assumed here that most observed objects (people) are in contact with the ground. Therefore, this disclosure also designs a 3D ground plane reconstruction module, which uses this module to estimate the 3D ground plane and place 3D people on the ground, adding geometric prior constraints to the 3D positions of the crowd in the image.
[0087] Alternatively or additionally, the 3D ground plane can be fitted to the true 3D human position by sampling sufficient reconstructed samples. Typically, gigapixel cameras used to capture large scenes are stationary during video capture. Therefore, for a sequence of images of the same scene, the 3D ground plane only needs to be estimated once. Therefore, process 500 further includes the following step 4.
[0088] In step 4 of process 500 , for a video / image sequence of the same scene, a representative large-scene high-resolution image 520 , such as an image with a relatively uniform distribution in the depth direction, is selected and the image 520 is also uniformly divided to obtain a plurality of uniformly divided image blocks 521 .
[0089] In step 5 of process 500, sample objects with high detection confidence are selected within each image block 521. Experiments have found that not all 3D position samples contribute to an accurate 3D ground plane, and some low-quality predictions can compromise the reconstruction results. Therefore, selecting sample objects in step 5 ensures a uniform distribution of high-quality samples across depth.
[0090] In step 6 of process 500 , the feet of these sampled samples should be in contact with the ground plane, and the reconstruction results of the 3D ground plane are obtained by fitting the 3D positions of the foot joints of the sampled samples.
[0091] As an example, referring to FIG6 , a process 600 for reconstructing a 3D ground plane using sample fitting is shown. A sample object (a human body selected by a black frame) with a high-confidence reconstruction result is selected from sample image 600 , and the fitting process indicated by 620 is performed to obtain a 3D reconstruction result of the ground plane.
[0092] 5 , in step 7 of process 500 , after obtaining the 3D ground plane, the depth values of all objects (human bodies) reconstructed in the precise reconstruction result 514 are adjusted so that the human bodies stand on the ground, thereby obtaining a 3D human body reconstruction result 540 for all persons in the entire image.
[0093] Continuing with Figure 5, as an example, during the training phase, the multi-person 3D human reconstruction network can be jointly optimized 530 by combining the 3D information from the precise reconstruction results 514 of the large-scale, high-resolution image used as the training image with the 2D projections projected onto the training image. Specifically, the model is trained by combining the 3D pose and shape of the human body, its 3D spatial position, its 2D pose and 2D position projected into the image, and the distance between the 2D projection and the ground plane. The loss function can, for example, be the L2-loss between the true and predicted values of each metric. Thus, when the model is applied for prediction, a single RGB image can be input to obtain the 3D pose, shape, and 3D relative position of each person in the image.
[0094] This paper proposes a 3D human body reconstruction method for dense crowds in large scenes. It designs a human-size-aware dynamic windowing module and a 3D ground plane reconstruction module. The former automatically cuts large scene images into appropriate image windows and inputs them into a single-step network for accurate 3D human body reconstruction of dense crowds. The latter eliminates the relative position ambiguity of the reconstruction results of multiple people in large scenes by adding ground plane geometric constraints.
[0095] Example Process
[0096] 7 shows a flow chart of a process 700 for object reconstruction according to some embodiments of the present disclosure. The process 700 may be implemented at the model application device 120.
[0097] At block 710 , the model application apparatus 120 determines respective reference sizes of a plurality of objects to be reconstructed in a target image.
[0098] At block 720 , the model application device 120 divides the target image into a plurality of first image blocks based on the respective reference sizes of the plurality of objects, wherein the size of each of the first image blocks is positively correlated with the reference size of at least one object to be reconstructed in the first image block.
[0099] In block 730 , the model application device 120 generates corresponding object reconstruction results for the plurality of first image blocks respectively using the trained object reconstruction model.
[0100] In block 740 , the model application device 120 generates a first object reconstruction result of the target image based on the corresponding object reconstruction results of the first image blocks.
[0101] In some embodiments, determining the corresponding reference sizes of multiple objects to be reconstructed in a target image includes: dividing the target image into multiple second image blocks of the same size; generating a second object reconstruction result of the target image based on the multiple second image blocks using a trained object reconstruction model; determining corresponding contour information of the multiple objects based on the second object reconstruction result; and determining the corresponding reference sizes of the multiple objects based on the corresponding contour information of the multiple objects.
[0102] In some embodiments, dividing the target image into multiple first image blocks includes: determining, for the area to be divided in the target image, a reference object in the area to be divided based on a corresponding reference size of an object to be reconstructed in the area to be divided; determining an image block size based on a reference size of the reference object; and determining a sub-area in the area to be divided that has the image block size and includes the reference object as one of the multiple first image blocks.
[0103] In some embodiments, the reference object is the object to be reconstructed with the largest reference size in the area to be divided.
[0104] In some embodiments, the plurality of first image blocks are divided along a depth direction of the target image.
[0105] In some embodiments, respectively generating corresponding object reconstruction results for multiple first image blocks includes: for a given first image block among the multiple first image blocks, expanding the given first image block by adding pixels with preset pixel values in at least one direction; and generating an object reconstruction result for the given first image block by processing the expanded given first image block using a trained object reconstruction model.
[0106] In some embodiments, process 700 also includes determining three-dimensional position information of a reference surface in the target image; and generating a first object reconstruction result of the target image includes: updating corresponding object reconstruction results of multiple first image blocks based on the three-dimensional position information of the reference surface, wherein a preset part of the reconstructed object in the updated object reconstruction result of each first image block contacts the reference surface; and combining the updated corresponding object reconstruction results into a first object reconstruction result based on the relative positions of the multiple first image blocks in the target image.
[0107] In some embodiments, determining the three-dimensional position information of a reference surface in a target image includes: acquiring a sample image including the reference surface, the sample image and the target image having the same field of view with respect to the reference surface; selecting a plurality of sample objects from a plurality of reconstructed objects indicated by an object reconstruction result of the sample image; determining the three-dimensional position information of the corresponding preset parts of the plurality of sample objects based on the object reconstruction result of the sample image; and fitting the three-dimensional position information of the corresponding preset parts of the plurality of sample objects into the three-dimensional position information of the reference surface.
[0108] In some embodiments, the object reconstruction result of the sample image includes corresponding reconstruction confidences of multiple reconstructed objects, and selecting multiple sample objects from the multiple reconstructed objects indicated by the object reconstruction result of the sample image includes: selecting corresponding reconstructed objects with reconstruction confidences higher than a preset threshold from multiple areas with different depths in the sample image as the sample objects.
[0109] In some embodiments, the object reconstruction model is trained in the following manner: using the object reconstruction model, based on the training image, an object reconstruction result of the training image is generated; based on the object reconstruction result of the training image, the spatial information of the object in the training image and the plane information projected to the training image are determined; and using the spatial information, plane information and labeling information for the training image to train the object reconstruction model.
[0110] Example device
[0111] 8 shows a block diagram of an object reconstruction apparatus 800 according to some embodiments of the present disclosure. The apparatus 800 may be implemented in the model application device 120. Each module / component in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0112] The apparatus 800 includes a determination module 810 configured to determine the corresponding reference sizes of multiple objects to be reconstructed in a target image. The apparatus 800 also includes a division module 820 configured to divide the target image into multiple first image blocks based on the corresponding reference sizes of the multiple objects, wherein the size of each first image block in the multiple first image blocks is positively correlated with the reference size of at least one object to be reconstructed in the first image block. The apparatus 800 also includes an image block reconstruction module 830 configured to generate corresponding object reconstruction results for the multiple first image blocks using a trained object reconstruction model; and a target image reconstruction module 840 configured to generate a first object reconstruction result of the target image based on the corresponding object reconstruction results of the multiple first image blocks.
[0113] In some embodiments, the determination module 810 is further configured to: divide the target image into multiple second image blocks of the same size; generate a second object reconstruction result of the target image based on the multiple second image blocks using a trained object reconstruction model; determine corresponding contour information of the multiple objects based on the second object reconstruction result; and determine corresponding reference sizes of the multiple objects based on the corresponding contour information of the multiple objects.
[0114] In some embodiments, the division module 820 is further configured to: determine, for the area to be divided in the target image, a reference object in the area to be divided based on the corresponding reference size of the object to be reconstructed in the area to be divided; determine the image block size based on the reference size of the reference object; and determine a sub-area in the area to be divided having the image block size and including the reference object as one of multiple first image blocks.
[0115] In some embodiments, the reference object is the object to be reconstructed with the largest reference size in the area to be divided.
[0116] In some embodiments, the plurality of first image blocks are divided along a depth direction of the target image.
[0117] In some embodiments, the image block reconstruction module 830 is further configured to: for a given first image block among multiple first image blocks, expand the given first image block by adding pixels with preset pixel values in at least one direction; and generate an object reconstruction result for the given first image block by processing the expanded given first image block using a trained object reconstruction model.
[0118] In some embodiments, the determination module 810 is further configured to determine the three-dimensional position information of the reference surface in the target image; and generating a first object reconstruction result of the target image includes: updating the corresponding object reconstruction results of multiple first image blocks based on the three-dimensional position information of the reference surface, wherein a preset part of the reconstructed object in the updated object reconstruction result of each first image block contacts the reference surface; and combining the updated corresponding object reconstruction results into a first object reconstruction result based on the relative positions of the multiple first image blocks in the target image.
[0119] In some embodiments, the determination module 810 is further configured to obtain a sample image including a reference surface, the sample image and the target image having the same field of view with respect to the reference surface; select multiple sample objects from multiple reconstructed objects indicated by the object reconstruction result of the sample image; determine the three-dimensional position information of the corresponding preset parts of the multiple sample objects based on the object reconstruction result of the sample image; and fit the three-dimensional position information of the corresponding preset parts of the multiple sample objects into the three-dimensional position information of the reference surface.
[0120] In some embodiments, the object reconstruction result of the sample image includes corresponding reconstruction confidences of multiple reconstructed objects, and the determination module 810 is further configured to select reconstructed objects with corresponding reconstruction confidences higher than a preset threshold from multiple areas with different depths in the sample image as the sample objects.
[0121] In some embodiments, the object reconstruction model is trained in the following manner: using the object reconstruction model, based on the training image, an object reconstruction result of the training image is generated; based on the object reconstruction result of the training image, the spatial information of the object in the training image and the plane information projected to the training image are determined; and using the spatial information, plane information and labeling information for the training image to train the object reconstruction model.
[0122] The units included in the device 800 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 800 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0123] FIG9 shows a block diagram of an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 900 shown in FIG7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 900 shown in FIG9 can be used to implement the model application device 120 or the model training device 125 of FIG1 .
[0124] As shown in FIG9 , electronic device 900 is a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, a storage unit 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a real or virtual processor and is capable of performing various processes according to programs stored in storage unit 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 900.
[0125] The electronic device 900 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. The storage unit 920 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 930 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 900.
[0126] The electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The storage unit 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0127] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 900 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 900 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0128] The input device 950 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 960 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 900 may also communicate with one or more external devices (not shown) via the communication unit 940 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 900, or with any device that allows the electronic device 900 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0129] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, wherein the one or more computer instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0130] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0131] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0132] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0133] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0134] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. An object reconstruction method, comprising: Determining corresponding reference sizes of a plurality of objects to be reconstructed in a target image; Based on the corresponding reference sizes of the plurality of objects, dividing the target image into a plurality of first image blocks, wherein the size of each first image block in the plurality of first image blocks is positively correlated with the reference size of at least one object to be reconstructed in the first image block; Using a trained object reconstruction model to respectively generate corresponding object reconstruction results of the plurality of first image blocks; And Based on the corresponding object reconstruction results of the plurality of first image blocks, generating a first object reconstruction result of the target image.
2. The method according to claim 1, wherein determining the corresponding reference sizes of the plurality of objects to be reconstructed in the target image comprises: Dividing the target image into a plurality of second image blocks having the same size; Using the trained object reconstruction model to generate a second object reconstruction result of the target image based on the plurality of second image blocks; Based on the second object reconstruction result, determining corresponding contour information of the plurality of objects; And Based on the corresponding contour information of the plurality of objects, determining the corresponding reference sizes of the plurality of objects.
3. The method according to claim 1, wherein dividing the target image into a plurality of first image blocks comprises: For a region to be divided in the target image, based on the corresponding reference size of the object to be reconstructed in the region to be divided, determining a reference object in the region to be divided; Based on the reference size of the reference object, determining an image block size; And Determining a sub-region in the region to be divided that has the image block size and includes the reference object as one of the plurality of first image blocks.
4. The method according to claim 3, wherein the reference object is the object to be reconstructed with the largest reference size in the region to be divided.
5. The method according to claim 1, wherein the plurality of first image blocks are divided along the depth direction of the target image.
6. The method according to any one of claims 1-5, wherein respectively generating the corresponding object reconstruction results of the plurality of first image blocks comprises: For a given first image block among the plurality of first image blocks, expanding the given first image block by adding pixels with a preset pixel value in at least one direction; And By using the trained object reconstruction model to process the expanded given first image block, generating an object reconstruction result of the given first image block.
7. The method according to claim 1, further comprising: 8. The method according to claim 7, wherein determining the three-dimensional position information of the reference plane in the target image comprises: Obtaining a sample image containing the reference plane, the sample image having the same field of view as the target image with respect to the reference plane; Selecting a plurality of sample objects from among the plurality of reconstructed objects indicated by the object reconstruction result of the sample image; Based on the object reconstruction result of the sample image, determining the three-dimensional position information of the corresponding preset parts of the plurality of sample objects; And Fitting the three-dimensional position information of the corresponding preset parts of the plurality of sample objects into the three-dimensional position information of the reference plane.
9. The method according to claim 8, wherein the object reconstruction result of the sample image includes the corresponding reconstruction confidence levels of the plurality of reconstructed objects, and selecting a plurality of sample objects from among the plurality of reconstructed objects indicated by the object reconstruction result of the sample image comprises: Selecting, from within a plurality of regions having different depths in the sample image, the reconstructed objects with corresponding reconstruction confidence levels higher than a preset threshold as the sample objects.
10. The method according to claim 1, wherein the object reconstruction model is trained in the following manner: Using the object reconstruction model, generating an object reconstruction result of the training image based on the training image; Based on the object reconstruction result of the training image, determining the spatial information of the objects in the training image and the planar information projected onto the training image; and Using the spatial information, the planar information, and the labeling information for the training image, training the object reconstruction model.
11. An object reconstruction device, comprising: A determination module configured to determine the corresponding reference sizes of a plurality of objects to be reconstructed in a target image; A division module configured to divide the target image into a plurality of first image blocks based on the corresponding reference sizes of the plurality of objects, the size of each first image block in the plurality of first image blocks being positively correlated with the reference size of at least one object to be reconstructed in the first image block; An image block reconstruction module configured to respectively generate the corresponding object reconstruction results of the plurality of first image blocks by using a trained object reconstruction model; And A target image reconstruction module configured to generate a first object reconstruction result of the target image based on the corresponding object reconstruction results of the plurality of first image blocks.
12. An electronic device, comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.