Method and apparatus for constructing a three-dimensional (3D) human body model
Patent Information
- Application Number
- CN202210918430.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-28
- Filing Date
- 2022-08-01
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-08-01
AI Technical Summary
然而,实际上,患者的 身体可能至少部分地被医疗设备和/或衣物(例如,病号服、覆盖片等)遮盖,因此基于单个视图的HMR系统和方法可能不能产生令人满意的结果
Smart Images

Figure CN115272581B_ABST
Abstract
Description
Technical Field
[0001] This application pertains to the field of 3D modeling. Background Technology
[0002] Three-dimensional (3D) patient models that realistically reflect a patient's body shape and posture can be used in a variety of medical applications, including patient localization, treatment planning, and surgical navigation. For example, in radiotherapy and medical imaging, success often depends on the ability to position and maintain the patient in the desired pose so that procedures can be performed with better accuracy and faster speed. Therefore, knowledge about the patient's physical characteristics in these situations (e.g., visual knowledge) is crucial for obtaining optimal patient outcomes. Conventional human body mesh reconstruction (HMR) systems and methods utilize a single image of the patient and expect the patient's body to be visible (e.g., unoccluded) in these images. However, in reality, the patient's body may be at least partially obscured by medical devices and / or clothing (e.g., hospital gowns, covers, etc.), so single-view-based HMR systems and methods may not produce satisfactory results. Therefore, new and / or improved patient modeling and mesh construction techniques that can accurately estimate patient body information even when one or more parts of the patient's body are obscured are highly desirable. Summary of the Invention
[0003] This document describes systems, methods, and apparatuses associated with constructing multi-view patient models (e.g., 3D human body mesh models) based on multiple single-view models of the patient. The apparatuses described herein may include one or more processors configured to acquire first information associated with a first 3D model of the human body (e.g., a first single-view model) and second information associated with a second 3D model of the human body (e.g., a second single-view model). The first 3D model may be determined based on at least a first image of the human body captured by a first sensing device (e.g., a first camera). The second 3D model may be determined based on at least a second image of the human body captured by a second sensing device (e.g., a second camera). The first information may instruct the first 3D model to cover (e.g., depict with high accuracy) at least a first body key point of the human body (e.g., a first portion such as the left shoulder), and the second information may instruct the second 3D model to cover at least a second body key point of the human body (e.g., a second portion such as the right shoulder). One or more processors of the device may also be configured to: determine a first region (e.g., a first cube) of a first 3D model associated with a first body keypoint and a second region (e.g., a second cube) of a second 3D model associated with a second body keypoint, and generate a third 3D model of the human body (e.g., a multi-view model) based at least on the first region of the first 3D model and the second region of the second 3D model. Figure 3 (D model).
[0004] In the example, one or more processors can be configured to: determine a first set of vertices associated with a first region of the first 3D model based on a first 3D model, and determine a second set of vertices associated with a second region of the second 3D model based on a second 3D model, wherein a third 3D model of the human body is generated based at least on the first and second sets of vertices. In the example, one or more processors can also be configured to: determine the corresponding positions of the first and second sets of vertices in a coordinate system associated with the third 3D model, and generate the third 3D model based at least on the corresponding positions of the first and second sets of vertices in the coordinate system associated with the third 3D model. The corresponding positions of the first and second sets of vertices in the coordinate system can be determined, for example, based on the corresponding positions of the first and second sensing devices in the coordinate system.
[0005] In the example, one or more processors may also be configured to: determine a second body keypoint of the human body covered by both the first 3D model and the second 3D model based on the first information and the second information, and determine that the second 3D model will be used to obtain information associated with the second body keypoint based on a confidence indication associated with at least one of the first 3D model or the second 3D model regarding the second body keypoint. The confidence indication may be determined, for example, based on the position of the first sensing device relative to the second body keypoint and the position of the second sensing device relative to the second body keypoint.
[0006] In the example, the first 3D model and the second 3D model described herein can be constructed by a first sensing device and a second sensing device, respectively, and the device can be configured to receive the first 3D model from the first sensing device and the second 3D model from the second sensing device. In the example, the device may include the first sensing device or the second sensing device, and can be configured to construct either the first 3D model or the second 3D model in addition to constructing the third 3D model. Attached Figure Description
[0007] The examples disclosed herein can be understood in more detail from the following description, which is given by way of example in conjunction with the accompanying drawings.
[0008] Figure 1 This is a simplified block diagram illustrating an example environment associated with one or more embodiments described herein.
[0009] Figure 2 This is a simplified block diagram illustrating an example of generating a multi-view human body model based on multiple single-view models.
[0010] Figure 3 This is a simplified diagram illustrating the construction of a multi-view human body model based on multiple single-view human body models.
[0011] Figure 4 This is a flowchart illustrating the operations that can be associated with the construction of a multi-view human body model.
[0012] Figure 5 This is a block diagram illustrating an example device that can be configured to perform the single-view and / or multi-view model building operations described herein. Detailed Implementation
[0013] The present disclosure is illustrated by way of example rather than limitation in the figures.
[0014] Figure 1 This is a diagram illustrating an example environment 100 associated with one or more embodiments described herein. Environment 100 may be part of a medical facility such as a scanning room (e.g., magnetic resonance imaging (MRI), X-ray, computed tomography (CT), etc.) or operating room (OR), rehabilitation facility, fitness center, etc. Environment 100 may be equipped with one or more sensing devices (e.g., 102a, 102b, 102c), such as one or more digital cameras, configured to capture images (e.g., two-dimensional (2D) images) of patient 104 within environment 100. Sensing devices 102a-c may be communicatively coupled to processing unit 106 and / or other devices of environment 100 via communication network 108. Each sensing device 102a-c may include one or more sensors, such as one or more 2D vision sensors (e.g., 2D cameras), one or more 3D vision sensors (e.g., 3D cameras), one or more red, green and blue (RGB) sensors, one or more depth sensors, one or more RGB plus depth (RGB-D) sensors, one or more thermal sensors (e.g., infrared (FIR) or near-infrared (NIR) sensors), one or more radar sensors and / or other types of image capturing devices or circuitry.
[0015] When inside environment 100, the patient 104's body may be partially obstructed (e.g., partially obstructed) by other equipment or objects (e.g., the C-arm of an X-ray machine, a smear, etc.) and thus cannot enter the field of view (FOV) of the sensing device. For example, depending on the installation location of sensing devices 102a-c, the patient 104's pose and / or posture, and / or the location of other equipment or objects in environment 100, the FOV of the first sensing device (e.g., 102a) may include a first set of regions or areas of the patient's body (e.g., left shoulder, left arm, etc.), and the FOV of the second sensing device (e.g., 102b) may include a second set of regions or areas of the patient's body (e.g., right shoulder, right arm, etc.). Therefore, the first set of regions or areas is visible in the image captured by the first sensing device but not in the image captured by the second sensing device. Similarly, the second set of regions or areas is visible in the image captured by the second sensing device but not in the image captured by the first sensing device.
[0016] Each of the sensing devices 102a-c may include a functional unit (e.g., a processor) configured to process images captured by the sensing device and generate (e.g., construct) a human body model, such as a 3D human body mesh model of a patient, based on the images. Such a human body model may be referred to herein as a single-view model (e.g., because the model may be based on the field of view of one sensing device) and may include multiple parameters indicating the patient's body shape and / or posture during a medical procedure (e.g., an MRI, X-ray, or CT procedure). For example, the parameters of the human body model may be used to determine multiple vertices associated with the patient's body (e.g., 6890 vertices based on 82 body shape and posture parameters), and each vertex may include its own position, normal, texture, and / or shading information. Using these vertices, a 3D mesh of the patient may be created, for example, by connecting multiple vertices to edges to form polygons (e.g., triangles), connecting multiple polygons to form surfaces, using multiple surfaces to determine a 3D shape, and applying texture and / or shading to the surfaces and / or shape.
[0017] The single-view model described herein can also be generated by processing unit 106. For example, processing unit 106 can be communicatively coupled to sensing devices 102a-c and can receive images of the patient captured by the sensing devices (e.g., in real time or based on a predetermined schedule). Using the corresponding images received from each sensing device, processing unit 106 can construct a corresponding single-view model for the patient that reflects the field of view of the sensing devices. Similarly, the single-view model described herein can also be generated by one sensing device based on images captured by another sensing device. For example, sensing devices 102a-c can be interconnected via communication link 108 and exchange images with each other. One sensing device can be configured to perform processing tasks of other sensing devices (e.g., constructing the single-view model described herein) based on images received from other sensing devices.
[0018] Any of the sensing devices 102a-c and / or processing unit 106 may also be configured to generate (e.g., construct) a multi-view model (e.g., multi-view model) of the patient based on multiple single-view models (e.g., the single-view models described herein) constructed for the patient. Figure 3 (D-type human body mesh model). As described herein, each of these single-view models may correspond to a corresponding sensing device and may represent the patient's body shape and / or posture as observed from the field of view of the sensing device. Since the field of view of one sensing device may differ from that of another, single-view models associated with different sensing devices may have varying degrees of accuracy regarding different regions or areas of the patient's body. For example, a first single-view model of a patient associated with a first sensing device may have higher accuracy regarding the left side of the patient's body than regarding the right side, because the right side of the patient's body is occluded or obscured in the FOV of the first sensing device. On the other hand, a second single-view model of a patient associated with a second sensing device may have higher accuracy regarding the right side of the patient's body than regarding the left side, because the left side of the patient's body is occluded or obscured in the FOV of the second sensing device. Therefore, a multi-view model of a patient can be derived based on multiple (e.g., two or more) single-view models. In the above scenario, for example, a multi-view model can be derived by combining a portion of a first single-view model (e.g., corresponding to the left side of the patient's body) with a portion of a second single-view model (e.g., corresponding to the right side of the patient's body), such that the derived multi-view model has high accuracy about both sides of the patient's body. The techniques used to construct such a multi-view model will be described in more detail below, and are not limited to two single-view models (e.g., the description associated with two single-view models can also be applied to three or more single-view models).
[0019] The models (e.g., single-view and / or multi-view models) generated by sensing devices 102a-c and / or processing unit 106 can be used in various clinical applications, including, for example, patient positioning (e.g., during scanning or treatment), medical device control (e.g., automatically adjusting the height of a scanner bed or operating table based on the patient's body shape), and unified medical record review (e.g., aligning scans from different imaging modalities onto the same patient model). The models can be provided to downstream applications or devices in real time, or they can be saved to a storage library (e.g., database 110) communicatively coupled to sensing devices 102a-c and / or processing unit 106 for later use.
[0020] Figure 2 A diagram illustrates an example of generating a multi-view human body model 206 based on multiple single-view models (e.g., 204a, 204b, 204c, etc.). In this example, single-view models 204a-c are shown as being generated by sensing devices 202a, 202b, 202c (e.g., ...). Figure 1 The single-view models 204a-c shown are generated by the sensing devices 102a-c. However, it should be noted that one or more (e.g., all) of the single-view models 204a-c can also be generated by other devices (such as the processing unit 106 or any of the sensing devices 202a-c described above), and therefore the following description is also applicable to these other devices.
[0021] Using sensing device 202a as an example, the sensing device can be configured to capture images (e.g., 2D images) of a patient that reflect the field of view of the sensing device. Sensing device 202a can also be configured to generate a human body model 204a (e.g., a 3D human mesh model) based on one or more images captured by the sensing device. Sensing device 202a can perform this task, for example, using a pre-trained artificial neural network (ANN) (such as an autoencoder). The ANN can be trained to extract features from the images captured by sensing device 202a and infer (e.g., estimate or predict) parameters from the extracted features to recover the human body model. The inferred parameters may include, for example, one or more pose parameters Θ and one or more body shape parameters β, which may respectively indicate the patient's pose and body shape. The ANN may include multiple layers, such as an input layer, one or more convolutional layers, one or more pooling layers, one or more fully connected layers, and / or an output layer. Each layer may correspond to multiple filters (e.g., kernels), and each filter may be designed to detect (e.g., learn) corresponding features or patterns present in the input image. Filters can be associated with corresponding weights, which, when applied to the input, produce an output indicating whether certain visual features or patterns have been detected. The weights associated with the filters can be learned by the ANN through a training process that may include: inputting a large number of images from one or more training datasets into the ANN, computing the difference or loss associated with the current prediction or estimate (e.g., based on an objective function such as mean squared error or L1 norm, a loss function based on residuals, etc.), and updating the weights assigned to the filters to minimize the difference or loss (e.g., stochastic gradient descent based on the loss function). Once trained, the ANN can take images at the input layer, extract visual features or patterns from the images and / or classify them, and provide an indication at the output layer of the presence of the identified features or patterns. The identified features can be indicated, for example, by feature maps or feature vectors.
[0022] An ANN can also be trained to infer pose and body shape parameters for reconstructing a 3D human model, for example, based on features extracted from input images. For instance, an ANN can be trained to determine joint angles of a patient, such as those depicted in patient images, based on a training dataset covering a wide range of human subjects, human activity, background noise, body shape and / or pose variations, camera motion, etc. Multiple joints can include, for example, the 23 joints included in a skeletal apparatus, plus the root joint, and the pose parameters derived from them can include 72 parameters (e.g., 3 parameters for each of the 23 joints and 3 parameters for the root joint, where each parameter corresponds to an axial rotation starting from the root orientation). An ANN can learn to determine one or more body shape parameters based on the training dataset, which are used to predict a patient's mixed body shape based on patient images. For example, an ANN can learn to determine body shape parameters by performing principal component analysis (PCA) on the images, and the determined body shape parameters can include multiple coefficients in the PCA space (e.g., the first 10 coefficients). Once the pose and body shape parameters are determined, multiple vertices (e.g., 6890 vertices based on 82 body shape and pose parameters) can be obtained to construct a model of the patient's body (e.g., a 3D mesh). Each vertex can include its own position, normal, texture, and / or shading information. Using these vertices, the 3D mesh of the patient can be created, for example, by connecting multiple vertices to edges to form polygons (e.g., triangles), connecting multiple polygons to form surfaces, using multiple surfaces to determine the 3D shape, and applying textures and / or shading to the surfaces and / or shapes.
[0023] As part of the aforementioned model building process, sensing device 202a can determine multiple body key points (e.g., 2D anatomical body key points) of the patient's body, which are visible in images captured by the sensing device and thus covered (e.g., depicted) by the human body model 204a generated therefrom (e.g., with high accuracy and / or confidence). Body key points (e.g., anatomical body key points) may correspond to different regions or areas of the patient's body and may be determined based on features extracted from the images. Body key points captured by the sensing device and / or covered by the human body model 204a (e.g., with high accuracy and / or confidence levels) may also be determined based on the spatial relationship between sensing device 202a, the patient, and / or other objects near the sensing device or the patient. For example, body key points captured by the sensing device (e.g., within the FOV of the sensing device) may be determined based on the installation location and / or angle of the sensing device 202a and / or the location, height, and / or movement trajectory of one or more devices obstructing the FOV of the sensing device for the patient. This spatial information can be predetermined (e.g., during device installation or calibration) and used to determine (e.g., calculate) which key body points of the patient's body are visible in the images captured by the sensing device 202a, and thus can be covered by the human body model 204a with high accuracy or confidence.
[0024] Accuracy or confidence indicators, such as accuracy or confidence scores, can be determined and / or stored to indicate which body keypoints covered by the sensing device or single-view model can have a high level of accuracy or a high confidence score (e.g., fully visible to the sensing device) and which body keypoints covered by the sensing device or single-view model can have a low level of accuracy or a low confidence score (e.g., partially visible or invisible to the sensing device). Such accuracy or confidence indicators can, for example, be determined by the individual sensing devices during the process of analyzing images captured by the sensing devices and generating a single-view model as described herein. Accuracy or confidence indicators can also be determined by, for example, […]. Figure 1 The processing unit 106's separate computing device determines this. The computing device can determine the accuracy or confidence level of each body keypoint and / or sensing device based, for example, on the spatial relationship between the sensing device and the patient. For instance, if the computing device determines that a body keypoint is not within the FOV of the sensing device or is partially obscured within the FOV, the computing device can assign a low accuracy or confidence score to that particular body keypoint / sensing device. Conversely, if the computing device determines that the body keypoint is entirely within the FOV of the sensing device, the computing device can assign a high accuracy or confidence score to that particular body keypoint / sensing device.
[0025] The multi-view model builder 208 can be configured to receive all or a subset of information about the single-view human body models 204a-c, and construct the multi-view model 206 based on the received information. It should be noted that the multi-view model builder 208 can be a computing device independent of the sensing devices 202a-c, or the multi-view model builder 208 can be one of the sensing devices 202a-c. Therefore, Figure 2 The fact that the multi-view model builder 208 is depicted as separate from the sensing devices 202a-c should not be interpreted as indicating that the sensing devices 202a-c are incapable of assuming the role of the multi-view model builder 208. In any case, the multi-view model builder 208 may be configured to obtain (e.g., receive) relevant information about two or more single-view models generated by the sensing devices 202a-c, determine relevant body key points (e.g., 2D body key points) of the patient's body that can be covered (e.g., depicted) by the respective single-view models at a specific level of accuracy or confidence, and use the corresponding portions of the single-view models corresponding to these body key points to generate a multi-view model 206. For example, the multi-view model builder 208 may receive first information associated with a first single-view model of the patient (e.g., model 204a) and second information associated with a second single-view model of the patient (e.g., model 204b). The first information may instruct the first model to cover (e.g., depict) at least a first body keypoint (e.g., left shoulder) of the patient's body (e.g., with a satisfactory level of accuracy or confidence), and the second information may instruct the second model to cover at least a second body keypoint (e.g., left leg) of the patient's body (e.g., with a satisfactory level of accuracy or confidence). Based on the first and second information, the multi-view model builder 208 may determine a first region of the first single-view model associated with the first body keypoint and a second region of the second single-view model associated with the second body keypoint, and generate a multi-view model 206 (e.g., a third model) based at least on the first region of the first single-view model and the second region of the second single-view model.
[0026] The multi-view model builder 208 can determine and / or store the correlation between individual body keypoints (e.g., 2D body keypoints) of the patient's body and regions (e.g., 3D regions) in each of the single-view models 204a-c. Such regions may, for example, include a set of vertices corresponding to the single-view model. For example, the region corresponding to the right shoulder body keypoint may include a cube of certain dimensions surrounding the right shoulder and encompassing multiple vertices of the single-view model. The regions (e.g., cubes) and their dimensions may be pre-configured or predefined by or for the multi-view model builder 208. Therefore, the multi-view model builder 208 can construct the multi-view model 206 by traversing the list of body keypoints associated with the patient's body, selecting regions corresponding to the individual body keypoints from the multiple single-view models (e.g., 204a-c), and using those regions (e.g., the vertices and / or model parameters associated with those regions) to construct a portion of the multi-view model 206. If multiple single-view models cover body keypoints at a specific accuracy or confidence level (e.g., the right shoulder is visible in the FOV of both sensing devices 202a and 202b and is covered by both single-view models 204a and 204b), the multi-view model builder 208 can select the corresponding region from the single-view models that has the highest accuracy or confidence score for the body keypoint (e.g., the right shoulder).
[0027] For each selected region, the multi-view model builder 208 can determine a set of vertices that can be used to construct the multi-view model 206 from the corresponding single-view model. In an example, the multi-view model builder 208 can project this set of vertices from the coordinate system associated with the single-view model to the coordinate system associated with the multi-view model 206 (e.g., the "world" coordinate system) (e.g., transforming the position of the vertices). The "world" coordinate system can be defined in various ways, for example, based on the requirements of downstream applications. For example, the "world" coordinate system can be defined based on the viewpoint of the multi-view model builder 208, the viewpoint of a medical expert, the viewpoint of the display device used to present the multi-view model 206, etc. The projection (e.g., transformation) can be performed, for example, based on the rotation matrix and / or translation vector from the respective cameras associated with the respective sensing devices 202a, 202b, 202c to the world. Such rotation matrix and / or translation vector can be determined, for example, based on the respective position of the sensing device relative to the "world" (e.g., the origin position of the "world" coordinate system) (e.g., predetermined during system calibration).
[0028] Once the vertices corresponding to different regions of the single-view model have been transformed (e.g., unified in the "world" coordinate system), the multi-view model builder 208 can combine them and derive the multi-view model 206 based on this combination (e.g., by connecting vertices with edges to form polygons, connecting multiple polygons to form surfaces, using multiple surfaces to determine 3D shapes, and / or applying textures and / or shading to surfaces and / or shapes). In this way, the patient's views from the various sensing devices (e.g., as reflected in the various single-view models 204a, 204b, 204c) can be transformed and combined to form a patient's "world" view covering all body key points of the patient's body (e.g., as reflected in the multi-view model 206).
[0029] Figure 3 An example is shown based on multiple single-view human body models (e.g., Figure 2 Constructing a multi-view human body model 300 from the single-view model 204a-c (e.g., Figure 2 The figure shows a multi-view human body model 300 (206). As shown, a first part of the multi-view human body model 300, including these body key points, can be constructed based on a single-view model 1 (e.g., based on determining that single-view model 1 covers the head and left arm with a satisfactory level of accuracy or confidence (e.g., this level can be predefined)). Similarly, based on determining that single-view models 2 and 3 cover the corresponding body key points with a satisfactory level of accuracy or confidence, a second part of the multi-view human body model 300, including the back and right arm, can be constructed based on single-view model 2, and a third part of the multi-view human body model 300, including the left and right legs, can be constructed based on single-view model 3.
[0030] Figure 4 An example is shown that can be used with a multi-view human body model (e.g., Figure 2 The flowchart of the associated operations 400 in the construction of the multi-view model 206 in the model. These operations can be performed by, for example... Figure 2The multi-view model builder 208 shown is used to perform this action. As shown, the multi-view model builder can obtain (e.g., receive) first information associated with a first single-view model of the patient at 402. The first information obtained by the multi-view model builder may include, for example, parameters of the first single-view model (e.g., multiple vertices included in the first model), indications of one or more body keypoints covered by the first single-view model, and / or indications of the corresponding accuracy or confidence level of the covered body keypoints. At 404, the multi-view model builder can obtain (e.g., receive) second information associated with a second single-view model of the patient. The second information received by the multi-view model builder may include, for example, parameters of the second single-view model (e.g., multiple vertices included in the second model), indications of one or more body keypoints covered by the second single-view model, and / or indications of the corresponding accuracy or confidence level of the covered body keypoints. At 406, the multi-view model builder may select a first region (e.g., a first cube) of a first single-view model associated with a first body keypoint (e.g., left shoulder) and a second region (e.g., a second cube) of a second single-view model associated with a second body keypoint (e.g., right shoulder) based on first information and / or second information. The multi-view model builder may select a first single-view model for the first region based on determining that the first information indicates the first single-view model covers the first body keypoint with a satisfactory accuracy or confidence level (e.g., the accuracy or confidence level of the first body keypoint exceeds a predetermined threshold). Similarly, the multi-view model builder may select a second single-view model for the second region based on determining that the second information indicates the second single-view model covers the second body keypoint with a satisfactory accuracy or confidence level (e.g., the accuracy or confidence level of the second body keypoint exceeds a predetermined threshold).
[0031] At 408, the multi-view model builder obtains parameters associated with a first region from a first single-view model and parameters associated with a second region from a second single-view model. Parameters may include, for example, corresponding vertices of the first and second single-view models corresponding to the first and second regions. At 410, the multi-view model builder transforms the vertices from the corresponding coordinate systems associated with the first and second single-view models (e.g., associated with the first and second sensing devices that generated the single-view models) to a target coordinate system (e.g., the world coordinate system) associated with the target multi-view human body model. At 412, the multi-view model builder generates the target multi-view human body model using at least the transformed vertices of the first and second single-view models, and then at 414 passes the multi-view human body model to a downstream application for use or display.
[0032] For the sake of simplicity, the operations 400 of the multi-view model builder are depicted and described in a specific order throughout this document. However, it should be noted that these operations can occur in various orders, simultaneously, and / or with other operations not presented or described herein. Furthermore, it should be noted that... Figure 4 Not all operations that the Multi-View Model Builder can perform are described in this document. It should also be noted that not all of the illustrated operations require the Multi-View Model Builder to perform.
[0033] The systems, methods, and / or apparatuses described herein may be implemented using one or more processors, one or more storage devices, one or more sensors, and / or other suitable auxiliary devices (such as display devices, communication devices, input / output devices, etc.). Figure 5 This is a block diagram illustrating an example device 500 that can be configured to perform the single-view and / or multi-view model building operations described herein. As shown, device 500 may include a processor (e.g., one or more processors) 502, which may be a central processing unit (CPU), graphics processing unit (GPU), microcontroller, reduced instruction set computer (RISC) processor, application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), physical processing unit (PPU), digital signal processor (DSP), field-programmable gate array (FPGA), or any other circuitry or processor capable of performing the functions described herein. Device 500 may also include communication circuitry 504, memory 506, mass storage device 508, input device 510, sensor 512, and / or communication link 514 (e.g., communication bus) through which one or more components shown in the figure exchange information.
[0034] Communication circuitry 504 can be configured to send and receive information using one or more communication protocols (e.g., TCP / IP) and one or more communication networks, including local area networks (LANs), wide area networks (WANs), the Internet, and wireless data networks (e.g., Wi-Fi, 3G, 4G / LTE, or 5G networks). Memory 506 may include a storage medium (e.g., a non-transitory storage medium) configured to store machine-readable instructions that, when executed, cause processor 502 to perform one or more functions described herein. Examples of machine-readable media may include volatile or non-volatile memory, including but not limited to semiconductor memory (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), flash memory, etc.). Mass storage device 508 may include one or more disks, such as one or more internal hard disks, one or more removable disks, one or more magneto-optical disks, one or more CD-ROMs or DVD-ROMs, etc., on which instructions and / or data may be stored for operation of processor 502. Input device 510 may include a keyboard, mouse, voice-controlled input device, touch-sensitive input device (e.g., touchscreen), etc., for receiving user input from device 500. If, for example, device 500 is configured to perform one or more image capture functions as described herein (e.g., if device 500 is a sensing device as described herein), then sensor 512 may be included in device 500. Sensor 512 may include a 2D vision sensor (e.g., a 2D camera), a 3D vision sensor (e.g., a 3D camera), an RGB sensor, a depth sensor, an RGB-D sensor, a thermal sensor (e.g., an infrared (FIR) or near-infrared (NIR) sensor), a radar sensor, and / or other types of image capture devices or circuitry.
[0035] It should be noted that device 500 can operate as a standalone device or can be connected to other computing devices (e.g., networked or grouped) to perform the functions described herein. And even in Figure 5 Only one example of each component is shown in the figure, and those skilled in the art will understand that the device 500 may include multiple instances of one or more components shown in the figure.
[0036] Although this disclosure has been described according to certain embodiments and generally associated methods, changes and variations of the embodiments and methods will be apparent to those skilled in the art. Therefore, the above description of exemplary embodiments does not limit this disclosure. Other changes, substitutions, and modifications are possible without departing from the spirit and scope of this disclosure. Furthermore, unless specifically stated otherwise, discussions using terms such as “analyze,” “determine,” “enable,” “identify,” and “modify” refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data representing physical (e.g., electronic) quantities within the registers and memories of the computer system into other data representing physical quantities within the computer system's memory or other such information storage, transmission, or display devices.
[0037] It should be understood that the above description is intended to be illustrative and not restrictive. Many other embodiments will become apparent to those skilled in the art upon reading and understanding the above description. Therefore, the scope of this disclosure should be determined by reference to the appended claims and the full scope of their equivalents.
Claims
1. A method for constructing a three-dimensional (3D) human body model, the method comprising: First information associated with a first 3D model of a human body is obtained, wherein the first 3D model is determined based on at least a first image of the human body captured by a first sensing device, and wherein the first information indicates that the first 3D model covers at least a first body key point of the human body. Second information is obtained associated with a second 3D model of the human body, wherein the second 3D model is determined based on at least a second image of the human body captured by a second sensing device, and wherein the second information indicates that the second 3D model covers at least a second body key point of the human body; Determine a first region of the first 3D model associated with the first body keypoint and a second region of the second 3D model associated with the second body keypoint, wherein the first body keypoint corresponds to a first portion of the human body visible in the first image but invisible in the second image, and wherein the second body keypoint corresponds to a second portion of the human body visible in the second image but invisible in the first image; and A third 3D model of the human body is generated based at least on the first region of the first 3D model and the second region of the second 3D model.
2. The method according to claim 1, further comprising: A first set of vertices associated with the first region of the first 3D model is determined based on the first 3D model, and a second set of vertices associated with the second region of the second 3D model is determined based on the second 3D model, wherein the third 3D model of the human body is generated based at least on the first set of vertices and the second set of vertices.
3. The method according to claim 2, further comprising: Determine the corresponding positions of the first group of vertices in the coordinate system associated with the third 3D model; as well as Determine the corresponding positions of the second group of vertices in the coordinate system associated with the third 3D model; The third 3D model of the human body is generated at least based on the corresponding positions of the first set of vertices in the coordinate system and the corresponding positions of the second set of vertices in the coordinate system.
4. The method according to claim 3, wherein, The corresponding positions of the first set of vertices in the coordinate system are determined at least based on the position of the first sensing device in the coordinate system, and the corresponding positions of the second set of vertices in the coordinate system are determined at least based on the position of the second sensing device in the coordinate system.
5. The method according to claim 1, wherein, The first region corresponds to the first cube in the first 3D model, and the second region corresponds to the second cube in the second 3D model.
6. The method according to claim 1, further comprising: Based on the first information and the second information, it is determined that both the first 3D model and the second 3D model cover the second body key points of the human body; as well as Based on the confidence indices associated with the second body keypoints, which are associated with at least one of the first 3D model or the second 3D model, it is determined that information associated with the second body keypoints will be obtained from the second 3D model.
7. The method according to claim 6, further comprising: The confidence level indication is determined based on the position of the first sensing device relative to the second body key point and the position of the second sensing device relative to the second body key point.
8. The method according to claim 1, wherein, The first 3D model is constructed by the first sensing device, and the second 3D model is constructed by the second sensing device, wherein the method further includes: receiving the first 3D model from the first sensing device and receiving the second 3D model from the second sensing device.
9. An apparatus for constructing a three-dimensional (3D) human body model, comprising: One or more processors are configured to execute any of the methods in 1-8.
Citation Information
Patent Citations
Systems and methods for scanning three-dimensional objects
US20170372527A1