Three-dimensional model monitoring apparatus and program of the same
The 3D model monitoring device addresses the inefficiencies of manual virtual camera operation by automatically focusing on and highlighting areas of interest in 3D models, enhancing defect detection and reducing time consumption.
Patent Information
- Application Number
- JP2024094195
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-23
AI Technical Summary
Existing methods for checking live-action 3D models created using volumetric capture technology are time-consuming and prone to missing defects due to the need for manual operation of a virtual camera to view multiple directions and focus on specific parts, with no device capable of automatically highlighting areas of interest.
A 3D model monitoring device equipped with posture information estimation, capture area determination, target part viewpoint calculation, and display units to automatically focus on and highlight areas of interest within the 3D model, using virtual cameras to provide enhanced viewing and defect detection.
Enables efficient and automated monitoring of 3D models by focusing on specific parts of interest, reducing the time required to identify and display potential defects, thereby improving the quality assurance process.
Smart Images

Figure 2025185797000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a three-dimensional model monitoring device and a program therefor. [Background technology]
[0002] Live-action three-dimensional (3D) models created using volumetric capture technology can have defects due to various factors, such as calibration errors in the camera used to capture the image, the use of mirrored or thin objects, overexposure due to lighting, and the subject extending beyond the capture area. Currently, the only way to check for defects in live-action 3D models is for an observer to visually check the created 3D model. In this case, the observer places the 3D model in a three-dimensional virtual space and checks the quality based on the rendering results.
[0003] Furthermore, 3D models created using volumetric capture technology are called volumetric videos, and are not static 3D models but dynamic 3D models consisting of multiple frames. Therefore, when an observer checks a 3D model, they must move the virtual camera to the area of interest in the 3D virtual space each time a frame changes, which is time-consuming. Therefore, a device that can automatically display the area of interest is needed. Non-Patent Document 1 discloses a technology for placing a real-life 3D model created using volumetric capture technology in a three-dimensional virtual space and detecting the optimal viewpoint for positioning a virtual camera to display the 3D model. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Jingwu He etc, “Viewpoint Assessment and Recommendation foe Photographing Architectures”, IEEE Visualization 2019, vol.25, No.8, pp.2636-2649 Summary of the Invention [Problem to be solved by the invention]
[0005] When checking a live-action 3D model created using volumetric capture technology, the observer must check the image from multiple directions by operating a virtual camera in 3D space, rather than viewing the image from a single viewpoint. Operating the virtual camera is complicated, requiring multiple operations such as rotating and translating the camera and changing the viewing angle, making it time-consuming to check the quality of the model, and in some cases there is a risk that parts with quality issues may be overlooked.
[0006] Furthermore, when there is a specific part of a dynamic 3D model that the viewer wishes to focus on, in order to track that part, the viewer must operate and move the virtual camera for each frame, which is time-consuming. Furthermore, the technology described in Non-Patent Document 1 detects the optimal viewpoint for positioning a virtual camera to display a 3D model, but there is no device that can operate the virtual camera while focusing on any desired part, such as a missing part.
[0007] The present invention was made in consideration of such problems, and its objective is to provide a 3D model monitoring device and a program therefor that are capable of monitoring 3D models created using volumetric capture technology by focusing on areas of interest. [Means for solving the problem]
[0008] In order to solve the above problems, the three-dimensional model monitoring device of the present invention is a three-dimensional model monitoring device that monitors a three-dimensional model created using volumetric capture technology, and is configured to include a posture information estimation unit, a capture area determination unit, a target part viewpoint information calculation unit, a selection unit, and a display unit.
[0009] In this configuration, the 3D model monitoring device uses a posture information estimation unit to estimate the 3D coordinate positions of each predetermined part of the subject as posture information. The posture information can be estimated from multiple camera images captured by posture estimation cameras placed at multiple positions.
[0010] The three-dimensional model monitoring device then uses a capture area determination unit to determine a predetermined set of the subject's body parts as parts of interest, and, based on the posture information, determines whether each part of each part of interest is inside or outside the capture areas of all model acquisition cameras in the volumetric studio. In the three-dimensional model monitoring device, the part-of-interest viewpoint information calculation unit calculates, for each part of interest, viewpoint information of a part-of-interest virtual camera that virtually photographs the part of interest.
[0011] Furthermore, the 3D model monitoring device uses a selection unit to display buttons for selecting the part of interest on the screen along with the result of the inside / outside determination, and accepts the selection. In this way, the 3D model monitoring device visually displays the result of the inside / outside determination using colors, patterns, etc., thereby making it possible for the observer to know that there are parts of the part of interest that have not been photographed.
[0012] The 3D model monitoring device then causes the display unit to virtually photograph the 3D model with the virtual camera for the part of interest based on the viewpoint information corresponding to the part of interest selected by the selection unit, and displays the photograph on the screen. This allows the 3D model monitoring device to display the 3D model while focusing on the part of interest that the observer wants to check.
[0013] In order to achieve the above object, the three-dimensional model monitoring device according to the present invention may further include an attention area determination unit and an attention area viewpoint information calculation unit.
[0014] In this configuration, the three-dimensional model monitoring device determines, for each model acquisition camera, a pixel group that is a missing pixel in the three-dimensional model based on a plurality of determination items from the camera image as a caution area by the caution area determination unit. Then, the three-dimensional model monitoring device calculates, for each attention area, viewpoint information of a virtual camera for attention area that virtually photographs the attention area, by the attention area viewpoint information calculation unit.
[0015] At this time, the selection unit further displays buttons for selecting the judgment item and the model acquisition camera on the screen and accepts the selection. Furthermore, by selecting the judgment item and the model acquisition camera, the display unit virtually photographs the 3D model with the attention area virtual camera based on the viewpoint information of the corresponding attention area, and displays it on the screen. This allows the 3D model monitoring device to present areas indicative of possible defects in the 3D model. The three-dimensional model monitoring device can be operated by a program that causes a computer to function as a three-dimensional model monitoring device. [Effects of the Invention]
[0016] According to the present invention, in a 3D model created by volumetric capture technology, it is possible to monitor the 3D model by focusing on any designated part. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a block diagram showing the configuration of a three-dimensional model monitoring device according to an embodiment of the present invention. [Figure 2] 1A and 1B are schematic diagrams showing an example of camera arrangement in a volumetric capture studio, where (a) is an XZ plan view and (b) is an XY plan view. [Figure 3] FIG. 10 is a diagram illustrating an example of joint points used as posture information. [Figure 4] FIG. 10 is a diagram illustrating an example of a target part. [Figure 5] 10 is an explanatory diagram for explaining the field of view of a virtual camera with respect to one three-dimensional coordinate point in the attention part viewpoint information calculation unit. FIG. [Figure 6] 10 is an explanatory diagram for explaining the field of view of a virtual camera with respect to a plurality of three-dimensional coordinate points in the attention part viewpoint information calculation unit. FIG. [Figure 7] FIG. 10 is a diagram showing an example of an attention area. [Figure 8] 10 is an explanatory diagram for explaining an example of viewpoint information of a virtual camera in an attention area viewpoint information calculation unit. FIG. [Figure 9] 10 is a diagram showing a selection button row showing an example of selection buttons displayed on a screen by a selection unit. FIG. [Figure 10] FIG. 2 is a diagram showing an example of a graphical user interface (GUI) screen displayed on a screen by a display unit. [Figure 11] 4 is a flowchart showing the operation of the three-dimensional model monitoring device according to the embodiment of the present invention. [Figure 12] 12 is a flowchart showing the operation of the attention part identification process of FIG. 11. [Figure 13] 12 is a flowchart showing the operation of the attention area identification process of FIG. 11. [Figure 14] FIG. 10 is a block diagram showing the configuration of a three-dimensional model monitoring device according to a first modified example of the present invention. [Figure 15] FIG. 10 is a block diagram showing the configuration of a three-dimensional model monitoring device according to a second modified example of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [Configuration of 3D model monitoring device] The configuration of a three-dimensional model monitoring device 1 according to an embodiment of the present invention will be described with reference to FIG.
[0019] The three-dimensional model monitoring device 1 is a device that displays a region of interest in a real-life three-dimensional (3D) model created using volumetric capture technology as an image captured by a virtual camera. The region of interest here is a region that the observer wants to focus on in the 3D model, a region where there may be a defect in the 3D model, etc. The three-dimensional model monitoring device 1 receives as input a 3D model, a camera image for pose estimation, and a camera image for model acquisition.
[0020] The 3D model was captured by multiple model acquisition cameras C in the volumetric capture studio (hereafter referred to as the studio) shown in Figure 2. M It is three-dimensional data of the subject S generated by volumetric capture technology from multiple two-dimensional images of the subject S captured by a camera. Here, as an example, the subject S is a person. Figure 2(a) is a schematic diagram of the studio viewed from above, corresponding to a cross section parallel to the XZ plane. Figure 2(b) is a schematic diagram of the studio viewed from the front, corresponding to a cross section parallel to the XY plane. The coordinate system used here is a studio coordinate system (world coordinate system) with the floor surface at the center of the studio as the origin and the positive direction of the Y axis pointing upward from the studio floor surface.
[0021] The posture estimation camera image is an image for estimating the posture of the subject S in the studio shown in FIG. 2, and is captured by the posture estimation main camera C PM and sub-camera C for posture estimation PS The video was taken with the main camera C for posture estimation. PM The sub-camera for posture estimation C is placed in front of the studio. PS is the main camera C for posture estimation PM In Figure 2, the sub-camera C for posture estimation is placed at a position rotated 45 to 90 degrees around the origin. PS The main camera C for posture estimation PMThe example shows the position rotated 90 degrees clockwise from the original position.
[0022] The posture estimation camera does not have to be a camera actually installed in the studio. For example, the posture estimation camera may be a virtual camera that virtually captures a 3D model. Even when a virtual camera is used in this way, the positional relationship between the main camera and the sub-camera is the same as the positional relationship between the main camera and the sub-camera actually installed in the studio.
[0023] The model acquisition camera images were taken by multiple model acquisition cameras C in the studio shown in Figure 2. M The video was taken by multiple model acquisition cameras C M Each model acquisition camera C M A number for identifying the item is associated with the item. The 3D model, posture estimation camera footage, and model acquisition camera footage can be pre-created and pre-recorded. However, if the 3D model can be acquired in real time, the posture estimation camera footage and model acquisition camera footage taken during shooting, and the 3D model created from them, can also be used.
[0024] As shown in FIG. 1, the three-dimensional model monitoring device 1 includes a posture information estimation unit 10, a virtual camera identification unit for a target part 11, a virtual camera identification unit for an attention area 12, a selection unit 13, and a display unit 14.
[0025] The posture information estimation unit 10 estimates the three-dimensional coordinate position of each predetermined part of the subject as posture information. Here, the posture information estimation unit 10 estimates the posture of the subject for each frame from the posture estimation camera video.
[0026] Figure 3 shows 30 body parts P1,...,P2, etc. that characterize a person's body, such as the head, neck, right shoulder, and left shoulder. 30Here, all parts for specifying the posture, including parts that are not joints themselves, are called joint points. The coordinates of the joint points are three-dimensional coordinates in the world coordinate system set as shown in Figure 2. Below, each part P of the subject is called j The three-dimensional coordinates where is located are expressed as (P jx ,P jy ,P jz ) where j indicates the identification number of the subject part (j=1, 2, ..., M). Note that M is the total number of parts (here, M=30).
[0027] For estimating posture information, an existing posture estimation algorithm capable of detecting the joint points of a person can be used. For example, OpenPose or VisionPose (registered trademark of NextSystems) shown in Reference 1 below can be used as a posture estimation algorithm.
[0028] (Reference 1) Zhe Cao etc, “OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields”, arXiv preprint arXiv:1812.08008(2018)
[0029] Here, an example using VisionPose will be explained. VisionPose uses two cameras for pose estimation (the main camera C for pose estimation). PM , Sub-camera C for posture estimation PS ) position information, the 3D coordinates of the subject's joint points can be obtained. The 3D coordinates estimated by VisionPose are obtained by the main camera C for posture estimation. PM The coordinate system has the origin at . Therefore, the posture information estimation unit 10 calculates the posture information of the posture estimation main camera C PM The 3D coordinates estimated by the VisionPose algorithm are converted into coordinates in the world coordinate system by translating the origin of the to the origin of the studio, which is the origin of the world coordinate system. The posture information estimation unit 10 outputs the estimated posture information in the world coordinate system (three-dimensional coordinates of each part) to the target part virtual camera identification unit 11 and the attention area virtual camera identification unit 12.
[0030] The posture information estimation unit 10 sets the three-dimensional coordinates of the part for which posture information cannot be estimated to specific coordinates (0,0,0) or coordinates outside the valid range in which the subject exists, for example, (-1000,-1000,-1000). PM When the subject's arms are hidden by the body, or when the subject is outside the shooting area of the main camera C for posture estimation, PM and sub-camera C for posture estimation PS If the location cannot be seen from both sides, set the coordinates to (0,0,0) or similar.
[0031] The virtual camera for a part of interest specifying unit 11 specifies a virtual camera that virtually captures an image of a part of interest. The part of interest is a part of a 3D model that an observer focuses on, and refers to a set of multiple joint points in the posture information of the subject.
[0032] Here, examples of the target part will be described with reference to Figs. 3 and 4. As shown in Fig. 4, the target part is a part of the subject, which is made up of 30 joint points (P1, ..., P 30 ) are assembled together. Here, the parts of interest are a whole body part A1, a face part A2, a left arm part A3, a right arm part A4, a torso part A5, a left leg part A6, and a right leg part A7.
[0033] The face part A2 consists of seven parts (seven joints): the head (Head: P1), left eye (Eye Left: P2), right eye (Eye Right: P3), left ear (Ear Left: P4), right ear (Ear Right: P5), nose (Nose: P6), and neck (Neck: P7). Left Arm part A3 consists of left shoulder (Shoulder Left: P8), left elbow (Elbow Left: P9), left wrist (Wrist Left: P 10 ), Left Hand (Hand Left:P 11 ), left thumb (Thumb Left:P 12 ), left hand (Top Left:P 13 ) and consists of six parts (six joint points). Right Arm Part A4 is the right shoulder (Shoulder Right:P 14 ), Right Elbow (Elbow Right:P 15 ), Right Wrist (P 16 ), Right hand (Hand Right:P 17 ), Right Thumb (Thumb Right:P 18 ), Top Right:P 19 ) and consists of six parts (six joint points). Body part A5 is the center of the shoulder (Spine Shoulder:P 20 ), Shoulder Left:P8, Shoulder Right:P 14 ), spine (Spine Mid:P 21 ), waist center (Spine Base:P 22 ), left hip (Hip Left:P 23 ), Right Hip (Hip Right:P 24 ) and consists of six parts (six joint points).
[0034] The left leg part A6 is attached to the center of the waist (Spine Base:P 22 ), left hip (Hip Left:P 23 ), left knee (Knee Left:P 25 ), Left Ankle (Ankle Left:P 26 ), left foot (Foot Left:P 27 ) and consists of five parts (five joint points). Right Leg part A7 is attached to the center of the waist (Spine Base:P 22 ), Right Hip (Hip Right:P24 ), Right Knee (Knee Right:P 28 ), Right Ankle (Ankle Right:P 29 ), Right Foot (Foot Right:P 30 ) and consists of five parts (five joint points). Full Figure Part A1 includes all parts (P1~P 30 ) is composed of
[0035] Returning to FIG. 1, the configuration of the three-dimensional model monitoring device 1 will be further described. The target part virtual camera specifying unit 11 specifies the gaze point and the viewing angle of a virtual camera that virtually photographs each target part. Here, the target part virtual camera specifying unit 11 includes a capture area determining unit 110 and a target part viewpoint information calculating unit 111.
[0036] The capture area determination unit 110 determines the position of all the model acquisition cameras C for each part (joint point) of each target part based on the posture information estimated by the posture information estimation unit 10. M (See Figure 2) to determine whether the camera is inside or outside the capture area (shooting range).
[0037] Specifically, a method for determining whether an object is inside or outside the capture area will be described using mathematical expressions. Main camera C for posture estimation PM (See Figure 2) M The horizontal tilt of the camera is ψ, the vertical tilt is θ, and the tilt in the camera axis direction is φ. M The position of the origin of the studio coordinate system as seen from the camera coordinate system is (t x ,t y ,t z ) At this time, the capture area determination unit 110 calculates the three-dimensional coordinates (P jx ,P jy ,P jz ) into the three-dimensional coordinates (A jx ,Ajy ,A jz )
[0038]
number
[0039] Next, the capture area determination unit 110 calculates the three-dimensional coordinates (A jx ,A jy ,A jz ) into two-dimensional coordinates (N jx ,N jy )
[0040]
number
[0041] Then, the capture area determination unit 110 calculates the two-dimensional coordinates (N jx ,N jy ) and the model acquisition camera C M The two-dimensional coordinates of the image coordinate system (u j ,v j )
[0042]
number
[0043] where f is the camera C for model acquisition M focal length (mm), δ u is the model acquisition camera C M The physical spacing of pixels in the horizontal direction (mm) in the image sensor, δ v is the model acquisition camera C M The physical spacing of pixels in the vertical direction of the image sensor (mm), (c u ,c v ) is the position of the intersection of the optical axis and the image plane in the image coordinate system (image center). MThese are the internal parameters. These internal parameters can be obtained through camera calibration. Since camera calibration can use common methods as shown in Reference 2 below, the description is omitted here.
[0044] (Reference 2) Zhengyou Zhang etc, “A Flexible New Technique for Camera Calibration”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.22, Issue:11, pp.1330-1334(2000)
[0045] The capture area determination unit 110 performs this conversion for all cameras C for model acquisition M at all the joint points of the captured subject, thereby projecting the joint points of the subject onto the image coordinates of each camera C for model acquisition M Then, if the capture area determination unit 110 projects the joint points within the captured images of all cameras C for model acquisition M , it determines that the joint point is within the capture area, and otherwise, it determines that the joint point is outside the capture area. That is, when the number of horizontal pixels of the camera C for model acquisition M is W and the number of vertical pixels is H, the capture area determination unit 110 determines that the joint point is within the capture area if the two-dimensional coordinates (u M , v j ) in the image coordinate system of the camera C for model acquisition j satisfy 0 ≦ u j < W and 0 ≦ v j < H, and otherwise, it determines that the joint point is outside the capture area.
[0046] The capture area determination unit 110 outputs the determination results of whether the caption area is inside or outside for each part (joint point) of the subject to the target part viewpoint information calculation unit 111. Also, the capture area determination unit 110 performs for all cameras C for model acquisition MThe two-dimensional coordinates of the joint points in the image coordinate system (see the above formula (3)) are output to the attention area virtual camera specifying unit 12.
[0047] The part-of-interest viewpoint information calculation unit 111 calculates, for each part of interest, viewpoint information of a virtual camera for the part of interest that virtually photographs the part of interest. Here, the viewpoint of the virtual camera that captures the target part is specified by a point within the target part to be gazed at (point of gaze), a predetermined distance from the point of gaze as the center, and a viewing angle that includes the target part in its field of view. In other words, the virtual camera is a camera that virtually captures images on a sphere having a radius that is a predetermined distance from the point of gaze as the center, at a viewing angle that includes the target part in its field of view. The image captured by this virtual camera is displayed as an enlarged image of the target part on the display unit 14, which will be described later.
[0048] The part-of-interest viewpoint information calculation unit 111 calculates the gaze point and the viewing angle as viewpoint information for each part of interest (A1, ..., A7) explained in Fig. 4. It is assumed that the parts of interest (A1, ..., A7) are numbered in advance.
[0049] Here, a specific method will be described in which the part-of-interest viewpoint information calculation unit 111 calculates the gaze point and the viewing angle as viewpoint information of the virtual camera corresponding to a certain part of interest. Here, the aspect of the screen captured by the virtual camera is assumed to be a square or a rectangle that is long in the horizontal direction, and the viewing angle is assumed to be the vertical viewing angle. However, if the aspect of the screen is assumed to be a rectangle that is long in the vertical direction, the viewing angle may be assumed to be the horizontal viewing angle.
[0050] If the number of parts of a part of interest that are determined to be within the capture area by the capture area determination unit 110 is "0", the part of interest viewpoint information calculation unit 111 does not calculate camera viewpoint information because there is no element that determines the viewpoint. In this case, there is no virtual camera that corresponds to the part of interest, and the part of interest cannot be enlarged and displayed.
[0051] When the number of parts determined to be within the capture area is "1", the attention part viewpoint information calculation unit 111 sets the gaze point q to the three-dimensional coordinates of the part determined to be within the capture area (here, p1), as shown in FIG. A Virtual camera C V Display area D included in the field of view of A radius (constant), d c The gaze point q and the virtual camera C V Let r be the distance (constant) between A and d c is set in advance. Virtual camera C V is the gaze point q(q x ,q y ,q z ) centered on r A A sphere with radius D is displayed. A The vertical viewing angle θ is calculated by the following equation (4) so that it is included in h Calculate.
[0052]
number
[0053] Furthermore, when the number of parts determined to be within the capture area is "2" or more, the attention part viewpoint information calculation unit 111 calculates the gaze point q by dividing the gaze point q by the center of gravity (x c ,y c ,z c ) and d pmax Let d be the distance from the gaze point q to the farthest part (here, p1). pmax +r B Virtual camera C V Display area D included in the field of view of B radius, d c The gaze point q and the virtual camera C V The distance (constant) between the B is display area D BFor example, if P1 is the head part shown in Figure 3, the joint point is located at the center of the head, so part of the head is in front of the virtual camera C. V Therefore, the preset adjustment value r B Display area D B Increase the radius of the Virtual Camera C V is the gaze point q(x c ,y c ,z c ) at the center, d pmax +r B A sphere with radius D is displayed. B The vertical viewing angle θ is calculated by the following equation (5) so that it is included in h Calculate.
[0054]
number
[0055] The viewpoint information of the target part is taken from the virtual camera C V Vertical viewing angle θ h is a constant, and the gaze point q and the virtual camera C V Distance d c In that case, d c If the number of parts determined to be within the capture area is "1", it can be calculated using the following formula (6), and if the number of parts determined to be within the capture area is "2" or more, it can be calculated using the following formula (7).
[0056]
number
number
[0057] Returning to FIG. 1, the configuration of the three-dimensional model monitoring device 1 will be further described. The target part viewpoint information calculation unit 111 outputs the calculated viewpoint information (gazing point and viewing angle) to the selection unit 13 in association with each part.
[0058] The attention area virtual camera specifying unit 12 specifies a virtual camera that virtually captures the attention area. The attention area is an area that indicates a possibility of a defect in the 3D model and is an area that calls the observer's attention. The attention area virtual camera specifying unit 12 includes an attention area determining unit 120 and an attention area viewpoint information calculating unit 121 .
[0059] The attention area determination unit 120 is a model acquisition camera C M For each frame (see FIG. 2), a set of pixels in the camera image that are defective pixels in the 3D model based on multiple criteria is determined as an attention area. Here, the attention area determination unit 120 determines, for each frame of the model acquisition camera image, an area that will be defective pixels in the 3D model based on predetermined conditions as an attention area. Causes of defects in the 3D model include, for example, blown-out highlights, color overlay, inability to acquire information, and insufficient pixel features.
[0060] Overexposure refers to a condition in which the gradation in bright areas is lost and the image turns white. Color cast refers to the bias of an image toward a specific color. For example, if a volumetric capture studio uses a floor or background of a single color (e.g., green) for background processing, that color will be cast on part of the subject. The inability to acquire information refers to a state in which information cannot be acquired using volumetric capture technology, such as a thin object. The lack of pixel features refers to a state in which there are few image features in a flat area of the image, for example.
[0061] A method for determining the pixels that cause these defects will be described below. When detecting blown-out highlights, the attention area determination unit 120 converts the input model acquisition camera image from RGB format to HSV format for each frame, and determines pixels whose saturation S is lower than a predetermined standard and whose brightness V is higher than a predetermined standard as pixels where blown-out highlights have occurred. Note that a common method can be used to convert from RGB format to HSV format. For example, the attention area determination unit 120 first normalizes each of the R, G, and B values of a pixel to a range of "0" to "1." Here, assuming that the maximum of the three R, G, and B values is MAX and the minimum is MIN, the attention area determination unit 120 calculates the hue H using the following equation (8).
[0062]
number
[0063] Here, if H<0, the attention area determination unit 120 adds "360" to H to keep H within the range of "0" to "360". The lightness V can be set to the same value as MAX. Also, the saturation S can be calculated by subtracting MAX from MIN. This conversion from RGB format to HSV format may be performed using an open source library called OpenCV.
[0064] When detecting color cast, the attention area determination unit 120 pre-sets ranges of the hue, saturation, and brightness of the background color in an HSV format image, and determines pixels within that range in the subject part extracted from the frame by background processing as color cast pixels.
[0065] When detecting when information cannot be acquired, the attention area determination unit 120 extracts the silhouette of the subject from the frame using background subtraction or chromakey processing, and determines that pixels in a line width area within the silhouette that is shorter than a predetermined width are pixels from which information cannot be acquired.
[0066] When detecting pixel feature deficiencies, the attention area determination unit 120 grayscales an image in which the background has been removed from the frame using background subtraction or chromakey processing, and extracts features using SIFT (Scale Invariant Feature Transform) or the like.The attention area determination unit 120 then determines pixels in areas with no feature (for example, pixels with no feature in the surrounding 20 pixels) as pixels with insufficient feature.Note that SIFT is a common technique as shown in Reference 3 below, and therefore a detailed description thereof will be omitted here.
[0067] (Reference 3) David G. Lowe. “Distinctive Image Features from Scale-Invariant Keypoints”, International Journal of Computer Vision, 2004
[0068] The attention area determination unit 120 detects all the model acquisition cameras C M For each frame (see Figure 2), the presence or absence of an attention area is determined for each judgment item (bloated highlights, color cast, etc.), and the judgment result and pixel information indicating the pixel position of the attention area are output to the attention area viewpoint information calculation unit 121.
[0069] The attention area viewpoint information calculation unit 121 calculates, for each attention area, viewpoint information of a virtual camera for attention area that virtually photographs the attention area. Here, the virtual camera that captures the attention area is specified by a point to be gazed at within the attention area (gazing point), a predetermined distance from the gaze point, a field of view angle that includes the attention area in its field of view, and a viewpoint angle that indicates the direction of the gaze point. In other words, the virtual camera is a camera that is separated a predetermined distance from the gaze point, includes the attention area in its field of view angle, and virtually captures images at the viewpoint angle. The image captured by this virtual camera is displayed as an image with an enlarged attention area on the display unit 14, which will be described later.
[0070] The attention area viewpoint information calculation unit 121 is a model acquisition camera C M The viewpoint information of the virtual camera is calculated from the camera image (frame) taken by the camera, pixel information of the attention area determined by the attention area determination unit 120, and the two-dimensional coordinates of the joint points calculated by the capture area determination unit 110.
[0071] Specifically, the attention area viewpoint information calculation unit 121 calculates the viewpoint information of a certain model acquisition camera C M A method for calculating the gaze point, the viewing angle, and the viewpoint angle as viewpoint information of a virtual camera corresponding to the above will be described. Here, the aspect of the screen captured by the virtual camera is assumed to be a square or a rectangle that is long in the horizontal direction, and the viewing angle is assumed to be the vertical viewing angle. However, if the aspect of the screen is assumed to be a rectangle that is long in the vertical direction, the viewing angle may be assumed to be the horizontal viewing angle.
[0072] First, the attention area viewpoint information calculation unit 121 removes the background from the camera image and extracts the foreground subject. Note that any common method can be used to extract the subject. One method is a semantic segmentation method, such as DeepLab v3, which is described in Reference 4 below.
[0073] (Reference 4) Liang-Chieh Chen, George. Papandreou, Florian. Schroff, Hartwig. Adam. “Rethinking Atrous Convolution for Semantic Image Segmentation”, European Conference on Computer Vision (ECCV), 2017
[0074] Other methods for extracting a subject include subject extraction using chromakey processing when the background of the subject is a uniform color, and subject extraction using background subtraction. Chromaki processing extracts the subject by determining that the pixels corresponding to the background color are background and deleting them from the image converted to HSV format.In addition, background subtraction extracts the subject by taking the difference for each pixel of the image between a background image without a subject and an image with the current subject.
[0075] The attention area viewpoint information calculation unit 121 extracts the subject using one or more of these multiple methods. As a method using multiple methods, for example, a chromakey process is performed on a region (sum region) that combines the subject region extracted by semantic segmentation and the subject region extracted by background subtraction, and the remaining part is extracted as the subject. In this way, by combining multiple methods, it is possible to extract the subject region with high accuracy.
[0076] Next, the attention area viewpoint information calculation unit 121 generates a binary image from the camera image, in which pixels determined to be attention areas within the subject area are colored a specific color (here, white). Here, the attention area viewpoint information calculation unit 121 performs morphology processing on the binary image, contracting white areas and then expanding the white areas. This processing can be performed using, for example, an open source library called OpenCV. This makes it possible to calculate viewpoint information for a virtual camera that removes very small white areas that become noise and focuses on large white areas. As a result, the attention area viewpoint information calculation unit 121 can generate a binary image IM in which the attention area W is colored white, as shown in Fig. 7. Note that in Fig. 7, the subject is illustrated with a dotted line to make it easier to understand, but in reality, the dotted line does not exist and the subject is included in the black area.
[0077] Then, the attention area viewpoint information calculation unit 121 calculates the center of gravity p of the two-dimensional coordinates of the pixels belonging to all the attention areas W. c (p cx ,p cy ) and calculate the center of gravity p c The furthest point p in the two-dimensional coordinates of the attention area W is f (p fx ,pfy ) is found.
[0078] Next, the attention area viewpoint information calculation unit 121 calculates the joint points of the subject calculated by the capture area determination unit 110 using the camera C for acquiring the model. M Using the 2D coordinates in the image coordinate system, all (M) joint points p j (j=1,2,…,M) from the center of gravity p c The joint point closest to the target object on the 2D coordinate system is obtained, and the 3D coordinate of that joint point is calculated as p n (p nx ,p ny ,p nz )
[0079] Then, the attention area viewpoint information calculation unit 121 calculates the viewpoint information of the model acquisition camera C in the world coordinate system (studio coordinate system). M Position (t x ,t y ,t z ) and the position of the articulation point (p nx ,p ny ,p nz ) and the distance d n is calculated using the following formula (9).
[0080]
number
[0081] Next, the attention area viewpoint information calculation unit 121 calculates the center of gravity p c (p cx ,p cy ) from the image coordinate system to the 2D coordinates (n cx ,n cy )
[0082]
number
[0083] where f is the camera C for model acquisition M focal length (mm), δ uis the model acquisition camera C M The physical spacing of pixels in the horizontal direction (mm) in the image sensor, δ v is the model acquisition camera C M The physical spacing of pixels in the vertical direction of the image sensor (mm), (c u ,c v ) is the position of the intersection of the optical axis and the image plane in the image coordinate system (image center). M These intrinsic parameters can be obtained by camera calibration.
[0084] Next, the attention area viewpoint information calculation unit 121 calculates the two-dimensional coordinates (n cx ,n cy ) is calculated using equation (9) n Using the three-dimensional coordinates of the camera coordinate system (A cx ,A cy ,A cz )
[0085]
number
[0086] Then, the attention area viewpoint information calculation unit 121 calculates the three-dimensional coordinates (A cx ,A cy ,A cz ) into three-dimensional coordinates in the world coordinate system (W cx ,W cy ,W cz )
[0087]
number
[0088] Here, the main camera C for posture estimation PM (See Figure 2) MThe horizontal tilt of the camera is ψ, the vertical tilt is θ, and the tilt in the camera axis direction is φ. M The position of the origin of the world coordinate system (studio coordinate system) as seen from the camera coordinate system is (t x ,t y ,t z ) As a result, the attention area viewpoint information calculation unit 121 calculates the center of gravity p c The gaze point W in the world coordinate system corresponds to c The three-dimensional coordinates (W cx ,W cy ,W cz ) is calculated. Similarly, the attention area viewpoint information calculation unit 121 calculates the center of gravity p c The furthest point p in the two-dimensional coordinates of the attention area W is f Regarding the farthest point W in the world coordinate system, f The three-dimensional coordinates (W fx ,W fy ,W fz ) is calculated.
[0089] Then, the attention area viewpoint information calculation unit 121 calculates the attention point W shown in FIG. c and the farthest point W f The distance between the virtual camera C and the V The radius r of the area to be photographed is calculated using the following equation (13).
[0090]
number
[0091] Then, the attention area viewpoint information calculation unit 121 calculates the viewpoint information of the virtual camera C shown in FIG. V Vertical viewing angle θ h is calculated using the following formula (14).
[0092]
number
[0093] where d c is the gaze point Wc and virtual camera C V The distance (constant) between the virtual camera C and the camera is set in advance. V The horizontal angle (horizontal angle) θ is the viewing angle of p is the corresponding model acquisition camera C M is equal to the horizontal gradient ψ.
[0094] In this way, the attention area viewpoint information calculation unit 121 calculates the attention point W as shown in FIGS. c (W cx ,W cy ,W cz ) and vertical viewing angle θ h and the viewpoint angle θ p and are calculated. Note that the viewpoint angle θ p is parallel to the XZ plane and is the virtual camera C V When facing the positive direction of the Z axis, it is set to 0 degrees. The attention area viewpoint information calculation unit 121 calculates the viewpoint information of all the model acquisition cameras C M In the step S100, the gaze point, the viewing angle, and the viewpoint angle are calculated for each of the above-mentioned judgment items, and are output to the selection unit 13 as viewpoint information.
[0095] The selection unit 13 displays selection buttons of a graphical user interface (GUI) for selecting a part of interest and an area of interest on the display screen via the display unit 14, and accepts the selection. The selection unit 13 changes the display state based on the viewpoint information of the virtual camera for each target part calculated by the target part virtual camera specification unit 11, and generates and displays a selection button. Furthermore, the selection unit 13 generates and displays selection buttons by changing the display state based on the viewpoint information of the virtual camera for each attention area calculated by the attention area virtual camera identification unit 12.
[0096] An example of selection buttons displayed on the screen by the selection unit 13 will now be described with reference to Fig. 9. Fig. 9 shows a selection button row BL in which selection buttons displayed on the screen are arranged.
[0097] First, the featured part P AT Select button B corresponding to PT This article explains: The selection unit 13 selects the target part P AT Select button B corresponding to PT As shown in Figure 4, the selection buttons B correspond to the target parts A1, ..., A7. PT For example, the "Full Figure" button corresponds to the full body part A1 in Figure 4, and the "Body" button corresponds to the torso part A5 in Figure 4. Similarly, each button below corresponds to a part of interest in Figure 4.
[0098] The selection unit 13 changes the display state according to the number of parts (joint points) that can be acquired for each target part, and selects the selection button B PT Display. For example, the left leg is made up of the left shoulder P8, left elbow P9, and left wrist P1 shown in FIG. 10 , left hand P 11 , left thumb P 12 , left hand P 13 However, if not all of these parts have been acquired, the virtual camera viewpoint information is not output from the target part virtual camera identification unit 11, and the virtual camera cannot be identified. In this case, the selection unit 13 displays the corresponding selection button ("Left Leg" button) in light gray (grays out), for example, so that the button cannot be pressed.
[0099] Furthermore, when all of the parts (joint points) that make up the part of interest have been acquired, that is, when the viewpoint information of the virtual cameras corresponding to all of the parts has been output from the virtual camera identification unit 11 for the part of interest, the selection unit 13 makes the selection button for the corresponding part of interest pressable.
[0100] Furthermore, if some of the parts (joint points) that make up the part of interest cannot be obtained, that is, if the virtual camera identification unit 11 for the part of interest outputs viewpoint information of a virtual camera corresponding to one or more parts that are less than the maximum number of parts that make up the part of interest, the selection unit 13 makes the selection button for the corresponding part of interest pressable, and displays the selection button in a specific color or pattern, for example, red, as a warning, since there is a high possibility that a defect will occur in the 3D model.
[0101] Next, attention area A NT Select button B corresponding to AR This article explains: The selection unit 13 selects the attention area A. NT Select button B corresponding to AR As a result, the selection button B corresponding to the attention area that can be determined by the attention area determination unit 120 is AR For example, in Fig. 9, the selection button B corresponding to each of the attention areas for the judgment items of blown out highlights, color cast, information not available, and pixel feature insufficiency is displayed. AR An example is shown in which the "Blown Out White" button, "Color Cast" button, "Cannot Acquire" button, and "Insufficient Features" button are displayed. Also, each selection button B AR Next to it is the model acquisition camera C M Switch button B to switch between virtual cameras corresponding to CG is displayed.
[0102] For example, when one or more pieces of viewpoint information corresponding to a judgment item (for example, blown-out highlights) are output from the attention area virtual camera specifying unit 12, the selecting unit 13 selects the selection button B AR ("Blown-out highlights" button) can be pressed, and because there is a high possibility of defects in the 3D model, a specific color or pattern, for example, red, is used to warn users. AR Display. At this time, if the viewpoint information is two or more, the selection unit 13 selects the corresponding switching button B CG When there is only one viewpoint, the corresponding switch button B CGis displayed in light gray, for example, to indicate that the button is not pressed.
[0103] Furthermore, when the attention area virtual camera specifying unit 12 does not output viewpoint information corresponding to the judgment item (for example, lack of feature amount), the selecting unit 13 selects the selection button B AR ("insufficient features" button) is displayed in light gray, for example, and the button is set to a state in which it cannot be pressed. At this time, the selection unit 13 selects the corresponding switching button B CG For example, the button is displayed in light gray, and is in a state where it cannot be pressed. The selection unit 13 outputs the generated various selection buttons to the display unit 14, which displays them on the screen as a selection button row BL. Furthermore, the selection unit 13 outputs the viewpoint information corresponding to the selected button to the display unit 14.
[0104] The display unit 14 displays a GUI screen including the selection buttons generated by the selection unit 13 and a 3D model virtually captured by a virtual camera corresponding to the specified viewpoint information. Fig. 10 shows an example of a GUI screen D displayed by the display unit 14. The GUI screen D has a selection button row BL and a confirmation image area D that displays the image of the 3D model taken by the virtual camera selected by the selection button. CV and In this example, the GUI screen D further includes an overall image of the 3D model M, a "rewind frame" button, and a "forward frame" button for switching the image frame by frame.
[0105] The display unit 14 includes a selection button B for the part of interest. PT When the mouse is pressed, the corresponding selection button B PT The image area D is a virtual image captured from a position a predetermined distance away from the fixation point at the fixation point and viewing angle of the virtual camera specified by the viewpoint information input from the selection unit 13 in response to the fixation point and viewing angle. CV Display in.
[0106] In this case, the display unit 14 displays, for example, a confirmation image area D CV By dragging the mouse up, down, left, or right, the virtual camera's position is moved on a sphere at a predetermined distance from the center of the gaze point, changing the viewpoint position. Alternatively, the virtual camera can be rotated at a constant angular velocity while being kept horizontal to the XZ plane (see Figure 2) in the studio. This allows the viewer to enlarge and check a part of the generated 3D model.
[0107] The display unit 14 also displays a selection button B for the area of interest. AR When the mouse is pressed, the corresponding selection button B AR The image area D is a virtual image captured from a position a predetermined distance away from the point of interest, with the point of interest, field of view angle, and viewpoint angle (horizontal angle) of the virtual camera specified by the viewpoint information input from the selection unit 13 in response to the CV Display in.
[0108] In this case, the confirmation image area D CV The image displayed on the screen is captured by one model acquisition camera C. M The display unit 14 displays the image of the area of interest selected by the selection button B AR Switch button B corresponding to CG By pressing the button, the camera C for model acquisition will be M Switch to display an enlarged image of the attention area. This allows the observer to use the model acquisition camera C to identify areas in the 3D model that may be missing. M can be checked every time. When the display unit 14 displays a 3D model in real time, the 3D model M and the confirmation image area D CV The images displayed will be updated as they are displayed.
[0109] Furthermore, when display unit 14 displays a previously created 3D model, it uses the "rewind frame" button and the "forward frame" button to move the input 3D model forward or backward in frame units. In this case, display unit 14 also moves the posture estimation camera image input to posture information estimation unit 10 forward or backward in frame units, and moves the model acquisition camera image input to attention area virtual camera identification unit 12 forward or backward in frame units.
[0110] With the above-described configuration, the 3D model monitoring device 1 can monitor a 3D model created using volumetric capture technology by focusing on any specified part of the 3D model. The 3D model monitoring device 1 can also display areas that indicate possible defects in the 3D model.
[0111] The three-dimensional model monitoring device 1 can be operated by a program (three-dimensional model monitoring program) that causes a computer (not shown) to function as each of the above-mentioned units.
[0112] [Operation of the 3D model monitoring device] Next, the operation of the 3D model monitoring device 1 according to the embodiment of the present invention will be described with reference to Figures 11 to 13 (see Figure 1 for the configuration as appropriate). Note that here, the 3D model is a created model, and the posture estimation camera video and model acquisition camera video are recorded videos that were used when the 3D model was created.
[0113] In step S1, the three-dimensional model monitoring device 1 sets an initial value "1" to a variable f, which is a variable f representing the number of frames. In step S2, the three-dimensional model monitoring device 1 inputs the f-th frame of the created 3D model. In step S3, the three-dimensional model monitoring device 1 inputs the f-th frame of the posture estimation camera image captured by the posture estimation camera in the studio when creating the 3D model input in step S1. In the example of FIG. 2, the posture estimation main camera C PM and sub-camera C for posture estimation PS The fth frame of footage shot with the two cameras is input.
[0114] In step S4, the three-dimensional model monitoring device 1 uses the model acquisition camera C in the studio when creating the 3D model input in step S1. M The fth frame of the model acquisition camera image is input. M There are multiple cameras, and the fth frame of video captured by each camera is input.
[0115] In step S5, the posture information estimation unit 10 acquires, as posture information, the three-dimensional coordinates of the joint points of the subject from the frame input in step S3 using a posture estimation algorithm. In step S6, the target part virtual camera specifying unit 11 specifies a virtual camera that virtually photographs the target part (target part specifying process).
[0116] Here, the operation of the target part identification process in step S6 will be described in detail with reference to FIG. In step S60, the capture area determination unit 110 determines whether the joint points of the subject, which are the posture information estimated in step S5, are located at the positions of the respective model acquisition cameras C M (See Figure 2) is determined to be inside or outside the range to be photographed (capture area). In step S61, the target part virtual camera specifying unit 11 sets a variable for counting the number of target parts to i, and sets the initial value of the variable i to 1. Note that the target parts are assumed to be numbered in advance starting from 1.
[0117] In step S62, the target part virtual camera specifying unit 11 starts a loop for calculating viewpoint information of a virtual camera that virtually photographs the target part. In step S63, the target part virtual camera specifying unit 11 determines whether or not all the parts (joint points) of the target part i (i-th target part) are outside the capture area based on the determination result in step S62. If all parts are outside the capture area (Yes in step S63), the target part virtual camera specifying unit 11 proceeds to step S66; otherwise (No in step S63), the operation proceeds to step S64.
[0118] In step S64, the part-of-interest viewpoint information calculation unit 111 calculates the gaze point and the viewing angle for photographing the part of interest i as viewpoint information. In step S65, the part-of-interest viewpoint information calculation unit 111 stores the viewpoint information of the part-of-interest i in a storage unit (not shown).
[0119] In step S66, the target part virtual camera specifying unit 11 determines whether i is the number of target parts, which is, for example, 7 for the face part A2 in FIG. If i is the number of parts of interest (Yes in step S66), the part of interest virtual camera specifying unit 11 proceeds to step S68; otherwise (No in step S66), the operation proceeds to step S67.
[0120] In step S67, the target part virtual camera specifying unit 11 adds "1" to i, and returns the operation to step S63. In step S68, the target part virtual camera specifying unit 11 ends the loop for calculating the viewpoint information of the virtual camera for the target part. Returning to FIG. 11, the explanation will be continued.
[0121] In step S7, the attention area virtual camera specifying unit 12 specifies a virtual camera that virtually captures the attention area (attention area specifying process).
[0122] Here, the operation of the attention area identification process in step S7 will be described in detail with reference to FIG. In step S70, the attention area virtual camera identification unit 12 sets a variable j for counting the number of attention areas, and sets the variable j to an initial value of 1. The number of attention areas is the number of attention area determination items (bloated highlights, color cast, inability to acquire information, insufficient pixel features, etc.), and each item is numbered in advance starting from 1. In step S71, the attention area virtual camera specifying unit 12 starts a loop for calculating viewpoint information of a virtual camera that virtually captures the attention area.
[0123] In step S72, the attention area virtual camera identification unit 12 sets a variable k for counting the number of model acquisition cameras, and sets the initial value of the variable k to 1. Note that the model acquisition cameras are assumed to be numbered in advance starting from 1.
[0124] In step S73, the attention area virtual camera specifying unit 12 starts a loop for calculating viewpoint information of a virtual camera that virtually captures the attention area j. In step S74, the attention area determination unit 120 determines whether or not there is an attention area j in the camera image captured by the k-th model acquisition camera. In step S75, if it is determined that there is an attention area (Yes), the attention area virtual camera specifying unit 12 proceeds to step S76, and if it is determined that there is no attention area (No), the operation proceeds to step S78.
[0125] In step S76, the attention area viewpoint information calculation unit 121 calculates the gaze point, the viewing angle, and the viewpoint angle (horizontal angle) for photographing the attention area j as viewpoint information. In step S77, the attention area viewpoint information calculation unit 121 stores the viewpoint information of the attention area j in a storage unit not shown.
[0126] In step S78, the attention area virtual camera specifying unit 12 determines whether k is the number of cameras for model acquisition. If k is the number of cameras (Yes in step S78), the attention area virtual camera specifying unit 12 proceeds to step S80; otherwise (No in step S78), the operation proceeds to step S79.
[0127] In step S79, the attention area virtual camera specifying unit 12 adds "1" to k, and returns the operation to step S74. In step S80, the attention area virtual camera specifying unit 12 ends the loop for calculating the viewpoint information of the virtual camera for the attention area j.
[0128] In step S81, the attention area virtual camera specifying unit 12 determines whether j is the number of attention areas. If j is the number of attention areas (Yes in step S81), the attention area virtual camera specifying unit 12 proceeds to step S83; otherwise (No in step S78), the operation proceeds to step S82.
[0129] In step S82, the attention area virtual camera specifying unit 12 adds "1" to j, and returns the operation to step S72. In step S83, the attention area virtual camera specifying unit 12 ends the loop for calculating the viewpoint information of the virtual camera for the attention area. Returning to FIG. 11, the explanation will be continued.
[0130] In step S8, the selection unit 13 displays a selection button corresponding to the target part and a selection button corresponding to the attention area on the GUI screen D (see FIG. 10). In step S9, the display unit 14 accepts the pressing of the select button via the GUI screen D.
[0131] In step S10, the display unit 14 displays an image of the 3D model captured by the virtual camera in the confirmation image area D on the GUI screen D based on the viewpoint information of the attention part or attention area selected in step S9. CV Display in.
[0132] In step S11, the display unit 14 determines whether pressing the "forward frame" button or the "rewind frame" button on the GUI screen D will advance the frame of the 3D model (next frame) or go back the frame (previous frame). Here, if the "Next Frame" button is pressed (next frame), the three-dimensional model monitoring device 1 proceeds to step S12, and if the "Previous Frame" button is pressed (previous frame), the three-dimensional model monitoring device 1 proceeds to step S13. Note that if f=1, the "Previous Frame" button cannot be pressed. Also, if neither the "Next Frame" button nor the "Previous Frame" button is pressed (no change), the three-dimensional model monitoring device 1 returns to step S9.
[0133] In step S12, the three-dimensional model monitoring device 1 adds "1" to f, and proceeds to step S14. In step S13, the three-dimensional model monitoring device 1 subtracts "1" from f, and returns the operation to steps S2, S3, and S4.
[0134] In step S14, the three-dimensional model monitoring device 1 determines whether f is the number of frames of the 3D model plus 1. The number of frames of the 3D model is the total number of frames of the 3D model to be used. Here, if f is the number of frames of the 3D model plus 1 (Yes in step S14), the 3D model monitoring device 1 ends the operation, otherwise (No in step S14), the operation returns to steps S2, S3, and S4.
[0135] By the above operations, the three-dimensional model monitoring device 1 can enlarge and present the attention part and attention area of the 3D model. This allows the observer to monitor the defect state of the generated 3D model with simple operations.
[0136] Although the embodiment of the present invention has been described above, the present invention is not limited to this embodiment. For example, here, the attention parts and attention areas of the 3D model are photographed with a virtual camera and displayed. However, it is also possible to display an image photographed from the recommended viewpoint of the 3D model. In that case, for example, as shown in FIG. 14, a recommended viewpoint information calculation unit 15 may be added to the configuration of the three-dimensional model monitoring device 1 (see FIG. 1) to form a three-dimensional model monitoring device 1B.
[0137] The recommended viewpoint information calculation unit 15 calculates viewpoint information of a virtual camera that is recommended when displaying a 3D model. The recommended viewpoint for viewing a 3D model can be determined, for example, based on the orientation of the subject's body. An example of a method for calculating the recommended viewpoint will be described below.
[0138] The recommended viewpoint information calculation unit 15 calculates the three-dimensional coordinate point L(L x ,L y ,L z ) and the three-dimensional coordinate point R(R x ,R y ,R z ) and the two-dimensional vector RL = (L x -R x ,L z -R z ) is found. Next, the recommended viewpoint information calculation unit 15 calculates a vector F=(F x ,F z ) is calculated using the following formula (15).
[0139]
number
[0140] where θ f =-90 degrees. The recommended viewpoint information calculation unit 15 calculates the angle between this vector F and the positive direction of the X axis. Here, the angle is calculated by dividing the vector F by the coordinate (F x ,F z ) is the angle (deflection angle) that the vector to the left makes with the positive direction of the X axis, and its range is from 0 to 360 degrees. The angle calculated here is the horizontal angle (viewpoint angle) θ of the virtual camera in virtual space. p This becomes: In addition, θ f By changing the value of , the recommended viewpoint can be set to a viewpoint from an angle other than the front of the subject's body.
[0141] In addition, the recommended viewpoint information calculation unit 15 calculates the gaze point W as the center of gravity of all the joint points. c (W cx ,W cy ,W cz ) is calculated. Further, the recommended viewpoint information calculation unit 15 calculates the difference r between the Y coordinate of the point farthest from the gaze point when viewed in the Y-axis direction among the joint points of the subject and the Y coordinate of the gaze point, and calculates the vertical viewing angle θ h Calculate.
[0142]
number
[0143] where d c is the gaze point W c and the distance (constant) between the virtual camera and the object, which is set in advance. The recommended viewpoint information calculation unit 15 outputs the gaze point, the viewing angle, and the viewpoint angle to the selection unit 13 as viewpoint information of the recommended viewpoint. 9, the selection unit 13 displays a button (not shown) for selecting a recommended viewpoint. When the button is pressed, the display unit 14 displays an image of a 3D model virtually captured at the gaze point, viewing angle, and viewpoint angle calculated by the recommended viewpoint information calculation unit 15.
[0144] In this example, the target part and the attention area of the 3D model are photographed by a virtual camera and displayed. However, it is also possible to monitor only the target part. In that case, for example, as shown in FIG. 15, the attention area virtual camera specifying unit 12 may be omitted from the configuration of the three-dimensional model monitoring device 1 (see FIG. 1) to form a three-dimensional model monitoring device 1C.
[0145] In this case, the capture area determination unit 110 determines whether the model acquisition camera C M When determining whether or not the object (see FIG. 2) is inside or outside the range to be photographed (capture area), a preset range may be used as the determination criterion. The capture area varies depending on the shape of the volumetric studio, the camera placement, the camera performance, etc., but is set, for example, to a range of radius r meters and height h meters from the center of the studio. In this case, the capture area determination unit 110 determines each part P of the subject, which is a three-dimensional point. j (See Figure 3) jx ,P jy ,P jz ) but P jx 2 +P jz 2 ≦r 2 , and P iy If ≦y, it is determined that the point is inside the capture area, and if not, it is determined that the point is outside the capture area. [Explanation of symbols]
[0146] 1. 3D model monitoring device 10 Posture information estimation section 11 Virtual camera identification for target parts 110 Capture area determination unit 111 Part of interest viewpoint information calculation unit 12 Virtual camera identification unit for attention area 120 Attention area determination unit 121 Attention area viewpoint information calculation unit 13 Selection section 14 Display section 15 Recommended viewpoint information calculation unit
Claims
1. A three-dimensional model monitoring device for monitoring a three-dimensional model created by volumetric capture technology, comprising: a posture information estimation unit that estimates a three-dimensional coordinate position of each predetermined part of the subject as posture information; a capture area determination unit that determines whether a predetermined set of the parts is a part of interest and whether each part is inside or outside a capture area of all model acquisition cameras in a volumetric studio based on the posture information, for each part of interest; a part-of-interest viewpoint information calculation unit that calculates, for each part of interest, viewpoint information of a part-of-interest virtual camera that virtually photographs the part of interest; a selection unit that displays a button for selecting the target part on a screen together with the result of the inside / outside determination and accepts the selection; a display unit that virtually photographs the three-dimensional model with the virtual camera for the part of interest based on the viewpoint information corresponding to the part of interest selected by the selection unit and displays the photograph on a screen; A three-dimensional model monitoring device comprising:
2. an attention area determination unit that determines, for each of the model acquisition cameras, a pixel set that is a missing pixel of a three-dimensional model based on a plurality of determination items from the camera image as an attention area; an attention area viewpoint information calculation unit that calculates, for each of the attention areas, viewpoint information of a virtual camera for the attention area that virtually photographs the attention area; the selection unit further displays buttons for selecting the judgment item and the model acquisition camera on a screen and accepts the selection; The three-dimensional model monitoring device of claim 1, characterized in that the display unit virtually photographs the three-dimensional model with the attention area virtual camera based on the viewpoint information of the corresponding attention area, depending on the selection of the judgment item and the model acquisition camera, and displays the photograph on a screen.
3. The three-dimensional model monitoring device described in claim 1, characterized in that the part of interest viewpoint information calculation unit calculates, for each part of interest, the gaze point which is the center of gravity of each part within the capture areas of all the model acquisition cameras, and the field of view angle of the virtual camera for the part of interest which virtually photographs the area from the gaze point to the farthest part at a predetermined distance, as the viewpoint information of the part of interest.
4. The three-dimensional model monitoring device described in claim 1, characterized in that the part of interest viewpoint information calculation unit calculates, for each part of interest, the gaze point which is the center of gravity of each part for parts within the capture areas of all the model acquisition cameras, and the distance between the gaze point and the virtual camera for the part of interest, which virtually photographs the area from the gaze point to the farthest part at a predetermined field of view angle, as the viewpoint information of the part of interest.
5. The three-dimensional model monitoring device according to claim 3 or 4, characterized in that the display unit virtually photographs the three-dimensional model with the virtual camera for the part of interest, displays the photograph on the screen, and, in response to an input operation, moves the virtual camera for the part of interest on a sphere having a radius equal to the predetermined distance and a focus point that is the center of gravity of each part as its center.
6. The three-dimensional model monitoring device described in claim 2, characterized in that the attention area viewpoint information calculation unit calculates the three-dimensional coordinate position of the center of gravity of the attention area as the gaze point of the attention area, using the distance between the three-dimensional coordinate position of the part whose center of gravity of the image coordinate system of the attention area and its position projected onto the image coordinate system are closest to the three-dimensional coordinate position of the model acquisition camera as the depth, and calculates the field of view angle and viewpoint angle of a virtual camera for the attention area that virtually photographs the area from the gaze point of the attention area to the farthest pixel of the attention area at a predetermined distance as the viewpoint information of the attention area.
7. A program for causing a computer to function as the three-dimensional model monitoring device according to claim 1.