Image processing device, image processing method and program
The image processing device effectively generates virtual viewpoint images by identifying and highlighting moving parts of an object as afterimages, addressing the limitation of existing methods that include motionless parts in the representation.
Patent Information
- Application Number
- JP2021126787
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-02
- Publication Date
- 2025-10-23
- Estimated Expiration
- 2041-08-02
AI Technical Summary
Existing methods fail to express only the moving parts of an object as an afterimage in virtual viewpoint images, as they often include motionless parts in the representation.
An image processing device that acquires shape data of an object over time, determines specific moving parts, and generates a virtual viewpoint image by combining projection images of these parts with background images, using techniques like alpha blending or color changes to represent motion as an afterimage.
Enables the creation of virtual viewpoint images where only the moving parts of an object are clearly represented as an afterimage, enhancing the visual representation of object motion.
Smart Images

Figure 0007759160000001 
Figure 0007759160000002 
Figure 0007759160000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for generating a virtual viewpoint image based on a plurality of captured images. [Background technology]
[0002] There is a technology that generates a virtual viewpoint image that shows the view from a virtual viewpoint by using a plurality of captured images obtained by placing a plurality of imaging devices at different positions and synchronously capturing images of a subject (object).Patent Document 1, which applies this virtual viewpoint image technology, discloses a method of generating virtual viewpoint images of a specific object viewed from a desired virtual viewpoint at different times and displaying them in a superimposed manner to show the trajectory of the object. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-13095 [Non-patent literature]
[0004] [Non-Patent Document 1] Koichi Ogawara, Xiaolu Li, Katsushi Ikeuchi, "Estimation of Human Motion Using a Flexible Deformable Model with Joint Structure," Proc. of MIRU2006, pp.994-999,2006. Summary of the Invention [Problem to be solved by the invention]
[0005] There are cases where it is desired to express only the moving part of an object (for example, only the arm of a person) as an afterimage in a virtual viewpoint image. In this regard, the method of Patent Document 1 above expresses the motionless parts other than the arm as an afterimage, and it is not possible to obtain a virtual viewpoint image in which only the moving arm is expressed as an afterimage. [Means for solving the problem]
[0006] The image processing device according to the present disclosure includes an acquisition means for acquiring shape data of an object corresponding to each of a plurality of times during a period in which images are continuously captured by a plurality of imaging devices, a determination means for determining a specific portion that is a part of the object, and a generation means for generating a virtual viewpoint image corresponding to a virtual viewpoint using the shape data of the object at a reference time among the plurality of times and the shape data of the specific portion of the object at a specific time among the plurality of times that is different from the reference time. the generating means generates the virtual viewpoint image by combining a first projection image obtained by projecting shape data of the object at the reference time onto the virtual viewpoint and a second projection image obtained by projecting shape data of a specific part of the object at the specific time onto the virtual viewpoint. The present invention is characterized by having the following. [Effects of the Invention]
[0007] According to the technology of the present disclosure, it is possible to obtain a virtual viewpoint image in which a part of an object is expressed like an afterimage. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of an image processing system. [Figure 2] FIG. 1 is a diagram showing an example of the hardware configuration of an image processing apparatus. [Figure 3] FIG. 2 is a diagram showing an example of the software configuration of the image processing apparatus. [Figure 4] 5 is a flowchart showing the flow of a virtual viewpoint image generation process according to the first embodiment. [Figure 5] 10(a) and 10(b) are diagrams for explaining detection of a specific site. [Figure 6] 10(a) to 10(d) are diagrams showing projected images of a 3D model. [Figure 7] 10A and 10B are diagrams illustrating a synthesis process of projected images. [Figure 8] 10(a) to 10(c) are diagrams showing an example of afterimage representation. [Figure 9] 10 is a flowchart showing the flow of a virtual viewpoint image generation process according to the second embodiment. [Figure 10] FIG. 10 is a diagram showing an example of a UI screen for specifying a region. [Figure 11] 10A and 10B are diagrams for explaining adjustments during synthesis. [Figure 12] 10A and 10B are diagrams showing an example of afterimage representation. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, modes for carrying out the present embodiments will be described with reference to the drawings, etc. Note that the following embodiments do not limit the technology of the present disclosure, and all of the configurations described in the following embodiments are not necessarily essential to the means for solving the problems.
[0010] [Embodiment 1] First, a brief overview of a virtual viewpoint image will be provided. A virtual viewpoint image is an image that represents the view from a virtual camera viewpoint (virtual viewpoint) that is different from the actual camera viewpoint, and is also called a free viewpoint image. The virtual viewpoint is set by a method such as a user directly specifying the virtual viewpoint by operating a controller, or by selecting from a plurality of pre-set virtual viewpoint candidates. Note that virtual viewpoint images include both moving images and still images, but the following embodiment will be described using moving images as an example. When generating a moving virtual viewpoint image, the virtual viewpoint may be fixed, or may change to follow the movement of an object.
[0011] <System configuration> 1 is a diagram showing an example of the configuration of an image processing system for generating a virtual viewpoint image according to this embodiment. The image processing system 100 includes a plurality of image capturing devices (cameras) 101, an image processing device 102, a controller 103, and a display device 104. In the image processing system 100, the image processing device 102 generates a virtual viewpoint image based on a plurality of captured images obtained by synchronous imaging of the plurality of cameras 101, and displays the generated virtual viewpoint image on the display device 104.
[0012] The cameras 101 are installed to surround the object, and each camera 101 captures an image of the object in a synchronized time sequence. However, when installation locations are limited, such as in a studio or concert hall, the multiple cameras 101 are installed in only a portion of the imaging target area. Each camera 101 is realized, for example, by a digital video camera equipped with a video signal interface such as a serial digital interface (SDI). Each camera 101 adds time information such as a time code to the output video signal and transmits it to the image processing device 102. At this time, a set of imaging parameters, including the camera's three-dimensional position (x, y, z), imaging direction (pan, tilt, roll), angle of view, and resolution, is also transmitted. The imaging parameter set is calculated and stored in advance for each camera 101 by performing known camera calibration.
[0013] The image processing device 102 generates a virtual viewpoint image in which a part of an object is expressed as an afterimage based on a plurality of captured images obtained by synchronously capturing images from a plurality of cameras 101. Specifically, it generates shape data of the object in the foreground, sets a moving part (specific part) of the object that is the target of afterimage expression, projects the shape data of the object onto the virtual viewpoint, and synthesizes the projected image. The functions of the image processing device 102 will be described in detail later.
[0014] The controller 103 is a control device operated by a user to cause the image processing device 102 to generate a virtual viewpoint image. The operator performs various settings and data inputs required for generating a virtual viewpoint image via input devices such as a joystick or keyboard provided in the controller 103. Specifically, the operator specifies the three-dimensional position of the virtual viewpoint (virtual camera) in the imaging space, the line of sight direction, the angle of view, the resolution, and a time code required for image generation. The time represented by the time code here includes the start and end times of virtual viewpoint image generation within the imaging period of the plurality of captured images, the time of a key frame (reference time for afterimage representation), and the like. Information for generating a virtual viewpoint image set based on the user's specifications (hereinafter referred to as "virtual viewpoint information") is transmitted to the image processing device 102.
[0015] The display device 104 acquires and displays image data (UI screen data for a graphical user interface and virtual viewpoint image data) sent from the image processing device 102. The display device 104 is realized by, for example, a liquid crystal display, a projector, a head-mounted display, or the like.
[0016] <Hardware configuration> 2 is a diagram showing an example of the hardware configuration of the image processing device 102. The image processing device 102, which is an information processing device, includes a CPU 211, a ROM 212, a RAM 213, an auxiliary storage device 214, an operation unit 215, a communication I / F 216, and a bus 217.
[0017] The CPU 211 controls the entire image processing device 102 using computer programs and data stored in the ROM 212 or the RAM 213, thereby realizing each function of the image processing device 102. The image processing device 102 may have one or more dedicated hardware or a GPU (Graphics Processing Unit) different from the CPU 211. At least a part of the processing by the CPU 211 may be performed by the GPU or dedicated hardware. Examples of dedicated hardware include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor).
[0018] The ROM 212 stores programs that do not require modification. The RAM 213 temporarily stores programs and data supplied from the auxiliary storage device 214, and data supplied from the outside via the communication I / F 217. The auxiliary storage device 214 is configured, for example, with a hard disk drive or the like, and stores various data such as image data and volume data.
[0019] The operation unit 215 is composed of, for example, a keyboard, a mouse, etc., and receives operations from a user to input various instructions to the CPU 211. The CPU 211 operates as a display control unit that controls the display device 104 and an operation control unit that controls the operation unit 215. The communication I / F 216 is used for communication with devices external to the image processing device 102. For example, when the image processing device 102 is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 216. When the image processing device 102 has a function for wireless communication with an external device, the communication I / F 216 is equipped with an antenna.
[0020] The bus 217 connects the various components of the image processing device 102 to transmit information. In this embodiment, the controller 103 and the display device 104 are provided as external devices, but they may also be provided as internal functional components of the image processing device 102.
[0021] <Software configuration> 3 is a diagram showing an example of the software configuration of the image processing device 102. The image processing device 102 has a data acquisition unit 300, a model generation unit 301, a model analysis unit 302, and a virtual viewpoint image generation unit 303. The virtual viewpoint image generation unit 303 has a projection unit 304 and a synthesis unit 305. The function of each unit will be described below.
[0022] The data acquisition unit 300 acquires various data used to generate a virtual viewpoint image, specifically, a plurality of captured image data captured synchronously by each camera 101, an imaging parameter set for each camera 101, and virtual viewpoint information. The acquired data is stored in the auxiliary storage device 214.
[0023] The model generation unit 301 generates shape data (hereinafter referred to as a "foreground model") representing the three-dimensional shape of a foreground object, such as a stage performer, using, for example, a volume intersection method, based on multiple captured images and the imaging parameter sets of each camera 101. In this embodiment, the foreground model is represented by a textured polygon mesh or a three-dimensional point cloud with each point colored, and is generated in a general-purpose format file, such as Stanford PLY or Wavefront OBJ. Note that the foreground model may also be shape data without color information. In this case, the projection unit 304, described below, performs a process of applying colors corresponding to the virtual viewpoint (texture mapping). The generated foreground model is stored in the auxiliary storage device 214.
[0024] The model analysis unit 302 analyzes the shape changes (movement) of the foreground model generated by the model generation unit 301 over multiple time periods, detects parts of the foreground model that have undergone significant changes, and determines specific parts that will be the subject of afterimage representation.
[0025] The projection unit 304 projects a foreground model onto a virtual viewpoint based on virtual viewpoint information input from the controller 103, generating a projection image corresponding to the virtual viewpoint. Here, the projection target includes not only the foreground model described above, but also shape data corresponding to a specific portion of the foreground model (hereinafter referred to as a "partial model") and shape data of a background object (hereinafter referred to as a "background model"). The background model is shape data with color information that represents the three-dimensional shape of a background object such as stage equipment, and is generated in advance and stored in the auxiliary storage device 214. The background model may be generated using design data such as CAD, or shape and color data scanned with a laser scanner or the like. Alternatively, the background model may be generated using computer vision techniques such as Structure from Motion from images captured from multiple viewpoints.
[0026] The synthesis unit 305 synthesizes the projected images obtained by the projection unit 304 projecting each 3D model (foreground model, partial model, background model) onto the virtual viewpoint, and generates a virtual viewpoint image in which a specific part is expressed like an afterimage. Data of the generated virtual viewpoint image is transmitted to the display device 104.
[0027] The above-mentioned functional units may be distributed among multiple image processing devices. For example, a first image processing device may have the data acquisition unit 300 and the model generation unit 301, and a second image processing device may have the model analysis unit 302 and the virtual viewpoint image generation unit 303.
[0028] <Virtual viewpoint image generation flow> Next, a flow of generating a virtual viewpoint image in the image processing device 102 will be described with reference to the flowchart in Fig. 4. A series of processes shown in the flowchart in Fig. 4 is realized by the CPU 211 reading out a control program stored in the ROM 212 or the auxiliary storage device 214, expanding it in the RAM 213, and executing it. In the following description, the symbol "S" means a step.
[0029] In S401, the data acquisition unit 300 acquires a plurality of captured image data for a continuous period (for example, 5 seconds), an imaging parameter set for each camera 101, and virtual viewpoint information. The acquired data is expanded and stored in the RAM 213.
[0030] In S402, the model generation unit 301 uses the captured image data and the imaging parameter set acquired in S401 to generate foreground models corresponding to each of a plurality of times in the past or future from a reference time t of interest within the continuous period. Here, the time of interest is, for example, a time corresponding to a specific key frame designated by a time code within the acquired continuous period. Furthermore, the plurality of times in the past or future for which foreground models are to be generated may be set in advance or may be designated in the virtual viewpoint information.
[0031] In S403, the model analysis unit 302 analyzes the change over time in the shape of the foreground model generated in S402 and determines a specific portion to be processed for afterimage rendering. Specifically, based on the foreground model generated for multiple time points, a portion that has undergone significant movement over time is detected, and this portion is determined as the specific portion to be processed for afterimage rendering. FIG. 5(a) is a diagram illustrating the detection of a specific portion, showing foreground models corresponding to different time points superimposed on the y-axis in the same space. FIG. 5(b) is a diagram illustrating the relationship between time points. In FIG. 5(b), a thin bidirectional arrow 501 indicates the range in which a virtual viewpoint image is generated, and a thick bidirectional arrow 502 indicates the range to be rendered as an afterimage. Time tn is a time within the range 502 other than the reference time t, and the variable n is an arbitrary integer corresponding to the above-mentioned "multiple times in the past or future." For example, "n = 1" represents one time period past the reference time t, "n = 2" represents two time periods past the reference time t, and "n = -1" represents one time period future the reference time t. In Figure 5(a), the dashed-dotted line represents the left arm of the foreground model at time t-1, and the dashed-two-dotted line represents the left arm of the foreground model at time t-2. This means that only the left arm moves from time t-2 to the reference time t. This means that only the part of the shape represented by the foreground model that corresponds to the left arm changes significantly between different times, as indicated by range 502. In this way, by calculating the amount of change in the foreground model between different times, moving parts of the foreground object can be detected. For example, if the foreground model is in polygon mesh format, the distance between polygons can be used to calculate the amount of change. That is, from each vertex of the foreground model at one time to each vertex of the foreground model at another time, the nearest point is searched for in the normal direction, and the point of contact with the mesh of the foreground model at another time is determined as the nearest point. The distance between the vertex and the nearest point is then calculated as a signed distance. The signed distance is a distance that is positive if the nearest point from the vertex is outside the mesh, and negative if it is inside. This signed distance is calculated for all vertices of the foreground model at a given time, and vertices whose absolute value is greater than the standard deviation (σ) of the distribution are extracted. The part of the foreground model made up of the extracted vertices is then detected.In the example of FIG. 5(a), each vertex of the part corresponding to the left arm of the foreground model has a large absolute value of the signed distance, and therefore is extracted using the above procedure. A part of the foreground model consisting of the group of vertices extracted in this way is determined as the specific part to be processed for afterimage representation. The above process is then performed for all multiple time points corresponding to the foreground model generated in S402, and the specific part at each time point is determined. Shape data corresponding to the determined specific part (shape data corresponding to a part of the foreground model; hereinafter referred to as the "partial model") is stored in RAM 213 separately from the foreground model.
[0032] In S404, the projection unit 304 performs a process of projecting the foreground model at the reference time, the partial model of the specific part determined in S403, and a previously prepared background model based on virtual viewpoint information. The projection process may use well-known model-based rendering or image-based rendering. Figure 6(a) shows a projected image of the foreground model 501 at the reference time s, while Figure 6(b) and Figure 6(c) show projected images of the partial model of the left arm of the foreground model 501 at times s-1 and s-2, respectively. Figure 6(d) shows a projected image of the background model. These projected images are, for example, four-channel images in which each pixel has an 8-bit RGB value (0 to 255) plus an alpha value. Here, the alpha value is a value that represents transparency, with a maximum value of 255 indicating opacity, a smaller value indicating increased transparency, and a minimum value of 0 indicating complete transparency. In other words, the alpha value of a pixel onto which shape data of each model is projected is greater than 0, and the alpha value of a pixel onto which nothing is projected is 0. Therefore, for the projected image of the background model, which is completely opaque, and the projected image of the foreground model at current time s, the alpha value of each pixel is set to 255. On the other hand, for the projected image of the partial model, which is to be transparent to express an afterimage, different alpha values are set, such as 255 × 0.6 = 153 for time s-1 and 255 × 0.3 = 77 for time s-2. In this case, "0.6" and "0.3" are coefficients for controlling the transmittance, and values less than 1.0 corresponding to the difference from the reference time are entered. That is, in the examples of (b) and (c) of Figure 6 above, the transmittance increases (the alpha value decreases) as the corresponding specific time moves further in the past than the reference time. The transmittance control method described here is merely an example; for example, the transmittance may increase as the specific time corresponding to the partial model moves further in the future than the reference time.
[0033] In S405, the composition unit 305 composes all of the projection images generated in S404 to generate a virtual viewpoint image. Fig. 7(a) is a diagram showing how the four projection images shown in Figs. 6(a) to 6(d) are alpha-composited. By layering and alpha-compositing the four projection images 701 to 704 as shown in Fig. 7(a), a virtual viewpoint image at the reference time t is obtained, in which the left arm is expressed like an afterimage, as shown in Fig. 7(b).
[0034] In S406, it is determined whether processing has been completed for all reference times specified in the virtual viewpoint information. If there are any unprocessed reference times, the process proceeds to S407, where the reference time of interest is updated, and the process returns to S402 to continue. At this time, if the virtual viewpoint information specifies that the virtual viewpoint moves over time, a virtual viewpoint image projected at any camera angle can be obtained. If processing has been completed for all reference times, this flow ends.
[0035] The above is the flow of generating a virtual viewpoint image according to this embodiment. Data on the virtual viewpoint image thus obtained is transmitted to the display device 104 and provided for viewing by the user. Note that the target foreground object is not limited to a dynamic object such as a person, but may also be a static object such as a building. For example, based on captured images obtained by continuously capturing images of the construction process of a high-rise building or the like over a certain period of time, newly constructed portions are made opaque and existing building portions are made transparent, thereby obtaining virtual viewpoint images corresponding to each reference time as shown in FIGS. 8(a) to 8(g), which represent daily changes. In this case, for example, the building is captured daily (S401), a foreground model is generated from the captured images (S402), and portions with significant shape changes compared to the previous day (i.e., portions constructed on that day) are detected and identified (S403). Then, a partial model corresponding to the construction portion of the current day and the foreground model of the previous day are projected onto the virtual viewpoint (S404), and alpha compositing is performed using the obtained projected images (S405). In this case, by setting the alpha value to "127" for the current day and "255" for the previous day, a virtual viewpoint image can be obtained that allows the difference from the previous day to be intuitively grasped.
[0036] <Variation 1> In the above-described embodiment, the specific part is determined based on the change in shape of the foreground model between multiple time points, but the method of determining the specific part is not limited to this. That is, prior to determining the specific part, a projection process onto a virtual viewpoint for multiple time points may be performed, differences between the obtained projection images between multiple time points may be calculated, and the part corresponding to the image area with a large difference value may be determined as the specific part. In this case, a partial model corresponding to the determined specific part may be generated from the foreground model, and a projection process of the partial model onto the virtual viewpoint may be performed again to obtain a projection image corresponding to Figure 7(b) or (c), and then the above-described synthesis process may be performed.
[0037] <Variation 2> Instead of alpha blending, which changes the transmittance, afterimages may be expressed by color change processing, for example, by changing the brightness or saturation. In this case, each pixel of the projected image becomes a three-channel RGB projected image without an alpha value. Then, for example, the brightness or saturation of the specific time corresponding to the partial model may be lowered as it moves further in the past than the reference time, or the brightness or saturation may be lowered as it moves further in the future than the reference time. This allows the movement of the object over time to be visually expressed. In addition, highlight processing may be performed to make a specific part stand out at a specific time.
[0038] <Variation 3> In the above-described embodiment, projection images corresponding to multiple times are layered and composited, but in order to accurately represent the occlusion relationship of the foreground object at the multiple times, compositing using depth information may also be performed. That is, when generating projection images for each model in S404, a depth image is also generated that records the distance from the virtual viewpoint to the foreground object for each pixel. Then, when compositing in S405, projection images are selected for each pixel so that the smaller the depth (i.e., the closer the distance to the virtual viewpoint), the higher the layer. This enables a more three-dimensional image representation that appropriately represents the actual occlusion relationship.
[0039] <Variation 4> In the above-described embodiment, the specific region is determined by detecting the movement of the foreground object. However, the specific region may also be determined based on its positional relationship with the specific region in the three-dimensional space where the image was captured. Specifically, a bounding box is set in the three-dimensional space such that the person's shoulders are located at the boundary and all parts of the body except the arms are included within the bounding box. The specific region is determined by detecting the part inside or outside the bounding box. The shape of the bounding box can be any shape, such as a rectangular parallelepiped, cylinder, sphere, or oval sphere, as long as it can be determined as an inside or outside in the three-dimensional space. Furthermore, the position, orientation, and size of the bounding box may be changed every time (every frame). Information such as the shape and size of the bounding box, whether it has changed every time, and whether it is detected as inside or outside can be stored in the auxiliary storage device 214 in advance and read and used as needed. This makes it easier to determine the specific region.
[0040] As described above, according to this embodiment, it is possible to easily obtain a virtual viewpoint image in which the moving parts of the foreground object are expressed by an afterimage.
[0041] [Embodiment 2] In the first embodiment, the specific part to be subjected to afterimage rendering was automatically determined by analyzing the change in the shape of the foreground model over time and detecting parts with large movement. Next, a second embodiment will be described in which the structure of the generated foreground model is analyzed, the results are presented to the user, and the specific part is determined based on the user's explicit instructions. Furthermore, while the first embodiment was based on the assumption that parts other than the specific part do not change over multiple time periods, this embodiment will also describe adjustments when the position, posture, etc. of the foreground object change over multiple time periods. Note that a description of the basic system configuration and other aspects common to the first embodiment will be omitted, and the following description will focus on the differences.
[0042] <Virtual viewpoint image generation flow> A flow of generating a virtual viewpoint image according to this embodiment will be described with reference to the flowchart in Fig. 9. A series of processes shown in the flowchart in Fig. 9 is realized by the CPU 211 reading out a control program stored in the ROM 212 or the auxiliary storage device 214, expanding it in the RAM 213, and executing it. In the following description, the symbol "S" denotes a step.
[0043] 4 according to embodiment 1. That is, captured image data, an imaging parameter set, and virtual viewpoint information for a continuous period are acquired (S901), and then foreground models corresponding to each of a plurality of times are generated (S902).
[0044] In S903, the model analysis unit 302 analyzes the structure of the foreground model generated in S902. A known method may be used for the structural analysis. For example, a known method is to minimize the distance between the surfaces of the reference bone model and the foreground model to superimpose the reference bone model on the foreground object and associate the bone structure (Non-Patent Document 1). Other methods that may be used include the "OpenPose" method, which uses a deep learning model with an input video of the object to infer the bone structure, and motion capture methods. In the case of a foreground model in polygon mesh format, the structural analysis results are obtained as metadata that associates a part name and part ID with each mesh polygon, which is a component of the model. The structural analysis results obtained in this manner are stored in RAM 213.
[0045] In S904, a specific body part is determined based on a user input via a UI screen that reflects the structural analysis results obtained in S903. More specifically, first, data for a UI screen (a UI screen for specifying a specific body part for which afterimage rendering is desired) is transmitted to the display device 104 and displayed on the display device 104. FIG. 10 shows an example of the UI screen for specifying a body part. In the UI screen 1000 of FIG. 10, a foreground model for which the specific body part is to be specified is displayed in the center of the screen, and to the right of the UI screen 1000, a list 1001 of all body parts obtained by the structural analysis and a seek bar 1002 for specifying a desired time from among the multiple times described above are displayed. The user specifies the specific body part for which afterimage rendering is desired by manipulating a pointer 1003 using a mouse or the like. For example, to specify the left arm, the user moves the pointer 1003 over the left shoulder and then drags it to the vicinity of the fingertips of the left hand. For the body part specified by the user in this way, the corresponding body part name in the list 1001 is highlighted, for example. This allows the user to confirm that the body part specified by the user has been selected as the specific body part. The above method of designation is one example, and for example, buttons corresponding to the parts obtained by structural analysis may be provided, and the desired part may be designated by pressing the button corresponding to the part.
[0046] S905 corresponds to S404 in the first embodiment. That is, a process is performed to project the foreground model at the reference time s, the partial model of the specific part determined in S904, and a background model prepared in advance based on virtual viewpoint information. In this embodiment, the position, orientation, etc. of the partial model corresponding to the specific part determined in S904 are adjusted as needed. FIG. 11A is a diagram illustrating how the position and orientation of the foreground model at the reference time are adjusted in accordance with the shape data of the foreground model so that the base of the left arm determined as the specific part coincides at each time. Now, a foreground object of a person is moving, moving its left arm, from time t-1 to time t. Foreground model 1100 corresponds to time t, and foreground model 1100' corresponds to time t-1. In the adjustment, a geometric transformation is performed so that the base of the left arm 1101' of foreground model 1100' at time t-1 overlaps with the base of the left arm 1101 of foreground model 1100 at time t. At this time, if the scale of the left arm differs between time t-1 and time t, for example, a transformation process may be performed to match the scale. Figure 11(b) shows the result of the adjustment, in which the foreground model 1100' at time t-1 and the foreground model at time t are three-dimensionally overlapped with the base of the left arm, designated as a specific part, aligned. Such adjustments are made to the partial models as necessary, and the foreground model and partial model are projected onto the virtual viewpoint.
[0047] S906 to S908 correspond to S405 to S407 in the first embodiment, respectively, and there is no particular difference therebetween, so a description thereof will be omitted.
[0048] The above is the flow for generating a virtual viewpoint image according to this embodiment. Note that, prior to the start of the flow of Fig. 9, a structural analysis of the foreground model at each time may be performed by a separate device, and the results may be acquired in S903. Furthermore, if adjustments are made in S905, the adjustment results may be displayed on the UI screen so that the user can make fine adjustments to the position, orientation, etc.
[0049] <Variation 1> Instead of changing the transparency of a specific part designated by a user depending on time, the transparency of the area surrounding the specific part may be changed depending on the distance from the specific part. This type of afterimage representation is useful, for example, for analyzing a golf swing. Specifically, the head of a golf club is determined as the specific part, and transparency processing is performed when projecting a foreground model of the golf club so that the head becomes opaque and the further away from the head the more transparent the object becomes. This transparency processing refers to the results of structural analysis, and the alpha value decreases as the distance from the designated head increases, making the golf club gradually transparent. The foreground model at each time is then projected onto the virtual viewpoint. In this case, the relationship between distance and alpha value can be determined by pre-setting a transparency start distance and a transparency end distance. Alternatively, the distance from the head to the farthest part can be detected and the transparency end distance can be set as the distance from the head to the closest part. In this case, the transparency start distance can be set as the distance from the head to the closest part. Figure 12(a) shows an example of a virtual viewpoint image obtained by the synthesis processing of this modification. The projection image of the foreground model at the most recent time is combined with projection images of the foreground model at each of the earlier times, which have been processed so that the image becomes gradually more transparent from the base of the golf club head to the handle. This modification allows for a virtual viewpoint image that emphasizes the trajectory of a specific part designated by the user.
[0050] By applying the method of this modified example and then superimposing the method of the above-described embodiment, which makes the image more transparent as the time progresses, a virtual viewpoint image like the one shown in Figure 12(b) can be obtained. At this time, the virtual viewpoint may be moved (for example, starting from a position capturing the person from the front and moving at a constant speed to a position capturing the person from behind at the end of the shot). This allows a virtual viewpoint image to be obtained in which the movement of the club head is projected at any camera angle.
[0051] As described above, according to this embodiment, a virtual viewpoint image in which only a portion of a foreground object designated by the user is represented by an afterimage can be easily obtained. Furthermore, by adjusting the position and orientation of the foreground model as necessary, a virtual viewpoint image that is less unnatural can be obtained.
[0052] (Other Examples) The present disclosure can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0053] 102 Image processing device 300 Data Acquisition Unit 302 Model Analysis Department 303 Virtual viewpoint image generation unit 304 Projection section 305 Synthesis Section
Claims
1. an acquisition means for acquiring shape data of an object corresponding to each of a plurality of times during a period in which images are continuously captured by a plurality of image capturing devices; a determining means for determining a specific portion that is a part of the object; a generating means for generating a virtual viewpoint image corresponding to a virtual viewpoint using shape data of the object at a reference time among the plurality of times and shape data of a specific part of the object at a specific time among the plurality of times that is different from the reference time; and the generating means generates the virtual viewpoint image by combining a first projection image obtained by projecting shape data of the object at the reference time onto the virtual viewpoint and a second projection image obtained by projecting shape data of a specific part of the object at the specific time onto the virtual viewpoint.
1. An image processing device comprising:
2. the first projection image and the second projection image are four-channel images in which each pixel has an alpha value in addition to an RGB value; the generating means generates the virtual viewpoint image by alpha blending using the first projection image and the second projection image.
2. The image processing device according to claim 1, wherein:
3. 3. The image processing device according to claim 2, wherein in the alpha blending, the transmittance of a pixel corresponding to the specific part at the specific time in the second projection image is different from the transmittance of a pixel corresponding to the specific part at the reference time.
4. The image processing device according to claim 3, characterized in that the generation means performs the alpha blending so that, when the specific time is a time earlier than the reference time, the transmittance of pixels corresponding to the specific part at the specific time is higher than the transmittance of pixels corresponding to the specific part at the reference time.
5. 5. The image processing device according to claim 4, wherein the generating means performs the alpha blending so that the transmittance of pixels corresponding to the specific part in the second projection image becomes higher as the specific time becomes earlier than the reference time.
6. The image processing device according to claim 3, characterized in that the generation means performs the alpha blending so that, when the specific time is a time in the future than the reference time, the transmittance of pixels corresponding to the specific part at the specific time is higher than the transmittance of pixels corresponding to the specific part at the reference time.
7. 7. The image processing device according to claim 6, wherein the generating means performs the alpha blending so that the transmittance of pixels corresponding to the specific part in the second projection image becomes higher as the specific time becomes further in the future than the reference time.
8. 3. The image processing device according to claim 2, wherein in the alpha blending, the transmittance of the pixel corresponding to the specific portion in the second projection image is different from the transmittance of the pixel corresponding to the periphery of the specific portion.
9. 9. The image processing device according to claim 8, wherein the generating means performs the alpha blending so that the transmittance of pixels corresponding to the specific portion is higher than the transmittance of pixels corresponding to a periphery of the specific portion.
10. 10. The image processing device according to claim 9, wherein the generating means performs the alpha blending so that the transmittance of pixels corresponding to the periphery of the specific portion in the second projection image increases as the distance from the specific portion increases.
11. the first projection image and the second projection image are three-channel images in which each pixel has an RGB value; the generating means generates the virtual viewpoint image by combining the first projection image with an image obtained by changing the brightness or saturation of the second projection image.
2. The image processing device according to claim 1, wherein:
12. 12. The image processing device according to claim 11, wherein in the synthesis, the brightness or saturation of a pixel corresponding to the specific part at the specific time in the second projection image is different from the brightness or saturation of a pixel corresponding to the specific part at the reference time.
13. The image processing device according to claim 12, characterized in that the generation means performs the synthesis so that, when the specific time is a time earlier than the reference time, the brightness or saturation of the pixels corresponding to the specific part at the specific time is lower than the brightness or saturation of the pixels corresponding to the specific part at the reference time.
14. 14. The image processing device according to claim 13, wherein the generating means performs the synthesis so that the brightness or saturation of pixels corresponding to the specific part at the specific time in the second projection image becomes lower as the specific time becomes earlier than the reference time.
15. The image processing device according to claim 12, characterized in that the generation means performs the synthesis so that, when the specific time is a time in the future than the reference time, the brightness or saturation of the pixels corresponding to the specific part at the specific time is lower than the brightness or saturation of the pixels corresponding to the specific part at the reference time.
16. 16. The image processing device according to claim 15, wherein the generating means performs the synthesis so that the brightness or saturation of pixels in the second projection image corresponding to the specific part at the specific time becomes lower as the specific time becomes further in the future than the reference time.
17. The generating means generating a depth image in which the distance from the virtual viewpoint to the object is recorded for each pixel; The second projection image is selected for each pixel so that the smaller the depth, the higher the layer is selected, and the synthesis is performed.
2. The image processing device according to claim 1, wherein:
18. 2. The image processing apparatus according to claim 1, wherein the generating means performs the synthesis by performing a geometric transformation on the shape data corresponding to the specific portion.
19. 19. The image processing device according to claim 18, wherein the geometric transformation is a process of adjusting at least one of a position, an orientation, and a scale of the shape data corresponding to the specific portion in accordance with the shape data of the object at the reference time.
20. The determining means analyzing the shape changes of the object at the plurality of times using the shape data; determining a portion with a large amount of change as the specific portion based on the results of the analysis; 20. The image processing device according to claim 1, wherein the image processing device is a processor.
21. The determining means analyzing the shape changes of the object at the plurality of times using projected images obtained by projecting shape data of the object corresponding to each of the plurality of times onto a virtual viewpoint; determining a portion with a large amount of change as the specific portion based on the results of the analysis; 20. The image processing device according to claim 1, wherein the image processing device is a processor.
22. The determining means analyzing, using the shape data, a positional relationship between the object at the plurality of times and a specific region in the three-dimensional space where the images were captured; determining a portion inside or outside the specific region as the specific site based on the results of the analysis; 20. The image processing device according to claim 1, wherein the image processing device is a processor.
23. 23. The image processing device according to claim 22, wherein the specific region is a bounding box.
24. further comprising a graphical user interface for receiving designation of the specific portion from a user; The determining means determines the specific part based on the received specification.
20. The image processing device according to claim 1, wherein the image processing device is a processor.
25. the graphical user interface displays the results of a structural analysis of the object; the user specifies the specific portion based on the displayed result of the structural analysis; 25. The image processing device according to claim 24.
26. an acquisition step of acquiring shape data of an object corresponding to each of a plurality of times during a period in which images are continuously captured by a plurality of image capturing devices; a determining step of determining a specific portion that is a part of the object; a generating step of generating a virtual viewpoint image corresponding to a virtual viewpoint using shape data of the object at a reference time among the plurality of times and shape data of a specific part of the object at a specific time among the plurality of times that is different from the reference time; Including, In the generating step, a first projection image obtained by projecting shape data of the object at the reference time onto the virtual viewpoint and a second projection image obtained by projecting shape data of a specific part of the object at the specific time onto the virtual viewpoint are combined to generate the virtual viewpoint image. An image processing method comprising:
27. A program for causing a computer to function as the image processing device according to any one of claims 1 to 25.
Citation Information
Patent Citations
Game device
JP1995328228A
Image processing method
JP2003036450A
Image formation information, information storage medium and image formation device
JP2004298305A
Image processing for game
JP2006065877A
Image processing device, display method, and program
JP2021013095A