Image processing device, control method, and program
The image processing device uses three-dimensional shape data to detect and display athlete movements, addressing the challenge of conveying rapid movements in virtual viewpoint videos, thereby enhancing viewer experience and coaching tools.
Patent Information
- Application Number
- JP2023097066
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Existing systems struggle to effectively display the rapid movements of athletes in virtual viewpoint videos without disrupting the viewer's experience, as numerical information about speed and direction can be difficult to grasp and interfere with the viewing experience.
An image processing device that utilizes three-dimensional shape data to detect and display the movements of subjects, generating a virtual viewpoint image with additional information such as speed and direction through a graphical user interface, allowing for seamless integration with the video.
Enables the provision of athlete information in virtual viewpoint videos without interfering with the viewing experience, enhancing the viewer's understanding of player movements and improving coaching and commentary capabilities.
Smart Images

Figure 0007746331000001 
Figure 0007746331000002 
Figure 0007746331000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device that generates a virtual viewpoint video. [Background technology]
[0002] There is a virtual viewpoint video generation system that generates a virtual viewpoint video, which is an image seen from a virtual viewpoint specified by a user, based on images captured by a shooting system using multiple cameras. Patent Document 1 describes a system in which, after transmitting images captured by multiple cameras, an image computing server (image processing device) extracts images with large changes from the captured images as foreground images and images with small changes as background images.
[0003] In sports, the position information of athletes is detected using sensors attached to athletes or from images captured from multiple directions. The position information of athletes is used, for example, for coaching athletes and for commentary during broadcasts. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2017-211828 Summary of the Invention [Problem to be solved by the invention]
[0005] However, information such as the speed and direction of a player's movements fluctuates rapidly. Therefore, if player information is displayed numerically, it can be difficult for viewers to grasp at a glance. This can also disrupt the user's viewing experience. [Means for solving the problem]
[0006] The image processing device an acquisition means for acquiring three-dimensional shape data; Position multiple subjects Based on the three-dimensional shape data corresponding to each of the plurality of subjects, A detection means for detecting and using the three-dimensional shape data to display information regarding the movements of the plurality of subjects at positions based on the positions of the subjects detected by the detection means. Image generation means for generating a virtual viewpoint image and Has. [Effects of the Invention]
[0007] According to the present invention, information about players in a virtual viewpoint video can be provided to a viewer without interfering with the viewing experience of the virtual viewpoint video. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram illustrating an example of an image processing system. [Figure 2] (a) is a diagram showing an example of the position of a subject; (b) is a diagram showing an example of a subject extracted by a shape extraction unit; (c) is a diagram showing an example of the position of a subject in a different state; (d) is a diagram showing an example of a subject extracted by a shape extraction unit; and (e) is a diagram showing an example of an extracted shape. [Figure 3] 1A is an example of an extracted shape with an identifier attached, and FIG. 1B is an example of a graphical user interface for displaying the extracted shape and the identifier. [Figure 4] FIG. 10 is a diagram illustrating an example of a representative position. [Figure 5] 10 is a flowchart illustrating an example of a tracking analysis process performed by a tracking unit. [Figure 6] 10 is a flowchart showing an example of a process of superimposing subject information by a video generating unit. [Figure 7] 10 shows an example of an image generated by the image generation unit. [Figure 8] An example of a marking that indicates the subject's speed and direction of travel. [Figure 9] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0009] [First embodiment] (System configuration and operation of image processing device) FIG. 1 shows an example of the configuration of an image processing system for generating a virtual viewpoint video according to this embodiment. The image processing system includes, for example, an imaging unit 1, a synchronization unit 2, a three-dimensional shape estimation unit 3, a storage unit 4, a viewpoint instruction unit 5, a video generation unit 6, a display unit 7, and a subject position detection unit 8. The video generation unit 6 includes a foreground video generation unit 61, a background video generation unit 62, a subject information video generation unit 63, and a video synthesis unit 64, and the subject position detection unit 8 includes a shape extraction unit 11, a tracking unit 12, a subject position calculation unit 13, and an identification setting unit 14. The image processing system may be configured by one image processing device, or may be a system configured by multiple image processing devices. In the following description, the image processing system will be described as being one image processing device.
[0010] An overview of the operation of each component in an image processing device that generates virtual viewpoint images using this system is provided below. First, multiple image capture units 1 capture images in synchronization with one another based on a synchronization signal from a synchronization unit 2. The image capture units 1 output the captured images to a three-dimensional shape estimation unit 3. Note that the image capture units 1 are installed so as to surround the capture area including the subject, so that the subject can be captured from multiple directions. The three-dimensional shape estimation unit 3 uses the input captured images from multiple viewpoints to extract, for example, the silhouette of the subject, and then generates the three-dimensional shape of the subject using a volume intersection method or the like. The three-dimensional shape estimation unit 3 also outputs the generated three-dimensional shape of the subject and the captured images to a storage unit 4. Here, the subject refers to the object for which the three-dimensional shape is to be generated, and includes people and items handled by people.
[0011] The subject position detection unit 8 detects the position of the subject within the shooting area, and outputs the detected subject position information to the storage unit 4, as will be described in detail later.
[0012] The storage unit 4 saves and stores a group of data (material data) used to generate the virtual viewpoint video. Specifically, the data used to generate the virtual viewpoint video includes the captured image and the three-dimensional shape of the subject input from the three-dimensional shape estimation unit 3, camera parameters such as the position, orientation, and optical characteristics of each imaging unit, and subject position information acquired by the subject position detection unit 8. Note that a background model and a background texture image are saved (recorded) in advance in the storage unit 4 as data used to generate the background of the virtual viewpoint video.
[0013] The viewpoint instruction unit 5 comprises a viewpoint operation unit, which is a physical user interface such as a joystick or jog dial (not shown), and a display unit for displaying a virtual viewpoint video. The virtual viewpoint of the virtual viewpoint video displayed here can be changed by the viewpoint operation unit. In response to changes in the virtual viewpoint by the viewpoint operation unit, a virtual viewpoint video is generated as needed by the video generation unit 6 (described later) and displayed on the display unit. This display unit may share the display unit 7 (described later) or may be provided with a separate display device. The viewpoint instruction unit 5 generates virtual viewpoint information based on input from the viewpoint operation unit and outputs the generated virtual viewpoint information to the video generation unit 6. The virtual viewpoint information includes information corresponding to external camera parameters such as the position and orientation of the virtual viewpoint, information corresponding to internal camera parameters such as the focal length and angle of view, and time information specifying the shooting time to be played back.
[0014] Based on the time information included in the input virtual viewpoint information, the video generation unit 6 acquires the material data for the shooting time from the storage unit 4. The video generation unit 6 generates a virtual viewpoint video from the set virtual viewpoint using the three-dimensional shape of the subject and the shot image from the acquired material data, and outputs the generated video to the display unit 7.
[0015] The display unit 7 is a display means for displaying the image input from the image generation unit 6. The display unit 7 is configured by a display, a head-mounted display (HMD), or the like.
[0016] (Subject position tracking method) Next, a method for tracking the subject position in this embodiment will be described.
[0017] First, the three-dimensional shape estimation unit 3 generates a three-dimensional shape of the subject, and outputs the generated three-dimensional shape to the storage unit 4, and also outputs the generated three-dimensional shape to the shape extraction unit 11.
[0018] The shape extraction unit 11 cuts out the lower part of the three-dimensional shape of the subject as shown in FIG. 2(b) from the three-dimensional shape of the subject as shown in FIG. 2(a). In this embodiment, the cutout is from the bottom of the circumscribed rectangular parallelepiped of the three-dimensional shape of the subject to a predetermined height (e.g., a height equivalent to 50 cm). For example, as shown in FIG. 2(c), if one of the subjects is standing and the other subject has jumped or otherwise moved away from the floor of the shooting area, the three-dimensional shape of the subject is cut out as shown in FIG. 2(d). That is, for both of the three-dimensional shapes of the subjects, the three-dimensional shape is cut out from the part corresponding to the feet to a predetermined height.
[0019] Next, as shown in FIG. 2(e), the shape extraction unit 11 generates a two-dimensional image by projecting the cut-out three-dimensional shape onto a plane that is a view of the three-dimensional shape of the subject from directly above. In this embodiment, the shape extraction unit 11 parallel projects the cut-out three-dimensional shape onto a two-dimensional plane that corresponds to the feet (floor). In this embodiment, the projected image is a binary image in which the cut-out three-dimensional shape is white and the other parts are black. The shape extraction unit 11 divides this two-dimensional image into individual regions and obtains circumscribing rectangles 201 to 204 shown in FIG. 2(e). The shape extraction unit 11 outputs the vertex information of this circumscribing rectangle as the extracted three-dimensional shape (extracted shape). Here, the shape extraction unit 11 converts the vertex information of the circumscribing rectangle into the same coordinate system and units as the three-dimensional space of the captured region before outputting it. In addition, to determine the individual shapes, the shape extraction unit 11 uses a technique such as continuous component analysis on the projected two-dimensional image. By using such a method, the shape extraction unit 11 can divide the three-dimensional shape into individual regions.
[0020] The identification setting unit 14 assigns an identifier to the extracted shape output by the shape extraction unit 11. Specifically, the identification setting unit 14 calculates the distance between each extracted shape and assigns an identifier according to the distance between the extracted shapes. For example, as shown in FIG. 3(a), the identification setting unit 14 assigns the same identifier to extracted shapes whose distance between them is less than a predetermined distance (solid arrow), and assigns a different identifier to extracted shapes whose distance between them is equal to or greater than the predetermined distance (dashed arrow). The predetermined distance threshold used as the criterion for judgment is preferably a distance equivalent to the distance between the feet of the subject when standing. In this embodiment, the predetermined distance threshold is set to 50 cm.
[0021] The identification setting unit 14 displays the assigned identifiers on a display unit provided in the identification setting unit 14 using a graphical interface (GUI) such as that shown in FIG. 3(b). The user operates the image processing system while viewing this GUI. Specifically, the identification setting unit 14 displays the current identifier assignments (initial identifier assignments) on the graphical interface, distinguishing them by at least one of text and color coding. In FIG. 3(b), the identification setting unit 14 displays each identifier using both text and color coding. The user checks the GUI to confirm whether the desired identifiers have been assigned as the initial state. If the desired identifiers have not been assigned, the user instructs the subject to change their position or close their legs, repeating this process until the desired identifier assignment is achieved. Alternatively, the user operates the image processing system via the GUI to instruct changes to the desired identifier assignment. If the desired identifiers have been assigned, the user presses, for example, a confirmation button (initial identifier confirmation button) on the graphical interface as shown in FIG. 3(b). In response to this operation, the identification setting unit 14 determines the initial identifiers. Then, the identification setting unit 14 outputs the identifier assigned to each extracted shape to the tracking unit 12.
[0022] In response to receiving an identifier from the identification setting unit 14, the tracking unit 12 assigns the identifier to each extracted shape as an initial state. Thereafter, the tracking unit 12 tracks the extracted shape to which this identifier has been assigned. Note that the identifier assigned to the extracted shape during tracking is not an identifier determined by the identification setting unit 14, but an identifier determined based on the tracking results of the tracking unit 12 for the position of each extracted shape. In tracking (tracking analysis) of the extracted shape, the tracking unit 12 tracks the extracted shape based on the position of each extracted shape at the time immediately before the extracted shape was photographed, the identifier of each extracted shape, and information on the subject position input from the subject position calculation unit (described later). Note that specific tracking processing by the tracking unit 12 will be described later. The tracking unit 12 assigns an identifier to each extracted shape at that time based on the results of the tracking analysis, and outputs each extracted shape to the subject position calculation unit 13.
[0023] The subject position calculation unit 13 calculates a representative position for each extracted shape to which an identifier has been assigned, input by the tracking unit 12. For example, as shown in Fig. 4, the subject position calculation unit 13 calculates a position indicating each extracted shape group for each extracted shape group to which the same identifier has been assigned, such as representative positions 401 and 402. In this embodiment, the representative position is the center position of the extracted shape group.
[0024] However, because this representative position is affected by shape estimation errors and fluctuations in boundary portions when the shape is extracted by the shape extraction unit 11, the position may fluctuate at each time even if the subject is stationary. Therefore, in this embodiment, the subject position calculation unit 13 performs processing such as low-pass filtering or moving averaging in the time direction on the center position information at each time to generate position information in which high-frequency components are suppressed. The subject position calculation unit 13 then outputs the position information of the representative position together with an identifier as the position of the subject to the tracking unit 12. The subject position calculation unit 13 also records (stores) information in the storage unit 4 as subject position information (subject position information), in which information on the time when the three-dimensional shape that served as the basis for tracking analysis was photographed is added to the position information of the representative position.
[0025] (Tracking analysis processing by the tracking unit 12) Next, an example of the tracking and analysis process of the extracted shape position in the tracking unit 12 will be described with reference to the flowchart of FIG.
[0026] In step S501, the tracking unit 12 performs initialization processing upon receiving input from the identification setting unit 14. Specifically, the tracking unit 12 acquires the identifiers of the extracted shapes input from the identification setting unit 14.
[0027] In step S502, the tracking unit 12 acquires the next extracted shape input from the shape extraction unit 11.
[0028] In step S503, the tracking unit 12 assigns the identifiers acquired from the identification setting unit 14 to the acquired extracted shapes, and outputs the extracted shapes to which the identifiers have been assigned to the subject position calculation unit 13.
[0029] In step S504 , the subject position calculation unit 13 calculates the subject position from the group of extracted shapes assigned the same identifier, and outputs the subject position to the tracking unit 12 .
[0030] The above processing from steps S501 to S504 corresponds to the initialization processing.
[0031] The subsequent steps S505 to S509 are executed every time while the imaging unit 1 is imaging the subject. When the imaging unit 1 has finished imaging the subject, the process of step S509 is completed, and the process of this flowchart ends.
[0032] In step S505, the tracking unit 12 acquires the extracted shape input from the shape extraction unit 11 and the subject position at the previous time (previous time) calculated by the subject position calculation unit 13. The previous time is, for example, the shooting time of the extracted shape generated one frame before the extracted shape currently being processed. Here, for comparison, the current time is also referred to as the current time. Here, the current time refers to the shooting time of the image used to generate the extracted shape currently being processed.
[0033] In step S506, if the subject position at the previous time overlaps with the representative position of each extracted shape at the current time, the tracking unit 12 assigns to the extracted shape an identifier that was assigned to the subject position that overlaps with that representative position. Here, in step S506, if the representative position of one extracted shape overlaps with multiple subject positions, the tracking unit 12 assigns an identifier indicating "undeterminable" to the extracted shape at the current time. This is because, for example, multiple extracted shapes with different identifiers assigned may overlap at the current time, such as when two subjects are close to each other, and therefore an identifier indicating "undeterminable" is assigned in the processing of this step. Extracted shapes assigned an identifier including an identifier indicating "undeterminable" are subjected to the processing of S509, described below.
[0034] In step S507, if the representative position of an extracted shape to which an identifier has not yet been assigned overlaps with the extracted shape of the previous time, the tracking unit 12 assigns the identifier assigned to the extracted shape of the previous time to the extracted shape of the current time.
[0035] In step S508, if there is another extracted shape to which an identifier has already been assigned at the current time within a predetermined range from an extracted shape to which no identifier has yet been assigned, the tracking unit 12 assigns the assigned identifier to the other extracted shape. The predetermined range is preferably a range equivalent to the distance between the subject's feet when standing. For example, the predetermined range is a 50 cm radius from the center of the extracted shape. Here, if there are multiple other extracted shapes to which identifiers have been assigned within a predetermined range from a certain extracted shape, the tracking unit 12 assigns the identifier of the closest extracted shape to the extracted shape. For extracted shapes to which no identifier has been assigned after completing the processing up to step S508, the tracking unit 12 determines that the extracted shape is not a tracking target. In this case, the tracking unit 12 does not output the extracted shape determined to be not a tracking target to the subject position calculation unit 13.
[0036] In step S509, the tracking unit 12 outputs to the subject position calculation unit 13 the extracted shape to which an identifier has been assigned in the processes of steps S506 to S508, and the identifier assigned thereto.
[0037] In step S510, a control unit (not shown) determines whether or not the imaging process of the subject by the imaging unit 1 has ended. If it is determined that the imaging process of the subject by the imaging unit 1 has not ended, the process of step S508 is executed. If it is determined that the imaging process of the subject by the imaging unit 1 has ended, the process of this flowchart ends.
[0038] Note that the processes from step S506 to step S508 are performed for each extracted shape. By repeating the processes from step S506 to step S509, the identifier set by the identification setting unit 14 is associated with the extracted shape at each time. Using this identifier, the subject position calculation unit 13 can distinguish between each subject and calculate the subject position.
[0039] Furthermore, if the tracking unit 12 assigns an "undeterminable" identifier to an extracted shape, it is possible that some of the identifiers initially set will not be assigned at a certain time. In such a case, the subject position calculation unit 13 does not update subject position information that has the same identifier as an identifier not assigned to the extracted shape. As a result, even if multiple extracted shapes overlap due to multiple subjects approaching each other, the multiple subject position information will not be in the same position. In this case, the multiple subject positions are maintained at their respective positions up to the previous time. If the overlapping multiple extracted shapes subsequently separate again due to the subjects moving away from each other, an identifier is assigned to each extracted shape based on the most recent subject position. In other words, when the overlapping of the multiple extracted shapes is resolved, updating of each subject position information is resumed.
[0040] By the above processing, the image processing system can track each subject and obtain the position information of each subject even when multiple subjects are in the shooting area. Furthermore, by the above processing, the image processing system can track each subject even when the subjects move closer to or further away from each other, causing overlapping or separation in the generated 3D shape model.
[0041] (Display processing of subject information) Next, a method for calculating and displaying subject information in this embodiment will be described with reference to the flowchart of FIG.
[0042] For the sake of explanation, the image generation unit 6 will be described here as being composed of four units: a foreground image generation unit 61, a background image generation unit 62, a subject information image generation unit 63, and an image synthesis unit 64. However, the image generation unit 6 does not necessarily have to be composed of these four units, and may be substantially composed as a single image generation unit.
[0043] First, in S601, the foreground image generation unit 61 acquires material data (virtual viewpoint material) stored in the storage unit 4 based on the time information included in the virtual viewpoint information. In S602, a foreground (subject) image is generated based on the acquired virtual viewpoint material. FIG. 7(a) is an example of the generated foreground image. Next, in S603, the background image generation unit 62 generates a background image. In this embodiment, the background to be drawn is generated using pre-stored 3D model data (mesh) and texture data. FIG. 7(b) is an example of the generated background image. As will be described in detail later, in the processing from S604 to S610, the subject information image generation unit 63 generates an image in which information about the subject is drawn at a position based on the subject position information. FIG. 7(c) is an example of the subject information image generated at this time. In S611, the image synthesis unit 64 synthesizes the generated foreground image, background image, and subject information image based on the depth information of each image to create a single image and output it to the display unit 7. 7(d) is an example of the virtual viewpoint image generated after synthesis. By performing the processes of S601 to S611 every time, an image on which subject information is superimposed is output. In other words, if the virtual viewpoint image is a moving image, the processes of S601 to S611 are performed for each frame (image) that constitutes the moving image.
[0044] The specific operation of the subject information video generating unit 63 for generating the subject information video at this time will be described.
[0045] First, in S604, the subject information video generation unit 63 acquires position information of all subjects included in the shooting area from the storage unit 4. At this time, the subject information video generation unit 63 acquires not only the subject position information at the time included in the virtual viewpoint information, but also the subject position information for several frames to several tens of frames before and after that time. From S605 to S610, processing is performed for each subject. First, in S605, the subject information video generation unit 63 determines whether to perform subsequent processing based on an identifier for identifying the subject included in the position information and display target information (described later). The display target information is data used to specify the target for displaying subject information, and is data linked to the subject identifier. For example, the display target information is generated as data to be displayed for a participating player based on information about the participating players, and is recorded in the storage unit 4. As a result, the subject information is not displayed for referees and other individuals not included in the display target information. Alternatively, a configuration may be adopted in which subjects to be displayed are selected from an interface (not shown) to generate display target information. In S606, filtering processing is performed on the position information of subjects determined to be display targets. Specifically, the position information for multiple frames acquired in S604 is averaged to calculate the position information, thereby reducing blurring due to detection errors of the subject position detection unit 8.
[0046] Next, in S607, the subject information video generation unit 63 acquires subject position information of the subject to be displayed at a time in the past, for example, one second ago, and several to several tens of frames before and after that time, and calculates the average value. Note that in S607, if the average value of the subject position information at the past time calculated in this step corresponds to the averaged position information previously calculated in S606, the subject information video generation unit 63 may perform processing to read that information. This can reduce the processing load in S607.
[0047] Next, in S608, the averaged position information for the previous time is subtracted from the averaged position information for the current time, and the difference in position information between the current time and the previous time (here, 1 second) is divided to find the position deviation for the current time unit. The magnitude of this position deviation for the current time unit is the speed, and the direction of the vector on the two-dimensional plane is the direction of travel of the subject.
[0048] In S609, the subject information video generation unit 63 generates a subject information video by drawing a triangular traveling direction mark, for example, as shown in FIG. 8(a), based on the calculated speed and traveling direction of each subject. In S610, the subject information video generation unit 63 determines whether or not the processes from S605 to S609 have been completed for all subjects. If the processes from S605 to S609 have not been completed for all subjects, the processes from S605 onwards are executed again. If the processes from S605 to S609 have been completed for all subjects, the process of S611 is executed.
[0049] Finally, in S611, the image synthesis unit 64 synthesizes the generated foreground image, background image, and subject information image based on the depth information of each image, creates a single image, and outputs it to the display unit 7.
[0050] As shown in FIG. 8, a traveling direction mark 81 displayed in the subject information video is drawn in the direction of travel of the subject. The drawing size of this traveling direction mark varies depending on the subject's speed. For example, the traveling direction mark 82 is drawn longer in the traveling direction as the subject's speed increases. Furthermore, if this traveling direction mark is placed in the center of the subject, it will be obscured by the subject's feet, making it difficult to determine the traveling direction. Therefore, in this embodiment, it is desirable to draw the traveling direction mark 83 at a fixed distance from the subject position and rotate its drawing direction around the subject position to match the traveling direction, as shown in FIG. 8(a). Furthermore, as shown in FIG. 8(a), since the traveling direction mark 81 is drawn at a fixed distance from the subject position, it is drawn on a circle centered on the subject position. Therefore, drawing a circular mark 84 or the like concentric with the drawing position of the traveling direction mark makes it easier for the viewer to visually recognize the correspondence between the subject's traveling direction mark and the subject, which is desirable. Furthermore, as shown in FIG. 8(a), speed information may also be displayed as a numerical value.
[0051] Additionally, information based on past data such as shooting success rate corresponding to the subject's position on the court or field may be displayed around the player.
[0052] With the above system configuration, even when the viewpoint changes significantly in virtual viewpoint video, it is possible to display additional information such as speed and direction of travel near the position of the subject displayed within the field of view. This allows the viewer to see the direction and speed of each subject moving, even when changing the viewpoint while the time in the virtual viewpoint video is stopped, improving the viewer experience. It can also be used for coaching and commentary.
[0053] (Other aspects of the first embodiment) In this embodiment, the subject position detection unit 8 has been described as a means for detecting the subject position based on the shape estimation results, but the present invention is not limited to this method of detecting the subject position. For example, a configuration in which a position sensor such as a GPS is attached to the player and the sensor value is acquired may be used, or alternatively, a configuration in which the subject position is detected using image recognition technology from images obtained by multiple imaging means may be used.
[0054] In this embodiment, the marks indicating the speed and direction of travel are triangular icons (marks), but the present invention is not limited to this. They may be arrow-shaped as shown in FIG. 8(b), or may have other shapes. For example, in this embodiment, the marks are displayed near the floor, but cone-shaped marks may be displayed at waist height of the subject, as shown in FIG. 8(c). Also, in FIG. 8(a), only the size (length) of the mark in the direction of travel is changed, but the size of the mark itself may be changed, and the color may also be changed, for example, blue for low speed and red for high speed. Furthermore, the mark may have multiple pieces of information, for example, by changing the shape of the mark according to speed and changing the color according to acceleration.
[0055] In this embodiment, if there is a subject among multiple subjects that is not to be displayed, the position information and speed of that subject are not calculated, but it is also possible to perform the calculations themselves but not display them.
[0056] Furthermore, in this embodiment, the subject information video generation unit 63 is configured to calculate the speed and traveling direction for each frame when generating the virtual viewpoint video, but this is not necessarily limited to this. For example, the subject position detection unit 8 may detect the position, calculate the speed and traveling direction, and record them in the storage unit 4. In this case, the subject information video generation unit 63 may obtain information on the speed and traveling direction as well as the position information, and use this information to draw a mark.
[0057] In this embodiment, averaging processing is performed as the filtering process for the position information, but this is not limited to this. For example, a low-pass filter such as an IIR filter or an FIR filter may be used. However, in a configuration in which the speed or the like is calculated each time, if a low-pass filter is used, the value will become incorrect if the time at which the virtual viewpoint video is played back is discontinuously changed. Therefore, in such a configuration, it is desirable to obtain information from nearby times and then average the information.
[0058] In addition, when speed information is calculated in advance and stored in the storage unit 4 in this way, a graph of the speed information of a specific player may be displayed. Also, for example, acceleration may be calculated from the speed history, and it may be determined that the player's acceleration is decreasing, and the player's fatigue level may be simply calculated and displayed.
[0059] (Other configurations) In the above embodiment, each processing unit shown in Fig. 1 is described as being configured by hardware, but the processing performed by each processing unit shown in these figures may be configured by a computer program.
[0060] FIG. 9 is a block diagram showing an example of the hardware configuration of a computer applicable to the indirect position estimation device according to each of the above embodiments.
[0061] The CPU 901 controls the entire computer using computer programs and data stored in the RAM 902 and the ROM 903, and also executes the processes described above as being performed by the indirect position estimation device according to each of the above embodiments. That is, the CPU 901 functions as each processing unit shown in FIG. 1.
[0062] The RAM 902 has an area for temporarily storing computer programs and data loaded from an external storage device 906, data acquired from the outside via an I / F (interface) 907, etc. The RAM 902 also has a work area used when the CPU 901 executes various processes. That is, the RAM 902 can be allocated as a frame memory, for example, or can provide various other areas as needed.
[0063] The ROM 903 stores setting data for the computer, a boot program, and the like. The operation unit 904 is made up of a keyboard, a mouse, and the like, and can be operated by a user of the computer to input various instructions to the CPU 901. The output unit 905 displays the results of processing by the CPU 901. The output unit 905 is made up of, for example, a liquid crystal display. For example, the viewpoint instruction unit 5 is made up of the operation unit 904, and the display unit 7 is made up of the output unit 905.
[0064] The external storage device 906 is a large-capacity information storage device, such as a hard disk drive. The external storage device 906 stores an operating system (OS) and computer programs for causing the CPU 901 to implement the functions of the various units shown in Fig. 1. Furthermore, the external storage device 906 may also store image data to be processed.
[0065] Computer programs and data stored in the external storage device 906 are loaded into the RAM 902 as appropriate under the control of the CPU 901, and become the subject of processing by the CPU 901. An I / F 907 can be connected to a network such as a LAN or the Internet, or other devices such as a projector or display device, and the computer can acquire and send various information via this I / F 907. In the first embodiment, the imaging unit 1 is connected to this, and captured images are input and controlled. 908 is a bus connecting the above-mentioned units.
[0066] The operation of the above-described configuration is controlled mainly by the CPU 901 as explained in the previous embodiment.
[0067] In another configuration, the above-described functions can be achieved by supplying a storage medium containing computer program code for implementing the functions to a system, and the system then reading and executing the computer program code. In this case, the computer program code itself read from the storage medium implements the functions of the above-described embodiments, and the storage medium containing the computer program code constitutes the present invention. Also included is a case in which an operating system (OS) running on a computer performs some or all of the actual processing based on the instructions of the program code, thereby implementing the above-described functions.
[0068] Furthermore, the present invention may be realized in the following form: That is, computer program code read from a storage medium is written to memory in a function expansion card inserted into a computer or in a function expansion unit connected to the computer, and the CPU in the function expansion card or function expansion unit then performs some or all of the actual processing based on the instructions of the computer program code, thereby realizing the above-mentioned functions.
[0069] When the present invention is applied to the storage medium, the storage medium stores computer program code corresponding to the processes described above.
[0070] The disclosure of this embodiment includes the following configurations and methods.
[0071] (Configuration 1) a detection means for detecting the position of a subject; an image generation means for generating a virtual viewpoint image from the virtual viewpoint material; An image processing device comprising: a display means for displaying information relating to the movement of the subject as subject information at a position based on the subject position information detected by said detection means.
[0072] (Configuration 2) The image processing device according to configuration 1, wherein the display means comprises a calculation means for calculating the speed and direction of travel of the subject from the subject position information at multiple times, and displays the speed and direction of travel as the subject information.
[0073] (Configuration 3) 3. The image processing device according to configuration 2, wherein the display means indicates the subject's moving direction by using an arrow, a triangle, or an icon similar thereto near the subject position as the subject information.
[0074] (Configuration 4) 4. The image processing device according to configuration 3, wherein the display means changes at least one of the size, length, and color of the icon indicating the traveling direction based on the calculated speed of the subject as the subject information.
[0075] (Configuration 5) 4. The image processing device according to configuration 3, wherein the display means displays, as the subject information, a circle or other icon surrounding the subject near the feet of the subject based on the subject position information.
[0076] (Configuration 6) 2. The image processing device according to configuration 1, wherein when the virtual viewpoint material includes a plurality of subjects, the display means displays information based on the subject position information for each of the subjects.
[0077] (Configuration 7) 7. The image processing device according to configuration 6, wherein the display means does not display subject information for a subject that is not designated in advance, even if the subject is included in the virtual viewpoint material.
[0078] (Configuration 8) a detection step of detecting the position of the subject; an image generation process for generating a virtual viewpoint image from the virtual viewpoint material; A control method comprising a display step of displaying information relating to movement of the subject as subject information at a position based on the subject position information detected in the detection step.
[0079] (Configuration 9) a detection step of detecting the position of the subject; an image generation process for generating a virtual viewpoint image from the virtual viewpoint material; a display step of displaying information relating to the movement of the subject as subject information at a position based on the subject position information detected in the detection step;
Claims
1. An acquisition means for acquiring three-dimensional shape data; a detecting means for detecting positions of a plurality of subjects based on the three-dimensional shape data corresponding to each of the plurality of subjects; an image generating means for generating a virtual viewpoint image using the three-dimensional shape data so as to display information relating to the movements of the plurality of subjects at positions based on the positions of the subjects detected by the detecting means; An image processing device having:
2. further comprising a calculation means for calculating a speed and a direction of travel of the subject based on positions of the plurality of subjects at a plurality of times; 2. The image processing device according to claim 1, wherein said image generating means displays a speed and a direction of travel.
3. 3. The image processing device according to claim 2, wherein the image generating means indicates the direction of travel of the subjects by using an icon such as an arrow, a triangle, or a similar icon near the position of the subject as information relating to the movement of the plurality of subjects.
4. 4. The image processing device according to claim 3, wherein the image generating means changes at least one of the size, length, and color of the icon indicating the direction of travel based on the calculated speed of the subject as information regarding the movement of the multiple subjects.
5. 4. The image processing device according to claim 3, wherein the image generating means displays, as information relating to the movement of the plurality of subjects, a circle or other icon surrounding the subject near the feet of the subject based on the position of the subject.
6. 2. The image processing device according to claim 1, wherein the image generating means displays information relating to the movement of a pre-designated subject among the plurality of subjects.
7. The image processing device described in Claim 1, characterized in that the image generation means displays further information around the subject from data of the subject based on the position of the subject.
8. The image processing device described in Claim 1, characterized in that the image generation means displays a graph of the velocity information of the multiple subjects.
9. An acquisition step of acquiring three-dimensional shape data; a detecting step of detecting positions of a plurality of subjects based on the three-dimensional shape data corresponding to each of the plurality of subjects; an image generating step of generating a virtual viewpoint image using the three-dimensional shape data so that information relating to the movements of the plurality of subjects is displayed at positions based on the positions of the subjects detected in the detecting step; A control method comprising:
10. An acquisition step of acquiring three-dimensional shape data; a detecting step of detecting positions of a plurality of subjects based on the three-dimensional shape data corresponding to each of the plurality of subjects; an image generating step of generating a virtual viewpoint image using the three-dimensional shape data so that information relating to the movements of the plurality of subjects is displayed at positions based on the positions of the subjects detected in the detecting step; A program that causes a computer to execute the following.
Citation Information
Patent Citations
Display device, display method, and program
JP2017135695A
Image processing system, image processor, control method, and program
JP2017211828A
Recognition device, recognition method, and recognition program
JP2022062374A
Work machine surroundings monitoring system
WO2018151280A1
Vehicle control device
WO2019239725A1