Image processing device, image processing method, and program
The image processing device tracks subjects in diverse environments by generating and using distance information to set consistent tracking points, addressing the challenge of varying subject positions and postures.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-10
AI Technical Summary
Existing systems struggle to accurately track subjects in diverse shooting environments, particularly when the subject's position changes significantly due to varying heights or complex postures, such as in bicycle races or movies with wire action.
An image processing device that generates and tracks three-dimensional shape information from multiple captured images, setting tracking points with identifiers based on distance information in virtual space, allowing for consistent identification of subjects across different capturing times.
Enables reliable tracking of subjects in various shooting environments, including sloped areas and complex postures, by using distance information to maintain consistent tracking points, even when subjects change position or posture.
Smart Images

Figure 2026042070000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing device that tracks a subject. [Background technology]
[0002] There is a technology that uses multiple captured images taken by an imaging system consisting of multiple imaging devices to generate a virtual viewpoint image, which is an image captured from a virtual viewpoint specified by a user. This technology can provide a virtual viewpoint image captured from a position where a real imaging device cannot be placed, for example, in sports such as soccer or basketball.
[0003] In recent years, there has been a demand for tracking the position of a subject included in a video to analyze the movement of the subject and utilize the analysis results. For example, in sports, it is desirable to track the position information of players and display it in conjunction with stats information including team and player information, thereby utilizing it for coaching players and commentary in broadcasts. As a method for tracking the position of a subject in technology for generating virtual viewpoint images, Patent Document 1 proposes a method for estimating the position of a subject using a part of an estimated three-dimensional shape that is included at a predetermined height. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2024-55093 Summary of the Invention [Problem to be solved by the invention]
[0005] While previous systems were only capable of tracking subjects as desired at the time, in recent years, there has been a growing demand for tracking subjects in an even wider variety of shooting environments. For example, when filming bicycle races, the course at a bicycle racing track has sloped areas, which causes the three-dimensional position of the cyclist to change significantly depending on his or her running position, making it difficult to track the subject. Also, when filming a movie featuring wire action, it can be difficult to track the subject because the actors move around in three-dimensional space in various postures.
[0006] The present disclosure aims to facilitate tracking of a subject in a wide variety of shooting environments. [Means for solving the problem]
[0007] In order to solve the above problem, an image processing device according to the present disclosure has the following configuration: an acquisition means for acquiring first shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a first capturing time and second shape information indicating the three-dimensional shape of the subject generated based on a plurality of captured images at a second capturing time, a generation means for generating first distance information indicating a distance from a first position in a virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information and second distance information indicating a distance from a second position in the virtual space to the three-dimensional shape corresponding to the second shape information based on the second shape information, and a setting means for setting a first tracking point of the three-dimensional shape corresponding to the first shape information at the first capturing time based on the first distance information, setting a second tracking point of the three-dimensional shape corresponding to the second shape information at the second capturing time based on the second distance information, and setting the first tracking point with the same identifier as the second tracking point if the distance between the position of the first tracking point and the position of the second tracking point is equal to or less than a predetermined value. [Effects of the Invention]
[0008] According to the present disclosure, it is possible to easily track a subject in a wide variety of shooting environments. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram illustrating an example of an image processing system. [Figure 2] FIG. 10 is a flowchart showing the process performed by the subject position detection unit 8 to set a detection point of the subject. [Figure 3] FIG. 10 is a diagram showing the processing content of the process in which subject position detection unit 8 sets a detection point of the subject. [Figure 4] FIG. 1 is a diagram showing a coordinate system used in track events. [Figure 5] FIG. 10 is a diagram showing an image for superimposition indicating the speed of a subject. [Figure 6] FIG. 10 is a diagram showing an image for superimposition showing the trajectory of a subject. [Figure 7] FIG. 10 is a diagram showing a display screen when a subject identifier is assigned to a detection point based on a user operation. [Figure 8] 10 is a diagram showing a process in which the viewpoint instruction unit 5 sets a virtual viewpoint using tracking information and subject identification information. [Figure 9] FIG. 1 illustrates an example of the hardware configuration of an image processing apparatus. [Figure 10] FIG. 10 is a flowchart showing a process in which a tracer 9 generates trace information. [Figure 11] FIG. 10 is a flowchart showing the process performed by the smoothing unit 11 to associate a tracking identifier with a subject identifier. [Figure 12] FIG. 10 is a diagram showing a correspondence table showing combinations of tracking identifiers and subject identifiers. DETAILED DESCRIPTION OF THE INVENTION
[0010] <Embodiment> According to a preferred embodiment of the present disclosure, an image processing device includes an acquisition means for acquiring first shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a first capturing time. The acquisition means also acquires second shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a second capturing time. The image processing device also includes a generation means for generating first distance information indicating a distance from a first position in a virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information. The generation means also generates second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information based on the second shape information. The image processing device also includes a setting means for setting a first tracking point of the three-dimensional shape corresponding to the first shape information at the first capturing time based on the first distance information. The setting means also sets a second tracking point of the three-dimensional shape corresponding to the second shape information at the second capturing time based on the second distance information. If the distance between the position of the first tracking point and the position of the second tracking point is equal to or less than a predetermined value, the first tracking point is assigned the same identifier as the second tracking point. The identifier here is an identifier that identifies the subject, and may be, for example, an ID assigned to each subject or the name of the subject. The identifier that identifies the subject may also be acquired from stats information. The first shape information and the second shape information are information that stores the three-dimensional coordinates of multiple components that make up a three-dimensional shape. The positions of the first tracking point and the second tracking point are three-dimensional coordinates in virtual space. Specifically, they are sets of coordinates on the X-axis, Y-axis, and Z-axis of the coordinate system of the virtual space. The first position and the second position are three-dimensional coordinates in virtual space. Therefore, the first position may be referred to as a predetermined point, and the second position may be referred to as a predetermined point.
[0011] This embodiment allows easy tracking of a subject in a wide variety of shooting environments. For example, the subject can be easily tracked in a shooting environment where the height of the subject changes significantly depending on the situation, such as a bicycle race. Or, in a movie where an actor uses wire action to rotate so that his head is facing downwards, that is, even when the subject's posture changes significantly, the subject can be easily tracked.
[0012] The first distance information indicates the distance from the first position to a plurality of first components constituting the three-dimensional shape corresponding to the first shape information. For example, if the three-dimensional shape is point cloud data composed of a plurality of points, the first components are points. Alternatively, if the three-dimensional shape is a mesh model, the first components are polygons constituting the mesh model. Alternatively, if the three-dimensional shape is represented by voxels, the first components are voxels. The first distance information may be created by placing a virtual camera at the first position and determining the orientation of the virtual camera so that it faces the three-dimensional shape, thereby generating a distance image. In this case, each pixel of the distance image has distance information from the virtual camera to the corresponding first component. Note that distance information to the components constituting the surface of the three-dimensional shape as viewed from the virtual camera is stored in each pixel. The orientation of the virtual camera is defined by pan, tilt, and roll. The second distance information indicates the distance from the second position to a plurality of second components constituting the three-dimensional shape corresponding to the second shape information.
[0013] The first tracking point is the first component with the shortest distance or the longest distance among the plurality of first components. The second tracking point is the second component with the shortest distance or the longest distance among the plurality of second components. Whether the component with the shortest distance or the component with the longest distance is set as the first tracking point is determined based on the relative position of the first position with respect to the three-dimensional shape. When the first position is set at a low position with respect to the three-dimensional shape in virtual space, the component with the longest distance is set as the first tracking point. For example, in the case of filming a bicycle race, when the first position is set at a position in virtual space that is lower than the three-dimensional shape representing the cyclist, corresponding to the ground in real space, the component with the longest distance is the component that constitutes the cyclist's head or back. When the first position is set at a high position with respect to the three-dimensional shape, the component with the shortest distance is set as the first tracking point. When the first position is set at a high position relative to the three-dimensional shape, the component with the longest distance may be set as the first tracking point, or when the first position is set at a low position relative to the three-dimensional shape, the component with the shortest distance may be set as the first tracking point. The relative positional relationship between the three-dimensional shape and the first position and the combination that determines whether the distance is longest or shortest may be set in advance depending on the subject to be photographed. Alternatively, the operator may set the combination at the start of photographing. Whether the component with the shortest distance or the component with the longest distance is set as the second tracking point is the same as for the first tracking point, and therefore a description thereof will be omitted.
[0014] This feature makes it easier to track subjects with complex shapes. For example, bikes used in bicycle racing are thin, and it may be difficult to reliably generate a three-dimensional shape. In such cases, it is possible to set up tracking using distance information on the components of the three-dimensional shape of the rider on the bike, rather than the bike itself.
[0015] The first tracking point may be a first component included in a predetermined area among the plurality of first components. For example, a region corresponding to the foreground may be detected in a distance image, and the first tracking point may be set from a component included in the detected region. The region corresponding to the foreground is determined based on a distance value. Similarly, the second tracking point may be a second component included in a predetermined area among the plurality of second components.
[0016] Further, the plurality of first components and the plurality of second components are classified into a plurality of regions, The first tracking point and the second tracking point may be set for each of the plurality of regions.
[0017] According to this aspect, a plurality of first tracking points and a plurality of second tracking points can be automatically set from the distance information.
[0018] The setting means also collectively sets, as a single first tracking point, a plurality of first tracking points that fall within a predetermined range among the plurality of first tracking points. Specifically, the setting means collectively sets a single first tracking point, which is a collection of the plurality of first tracking points, at the center of gravity of the plurality of first tracking points that fall within the predetermined range. The predetermined range is a range centered on the first tracking point. The predetermined range differs for each subject. For example, the subject may be a bicycle race, and the predetermined range may be set along the course of the bicycle race track. The direction of movement of the cyclist can be estimated based on the course of the bicycle race track. Therefore, the predetermined range may be determined based on the position of the cyclist on the course of the bicycle race track. Specifically, the predetermined range is set as an ellipse with its major axis aligned with the direction of travel of the cyclist. The lengths of the major and minor axes are set to encompass one cyclist. By setting the first position above the three-dimensional shape, a range image with a bird's-eye view of the three-dimensional shape can be generated. In the overhead image, a predetermined range that surrounds the cyclist may be set, and this range is set in advance for each subject. For example, if the subject is a bicycle race, the cyclist will race in a foreground position, so an ellipse is set as the predetermined range in the overhead image. Note that, although the above describes a method of grouping multiple first tracking points into one first tracking point, multiple second tracking points may also be grouped into one first tracking point.
[0019] The shape information indicates three-dimensional shapes of a plurality of subjects, and the setting means sets the first tracking points in the same number as the number of the plurality of subjects.
[0020] This aspect enables tracking of subjects even when they are touching or located close to each other. For example, if multiple subjects are holding hands, a single three-dimensional shape is generated. It is difficult to determine whether multiple subjects exist based on the three-dimensional shape alone. Therefore, by dividing the three-dimensional shape into multiple regions and setting multiple first tracking points, first tracking points can be set for multiple subjects even when multiple subjects exist in the three-dimensional shape. However, simply dividing the three-dimensional shape into multiple regions and setting first tracking points for each region may result in multiple first tracking points being set for a three-dimensional shape representing a single subject, depending on the method for setting the regions. Therefore, by collectively setting multiple first tracking points within a predetermined range as one tracking point, one tracking point can be set for one subject.
[0021] Furthermore, the first position is generated based on a bounding box that encloses the three-dimensional shape. The method for setting the bounding box that encloses the three-dimensional shape is not particularly limited. A trained model may be provided that inputs multiple three-dimensional shapes representing multiple subjects and outputs a bounding box that encloses each of the three-dimensional shapes. Alternatively, a virtual space containing multiple three-dimensional shapes may be divided into multiple regions, and each region may be determined to determine whether it contains a three-dimensional shape. By first dividing into larger regions and then gradually determining the smaller regions, that is, by classifying using an octree, multiple bounding boxes that enclose multiple three-dimensional shapes may be set.
[0022] Specifically, the first position is set to a position that is a predetermined distance away from the center of the upper surface of the bounding box.
[0023] This aspect makes it possible to generate distance information for each of a plurality of three-dimensional shapes.
[0024] The first position may be set based on a three-dimensional shape of a background identified based on the position of the three-dimensional shape. For example, if the three-dimensional shape of the background is a three-dimensional shape representing a velodrome, the line of sight of the virtual camera corresponding to the first position is set in a direction perpendicular to the course of the velodrome. As a result, the distance information becomes information indicating the distance from the first position to the three-dimensional shape in a direction perpendicular to the course of the velodrome on which the three-dimensional shape is located.
[0025] This embodiment reduces the risk that when multiple subjects are present and their three-dimensional shapes are captured from the first position, the multiple subjects will overlap, making it impossible to obtain appropriate distance information. In other words, when multiple subjects are present, the multiple subjects can be tracked more stably. It also becomes applicable to shooting in a wide variety of venues. For example, if the first position is set to face the Z-axis direction of the virtual space, and the stadium has an incline, multiple athletes will overlap when viewed from the first position. In such cases, setting the first position to match the venue can reduce the risk of multiple subjects overlapping and obtaining inappropriate distance information.
[0026] The image processing device also has an output unit that outputs position information indicating the position of the first tracking point and identifier information indicating the identifier. For example, the position information and the identifier information may be output to an external recording medium in association with each other.
[0027] This allows the tracking results to be reused in other devices, and can be used for a variety of purposes, such as displaying the tracked target's trajectory or analyzing the tracking target's movement trends.
[0028] According to another preferred embodiment of this embodiment, an image processing method includes an acquisition step of acquiring first shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a first capturing time. The acquisition step also includes acquiring second shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a second capturing time. The image processing method also includes a generation step of generating first distance information indicating a distance from a first position in a virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information. The generation step also includes generating second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information based on the second shape information. The image processing method also includes a setting step of setting a first tracking point of the three-dimensional shape corresponding to the first shape information at the first capturing time based on the first distance information. The setting step also includes setting a second tracking point of the three-dimensional shape corresponding to the second shape information at the second capturing time based on the second distance information. If the distance between the position of the first tracking point and the position of the second tracking point is equal to or less than a predetermined value, the first tracking point is assigned the same identifier as the second tracking point.
[0029] According to another preferred embodiment of this embodiment, a program causes a computer to execute the image processing method described above. By executing this program, the computer preferably functions as the image processing device described above.
[0030] <Example> In this embodiment, a distance image from a predetermined point in three-dimensional space is generated for each three-dimensional shape, and tracking points used to track the object are set using the distance image, thereby facilitating tracking of the object. In one example, a virtual camera is positioned perpendicular to the top surface of a bounding box surrounding the three-dimensional shape, and a distance image showing the distance from the virtual camera to the three-dimensional shape is generated. The distance image is then divided into predetermined regions, and distance minima are extracted for each region. Of the extracted minimum points, those that fall within a predetermined range are grouped together to set tracking points.
[0031] The image processing system generates a virtual viewpoint image representing a scene from a specified virtual viewpoint based on multiple images captured by multiple imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint image, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user. For example, the virtual viewpoint image also includes an image corresponding to a viewpoint selected by a user from multiple candidates. Furthermore, this embodiment will mainly describe a case where the virtual viewpoint is specified by a user operation, but the virtual viewpoint may also be specified automatically based on the results of image analysis, etc. Furthermore, this embodiment will mainly describe a case where the virtual viewpoint image is a video, but the virtual viewpoint image may also be a still image.
[0032] The viewpoint information used to generate a virtual viewpoint image is information indicating the position and orientation (line of sight direction) of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set serving as viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. Furthermore, the viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a moving image of the virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.
[0033] The image processing system has multiple imaging devices that capture images of an imaging area from multiple directions. The imaging area may be, for example, a stadium where sports such as soccer or karate are held, or a stage where a concert or play is held. The multiple imaging devices are installed at different positions surrounding the imaging area and capture images synchronously. Note that the multiple imaging devices do not need to be installed around the entire periphery of the imaging area; depending on installation space restrictions, they may be installed only around a portion of the periphery of the imaging area. Furthermore, the number of imaging devices is not limited to the example shown in the figure. For example, if the imaging area is a soccer stadium, approximately 30 imaging devices may be installed around the stadium. Furthermore, imaging devices with different functions, such as telephoto cameras and wide-angle cameras, may be installed.
[0034] In this embodiment, the multiple image capturing devices are cameras each having an independent housing and capable of capturing images from a single viewpoint. However, this is not limiting, and two or more image capturing devices may be configured in the same housing. For example, a single camera equipped with multiple lens groups and multiple sensors and capable of capturing images from multiple viewpoints may be installed as the multiple image capturing devices.
[0035] A virtual viewpoint image is generated, for example, by the following method. First, multiple images (multiple viewpoint images) are acquired by capturing images from different directions using multiple imaging devices. Next, a foreground image in which a foreground region corresponding to a predetermined object, such as a person or a ball, is extracted, and a background image in which a background region other than the foreground region is extracted are acquired from the multiple viewpoint images. Furthermore, a foreground model representing the three-dimensional shape of the predetermined object and texture data for coloring the foreground model are generated based on the foreground image, and texture data for coloring a background model representing the three-dimensional shape of a background, such as a stadium, is generated based on the background image. The texture data is then mapped to the foreground model and background model, and rendering is performed according to the virtual viewpoint indicated by the viewpoint information, thereby generating a virtual viewpoint image. However, the method for generating a virtual viewpoint image is not limited to this, and various methods can be used, such as a method of generating a virtual viewpoint image by projective transformation of captured images without using a three-dimensional model.
[0036] A foreground image is an image in which an object region (foreground region) is extracted from an image captured by an imaging device. An object extracted as a foreground region is a dynamic object (moving body) that moves (its absolute position and shape can change) when images are captured from the same direction in a time series. Examples of objects include players, referees, and other people on the field where a sport is being played, such as a ball in a ball game, or singers, musicians, performers, and presenters in a concert or entertainment event.
[0037] A background image is an image of at least a region (background region) different from the foreground object. Specifically, a background image is an image in which the foreground object has been removed from the captured image. Furthermore, the background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Examples of such imaged objects include a stage for a concert, a stadium where an event such as a sport is held, a structure such as a goal used in a ball game, or a field. However, the background is at least a region different from the foreground object, and the imaged object may include other objects in addition to the object and background.
[0038] A virtual camera is a virtual camera that is different from the multiple imaging devices actually installed around the imaging area, and is a concept for conveniently explaining a virtual viewpoint related to the generation of a virtual viewpoint image. That is, a virtual viewpoint image can be considered to be an image captured from a virtual viewpoint set in a virtual space associated with the imaging area. The position and orientation of the viewpoint in the virtual image capture can be expressed as the position and orientation of the virtual camera. In other words, a virtual viewpoint image can be considered to be an image simulating an image captured by a camera assuming that the camera exists at the position of the virtual viewpoint set in space. In this embodiment, the content of the transition of the virtual viewpoint over time is referred to as a virtual camera path. However, the concept of a virtual camera is not required to realize the configuration of this embodiment. That is, it is sufficient if at least information representing a specific position and information representing a direction in space are set, and a virtual viewpoint image is generated according to the set information.
[0039] (Configuration and operation of virtual viewpoint image generation) FIG. 1 shows an example of the configuration of an image processing system for generating a virtual viewpoint image according to this embodiment. The image processing system includes an imaging unit 1, a synchronization unit 2, a three-dimensional shape estimation unit 3, a storage unit 4, a viewpoint instruction unit 5, a video generation unit 6, a display unit 7, a smoothing unit 11, a superimposed image generation unit 12, and an image processing unit 13. The image processing unit 13 includes a subject position detection unit 8, a tracking unit 9, and an identification setting unit 10. The image processing unit 13 may also include the smoothing unit 11. The image processing system may be configured by one image processing device, or may be a system configured by multiple image processing devices. In the following description, the image processing unit 13 is assumed to be one image processing device, and the remaining devices are assumed to be individual devices.
[0040] The multiple imaging units 1 capture images in synchronization with one another based on a synchronization signal from the synchronization unit 2. The imaging units 1 output the captured images to the three-dimensional shape estimation unit 3. The imaging units 1 are installed to surround a capture area including the subject so that the subject can be captured from multiple directions.
[0041] The synchronization unit 2 outputs a synchronization signal to the plurality of image capture units 1 .
[0042] The three-dimensional shape estimation unit 3 uses the input captured images to generate, for example, a silhouette image of the subject, and then generates the three-dimensional shape of the subject using a volume intersection method or the like. The three-dimensional shape estimation unit 3 also associates the generated three-dimensional shape of the subject with the captured images and the shooting times of the captured images, and outputs them to the storage unit 4. That is, the three-dimensional shape of the subject is generated for each shooting time, and the generated three-dimensional shape of the subject is associated with the captured images and the shooting times and output to the storage unit 4. The format of the association is not limited, and for example, a single file may contain information indicating the three-dimensional shape of the subject, the captured images, and the shooting times. Alternatively, a file whose file name includes the shooting time and contains information indicating the three-dimensional shape of the subject, and a file whose file name includes the shooting time and contains the captured images may be output to the storage unit 4. Here, the subject refers to an object for which a three-dimensional shape is to be generated, and includes people and items handled by people.
[0043] The storage unit 4 saves and accumulates the following data group as data (material data) used to generate a virtual viewpoint image. Specifically, the data used to generate a virtual viewpoint image includes the three-dimensional shape of the subject and the captured image input from the three-dimensional shape estimation unit 3, and the shooting time of the captured image. The data used to generate a virtual viewpoint image also includes camera parameters such as the position, orientation, and optical characteristics of each image capture unit. Note that a background model and background texture image are saved (recorded) in advance in the storage unit 4 as data used to generate the background of the virtual viewpoint image. The storage unit 4 also acquires tracking information from the tracking unit 9 and subject identification information from the identification setting unit 10, and records them respectively. The storage unit 4 also records information indicating the combination of the tracking identifier and subject identifier acquired from the smoothing unit 11.
[0044] The viewpoint instruction unit 5 comprises a viewpoint operation unit, which is a physical user interface such as a joystick or jog dial (not shown), and a display unit for displaying a virtual viewpoint image. The virtual viewpoint of the virtual viewpoint image displayed here can be changed by the viewpoint operation unit. In response to changes in the virtual viewpoint by the viewpoint operation unit, a virtual viewpoint image is generated as needed by the image generation unit 6 (described later) and displayed on the display unit. This display unit may share the display unit 7 (described later) or may be provided with a separate display device. The viewpoint instruction unit 5 generates virtual viewpoint information based on input from the viewpoint operation unit and outputs the generated virtual viewpoint information to the image generation unit 6. The virtual viewpoint information includes information corresponding to external parameters of the camera, such as the position and orientation of the virtual viewpoint, information corresponding to internal parameters of the camera, such as the focal length and angle of view, and time information specifying the shooting time of the captured image used to generate the virtual viewpoint image.
[0045] Based on the time information included in the input virtual viewpoint information, the video generation unit 6 acquires the material data for the shooting time from the storage unit 4. The video generation unit 6 generates a virtual viewpoint image for the set virtual viewpoint using the three-dimensional shape of the subject and the captured image from the acquired material data, and outputs the generated image to the display unit 7.
[0046] The display unit 7 is a display means for displaying the image input from the image generation unit 6. The display unit 7 is configured with a display or the like.
[0047] The subject position detection unit 8 acquires the three-dimensional shape generated by the three-dimensional shape estimation unit 3. When multiple subjects are photographed, multiple subjects may be included in one three-dimensional shape. For example, when multiple subjects are touching, one three-dimensional shape is generated. Therefore, when multiple subjects are included in one three-dimensional shape, the subject position detection unit 8 separates the multiple subjects and sets detection points (tracking points) for each of the multiple subjects. These detection points are points that indicate three-dimensional coordinates in virtual space. Note that when one subject is included in one three-dimensional shape, one detection point is set. Specific processing will be described later. The set detection points are output to the tracking unit 9.
[0048] The tracking unit 9 assigns an individual tracking identifier to each detection point acquired from the subject position detection unit 8. If this is the first detection point acquired after shooting begins, the tracking unit 9 assigns an individual tracking identifier to the detection point. The method of assignment is not particularly limited. For example, tracking identifiers may be assigned randomly to detection points, or tracking identifiers may be assigned in order of proximity to the origin of the virtual space. If multiple detection points are acquired, different tracking identifiers are assigned to each detection point. For example, if detection points A and B are acquired, tracking identifiers A and B are assigned to each detection point. For subsequently acquired detection points, the position of the detection point at the shooting time (the shooting time of the processing target) corresponding to the acquired three-dimensional shape is compared with the position of the detection point at the previous shooting time. If the position of the detection point at the previous shooting time is within a predetermined range of the position of the detection point at the shooting time of the processing target, it is determined to be a detection point of the same subject. Then, the tracking identifier associated with the detection point at the previous shooting time is acquired, and this tracking identifier is assigned to the detection point at the shooting time of the processing target. In other words, the position of the detection point at the shooting time of the processing target is assigned the same tracking identifier as the detection point at the previous shooting time that is within a predetermined range. The tracking unit 9 repeatedly executes the above process in the order of shooting times, thereby generating information that associates the position information of the detection point with the tracking identifier at each shooting time. This information indicates the combination of the position information of the detection point and the tracking identifier for each shooting time. Then, the information that associates the position information of the detection point with the tracking identifier is output to the storage unit 4 as tracking information. When using the tracking information, the position information of the detection point is obtained based on the shooting time and the tracking identifier assigned to the subject to be tracked.
[0049] The identification setting unit 10 assigns individual subject identifiers to the detection points acquired by the subject position detection unit 8. The method for generating subject identifiers is not particularly limited. For example, subject identification information may be generated using captured images stored in the storage unit 4. Specifically, the positions of each detection point detected by the subject position detection unit 8 are projected onto the captured images of the multiple image capture devices based on the internal and external parameters of the multiple image capture devices. This projection allows the pixels corresponding to the detection points in the captured images to be identified. Color information about the vicinity of the identified pixels is then acquired. The reason for acquiring color information about the vicinity of the identified pixels is to prevent erroneous subject identification in the subsequent process when noise is present in the captured images. At this time, color information outside the subject's silhouette is not acquired. For example, in the case of bicycle racing, each subject (athlete) wears a different uniform color, so a subject identifier is generated in advance for each color. Subject identification information is generated by associating a subject identifier with a red uniform, such as Athlete A, and a blue uniform, such as Athlete B. Stats information, for example, may be used to generate this subject identifier in advance. Then, based on the generated subject identification information and color information acquired from the captured image, a subject identifier corresponding to the detection point is determined and assigned. Note that information other than color information, such as hue, saturation, and brightness, may be used. Note that when the identification setting unit 10 acquires color information from the captured image, if multiple subjects are present, occlusion by the multiple subjects may occur. For this reason, color information may be acquired from multiple captured images, a majority vote may be performed, and subjects with clearly different color information may be excluded, before assigning a subject identifier to the detection point. The identification setting unit 10 outputs information associating the position information of the detection point with the subject identifier to the storage unit 4 as subject identification information.
[0050] Because the process of generating subject identification information involves a large processing load, it is not necessary to perform the process for every shooting time. For example, the process may be performed once every few seconds. Alternatively, the process may be performed when conditions are met, depending on the method of assigning subject identifiers. For example, the identifiers may be assigned using position information of the detection points. Specifically, when shooting a baseball game, the positions of each player are roughly determined for each position just before the pitcher throws the ball. Therefore, a predetermined area is set for each position, and a subject identifier is assigned when a detection point is located within the predetermined area. Note that, since information on the participating players and their positions at the time of shooting can be extracted from the stats information, a detection point located within a predetermined area corresponding to a position can be determined to be that of that participating player. Therefore, it is possible to easily assign subject identifiers for participating players.
[0051] The smoothing unit 11 acquires the tracking information and subject identification information recorded in the storage unit 4, and generates a correspondence table showing the correspondence between the tracking identifiers assigned to the tracking information and the subject identifiers. Note that the identification setting unit 10 may assign an identifier once every few seconds, or may not have subject identification information in the storage unit 4 depending on the shooting time, such as when a subject cannot be properly identified due to the overlap of multiple subjects. Therefore, the smoothing unit 11 identifies the combination of tracking identifiers and subject identifiers for each shooting time based on the correspondence table showing the correspondence between tracking identifiers and subject identifiers. Details will be described later with reference to FIG. 10.
[0052] Furthermore, the smoothing unit 11 smoothes the position information included in the tracking information. The reason for smoothing is that the positions of the detection points detected by the subject position detection unit 8 may contain errors due to factors such as the subject's posture and the accuracy of shape estimation. Therefore, smoothing is performed because the information contains inappropriate information, including small fluctuations, when performing operations such as virtual viewpoint control and calculation of trajectory and speed information, as described below. Here, we will explain the smoothing process specifically for track-based competitions. Specifically, as shown in Figure 4, smoothing is performed separately for corner and straight sections. First, for the straight sections, a Cartesian coordinate system of the X-, Y-, and Z-axes is used, and processing such as low-pass filtering and moving average is performed in the time direction for each value of the X-, Y-, and Z-axes to generate smoothed position information with suppressed high-frequency components. For the corner sections, the Cartesian coordinate system of the X-, Y-, and Z-axes is first converted to a cylindrical coordinate system with the corner center 401 as the origin, as shown in Figure 4. Then, smoothing is performed in the time direction for each value of the radius r, angle θ, and height z, as with the straight sections, and the smoothed cylindrical coordinate position information is reconverted to a Cartesian coordinate system of the X-, Y-, and Z-axes. The reason for converting the corner portions into cylindrical coordinates before smoothing is that if smoothing is performed while the coordinates are still in Cartesian, the output result will be biased toward the inside of the corner, resulting in an incorrect smoothed result. The smoothing unit 11 is configured to include a velocity calculation means, and after smoothing the position information, it also calculates velocity information from the smoothed position information (smoothed position information). The smoothing unit 11 records the smoothed position information, velocity information, and subject identification information together in the storage unit 4 for each shooting time.
[0053] The superimposed image generation unit 12 acquires the smoothed position information, speed information, and subject identification information recorded in the storage unit 4 and generates an image for superimposition. An image for superimposition is, for example, a superimposed image (speed display image 501) that displays the speed of each player as shown in FIG. 5. When generating the speed display image 501, the speed information of each player at the corresponding time is acquired from the storage unit 4 based on the time information input from the viewpoint instruction unit 5, and the numerical value is drawn. In addition, the smoothed position information at each time is plotted on the virtual viewpoint image or drawn as lines connecting the information on the virtual viewpoint image as shown in FIG. 6, thereby drawing trajectories 601 to 603 of the course taken by each player.
[0054] (Subject position tracking method) Next, the subject position tracking method in this embodiment will be described using bicycle racing as an example. The corners of bicycle racing courses have a sloped structure called a bank, and there is a maximum elevation difference of 3 meters or more between the inside and outside of the course. This disclosure can also be applied to shooting environments where there is an elevation difference in the area of the subject to be shot.
[0055] The method for tracking the subject position consists of a process of detecting the detection point of the subject, a process of setting a tracking identifier based on the detection point at the previous shooting time, and a process of setting a subject identifier that indicates which subject the set tracking identifier belongs to. Each of the above processes will be explained using Figures 2, 10, 11, and 12.
[0056] 2 is a flowchart showing the process by which the subject position detection unit 8 sets the detection point of the subject. This process is assumed to be performed for each captured image corresponding to a three-dimensional shape. In other words, it is performed for each three-dimensional shape corresponding to a set of multiple captured images captured synchronously. Furthermore, since it is also possible for multiple image capture units 1 to capture multiple videos synchronously and generate a three-dimensional shape showing a series of movements from the multiple videos, it can also be said that this process is performed for each frame of a video of captured images.
[0057] In step S201, a three-dimensional shape is acquired from the shape extraction unit 11, and a height image for the three-dimensional shape is generated. As shown in FIG. 3(a), multiple three-dimensional shapes (subjects 301 to 303) are acquired. At this time, information indicating the area (bounding box) surrounding the three-dimensional shape is acquired. Note that the area surrounding the three-dimensional shape may be identified by the subject position detection unit 8 without acquiring the area surrounding the three-dimensional shape. Note that a method for setting the area surrounding the three-dimensional shape may be space division using an octree, but since this is a well-known technology, its description will be omitted. A height image (FIG. 3(b)) is generated by parallel projection from directly above each area surrounding the three-dimensional shape. Specifically, a distance image is generated by calculating the distance from the bottom surface of the bounding box surrounding the three-dimensional shape to each component of the three-dimensional shape. Note that if multiple components correspond to the same pixel, the component with the largest distance value is associated with the pixel, and a distance image indicating the distance from the bottom surface to the component of the three-dimensional shape that is farthest from the bottom surface is generated. Therefore, the distance image corresponds to a height image. When this height image is displayed as an image that can be identified by the operator, it is generated so that higher heights are brighter and lower heights are darker, and areas that do not include three-dimensional shapes are treated as 0. Note that the height image contains distance information for each pixel and does not necessarily need to be displayed as an identifiable image. The image size can be determined based on the circumscribing rectangle of the three-dimensional shape to be detected. Here, we assume that objects 301 and 302 are close to each other, and therefore are estimated as the same three-dimensional shape. As described above, the height image is a distance image that indicates the distance from a specified point in virtual space to the three-dimensional shape. It can also be considered information indicating the height from the floor. Note that the image is not limited to a height image. A virtual camera may be set at a position a specified distance from the center of the top surface of the bounding box in a direction perpendicular to the top surface, and a distance image from the virtual camera to the three-dimensional shape may be generated. The following processing is performed for each area surrounding the three-dimensional shape.
[0058] In step S202, a process is performed to remove false shapes 310 from the height image that may be generated due to occlusion conditions during image capture or subject extraction errors. The false shapes 310 include, for example, noise called floating debris, which is generated as a three-dimensional shape from sand or dust, or three-dimensional shapes generated due to subject extraction errors. Specifically, the floating shapes are removed by shrinking a predetermined number of pixels in the image whose numerical value is not 0 in FIG. 3(b) and then expanding the predetermined number of pixels (311 in FIG. 3(c)). In this embodiment, this process is referred to as shrinkage / expansion processing. Note that the process for removing false shapes 310 is not limited to this, and known techniques may be used. Techniques for removing false shapes and noise from captured images are known, so other processing methods will not be described here.
[0059] In step S203, the point where the height is maximum (the point where the local maximum value is) within a predetermined region in the height image is detected (identified) as the detection point. Specifically, the point where the height is maximum within an area of approximately 20 cm square is detected. This makes it possible to identify the top of each subject's head even when multiple people are walking hand in hand. In a bicycle race, since athletes are in a forward-leaning posture, their heads or backs may be detected as detection points 320 to 322 (FIG. 3(d)). The predetermined region is set by dividing the height image into multiple regions. Furthermore, the shape and size of the predetermined region may be set for each subject being photographed.
[0060] In step S204, multiple detection points are integrated. This process is performed because, when multiple regions are set in step S203, multiple detection points 320 and 321 may be detected for a single subject. When multiple detection points are detected, they are each classified into multiple regions. To set one detection point for a single subject, the multiple detection points are integrated to set one representative detection point. Specifically, as shown in FIG. 3(e), a determination is made as to whether there are any other detection points within a predetermined range centered on the detected detection point 321. For example, when photographing a bicycle race, a search is made for other detection points within a range of 70 cm in the general direction of travel (determining whether they are included within the dash-dotted line 330 in the figure), and any points found within the predetermined range are integrated. Here, the size of the predetermined range is set to 70 cm because, in bicycle races, the subject (cyclist) is riding in a forward-leaning position as shown in FIG. 3(a), and their head and back may be detected as detection points. Therefore, 70 cm is set as the approximate distance encompassing the head and back. In this example, the detection point 320 corresponds to this. For example, the midpoint 340 (center of gravity) between the detection points 320 and 321 is set as the detection point after integration. The traveling direction will be described later. Because there are no pixels with a pixel value of 0 between the detection points 320 and 321, they are treated as detection points of the same subject, and the detection points are integrated. In this embodiment, it is assumed that no other detection points are included in the predetermined range centered on the detection point 322. However, depending on the setting of the predetermined range, other detection points may be included. In such a case, one detection point is set for two subjects, making it difficult to properly track the subjects. Therefore, even if other detection points exist within the predetermined range centered on the detection point, if there is an area with a pixel value of 0 on the line connecting the detection point and the other detection points in the predetermined range, it may be possible to consider that the three-dimensional shapes are not connected and not integrate the detection points. For example, because there is an area with a pixel value of 0 between the detection point 322 and the detection point 320, the detection point 322 is treated as a detection point of a subject different from the subjects detected by the detection points 320 and 321, and integration is not performed.
[0061] Through the processing of steps S201 to S204, the subject position detection unit 8 detects one detection point for each subject. By repeatedly performing the above processing for each capture time of a captured image corresponding to a three-dimensional shape, it is possible to detect a detection point at each capture time. This detection point is three-dimensional position information on the X-axis, Y-axis, and Z-axis, and this is output to the tracking unit 9 and the identification setting unit 10. As a result, the tracking unit 9 can generate tracking information, and the identification setting unit 10 can generate subject identification information.
[0062] In this embodiment, the direction of travel is determined based on a spatial position. Specifically, in the case of a track event, the direction is the tangent direction of the course (generally, counterclockwise is positive) as shown in FIG. 4. Therefore, the direction of travel is determined based on the position of the subject on the stadium. Alternatively, the direction of travel may be determined based on the speed of the subject. Information on the direction of travel is associated with the subject to be photographed and is recorded in advance in the subject position detection unit 8.
[0063] In the above, the subject position detection unit 8 uses a method of performing contraction / expansion processing on the height image to remove the false shape 310, but the present invention is not limited to this. For example, a configuration may be adopted in which area division processing (segmentation) is performed on the effective pixels of the image, and any divided area whose area is less than a predetermined size (for example, less than 1,000 pixels) is removed from the detection target.
[0064] In the above description, when multiple detection points are detected for the same subject, they are integrated and the midpoint of the integrated points is used as a new detection point, but this is not limited to this. For example, an integration method may be used in which one of the multiple detection points to be integrated is used and the others are not used. In this case, it is desirable to use the detection point that was also detected at the previous time point as the detection point to be used in order to ensure data continuity.
[0065] 10 is a flowchart showing the process of generating tracking information by the tracking unit 9. It is assumed that this process is performed for each captured image corresponding to the three-dimensional shape.
[0066] In step S1001, the tracking unit 9 acquires the detection points from the subject position detection unit 8.
[0067] In step S1002, the tracking unit 9 determines whether the shooting time associated with the detection point to be processed is a detection point of the shooting time at the start of shooting. The method of determination is not particularly limited. The shooting start time may be set in advance, and a determination may be made as to whether it matches the shooting time corresponding to the acquired detection point. Alternatively, a variable N=0 may be set when shooting starts, and N=N+1 may be set after the processing of step S1007 described later, thereby counting the number of repetitions of the processing described in FIG. 10, and making a determination based on the number of repetitions. If it is a detection point of the shooting time at the start of shooting, proceed to step S1005. If it is not a detection point of the shooting time at the start of shooting, proceed to step S1003.
[0068] In step S1003, the tracking unit 9 acquires the tracking information of the previous shooting time from the storage unit 4. However, this is not limited to the above, and the immediately preceding tracking information may be retained.
[0069] In step S1004, the tracking unit 10 compares the three-dimensional position of the detection point acquired in step S1001 with the three-dimensional position of the detection point included in the tracking information acquired in step S1003. If the detection point at the previous image capture time is within a predetermined range of the detection point acquired in step S1001, the tracking unit 10 assigns the same tracking identifier to the detection point at the previous image capture time as the detection point at the previous image capture time. Note that the predetermined range may be set based on the traveling direction, as in the process of integrating multiple detection points. Furthermore, instead of being limited to the predetermined range, the tracking identifier of the detection point at the previous image capture time that is closest to the detection point at step S1001 may be assigned.
[0070] In step S1005, the tracking unit 9 randomly assigns tracking identifiers to the detection points acquired in step S1001. If multiple detection points are acquired, a different tracking identifier is assigned to each of them.
[0071] In step S1006, the tracking unit 9 generates tracking information including the detection points acquired in step S1001 and the tracking identifiers assigned to the detection points.
[0072] In step S 1007 , the tracking unit 9 outputs the tracking information to the storage unit 4 .
[0073] By the above process, tracking information at each photographing time can be generated.
[0074] FIG. 11 is a flowchart showing the process by which the smoothing unit 11 associates tracking identifiers with object identifiers. It is assumed that this process is performed sequentially for each shooting time of captured images corresponding to a three-dimensional shape. The smoothing unit 11 also stores a correspondence table indicating combinations of tracking identifiers and object identifiers in advance. The correspondence table can be created by acquiring tracking identifiers by obtaining tracking information corresponding to the shooting time at the start of shooting. The object identifiers in the correspondence table are updated each time object identification information is acquired. The correspondence table is then used to identify the object identifiers corresponding to the tracking identifiers. The reason for using the correspondence table to identify the object identifiers is that there may be shooting times for which there is no object identification information. Even if there is no object identification information at the shooting time to be processed, the correspondence table can be used to identify the object identifiers corresponding to the tracking identifiers.
[0075] In step S1101, the smoothing unit 11 acquires the tracking information from the storage unit 4.
[0076] In step S1102, the smoothing unit 11 determines whether or not there is subject identification information in the storage unit 4. If there is subject identification information, the process proceeds to step S1103. If there is no subject identification information, the process proceeds to step S1106.
[0077] In step S1103, the smoothing unit 11 acquires the subject identification information from the storage unit 4. Note that this process may be combined with the process of step S1102 into one process.
[0078] In step S1104, the smoothing unit 11 compares the tracking information acquired in step S1101 with the subject identification information acquired in step S1103. Since the tracking information and the subject identification information contain detection points, a pair of tracking identifier and subject identifier that contain the same detection points is identified.
[0079] In step S1105, the smoothing unit 11 updates the correspondence table using the pair of tracking identifier and subject identifier identified in step S1104.
[0080] In step S1106, the smoothing unit 11 identifies the object identifier corresponding to the tracking identifier included in the tracking information acquired in step S1101 based on the correspondence table.
[0081] In step S1107, the smoothing unit 11 outputs to the storage unit 4 information indicating the combination of the tracking identifier and the object identifier identified in step S1106.
[0082] By the above process, the correspondence table is updated in the order of the photographing time, so that even at a photographing time for which there is no photographic subject identification information, the photographic subject identifier can be identified using the correspondence table that was updated immediately before.
[0083] 12 is a diagram showing a correspondence table showing combinations of tracking identifiers and subject identifiers. Tracking identifiers Tracking A, Tracking B, and Tracking C correspond to subject identifiers Subject A, Subject B, and Subject C. Note that the above combinations are just an example, and it is sufficient that one tracking identifier corresponds to one subject identifier, for example, Tracking A may correspond to Subject B. For shooting times with subject identification information, the correspondence table is updated using the tracking information and subject identification information.
[0084] In this embodiment, an example of updating the correspondence table has been described, but the present invention is not limited to this. For example, the smoothing unit 11 may store the subject identification information acquired immediately before the photographing time of the subject to be processed, which is the subject identification information at a photographing time before the photographing time of the subject to be processed. In this case, the smoothing unit 11 records which tracking identifier the subject identification information acquired immediately before is associated with, and uses this information to identify the subject identifier that corresponds to the tracking identifier at the photographing time of the subject to be processed.
[0085] The above configuration makes it easy to track a subject in a wide variety of shooting environments. As described in the examples, a subject can be tracked even in shooting environments where the floor is inclined and there are differences in elevation depending on the position. As a result, based on the tracking information and subject identification information, it is possible to superimpose the subject's speed change and trajectory, as described above, and provide a virtual viewpoint image with high added value for analysis and viewing experience.
[0086] In this embodiment, the tracking information and the subject identification information are recorded in the storage unit 4, but this is not limiting. The tracking information and the subject identification information may be recorded together as one piece of information. Specifically, the position information of the detection point, the tracking identifier, and the subject identifier may be recorded in association with each other.
[0087] (Other forms) In the above embodiment, a specific example of the image processing system is shown, but the present invention is not necessarily limited to the embodiment shown here.
[0088] For example, the subject position detection unit 8 may obtain the shape estimation result of the three-dimensional shape estimation unit 3 from the shape estimation result that the three-dimensional shape estimation unit 3 has accumulated in the accumulation unit 4.
[0089] In the above embodiment, the highest point within a predetermined area of the three-dimensional shape is selected, but the lowest point may also be selected. Specifically, a distance image is created as if the three-dimensional shape were viewed from below, and the point where the distance is smallest (the point where the minimum value is reached) within a predetermined range is detected as the detection point. By processing in this manner, the tire contact surface can be detected in bicycle racing. This makes it possible to detect the subject position with little error, regardless of the rider's posture, etc. However, in this case, since the detection point is near the tire contact surface, when acquiring color information in the identification setting unit 10, it is desirable to acquire color information at a predetermined height position from the position of the detection point.
[0090] In the above embodiment, the tracking unit 9 performs tracking by referencing the detection points from the previous time. However, there may be cases where the number of detection points output by the subject position detection unit 8 is incorrect. For example, due to an estimation error in the three-dimensional shape estimation unit 3, the subject position detection unit 8 may detect incorrectly, resulting in no detection points for a certain subject at a particular shooting time or an increased number of detection points. Another possible scenario is the detection of a false shape generated by floating dust. To address such cases, the tracking unit 9 preferably performs the following processing. If a detection point disappears at the shooting time immediately before the target shooting time, the tracking unit 9 interpolates the detection point by assuming that the detection point at the previous shooting time is moving in the same direction as the previous speed, and records the tracking information in the storage unit 4. Furthermore, if a detection point that was not present at the previous shooting time appears, there is a possibility of false detection. Therefore, whether or not the object is a subject is determined based on whether or not it continues to appear continuously (for example, for 10 frames). If the object is determined to be a subject, a new tracking identifier is assigned as the starting point for tracking and recorded in the storage unit 4. By performing these processes, the tracking unit 9 can perform appropriate processing in response to an increase or decrease in the number of detection points.
[0091] In the above embodiment, the smoothing unit 11 is configured to acquire the tracking information and subject identification information recorded in the storage unit 4 and check their correspondence, but this is not necessarily limited to this. For example, when the identification setting unit 10 generates subject identification information, it may be configured to correspond to the tracking identification information and re-record it as tracking information with a tracking identifier assigned. This configuration simplifies the processing of the smoothing unit 11 that handles information, but if an abnormality occurs in the tracking information or the subject identification information, it may be difficult to determine which one has the abnormality and to identify the cause. For this reason, it is desirable to record each piece of information separately and have the smoothing unit 11, which uses this information, check and correct the abnormality.
[0092] In the above embodiment, the smoothing unit 11 is configured to have a means for smoothing the positions of the detection points and a means for calculating the velocity, but the smoothing unit 11 does not necessarily need to include the means for calculating the velocity, and may have a separate means for calculating the velocity.
[0093] In the above embodiment, bicycle racing has been used as an example, but the subject of photography is not limited to bicycle racing. It can also be applied to photography environments such as other sporting events and concerts. In particular, since track and field hurdle races and obstacle races are events in which the subject is high above the floor, there is a high possibility that desirable results can be obtained by utilizing this embodiment.
[0094] In the above embodiment, the subject identifiers are assigned automatically by the identification setting unit 10, but this is not necessarily limited to this. The identification setting unit 10 may be provided with a user interface, and may assign subject identifiers to detection points based on user operation, and record the identifiers as subject identification information in the storage unit 4. For example, a method of assignment when there are multiple subjects will be described.
[0095] FIG. 7 is a diagram showing a display screen when object identifiers are assigned to detection points based on a user operation. For example, a user instructs a transition to an object identifier assignment mode from a graphical user interface such as that shown in FIG. 7(a). Then, the user specifies the object identifiers to be assigned by, for example, sequentially clicking on detection points 701-703, which are displayed on a screen such as that shown in FIG. 7(a). Specifically, when assigning object identifiers A to C to the detection points 701-703 in that order, the user clicks on the detection points 701-703 in that order. The identification setting unit 10 uses the clicked order as input and assigns identifiers based on that order, as shown in FIG. 7(b). The identification setting unit 10 then records information associating the object identifiers with the detection points in the storage unit 4 as object identification information. This allows the user to manually set identifiers even when the identification setting unit 10 does not automatically identify objects or when objects cannot be identified based on simple color information, for example.
[0096] In the above embodiment, the superimposed image generating unit 12 generates the images to be superimposed, and then the images are synthesized by the video generating unit 6. However, this is not necessarily limited to this, and the functions of the superimposed image generating unit 12 may be provided in the video generating unit 6.
[0097] For example, the viewpoint instruction unit 5 may acquire and use the tracking information and subject identification information stored in the storage unit 4. In this case, the viewpoint instruction unit 5 may generate a virtual viewpoint that can always orbit around the subject even if the subject moves, by setting the position of the rotation center of the virtual viewpoint to the position 800 of the subject's detection point, as shown in FIG. 8(a), for example. The viewpoint instruction unit 5 may also be configured to set the line of sight of the virtual viewpoint to the position 800 of the subject's detection point, as shown in FIG. 8(b), for example. In this case, the image processing system can generate a virtual viewpoint image of a viewpoint in which the virtual viewpoint, placed in a semi-fixed position, automatically rotates horizontally as the subject moves.
[0098] (Other configurations) In the above embodiment, each processing unit shown in Fig. 1 is configured as hardware, but the processing performed by each processing unit shown in these figures may be configured as a computer program.
[0099] FIG. 9 is a block diagram showing an example of the hardware configuration of a computer that can be applied to the image processing apparatus according to each of the above embodiments.
[0100] The CPU 901 controls the entire computer using computer programs and data stored in the RAM 902 and the ROM 903, and also executes the processes described above as being performed by the indirect position estimation device according to each of the above embodiments. That is, the CPU 901 functions as each processing unit shown in FIG.
[0101] The RAM 902 has an area for temporarily storing computer programs and data loaded from an external storage device 906, data acquired from the outside via an I / F (interface) 907, etc. The RAM 902 also has a work area used when the CPU 901 executes various processes. That is, the RAM 902 can be allocated as a frame memory, for example, or can provide various other areas as needed.
[0102] The ROM 903 stores setting data for the computer, a boot program, and the like. The operation unit 904 is composed of a keyboard, a mouse, and the like, and allows a user to operate the computer to input various instructions to the CPU 901. The output unit 905 displays the results of processing by the CPU 901. The output unit 905 is composed of, for example, a liquid crystal display. For example, the viewpoint instruction unit 5 is composed of the operation unit 904, and the display unit 7 is composed of the output unit 905.
[0103] The external storage device 906 is a large-capacity information storage device, such as a hard disk drive. The external storage device 906 stores an operating system (OS) and computer programs for causing the CPU 901 to implement the functions of the various units shown in Fig. 1. Furthermore, the external storage device 906 may also store image data to be processed.
[0104] Computer programs and data stored in the external storage device 906 are loaded into the RAM 902 as appropriate under the control of the CPU 901, and become the subject of processing by the CPU 901. The I / F 907 can be connected to a network such as a LAN or the Internet, or to other devices such as a projector or display device, and the computer can acquire and send various information via this I / F 907. In the first embodiment, the imaging unit 1 is connected to this, and captured images are input and controlled. 908 is a bus connecting the above-mentioned units.
[0105] The operation of the above-described configuration is controlled mainly by the CPU 901 as explained in the previous embodiment.
[0106] In another configuration, the above-described functions can be achieved by providing a storage medium containing computer program code for implementing the functions to a system, and the system then reading and executing the computer program code. In this case, the computer program code itself read from the storage medium implements the functions of the above-described embodiments, and the storage medium storing the computer program code constitutes the present disclosure. Also included is a case in which an operating system (OS) running on a computer performs some or all of the actual processing based on the instructions of the program code, thereby implementing the above-described functions.
[0107] The disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) an acquisition means for acquiring first shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a first photographing time, and second shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a second photographing time; a generating means for generating, based on the first shape information, first distance information indicating a distance from a first position in the virtual space to a three-dimensional shape corresponding to the first shape information, and second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information, based on the second shape information; a setting means for setting a first tracking point of a three-dimensional shape corresponding to the first shape information at the first imaging time based on the first distance information, setting a second tracking point of a three-dimensional shape corresponding to the second shape information at the second imaging time based on the second distance information, and setting the same identifier as the second tracking point to the first tracking point when a distance between a position of the first tracking point and a position of the second tracking point is equal to or smaller than a predetermined value; An image processing device having: (Configuration 2) the first distance information is information indicating distances from the first position to a plurality of first components that form a three-dimensional shape corresponding to the first shape information, 2. The image processing device according to claim 1, wherein the second distance information is information indicating distances from the second position to a plurality of second components that constitute a three-dimensional shape corresponding to the second shape information. (Configuration 3) the first tracking point is a first component having the smallest distance or a first component having the largest distance among the plurality of first components; 3. The image processing device according to configuration 2, wherein the second tracking point is the second component with the smallest distance or the second component with the largest distance among the plurality of second components. (Configuration 4) the first tracking point is a first component included in a predetermined area among the plurality of first components; 4. The image processing device according to configuration 2 or 3, wherein the second tracking point is a second component element included in a predetermined area among the plurality of second components. (Configuration 5) the plurality of first components and the plurality of second components are classified into a plurality of regions; 4. The image processing device according to configuration 2 or 3, wherein the first tracking point and the second tracking point are set for each of the plurality of regions. (Configuration 6) 6. The image processing device according to configuration 5, wherein the setting means collectively sets a plurality of first tracking points that are included in a predetermined range among the plurality of first tracking points as one first tracking point. (Configuration 7) 7. The image processing device according to configuration 6, wherein the setting means sets one first tracking point that combines the plurality of first tracking points at a center of gravity position of the plurality of first tracking points included in the predetermined range. (Configuration 8) 8. The image processing device according to configuration 6 or 7, wherein the predetermined range is a range centered on the first tracking point. (Configuration 9) 9. The image processing device according to any one of configurations 6 to 8, wherein the predetermined range differs for each subject to be photographed. (Configuration 10) The subject of the photograph is a bicycle race, 10. The image processing device according to configuration 9, wherein the predetermined range is set along the course of a bicycle racing track. (Configuration 11) the first shape information indicates three-dimensional shapes of a plurality of subjects; 11. The image processing device according to any one of configurations 1 to 10, wherein the setting means sets the first tracking points in the same number as the number of the plurality of subjects. (Configuration 12) 12. The image processing device according to any one of configurations 1 to 11, wherein the identifier is an identifier that indicates a subject. (Configuration 13) 13. The image processing device according to any one of configurations 1 to 12, wherein the first position is generated based on a bounding box that encloses a three-dimensional shape corresponding to the first shape information. (Configuration 14) 14. The image processing device according to configuration 13, wherein the first position is set at a position that is a predetermined distance away from the center of the upper surface of the bounding box. (Configuration 15) The image processing device according to any one of configurations 1 to 14, characterized in that the first position is set based on a three-dimensional shape of a background identified based on the position of a three-dimensional shape corresponding to the first shape information. (Configuration 16) the three-dimensional shape of the background represents a velodrome; The image processing device described in configuration 15, characterized in that the first distance information is information indicating the distance from the first position to the three-dimensional shape corresponding to the first shape information in a direction perpendicular to the course of a velodrome on which the three-dimensional shape corresponding to the first shape information is located. (Configuration 17) 17. The image processing device according to any one of configurations 1 to 16, further comprising an output unit that outputs position information indicating the position of the first tracking point and information indicating the identifier. (method) an acquisition step of acquiring first shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a first photographing time, and second shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a second photographing time; a generation process of generating first distance information indicating a distance from a first position in the virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information, and second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information based on the second shape information; a setting step of setting a first tracking point of the three-dimensional shape corresponding to the first shape information at the first imaging time based on the first distance information, setting a second tracking point of the three-dimensional shape corresponding to the second shape information at the second imaging time based on the second distance information, and setting the same identifier as the second tracking point to the first tracking point when a distance between a position of the first tracking point and a position of the second tracking point is equal to or less than a predetermined value; An image processing method comprising: (program) A program for causing a computer to function as each of the means of the image processing device according to any one of configurations 1 to 17. [Explanation of symbols]
[0108] 3 Three-dimensional shape estimation section 4. Storage unit 11 Subject position detection unit 12 Tracking Department 13 Identification setting section 14 Smooth section
Claims
1. an acquisition means for acquiring first shape information indicating a three-dimensional shape of a subject generated based on a plurality of captured images at a first photographing time, and second shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a second photographing time; a generating means for generating first distance information indicating a distance from a first position in the virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information, and second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information based on the second shape information; a setting means for setting a first tracking point of a three-dimensional shape corresponding to the first shape information at the first imaging time based on the first distance information, setting a second tracking point of a three-dimensional shape corresponding to the second shape information at the second imaging time based on the second distance information, and setting the same identifier as the second tracking point to the first tracking point when a distance between a position of the first tracking point and a position of the second tracking point is equal to or smaller than a predetermined value; An image processing device having:
2. the first distance information is information indicating distances from the first position to a plurality of first components that configure a three-dimensional shape corresponding to the first shape information, 2 . The image processing apparatus according to claim 1 , wherein the second distance information is information indicating distances from the second position to a plurality of second components that form the three-dimensional shape corresponding to the second shape information.
3. the first tracking point is a first component having the smallest distance or a first component having the largest distance among the plurality of first components; The image processing device according to claim 2 , wherein the second tracking point is a second component with the smallest distance or a second component with the largest distance among the plurality of second components.
4. the first tracking point is a first component included in a predetermined area among the plurality of first components, 3. The image processing apparatus according to claim 2, wherein the second tracking point is a second component element included in a predetermined area among the plurality of second components.
5. the plurality of first components and the plurality of second components are classified into a plurality of regions; The image processing device according to claim 2 , wherein the first tracking point and the second tracking point are set for each of the plurality of regions.
6. 6. The image processing apparatus according to claim 5, wherein said setting means collectively sets a plurality of first tracking points included in a predetermined range as one first tracking point.
7. 7. The image processing apparatus according to claim 6, wherein the setting means sets one first tracking point, which is a group of the first tracking points, at a center of gravity of the plurality of first tracking points included in the predetermined range.
8. 7. The image processing apparatus according to claim 6, wherein the predetermined range is a range centered on the first tracking point.
9. 7. The image processing apparatus according to claim 6, wherein the predetermined range differs for each subject to be photographed.
10. The subject of the photograph is a bicycle race, 10. The image processing device according to claim 9, wherein the predetermined range is set along a course of a bicycle racing track.
11. the first shape information indicates three-dimensional shapes of a plurality of subjects; 2. The image processing apparatus according to claim 1, wherein said setting means sets the first tracking points in the same number as the number of the plurality of subjects.
12. 2. The image processing device according to claim 1, wherein the identifier is an identifier that indicates a subject.
13. The image processing device according to claim 1 , wherein the first position is generated based on a bounding box that encloses a three-dimensional shape corresponding to the first shape information.
14. The image processing device according to claim 13 , wherein the first position is set at a position that is a predetermined distance away from the center of the upper surface of the bounding box.
15. The image processing device according to claim 1 , wherein the first position is set based on a three-dimensional shape of a background that is specified based on a position of the three-dimensional shape that corresponds to the first shape information.
16. the three-dimensional shape of the background represents a velodrome; 16. The image processing device according to claim 15, wherein the first distance information is information indicating a distance from the first position to the three-dimensional shape corresponding to the first shape information in a direction perpendicular to a course of a bicycle racing track on which the three-dimensional shape corresponding to the first shape information is located.
17. 2. The image processing apparatus according to claim 1, further comprising an output unit for outputting position information indicating the position of the first tracking point and information indicating the identifier.
18. an acquisition step of acquiring first shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a first photographing time, and second shape information indicating a three-dimensional shape of the subject generated based on a plurality of captured images at a second photographing time; a generation process of generating first distance information indicating a distance from a first position in the virtual space to a three-dimensional shape corresponding to the first shape information based on the first shape information, and second distance information indicating a distance from a second position in the virtual space to a three-dimensional shape corresponding to the second shape information based on the second shape information; a setting step of setting a first tracking point of the three-dimensional shape corresponding to the first shape information at the first imaging time based on the first distance information, setting a second tracking point of the three-dimensional shape corresponding to the second shape information at the second imaging time based on the second distance information, and setting the same identifier as the second tracking point to the first tracking point when a distance between a position of the first tracking point and a position of the second tracking point is equal to or smaller than a predetermined value; An image processing method comprising:
19. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image processing apparatus, control method, and program
JP2024055093A