Information Processing Apparatus, Information Processing Method, and Storage Medium

The subject information is acquired through multiple imaging devices, distinguish visible and obstructed areas, and generate virtual viewpoint images, solving the problem of inaccurate display of blind spot areas in the prior art, and realizing more accurate virtual viewpoint image display.

CN113347348BActive Publication Date: 2025-07-08CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110171250.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-18
Filing Date
2021-02-08
Publication Date
2025-07-08
Estimated Expiration
2041-05-12

AI Technical Summary

Technical Problem

The prior art cannot effectively reflect the invisible area in the imaging area due to the subject blocking, resulting in inaccurate display of the blind spot area.

Method used

By taking images from different directions by multiple imaging devices, the position and shape information of the subject is obtained, the visible area and the occlusion area are distinguished, and a virtual viewpoint image is generated to display the visibility information.

Benefits of technology

It is able to accurately distinguish areas that can be seen from specific locations and areas that cannot be seen due to occlusion, improving the display accuracy of virtual viewpoint images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113347348B_ABST
    Figure CN113347348B_ABST
Patent Text Reader

Abstract

The present invention relates to an information processing apparatus, an information processing method, and a storage medium. The information processing apparatus obtains subject information indicating the position and shape of a subject in a shooting area, where the shooting area is shot from different directions by a plurality of imaging devices; based on the subject information, a visible area that can be seen from a specific position in the shooting area and an occlusion area that cannot be seen from the specific position due to being occluded by the subject are distinguished; a virtual viewpoint image for displaying information based on the discrimination result is generated, and the virtual viewpoint image is based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating the position and direction of the virtual viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a storage medium. In particular, the present disclosure relates to processing of information obtained by photographing a subject from multiple directions using multiple imaging devices. Background Art

[0002] There is a technique of synchronously photographing from multiple viewpoints using multiple imaging devices arranged at different positions, and generating a virtual viewpoint image in which the viewpoint can be freely changed using multiple images obtained by photographing. For example, based on multiple captured images of games such as soccer and basketball, a virtual viewpoint image corresponding to a viewpoint set by a user is generated, so that the user can watch the game from different viewpoints. The position and direction of the virtual viewpoint are set according to the position and direction of a specific subject (such as a soccer player and a referee, etc.) in the photographed area, so that a virtual viewpoint image that reproduces the field of view of the player or the referee can be generated.

[0003] Japanese Patent No. 6555513 discusses providing video content corresponding to the viewpoint of a specific player using data obtained by photographing a soccer game from multiple directions. Japanese Patent No. 6555513 also discusses estimating a field of view defined by the viewpoint coordinates, line-of-sight direction, and viewing angle of other players, and displaying a blind spot area that does not exist in the field of view of other players in the video content.

[0004] However, the blind spot area specified by the method discussed in Japanese Patent No. 6555513 does not reflect the invisible area blocked by the subject in the photographed area. For example, even if the area is in the direction close to the line of sight of the player, the player cannot see the area blocked by other players from the viewpoint of the player. Therefore, even an area that is not included in the blind spot area displayed in the video content discussed in Japanese Patent No. 6555513 may include an area representing an actual blind spot. Summary of the Invention

[0005] According to one aspect of the present disclosure, a technique is provided that can easily distinguish an area that can be seen from a specific position and an area that cannot be seen due to occlusion.

[0006] According to an aspect of the present disclosure, there is provided an information processing apparatus including: an acquisition unit configured to acquire subject information representing a position and a shape of a subject in a shooting area, where the shooting area is shot from different directions by a plurality of imaging devices; a discrimination unit configured to discriminate, based on the subject information acquired by the acquisition unit, a visible area visible from a specific position in the shooting area and an occluded area that cannot be seen from the specific position due to being occluded by the subject; and a generation unit configured to generate a virtual viewpoint image displaying information based on a discrimination result of the discrimination unit, the virtual viewpoint image being based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating a position and a direction of a virtual viewpoint.

[0007] According to another aspect of the present disclosure, there is provided an information processing apparatus including: an acquisition unit configured to acquire subject information representing a position and a shape of a subject in a shooting area, the subject information being generated based on image data obtained by shooting the shooting area from different directions by a plurality of imaging devices, a generation unit configured to generate information capable of discriminating a visible area visible from a specific position corresponding to a designated subject in the shooting area and an occluded area that cannot be seen from the specific position due to being occluded by other subjects based on the subject information acquired by the acquisition unit; and an output unit configured to output the information generated by the generation unit.

[0008] According to another aspect of the present disclosure, there is provided an information processing method including: acquiring subject information representing a position and a shape of a subject in a shooting area, where the shooting area is shot from different directions by a plurality of imaging devices; discriminating, based on the subject information acquired through the acquisition, a visible area visible from a specific position in the shooting area and an occluded area that cannot be seen from the specific position due to being occluded by the subject; and generating a virtual viewpoint image displaying information based on a result of the discrimination, the virtual viewpoint image being based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating a position and a direction of a virtual viewpoint.

[0009] According to another aspect of the present disclosure, there is provided an information processing method including: acquiring subject information representing a position and a shape of a subject in a shooting area, the subject information being generated based on image data obtained by shooting the shooting area from different directions by a plurality of imaging devices, generating information capable of discriminating a visible area visible from a specific position corresponding to a designated subject in the shooting area and an occluded area that cannot be seen from the specific position due to being occluded by other subjects based on the subject information acquired through the acquisition; and outputting the information generated through the generation.

[0010] According to another aspect of the present disclosure, a storage medium is provided which stores a program for causing a computer to function as the above-described information processing apparatus.

[0011] Other features of the present disclosure will become clear from the following description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 An example of the configuration of the image processing system is illustrated.

[0013] Figure 2 An example of the hardware configuration of the information processing apparatus is illustrated.

[0014] Figure 3A and Figure 3B are flowcharts illustrating operation examples of the information processing apparatus and the display apparatus.

[0015] Figure 4A and Figure 4B illustrate processing for visibility determination.

[0016] Figure 5A and Figure 5B illustrate an example of visibility information.

[0017] Figure 6 is a flowchart illustrating an operation example of the information processing apparatus.

[0018] Figure 7 illustrate an example of visibility information.

[0019] Figure 8 illustrate an example of output of visibility information. DETAILED DESCRIPTION

[0020] [System Configuration]

[0021] Figure 1An example of the configuration of the image processing system 1 is illustrated. The image processing system 1 generates a virtual viewpoint image representing a view from a specified virtual viewpoint. The image processing system 1 generates the virtual viewpoint image based on the specified virtual viewpoint and a plurality of images (multi-viewpoint images) obtained by imaging with a plurality of imaging devices. The virtual viewpoint image according to the present exemplary embodiment is also referred to as a free viewpoint video. However, the virtual viewpoint image is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by a user, but includes, for example, an image corresponding to a viewpoint selected by the user from a plurality of candidates. According to the present exemplary embodiment, the case of specifying the virtual viewpoint by a user operation is mainly described, but the virtual viewpoint can be automatically specified based on, for example, the result of image analysis. Further, according to the present exemplary embodiment, the case where the virtual viewpoint image is a still image is mainly described, but the virtual viewpoint image can be a moving image. In other words, the image processing system 1 can process both still images and moving images. In the following description, unless otherwise specified, the term "image" will be described as a concept including both moving images and still images.

[0022] The image processing system 1 includes a plurality of imaging devices 100, an image processing device 110, an information processing device 200, and a display device 300. The plurality of imaging devices 100 image a subject to be imaged in an imaging area from a plurality of directions. Examples of the imaging area include a stadium where a football game and a karate game are held, and a stage where a concert and a performance are held. According to the present exemplary embodiment, the case where the imaging area is a football field is mainly described. The plurality of imaging devices 100 are each disposed at different positions to surround the imaging area and image synchronously with each other. The plurality of imaging devices 100 may be arranged not to cover the entire periphery of the imaging area, but only in the directions of a part of the imaging area depending on restrictions such as the installation position. The number of imaging devices is not limited to Figure 1 the illustrated example. For example, in the case where the imaging area is a football field, about 30 imaging devices may be arranged around the field. Further, imaging devices having different functions, such as a telephoto camera and a wide-angle camera, may also be provided.

[0023] Each of the plurality of imaging devices 100 according to the present exemplary embodiment is a camera having an independent housing and can image an image from a single viewpoint. However, the present exemplary embodiment is not limited to such a configuration, and two or more imaging devices may be included in the same housing. For example, a single camera including a plurality of lens groups and a plurality of sensors and imaging from a plurality of viewpoints may be provided as the plurality of imaging devices 100.

[0024] The image processing device 110 generates subject information 101 based on the image data obtained from a plurality of imaging devices 100. The subject information 101 includes foreground subjects (such as people and balls) and background subjects (such as stadiums and fields) in the imaging area. The subject information 101 represents the three-dimensional positions and three-dimensional shapes of the respective subjects. The subject information 101 of each subject can be stored individually, or the subject information 101 of a plurality of subjects in the imaging area can be stored collectively.

[0025] The information processing device 200 includes a subject information acquisition unit 201 that acquires the subject information 101 from the image processing device 110, a field of view information acquisition unit 202 that acquires the field of view information 102, and a virtual viewpoint acquisition unit 203 that acquires the virtual viewpoint information 103. The information processing device 200 further includes a visibility determination unit 204 and an information generation unit 205.

[0026] The field of view information 102 represents the field of view corresponding to a predetermined viewpoint in the imaging area. The field of view information 102 is a reference for the following visibility determination. The field of view information 102 includes information indicating the position and direction (the line-of-sight direction from the viewpoint position) of the viewpoint in the horizontal and vertical directions in the three-dimensional space, and the viewing angle corresponding to the size of the field of view. The content of the field of view information 102 is not limited to the above content. The size of the field of view can be represented by the focal length or the zoom value. In addition, the field of view information 102 can include information indicating the distances from the viewpoint position to the near plane and the far plane that are the boundaries of the field of view. The field of view information 102 can be specified based on a user operation. At least a part of the field of view information 102 can be specified automatically. For example, the image processing system 1 can receive a user operation for specifying the position, direction, and viewing angle of the viewpoint, and generate the field of view information 102.

[0027] For example, the image processing system 1 can receive a user operation for designating a specific person as a subject in the imaging area, and generate the field of view information 102 of the designated subject, such as the field of view information 102 representing the field of view of the specific person. In this case, the viewpoint position indicated by the field of view information 102 corresponds to the position of the head or eyes of the designated subject, and the line-of-sight direction indicated by the field of view information 102 corresponds to the direction of the head or eyes of the designated subject.

[0028] The field of view information 102 representing the field of view of a person can be generated based on data obtained from sensors (e.g., action cameras, electronic compasses, and global positioning systems (GPS)) for detecting the position and orientation of the head or eyes of the person. The field of view information 102 can be generated based on the analysis result of the image data obtained from the imaging device 100 or the subject information 101 generated by the image processing device 110. For example, the coordinate position of the head of the person (the viewpoint position indicated by the field of view information 102) can be a three-dimensional coordinate position that is offset a predetermined distance forward and downward from the position of the top of the head (detected by the sensor or through image analysis) to the front of the face. In addition, for example, the coordinate position of the head can be a three-dimensional coordinate position that is offset a predetermined distance forward from the detected center position of the head to the front of the face. In addition, the coordinate position of the eyes of the person, which is the viewpoint position indicated by the field of view information 102, can be a three-dimensional coordinate position between the left and right eyes detected by the sensor or through image analysis. The image processing system 1 can obtain the field of view information 102 from the outside.

[0029] The virtual viewpoint information 103 indicates the position and orientation (line-of-sight direction) of the virtual viewpoint corresponding to the virtual viewpoint image. Specifically, the virtual viewpoint information 103 includes a parameter set including parameters representing the position of the virtual viewpoint in the three-dimensional space and parameters representing the orientation of the virtual viewpoint in the pan, tilt, and roll directions. The content of the virtual viewpoint information 103 is not limited to the above information. For example, the parameter set of the virtual viewpoint information 103 can include parameters representing the size of the field of view corresponding to the virtual viewpoint, such as the zoom factor and focal length, and parameters representing time. The virtual viewpoint information 103 can be specified based on a user operation, or at least a part of the virtual viewpoint information 103 can be automatically specified. For example, the image processing system 1 can receive a user operation for specifying the position, orientation, and viewing angle of the virtual viewpoint and generate the virtual viewpoint information 103. The image processing system 1 can obtain the virtual viewpoint information 103 from the outside.

[0030] The visibility determination unit 204 performs visibility determination using the subject information 101 obtained by the subject information acquisition unit 201, the field of view information 102 obtained by the field of view information acquisition unit 202, and the virtual viewpoint information 103 obtained by the virtual viewpoint acquisition unit 203. The visibility determination performed by the visibility determination unit 204 is a process for distinguishing, in the imaging area, the area visible from the viewpoint indicated by the field of view information 102 from the area not visible from the viewpoint indicated by the field of view information 102. The information generation unit 205 generates information corresponding to the result of the visibility determination generated by the visibility determination unit 204 and outputs this information to the display device 300. The information generation unit 205 can generate and output information corresponding to the virtual viewpoint using the visibility determination result and the virtual viewpoint information 103. The visibility determination process and the information generated by the information generation unit 205 will be described in detail later.

[0031] The display device 300 is a computer having a display, and the display device 300 combines the information obtained from the information generation unit 205 with the virtual viewpoint image 104 for display. Computer graphics (CG) representing the field of view information 102 can be included in the image to be displayed by the display device 300. The display device 300 can include a visual presentation device such as a projector instead of a display.

[0032] The image processing system 1 generates the virtual viewpoint image 104 based on the image data obtained by the imaging device 100 and the virtual viewpoint information 103. For example, the virtual viewpoint image 104 is generated by the following method. A foreground image and a background image are obtained by extracting a foreground region corresponding to a predetermined subject (such as a person and a ball) and a background region other than the foreground region from multi-viewpoint images respectively obtained by imaging with a plurality of imaging devices 100 from different directions. A foreground model representing the three-dimensional shape of the predetermined subject and texture data for coloring the foreground model are generated based on the foreground image. Texture data for coloring a background model representing the three-dimensional shape of the background (such as a stadium) is generated based on the background image. Then, the texture data is mapped onto the foreground model and the background model and rendered according to the virtual viewpoint indicated by the virtual viewpoint information 103, thereby generating the virtual viewpoint image 104. However, the method for generating the virtual viewpoint image 104 is not limited to the above process, and various methods can be used, such as a method for generating the virtual viewpoint image 104 by performing projective transformation on the captured images without using a three-dimensional model. The image processing system 1 can obtain the virtual viewpoint image 104 from the outside.

[0033] The configuration of the image processing system 1 is not limited to Figure 1The example shown. For example, the image processing device 110 and the information processing device 200 can be integrally constructed. Additionally, the field of view information 102, the virtual viewpoint information 103, and the virtual viewpoint image 104 can be generated by the information processing device 200. Alternatively, the information generated by the information generation unit 205 and the virtual viewpoint image 104 can be combined and output to the display device 300. The image processing system 1 includes a storage unit that stores the subject information 101, the field of view information 102, the virtual viewpoint information 103, and the virtual viewpoint image 104, and an input unit that receives user operations.

[0034] [Hardware Configuration]

[0035] Figure 2 An example of the hardware configuration of the information processing device 200 is illustrated. The hardware configurations of the image processing device 110 and the display device 300 are similar to the configuration of the information processing device 200 described below. The information processing device 200 includes a central processing unit (CPU) 211, a read-only memory (ROM) 212, a random access memory (RAM) 213, an auxiliary storage device 214, a display unit 215, an operation unit 216, a communication interface (I / F) 217, and a bus 218.

[0036] The CPU 211 uses the computer programs and data stored in the ROM 212 and the RAM 213 to control the entire information processing device 200, thereby implementing Figure 1 each function of the information processing device 200 shown. The information processing device 200 can include one or more pieces of dedicated hardware different from the CPU 211, and the dedicated hardware can execute at least a part of the processing executed by the CPU 211. Examples of the dedicated hardware include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and a digital signal processor (DSP). The ROM 212 stores programs that do not need to be changed, for example. The RAM 213 temporarily stores programs and data provided from the auxiliary storage device 214, for example, and data provided from the outside via the communication I / F 217. The auxiliary storage device 214 includes, for example, a hard disk drive, and stores various types of data such as image data and audio data.

[0037] The display unit 215 includes, for example, a liquid crystal display and a light emitting diode (LED), and displays, for example, a graphical user interface (GUI) for a user to operate the information processing apparatus 200. The operation unit 216 includes, for example, a keyboard, a mouse, a joystick, and a touch panel, and inputs various instructions to the CPU 211 when a user operation is received. The CPU 211 serves as a display control unit for controlling the display unit 215 and an operation control unit for controlling the operation unit 216. The communication I / F 217 is used for communication between the information processing apparatus 200 and an external apparatus. For example, when the information processing apparatus 200 is wired-connected to an external apparatus, a communication cable is connected to the communication I / F 217. When the information processing apparatus 200 has a function of wireless communication with an external apparatus, the communication I / F 217 includes an antenna. The bus 218 connects the units included in the information processing apparatus 200 to each other and transmits information therebetween. According to the present exemplary embodiment, the display unit 215 and the operation unit 216 are present in the information processing apparatus 200. Alternatively, at least one of the display unit 215 and the operation unit 216 may exist outside the information processing apparatus 200 as an independent apparatus.

[0038] [Operation Flow]

[0039] Figure 3A is a flowchart illustrating an operation example of the information processing apparatus 200. The processing illustrated in Figure 3A is implemented by loading a program stored in the ROM 212 into the RAM 213 by the CPU 211 included in the information processing apparatus 200 and executing the program. Figure 3A At least a part of the processing illustrated in Figure 6 may be implemented by one or more dedicated hardware different from the CPU 211. This also applies to the processing performed by the flowchart shown in Figure 3A . The processing illustrated in Figure 3A starts at the timing when an instruction for generating information for visibility determination is input to the information processing apparatus 200. However, Figure 3A the timing at which the processing illustrated in

[0040] In step S2010, the subject information acquisition unit 201 acquires subject information 101. In step S2020, the field of view information acquisition unit 202 acquires field of view information 102. In step S2030, the virtual viewpoint acquisition unit 203 acquires virtual viewpoint information 103. The order of the processes from step S2010 to S2030 is not limited to the above order, and at least a part of the processes can be executed in parallel. In step S2040, the visibility determination unit 204 performs visibility determination. Details of the visibility determination will be described with reference to Figure 4A and Figure 4B The details of the visibility determination will be described.

[0041] Figure 4A An example is shown in which two athletes 20 and 21 are present in the playing field 10 of the imaging area, and the ball 15 is located near the left hand of athlete 20. The positions and shapes of athletes 20 and 21 and the ball 15 are represented by the subject information 101. In Figure 4A and Figure 4B In the example shown, the field of view information 102 is specified to correspond to the field of view of the referee 22, who views athlete 20 from a position within the imaging area and is located outside the playing field 10. The quadrangular pyramid 30 indicates the position, direction, and angle (angle range) of the field of view of the viewpoint of the referee 22.

[0042] Figure 4B An example of the result of the visibility determination is shown. In Figure 4B , the visible area 51 with the additional flag "1" is the area that can be seen from the viewpoint corresponding to the field of view information 102. The out-of-range area 52 with the additional flag "-1" is the area that cannot be seen from the viewpoint corresponding to the field of view information 102. This is because this area is located at a position outside the field of view angle corresponding to the field of view information 102. The occluded area 50 with the additional flag "0" refers to the area that is located within the field of view angle corresponding to the field of view information 102 but cannot be seen because it is occluded by athlete 21. The area combining the above out-of-range area 52 and occluded area 50 is the invisible area that cannot be seen from the viewpoint corresponding to the field of view information 102. The occluded area 50 includes the area 60 on the area corresponding to athlete 20 that is occluded by athlete 21. The out-of-range area 52 includes the area 72 on the area corresponding to athlete 21 that is outside the field of view angle corresponding to the field of view information 102.

[0043] For example, visibility determination is performed as follows. In the visibility determination process, virtual light beams 40 are emitted in each viewing angle direction from the position of the viewpoint corresponding to the field of view information 102. When the bottom surface of the quadrangular pyramid 30 is regarded as an image of 1920 * 1080 pixels, the light beams 40 are emitted in the direction of each pixel. A flag "1" is attached to the surface of the subject that is first hit (intersected) by the light beam 40, and a flag "0" is attached to the surface of the subject that is hit by the light beam 40 the second time and later. A flag "-1" is attached to the surface of the subject that is outside the viewing angle and has never been hit by the light beam 40. When the near plane and the far plane of the field of view are specified by the field of view information 102, if the distance between the position of the field of view and the position where the light beam hits the subject is closer than the near plane or farther than the far plane, the flag "-1" is attached.

[0044] The method of visibility determination is not limited to the above processing. Visibility determination can be performed based on the field of view information 102 and the subject information 101. For example, the method of emitting the light beams 40 is not limited to the above processing, and the light beams 40 can be emitted at each predetermined viewing angle in the viewing angle. The flag can be attached to each grid for constructing the three-dimensional model of the subject, or can be attached to each pixel of the texture to be mapped to the three-dimensional model. Alternatively, the three-dimensional coordinate values (x, y, and z) of the intersection point of the light beam 40 and the subject surface and the flag value can be stored in association with each other. For example, the result of determining occlusion using a depth buffer can be used to determine the flag value. Specifically, the distance between the intersection point of the light beam emitted from the virtual viewpoint indicated by the virtual viewpoint information 103 and the subject surface and the position of the viewpoint indicated by the field of view information 102 is calculated. Then, this distance is compared with the depth value corresponding to the viewpoint indicated by the field of view information 102, so that occlusion can be determined.

[0045] In step S2050, the information generation unit 205 generates visibility information corresponding to the visibility determination result. As an example of the visibility information, the case of generating a flag image is described. By projecting the flags attached to each position on the subject to the virtual viewpoint indicated by the virtual viewpoint information 103, and by replacing the flag values of each position viewed from the virtual viewpoint with predetermined luminance values, a flag image is created. For example, the flag values "1", "0", and "-1" are replaced with the luminance values "0", "50", and "100", respectively. The luminance value corresponding to the flag value "-1" and the luminance value corresponding to the flag value "0" can be assigned the same value. In this case, the out-of-range area 52 and the occlusion area 50 cannot be distinguished in the flag image. In step S2060, the information generation unit 205 outputs the visibility information (flag image) to the display device 300.

[0046] Figure 3BThis is a flowchart illustrating an operation example of the display device 300. The processing illustrated in Figure 3B is implemented by loading a program stored in the ROM 212 into the RAM 213 by the CPU 211 in the display device 300 and executing the program. Figure 3B The processing illustrated in Figure 3B . Figure 3B At least a part of the processing illustrated in Figure 3B can be implemented by one or more dedicated hardware different from the CPU 211. The processing illustrated in Figure 3B starts at a timing when the information processing device 200 and the display device 300 are connected so as to be able to communicate with each other. However, Figure 3B The processing illustrated in Figure 3B . However, Figure 3B The timing at which the processing illustrated in Figure 3B starts is not limited to this timing. In the case where the virtual viewpoint image is a moving image, the processing in Figure 3B can be executed for each frame included in the moving image of the virtual viewpoint image. Figure 3B The processing in Figure 3B .

[0047] In step S3010, the display device 300 acquires the virtual viewpoint image 104. The virtual viewpoint image 104 acquired in step S3010 is an image corresponding to the virtual viewpoint used by the information processing device 200 to generate the visibility information. In step S3020, the display device 300 acquires the visibility information output from the information processing device 200. The order of the processing illustrated in steps S3010 and S3020 is not limited to the above order, and at least a part of the processing can be executed in parallel.

[0048] In step S3030, the display device 300 superimposes the flag image acquired from the information processing device 200 on the virtual viewpoint image 104. The display device 300 can acquire the field of view information 102 used by the information processing device 200 when generating the visibility information, and superimpose the CG representing the viewpoint corresponding to the field of view information 102 on the virtual viewpoint image 104.

[0049] Figure 5A An example of the superimposed image 95, that is, a new virtual viewpoint image generated by superimposing (superimposing and displaying) the flag image and the CG corresponding to the field of view information 102 on the virtual viewpoint image 104, is illustrated. The flag image is superimposed semi-transparently, and in the superimposed image 95, the occlusion area 80 corresponding to the flag "0" is covered with a predetermined semi-transparent color (for example, red). The camera model 90 in the superimposed image 95 represents the position of the viewpoint (the viewpoint of the referee 22) corresponding to the field of view information 102, and the boundary 85 represents the angular edge of the field of view corresponding to the field of view information 102.

[0050] The method of representing the visibility determination result is not limited to the above method. The contrast, the luminance value and color of the occlusion area in the virtual viewpoint image, and the rendering method of the texture can be changed, or the boundary of the occlusion area can be displayed. According to Figure 5AIn the example, by processing the part corresponding to the occlusion area 80 in the virtual viewpoint image, the occlusion area 80 and other areas can be distinguished. However, the method of displaying visibility information is not limited to this. For example, the visible area in the virtual viewpoint image (overlaying an image for covering the visible area) or the out-of-range area can be processed. Two or more of the visible area, the out-of-range area, and the occlusion area can be processed separately in different aspects so that these areas can be distinguished from each other. In order to distinguish and indicate the invisible area including the out-of-range area and the occlusion area, the entire invisible area can be uniformly processed (overlaying an image for covering the invisible area).

[0051] As Figure 5B shown, instead of the camera model 90, an arrow icon 89 indicating the position and direction of the viewpoint corresponding to the field-of-view information 102 can be displayed in the virtual viewpoint image. According to this exemplary embodiment, the image processing system 1 obtains a virtual viewpoint image 104 that does not display visibility information and displays the visibility information by overlaying it on the virtual viewpoint image 104. Alternatively, the image processing system 1 can directly generate a virtual viewpoint image that displays visibility information based on the image data obtained by the imaging device 100, the virtual viewpoint information 103, and the visibility information without obtaining the virtual viewpoint image 104 that does not display visibility information.

[0052] As described above, the viewer of the virtual viewpoint image can easily confirm the blind spot of the referee 22 by identifying the area that cannot be seen from the viewpoint of the referee 22 (i.e., the blind spot) and displaying the blind spot in a distinguishable manner on the virtual viewpoint image seen from an aerial view. For example, the viewer can easily identify that a part of the ball 15 and the athlete 20 is hidden in the blind spot of the referee 22 by viewing Figure 5B the overlaid image 95 shown.

[0053] According to the above example, based on the assumption that the field-of-view information 102 represents the field of view of the referee 22, the visibility from the viewpoint of the referee 22 is determined. However, the field-of-view information 102 is not limited to this. For example, the visibility of each area from the field of view of the athlete 21 can be distinguished by specifying the field-of-view information 102 to correspond to the field of view of the athlete 21. In addition, the field-of-view information 102 is specified so that the position where there is no athlete is set as the position of the viewpoint, so that it is possible to distinguish which area can be seen when there is an athlete at this position.

[0054] In addition, for example, the position of the viewpoint indicated by the field of view information 102 can be specified as the position of the ball 15, so that it is possible to easily distinguish which area can be seen from the position of the ball 15. In other words, from which area the ball 15 can be seen. If the position of the viewpoint indicated by the field of view information 102 is specified as, for example, the goal instead of the ball 15, it can be determined from which area the goal can be seen. The field of view information 102 may not include information indicating the line of sight direction and the field of view size. In other words, the field of view may include all directions viewed from the position of the viewpoint.

[0055] [Other examples of visibility information]

[0056] As described above with reference to Figure 4A 、 Figure 4B 、 Figure 5A and Figure 5B the examples are cases where the visibility information corresponding to the result of visibility determination output from the information processing device 200 is information that indicates the visible area and the invisible area in the imaging area in a distinguishable manner. Now, another example of visibility information will be described.

[0057] First, a case where the comparison result of the visibility of multiple viewpoints is used as the visibility information will be described. In this case, the information processing device 200 executes the Figure 6 shown processing flow, rather than the Figure 3A shown processing flow. The processing exemplified in Figure 6 starts at the same timing as the processing exemplified in Figure 3A .

[0058] In step S2010, the subject information acquisition unit 201 acquires the subject information 101. In step S2070, the field of view information acquisition unit 202 acquires a plurality of field of view information 102. At least one of the position, direction, and angle of the field of view of the plurality of field of view information 102 is different from each other. In this case, the positions of the viewpoints indicated by the plurality of field of view information 102 are different, and the following information is acquired: the field of view information 102 representing the field of view of the referee 22, and the field of view information 102 corresponding to four viewpoints obtained by moving 5 meters forward, backward, left, and right respectively from the position of the viewpoint of the referee 22. However, the number of the field of view information 102 to be acquired and the relationship between the respective field of view information 102 are not limited to the above number and relationship.

[0059] In step S2040, the visibility determination unit 204 performs visibility determination on each field of view information 102 and calculates the area of the occluded area of each field of view information 102. The area of the occluded area can be calculated by adding up the areas of the meshes with the flag "0" attached among the meshes constructing the subject.

[0060] In step S2050, the information generation unit 205 generates visibility information corresponding to the result of visibility determination. Specifically, the information generation unit 205 generates a flag image indicating an occlusion area corresponding to each of the plurality of field-of-view information 102. In addition, the information generation unit 205 generates information indicating the result of sorting (ranking) the field-of-view information 102 based on the area of the occlusion area calculated by the visibility determination unit 204 for each field-of-view information 102. Quick sort can be used as the sorting algorithm, but other algorithms such as bubble sort can also be used, not limited to quick sort. In step S2060, the information generation unit 205 outputs the flag image and the information indicating the sorting result as visibility information to the display device 300. The display device 300 uses the processing method referred to Figure 3B to generate and display a virtual viewpoint image with the visibility information superimposed.

[0061] Figure 7 An example of the superimposed image 96 as a virtual viewpoint image with the visibility information superimposed is illustrated. Icons 91 indicating four different viewpoints are displayed in front of, behind, to the left, and to the right of the camera model 90 indicating the viewpoint of the referee 22. On each icon 91, a numerical value indicating the area of the occlusion area viewed from each viewpoint is displayed. The icon 91 corresponding to the viewpoint with the smallest occlusion area is highlighted together with the letter "Best" ( Figure 7 the color of the arrow changes in the example). The method for emphasizing the icon 91 is not limited to this, and the rendering method of the icon 91 can be changed.

[0062] The viewpoint located behind the referee 22 is outside the range of the superimposed image 96. Therefore, the camera model corresponding to this viewpoint is not displayed, and only the arrow indicating the direction of the viewpoint and the area of the occlusion area are displayed. The occlusion area 80 indicates the occlusion area viewed from the viewpoint of the referee 22 as a reference. However, the occlusion areas corresponding to each of the plurality of viewpoints can be displayed in a distinguishable manner. The display of the occlusion area with the smallest area is not limited to the above method. This also applies to the display of the boundary 85.

[0063] As described above, the blind spots from a plurality of viewpoints near the referee 22 are identified, and the viewpoints with smaller blind spots are displayed, so that the viewer of the virtual viewpoint image can easily confirm the positions where the referee 22 should move to make the blind spots smaller.

[0064] Based on the above example, a display indicating a viewpoint with a small occlusion area was performed. However, not limited to the above example, for example, a display for indicating the size order of occlusion areas corresponding to respective viewpoints and a display for indicating a viewpoint with a larger occlusion area can be performed. A display for indicating a viewpoint with an occlusion area size smaller than a threshold value can be performed, or conversely, a display for indicating a viewpoint with an occlusion area of the threshold value or larger can be performed. Based on the above example, the comparison results of the sizes of occlusion areas corresponding to a plurality of viewpoints are displayed. However, not limited to the above example, the comparison results of the sizes of visible areas corresponding to a plurality of viewpoints can be displayed, or the comparison results of the sizes of invisible areas corresponding to a plurality of viewpoints can be displayed.

[0065] Based on the above example, the area is used as an index for indicating the size of the area. However, the volume can also be used as an index for indicating the size of the area. In this case, the light beam 40 is emitted from the viewpoint indicated by the field of view information 102 to each voxel in the three-dimensional space associated with the imaging area, and the flag "1" is attached to the voxel existing in the space until the light beam 40 first hits the subject. The three-dimensional position to which the flag "1" is attached in the above manner refers to the position that can be seen from the viewpoint indicated by the field of view information 102. The volume of the visible area can be calculated by counting the number of voxels to which the flag "1" is attached. The volumes of the occlusion area and the out-of-range area can also be calculated in a similar manner. In this case, the volume of the occlusion area corresponding to each viewpoint can be displayed as part of the visibility information in the superimposed image 96.

[0066] Based on the above example, a plurality of viewpoints to be compared are positioned at equal intervals. However, not limited to the above example, for example, a plurality of pieces of field of view information 102 corresponding to the viewpoints of a plurality of referees can be obtained, and these viewpoints can be compared with each other. Thus, it can be confirmed which referee's viewpoint includes the smallest blind spot. The plurality of viewpoints to be compared can be arbitrarily specified based on a user operation.

[0067] Next, a case where information indicating the visibility of a specific subject viewed from a specified viewpoint is used as the visibility information will be described. In this case, in Figure 3A the step S2040 shown, the visibility determination unit 204 determines whether a specific subject in the imaging area can be seen from the viewpoint indicated by the field of view information 102. For example, it is assumed that the ball 15 is selected as the subject for visibility determination. The visibility determination unit 204 calculates the ratio of the meshes to which the flag "1" indicating "visible" is attached and the meshes to which the flag "0" or "-1" indicating "invisible" is attached among the plurality of meshes included in the three-dimensional model of the ball 15.

[0068] In the case where the ratio of the signs "0" or "-1" exceeds a predetermined ratio, for example, 50%, the information generation unit 205 outputs, as visibility information, information indicating that the ball 15 cannot be seen from the referee 22 to the display device 300. The above-mentioned predetermined ratio can be set based on a user operation on the information processing device 200, or can be a ratio corresponding to a parameter obtained from outside the information processing device 200. Figure 8 An example of the imaging area and the display on the display device 300 is illustrated. In the case where most of the meshes or a predetermined ratio of the three-dimensional model of the construction ball 15 cannot be seen from the viewpoint of the referee 22 indicated by the field-of-view information 102, an icon 94 indicating the result of the visibility determination is displayed as visibility information on the display device 300. The icon 94 can be displayed superimposed on the virtual viewpoint image.

[0069] As described above, by specifying the blind spot from the viewpoint of the referee 22 and displaying an indication of whether the ball 15 is within the blind spot (whether the ball 15 is located in an invisible area), the viewer of the virtual viewpoint image can easily recognize that the ball 15 cannot be seen from the referee 22.

[0070] According to the above example, in the case where a predetermined ratio or more of the ball 15 cannot be seen from the viewpoint corresponding to the field-of-view information 102, the visibility information is displayed. On the contrary, in the case where a predetermined ratio or more of the ball 15 can be seen, the visibility information indicating this situation can be displayed. In other words, the visibility information is information indicating whether the specific subject ball 15 is located in the visible area. In addition, depending on whether the ball 15 is visible, different visibility information can be displayed. Also, depending on whether the ball 15 is located in the occlusion area or the out-of-range area, different visibility information can be displayed. According to the above example, the ball 15 is described as the subject for which the visibility determination is to be made. However, the subject can be a specific athlete and a specific position in the imaging area, and is not limited to the above example. The subject for which the visibility determination is to be made can be set based on a user operation received by the information processing device 200, and there can be multiple subjects as the targets for the visibility determination.

[0071] The information processing device 200 can determine the visibility of a specific subject viewed from a plurality of viewpoints corresponding to the plurality of field-of-view information 102. For example, the information processing device 200 can determine whether the ball 15 can be seen from the viewpoints of each of the plurality of referees, and display information indicating from which referee the ball 15 can be seen as visibility information.

[0072] As described above with reference to a plurality of examples, the image processing system 1 obtains subject information 101 indicating the position and shape of a subject in a captured area captured by a plurality of imaging devices 100 from different directions. The image processing system 1 distinguishes a visible area that can be seen from a specific position in the captured area and an occluded area that cannot be seen from the specific position because the subject is occluded, based on the subject information 101. The image processing system 1 generates a virtual viewpoint image that displays information based on the discrimination result. The virtual viewpoint image is based on the image data captured by the plurality of imaging devices 100 and virtual viewpoint information 103 indicating the position and direction of the virtual viewpoint. According to the above configuration, it is possible to easily distinguish the area that can be seen from a specific position and the area that cannot be seen from the specific position due to occlusion.

[0073] According to the above exemplary embodiment, the following situation has been mainly described, that is, based on the field-of-view information 102 and the subject information 101 regarding the viewpoint corresponding to the referee 22, visibility information indicating whether each area in the captured area can be seen from the referee 22 is generated. However, for example, the information processing device 200 can obtain the field-of-view information 102 corresponding to the left eye and the right eye of the referee 22 respectively, and generate visibility information indicating the area that can be seen only by the left eye, the area that can be seen only by the right eye, and the area that can be seen by both eyes in a distinguishable manner. The field-of-view information 102 can include not only information about the position and direction of the viewpoint and the size of the field of view, but also information about the visual acuity of a person (for example, the referee 22) or information about the environment of the captured area (for example, information about the weather, brightness, and the position of the light source). The information processing device 200 can use the above field-of-view information 102 to generate visibility information indicating the area that can be clearly seen from a specific viewpoint and the area that cannot be clearly seen from a specific viewpoint in a distinguishable manner. According to the above configuration, it is possible to more detailedly verify the visibility from a specific viewpoint in a certain scene.

[0074] According to the above exemplary embodiment, an example in which the information processing device 200 outputs the result of visibility determination from the viewpoint at a specific time has been described. However, the information processing device 200 can aggregate and output the results of visibility determination at a plurality of times when the viewpoint changes. For example, the area and volume of the invisible area viewed each time from the viewpoint of the referee 22 moving around can be plotted and output. Alternatively, the duration during which the area and volume of the invisible area exceed a threshold or the ball 15 enters the invisible area can be measured and output. The duration of the entire game can be used to score the actions of the referee 22. In addition, for example, information indicating the overlapping portion of the visible areas (the areas visible at any time) and the overlapping portion of the invisible areas (the areas invisible at some times) at a plurality of times can be output as visibility information.

[0075] According to the present disclosure, it is possible to easily distinguish between a region visible from a specific position and a region that cannot be seen from that position due to occlusion.

[0076] <Other embodiments>

[0077] Embodiments of the present disclosure can also be implemented by a computer of a system or apparatus that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a "non-transitory computer-readable storage medium") to perform the functions of one or more of the above embodiments, and / or includes one or more circuits (e.g., an application specific integrated circuit (ASIC)) for performing the functions of one or more of the above embodiments. Moreover, embodiments of the present disclosure can be implemented by a method in which a computer of a system or apparatus, such as by reading and executing computer-executable instructions from a storage medium, performs the functions of one or more of the above embodiments and / or controls one or more circuits to perform the functions of one or more of the above embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)), and may include a network of separate computers or separate processors to read and execute the computer-executable instructions. The computer-executable instructions may be supplied to the computer, for example, from a network or a storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read only memory (ROM), a memory of a distributed computing system, an optical disk (such as a compact disk (CD), a digital versatile disk (DVD), or a Blu-ray disk (BD)), a flash device, and a memory card, etc.

[0078] Other embodiments

[0079] Embodiments of the present invention can also be implemented by the following method, that is, by providing software (program) that performs the functions of the above embodiments to a system or apparatus through a network or various storage media, and the method in which a computer or a central processing unit (CPU), a microprocessing unit (MPU) of the system or apparatus reads and executes the program.

[0080] Although the present disclosure has been described with respect to exemplary embodiments, the scope of the claims should be given the broadest interpretation to cover all such variations as well as equivalent structures and functions.

Claims

1. An information processing apparatus, comprising: an acquisition unit configured to acquire object information representing positions and shapes of a plurality of objects in an imaging area, wherein the imaging area is imaged by a plurality of imaging devices from different directions; a determination unit configured to determine, based on the object information acquired by the acquisition unit, a visible area visible from a specific object included in the plurality of objects and an occlusion area that is not visible from the specific object due to occlusion by other objects; and a generation unit configured to generate a virtual viewpoint image that displays the visible area and the occlusion area based on the determination result of the determination unit, the virtual viewpoint image being generated based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating positions and directions of the virtual viewpoint.

2. The information processing apparatus according to claim 1, further comprising a field-of-view information acquisition unit configured to acquire field-of-view information related to a viewpoint of a specific object, the field-of-view information indicating a position of the specific object as the viewpoint position, a line-of-sight direction of the viewpoint position, and a field-of-view size centered on the line-of-sight direction, Among them, wherein the determination unit determines, in an area included in a range of the field of view specified by the field-of-view information acquired by the field-of-view information acquisition unit, a visible area where there is no object between the positions of the viewpoints, and determines an occlusion area where there is an object between the positions of the viewpoints.

3. The information processing apparatus according to claim 1, Among them, wherein the visible area is included within a predetermined angular range viewed from the specific object, and wherein the determination unit determines the visible area and an out-of-range area that is not included within the predetermined angular range viewed from the specific object.

4. The information processing apparatus according to claim 3, wherein, The virtual viewpoint image displays an invisible area including the occlusion area and the out-of-range area.

5. The information processing apparatus according to claim 4, wherein, The information based on the determination result displayed on the virtual viewpoint image includes an image covering the invisible area in the virtual viewpoint image.

6. The information processing apparatus according to claim 1, wherein, The information based on the determination result displayed on the virtual viewpoint image includes information indicating whether a specific object in the imaging area is located in the visible area.

7. The information processing apparatus according to claim 1, wherein, The information based on the determination result displayed on the virtual viewpoint image includes information indicating an area or volume of the occlusion area.

8. The information processing apparatus according to claim 1, wherein, The information based on the determination result displayed on the virtual viewpoint image includes information indicating a specific position where the area occluded and not visible is small among a plurality of specific positions in the imaging area.

9. The information processing apparatus according to claim 8, Among them, wherein the specific position corresponds to a designated object in the imaging area, and wherein the occlusion area refers to an area that is not visible from the specific position due to occlusion by other objects.

10. The information processing apparatus according to claim 9, Among them, wherein the designated object is a person, and wherein the specific position is a position of the head or eyes of the person.

11. The information processing apparatus according to claim 1, wherein, The object information represents three-dimensional positions and three-dimensional shapes of the objects.

12. An information processing method, comprising: Obtain subject information representing the respective positions and shapes of a plurality of subjects in a imaging region, wherein the imaging region is imaged from different directions by a plurality of imaging devices; Based on the subject information obtained by the obtaining, determine a visible region that can be seen from a specific subject included in the plurality of subjects and an occlusion region that cannot be seen from the specific subject due to being occluded by other subjects; and Generate a virtual viewpoint image that displays the visible region and the occlusion region based on the determination result, the virtual viewpoint image being generated based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating the position and direction of the virtual viewpoint.

13. The information processing method according to claim 12, wherein, The information based on the determination result displayed on the virtual viewpoint image includes at least any one of information indicating the visible region and information indicating the occlusion region.

14. A non-transitory computer-readable storage medium storing a program that causes a computer to execute an information processing method, the information processing method including: Obtain subject information representing the respective positions and shapes of a plurality of subjects in a imaging region, wherein the imaging region is imaged from different directions by a plurality of imaging devices; Based on the subject information obtained by the obtaining, determine a visible region that can be seen from a specific subject included in the plurality of subjects and an occlusion region that cannot be seen from the specific subject due to being occluded by other subjects; and Generate a virtual viewpoint image that displays the visible region and the occlusion region based on the determination result, the virtual viewpoint image being generated based on image data captured by the plurality of imaging devices and virtual viewpoint information indicating the position and direction of the virtual viewpoint.

15. A computer program product including a program that causes a computer to execute the steps of the information processing method according to claim 12.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method and storage medium

    CN109816766A

  • Display control device, display control method, and program

    US20160055676A1