Image Processing System

By segregating image processing tasks in the image processing system, with one group capturing foreground images and another with wider views capturing background images, the system addresses the issue of mottled backgrounds and processing load, resulting in high-quality virtual viewpoint images with improved user experience and efficiency.

JP7683091B2Active Publication Date: 2025-05-26CANON KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024096380
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-05-26
Estimated Expiration
2039-05-23

AI Technical Summary

Technical Problem

Existing image processing systems that generate virtual viewpoint images from multiple camera viewpoints often produce background images with mottled colors due to varying brightness and color across different camera positions, leading to an uncomfortable viewing experience. Additionally, the increased processing load on image processing devices when generating both foreground and background images is a concern.

Method used

The system separates the image processing tasks by using one group of cameras to capture foreground images and another group with wider imaging ranges to capture background images. This allows for the generation of high-quality virtual viewpoint images by minimizing the number of cameras needed for background capture and distributing the processing load.

Benefits of technology

This approach results in virtual viewpoint images with reduced mottling in the background, enhancing user comfort, and reduces the processing load on image processing devices, improving overall system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683091000001
    Figure 0007683091000001
  • Figure 0007683091000002
    Figure 0007683091000002
  • Figure 0007683091000003
    Figure 0007683091000003
Patent Text Reader

Abstract

To solve the problem occurring when generating a background image.SOLUTION: An image processing system generates a foreground image including a foreground object based on an image captured by an imaging apparatus included in a first imaging apparatus group. Based on an image captured by an imaging apparatus included in a second imaging apparatus group different from the first imaging apparatus group, the image processing system generates a background image not containing the foreground object. A virtual viewpoint image is generated based on the generated foreground image and the background image.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing system that generates a virtual viewpoint image according to a virtual viewpoint. [Background technology]

[0002] Recently, a technology that uses multiple cameras (imaging devices) installed at different positions to synchronously capture images from multiple viewpoints and generate a video (virtual viewpoint video) from an arbitrary virtual camera (virtual viewpoint) using the multiple viewpoint images obtained by the capture has been attracting attention (Patent Document 1). With such a technology, for example, it becomes possible to watch highlight scenes of soccer or basketball from various angles, and it becomes possible to give the user a high sense of realism compared to normal video content.

[0003] Patent Document 1 discloses a method of generating a foreground image including a moving object such as a person and a background image not including the object from images captured by each of a plurality of cameras, and generating a virtual viewpoint image from the generated foreground image and background image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2017-212593 A Summary of the Invention [Problem to be solved by the invention]

[0005] In order to generate a high-quality virtual viewpoint image, it is desirable to capture images from various directions using many cameras. On the other hand, when a background image is generated based on images captured from various directions by many cameras, an undesirable image may be generated. For example, when capturing images of a grass field in a stadium from different directions, the same place may appear bright or dark depending on the camera position, resulting in captured images with different brightness and color. When background regions are cut out from each of the captured images with different brightness and color and then synthesized, a background image with a mottled color as a whole is generated. As a result, the background color of the generated virtual viewpoint video may be mottled, which may cause the user to feel uncomfortable.

[0006] Furthermore, if each image processing device connected to each of the multiple cameras generates both a foreground image and a background image, the processing load of each image processing device increases.

[0007] An object of the present invention is to solve at least one of the above problems. [Means for solving the problem]

[0008] In order to solve the above problems, the present invention , No. A foreground image including a foreground object is generated based on an image captured by an imaging device included in one imaging device group. foreground Generation means and , No. Based on an image captured by an imaging device included in the imaging device group of 2, The above Generate a background image that does not contain any foreground objects background Generation means and ,before Foreground image and before Background image And based on , depending on the virtual viewpoint An image generating means for generating a virtual viewpoint image; The imaging device included in the first imaging device group and the imaging device included in the second imaging device group perform imaging in synchronization with each other, and the imaging range of the imaging device included in the second imaging device group is wider than the imaging range of the imaging device included in the first imaging device group. It is characterized by: Effect of the Invention

[0009] According to the present invention, at least one of the above-mentioned problems can be solved. [Brief description of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a multi-view camera arrangement. [Diagram 2] FIG. 13 is a diagram showing an example of a shooting range. [Diagram 3] FIG. 1 illustrates an example of a system configuration of an image processing system. [Figure 4] 1 is a functional block diagram showing a configuration of a camera adapter in a first embodiment. [Diagram 5] FIG. 2 is a functional block diagram showing the configuration of a background processing server and a foreground processing server. [Figure 6] 11 is a diagram showing the relationship between a background image transmitted from a camera adapter and processing in a background processing server. [Figure 7] FIG. 13 is a diagram illustrating an example of a system configuration of an image processing system according to a second embodiment. [Figure 8] FIG. 13 is a functional block diagram showing the configuration of a camera adapter in a third embodiment. [Figure 9] Background processing server hardware configuration diagram [Figure 10] Functional block diagram showing the configuration of the virtual viewpoint video generation server [Figure 11] 1 is a flowchart showing the process executed by the camera adapter. [Figure 12] A flowchart showing the process executed by the background processing server and the foreground processing server. [Figure 13] Flowchart showing the process executed by the virtual viewpoint video generation server DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the configurations described in the following embodiments are merely examples, and the present invention is not intended to be limited to the described embodiments.

[0012] (First embodiment) In this embodiment, an example of an image processing system will be described in which a plurality of cameras (imaging devices) are installed in a facility such as a stadium (a stadium) to capture images, and a virtual viewpoint image is generated using the multiple viewpoint images obtained by the capture. This image processing system is a system that generates a virtual viewpoint image that represents the appearance from a specified virtual viewpoint based on a plurality of images based on images captured by the plurality of imaging devices and a specified virtual viewpoint. The virtual viewpoint image in this embodiment is also called a free viewpoint video, but is not limited to an image corresponding to a viewpoint freely (arbitrarily) specified by the user, and for example, an image corresponding to a viewpoint selected by the user from a plurality of candidates is also included in the virtual viewpoint image. In addition, the specification of the virtual viewpoint is not limited to the case where it is performed by a user operation, and may be automatically performed based on the result of image analysis, etc. In addition, in this embodiment, the case where the virtual viewpoint image is a video will be mainly described, but the virtual viewpoint image may be a still image.

[0013] 1 is a diagram showing an example of a multi-view camera arrangement for shooting soccer in the image processing system of this embodiment. Although soccer is used as an example in this embodiment, soccer is just one example of a subject to be shot, and other sporting events, concerts, plays, etc. may also be used.

[0014] Reference numeral 100 denotes a stadium, and 22 cameras 111a to 111v are arranged to surround a field 101 where the sport takes place. These multiple cameras 111a to 111v are installed at different positions to surround such an imaging area, and capture images in a synchronized manner.

[0015] Of the 22 cameras, 111a to 111p are foreground cameras for capturing images of foreground objects such as players, referees, and balls, and are installed so that they can capture images of the field on which the objects are located from multiple directions. 111q to 111v are background cameras, and are installed so that they can capture images of background areas such as spectator stands and the field. Each of the background cameras 111q to 111v is set to a wide angle of view so that it can capture a wider range than each of the foreground cameras 111a to 111v. This makes it possible to capture the entire background area without missing anything even with a small number of cameras.

[0016] Here, a foreground image is an image in which a foreground object area (foreground region) is extracted from an image captured by a camera. A foreground object extracted as a foreground region refers to a dynamic object (moving body) that moves (its absolute position or shape can change) when images are captured from the same direction in chronological order. Foreground objects include people on the field, such as players and referees, and the ball.

[0017] A background image is an image of at least an area (background area) different from the foreground object. Specifically, a background image is an image in a state where the foreground object has been removed from the captured image. The background refers to an imaged object that remains stationary or nearly stationary when images are captured from the same direction in chronological order. Such imaged objects include the field 101 on which the game is played, spectator seats, goals used in ball games, and other structures. However, the background is at least an area different from the foreground object, and the imaged object may include other objects in addition to the foreground object and the background.

[0018] Reference numerals 111q to 111t denote cameras for photographing the spectator seats, and as an example, the photographing range of 111t is shown in FIG. 2(a) 201. Reference numeral 111u denotes a camera for photographing the field 101, and the photographing range is shown in FIG. 2(b) 202. Reference numeral 111v denotes a camera for photographing the sky.

[0019] In this way, the camera parameters of each camera for capturing an image of the background region are set so that the angle of view is wider than that of each camera for capturing an image of the foreground region, thereby widening the capturing range. This allows the entire background region to be captured with fewer cameras than the number of cameras for capturing an image of the foreground region. In addition, for different types of regions such as the spectator seats, the field, and the sky, the number of cameras responsible for capturing images of each region is minimized. This reduces the possibility of a background image with a mottled color overall being generated when a background image generated based on images captured by each camera is synthesized.

[0020] FIG. 3 is a diagram showing an example of a system configuration of an image processing system according to this embodiment.

[0021] Reference numerals 111a to 111v denote the above-mentioned cameras, which are connected to camera adapters 120a to 120v, respectively. In this embodiment, the frame rate of the cameras 111a to 111v is 60 fps (frames per second). That is, each camera captures 60 frames of images per second.

[0022] Camera adapters 120a to 120v are image processing devices that process the captured images captured by cameras 111a to 111v, respectively. Camera adapters 120a to 120p extract foreground images from the input captured images and transmit the foreground image data to foreground processing server 302 connected via network 301. Camera adapters 120q to 120v extract background images from the input captured images and transmit the background image data to background processing server 303 connected via network 301. In the example of FIG. 3, camera adapters 120a to 120p for foreground processing and camera adapters 120q to 120v for background processing are on the same network topology. For example, camera adapters 120a to 120d for foreground processing and camera adapter 120q for background processing are bus-connected to the same network cable.

[0023] Based on the multiple pieces of foreground image data received, the foreground processing server 302 performs generation processing of a foreground model, which is shape data representing the three-dimensional shape of the foreground object, and outputs it to the virtual viewpoint video generation server 304 together with foreground texture data for coloring the foreground model.

[0024] Based on the multiple background image data received, the background processing server 303 performs processing such as generating background texture data for coloring a background model representing the three-dimensional shape of the background, such as a stadium, and outputs the processing results to the virtual viewpoint video generation server 304.

[0025] The virtual viewpoint video generation server 304 generates and outputs a virtual viewpoint video based on data input from the foreground processing server 302 and the background processing server 303. A virtual viewpoint image is generated by mapping texture data to a foreground model and a background model and performing rendering according to a virtual viewpoint indicated by viewpoint information. Here, the viewpoint information used for generating the virtual viewpoint image is information indicating the position and orientation of the virtual viewpoint. Specifically, the viewpoint information is a parameter set including a parameter indicating the three-dimensional position of the virtual viewpoint and a parameter indicating the orientation of the virtual viewpoint in the pan, tilt, and roll directions. Note that the content of the viewpoint information is not limited to the above. For example, the parameter set as the viewpoint information may include a parameter indicating the size of the field of view (angle of view) of the virtual viewpoint. The viewpoint information may have multiple parameter sets. For example, the viewpoint information may have multiple parameter sets corresponding to multiple frames constituting a video of a virtual viewpoint image, and may be information indicating the position and orientation of the virtual viewpoint at each of multiple consecutive time points.

[0026] In the example of Fig. 3, the camera and the camera adapter are separated, but they may be one device included in the same housing. In addition, an example is shown in which the foreground processing server 302 and the background processing server 303 are different devices, but they may be configured as an integrated device. In addition, the foreground processing server 302, the background processing server 303, and the virtual viewpoint video generation server 304 may be configured as one device. In this way, the multiple functional units shown in Fig. 3 may be realized by one device, or one functional unit may be realized by multiple devices working together.

[0027] Next, the hardware configuration of the background processing server 303 will be described with reference to Fig. 9. Note that the hardware configurations of the camera adapters 120a to 120v, the foreground processing server 302, and the virtual viewpoint video generation server 304 are the same as the configuration of the background processing server 303.

[0028] The background processing server 303 includes a CPU 901 , a ROM 902 , a RAM 903 , an auxiliary storage device 904 , a display unit 905 , an operation unit 906 , a communication I / F 907 , and a bus 908 .

[0029] The CPU 901 realizes each function described below by controlling the background processing server 303 as a whole using computer programs and data stored in the ROM 902 and the RAM 903. The background processing server 303 may have one or more dedicated hardware components different from the CPU 901, and at least a part of the processing by the CPU 901 may be executed by the dedicated hardware components. Examples of the dedicated hardware components include an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), and a DSP (Digital Signal Processor). The ROM 902 stores programs that do not require modification. The RAM 903 temporarily stores programs and data supplied from the auxiliary storage device 904, and data supplied from the outside via the communication I / F 907. The auxiliary storage device 904 is composed of, for example, a hard disk drive, and stores various data such as image data and audio data.

[0030] Display unit 905 is composed of, for example, a liquid crystal display or LEDs, and displays a GUI (Graphical User Interface) and the like for the user to operate camera operation terminal 130. Operation unit 906 is composed of, for example, a keyboard, mouse, joystick, touch panel, and the like, and receives operations by the user to input various instructions to CPU 901. CPU 901 operates as a display control unit that controls display unit 905, and as an operation control unit that controls operation unit 906.

[0031] A communication I / F 907 is used for communication with external devices such as camera adapter 120 and virtual viewpoint video generation server 304. A bus 908 connects each unit of camera operation terminal 130 to transmit information.

[0032] Fig. 4 is a functional block diagram showing the configuration of camera adapters 120a to 120v, and Fig. 11 is a flowchart showing the flow of processing executed by camera adapters 120a to 120v. Note that the processing executed by camera adapter 120 is realized, for example, by CPU 901 of camera adapter 120 executing a program stored in ROM 902 or RAM 903.

[0033] FIG. 4(a) shows camera adaptors 120q to 120t for processing background images of the audience area, and FIG. 11(a) shows the flow of processing executed by the camera adaptors 120q to 120t.

[0034] 401 is a camera adapter for cutting out a background image of the seating area. The image data input unit 402 receives captured image data output from the camera (S1101). The seating area cutting unit 403 cuts out an image of the seating area from the input captured image (S1102). The seating area may be determined by determining a predetermined area as the seating area, or may be automatically determined from the characteristics of the image. In this embodiment, the seating area storage unit 404 stores coordinate data of the seating area, which is a predetermined area. The seating area can be cut out by comparing the stored coordinates with the coordinates in the captured image. The transmission unit 405 transmits the cut-out background image data of the seating area to the background processing server 303 (S1103). It is assumed that the transmission unit 405 transmits the background image data of the seating area at 30 fps.

[0035] FIG. 4(b) shows a camera adaptor 120u for processing an image of a field area, and FIG. 11(b) shows the flow of processing executed by the camera adaptor 120u.

[0036] In FIG. 4(b), 411 is a camera adapter for cutting out a background image of the field 101. The image data input unit 412 receives captured image data output from the camera (S1111). The field area cutting unit 413 cuts out an image of the field 101 area from the input captured image (S1112). The field area may be determined by determining a predetermined area as the field area, or may be automatically determined from the characteristics of the image. In this embodiment, the field area storage unit 414 stores coordinate data of the field area, which is a predetermined area, and the field area is cut out by comparing the stored coordinates with coordinates in the captured image.

[0037] The foreground / background separation unit 415 performs processing to separate the foreground and background of the field area image using multiple frames of field area images, for example, regarding moving objects as the foreground and stationary objects as the background (S1113). The background cutout unit 416 cuts out a background image based on the processing result by the foreground / background separation unit 415. The transmission unit 417 transmits the background image data of the cut-out field area to the background processing server 303 (S1114). It is assumed that the transmission unit 417 transmits the background image data of the field area at 1 fps.

[0038] FIG. 4(c) shows a camera adaptor 120v for processing images of the sky area, and FIG. 11(c) shows the flow of processing executed by the camera adaptor 120v.

[0039] 421 is a camera adapter for cutting out a background image of a sky area. The image data input unit 422 receives captured image data output from the camera (S1121). The sky area cutting unit 423 cuts out an image of a sky area from the input captured image (S1122). The sky area may be determined by determining a predetermined area as the sky area, or may be automatically determined from the characteristics of the image. In this embodiment, the sky area is cut out by storing coordinate data of a sky area, which is a predetermined area, in the sky area storage unit 424 and comparing the stored coordinates with coordinates in the captured image. The transmission unit 425 transmits background image data of the cut out sky area to the background processing server 303 (S1123). The transmission unit 425 transmits background image data of the sky area at a frame rate lower than 1 fps. For example, only one image data is transmitted per game.

[0040] FIG. 4(d) shows camera adapters 120a to 120p for processing foreground images, and FIG. 11(c) shows the flow of processing executed by camera adapters 120a to 120p.

[0041] 431 is a camera adapter for cutting out a foreground image. The image data input unit 432 receives captured image data output from the camera (S1131). The foreground / background separation unit 433 performs processing for separating the foreground and background of an image by using images of multiple frames, for example, regarding moving objects as the foreground and stationary objects as the background (S1132). However, this is not limited to this, and the foreground and background may be separated using a technique such as machine learning. The foreground cutout unit 434 cuts out a foreground image based on the processing result by the foreground / background separation unit 433. The transmission unit 434 transmits the cut out foreground image data to the foreground processing server 302 (S1133). It is assumed that the transmission unit 435 transmits background image data of the field area at 60 fps.

[0042] Here, the transmission frame rate of the background image is set lower than that of the foreground image. This is because foreground images, which include moving subjects such as athletes, change significantly over time, so it is necessary to update them at a high frame rate to prevent deterioration of image quality. On the other hand, background images do not change much over time, so even if they are updated at a low frame rate, the impact on image quality is small and the amount of transmitted data can be reduced.

[0043] Also, the transmission frame rate is changed depending on the type of background image because spectators in the seats may move, so it is desirable to update them relatively frequently to generate a highly realistic image, whereas fields and skies do not change much, so a lower update frequency is acceptable. In this way, by changing the frame rate depending on the type of background area, the background can be updated appropriately, and the sense of discomfort felt by the user viewing the virtual viewpoint image can be reduced.

[0044] Fig. 5(a) is a functional block diagram showing the configuration of the background processing server 303, and Fig. 12(a) is a flowchart showing the flow of processing executed by the background processing server 303. Note that the processing executed by the background processing server 303 is realized, for example, by the CPU 901 of the background processing server 303 executing a program stored in the ROM 902 or RAM 903.

[0045] Background image data receiving unit 501 receives background image data transmitted from camera adapters 120q to 120v (S1201). When each camera adapter transmits foreground image data or background image data, it transmits the data with a foreground / background identifier added to the header or the like of the data. Foreground processing server 302 and background processing server 303 can receive the necessary data as appropriate by referring to these identifiers.

[0046] The camera parameter storage unit 502 stores in advance parameters representing the positions, attitudes, angles of view, etc. of the cameras 111q to 111v. The camera parameters may be determined in advance, or may be acquired by detecting the state of the camera.

[0047] The background model data storage unit 503 stores in advance background model data that indicates the three-dimensional shapes of background objects such as spectator seats and a field.

[0048] Based on the received background image data, the background texture generation unit 504 uses the camera parameters stored in the camera parameter storage unit 503 to generate background texture data to be applied to the background model stored in the background model data storage unit 504 (S1202).

[0049] The background image data output unit 505 outputs the generated background texture data and the pre-stored background model data to the virtual viewpoint video generation server 304 (S1203).

[0050] 6 is a diagram showing the relationship between background image data transmitted from camera adaptors 120q-v and processing in background processing server 303. It shows how background texture data is generated in background processing server 303 based on background image data 601-606 captured by cameras 111q-111v and cut out by camera adaptors 120q-120v. Background processing server 303 may, for example, extract feature points of each input image, align the images based on the feature points, and stitch the images together to synthesize the background image data.

[0051] Fig. 5(b) is a functional block diagram showing the configuration of foreground processing server 302, and Fig. 12(b) is a flowchart showing the flow of processing executed by foreground processing server 302. Note that the processing executed by foreground processing server 302 is realized, for example, by CPU 901 of foreground processing server 302 executing a program stored in ROM 902 or RAM 903.

[0052] Foreground image data receiving section 511 receives foreground image data transmitted from camera adapters 120a to 120p (S1211). The foreground image data received here includes a foreground mask image and a foreground texture image.

[0053] The camera parameter storage unit 512 stores in advance parameters representing the positions, attitudes, angles of view, etc. of the cameras 111a to 111p. The camera parameters may be determined in advance, or may be acquired by detecting the state of the camera.

[0054] Foreground model data generating unit 513 generates foreground model data indicating the three-dimensional shape of the foreground object, based on the plurality of received foreground mask images and the camera parameters stored in camera parameter storage unit 512 (S1212).

[0055] Foreground image data output unit 514 outputs the generated foreground model data and the plurality of foreground texture data received from camera adapters 120a to 120p to virtual viewpoint video generation server 304 (S1213).

[0056] Fig. 10 is a functional block diagram showing the configuration of the virtual viewpoint video generation server 304, and Fig. 13 is a flowchart showing the flow of processing executed by the virtual viewpoint video generation server 304. Note that the processing executed by the virtual viewpoint video generation server 304 is realized, for example, by a CPU 901 of the virtual viewpoint video generation server 304 executing a program stored in a ROM 902 or a RAM 903.

[0057] Foreground image data receiving unit 1001 receives foreground image data transmitted from foreground processing server 302 (S1301). Background image data receiving unit 1002 receives background image data transmitted from background processing server 303 (S1302). Virtual viewpoint setting unit 1003 sets viewpoint information indicating a virtual viewpoint based on a user operation or the like (S1303).

[0058] The virtual viewpoint image generating unit 1004 generates a virtual viewpoint image based on the received foreground image data, background image data, and the set virtual viewpoint (S1304). As described above, the virtual viewpoint image is generated by mapping texture data to the foreground model data and background model data and performing rendering according to the virtual viewpoint indicated by the viewpoint information. The virtual viewpoint image generating unit 1004 generates and updates the virtual viewpoint image at 60 fps. As described above, the transmission frame rate of the background image is lower than that of the foreground image, so the update frequency is also lower. For example, when generating 60 frames of virtual viewpoint images per second, 60 frames of images are used for the foreground image, whereas one frame of images is used for the background image of the field area. In other words, the background image of the field area for that one frame is also used as the background image of the field area for the remaining 59 frames.

[0059] As described above, in this embodiment, some of the cameras are used as cameras dedicated to capturing images of the background, a background image is generated based on the captured images, and a foreground image is generated based on images captured by the remaining cameras. Then, a virtual viewpoint image corresponding to the designated viewpoint information is generated based on the generated foreground image and background image. At that time, the camera for capturing the background area is set to a wide angle of view so that it can capture a wider range than the camera for capturing the foreground area. Accordingly, the number of cameras for capturing the background area is made smaller than the number of cameras for capturing the foreground area, and the number of cameras capturing each area, such as the spectator seats, field, and sky, is minimized. This makes it possible to generate a background image that is not mottled even when the background images generated based on the images captured by each camera are combined, and the virtual viewpoint image finally generated can be an image that is less strange to the user. In addition, each camera adapter does not need to generate both the foreground image and the background image, and only needs to perform the generation process of one of them, so that the processing load of each camera adapter can be distributed and reduced.

[0060] Second embodiment In the first embodiment, the case where the camera adapters 120a to 120p for foreground processing and the camera adapters 120q to 120v for background processing are on the same network topology has been described. On the other hand, it is also possible to place the camera adapters 120a to 120p and 120q to 120v on different network topologies.

[0061] FIG. 7 is a diagram showing an example of a system configuration in which a camera adapter for foreground processing and a camera adapter for background processing are arranged on different network topologies.

[0062] Compared with the configuration diagram of FIG. 3, the cameras 111a to 111v and the camera adapters 120a to 120v are equivalent processing modules, so they are given the same reference numerals and their description will be omitted.

[0063] Foreground processing camera adapters 120a-120p are connected to network 701 and transmit foreground image data to foreground processing server 302 connected via network 701. Background processing camera adapters 120q-120v are connected to network 702 and transmit background image data to background processing server 303 connected via network 702. This makes it possible to adopt the most suitable network protocol for each of the foreground processing camera adapters and background processing camera adapters. For example, when foreground image data is transmitted via UDP / IP and background image data is transmitted via TCP / IP, it is possible to transmit UDP / IP without being affected by TCP / IP. This makes it possible to generate a background image with reduced transmission loss of background image data.

[0064] Also, for example, the foreground processing camera adaptors 120a to 120p may be daisy-chain connected, and the background processing camera adaptors 120q to 120v may be star connected.

[0065] (Third embodiment) In the first embodiment, an example was described in which the transmission and update frame rates of background images are made different depending on the characteristics of background areas such as the seating area and the field. Here, the image quality required may differ depending on the destination of the virtual viewpoint video and the viewing form, and it may be necessary to change the update frame rate and resolution of each background image accordingly. On the other hand, if the update frame rate of each background image can be changed unlimitedly, it may lead to tight network bandwidth and increased processing load on each server. For example, as described in the second embodiment, when the camera adapters 120q to 120v are connected to the same network, if the transmission rate is uniformly set to 30 fps, it is expected that the network bandwidth will be tight and the transmission will not be completed within the time. Therefore, in this embodiment, an example will be described in which the transmission and update frame rates of the background image data of each background area are adjusted while taking into consideration the overall balance. Here, an example will be described in which the update rate is prioritized over image quality for the seating area because movement is important, and the image quality is prioritized for the field area image because the resolution of the goal line and the grass grain is important.

[0066] Fig. 8 is a functional block diagram showing the configuration of camera adapter 120. Fig. 8(a) shows camera adapter 801 for processing images of the audience area, and Fig. 8(b) shows camera adapter 811 for processing images of the field area.

[0067] Camera adapter 801 has a configuration in which an update rate setting unit 802 and a data amount changing unit 803 are added to camera adapter 401 described in Fig. 4. The same processing blocks as those in camera adapter 401 are given the same reference numerals and descriptions thereof will be omitted.

[0068] The update rate setting unit 802 sets the update frame rate of the background image of the audience area. The transmission frame rate in the transmission unit 405 is also changed according to the update frame rate set here. The update frame rate may be set by a user operation or dynamically set according to the status of the communication line.

[0069] The data amount change unit 803 performs a process of changing the data amount of the image cut out by the audience area cutout unit 403 as necessary based on the update frame rate set by the update rate setting unit 802. For example, when the update frame rate is increased to 60 fps, a process of reducing the amount of data to be transmitted is performed. The data amount reduction process may be a compression process such as JPEG, or a process of reducing the resolution or gradation. The compression rate, etc. may be determined by the data amount change unit 802 calculating a compression rate, etc. that satisfies the update rate set by the update rate setting unit 802, or may be directly instructed by the update rate setting unit 802. Alternatively, the bit depth of the image may be limited. For example, the amount of data may be reduced by quantizing 12-bit image data to 8 bits. Furthermore, these methods may be used in combination.

[0070] Camera adapter 811 has a configuration in which an image quality parameter setting unit 811 and an update rate determination unit 812 are added to camera adapter 411 described in Fig. 4. The same processing blocks as those in 411 are given the same reference numerals and descriptions thereof will be omitted.

[0071] Image quality parameter setting section 812 sets image quality parameters for the image in the field region. For example, it may be possible to select from three levels of image quality: high image quality, standard image quality, and low image quality, and to allow the user to set a desired image quality from among these levels based on the user's operation.

[0072] The update rate determination unit 813 determines an update frame rate for the image cut out by the background cut out unit 416, based on the image quality parameters set by the image quality parameter setting unit 812. The transmission frame rate in the transmission unit 417 is also changed according to the update frame rate determined here. For example, when high image quality is set as the image quality parameter, the transmission / update frame rate is determined to be 1 fps in order to transmit uncompressed background image data without delay. The update frame rate may be determined by the update rate determination unit 813 calculating it based on the image quality set by the image quality parameter setting unit 812, or may be directly instructed by the image quality parameter setting unit 812.

[0073] In this way, according to this embodiment, it is possible to adaptively adjust the transmission / update frame rate of the background image data for each background area according to the required image quality.

[0074] (Other embodiments) In the above description, an example was described in which a foreground image is generated from images captured by 16 cameras 111a to 111p, and a background image is generated from images captured by six cameras 111q to 111v, but the number of cameras is not limited to these. By increasing the number of cameras for capturing foreground images and capturing images of the capture area from more directions, more accurate three-dimensional shape data of the foreground object is generated, and the resulting virtual viewpoint image also has high image quality. On the other hand, the more cameras for capturing background images are increased, the more likely the brightness and color of the background image obtained by synthesis becomes mottled. Therefore, it is preferable to use the minimum number of cameras for capturing background area images that can maintain the required image quality.

[0075] In addition to simply reducing the number, it is also advisable to provide dedicated cameras for capturing images of different background areas such as the spectator stands, the field, and the sky, as described in the above embodiment. In this way, even if the images captured by each camera have different brightness or color, the overall background image obtained by combining the images will be an image that is less strange to the user. In other words, it is advisable to capture images of multiple different types of background areas using the minimum number of cameras that can maintain the required image quality.

[0076] In addition, in the above explanation, an example was described in which the camera parameters are set so that the angle of view of the camera for capturing images of the background area is wider than that of the camera for capturing images of the foreground area, but a dedicated wide-angle camera may also be used for capturing images of the background area.

[0077] In the above description, the seating, the field, and the sky are captured as background regions, but other background regions may be added, and some regions may not be captured. For example, as for the sky, if it is sufficient to know whether the event is held during the day or night, and if image data prepared in advance can be used instead, actual capture of the sky may not be performed.

[0078] In the above description, an example of configuring a system by providing a camera dedicated to capturing a foreground image and a camera dedicated to capturing a background image has been described. However, some of the cameras may be used to generate both a foreground image and a background image. For example, not only a background image but also a foreground image may be generated from an image captured by the camera 111u for capturing an image of a field area. Also, without providing a camera dedicated to the background image, a foreground image and a background image may be generated only for the imaging devices of a small number of cameras among a plurality of cameras. In this way, by reducing the number of background images to be less than the number of foreground images, it is possible to suppress the generation of a mottled background image caused by the composition of background images. Also, since both a foreground image and a background image are not generated for all captured images, the processing load on the camera adapter can be distributed and reduced.

[0079] As described above, in the above embodiment, a second imaging device group is installed for capturing images of a background area in addition to a first imaging device group for capturing images of foreground objects such as athletes and referees from multiple directions in a stadium, which is an imaging target area. The image processing system in this embodiment generates a foreground image including a foreground object based on an image captured by an imaging device included in the first imaging device group. On the other hand, the imaging range of each imaging device included in the second imaging device group is wider than that of each imaging device included in the first imaging device group, and generates a background image not including a foreground object based on an image captured by each imaging device included in the second imaging device group. Then, a virtual viewpoint image according to a virtual viewpoint is generated based on the generated foreground image and background image.

[0080] The present invention can also be realized by supplying a program for implementing one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in a computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC, etc.) that implements one or more of the functions. The program may also be provided by recording it on a computer-readable storage medium. [Explanation of symbols]

[0081] 111 Camera 120 Camera adapter 302 Foreground Processing Server 303 Background Processing Server 304 Virtual Viewpoint Video Generation Server 401 Camera adapter for cutting out background images of the seating area 411 Camera adapter for cutting out background images of field areas 421 Camera adapter for extracting background images from sky regions

Claims

1. a foreground generating means for generating a foreground image including a foreground object based on an image captured by an imaging device included in the first imaging device group; a background generating means for generating a background image not including the foreground object based on an image captured by an imaging device included in a second imaging device group; an image generating means for generating a virtual viewpoint image according to a virtual viewpoint based on the foreground image and the background image; having an imaging device included in the first imaging device group and an imaging device included in the second imaging device group perform imaging in synchronization with each other; An image processing system, wherein an imaging range of an imaging device included in the second imaging device group is wider than an imaging range of an imaging device included in the first imaging device group.

2. The method further comprises: a data generating means for generating shape data indicating a three-dimensional shape of the foreground object based on the foreground image; 2. The image processing system according to claim 1, wherein said image generating means generates said virtual viewpoint image based on said shape data and said background image.

3. 3. The image processing system according to claim 2, wherein said data generating means receives the foreground image and does not receive the background image.

4. 4. The image processing system according to claim 1, wherein the background generating means transmits the background image at a predetermined frame rate.

5. the foreground generating means transmits a foreground image at a first frame rate; 5. The image processing system according to claim 1, wherein the background generating means transmits the background image at a second frame rate lower than the first frame rate.

6. the second imaging device group includes an imaging device that images a first area and an imaging device that images a second area, The image processing system according to any one of claims 1 to 5, characterized in that the background generation means includes a first background generation means that generates a background image corresponding to the first area based on an image captured by an imaging device that captures the first area, and a second background generation means that generates a background image corresponding to the second area based on an image captured by an imaging device that captures the second area.

7. 7. The image processing system according to claim 6, wherein the first background generating means transmits background images at a frame rate lower than that of the second background generating means.

8. the first region corresponds to a field in which the foreground object exists; 8. The image processing system according to claim 7, wherein the second area corresponds to an audience seat.

9. The first region is a region corresponding to the sky, 9. The image processing system according to claim 8, wherein the second area corresponds to an audience seat.

10. The image processing system according to claim 1 , wherein the background image corresponds to an audience seat.

11. 11. The image processing system according to claim 1, wherein the foreground object includes at least one of a player, a ball, and a referee.

Citation Information

Patent Citations

  • Information processing device, image processing system, information processing method, and program

    JP2017212593A

  • Imaging system, image processing device, image processing method, and program

    JP2018056971A

  • Image processing system, image processing method and program

    JP2019003325A

  • Image processing apparatus and method for controlling the same, and program, and image processing system

    JP2019050451A

  • Multiple camera controlling and image storing apparatus for synchronized multiple image acquisition and method thereof

    US20100157020A1