Image processing device, image processing method and program

The image processing apparatus addresses the challenge of continuous object observation by switching between images captured by multiple cameras with different positions, ensuring uninterrupted viewing even when the object is obscured.

JP2025084941APending Publication Date: 2025-06-03FUJIFILM CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025033133
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-27
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing image processing systems struggle to continuously provide an image of an object within an imaging region to a viewer, especially when the object is partially or completely obscured by other objects.

Method used

An image processing apparatus that uses a processor to detect object images from multiple images captured by cameras with different positions, and switches between outputting different images based on the presence or absence of the object in the current image, ensuring continuous observation of the object.

Benefits of technology

The system ensures continuous and uninterrupted observation of the object by seamlessly switching between images, even when the object is partially or completely obscured, thereby maintaining viewer engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084941000001_ABST
    Figure 2025084941000001_ABST
Patent Text Reader

Abstract

To continuously provide an image in which an object in an imaging region can be observed for a viewer who views an image obtained by imaging the imaging region.SOLUTION: An image processing device performs detection processing to detect an object image showing an object among a plurality of images obtained by imaging an imaging region by a plurality of cameras at different positions, outputs a first image among the plurality of images, and also outputs a second image in which the object image is detected through the detection processing among the plurality of images on transition from a detection state in which the object image is detected from the first image through the detection processing to a non-detection state in which the object image is not detected from the first image through the detection processing.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] Japanese Unexamined Patent Application Publication No. 2019-114147 discloses an information processing apparatus that determines the position of a viewpoint related to a virtual viewpoint image generated using a plurality of images captured by a plurality of imaging devices. The information processing apparatus described in Japanese Unexamined Patent Application Publication No. 2019-114147 includes a first acquisition unit that acquires position information indicating a position within a predetermined range from an imaging target of a plurality of imaging devices, and based on the position information acquired by the first acquisition unit, determines the position of a viewpoint related to a virtual viewpoint image for imaging the imaging target with a viewpoint at a position different from the position indicated by the position information acquired by the first acquisition unit, and is characterized by having a determination unit.

[0003] Japanese Unexamined Patent Application Publication No. 2019-118136 discloses an information processing apparatus characterized by having a storage unit that stores a plurality of captured video data, and an analysis unit that detects a blind spot from the plurality of captured video data stored in the storage unit, generates an instruction signal to prevent the blind spot, and outputs the instruction signal to a camera that generates the captured video data.

Summary of the Invention

[0004] One embodiment of the technology according to the present disclosure provides an image processing apparatus, an image processing method, and a program capable of continuously providing an image in which an object within an imaging region can be observed to a viewer of an image obtained by imaging the imaging region.

Means for Solving the Problems

[0005] A first aspect of the technology according to the present disclosure includes a processor and a memory built in or connected to the processor. The processor performs a detection process of detecting an object image indicating an object from a plurality of images obtained by imaging an imaging region by a plurality of cameras having different positions, outputs a first image among the plurality of images, and when transitioning from a detection state in which an object image is detected from the first image by the detection process to a non-detection state in which an object image is not detected from the first image by the detection process, the processor outputs a second image among the plurality of images in which an object image has been detected by the detection process. This is an image processing apparatus.

[0006] A second aspect of the technology according to the present disclosure is the image processing apparatus according to the first aspect, in which at least one of the first image and the second image is a virtual viewpoint image.

[0007] A third aspect of the technology according to the present disclosure is the image processing apparatus according to the first aspect or the second aspect, in which when the processor transitions from a detection state to a non-detection state in a situation where the first image is being output, the processor switches from outputting the first image to outputting the second image.

[0008] A fourth aspect of the technology according to the present disclosure is the image processing apparatus according to any one of the first aspect to the third aspect, in which the image is a plurality of frame images composed of a plurality of frames.

[0009] A fifth aspect of the technology according to the present disclosure is the image processing apparatus according to the fourth aspect, in which the plurality of frame images are moving images.

[0010] A sixth aspect of the technology according to the present disclosure is the image processing apparatus according to the fourth aspect, in which the plurality of frame images are consecutive shooting images.

[0011] A seventh aspect of the technology according to the present disclosure is the image processing apparatus according to any one of the fourth aspect to the sixth aspect, in which the processor outputs a plurality of frame images as the second image, and starts outputting the plurality of frame images as the second image from a timing earlier than the timing when the non-detection state is reached.

[0012] The eighth aspect according to the technology of the present disclosure is an image processing apparatus according to any one of the fourth to seventh aspects, in which a processor outputs a plurality of frame images as a second image, and ends the output of the plurality of frame images as the second image at a timing later than the timing when the non-detection state is reached.

[0013] The ninth aspect according to the technology of the present disclosure is that when a plurality of images include a third image in which an object image is detected by a detection process, and a plurality of frame images as a second image include a detection frame in which an object image is detected by the detection process and a non-detection frame in which an object image is not detected by the detection process, a processor determines a distance between a position of a second-image camera used for imaging to obtain a second image among a plurality of cameras and a position of a third-image camera used for imaging to obtain a third image among the plurality of cameras, and selectively outputs the non-detection frame and the third image according to the time of the non-detection state. The image processing apparatus is according to any one of the fourth to eighth aspects.

[0014] The tenth aspect according to the technology of the present disclosure is an image processing apparatus according to the ninth aspect, in which when a processor satisfies a non-detection frame output condition that the distance exceeds a threshold value and the time of the non-detection state is less than a predetermined time, the processor outputs a non-detection frame, and when the non-detection frame output condition is not satisfied, the processor outputs a third image instead of the non-detection frame.

[0015] The eleventh aspect according to the technology of the present disclosure is an image processing apparatus according to any one of the first to tenth aspects, in which a processor resumes the output of the first image on the condition that the processor returns from the non-detection state to the detection state.

[0016] The twelfth aspect of the technology according to the present disclosure is an image processing apparatus according to any one of the first to eleventh aspects, in which a plurality of cameras includes at least one virtual camera and at least one physical camera, and a plurality of images includes a virtual viewpoint image obtained by imaging an imaging region with the virtual camera and an imaging image obtained by imaging the imaging region with the physical camera.

[0017] The thirteenth aspect of the technology according to the present disclosure is an image processing apparatus according to any one of the first to twelfth aspects, in which, during a period when a processor switches from output of a first image to output of a second image, a plurality of virtual viewpoint images obtained by imaging with a plurality of virtual cameras that continuously connect from the position, orientation, and field of view angle of the camera used for imaging to obtain the first image to the position, orientation, and field of view angle of the camera used for imaging to obtain the second image are output.

[0018] The fourteenth aspect of the technology according to the present disclosure is an image processing apparatus according to any one of the first to thirteenth aspects, in which the object is a person.

[0019] The fifteenth aspect of the technology according to the present disclosure is an image processing apparatus according to the fourteenth aspect, in which a processor detects an object image by detecting a face image showing a person's face.

[0020] The sixteenth aspect of the technology according to the present disclosure is an image processing apparatus according to any one of the first to fifteenth aspects, in which, among a plurality of images, an image in which at least one of the position and size of the object image in the image satisfies a predetermined condition and the object image is detected by a detection process is output as a second image.

[0021] The seventeenth aspect of the technology according to the present disclosure is an image processing apparatus according to any one of the first to sixteenth aspects, in which the second image is an overhead image showing an overhead view of the imaging region.

[0022] The 18th aspect according to the technology of the present disclosure is an image processing apparatus according to any one of the 1st to 17th aspects, wherein the first image is an image for television broadcast.

[0023] The 19th aspect according to the technology of the present disclosure is an image processing apparatus according to any one of the 1st to 18th aspects, wherein the first image is an image obtained by being captured by a camera installed at an observation position for observing an imaging region or in the vicinity of the observation position among a plurality of cameras.

[0024] The 20th aspect according to the technology of the present disclosure is an image processing method including performing a detection process of detecting an object image indicating an object from a plurality of images obtained by imaging an imaging region by a plurality of cameras having different positions, outputting a first image among the plurality of images, and when transitioning from a detection state in which an object image is detected from the first image by the detection process to a non-detection state in which an object image is not detected from the first image by the detection process, outputting a second image among the plurality of images in which an object image is detected by the detection process.

[0025] The 21st aspect according to the technology of the present disclosure is a program for causing a computer to execute a process including performing a detection process of detecting an object image indicating an object from a plurality of images obtained by imaging an imaging region by a plurality of cameras having different positions, outputting a first image among the plurality of images, and when transitioning from a detection state in which an object image is detected from the first image by the detection process to a non-detection state in which an object image is not detected from the first image by the detection process, outputting a second image among the plurality of images in which an object image is detected by the detection process.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14A

Figure 14B

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29A

Figure 29B

Figure 29C

Figure 30

Figure 31

Embodiments for Carrying Out the Invention

[0027] An example of an embodiment related to an image processing apparatus, an image processing method, and a program of the technology of the present disclosure will be described with reference to the accompanying drawings.

[0028] First, the terms used in the following description will be described.

[0029] CPU refers to the abbreviation of "Central Processing Unit". RAM refers to the abbreviation of "Random Access Memory". SSD refers to the abbreviation of "Solid State Drive". HDD refers to the abbreviation of "Hard Disk Drive". EEPROM refers to the abbreviation of "Electrically Erasable and Programmable Read Only Memory". I / F refers to the abbreviation of "Interface". IC refers to the abbreviation of "Integrated Circuit". ASIC refers to the abbreviation of "Application Specific Integrated Circuit". PLD refers to the abbreviation of "Programmable Logic Device". FPGA refers to the abbreviation of "Field-Programmable Gate Array". SoC refers to the abbreviation of "System-on-a-chip". CMOS refers to the abbreviation of "Complementary Metal Oxide Semiconductor". CCD refers to the abbreviation of "Charge Coupled Device". EL refers to the abbreviation of "Electro-Luminescence". GPU refers to the abbreviation of "Graphics Processing Unit". WAN refers to the abbreviation of "Wide Area Network". LAN refers to the abbreviation of "Local Area Network". 3D refers to the abbreviation of "3 Dimensions". USB refers to the abbreviation of "Universal Serial Bus". 5G refers to the abbreviation of "5th Generation". LTE refers to the abbreviation of "Long Term Evolution". WiFi refers to the abbreviation of "Wireless Fidelity". RTC refers to the abbreviation of "Real Time Clock". SNTP refers to the abbreviation of "Simple Network Time Protocol". NTP refers to the abbreviation of "Network Time Protocol". GPS refers to the abbreviation of "Global Positioning System".Exif refers to the abbreviation of "Exchangeable image file format for digital still cameras". Fps refers to the abbreviation of "frame per second". GNSS refers to the abbreviation of "Global Navigation Satellite System". In the following, for the convenience of explanation, as an example of the "processor" according to the technology of the present disclosure, a CPU is exemplified, but the "processor" according to the technology of the present disclosure may be a combination of a plurality of processing devices such as a CPU and a GPU. When a combination of a CPU and a GPU is applied as an example of the "processor" according to the technology of the present disclosure, the GPU operates under the control of the CPU and is responsible for executing image processing.

[0030] In the following description, "match" refers to a match including an error generally tolerated in the technical field to which the technology of the present disclosure belongs (a match including an error to the extent that it does not go against the spirit of the technology of the present disclosure) in addition to a complete match. Also, in the following description, "the same imaging time" refers to the same imaging time including an error generally tolerated in the technical field to which the technology of the present disclosure belongs (a match including an error to the extent that it does not go against the spirit of the technology of the present disclosure) in addition to a completely identical imaging time.

[0031] [First Embodiment] As shown in FIG. 1 as an example, the image processing system 10 includes an image processing device 12, a user device 14, and a plurality of physical cameras 16. The user device 14 is used by the user 18.

[0032] In the first embodiment, a smartphone is applied as an example of the user device 14. However, the smartphone is merely an example, and for example, a personal computer may be used, or a portable multifunctional device such as a tablet terminal or a head-mounted display may be used. Also, in the first embodiment, a server is applied as an example of the image processing device 12. The number of servers may be one or a plurality. The server is merely an example, and for example, at least one personal computer may be used, or a combination of at least one server and at least one personal computer may be used. Thus, the image processing device 12 may be at least one device capable of executing image processing.

[0033] The network 20 is configured to include, for example, a WAN and / or a LAN. In the example shown in FIG. 1, although not shown, the network 20 includes, for example, a base station. The base station is not limited to one location and may be plural. Further, the communication standards used in the base station include wireless communication standards such as the 5G standard, the LTE standard, the WiFi (802.11) standard, and the Bluetooth (registered trademark) standard. The network 20 establishes communication between the image processing device 12 and the user device 14 and transmits and receives various types of information between the image processing device 12 and the user device 14. The image processing device 12 receives a request from the user device 14 via the network 20 and provides a service corresponding to the request to the user device 14 that is the request source via the network 20.

[0034] Note that in the first embodiment, a wireless communication method is applied as an example of the communication method between the user device 14 and the network 20 and the communication method between the image processing device 12 and the network 20, but this is merely an example, and a wired communication method may be used.

[0035] The physical camera 16 actually exists as an object and is an imaging device that can be visually recognized. The physical camera 16 is an imaging device having a CMOS image sensor and is equipped with an optical zoom function and / or a digital zoom function. Note that, instead of the CMOS image sensor, other types of image sensors such as a CCD image sensor may be applied. Also, in the present first embodiment, a plurality of physical cameras 16 are equipped with a zoom function, but this is merely an example, and a zoom function may be equipped on a part of the plurality of physical cameras 16, or a plurality of physical cameras 16 may not be equipped with a zoom function.

[0036] The plurality of physical cameras 16 are installed inside the soccer stadium 22. The imaging positions (hereinafter, also simply referred to as "positions") of the plurality of physical cameras 16 are different from each other, and the imaging directions (hereinafter, also simply referred to as "directions") of each physical camera 16 can be changed. In the example shown in FIG. 1, each of the plurality of physical cameras 16 is arranged so as to surround the soccer field 24, and images an area including the soccer field 24 as an imaging area. Imaging by the physical camera 16 refers to, for example, imaging at an angle of view including the imaging area. Here, the concept of "imaging area" includes not only the concept of an area indicating the entire inside of the soccer stadium 22 but also the concept of an area indicating a part inside the soccer stadium 22. The imaging area is changed according to the imaging position, imaging direction, and angle of view.

[0037] Note that, here, a configuration example in which each of the plurality of physical cameras 16 is arranged so as to surround the soccer field 24 is given, but the technology of the present disclosure is not limited to this. For example, the plurality of physical cameras 16 may be arranged so as to surround a specific part inside the soccer field 24. The positions and / or directions of the plurality of physical cameras 16 can be changed and are determined according to the virtual viewpoint image required by the user 18 or the like.

[0038] Although illustration is omitted, at least one physical camera 16 is installed in an unmanned aircraft (e.g., a multi-rotor unmanned aircraft), and imaging may be performed in a state of overlooking from above an area including a soccer field 24 as an imaging area.

[0039] The image processing device 12 is installed in the control room 32. The plurality of physical cameras 16 and the image processing device 12 are connected via a LAN cable 30. The image processing device 12 controls the plurality of physical cameras 16 and acquires images obtained by imaging by each of the plurality of physical cameras 16. Here, although connection using a wired communication method via the LAN cable 30 is exemplified, the present disclosure is not limited to this, and connection using a wireless communication method may be used.

[0040] In the soccer stadium 22, spectator seats 26 are provided so as to surround the soccer field 24, and a user 18 is seated on the spectator seats 26. The user 18 has a user device 14, and the user device 14 is used by the user 18. Here, although a form example in which the user 18 exists inside the soccer stadium 22 is described, the technology of the present disclosure is not limited to this, and the user 18 may exist outside the soccer stadium 22.

[0041] As shown in FIG. 2 as an example, the image processing device 12 acquires imaging images 46B indicating imaging areas when observed from the positions of the plurality of physical cameras 16 from each of the plurality of physical cameras 16. The imaging image 46B is a frame image indicating an imaging area when observed from the position of the physical camera 16. That is, the imaging image 46B is obtained by imaging the imaging area by each of the plurality of physical cameras 16. In the imaging image 46B, physical camera identification information for identifying the physical camera 16 used for imaging and the time when imaging was performed by the physical camera 16 (hereinafter, also referred to as "physical camera imaging time") are given for each frame. In addition, physical camera installation position information capable of specifying the installation position (imaging position) of the physical camera 16 used for imaging is also given for each frame in the imaging image 46B.

[0042] The image processing device 12 generates an image using 3D polygons by synthesizing a plurality of captured images 46B obtained by capturing an imaging area with a plurality of physical cameras 16. Then, based on the image using the generated 3D polygons, the image processing device 12 generates, one frame at a time, a virtual viewpoint image 46C showing the imaging area when the imaging area is observed from an arbitrary position and an arbitrary direction.

[0043] Here, the captured image 46B is an image obtained by being captured by the physical camera 16, whereas the virtual viewpoint image 46C can be considered as an image obtained by being captured by a virtual imaging device, that is, a virtual camera 42, from an arbitrary position and an arbitrary direction. The virtual camera 42 is not actually present as an object and is a virtual camera that cannot be visually recognized. In the present embodiment, virtual cameras are installed at a plurality of locations within the soccer stadium 22 (see FIG. 3). All the virtual cameras 42 are installed at different positions from each other. Also, all the virtual cameras 42 are installed at positions different from those of all the physical cameras 16. That is, all the physical cameras 16 and all the virtual cameras 42 are installed at different positions from each other.

[0044] The virtual viewpoint image 46C is provided, for each frame, with virtual camera identification information for identifying the virtual camera 42 used for imaging and the time when the imaging was performed by the virtual camera 42 (hereinafter also referred to as the "virtual camera imaging time"). Also, the virtual viewpoint image 46C is provided with virtual camera installation position information that can identify the installation position (imaging position) of the virtual camera 42 used for imaging.

[0045] Hereinafter, for convenience of explanation, when there is no need to distinguish between the physical camera 16 and the virtual camera 42, they are simply referred to as "camera". Also, hereinafter, for convenience of explanation, when there is no need to distinguish between the captured image 46B and the virtual viewpoint image 46C, they are referred to as "camera image". Also, hereinafter, for convenience of explanation, when there is no need to distinguish between the physical camera identification information and the virtual camera identification information, they are referred to as "camera identification information". Also, hereinafter, for convenience of explanation, when there is no need to distinguish between the physical camera imaging time and the virtual camera imaging time, they are referred to as "imaging time". Also, hereinafter, for convenience of explanation, when there is no need to distinguish between the physical camera installation position information and the virtual camera installation position information, they are referred to as "camera installation position information". Note that the camera identification information, the imaging time, and the camera installation position information are attached to each camera image in, for example, the Exif format.

[0046] The image processing apparatus 12 holds, for example, camera images for a predetermined time period (e.g., several hours to several tens of hours). Therefore, for example, the image processing apparatus 12 acquires a camera image at a specified imaging time from a group of camera images for the predetermined time period, and processes the acquired camera image.

[0047] The position of the virtual camera 42 (hereinafter also referred to as "virtual camera position") 42A and the orientation (hereinafter also referred to as "virtual camera orientation") 42B are changeable. Also, the angle of view of the virtual camera 42 is changeable.

[0048] In the present first embodiment, it is referred to as the virtual camera position 42A, but generally, the virtual camera position 42A is also referred to as the viewpoint position. Also, in the present first embodiment, it is referred to as the virtual camera orientation 42B, but generally, the virtual camera orientation 42B is also referred to as the line-of-sight direction. Here, the viewpoint position means, for example, the position of a virtual person's viewpoint, and the line-of-sight direction means, for example, the direction of a virtual person's line of sight.

[0049] That is, in the present embodiment, for convenience of explanation, the virtual camera position 42A is used for the explanation, but it is not essential to use the virtual camera position 42A. "Installing a virtual camera" means determining the viewpoint position, the line-of-sight direction, and / or the angle of view for generating the virtual viewpoint image 46C. Therefore, for example, it is not limited to the mode of installing an object such as a virtual camera on the imaging area on a computer, and another method such as specifying the coordinates and / or direction of the viewpoint position numerically may be used. Further, "imaging by the virtual camera" means generating a virtual viewpoint image 46C corresponding to the case of viewing the imaging area from the position and direction where the "virtual camera is installed".

[0050] In the example shown in FIG. 2, as an example of the virtual viewpoint image 46C, a virtual viewpoint image showing the imaging area when observing the imaging area from the virtual camera position 42A and the virtual camera orientation 42B within the spectator seat 26 is shown. The virtual camera position and the virtual camera orientation are not fixed. That is, the virtual camera position and the virtual camera orientation can be changed according to an instruction from the user 18 or the like. For example, the image processing apparatus 12 can also set the position of a person (hereinafter also referred to as "target person") designated as a target subject among the soccer players and referees in the soccer field 24 as the virtual camera position, and set the line-of-sight direction of the target person as the virtual camera orientation.

[0051] As shown in FIG. 3 as an example, the virtual cameras 42 are installed at a plurality of locations within the soccer field 24 and at a plurality of locations around the soccer field 24. Note that the installation mode of the virtual cameras 42 shown in FIG. 3 is merely an example. For example, the virtual cameras 42 may not be installed within the soccer field 24 and may be installed only around the soccer field 24, or the virtual cameras 42 may not be installed around the soccer field 24 and may be installed only within the soccer field 24. Also, the number of installed virtual cameras 42 may be more or less than the example shown in FIG. 3. Further, each virtual camera position 42A and each virtual camera orientation 42B of the virtual cameras 42 can also be changed.

[0052] As an example, as shown in FIG. 4, the image processing apparatus 12 includes a computer 50, an RTC 51, a reception device 52, a display 53, a first communication I / F 54, and a second communication I / F 56. The computer 50 includes a CPU 58, a storage 60, and a memory 62. The CPU 58 is an example of the "processor" according to the technology of the present disclosure. The memory 62 is an example of the "memory" according to the technology of the present disclosure. The computer 50 is an example of the "computer" according to the technology of the present disclosure.

[0053] The CPU 58, the storage 60, and the memory 62 are connected via a bus 64. In the example shown in FIG. 4, for the sake of illustration, one bus is shown as the bus 64, but a plurality of buses may be used. Also, the bus 64 may include a serial bus or a parallel bus composed of a data bus, an address bus, a control bus, and the like.

[0054] The CPU 58 controls the entire image processing apparatus 12. The storage 60 stores various parameters and various programs. The storage 60 is a non-volatile storage device. Here, as an example of the storage 60, an EEPROM is applied. However, this is merely an example, and an SSD or an HDD may be used. The memory 62 is a storage device. Various information is temporarily stored in the memory 62. The memory 62 is used as a work memory by the CPU 58. Here, as an example of the memory 62, a RAM is applied. However, this is merely an example, and other types of storage devices may be used.

[0055] The RTC51 is powered by a power supply system separated from the power supply system for the computer 50 and continues to tick the current time (for example, year, month, day, hour, minute, second) even when the computer 50 is in a shutdown state. Each time the current time is updated, the RTC51 outputs the current time to the CPU58. The CPU58 uses the current time input from the RTC51 as the imaging time. Here, an example of the form in which the CPU58 obtains the current time from the RTC51 is given, but the technology of the present disclosure is not limited thereto. For example, the CPU58 may obtain the current time provided from an external device (not shown) via the network 20 (for example, obtain it using SNTP and / or NTP), or may obtain the current time from a built-in or connected GNSS device (for example, a GPS device).

[0056] The reception device 52 receives instructions from users of the image processing apparatus 12 and the like. Examples of the reception device 52 include a touch panel, hard keys, and a mouse. The reception device 52 is connected to the bus 64 and the like, and the instructions received by the reception device 52 are obtained by the CPU58.

[0057] The display 53 is connected to the bus 64 and displays various information under the control of the CPU58. An example of the display 53 includes a liquid crystal display. Note that not limited to a liquid crystal display, other types of displays such as an EL display (for example, an organic EL display or an inorganic EL display) may be adopted as the display 53.

[0058] The first communication I / F 54 is connected to the LAN cable 30. The first communication I / F 54 is realized by a device having, for example, an FPGA. The first communication I / F 54 is connected to the bus 64 and controls the exchange of various information between the CPU 58 and the plurality of physical cameras 16. For example, the first communication I / F 54 controls the plurality of physical cameras 16 according to the request of the CPU 58. Further, the first communication I / F 54 acquires the captured image 46B (see FIG. 2) obtained by being captured by each of the plurality of physical cameras 16, and outputs the acquired captured image 46B to the CPU 58. Here, although the first communication I / F 54 is illustrated as a wired communication I / F, it may be a wireless communication I / F such as a high-speed wireless LAN.

[0059] The second communication I / F 56 is wirelessly communicably connected to the network 20. The second communication I / F 56 is realized by a device having, for example, an FPGA. The second communication I / F 56 is connected to the bus 64. The second communication I / F 56 controls the exchange of various information between the CPU 58 and the user device 14 in a wireless communication method via the network 20.

[0060] Note that at least one of the first communication I / F 54 and the second communication I / F 56 can be configured with a fixed circuit instead of an FPGA. Further, at least one of the first communication I / F 54 and the second communication I / F 56 may be a circuit configured with an ASIC, an FPGA, and / or a PLD or the like.

[0061] As an example, as shown in FIG. 5, the user device 14 includes a computer 70, a gyro sensor 74, a reception device 76, a display 78, a microphone 80, a speaker 82, a physical camera 84, and a communication I / F 86. The computer 70 includes a CPU 88, a storage 90, and a memory 92, and the CPU 88, the storage 90, and the memory 92 are connected via a bus 94. In the example shown in FIG. 5, for the sake of illustration, one bus is shown as the bus 94, but the bus 94 is configured as a serial bus or includes a data bus, an address bus, a control bus, etc.

[0062] The CPU 88 controls the entire user device 14. The storage 90 stores various parameters and various programs. The storage 90 is a non-volatile storage device. Here, as an example of the storage 90, an EEPROM is applied. However, this is only an example, and an SSD, an HDD, or the like may also be used. Various information is temporarily stored in the memory 92, and the memory 92 is used as a work memory by the CPU 88. Here, as an example of the memory 92, a RAM is applied. However, this is only an example, and other types of storage devices may also be used.

[0063] The gyro sensor 74 measures the angle around the yaw axis of the user device 14 (hereinafter also referred to as the "yaw angle"), the angle around the roll axis of the user device 14 (hereinafter also referred to as the "roll angle"), and the angle around the pitch axis of the user device 14 (hereinafter also referred to as the "pitch angle"). The gyro sensor 74 is connected to the bus 94, and the angle information indicating the yaw angle, roll angle, and pitch angle measured by the gyro sensor 74 is acquired by the CPU 88 via the bus 94 or the like.

[0064] The reception device 76 receives an instruction from the user 18 (see FIGS. 1 and 2). Examples of the reception device 76 include a touch panel 76A and hard keys. The reception device 76 is connected to the bus 94, and the instruction received by the reception device 76 is acquired by the CPU 88.

[0065] The display 78 is connected to the bus 94 and displays various information under the control of the CPU 88. An example of the display 78 is a liquid crystal display. Note that not limited to the liquid crystal display, other types of displays such as an EL display (for example, an organic EL display or an inorganic EL display) may be adopted as the display 78.

[0066] The user device 14 is provided with a touch panel display, which is realized by a touch panel 76A and a display 78. That is, the touch panel display is formed by superimposing the touch panel 76A on the display area of the display 78, or by incorporating a touch panel function inside the display 78 (the "in-cell" type). Note that the "in-cell" type touch panel display is merely an example, and an "out-cell" type or "on-cell" type touch panel display may also be used.

[0067] The microphone 80 converts the collected sound into an electrical signal. The microphone 80 is connected to the bus 94. The electrical signal obtained by converting the sound collected by the microphone 80 is acquired by the CPU 88 via the bus 94.

[0068] The speaker 82 converts an electrical signal into sound. The speaker 82 is connected to the bus 94. The speaker 82 receives the electrical signal output from the CPU 88 via the bus 94, converts the received electrical signal into sound, and outputs the sound obtained by converting the electrical signal to the outside of the user device 14.

[0069] The physical camera 84 captures an object to obtain an image indicating the object. The physical camera 84 is connected to the bus 94. The image obtained by capturing the object by the physical camera 84 is acquired by the CPU 88 via the bus 94. Note that the image obtained by capturing the object by the physical camera 84 may also be used for generating the virtual viewpoint image 46C together with the captured image 46B.

[0070] The communication I / F 86 is wirelessly communicably connected to the network 20. The communication I / F 86 is realized by a device composed of, for example, a circuit (such as an ASIC, an FPGA, and / or a PLD, etc.). The communication I / F 86 is connected to the bus 94. The communication I / F 86 controls the exchange of various information between the CPU 88 and an external device in a wireless communication manner via the network 20. Here, as the "external device", for example, the image processing device 12 is mentioned.

[0071] Each of the plurality of physical cameras 16 (see FIGS. 1 to 4) generates a moving image (hereinafter also referred to as a "physical camera moving image") indicating the imaging area by imaging the imaging area. In the present first embodiment, any one of the plurality of physical cameras 16 is used as a reference physical camera. The physical camera moving image (hereinafter also referred to as a "reference physical camera moving image") obtained by being imaged by the reference physical camera is distributed to, for example, the user device 14 and displayed on the display 78 of the user device 14. Then, the user 18 views the reference physical camera moving image displayed on the display 78.

[0072] The physical camera moving image is obtained by being imaged by the physical camera 16 at a specific frame rate (for example, 60 fps). As shown in FIG. 6 as an example, the physical camera moving image is a plurality of frame images composed of a plurality of frames obtained according to the specific frame rate. That is, the physical camera moving image is configured by arranging a plurality of captured images 46B obtained at timings defined by the specific frame rate in time series.

[0073] In the example shown in FIG. 6, among the plurality of captured images 46B included in the physical camera moving image, three frames of captured images 46B1 to 46B3 including a target person image 96 indicating the target person are shown. Here, the target person is an example of the "object" according to the technology of the present disclosure, and the target person image 96 is an example of the "object image" according to the technology of the present disclosure.

[0074] The captured images 46B1 to 46B3 for three frames are roughly classified into the captured image 46B1 of the first frame, the captured image 46B2 of the second frame, and the captured image 46B3 of the third frame from the oldest frame to the latest frame. In the captured image 46B1 of the first frame, the entire target person image 96 appears at a visible position including the facial expression of the target person.

[0075] However, in the captured image 46B2 of the second frame and the captured image 46B3 of the third frame, most of the area including the face of the target person in the target person image 96 is blocked at a level where it cannot be visually recognized by the person images showing people other than the target person. When the physical camera moving image shown in FIG. 6 is displayed on the display 78 of the user device 14 as the reference physical camera moving image, the user 18 has difficulty grasping the overall state of the target person image 96 from at least the captured images 46B2 and 46B3 of the second and third frames. In particular, when the user 18 hopes to observe the facial expression of the target person, the facial expression of the target person cannot be observed from at least the captured images 46B2 and 46B3 of the second and third frames. Thus, in the example shown in FIG. 6, an image that allows the target person to be observed cannot be continuously provided to the user 18.

[0076] In view of such circumstances, as shown in FIG. 7 as an example, in the image processing apparatus 12, an output control program 100 is stored in the storage 60. Then, the CPU 58 executes an output control process (FIGS. 14A and 14B) described later according to the output control program 100.

[0077] The CPU 58 reads the output control program 100 from the storage 60 and executes the output control program 100 on the memory 62, thereby operating as a virtual viewpoint image generation unit 58A, an image acquisition unit 58B, a detection unit 58C, an output unit 58D, and an image selection unit 58E.

[0078] The storage 60 stores an image group 102. The image group 102 includes physical camera moving images and virtual viewpoint moving images. The physical camera moving images are roughly classified into a reference physical camera moving image and other physical camera moving images obtained by imaging with a physical camera 16 other than the reference physical camera (hereinafter also referred to as "other physical cameras"). In the present first embodiment, there are a plurality of other physical cameras. The reference physical camera moving image includes a plurality of captured images 46B obtained by imaging with the reference physical camera as reference physical camera images in time series. The other physical camera moving images include a plurality of captured images 46B obtained by imaging with other physical cameras as other physical camera images in time series.

[0079] The virtual viewpoint moving image is obtained by imaging with a virtual camera 42 (see FIGS. 2 and 3) at a specific frame rate. As shown in FIG. 7 as an example, the virtual viewpoint moving image is a plurality of frame images composed of a plurality of frames obtained according to a specific frame rate. That is, the virtual viewpoint moving image is configured by arranging a plurality of virtual viewpoint images 46C obtained at timings defined by a specific frame rate in time series. In the present first embodiment, as described above, there are a plurality of virtual cameras 42, and a virtual viewpoint moving image is obtained by each virtual camera 42 and stored in the storage 60.

[0080] In the following, for convenience of explanation, a camera image obtained by imaging with a camera other than the reference physical camera is referred to as an "other camera image". That is, the other camera image refers to a general term for other physical camera images and virtual viewpoint images.

[0081] In the present first embodiment, the detection unit 58C performs a detection process. The detection process is a process of detecting a target person image 96 from each of a plurality of camera images obtained by imaging with a plurality of cameras having different positions. In the detection process, the target person image 96 is detected when a face image showing the face of the target person is detected. Examples of the detection process include a first detection process (see FIG. 11) described later and a second detection process (see FIG. 12) described later.

[0082] Also, in this first embodiment, the output unit 58D outputs a reference physical camera image among a plurality of camera images. Further, when the output unit 58D transitions from a detection state in which the target person image 96 is detected from the reference physical camera image by the detection process to a non-detection state in which the target person image 96 is not detected from the reference physical camera image by the detection process, among the plurality of camera images, the output unit 58D outputs another camera image in which the target person image 96 is detected by the detection process. For example, when transitioning from the detection state to the non-detection state under the situation where the reference physical camera image is being output, the output unit 58D switches from outputting the reference physical camera image to outputting another camera image.

[0083] Here, the transition from the detection state to the non-detection state means that the reference physical camera image that is the output target by the output unit 58D switches from the reference physical camera image in which the target person is captured to the reference physical camera image in which the target person is not captured. More specifically, the transition from the detection state to the non-detection state means that among the plurality of reference physical camera images included in the reference physical camera moving image, between temporally adjacent frames, the output target by the output unit 58D switches from the frame in which the target person is captured to the frame in which the target person is not captured. For example, when the reference physical camera is imaging the same imaging area, as shown in the captured images 46B1 to 46B2 in FIG. 6, due to the movement of an object (for example, the target person or an object around the target person, etc.) in the imaging area, the target person image 96 changes from a state where it can be detected to a state where it is hidden by another person or the like and cannot be detected.

[0084] Note that in this first embodiment, the camera image is an example of the "image" according to the technology of the present disclosure. Also, the reference physical camera image is an example of the "first image" according to the technology of the present disclosure. Further, the other camera image is an example of the "second image" according to the technology of the present disclosure.

[0085] In this embodiment, the virtual viewpoint image generation unit 58A generates a plurality of virtual viewpoint moving images by causing each of all the virtual cameras 42 to perform imaging. As shown in FIG. 8 as an example, the virtual viewpoint image generation unit 58A acquires a physical camera moving image from the storage 60. Based on the physical camera moving image acquired from the storage 60, the virtual viewpoint image generation unit 58A generates a virtual viewpoint moving image corresponding to the virtual camera position, virtual camera orientation, and field of view angle set at the current time for each virtual camera 42. Then, the virtual viewpoint image generation unit 58A stores the generated virtual viewpoint moving images in the storage 60 in units of virtual cameras 42.

[0086] Here, the virtual viewpoint moving image corresponding to the virtual camera position, virtual camera orientation, and field of view angle set at the current time means, for example, a moving image showing the area observed at the field of view angle set at the current time from the virtual camera position and virtual camera orientation set at the current time.

[0087] Also, here, an example form is given in which the virtual viewpoint image generation unit 58A generates a plurality of virtual viewpoint moving images by causing each of all the virtual cameras 42 to perform imaging. However, it is not necessarily required to cause each of all the virtual cameras 42 to perform imaging. For example, depending on the performance of the computer or the like, generation of virtual viewpoint moving images by some of the virtual cameras 42 may not be performed.

[0088] As shown in FIG. 9 as an example, the output unit 58D acquires a reference physical camera moving image from the storage 60 and outputs the acquired reference physical camera moving image to the user device 14. As a result, the reference physical camera moving image is displayed on the display 78 of the user device 14.

[0089] As an example, as shown in FIG. 10, with the reference physical camera moving image being displayed on the display 78 of the user device 14, the user 18 designates, with a finger via the touch panel 76A, the area of interest (hereinafter also referred to as the "area of interest"). In the example shown in FIG. 10, the area of interest is the area including the target person image 96 within the reference physical camera moving image displayed on the display 78.

[0090] The user device 14 transmits, to the image acquisition unit 58B, area-of-interest information indicating the area of interest within the reference physical camera moving image. The image acquisition unit 58B receives the area-of-interest information transmitted from the user device 14. The image acquisition unit 58B performs image analysis (for example, image analysis by a cascade classifier and / or pattern matching, etc.) on the received area-of-interest information, and extracts the target person image 96 from the area of interest indicated by the area-of-interest information. The image acquisition unit 58B stores, as the target person image sample 98, the target person image 96 extracted from the area of interest in the storage 60.

[0091] As an example, as shown in FIG. 11, the image acquisition unit 58B acquires a reference physical camera image in units of one frame from the reference physical camera moving image within the storage 60. The detection unit 58C executes a first detection process. The first detection process is a process of detecting the target person image 96 from the reference physical camera image by performing image analysis on the reference physical camera image acquired by the image acquisition unit 58B using the target person image sample 98 within the storage 60. Examples of the image analysis include image analysis by a cascade classifier and / or pattern matching, etc.

[0092] The target person image 96 detected by the first detection process includes an image indicating a target person in a mode different from the mode of the target person indicated by the target person image 96 shown in FIG. 10. That is, the detection unit 58C executes the first detection process to determine whether or not the target person indicated by the target person image sample 98 is captured in the reference physical camera image.

[0093] When the target person image 96 is detected by the first detection process, the output unit 58D outputs the reference physical camera image that was the processing target of the first detection process, that is, the reference physical camera image including the target person image 96, to the user device 14. As a result, the reference physical camera image including the target person image 96 is displayed on the display 78 of the user device 14.

[0094] As an example, as shown in FIG. 12, when the target person image 96 is not detected by the first detection process, the image acquisition unit 58B acquires a plurality of other camera images at the same imaging time as the reference physical camera image that was the processing target of the first detection process from the storage 60. Hereinafter, for convenience of explanation, the plurality of other camera images are also referred to as an "other camera image group".

[0095] The detection unit 58C executes a second detection process for each of the other camera images included in the other camera image group acquired by the image acquisition unit 58B. The second detection process is different from the first detection process in that the other camera images are used as the processing target instead of the reference physical camera image.

[0096] When there are a plurality of other camera images in which the target person image 96 is detected by the second detection process, the image selection unit 58E selects an other camera image that satisfies the best imaging conditions from the other camera image group including the target person image 96 detected by the second detection process. The best imaging conditions refer to conditions such that, for example, among the other camera image group, the position of the target person image 96 within the other camera image is within a predetermined range, and the size of the target person image 96 within the other camera image is equal to or greater than a predetermined size. In the first embodiment, as an example of the best imaging conditions, a condition is used in which the entire target person indicated by the target person image 96 is most prominently shown within a central frame predetermined in the central portion of the frame. The shape and / or size of the central frame may be fixed, or may be changed according to a given instruction and / or condition. Not limited to the central frame, a frame may be provided at other positions.

[0097] Also, here, a condition that the entire target person is captured within the central frame is exemplified, but this is merely an example, and a condition that an area of a predetermined ratio (for example, 80%) or more including the face of the target person within the central frame is captured may also be used. Note that the predetermined ratio may be a fixed value or a variable value that is changed according to a given instruction and / or condition.

[0098] As an example, as shown in FIG. 13, the image selection unit 58E selects an other camera image that satisfies the best imaging conditions from the group of other camera images including the target person image 96 detected by the second detection process, and outputs the selected other camera image to the output unit 58D. Further, when there is one frame of the other camera image in which the target person image 96 is detected by the second detection process, the detection unit 58C outputs the other camera image in which the target person image 96 is detected to the output unit 58D.

[0099] The output unit 58D outputs the other camera image input from the detection unit 58C or the image selection unit 58E to the user device 14. As a result, an other camera image including the target person image 96 is displayed on the display 78 of the user device 14.

[0100] On the other hand, when the target person image 96 is not detected by the second detection process, as shown in FIG. 11 as an example, the output unit 58D outputs the reference physical camera image that is the processing target of the first detection process to the user device 14. In this case, the reference physical camera image in which the target person image 96 is not detected by the first detection process is output to the user device 14. As a result, the reference physical camera image in which the target person image 96 is not detected by the first detection process is displayed on the display 78 of the user device 14.

[0101] Next, the operation of the image processing system 10 will be described with reference to FIGS. 14A and 14B.

[0102] Figures 14A and 14B show an example of the flow of output control processing executed by the CPU 58. The flow of output control processing shown in Figures 14A and 14B is an example of the "image processing method" according to the technology of the present disclosure. Note that the following description of the output control processing assumes, for convenience of explanation, that the image group 102 is already stored in the storage 60. Also, the following description of the output control processing assumes, for convenience of explanation, that the target person image sample 98 is already stored in the storage 60.

[0103] In the output control processing shown in Figure 14A, first, in step ST10, the image acquisition unit 58B acquires an unprocessed reference physical camera image for one frame from the reference physical camera moving image in the storage 60, and then the output control processing proceeds to step ST12. Here, the unprocessed reference physical camera image refers to a reference physical camera image for which the processing in step ST12 has not yet been performed.

[0104] In step ST12, the detection unit 58C executes the first detection process on the reference physical camera image acquired in step ST10, and then the output control processing proceeds to step ST14.

[0105] In step ST14, the detection unit 58C determines whether the target person image 96 has been detected from the reference physical camera image by the first detection process. In step ST14, if the target person image 96 has not been detected from the reference physical camera image by the first detection process, the determination is negative, and the output control processing proceeds to step ST18 shown in Figure 14B. In step ST14, if the target person image 96 has been detected from the reference physical camera image by the first detection process, the determination is positive, and the output control processing proceeds to step ST16.

[0106] In step ST16, the output unit 58D outputs the reference physical camera image that was the processing target of the first detection process in step ST14 to the user device 14, and then the output control process proceeds to step ST32. When the reference physical camera image is output to the user device 14 by executing the process of step ST16, the reference physical camera image is displayed on the display 78 of the user device 14 (see FIG. 11).

[0107] In step ST18 shown in FIG. 14B, the image acquisition unit 58B acquires a group of other camera images at the same imaging time as the reference physical camera image that was the processing target of the first detection process from the storage 60, and then the output control process proceeds to step ST20.

[0108] In step ST20, the detection unit 58C executes a second detection process on the group of other camera images acquired in step ST18, and then the output control process proceeds to step ST22.

[0109] In step ST22, the detection unit 58C determines whether or not the target person image 96 has been detected from the group of other camera images acquired in step ST18. In step ST22, if the target person image 96 has not been detected from the group of other camera images acquired in step ST18, the determination is negative and the output control process proceeds to step ST16 shown in FIG. 14A. In step ST22, if the target person image 96 has been detected from the group of other camera images acquired in step ST18, the determination is positive and the output control process proceeds to step ST24.

[0110] In step ST24, the detection unit 58C determines whether or not there are a plurality of other camera images in which the target person image 96 has been detected by the second detection process. In step ST24, if there are a plurality of other camera images in which the target person image 96 has been detected by the second detection process, the determination is positive and the output control process proceeds to step ST26. In step ST24, if there is one frame of other camera image in which the target person image 96 has been detected by the second detection process, the determination is negative and the output control process proceeds to step ST30.

[0111] In step ST26, the image selection unit 58E selects, from the other camera image group in which the target person image 96 has been detected by the second detection process, other camera images that satisfy the best imaging conditions (see FIG. 12). After that, the output control process proceeds to step ST28.

[0112] In step ST28, the output unit 58D outputs the other camera images selected in step ST26 to the user device 14. After that, the output control process proceeds to step ST32 shown in FIG. 14A. When the other camera images are output to the user device 14 by executing the process of step ST28, the other camera images are displayed on the display 78 of the user device 14 (see FIG. 13).

[0113] In step ST30, the output unit 58D outputs the other camera images in which the target person image 96 has been detected by the second detection process to the user device 14. After that, the output control process proceeds to step ST32 shown in FIG. 14A. When the other camera images are output to the user device 14 by executing the process of step ST30, the other camera images are displayed on the display 78 of the user device 14 (see FIG. 13).

[0114] In step ST32 shown in FIG. 14A, the output unit 58D determines whether or not a condition for ending the output control process (hereinafter also referred to as the “output control process end condition”) is satisfied. As an example of the output control process end condition, there is a condition that an instruction to end the output control process is given to the image processing apparatus 12. The instruction to end the output control process is received, for example, by the reception device 52 or 76. In step ST32, if the output control process end condition is not satisfied, the determination is negative and the output control process proceeds to step ST10. In step ST32, if the output control process end condition is satisfied, the determination is affirmative and the output control process ends.

[0115] In this way, by executing the output control process, the reference physical camera image in which the target person image 96 is not blocked by an obstacle is output to the user device 14 by the output unit 58D. Also, when the target person image 96 is blocked by an obstacle in the reference physical camera image, instead of the reference physical camera image in which the target person image 96 is blocked by an obstacle, a virtual viewpoint image 46C in which the entire target person image 96 is visible is output to the user device 14 by the output unit 58D. Thereby, it is possible to continuously provide the user 18 with a camera image capable of observing the target person.

[0116] Also, when the output control process is executed, as shown in FIG. 15 as an example, when the target person image 96 in the reference physical camera image transitions from a state where it is not blocked by an obstacle to a state where it is blocked by an obstacle in the situation where the reference physical camera moving image is being output, the output switches from the output of the reference physical camera moving image to the output of the virtual viewpoint moving image. Thereby, it is possible to continuously provide the user 18 with a camera image capable of observing the target person.

[0117] Also, when the output control process is executed, as shown in FIGS. 15 and 16 as an example, the output unit 58D switches from the output of the reference physical camera moving image to the output of the virtual viewpoint moving image at the timing when the target person image 96 comes to be blocked by an obstacle in the reference physical camera image. Then, the output unit 58D ends the output of the virtual viewpoint moving image at a timing later than the timing when the target person image 96 comes to be blocked by an obstacle in the reference physical camera image. That is, the output of the virtual viewpoint moving image ends at a timing later than the timing when the target person image 96 is not detected by the first detection process. Thereby, it is possible to provide the user 18 with a virtual viewpoint moving image capable of observing the target person after the state where the target person image 96 is not detected by the first detection process.

[0118] Also, when the output control process is executed, as an example, as shown in FIG. 16, on the condition that the target person image 96 in the reference physical camera image has returned from a state of being blocked by an obstacle to a state of not being blocked by an obstacle, the output unit 58D resumes the output of the reference physical camera moving image. That is, on the condition that the target person image 96 has returned from a state where it is not detected from the reference physical camera image by the first detection process to a state where it is detected from the reference physical camera image by the first detection process, the output is switched from the virtual viewpoint moving image to the reference physical camera moving image. Thereby, compared with the case where the output of the virtual viewpoint moving image continues even though the target person image 96 in the reference physical camera image has returned from a state of being blocked by an obstacle to a state of not being blocked by an obstacle, the labor of switching the output from the virtual viewpoint moving image to the reference physical camera moving image can be reduced.

[0119] Also, when the output control process is executed, another camera image that satisfies the best imaging conditions is selected by the image selection unit 58E (see step ST26 shown in FIG. 14B), and the selected another camera image is output to the user device 14 by the output unit 58D (see step ST28 shown in FIG. 14B). Thereby, compared with the case where another camera image in which the target person image 96 is simply detected without considering the position and size of the target person image 96 in the other camera image is output, the user 18 can more easily find the target person image 96 in the other camera image.

[0120] Also, in the output control process, the target person image 96 is detected by detecting a face image showing the face of the target person by the first detection process and the second detection process. Therefore, the target person image 96 can be detected with higher accuracy compared with the case where the face image is not detected.

[0121] Also, when the output control process is executed, a plurality of frame images composed of a plurality of frames are output to the user device 14 by the output unit 58D. Examples of the plurality of frame images include a reference physical camera moving image and a virtual viewpoint moving image as shown in FIGS. 15 and 16. Therefore, according to this configuration, the user 18 who is viewing the reference physical camera moving image and the virtual viewpoint moving image can be continuously observed for the target person.

[0122] In the image processing system 10, the imaging area is imaged by a plurality of physical cameras 16 and also by a plurality of virtual cameras 42. Therefore, compared with the case where the imaging area is imaged only by the physical cameras 16 without using the virtual cameras 42, the user 18 can observe the target person from various positions and orientations. Here, a plurality of physical cameras 16 and a plurality of virtual cameras 42 are exemplified, but the technology of the present disclosure is not limited to this. The number of physical cameras 16 may be one, and the number of virtual cameras 42 may also be one.

[0123] In the first embodiment, an example of the form in which the output of the virtual viewpoint moving image is terminated at a timing later than the timing when the target person image 96 is not detected by the first detection process has been described. However, the technology of the present disclosure is not limited to this. For example, not only the output of the virtual viewpoint moving image is terminated at a timing later than the timing when the target person image 96 is not detected by the first detection process, but the output unit 58D may start the output of the virtual viewpoint moving image from a timing earlier than the timing when the target person image 96 is not detected by the first detection process. For example, in the case of a moving image that has already been imaged, since the timing when the target person image 96 is not detected in the reference physical camera moving image can be recognized, the virtual viewpoint moving image can be output from a timing earlier than the timing when the target person image 96 is not detected in the reference physical camera moving image. Thereby, it is possible to provide the user 18 with a virtual viewpoint moving image that can observe the target person before the state where the target person image 96 is not detected by the first detection process.

[0124] In addition, in the above-described first embodiment, as an example of a form in which, when there are a plurality of other camera images in which the target person image 96 has been detected by the second detection process, other camera videos that satisfy the best imaging conditions are output, it is not necessarily required to output other camera videos that satisfy the best imaging conditions. For example, if any of the other camera images in which the target person image 96 has been detected is output, the user 18 can visually recognize the target person image 96.

[0125] In addition, in the above-described first embodiment, as an example of the best imaging conditions, among the group of other camera images, the position of the target person image 96 in the other camera image is within a predetermined range, and the size of the target person image 96 in the other camera image is equal to or greater than a predetermined size. However, the technology of the present disclosure is not limited to this. For example, the best imaging conditions may be a condition that, among the group of other camera images, the position of the target person image 96 in the other camera image is within a predetermined range, or a condition that the size of the target person image 96 in the other camera image is equal to or greater than a predetermined size.

[0126] In addition, in the above-described first embodiment, as an example, as shown in FIG. 17, an example of a form in which the output of the reference physical camera image is directly switched to the output of the virtual viewpoint image 46C that satisfies the best imaging conditions has been described. However, the technology of the present disclosure is not limited to this. If the output is directly switched from the output of the reference physical camera image to the output of the virtual viewpoint image 46C that satisfies the best imaging conditions, there is a possibility that it will be difficult to grasp the position of the target person before and after the output is switched.

[0127] Therefore, as an example, as shown in FIG. 18, during the period when the output unit 58D switches from the output of the reference physical camera image to the output of the virtual viewpoint image 46C that satisfies the best imaging conditions, the output unit 58D outputs camera images obtained by imaging with a plurality of cameras that continuously connect the position, orientation, and field of view. The camera images obtained by imaging with a plurality of cameras that continuously connect the position, orientation, and field of view refer to, for example, a plurality of virtual viewpoint images 46C obtained by imaging with a plurality of virtual cameras 42 that continuously connect from the imaging position, imaging direction, and field of view of the reference physical camera to the virtual camera position, virtual camera orientation, and field of view of the virtual camera 42 used for imaging to obtain the virtual viewpoint image 46C that satisfies the best imaging conditions. Thereby, compared with the case where the output is directly switched from the reference physical camera image to the virtual viewpoint image 46C, it is possible to make it easier for the user 18 to grasp the position of the target person.

[0128] In the first embodiment, when the target person image 96 is blocked by an obstacle in the reference physical camera image, instead of the reference physical camera image in which the target person image 96 is blocked by the obstacle, a virtual viewpoint image 46C in which the entire target person image 96 is visible or another physical camera image can be output to the user device 14 by the output unit 58D. However, the technology of the present disclosure is not limited to this. For example, when the target person image 96 is blocked by an obstacle in the reference physical camera image, only the virtual viewpoint image 46C in which the entire target person image 96 is visible may be output instead of the reference physical camera image in which the target person image 96 is blocked by the obstacle. Thereby, when the target person image 96 is not detected by the first detection process, the user 18 can continue to observe the target person by providing a virtual viewpoint moving image.

[0129] Further, even if the virtual viewpoint image 46C or other physical camera image in which the entire target person image 96 is visible is not output, for example, a virtual viewpoint image 46C or other physical camera image in which only a specific part such as the face shown by the target person image 96 is visible may be output. This specific part may be set according to an instruction given from the user 18. For example, when the face shown by the target person image 96 is set according to an instruction given from the user 18, a virtual viewpoint image 46C or other physical camera image in which the face of the target person is visible is output. Also, for example, a virtual viewpoint image 46C or other physical camera image in which the target person image 96 is visible at a ratio larger than the ratio of the target person image 96 visible in the reference physical camera image may be output.

[0130] Also, when the virtual viewpoint image 46C is output, it is not always necessary to output the image in which the target person image 96 is detected by the above detection process. For example, when the three-dimensional positions of each object in the imaging region are recognized by triangulation or the like and the target person image 96 is blocked by an obstacle in the reference physical camera image, from the positional relationship between the target person, the obstacle, and other objects, a virtual viewpoint image 46C showing the aspect observed from the viewpoint position, direction, and viewing angle from which the target person is estimated to be visible may be output. The detection process in the technology of the present disclosure also includes a process based on such an estimation.

[0131] In addition, in the above-described first embodiment, an example of the form in which the reference physical camera moving image is output by the output unit 58D has been described. However, the technology of the present disclosure is not limited to this. For example, as shown in FIG. 19, instead of the reference physical camera moving image, a reference virtual viewpoint moving image composed of a plurality of time-series virtual viewpoint images 46C obtained by imaging with a specific virtual camera 42 may be output by the output unit 58D to the user device 14. In this case, by executing the output control process, the output can be switched from the output of the reference virtual viewpoint moving image to the output of other camera images (in the example shown in FIG. 19, virtual viewpoint moving images other than the reference virtual viewpoint moving image). Even in the case where the output is switched from the output of the reference virtual viewpoint moving image to the output of other camera images in this way, the target person can be continuously observed by the user 18 in the same manner as in the first embodiment.

[0132] In addition, in the above-described first embodiment, an example of the form in which the physical camera image and the virtual viewpoint image 46C are selectively output by the output unit 58D has been described. However, as an example, as shown in FIG. 19, only the virtual viewpoint image 46C may be output by the output unit 58D, whether it is before or after the switch of the output. Also in this case, the target person can be continuously observed by the user 18 in the same manner as in the first embodiment.

[0133] In the first embodiment described above, an example of a form in which the output unit 58D switches from outputting the reference physical camera image to outputting other camera images has been described. However, the technology of the present disclosure is not limited to this. For example, as shown in FIG. 20, as other camera images, an overhead image including the target person image 96 may be output to the user device 14 by the output unit 58D. The overhead image refers to an image showing an overhead view of the imaging area (in the example shown in FIG. 20, the entire soccer field 24). For example, when the image selection unit 58E does not select other camera images that satisfy the best imaging conditions (when there are no other camera images that satisfy the best imaging conditions), the output unit 58D may output an overhead image. Therefore, according to the example of the form in which the overhead image is output by the output unit 58D, compared with the case where a camera image obtained by imaging only a part of the imaging area is output, a camera image in which the target person is more likely to be captured can be provided to the user 18.

[0134] In the first embodiment described above, an example of a form in which the reference physical camera obtains a reference physical camera moving image has been described. However, the reference physical camera moving image may be an image for television broadcast. Examples of the image for television broadcast include a recorded moving image or a live relay moving image. Also, not limited to moving images, still images may be used. For example, when the user 18 is watching a television broadcast video (e.g., an image for television relay, etc.) with the user device 14, when the target person image 96 is blocked by an obstacle in the television broadcast video, the virtual viewpoint image 46C or other physical camera image in which the target person image 96 is visible is output to the user device 14 using the technology described in the first embodiment above. Therefore, according to the example of the form in which an image for television broadcast is used as the reference physical camera moving image, even when the user 18 is watching an image for television relay, the user 18 can be made to continuously observe the target person.

[0135] In the first embodiment described above, the installation position of the reference physical camera is not particularly defined. However, it is preferable that the reference physical camera is a physical camera 16 that is installed at an observation position for observing the imaging region (for example, the soccer field 24) or in the vicinity of the observation position among the plurality of physical cameras 16. Further, when a reference virtual viewpoint moving image is output by the output unit 58D instead of the reference physical camera moving image, the imaging region may be imaged by a virtual camera 42 installed at an observation position for observing the imaging region (for example, the soccer field 24) or in the vicinity of the observation position. As an example of the observation position, for example, the position of the user 18 sitting in the spectator seat 26 shown in FIG. 1 can be mentioned. As a camera installed in the vicinity of the observation position, for example, a camera (for example, a physical camera 16 or a virtual camera 42) installed at the position closest to the user 18 sitting in the spectator seat 26 shown in FIG. 1 can be mentioned.

[0136] Therefore, according to this configuration, even when the user 18 is viewing a camera image obtained by imaging with a camera installed at an observation position for observing the imaging region or in the vicinity of the observation position among the plurality of cameras, the user 18 can be made to continuously observe the target person. Further, according to this configuration, when the user 18 is directly looking at the imaging region, the reference physical camera is imaging the same region or a region close to the region that the user 18 is looking at. Therefore, when the user 18 is directly looking at the imaging region (when directly observing in the real space), it is possible to detect from the reference physical camera moving image that the target person was not visible to the user 18. Thereby, when the target person becomes invisible directly from the user 18, it is possible to output the virtual viewpoint image 46C or another physical camera image in which the target person image 96 is visible to the user device 14.

[0137] In the first embodiment described above, when the target person image 96 transitions from a state where it is detected by the first detection process to a state where it is not detected, an example of a form in which the output of the reference physical camera moving image is switched to the output of the virtual viewpoint moving image in which the target person image 96 can be observed has been described. However, the technology of the present disclosure is not limited to this. For example, as shown in FIG. 21, even when the target person image 96 transitions from a state where it is detected by the first detection process to a state where it is not detected, the output unit 58D may continue to output the reference physical camera moving image and also output the virtual viewpoint moving image in which the target person image 96 can be observed in parallel. In this case, for example, as shown in FIG. 22, on the display 78 of the user device 14 which is the output destination of the camera image, the reference physical camera moving image and the virtual viewpoint moving image are displayed in parallel on different screens. Thereby, the user 18 can continuously observe the target person through the reference physical camera moving image and the virtual viewpoint moving image while enjoying the reference physical camera moving image. Note that instead of the reference physical camera moving image, a reference virtual viewpoint moving image may be output to the user device 14 by the output unit 58D. Also, instead of the virtual viewpoint moving image, another physical camera moving image may be output to the user device 14 by the output unit 58D. Further, for example, in the case where the user 18 has a plurality of user devices 14, for example, the reference physical camera moving image and the virtual viewpoint moving image may be output to different user devices 14 (one is not shown).

[0138] Also, in the first embodiment described above, the target person image 96 has been exemplified. However, the technology of the present disclosure is not limited to this, and it may be an image showing a non-person (an object other than a human). Examples of non-persons include robots (for example, robots imitating organisms such as humans, animals, or insects) equipped with devices capable of recognizing objects (for example, devices including a physical camera and a computer connected to the physical camera), animals, and insects.

[0139] [Second Embodiment] In the above-described first embodiment, a form example in which another camera image including the target person image 96 is output by the output unit 58D was described. However, in the second embodiment, a form example in which another camera image not including the target person image 96 is output by the output unit 58D depending on conditions will be described. In the second embodiment, the same reference numerals are given to the same components as those in the first embodiment, and the description thereof is omitted. In the second embodiment, parts different from the first embodiment will be described. Also, hereinafter, for the sake of convenience of explanation, when there is no need to distinguish between other physical camera moving images and virtual viewpoint moving images, they are referred to as "other camera moving images".

[0140] In the second embodiment, among a plurality of cameras (for example, all the cameras shown in FIG. 3), any one camera other than the reference physical camera is defined as the specific camera, and among the plurality of cameras, cameras other than the reference physical camera and the specific camera are defined as non-specific cameras. As an example of the specific camera, there is a camera used for imaging to obtain another camera image output by the output unit 58D when the process of step ST28 or step ST30 shown in FIG. 14B is executed. Here, the specific camera is an example of the "second image camera" according to the technology of the present disclosure.

[0141] Also, in the second embodiment, as detection processes, in addition to the first detection process and the second detection process described above, a third detection process and a fourth detection process are performed.

[0142] The third detection process is a process of detecting the target person image 96 from a specific camera image which is another camera image obtained by imaging with the specific camera. The specific camera image is an example of the "second image" according to the technology of the present disclosure. Also, in the third detection process, similar to the first and second detection processes, the target person image 96 is detected when a face image showing the face of the target person is detected. The other camera image to be the detection target of the face image is the specific camera image.

[0143] The types of a plurality of frames constituting other-camera moving images obtained by being imaged by a specific camera are roughly classified into a detection frame in which a target person image 96 is detected by a third detection process and a non-detection frame in which the target person image 96 is not detected by the third detection process. Hereinafter, for convenience of explanation, other-camera moving images obtained by being imaged by a specific camera are also referred to as "specific-camera moving images".

[0144] The fourth detection process is a process of detecting a target person image 96 from a non-specific-camera image which is an other-camera image obtained by being imaged by a non-specific camera. Among non-specific-camera images, a non-specific-camera image in which the target person image 96 is detected by the fourth detection process is an example of a "third image" according to the technology of the present disclosure. A non-specific camera used for imaging to obtain a non-specific-camera image in which the target person image 96 is detected by the fourth detection process is an example of a "third-image camera" according to the technology of the present disclosure. Also in the fourth detection process, similar to the first to third detection processes, the target person image 96 is detected when a face image showing the face of the target person is detected. The camera image to be the detection target of the face image is a non-specific-camera image.

[0145] In the present second embodiment, when the specific-camera moving image includes a detection frame and a non-detection frame, the CPU 58 selectively outputs the non-detection frame and the non-specific-camera image according to the distance between the position of the specific camera and the position of the non-specific camera and the time of the non-detection state described in the first embodiment.

[0146] For example, when the CPU 58 satisfies a non-detection frame output condition that the distance between the position of the specific camera and the position of the non-specific camera exceeds a threshold value and the time of the non-detection state is less than a predetermined time, the CPU 58 outputs a non-detection frame, and when the non-detection frame output condition is not satisfied, the CPU 58 outputs a non-specific-camera image instead of the non-detection frame. Hereinafter, this configuration will be described in detail.

[0147] As an example, as shown in FIG. 23, the CPU 58 of the image processing apparatus 12 according to the second embodiment further operates as a setting unit 58F, a determination unit 58G, and a calculation unit 58H, which is different from the CPU 58 of the image processing apparatus 12 described in the first embodiment.

[0148] When the other camera image that is the detection target of the second detection process or the other camera image selected by the image selection unit 58E is output by the output unit 58D, the setting unit 58F sets the camera used for imaging to obtain the other camera image output by the output unit 58D as the specific camera. Further, the setting unit 58F acquires camera identification information from the other camera image output by the output unit 58D. Then, the setting unit 58F holds the camera identification information acquired from the other camera image as specific camera identification information that can identify the specific camera.

[0149] As an example, as shown in FIG. 24, when the specific camera is set by the setting unit 58F, the image acquisition unit 58B acquires the specific camera identification information from the setting unit 58F. Then, the image acquisition unit 58B acquires a specific camera image at the same imaging time as the reference physical camera image that is the processing target of the first detection process from the specific camera moving image obtained by imaging with the specific camera specified from the specific camera identification information.

[0150] The detection unit 58C executes a third detection process using the target person image sample 98 on the specific camera image acquired by the image acquisition unit 58B in the same manner as the first and second detection processes. When the target person image 96 is detected from the specific camera image by the third detection process, the output unit 58D outputs the specific camera image including the target person image 96 detected by the third detection process to the user device 14. As a result, the specific camera image including the target person image 96 detected by the third detection process is displayed on the display 78 of the user device 14.

[0151] As an example, as shown in FIG. 25, when the target person image 96 is not detected from the specific camera image by the third detection process, the determination unit 58G determines whether or not the non-detection continuous time is less than a predetermined time (for example, 3 seconds). Here, the non-detection continuous time refers to the time in the non-detection state, that is, the time during which the non-detection state continues. The predetermined time may be a fixed time or a variable time that is changed according to the given instruction and / or condition.

[0152] As an example, as shown in FIG. 26, in a situation where the specific camera is set by the setting unit 58F, after the determination unit 58G determines whether or not the non-detection continuous time is less than the predetermined time, the image acquisition unit 58B acquires the specific camera identification information from the setting unit 58F. The image acquisition unit 58B acquires, from the image group 102, all non-specific camera images (hereinafter, also referred to as "non-specific camera image group") other than the specific camera image among a plurality of other camera images at the same imaging time as the reference physical camera image that is the processing target of the first detection process, using the specific camera identification information. Then, the detection unit 58C executes the fourth detection process using the target person image sample 98 on the non-specific camera image group acquired by the image acquisition unit 58B in the same manner as the first to third detection processes.

[0153] As an example, as shown in FIG. 27, when the target person image 96 is detected by the fourth detection process, the calculation unit 58H acquires the camera specific information attached to the non-specific camera image in which the target person image 96 is detected by the fourth detection process as non-specific camera identification information that can identify the non-specific camera used in the imaging for obtaining the non-specific camera image.

[0154] The calculation unit 58H calculates the distance between the specific camera and the unspecified camera (hereinafter, also referred to as "camera distance") using the camera installation position information regarding the specific camera specified by the specific camera identification information held by the setting unit 58F and the camera installation position information regarding the unspecified camera specified by the unspecified camera identification information. The calculation unit 58H calculates the camera distance for each piece of unspecified camera identification information, that is, for each unspecified camera image in which the target person image 96 is detected by the fourth detection process.

[0155] The determination unit 58G acquires the shortest camera distance (hereinafter, also referred to as "shortest camera distance") among the camera distances calculated by the calculation unit 58H. Then, the determination unit 58G determines whether the shortest camera distance exceeds a threshold value. The threshold value may be a fixed value or a variable value that is changed according to a given instruction and / or condition.

[0156] When it is determined by the determination unit 58G that the shortest camera distance exceeds the threshold value, the output unit 58D outputs the specific camera image acquired by the image acquisition unit 58B, that is, the specific camera image in which the target person image 96 is not detected by the third detection process, to the user device 14. Also, when the target person image 96 is not detected from the group of unspecified camera images by the fourth detection process, the output unit 58D outputs the specific camera image acquired by the image acquisition unit 58B, that is, the specific camera image in which the target person image 96 is not detected by the third detection process, to the user device 14. As a result, a specific camera image not including the target person image 96 is displayed on the display 78 of the user device 14.

[0157] FIG. 28 shows an example of the processing content of the CPU 58 when the determination unit 58G determines that the shortest camera distance is equal to or less than the threshold value, and when the non-detection duration is equal to or longer than the predetermined time and the target person image 96 is detected by the fourth detection process. In the example shown in FIG. 28, the calculation unit 58H outputs the shortest distance unspecified camera identification information to the image acquisition unit 58B and the setting unit 58F. The shortest distance unspecified camera identification information refers to unspecified camera identification information that can identify the unspecified camera that is the calculation target of the shortest camera distance calculated by the calculation unit 58H. The image acquisition unit 58 acquires the shortest distance unspecified camera image, which is an unspecified camera image obtained by being captured by the unspecified camera specified by the unspecified camera identification information, from among the unspecified camera images in which the target person image 96 is detected by the fourth detection process.

[0158] The output unit 58D outputs the shortest distance unspecified camera image acquired by the image acquisition unit 58B to the user device 14. As a result, the shortest distance unspecified camera image is displayed on the display 78 of the user device 14. Note that since the target person image 96 is included in the shortest distance unspecified camera image, the user 18 can observe the target person through the display 78.

[0159] When the output of the shortest distance unspecified camera image is completed, the output unit 58D outputs output completion information to the setting unit 58F. When the output completion information is input from the output unit 58D, the setting unit 58F sets, instead of the specific camera currently set, the unspecified camera specified from the shortest distance unspecified camera identification information input from the calculation unit 58H (hereinafter also referred to as the "shortest distance unspecified camera") as the specific camera.

[0160] Next, an example of the flow of the output control process according to the second embodiment will be described with reference to FIGS. 29A to 29C. The flowcharts shown in FIGS. 29A to 29C are different in that they have steps ST100 to ST138 as compared with the flowcharts shown in FIGS. 14A and 14B. Hereinafter, the differences from the flowcharts shown in FIGS. 14A and 14B will be described.

[0161] In step ST14 shown in FIG. 14A, if the determination is negative, the output control process proceeds to step ST100 shown in FIG. 29A. In step ST100, the detection unit 58C determines whether or not a specific camera is not set. For example, here, if the setting unit 58F does not hold the specific camera identification information, the detection unit 58C determines that the specific camera is not set, and if the setting unit 58F holds the specific camera identification information, the detection unit 58C determines that the specific camera is not not set (the specific camera is set).

[0162] In step ST100, if the specific camera is not set, the determination is affirmative and the output control process proceeds to step ST18. In step ST100, if the specific camera is not not set, the determination is negative and the output control process proceeds to step ST104 shown in FIG. 29B.

[0163] In step ST102, the setting unit 58F sets the camera used for imaging to obtain the other camera image output in step ST28 or step ST30 as the specific camera, and then the output control process proceeds to step ST32 shown in FIG. 14A.

[0164] In step ST104 shown in FIG. 29B, a specific camera image at the same imaging time as the reference physical camera image that is the processing target of the first detection process is acquired from the specific camera moving image obtained by imaging with the specific camera, and then the output control process proceeds to step ST106.

[0165] In step ST106, the detection unit 58C executes the third detection process on the specific camera image acquired in step ST104 using the target person image sample 98, and then the output control process proceeds to step ST108.

[0166] In step ST108, the detection unit 58C determines whether or not the target person image 96 has been detected from the specific camera image by the third detection process. In step ST108, if the target person image 96 has not been detected from the specific camera image by the third detection process, the determination is negative and the output control process proceeds to step ST112. In step ST108, if the target person image 96 has been detected from the specific camera image by the third detection process, the determination is positive and the output control process proceeds to step ST110.

[0167] In step ST110, the output unit 58D outputs the specific camera image that is the detection target of the third detection process to the user device 14, and then the output control process proceeds to step ST32 shown in FIG. 14A.

[0168] In step ST112, the determination unit 58G determines whether or not the non-detection duration is less than the predetermined time. In step ST112, if the non-detection duration is equal to or longer than the predetermined time, the determination is negative and the output control process proceeds to step ST128 shown in FIG. 29C. In step ST112, if the non-detection duration is less than the predetermined time, the determination is positive and the output control process proceeds to step ST114.

[0169] In step ST114, the detection unit 58C executes the fourth detection process on the non-specific camera image group using the target person image sample 98, and then the output control process proceeds to step ST116.

[0170] In step ST116, the detection unit 58C determines whether or not the target person image 96 has been detected from the non-specific camera image group by the fourth detection process. In step ST116, if the target person image 96 has not been detected from the non-specific camera image group by the fourth detection process, the determination is negative and the output control process proceeds to step ST110. In step ST116, if the target person image 96 has been detected from the non-specific camera image group by the fourth detection process, the determination is positive and the output control process proceeds to step ST118.

[0171] In step ST118, first, the calculation unit 58H acquires, as non-specific camera identification information that can identify the non-specific camera used in the imaging for obtaining the non-specific camera image, the camera identification information attached to the non-specific camera image in which the target person image 96 has been detected by the fourth detection process in step ST114. Next, the calculation unit 58H calculates the camera distance using the camera installation position information regarding the specific camera specified by the specific camera identification information held by the setting unit 58F and the camera installation position information regarding the non-specific camera specified by the non-specific camera identification information. The camera distance is calculated for each non-specific camera image in which the target person image 96 has been detected by the fourth detection process in step ST114. After the process of step ST118 is executed, the output control process proceeds to step ST120.

[0172] In step ST120, the determination unit 58G determines whether the shortest camera distance among the camera distances calculated in step ST118 exceeds a threshold value. In step ST120, if the shortest camera distance is equal to or less than the threshold value, the determination is negative, and the output control process proceeds to step ST122. In step ST120, if the shortest camera distance exceeds the threshold value, the determination is positive, and the output control process proceeds to step ST110.

[0173] In step ST122, first, the image acquisition unit 58B acquires the shortest distance non-specific camera identification information from the calculation unit 58H. Then, the image acquisition unit 58B acquires the shortest distance non-specific camera image obtained by being imaged by the non-specific camera specified from the shortest distance non-specific camera identification information from at least one frame of non-specific camera images in which the target person image 96 has been detected by the fourth detection process in step ST114. After the process of step ST122 is executed, the output control process proceeds to step ST124.

[0174] In step ST124, the output unit 58D outputs the shortest distance non-specific camera image acquired in step ST122 to the user device 14, and then the output control process proceeds to step ST126.

[0175] In step ST126, the setting unit 58F acquires the shortest distance unspecified camera identification information from the calculation unit 58H. Then, the setting unit 58F sets the shortest distance unspecified camera specified from the shortest distance unspecified camera identification information as the specified camera in place of the currently set specified camera. After that, the output control process proceeds to step ST32 shown in FIG. 14A.

[0176] In step ST128 shown in FIG. 29C, the detection unit 58C executes a fourth detection process on the unspecified camera image group using the target person image sample 98. Then, the output control process proceeds to step ST130.

[0177] In step ST130, the detection unit 58C determines whether the target person image 96 has been detected from the unspecified camera image group by the fourth detection process in step ST128. In step ST130, if the target person image 96 has not been detected from the unspecified camera image group by the fourth detection process in step ST128, the determination is negative, and the output control process proceeds to step ST110 shown in FIG. 29B. In step ST130, if the target person image 96 has been detected from the unspecified camera image group by the fourth detection process in step ST128, the determination is positive, and the output control process proceeds to step ST132.

[0178] In step ST132, first, the calculation unit 58H acquires, as non-specific camera identification information that can identify the non-specific camera used in the imaging for obtaining the non-specific camera image, the camera identification information attached to the non-specific camera image in which the target person image 96 has been detected by the fourth detection process in step ST128. Next, the calculation unit 58H calculates the camera distance using the camera installation position information regarding the specific camera specified by the specific camera identification information held by the setting unit 58F and the camera installation position information regarding the non-specific camera specified by the non-specific camera identification information. The camera distance is calculated for each non-specific camera image in which the target person image 96 has been detected by the fourth detection process in step ST128. After the process of step ST132 is executed, the output control process proceeds to step ST134.

[0179] In step ST134, first, the image acquisition unit 58B acquires the shortest distance non-specific camera identification information from the calculation unit 58H. Then, the image acquisition unit 58B acquires the shortest distance non-specific camera image obtained by being imaged by the non-specific camera specified from the shortest distance non-specific camera identification information from at least one frame of non-specific camera images in which the target person image 96 has been detected by the fourth detection process in step ST128. After the process of step ST134 is executed, the output control process proceeds to step ST136.

[0180] In step ST136, the output unit 58D outputs the shortest distance non-specific camera image acquired in step ST134 to the user device 14, and then the output control process proceeds to step ST138.

[0181] In step ST138, the setting unit 58F acquires the shortest distance non-specific camera identification information from the calculation unit 58H. Then, the setting unit 58F sets the shortest distance non-specific camera specified from the shortest distance non-specific camera identification information as the specific camera in place of the specific camera currently set, and then the output control process proceeds to step ST32 shown in FIG. 14A.

[0182] In this way, when the specific camera moving image obtained by being captured by a specific camera includes a frame containing the target person image 96 and a frame not containing the target person image 96, the output unit 58D selectively outputs a frame not containing the target person image 96 in the specific camera moving image and a non-specific camera image containing the target person image 96 according to the camera distance and the non-detection continuous time. Therefore, according to this configuration, compared with the case of always outputting a non-specific camera image containing the target person image 96 during the period when the target person image 96 is not detected, the discomfort caused to the user by the steep change of other camera images can be suppressed.

[0183] Further, when the output control process according to the second embodiment is executed, when the condition that the shortest camera distance exceeds the threshold value and the non-detection continuous time is less than the preset time is satisfied, a frame not containing the target person image 96 in the specific camera moving image is output. Also, when the condition that the shortest camera distance exceeds the threshold value and the non-detection continuous time is less than the preset time is not satisfied, instead of the frame not containing the target person image 96 in the specific camera moving image, a non-specific camera image containing the target person image 96 is output. Therefore, according to this configuration, compared with the case of always outputting a non-specific camera image containing the target person image 96 during the period when the target person image 96 is not detected, the discomfort caused to the user by the steep change of other camera images can be suppressed.

[0184] Note that in the above second embodiment, the condition that the shortest camera distance exceeds the threshold value and the non-detection continuous time is less than the preset time is exemplified, but the technology of the present disclosure is not limited thereto. For example, the condition that the shortest camera distance coincides with the threshold value and the non-detection continuous time is less than the preset time may also be used. Also, the condition that the shortest camera distance exceeds the threshold value and the non-detection continuous time reaches the preset time may also be used. Also, the condition that the shortest camera distance coincides with the threshold value and the non-detection continuous time reaches the preset time may also be used.

[0185] Also, for the image processing apparatus 12 described in the above second embodiment, various forms described in the above first embodiment can be appropriately applied.

[0186] Also, in each of the above embodiments, as an example of a multi-frame image composed of a plurality of frames, a form example in which a moving image is output by the output unit 58D to the user device 14 has been given. However, the technology of the present disclosure is not limited to this, and instead of a moving image, a series of consecutive images may be output by the output unit 58D. In this case, as an example, as shown in FIG. 30, instead of the reference physical camera moving image, a reference physical camera series of consecutive images, instead of the other physical camera moving image, other physical camera series of consecutive images, and instead of the virtual viewpoint moving image, a virtual viewpoint series of consecutive images may be stored in the storage 60 as the image group 102. Thus, even when a series of consecutive images are output to the user device 14, the user 18 can be made to continuously observe the target person.

[0187] Also, in each of the above embodiments, a form example in which a moving image is displayed on the display 78 of the user device 14 has been described. However, among the plurality of time-series camera images constituting the moving image displayed on the display 78, the camera image intended by the user 18 may be selectively displayed on the display 78 by the user 18 performing a flick operation and / or a swipe operation on the touch panel 76A.

[0188] Also, in each of the above embodiments, the soccer stadium 22 has been exemplified, but this is merely an example. Any location is acceptable as long as a plurality of physical cameras 16 can be installed, such as a baseball field, a rugby field, a curling rink, an athletics stadium, a swimming pool, a concert hall, an outdoor music venue, and a theater venue.

[0189] Also, in each of the above embodiments, the computers 50 and 70 have been exemplified, but the technology of the present disclosure is not limited to this. For example, instead of the computers 50 and / or 70, a device including an ASIC, an FPGA, and / or a PLD may be applied. Also, instead of the computers 50 and / or 70, a combination of a hardware configuration and a software configuration may be used.

[0190] In addition, in each of the above embodiments, an example has been described in which the output control process is executed by the CPU 58 of the image processing apparatus 12. However, the technology of the present disclosure is not limited to this. Some of the processes included in the output control process may be executed by the CPU 88 of the user device 14. Further, instead of the CPU 88, a GPU may be employed, or a plurality of CPUs may be employed, and various processes may be executed by one processor or a plurality of physically separated processors.

[0191] In addition, in each of the above embodiments, the output control program 100 is stored in the storage 60. However, the technology of the present disclosure is not limited to this. As an example, as shown in FIG. 29, the output control program 100 may be stored in an arbitrary portable storage medium 200. The storage medium 200 is a non-transitory storage medium. Examples of the storage medium 200 include an SSD or a USB memory. The output control program 100 stored in the storage medium 200 is installed in the computer 50, and the CPU 58 executes the output control process according to the output control program 100.

[0192] Further, the output control program 100 may be stored in the program memory of another computer or server device connected to the computer 50 via a communication network (not shown), and the output control program 100 may be downloaded to the image processing apparatus 12 in response to a request from the image processing apparatus 12. In this case, the output control process based on the downloaded output control program 100 is executed by the CPU 58 of the computer 50.

[0193] As the hardware resources for executing the output control process, the following various processors can be used. As the processor, for example, as described above, a general-purpose processor that functions as a hardware resource for executing the output control process according to software, that is, a program, such as the CPU, can be mentioned.

[0194] In addition, examples of other processors include dedicated electric circuits, which are processors having a circuit configuration specifically designed to execute specific processes such as FPGAs, PLDs, or ASICs. A memory is incorporated or connected to any of these processors, and each processor executes output control processing by using the memory.

[0195] The hardware resources for executing the output control processing may be constituted by one of these various processors, or may be constituted by a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resources for executing the output control processing may be a single processor.

[0196] As an example of configuration with a single processor, first, as represented by computers such as clients and servers, one processor is configured by a combination of one or more CPUs and software, and this processor functions as the hardware resources for executing the output control processing. Second, as represented by SoCs, etc., there is a form in which a processor that realizes the functions of the entire system including a plurality of hardware resources for executing the output control processing is used in one IC chip. Thus, the output control processing is realized as hardware resources by using one or more of the above various processors.

[0197] Furthermore, as the hardware structure of these various processors, more specifically, an electric circuit combining circuit elements such as semiconductor elements can be used.

[0198] Also, the above-described output control processing is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope not departing from the gist.

[0199] The description and illustration shown above are detailed descriptions of the part related to the technology of the present disclosure and are merely examples of the technology of the present disclosure. For example, the descriptions regarding the above configurations, functions, operations, and effects are descriptions of examples of the configurations, functions, operations, and effects of the part related to the technology of the present disclosure. Therefore, it goes without saying that within the scope not departing from the gist of the technology of the present disclosure, the description and illustration shown above may be modified by deleting unnecessary parts, adding new elements, or making replacements. Also, in order to avoid complication and facilitate the understanding of the part related to the technology of the present disclosure, the description regarding common general knowledge in technology that does not particularly require explanation for implementing the technology of the present disclosure is omitted in the description and illustration shown above.

[0200] In this specification, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0201] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually stated to be incorporated by reference.

Claims

1. A processor; A memory included in or connected to the processor; The processor, performing a detection process for detecting an object image showing the object from a plurality of images obtained by capturing an image of an imaging area including the object using a plurality of cameras positioned at different positions; outputting a first image of the plurality of images; when a transition is made from a detection state in which the object image is detected from the first image by the detection process to a non-detection state in which the object image is not detected from the first image by the detection process, outputting a second image in which the object image is detected by the detection process among the plurality of images; The transition from the detection state to the non-detection state is a transition from a state in which the object image can be detected due to the movement of an object in the imaging area to a state in which the object image cannot be detected due to the object being hidden by the object. Image processing device.

2. At least one of the first image and the second image is a virtual viewpoint image. The image processing device according to claim 1 .

3. When the processor transitions from the detection state to the non-detection state while the first image is being output, the processor switches from outputting the first image to outputting the second image.

3. The image processing device according to claim 1 or 2.

4. The image is a multi-frame image made up of a plurality of frames. The image processing device according to any one of claims 1 to 3.

5. The multi-frame image is a moving image. The image processing device according to claim 4.

6. The multiple frame images are continuous shot images. The image processing device according to claim 4.

7. The processor, outputting the multi-frame image as the second image; The output of the multiple frame images as the second image is started from a timing before the non-detection state is reached. The image processing device according to any one of claims 4 to 6.

8. The processor, outputting the multi-frame image as the second image; The output of the multiple frame images as the second image is terminated at a timing after the timing at which the non-detection state is reached. The image processing device according to any one of claims 4 to 7.

9. The processor resumes output of the first image on condition that the non-detection state is returned to the detection state. The image processing device according to any one of claims 1 to 8.

10. the plurality of cameras includes at least one virtual camera and at least one physical camera; The plurality of images include a virtual viewpoint image obtained by capturing an image of the imaging area by the virtual camera, and a captured image obtained by capturing an image of the imaging area by the physical camera. The image processing device according to any one of claims 1 to 9.

11. The processor outputs, during a period in which output of the first image is switched to output of the second image, a plurality of virtual viewpoint images obtained by capturing images using a plurality of virtual cameras that continuously connect the position, orientation, and angle of view of the camera used in capturing the first image to the position, orientation, and angle of view of the camera used in capturing the second image. The image processing device according to any one of claims 1 to 10.

12. The object is a person The image processing device according to any one of claims 1 to 11.

13. The processor detects the object image by detecting a face image showing the face of the person. The image processing device according to claim 12.

14. The processor outputs, as the second image, an image among the plurality of images in which at least one of a position and a size of the object image in the image satisfies a predetermined condition and in which the object image is detected by the detection process. The image processing device according to any one of claims 1 to 13.

15. The second image is an overhead image showing an overhead view of the imaging area. The image processing device according to any one of claims 1 to 14.

16. The first image is a television broadcast image. The image processing device according to any one of claims 1 to 15.

17. The first image is an image obtained by capturing an image using a camera among the plurality of cameras that is installed at an observation position that observes the imaging area or in the vicinity of the observation position. The image processing device according to any one of claims 1 to 16.

18. performing a detection process for detecting an object image showing the object from a plurality of images obtained by capturing an image of an imaging area including the object using a plurality of cameras positioned at different positions; outputting a first image of the plurality of images; when a transition is made from a detection state in which the object image is detected from the first image by the detection process to a non-detection state in which the object image is not detected from the first image by the detection process, outputting a second image in which the object image is detected by the detection process among the plurality of images, The transition from the detection state to the non-detection state is a transition from a state in which the object image can be detected due to the movement of an object in the imaging area to a state in which the object image cannot be detected due to the object being hidden by the object. Image processing methods.

19. A program for causing a computer to execute a process, The process comprises: performing a detection process for detecting an object image showing the object from a plurality of images obtained by capturing an image of an imaging area including the object using a plurality of cameras positioned at different positions; outputting a first image of the plurality of images; when a transition is made from a detection state in which the object image is detected from the first image by the detection process to a non-detection state in which the object image is not detected from the first image by the detection process, outputting a second image in which the object image is detected by the detection process among the plurality of images, The transition from the detection state to the non-detection state is a transition from a state in which the object image can be detected due to the movement of an object in the imaging area to a state in which the object image cannot be detected due to the object being hidden by the object. program.

Citation Information

Patent Citations

  • Image processing system, image processing method and program

    JP2018055279A

  • Image processing apparatus and control method thereof

    JP2018055644A

  • Imaging apparatus and imaging method

    JP2018148483A

  • Information processing apparatus, information processing method, and program

    JP2019012533A