Image processing system and image processing method
The image processing system efficiently generates highlight videos from multiple camera angles by detecting and synchronizing scene occurrences, addressing the need for individual scene analysis in each feed.
Patent Information
- Application Number
- JP2024066325
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies struggle to efficiently create highlight videos of specific scenes from multiple camera angles without requiring multiple cameramen and individual analysis of each video.
An image processing system that acquires and processes multiple camera feeds, detects the occurrence area of a predetermined scene, and generates synchronized scene images from different angles using positional relationship information between cameras.
Facilitates easy output of highlight videos from multiple angles without needing individual scene analysis in each camera feed, ensuring consistent scene identification across multiple camera views.
Smart Images

Figure 2025162859000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing system and an image processing method. [Background technology]
[0002] Patent Document 1 discloses an information processing device. This information processing device performs a first control process to select an analysis engine for scene detection from a plurality of analysis engines based on scene detection information for detecting scenes in an input video. The information processing device also performs a second control process to select an analysis engine from the plurality of analysis engines for obtaining second result information related to a scene based on scene-related information about the scene obtained as first result information by the analysis engine selected in the first control process. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2021 / 241430 Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides an image processing system and the like that can easily output video of a predetermined scene at an event captured from a plurality of different angles. [Means for solving the problem]
[0005] An image processing system according to one aspect of the present disclosure includes an acquisition unit, a detection unit, a generation unit, and an output unit. The acquisition unit acquires multiple image data obtained by capturing images of a space where an event is to be held using multiple cameras from different angles. The detection unit detects an occurrence area of a predetermined scene in the event from first image data among the multiple image data. The generation unit generates a first scene image including the occurrence area of the predetermined scene in the first image data, and one or more second scene images, each of which is captured in the same time period as the occurrence of the predetermined scene and includes the occurrence area of the predetermined scene, in one or more second image data other than the first image data among the multiple image data. The output unit outputs a video including the first scene image and the one or more second scene images. The generation unit identifies the occurrence area of the predetermined scene in each of the one or more second image data based on information indicating the relative positional relationship between the multiple cameras. [Effects of the Invention]
[0006] The present disclosure has an advantage in that it is easy to output video of a predetermined scene at an event captured from a plurality of different angles. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a block diagram showing an overall configuration including an image processing system according to an embodiment. [Figure 2] FIG. 2 is an explanatory diagram of an example of use of the image processing system according to the embodiment. [Figure 3] FIG. 3 is a diagram showing an example of an original image and a cut-out image. [Figure 4] FIG. 4 is an explanatory diagram of an example of generation of a second scene image by the image processing system according to the embodiment. [Figure 5] FIG. 5 is an explanatory diagram of information indicating the relative positional relationships of a plurality of cameras in the image processing system according to the embodiment. [Figure 6]FIG. 6 is a flowchart showing an example of the operation of image processing according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] [1. Findings that formed the basis of this disclosure] First, the inventor's point of view will be explained below.
[0009] For example, in highlight footage of sports such as soccer matches, which have few scoring scenes, it is common to use footage of a single scoring scene captured from various angles. Conventionally, creating such footage requires multiple cameras capable of capturing images from different angles and multiple cameramen to operate each of the cameras.
[0010] For example, when filming a relatively large number of sports matches at a low cost, such as university sports, it is possible to reduce the number of cameramen by remotely operating one or more of the multiple cameras. In such cases, when creating highlight footage of the filmed sports, it is possible to achieve automation and labor savings by using technology such as that disclosed in Patent Document 1.
[0011] However, the technology disclosed in Patent Document 1 only considers creating a highlight video from a single video using an analysis engine. Therefore, when using the technology disclosed in Patent Document 1 to create a video of a specific scene at an event captured from multiple different angles, there is a problem in that the multiple videos captured by multiple cameras must each be analyzed individually using an analysis engine. In this case, there is also a problem in that the specific scene identified by analysis in each of the multiple videos may not be the same.
[0012] In view of the above, the inventors have come up with the present disclosure.
[0013] Hereinafter, embodiments will be described with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection forms, steps, step order, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components not described in independent claims will be described as optional components.
[0014] It should be noted that the drawings are schematic diagrams and are not necessarily strict illustrations. In addition, in the drawings, substantially the same components are denoted by the same reference numerals, and overlapping descriptions may be omitted or simplified.
[0015] Furthermore, in this specification, ordinal numbers such as "first" and "second" do not refer to the number or order of components unless otherwise specified, but are used for the purpose of avoiding confusion and distinguishing between components of the same type.
[0016] (Embodiment) [2. Configuration] The overall configuration including an image processing system 1 according to an embodiment will be described below. FIG. 1 is a block diagram showing the overall configuration including an image processing system 1 according to an embodiment. The image processing system 1 is a system that acquires a plurality of image data obtained by capturing an image of a space Sp1 (see FIG. 2) where an event is held from different angles using a plurality of cameras 2, and generates and outputs a video (here, a highlight video) from the acquired plurality of image data. In the embodiment, the plurality of image data obtained by capturing an image using the plurality of cameras 2 is transmitted to the image processing system 1 via a switcher 3.
[0017] Here, the event may include, for example, a sports game such as soccer, baseball, or American football, as well as a concert or a play. In the following description, the event is assumed to be a soccer game. Therefore, in the following description, the space Sp1 where the event is held is a ballpark where the soccer game is held.
[0018] FIG. 2 is an explanatory diagram of an example of use of the image processing system 1 according to the embodiment. In the example shown in FIG. 2, the multiple cameras 2 are a total of three cameras: a first camera 2A, a second camera 2B, and a third camera 2C. The first camera 2A is a camera that images the stadium from the side. The second camera 2B is a camera that images the stadium from behind one of the two goals. The third camera 2C is a camera that images the stadium from behind the other of the two goals. In this way, the multiple cameras 2 image the space Sp1 from different angles. Note that the stands, goals, etc. are not shown in FIG. 2.
[0019] Returning to FIG. 1 , camera 2 is an imaging device capable of capturing moving images or still images. In the embodiment, camera 2 is capable of capturing moving images or still images with a relatively high resolution, such as 6K or 8K. In the embodiment, camera 2 is installed in a manner fixed to a base, wall, ceiling, or the like at a video production site or a live broadcast site, for example. Camera 2 includes optical system 21, imaging unit 22, image processing unit 23, and transmission unit 24.
[0020] The optical system 21 includes, for example, a focus lens and a zoom lens. The focus lens is made up of a combination of one or more lenses. The zoom lens is made up of a combination of one or more lenses.
[0021] The imaging unit 22 generates image data from optical information input via the optical system 21. The imaging unit 22 has an image sensor that converts the optical information into an electrical signal (analog signal) and an A / D converter that converts the analog signal into a digital signal (image signal). The image sensor is, for example, a CCD (Charge-Coupled Device) image sensor or a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor.
[0022] The image processing unit 23 generates image data by performing appropriate processing on the image signal converted by the imaging unit 22. For example, the image processing unit 23 generates image data by encoding the image signal in a predetermined format. Metadata, which is information indicating the imaging conditions at the time of imaging, is added to the image data. The imaging conditions may include, for example, the focal length, exposure, shutter speed, or angle of view of the camera 2. The imaging conditions may also include, for example, the start time and end time of imaging by the camera 2.
[0023] Here, as shown in Fig. 3, the image data includes an original image P0 and a cut-out image P1. The original image P0 is an image captured by the imaging unit 22. The cut-out image P1 is an image obtained by cutting out a partial area from the original image P0. Fig. 3 is a diagram showing an example of the original image P0 and the cut-out image P1. In the example shown in Fig. 3, the original image P0 is an image captured by the imaging unit 22 of the first camera 2A. Furthermore, in the example shown in Fig. 3, the cut-out image P1 is an image obtained by cutting out an area of the original image P0 that includes a predetermined object (here, a soccer ball and a player kicking the soccer ball).
[0024] The image processing unit 23 recognizes a predetermined object from the original image P0 by executing an appropriate image recognition algorithm, and then generates a cut-out image P1 by cutting out an area including the recognized predetermined object from the original image P0.
[0025] Note that the image processing unit 23 may recognize a predetermined object from the original image P0 using a recognition model trained by machine learning to output a predetermined object in response to the input of the original image P0. The predetermined object recognized by the image processing unit 23 may change over time. For example, the image processing unit 23 may recognize a soccer ball and a player kicking the soccer ball as the predetermined object at one point in time, and may recognize spectators in a stadium as the predetermined object at another point in time.
[0026] 1, the transmitter 24 is a communication interface for communicating with the switcher 3 via a network such as the Internet or a LAN (Local Area Network), SDI (Serial Digital Interface) transmission, or ST2110 IP (Internet Protocol) transmission. The communication between the transmitter 24 and the switcher 3 may be wired communication or wireless communication. The transmitter 24 transmits the image data generated by the image processing unit 23 to the switcher 3 via a network, SDI transmission, ST2110 IP transmission, or the like.
[0027] The switcher 3 is a device for switching between videos to be broadcast or distributed. In the embodiment, the switcher 3 is realized by installing dedicated software for the switcher 3 in, for example, a server device or a general-purpose information terminal such as a desktop or laptop personal computer. The switcher 3 may also be realized by dedicated hardware. The switcher 3 includes a processing unit 31, a display unit 32, an image processing unit 33, a transmitting unit 34, and a receiving unit 35.
[0028] The processing unit 31 executes a process of determining a program image from among a plurality of cut-out images P1 included in a plurality of image data received by the receiving unit 35. Here, the program image refers to a cut-out image P1 obtained by cutting out a predetermined area from an original image P0 captured by one of the plurality of cameras 2, and is an image to be broadcast or distributed.
[0029] In the embodiment, the processing unit 31 determines a program image by receiving an input for selecting a program image from an operator operating the switcher 3. The processing unit 31 also transmits the program image from the transmitting unit 34 to the image processing system 1, and distributes or broadcasts the program image. Therefore, as the operator selects program images sequentially, the program images are broadcast or distributed sequentially.
[0030] The processing unit 31 may automatically determine a program image from among the plurality of cut-out images P1 using an appropriate determination algorithm. Alternatively, for example, the processing unit 31 may determine a program image from the plurality of cut-out images P1 using a determination model that has been trained by machine learning so as to output a program image in response to input of the plurality of cut-out images P1.
[0031] The display unit 32 is, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and displays the original image P0 and the multiple cut-out images P1 received by the receiving unit 35. The display unit 32 may display the original image P0 and the multiple cut-out images P1 all at once, or may switch between displaying the original image P0 and the multiple cut-out images P1 in response to an operation input by an operator. If the processing unit 31 automatically executes the process of determining the program image, the switcher 3 does not need to be equipped with the display unit 32.
[0032] The image processing unit 33 decodes the image data received by the receiving unit 35. The image (original image P0 or multiple cut-out images P1) obtained by decoding the image data is provided to the processing unit 31. The image obtained by decoding the image data is displayed on the display unit 32.
[0033] The transmitter 34 is a communication interface for communicating with the image processing system 1 via a network such as the Internet or a LAN, SDI transmission, or IP transmission of ST2110. The communication between the transmitter 34 and the image processing system 1 may be wired communication or wireless communication. The transmitter 34 transmits multiple pieces of image data (including image data of the program image determined by the processor 31) received from each of the multiple cameras 2 to the image processing system 1 via a network, SDI transmission, or IP transmission of ST2110.
[0034] The receiving unit 35 is a communication interface for communicating with each of the multiple cameras 2 via a network such as the Internet or a LAN, SDI transmission, or ST2110 IP transmission. The communication between the receiving unit 35 and each of the multiple cameras 2 may be wired communication or wireless communication. The receiving unit 35 receives multiple pieces of image data transmitted from each of the multiple cameras 2 via a network, SDI transmission, ST2110 IP transmission, or the like.
[0035] The image processing system 1 is a recorder that stores a plurality of image data transmitted from each of a plurality of cameras 2 via a switcher 3, and generates and outputs video based on the plurality of image data. In the embodiment, the video generated and output by the image processing system 1 is a highlight video edited from characteristic images of a broadcasted or distributed program image (program video).
[0036] The image processing system 1 is realized by installing dedicated software for the image processing system 1 on a general-purpose information terminal such as a server device or a desktop or laptop personal computer. The image processing system 1 may be realized by dedicated hardware or by cloud computing. The image processing system 1 includes an acquisition unit 11, a processing unit 12, an output unit 13, and a storage unit 14.
[0037] The acquisition unit 11 is a communication interface for communicating with the switcher 3 via a network such as the Internet or a LAN, SDI transmission, or ST2110 IP transmission. The communication between the acquisition unit 11 and the switcher 3 may be wired communication or wireless communication. The acquisition unit 11 receives multiple image data transmitted from the switcher 3 via a network, SDI transmission, or ST2110 IP transmission. In other words, the acquisition unit 11 acquires multiple image data obtained by capturing images of the space Sp1 where the event will be held from different angles using multiple cameras 2.
[0038] The processing unit 12 performs various processes on the image data acquired by the acquisition unit 11, such as decoding the image data acquired by the acquisition unit 11. In the embodiment, the processing unit 12 uses a detection unit 121 and a generation unit 122 when generating the above-mentioned highlight video. Both the detection unit 121 and the generation unit 122 are functions that the processing unit 12 can execute.
[0039] The detection unit 121 detects an occurrence area of a predetermined scene in an event from first image data among the plurality of image data (in other words, image data obtained by capturing an image by any one of the plurality of cameras 2). The first image data is, for example, image data of a program image.
[0040] The detection unit 121 detects an occurrence area of a predetermined scene in an event from the first image data using an appropriate detection algorithm. For example, the detection unit 121 detects an occurrence area of the predetermined scene by detecting a characteristic object included in the predetermined scene from the first image data. Specifically, if the event is a soccer match, the detection unit 121 detects a goal (occurrence area) in the scoring scene (predetermined scene) by detecting a goal and a soccer ball inside the goal, which are characteristic objects included in the scoring scene, from the first image data. Note that the detection unit 121 may detect an occurrence area of the predetermined scene using a detection model trained by machine learning so as to output an occurrence area of the predetermined scene in an event in response to input of the first image data.
[0041] The generation unit 122 generates a first scene image P2 (see FIG. 4) and one or more second scene images P3 (see FIG. 4). The first scene image P2 is an image including an occurrence area of a predetermined scene in the first image data. Specifically, when the event is a soccer match, the generation unit 122 generates, as the first scene image P2, an image including a goal (occurrence area) of a scoring scene (predetermined scene) detected by the detection unit 121 in the image data of the program image. The first scene image P2 is, for example, an image obtained by cutting out a predetermined area including the occurrence area from the image data of the program image.
[0042] The second scene image P3 is an image that is taken in the same time period as the time period in which a predetermined scene occurred in one or more pieces of second image data other than the first image data among the plurality of image data (in other words, image data obtained by capturing an image by one or more cameras 2 other than the camera 2 that captured the first image data among the plurality of cameras 2), and that includes an occurrence area of the predetermined scene. Specifically, when the event is a soccer match, the generation unit 122 generates, as the second scene image P3, an image that includes the goal (occurrence area) of the scoring scene (predetermined scene) detected by the detection unit 121 in the image data of the image captured by a camera 2 other than the camera 2 that captured the program image. Note that the second scene image P3 is, for example, an image obtained by cutting out a predetermined area that includes the occurrence area from the image data of the image captured by a camera 2 other than the camera 2 that captured the program image.
[0043] In the embodiment, the generation unit 122 can identify a time period in which a predetermined scene occurred in each of the one or more sets of second image data, based on the metadata attached to each of the plurality of image data. Specifically, when generating the first scene image P2, the generation unit 122 generates time information including the start time and end time of capturing the first scene image P2 by referring to the metadata attached to the image data (first image data) including the first scene image P2. Then, the generation unit 122 can identify a time period in which a predetermined scene occurred in each of the one or more sets of second image data by referring to the time information and the metadata attached to each of the one or more sets of second image data.
[0044] Furthermore, in the embodiment, the generation unit 122 identifies an occurrence area of a predetermined scene in each of the one or more second image data based on information indicating the relative positional relationship between the multiple cameras 2 (hereinafter simply referred to as "positional relationship information"). Specifically, when generating the first scene image P2, the generation unit 122 generates position information indicating the position of the occurrence area of the predetermined scene detected by the detection unit 121 in the program image. Then, the generation unit 122 can identify the occurrence area of the predetermined scene in each of the one or more second image data by referring to the position information and the positional relationship information. Note that the positional relationship information will be described in detail later.
[0045] Fig. 4 is an explanatory diagram of an example of generation of a second scene image P3 by the image processing system 1 according to the embodiment. Fig. 4(a) shows an example of a first scene image P2 generated from image data (first image data) captured by the first camera 2A. Fig. 4(b) shows a second scene image P31 generated from image data (second image data) captured by the third camera 2C. Fig. 4(c) shows a second scene image P32 generated from image data (second image data) captured by the second camera 2B.
[0046] Specifically, the first scene image P2 shown in (a) of Fig. 4 is an image including a goal (occurrence area) in a goal-scoring scene (predetermined scene) in a soccer match (event), generated from an image captured by the first camera 2A. The first scene image P2 is an image capturing the goal from an angle diagonally forward. Note that in the example shown in Fig. 4, players including the goalkeeper and stand seats are not shown.
[0047] Further, the second scene image P31 shown in FIG. 4(b) is an image generated from an image captured by the third camera 2C, captured in the same time period as the scoring scene, and including the same goal as the goal included in the first scene image P2. The second scene image P31 is an image capturing the goal from the front. Further, the second scene image P32 shown in FIG. 4(c) is an image generated from an image captured by the second camera 2B, captured in the same time period as the scoring scene, and including the same goal as the goal included in the first scene image P2. The second scene image P32 is an image capturing the goal from behind. In this way, the first scene image P2 and one or more second scene images P3 (here, two second scene images P3) are images capturing the same area where a specific scene occurs from different angles.
[0048] As already described, the generation unit 122 identifies an occurrence area of a predetermined scene in each of one or more sets of second image data based on information (positional relationship information) indicating the relative positional relationship between the multiple cameras 2. Fig. 5 is an explanatory diagram of information (positional relationship information) indicating the relative positional relationship between the multiple cameras 2 in the image processing system 1 according to the embodiment.
[0049] Fig. 5(a) is an explanatory diagram of a lookup table as positional relationship information. In the example shown in Fig. 5(a), the image on the left represents an image captured by one of the two cameras 2, and the image on the right represents an image captured by the other camera 2.
[0050] 5(a), the generation unit 122 generates in advance a lookup table in which the coordinates of a plurality of grid points in the image on the left correspond to the coordinates of a plurality of grid points in the image on the right. Then, the generation unit 122 calculates, for example, the coordinates of each grid point indicating an occurrence area of a predetermined scene in image data (first image data) of an image captured by one camera 2 by referring to the lookup table, and calculates the coordinates of each grid point indicating an occurrence area of a predetermined scene in image data (second image data) of an image captured by the other camera 2. This enables the generation unit 122 to identify an occurrence area of the predetermined scene in the second image data.
[0051] Fig. 5(b) is an explanatory diagram of a projective transformation matrix as positional relationship information. In the example shown in Fig. 5(b), similar to the example shown in Fig. 5(a), the image on the left represents an image captured by one of the two cameras 2, and the image on the right represents an image captured by the other camera 2.
[0052] 5(b), the generation unit 122 obtains a projective transformation matrix by selecting from the right-hand image the same feature points (four feature points in this example) as those in the left-hand image. Then, the generation unit 122 uses the projective transformation matrix to project coordinates indicating an occurrence area of a predetermined scene in image data (first image data) of an image captured by one of the cameras 2, for example, to calculate coordinates indicating an occurrence area of the predetermined scene in image data (second image data) of an image captured by the other camera 2. This enables the generation unit 122 to identify an occurrence area of the predetermined scene in the second image data.
[0053] Returning to FIG. 1, the output unit 13 outputs a video including the first scene image P2 and one or more second scene images P3 generated by the generation unit 122. As already mentioned, in the embodiment, the output unit 13 outputs a highlight video edited from characteristic images of the broadcasted or distributed program images (program video). Here, the first scene image P2 and the one or more second scene images P3 are both images including an occurrence area of a predetermined scene in the event, and therefore correspond to characteristic images. Furthermore, in the embodiment, the output unit 13 outputs program images other than the program image extracted from the first scene image P2, including them in the highlight video.
[0054] The storage unit 14 is a semiconductor memory such as an SSD (Solid State Drive) or a flash memory, and is a non-volatile storage device. The storage unit 14 stores a plurality of image data (including image data of program images). The storage unit 14 also stores various information necessary for processing by the processing unit 12, such as the above-mentioned time information, position information, and positional relationship information. The storage unit 14 may be a storage device other than a semiconductor memory, such as an HDD (Hard Disk Drive).
[0055] [3. Operation] The operation of the image processing system 1 according to the embodiment, that is, the image processing method according to the embodiment, will be described below. Fig. 6 is a flowchart showing an example of the operation of the image processing system 1 according to the embodiment.
[0056] First, the image processing system 1 acquires a plurality of image data transmitted from a plurality of cameras 2 via the switcher 3 (S1). Although not shown here, the image processing system 1 stores the acquired plurality of image data in the storage unit 14. Furthermore, when the image processing system 1 acquires image data of a program image from the switcher 3, it stores the image data of the program image in the storage unit 14.
[0057] Next, the image processing system 1 detects an area where the predetermined scene occurs (S2). In the embodiment, as already described, the image processing system 1 detects an area where the predetermined scene occurs by detecting a characteristic object included in the predetermined scene from the first image data, which is image data of a program image, for example.
[0058] Next, the image processing system 1 generates a first scene image P2 (S3). In the embodiment, as already described, the image processing system 1 generates the first scene image P2 by cutting out a predetermined area including an occurrence area of a predetermined scene detected by the detection unit 121 from the image data of the program image, for example.
[0059] Next, the image processing system 1 generates one or more second scene images P3 (S4). In the embodiment, as already described above, the image processing system 1 generates the second scene images P3 by cutting out a predetermined area including an occurrence area of the predetermined scene detected by the detection unit 121 from image data of an image captured by a camera 2 other than the camera 2 that captured the program image, for example.
[0060] Then, the image processing system 1 outputs the highlight video (S5). In the embodiment, as already described, the image processing system 1 outputs a video including the first scene image P2 generated by the generation unit 122, one or more second scene images P3, and a program image other than the program image extracted from the first scene image P2, as the highlight video. The highlight video is broadcast or distributed, for example, in the same way as the program video.
[0061] The above steps S1 to S5 may be executed in real time while the program image is being broadcast or distributed, or may be executed after the broadcast or distribution of the program image has ended.
[0062] [4. Advantages, etc.] The advantages of the image processing system 1 (image processing method) according to the embodiment will be described below. As described above, the image processing system 1 according to the embodiment identifies an occurrence area of a predetermined scene in each of one or more second image data based on information (positional relationship information) indicating the relative positional relationship between the multiple cameras 2. Therefore, if the occurrence area of the predetermined scene can be detected in the first image data, the image processing system 1 according to the embodiment can identify an occurrence area of the predetermined scene in each of the one or more second image data by referring to the positional relationship information. Therefore, the image processing system 1 according to the embodiment does not need to perform a process of individually detecting an occurrence area of the predetermined scene in each of the multiple image data, and therefore has the advantage of easily outputting video of a predetermined scene in an event captured from multiple different angles.
[0063] In other words, the image processing system 1 according to the embodiment does not have the problem of having to analyze multiple videos captured by multiple cameras individually using an analysis engine, as described in [1. Knowledge forming the basis of the present disclosure]. Furthermore, the image processing system 1 according to the embodiment is less likely to have the problem that the predetermined scene identified by analysis in each of the multiple videos may not be the same.
[0064] [5. Other embodiments] Although the embodiments have been described above, the present disclosure is not limited to the above-described embodiments.
[0065] For example, in the above embodiment, the image processing system 1 is realized by a device such as a recorder separate from the switcher 3, but this is not limiting. For example, the image processing system 1 may be mounted on the switcher 3.
[0066] For example, in the above embodiment, the image processing system 1 is realized by a single device, but this is not limiting, and the image processing system 1 may be realized by a plurality of devices.
[0067] In the above-described embodiment, the processing performed by a specific processing unit may be performed by another processing unit. The order of multiple processing operations may be changed, or multiple processing operations may be performed in parallel.
[0068] In the above-described embodiments, each component may be realized by executing a software program suitable for that component, or by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0069] Furthermore, each component may be realized by hardware. Each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0070] Furthermore, the general or specific aspects of the present disclosure may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.
[0071] The present disclosure may also be realized as an image processing method executed by a computer such as the image processing system of the above-described embodiment. The present disclosure may also be realized as a program (computer program product) for causing a computer to execute such an image processing method, or as a computer-readable non-transitory recording medium on which such a program is recorded.
[0072] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that a person skilled in the art would think of, or forms realized by arbitrarily combining the components and functions of each embodiment within the scope that does not deviate from the intent of this disclosure.
[0073] (summary) As described above, the image processing system 1 according to the first aspect includes an acquisition unit 11, a detection unit 121, a generation unit 122, and an output unit 13. The acquisition unit 11 acquires a plurality of image data obtained by capturing images of a space Sp1 where an event is to be held using a plurality of cameras 2 from different angles. The detection unit 121 detects an occurrence area of a predetermined scene in the event from first image data among the plurality of image data. The generation unit 122 generates a first scene image P2 and one or more second scene images P3. The first scene image P2 is an image including an occurrence area of the predetermined scene in the first image data. The one or more second scene images P3 are images that are captured in the same time period as the occurrence time period of the predetermined scene in each of one or more second image data other than the first image data among the plurality of image data and include an occurrence area of the predetermined scene. The output unit 13 outputs a video including the first scene image P2 and one or more second scene images P3. The generating unit 122 identifies an occurrence area of a predetermined scene in each of one or more second image data based on information indicating the relative positional relationship between the multiple cameras 2 (positional relationship information).
[0074] In such an image processing system 1, if an occurrence area of a predetermined scene can be detected in the first image data, the occurrence area of the predetermined scene in each of one or more second image data can be identified by referring to the positional relationship information. Therefore, such an image processing system 1 has the advantage that it is not necessary to individually detect an occurrence area of a predetermined scene in each of the multiple image data, and therefore it is easy to output video of a predetermined scene in an event captured from multiple different angles.
[0075] Also, for example, in the image processing system 1 according to the second aspect, in the first aspect, metadata indicating the imaging conditions at the time of imaging by each of the plurality of cameras 2 is added to the plurality of image data. The generation unit 122 identifies the time period in which a predetermined scene occurred in each of the one or more second image data, based on the metadata added to each of the plurality of image data.
[0076] Such an image processing system 1 has the advantage that it is not necessary to perform a process of individually detecting the time period in which a specific scene occurred for each of multiple image data, making it easy to output footage of a specific scene at an event captured from multiple different angles.
[0077] Also, for example, in the image processing system 1 according to the third aspect, in the first or second aspect, the video further includes a program image, which is an image to be distributed or broadcast, and is a cut-out image P1 obtained by cutting out a predetermined area from an original image P0 captured by one of the multiple cameras 2.
[0078] Such an image processing system 1 has the advantage that it is easy to output a highlight video in which characteristic images from among the program images (program images) that have been broadcast or distributed are edited.
[0079] Also, for example, the image processing system 1 according to a fourth aspect is any one of the first to third aspects, and further includes a plurality of cameras 2 that capture images of the space Sp1 from different angles. Each of the plurality of cameras 2 transmits image data obtained by capturing the image of the space Sp1.
[0080] In such an image processing system 1, if an occurrence area of a predetermined scene can be detected in the first image data, the occurrence area of the predetermined scene in each of one or more second image data can be identified by referring to the positional relationship information. Therefore, such an image processing system 1 has the advantage that it is not necessary to individually detect an occurrence area of a predetermined scene in each of the multiple image data, and therefore it is easy to output video of a predetermined scene in an event captured from multiple different angles.
[0081] Furthermore, for example, in an image processing method according to a fifth aspect, a space Sp1 where an event is to be held is imaged by a plurality of cameras 2 from different angles, and a plurality of image data are acquired (S1). Furthermore, in this image processing method, an occurrence area of a predetermined scene in the event is detected from first image data among the plurality of image data (S2). Furthermore, in this image processing method, a first scene image P2 and one or more second scene images P3 are generated (S3, S4). The first scene image P2 is an image including an occurrence area of the predetermined scene in the first image data. The one or more second scene images P3 are images that are captured in the same time period as the occurrence of the predetermined scene in each of one or more second image data other than the first image data among the plurality of image data, and include an occurrence area of the predetermined scene. Furthermore, in this image processing method, a video including the first scene image P2 and the one or more second scene images P3 is output (S5). Furthermore, in this image processing method, an occurrence area of the predetermined scene in each of the one or more second image data is identified based on information (positional relationship information) indicating the relative positional relationship between the plurality of cameras 2.
[0082] In such an image processing method, if the occurrence area of a predetermined scene can be detected in the first image data, the occurrence area of the predetermined scene in each of one or more second image data can be identified by referring to the positional relationship information. Therefore, this image processing method has the advantage that it is not necessary to individually detect the occurrence area of the predetermined scene in each of the multiple image data, and therefore it is easy to output video of a predetermined scene in an event captured from multiple different angles. [Industrial Applicability]
[0083] The image processing system and the like of the present disclosure can be used in a system that generates highlight videos from videos to be distributed or broadcast. [Explanation of symbols]
[0084] 1. Image processing system 11 Acquisition Department 12 Processing section 121 Detector 122 Generation part 13 Output section 14 Storage section 2 Cameras 2A 1st Camera 2B Second Camera 2C 3rd camera 21 Optical system 22 Imaging unit 23 Image processing section 24 Transmitter 3 Switcher 31 Processing section 32 Display section 33 Image processing section 34 Transmitter 35 Receiving unit P0 Original image P1 Cutout image P2 1st scene image P3, P31, P32 2nd scene image Sp1 space
Claims
1. an acquisition unit that acquires a plurality of image data obtained by capturing images of a space where an event is to be held from different angles using a plurality of cameras; a detection unit that detects an occurrence area of a predetermined scene in the event from first image data of the plurality of image data; a generating unit that generates a first scene image including the occurrence area of the predetermined scene in the first image data, and one or more second scene images in each of one or more second image data other than the first image data among the plurality of image data, the second scene images being in the same time period as the time period in which the predetermined scene occurred and including the occurrence area of the predetermined scene; an output unit that outputs a video including the first scene image and the one or more second scene images; the generation unit identifies the occurrence area of the predetermined scene in each of the one or more second image data based on information indicating a relative positional relationship between the plurality of cameras; Image processing system.
2. metadata indicating imaging conditions at the time of imaging by the plurality of cameras is added to each of the plurality of image data; the generation unit identifies a time period in which the predetermined scene occurred in each of the one or more second image data based on the metadata added to each of the plurality of image data. The image processing system according to claim 1 .
3. the video is a cut-out image obtained by cutting out a predetermined area from an original image captured by one of the plurality of cameras, and further includes a program image that is an image to be distributed or broadcast; 3. The image processing system according to claim 1.
4. The plurality of cameras each capture an image of the space from a different angle, Each of the plurality of cameras transmits image data obtained by capturing an image of the space.
3. The image processing system according to claim 1.
5. Acquire multiple pieces of image data by capturing images of the space where the event will be held from different angles using multiple cameras, detecting an occurrence area of a predetermined scene in the event from first image data among the plurality of image data; generating a first scene image including the occurrence area of the predetermined scene in the first image data, and one or more second scene images in each of one or more second image data other than the first image data among the plurality of image data, the second scene images being in the same time period as the time period in which the predetermined scene occurred and including the occurrence area of the predetermined scene; outputting a video including the first scene image and the one or more second scene images; identifying the occurrence area of the predetermined scene in each of the one or more second image data based on information indicating a relative positional relationship between the plurality of cameras; Image processing methods.
Citation Information
Patent Citations
Information processing device, information processing method, and program
WO2021241430A1