Video generation device, video generation method, and video generation program

The video generation device automatically extracts and aligns areas from still images to generate moving images focused on a detection target, addressing the limitations of manual frame setting and subject movement constraints in existing technologies.

JP7747171B2Active Publication Date: 2025-10-01NEC CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024507197
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-10-01
Estimated Expiration
2042-03-14

Smart Images

  • Figure 0007747171000001
    Figure 0007747171000001
  • Figure 0007747171000002
    Figure 0007747171000002
  • Figure 0007747171000003
    Figure 0007747171000003
Patent Text Reader

Abstract

In order to easily generate a moving image focused on a predetermined detection target, a moving image generating device (1) comprises: a moving image generating unit (11) for joining, in chronological order, partial images generated by extracting a region in which the predetermined detection target appears from each of a plurality of still images in a time series, to generate a moving image having the partial images as frame images; and an adjusting unit (12) for performing adjustment to align positions in which the detection target appears, between the frames of the generated moving image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video generation device that automatically generates video images. [Background technology]

[0002] It may be difficult to observe a specific object in detail from a moving image that captures a wide range. Patent Document 1 below, for example, is an example of a technology for solving this problem. Patent Document 1 discloses an image processing device that creates a moving image centered on a specific object by setting a frame that surrounds a pre-specified object image in each of multiple frame images that make up the moving image, cutting out and enlarging the image included in the frame to the size of the frame image, and stitching the enlarged images together in the order of the frame images. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-279894 Summary of the Invention [Problem to be solved by the invention]

[0004] The image processing device described in Patent Document 1 has the problem that the user must manually set the frame, which is time-consuming. Patent Document 1 also describes detecting a portion that has changed from the previous frame and setting a rectangular frame that surrounds the detected changed portion, but this method cannot be applied to subjects other than those that are constantly moving. As such, with conventional technology, it has not been easy to generate moving images that focus on a specific detection target.

[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and one example of its purpose is to provide a video generation device, etc. that can easily generate video images focused on a specified detection target. [Means for solving the problem]

[0006] A moving image generating device according to one aspect of the present invention comprises a moving image generating means that connects in chronological order partial images generated by extracting areas in which a specified detection target appears from each of a plurality of still images in a time series, thereby generating a moving image in which the partial images are frame images, and a correction means that performs correction to align the positions in which the detection target appears between frames of the moving image.

[0007] A moving image generation method according to one aspect of the present invention includes at least one processor extracting areas in which a predetermined detection target appears from each of a plurality of still images in a time series, connecting the generated partial images in chronological order to generate a moving image in which the partial images serve as frame images, and performing correction to align the positions in which the detection target appears between frames of the moving image.

[0008] A video generation program according to one aspect of the present invention causes a computer to function as a video generation means that connects partial images generated by extracting areas in which a specified detection target appears from each of a plurality of still images in a time series in chronological order to generate a video in which the partial images are frame images, and a correction means that performs correction to align the positions in which the detection target appears between frames of the video. [Effects of the Invention]

[0009] According to one aspect of the present invention, it is possible to easily generate a moving image that focuses on a predetermined detection target. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing a configuration of a moving image generating device according to a first exemplary embodiment of the present invention. [Figure 2] 1 is a flowchart showing the flow of a moving image generating method according to a first exemplary embodiment of the present invention. [Figure 3] FIG. 10 is a diagram illustrating an overview of a video generation system according to a second exemplary embodiment of the present invention. [Figure 4]FIG. 2 is a diagram showing an outline of a method for generating moving images in the moving image generation system. [Figure 5] FIG. 10 is a block diagram showing the configuration of a moving image generating device according to a second exemplary embodiment of the present invention. [Figure 6] 10A and 10B are diagrams illustrating an example of a masking process performed by a masking unit included in the moving image generating device. [Figure 7] FIG. 10 is a flowchart showing the flow of a moving image generating method according to a second exemplary embodiment of the present invention. [Figure 8] FIG. 1 is a diagram illustrating an example of a computer that executes instructions of a program, which is software that realizes the functions of each device according to each exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0011] Exemplary Embodiment 1 A first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described below.

[0012] (Configuration of video generation device) The configuration of a moving image generating device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the moving image generating device 1. As shown in the figure, the moving image generating device 1 includes a moving image generating unit 11 and a correcting unit 12.

[0013] The moving image generator 11 extracts areas in which a predetermined detection target appears from each of a plurality of still images in time series, generates partial images by connecting the generated partial images in time series order, and generates a moving image in which the partial images serve as frame images. The corrector 12 then performs correction to align the positions in which the detection target appears between frames of the generated moving image.

[0014] As described above, the moving image generating device 1 according to this exemplary embodiment is configured to include a moving image generating unit 11 that connects in chronological order partial images generated by extracting areas in which a predetermined detection target appears from each of a plurality of still images in a time series to generate a moving image in which the partial images serve as frame images, and a correction unit 12 that performs correction to align the positions in which the detection target appears between frames of the generated moving image. Therefore, the moving image generating device 1 according to this exemplary embodiment has the effect of easily generating a moving image that focuses on the predetermined detection target.

[0015] (Video generation program) The functions of the moving image generation device 1 described above can also be realized by a program. The moving image generation program according to this exemplary embodiment causes a computer to function as a moving image generation unit that extracts areas in which a predetermined detection target appears from each of a plurality of time-series still images, generates partial images by connecting the generated partial images in chronological order, and generates moving images using the partial images as frame images, and as a correction unit that corrects the positions in which the detection target appears between frames of the generated moving images. This moving image generation program has the effect of easily generating moving images that focus on the predetermined detection target.

[0016] (Video generation method flow) The flow of the moving image generation method according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the moving image generation method. Note that the execution entity of each step in this moving image generation method may be a processor provided in the moving image generation device 1, or a processor provided in another device, or each step may be executed by a processor provided in a different device.

[0017] In S11, at least one processor extracts areas in which a predetermined detection target appears from each of a plurality of still images in a time series, generates partial images by connecting the partial images in time series order to generate a moving image having the partial images as frame images. Subsequently, in S12, at least one processor performs correction to align the positions in which the detection target appears between frames of the moving image generated in S11.

[0018] In this way, the moving image generation method according to the present exemplary embodiment includes a configuration in which at least one processor extracts areas in which a predetermined detection target appears from each of a plurality of still images in a time series, connects the generated partial images in chronological order to generate a moving image in which the partial images serve as frame images, and performs correction to align the positions in which the detection target appears between frames of the generated moving image. Therefore, the moving image generation method according to the present exemplary embodiment has the effect of easily generating a moving image that focuses on the predetermined detection target.

[0019] Exemplary Embodiment 2 A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are given the same reference numerals, and their description will be omitted as appropriate.

[0020] (System Overview) An overview of the video generation system according to this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a diagram showing an overview of the video generation system 7. The video generation system 7 is a system that automatically generates moving images focusing on each racehorse. As shown in the figure, the video generation system 7 includes a video generation device 2, a shooting device 3, an edge server 4, a terminal device 5, and a terminal device 6.

[0021] The camera 3 captures moving images of racehorses. Specifically, as shown in FIG. 3, the camera 3 captures moving images of racehorses circling the paddock. Because the position of the camera 3 is fixed, the camera 3 captures moving images of a range determined by the camera 3 (for example, the range indicated by the dashed line in FIG. 3). It is preferable that the camera 3 be capable of wide-angle photography so that multiple racehorses can be photographed at once. Furthermore, it is preferable that the camera 3 photograph the horse's body from the side, as shown in FIG. 3, in order to facilitate detection and identification of the horse's body, as will be described later.

[0022] Although only one camera device 3 is shown in FIG. 3, multiple camera devices 3 may be installed to simultaneously photograph a larger number of racehorses. Furthermore, a camera device 3 capable of ultra-wide-angle photography, such as a 360-degree camera, may also be used. However, placing camera devices 3 too close to the racehorses or placing too many camera devices 3 is undesirable as it can irritate the racehorses. For this reason, it is preferable to place camera devices 3 capable of wide-angle photography at a distance from the racehorses, as in the example of FIG. 3.

[0023] The edge server 4 acquires the moving images captured by the image capturing device 3 and transfers them to the moving image generating device 2. At this time, the edge server 4 adjusts the image quality and frame rate of the moving images before transferring them to the moving image generating device 2. For example, the edge server 4 may transfer moving images at a predetermined frame rate (e.g., 30 fps) and with image quality above a certain level to the moving image generating device 2 in real time. Note that the moving images captured by the image capturing device 3 may also be transferred directly to the moving image generating device 2 without going through the edge server 4.

[0024] The video generation device 2 uses the video received from the edge server 4 to generate a video focusing on each racehorse, and makes the generated video public so that users of the video generation system 7 can access it. The method for generating the video will be described later with reference to FIG. 4.

[0025] Terminal devices 5 and 6 are terminal devices used by users of video generation system 7. Users of video generation system 7 can use any terminal device such as terminal devices 5 and 6 to view videos generated by video generation device 2, i.e., videos focusing on each racehorse. For example, terminal device 5 shown in FIG. 3 displays a video focusing on the racehorse with bib number 2, while terminal device 6 displays a video focusing on the racehorse with bib number 1.

[0026] It should be noted that the terminal devices usable with the video generation system 7 are not limited to tablet-type terminal devices or smartphones as shown in Fig. 3. For example, it is also possible to use a terminal device such as a personal computer to view the moving images generated by the video generation system 7. Furthermore, the number of terminal devices usable with the video generation system 7 is not particularly limited.

[0027] As described above, according to the video generation system 7, videos focusing on each racehorse can be generated from the videos captured by the filming device 3, and these can be viewed by the user of the video generation system 7. Because the videos are generated automatically without any user operation, the video can be generated more easily and in a shorter time than with the image processing device described in the aforementioned Patent Document 1. Furthermore, because the video generation system 7 detects the area in which the racehorse appears, it can detect and generate a video regardless of whether the racehorse is moving or not.

[0028] Furthermore, because the video images captured by the filming device 3 are taken at a wide angle, it is difficult to determine the condition of each racehorse from these video images, but the video images focused on each racehorse generated by the video generation system 7 make it easy to determine the condition of each racehorse. Furthermore, because the video generation system 7 does not need to film each racehorse individually, it is possible to minimize the amount of filming equipment and filming personnel required, and as mentioned above, there is also the advantage that unnecessary stimulation is not given to the racehorses.

[0029] The video generation system 7 is not limited to racehorses and can generate video of any detection target. For example, the video generation system 7 can generate video focusing on individual athletes from video of multiple athletes in public sports other than horse racing (such as boat racing or bicycle racing). Other examples include video focusing on individual athletes from video of a sports match, and video focusing on individual performers from video of a concert. The video generation system 7 can also generate video focusing on specific people, vehicles, etc., appearing in video or still images captured by a camera such as a surveillance camera or a dashcam. Therefore, the term "racehorse" in the following description can be interpreted as any detection target.

[0030] (Overview of video generation method) FIG. 4 is a diagram showing an overview of a video generation method (hereinafter referred to as this method) in this exemplary embodiment. In this method, first, a time-series of still images 211 is acquired, each of which depicts a racehorse, which is a predetermined detection target. The still images 211 may be frame images extracted from a video generated by photographing the racehorse with the camera device 3. The frame images may be extracted by the edge server 4 shown in FIG. 3. Furthermore, the camera device 3 may capture time-series still images 211 instead of capturing video images, in which case the still images 211 captured by the camera device 3 may be acquired as they are.

[0031] Next, this method detects areas in the still image 211 where racehorses are captured. In Figure 4, the detected areas are indicated by dashed rectangles. That is, in the example shown, an area where the racehorse with bib number 1 is captured and an area where the racehorse with bib number 2 is captured are extracted. This process is performed for each of the still images 211 in the time series.

[0032] Next, in this method, the region detected as described above is extracted from the still image 211 to generate a partial image 215. The partial image 215 generated in this manner includes an image of the racehorse with bib number 1 and an image of the racehorse with bib number 2.

[0033] Therefore, in this method, the generated plurality of partial images 215 are classified according to the detection target that appears in the partial images 215. In the example of Fig. 4, the generated plurality of partial images 215 are classified into a partial image 2151 that shows the racehorse with bib number 1 and a partial image 2152 that shows the racehorse with bib number 2.

[0034] Then, in this method, partial images 2151 are connected in chronological order to generate a moving image using partial image 2151 as a frame image. Similarly, a moving image is generated from partial image 2152. The moving image generated here may have blurred positions of the racehorses between frames due to factors such as the accuracy of region extraction. When such blurring occurs, the moving image may appear unnatural.

[0035] Therefore, in this method, after generating a video, correction is performed to align the positions of the detected objects between frames of the video, thereby completing the video. By performing this correction, it is possible to generate a natural video that focuses on each individual racehorse.

[0036] (Configuration of video generation device) The configuration of the moving image generation device 2 according to this exemplary embodiment will be described with reference to Fig. 5. Fig. 5 is a block diagram showing the configuration of the moving image generation device 2. The moving image generation device 2 is a device that generates moving images from moving images or a plurality of still images in a time series. As shown in the figure, the moving image generation device 2 includes a control unit 20 that controls each unit of the moving image generation device 2 in an integrated manner, and a storage unit 21 that stores various data used by the moving image generation device 2. The moving image generation device 2 also includes an input unit 22 that accepts user input operations to the moving image generation device 2, and an output unit 23 that outputs data from the moving image generation device 2. The moving image generation device 2 may be a device dedicated to moving image generation, or may be a general-purpose device that can be used for other purposes.

[0037] The control unit 20 also includes a data acquisition unit 201, a detection unit (detection means) 202, a partial image generation unit (partial image generation means) 203, a masking unit (masking means) 204, an image classification unit (image classification means) 205, a moving image generation unit (moving image generation means) 206, and a correction unit (correction means) 207. The storage unit 21 stores a still image 211, a detection model 212, a face detection model 213, an individual identification model 214, a partial image 215, and a moving image 216. The masking unit 204 and the face detection model 213 will be described later in the section "Masking Processing."

[0038] The data acquisition unit 201 acquires a plurality of still images 211 in time series that are the source of a moving image, and stores them in the storage unit 21. For example, the data acquisition unit 201 may acquire a moving image from the edge server 4 shown in FIG. 3, extract frame images from the acquired moving image, and use these frame images as the still images in time series that are the source of the moving image. Note that the process of extracting frame images from a moving image may be performed by the edge server 4, and in this case, the data acquisition unit 201 may acquire the frame images received from the edge server 4 and store them in the storage unit 21 as still images 211.

[0039] The detection unit 202 detects an area in which the detection target appears from the still image 211. More specifically, the detection unit 202 detects an area in which the racehorse appears from the still image 211 using a detection model 212 constructed by machine learning using images in which the racehorse, the detection target, appears as training data.

[0040] The training data for the detection model 212 may be information indicating the area in an image in which the racehorse, the detection target, is captured (for example, information indicating the representative coordinates of the area and the width and height of the area) that corresponds as correct answer data to an image in which the racehorse is captured. The machine learning algorithm is not particularly limited, and a convolutional neural network, for example, may be applied. Note that the detection model 212 does not need to be able to identify individual detection targets. In other words, the detection model 212 may be trained to detect any racehorse.

[0041] The partial image generation unit 203 extracts the area detected by the detection unit 202 from the still image 211 to generate a partial image 215, and stores the partial image 215 in the storage unit 21. At this time, the partial image generation unit 203 may adjust the size of the area extracted from the still image 211, such as by enlarging it, or adjust the aspect ratio of the image. If a person such as an audience member is captured in the partial image 215 generated by the partial image generation unit 203, the generated partial image 215 is subjected to a masking process by the masking unit 204 before being stored in the storage unit 21.

[0042] The image classification unit 205 classifies the partial images 215 according to the detection targets that appear in the partial images 215. An individual identification model 214 is used to classify the partial images 215. The individual identification model 214 is a model for identifying the detection targets that appear in the partial images 215, i.e., each individual racehorse. The image classification unit 205 associates information indicating the classification result with the partial images 215 and stores the information in the memory unit 21.

[0043] The individual identification model 214 may be a trained model that has been machine-learned to identify the bib numbers of racehorses. Such an individual identification model 214 can be constructed by machine learning using, for example, training data as partial images 215 in which information indicating the area in which the bib is captured (e.g., information indicating the representative coordinates of the area and the width and height of the area) and the bib number are associated as correct answer data. This configuration uses the bib number affixed to each racehorse as identification information for the racehorse.

[0044] In this way, the image classification unit 205 may classify the partial image 215 by detecting, from the partial image 215, identification information attached to the detection targets in order to identify the multiple detection targets. According to this configuration, since the identification information attached to the detection targets is used, in addition to the effect achieved by the moving image generation device 1 according to the first exemplary embodiment, an effect of being able to accurately identify the multiple detection targets can be obtained.

[0045] The moving image generating unit 206 connects the partial images 215 in chronological order to generate a moving image 216 in which the partial images 215 are used as frame images. In this case, the moving image generating unit 206 connects, in chronological order, the partial images 215 classified into the same classification by the image classifying unit 205 to generate the moving image 216.

[0046] The correction unit 207 performs correction to align the positions at which the detection target appears between frames of the moving image 216 generated by the moving image generation unit 206. Then, the correction unit 207 stores the corrected moving image 216 in the storage unit 21. The method of correction is not particularly limited. For example, the correction unit 207 may perform the above correction using an algorithm for camera shake correction of moving images. This makes it possible to easily perform correction to align the positions at which the detection target appears between frames of the moving image 216.

[0047] As described above, the moving image generating device 2 includes a moving image generating unit 206 that connects in chronological order the partial images 215 generated by extracting areas in which a predetermined detection target appears from each of a plurality of still images 211 in a time series, to generate a moving image 216 in which the partial images 215 are frame images, and a correction unit 207 that performs correction to align the positions in which the detection target appears between frames of the generated moving image 216.

[0048] According to this configuration, partial images 215 generated by extracting an area in which a predetermined detection target appears from each still image 211 are used, so that the detection target can be detected and a moving image 216 can be generated regardless of whether the detection target is moving or not.

[0049] However, the positions of the detection targets appearing in the partial images 215 are not necessarily aligned. If the moving image 216 is generated from partial images 215 in which the positions of the detection targets are not aligned, the positions of the detection targets will be shifted between frames, resulting in a moving image that is difficult to view. Therefore, with the above configuration, after the moving image 216 is generated once, correction is performed to align the positions of the detection targets appearing between frames of the generated moving image 216. This makes it possible to automatically generate the moving image 216 in which the positions of the detection targets are aligned between frames from the partial images 215 in which the positions of the detection targets are not aligned. Therefore, the moving image generation device 2 has the effect of easily generating the moving image 216 that focuses on a predetermined detection target.

[0050] As described above, the moving image generation device 2 includes the detection unit 202 that detects an area in which the detection target appears from the still image 211 using the detection model 212 constructed by machine learning using an image in which the detection target appears as training data, and the partial image generation unit 203 that extracts the area detected by the detection unit 202 from the still image 211 to generate the partial image 215. This provides the effect of being able to automatically generate the partial image 215 from the still image 211, in addition to the effect provided by the moving image generation device 1 according to the first exemplary embodiment.

[0051] As described above, the moving image generating device 2 includes the image classification unit 205 that classifies the partial images 215 according to the detection targets appearing in the partial images 215, and the moving image generating unit 206 generates the moving image 216 by connecting in chronological order the partial images 215 that the image classification unit 205 has classified into the same classification. This provides the effect of being able to automatically generate the moving image 216 that focuses on each detection target from the still image 211 that shows multiple detection targets, in addition to the effect provided by the moving image generating device 1 according to the first exemplary embodiment.

[0052] (Regarding detection target identification) As described above, the individual identification model 214 can be used to identify the detection target appearing in the partial image 215. As described above, racehorses are fitted with bibs bearing identification information called bib numbers, and therefore the individual identification model 214 may be machine-trained to identify these bib numbers.

[0053] However, some numbers, such as 1 and 7, have similar appearances, and numbers can be difficult to read depending on the lighting conditions, angle of view, etc., so the classification results from the individual identification model 214 are not necessarily correct. For this reason, the image classification unit 205 may determine where to classify the partial image 215 by taking into account information related to the identification of the detected object in addition to the output value of the individual identification model 214.

[0054] For example, because racehorses generally enter the paddock in a fixed order, the image classification unit 205 may verify whether the identification result identified from the output value of the individual identification model 214 is valid based on the order in which the racehorses enter or the time at which the racehorses were photographed. As a result of this verification, partial images 215 for which the identification result is determined to be invalid may be excluded from being made into moving images. Furthermore, such partial images 215 may be presented to the user of the video generation device 2, for example by being output to the output unit 23, so that the user can decide on the correct classification destination.

[0055] Alternatively, a simpler identification method may be adopted. For example, the image classification unit 205 may identify racehorses based on the order in which they enter the paddock. For example, the image classification unit 205 may identify the racehorse that appears first in the still image 211 as the first racehorse. In this case, the image classification unit 205 may identify the racehorse that appears first in the time-series still images 211 taken until the racehorse leaves the field of view of the imaging device 3 as the first racehorse. The image classification unit 205 may also identify the racehorse that appears next after the first racehorse as the second racehorse. In the same manner, the image classification unit 205 can identify all the racehorses up to the last racehorse.

[0056] In the paddock, excited racehorses may become violent and make it difficult to identify their bib numbers. For this reason, the image classification unit 205 may analyze the partial images 215 to detect such unusual movements of the racehorses. The image classification unit 205 may then exclude the partial images 215 in which unusual movements are detected from the targets for creation of moving images, or may allow the user to decide the correct classification destination.

[0057] In addition to the above, it is also possible to identify racehorses based on, for example, their coat color, jockeys, etc. Furthermore, if the detection target is not a racehorse but, for example, a person, the image classification unit 205 may identify the person by recognizing their face.

[0058] (About masking process) The masking unit 204 and the face detection model 213 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of masking processing by the masking unit 204. Fig. 6 shows a partial image 215A before the masking processing is performed and a partial image 215B after the masking processing.

[0059] Masking unit 204 detects areas in partial image 215 where people's faces appear, and performs masking on the detected areas. The masking process is a process that makes people unrecognizable, and may be, for example, mosaic processing or blurring. In the example of Fig. 6, masking unit 204 detects areas in partial image 215A where the faces of people A and B appear, and performs blurring on those areas to generate partial image 215B.

[0060] The video generating device 2 according to this exemplary embodiment is equipped with a masking unit 204, and thus in addition to the effects of the video generating device 1 according to exemplary embodiment 1, it can automatically generate a video 216 that takes into consideration the privacy and portrait rights of people who appear in the video.

[0061] A face detection model 213 is used to detect an area in which a person's face is captured. The face detection model 213 may be constructed by machine learning using, as training data, a partial image 215 associated with information indicating an area in which a person's face is captured (e.g., information indicating representative coordinates of the area and the width and height of the area) as correct answer data. The machine learning algorithm is not particularly limited, and, for example, a convolutional neural network or the like may be applied.

[0062] 6 also shows a jockey guiding a racehorse in addition to persons A and B, but the jockey's face has not been blurred. In this way, masking unit 204 may not mask the face of a specific person appearing in partial image 215B, but may mask the faces of other people.

[0063] In an image of a paddock, as shown in partial images 215A and 215B in Figure 6, the faces of spectators A and B are captured from the front, while the jockey's profile is often captured. For this reason, by using face detection model 213 constructed by machine learning using training data in which faces captured from the front are used as correct answer data, it is possible to detect the faces of persons A and B but not the face of the jockey. Therefore, by using such face detection model 213, it is possible to automatically mask the faces of spectators but not the faces of the jockey, thereby preventing a decrease in the visibility of the jockey and the horse.

[0064] Alternatively, for example, a discrimination model constructed by machine learning to be able to distinguish between spectators and jockeys may be used. In this case, the masking unit 204 may perform masking processing only on the facial regions of spectators and jockeys identified using the discrimination model.

[0065] Furthermore, in video images of the paddock captured from a fixed position, adjusting the capture position in advance allows spectators and jockeys to appear in different areas. For example, in partial images 215A and 215B of FIG. 6, spectators A and B appear in a strip-shaped area (spectator seating area) at the top of the image, and the jockey appears in a lower area. Therefore, masking unit 204 may perform masking processing on face areas detected in the strip-shaped area (spectator seating area) at the top of partial image 215, but may not perform masking processing on face areas detected in other areas. Alternatively, masking unit 204 may perform face detection processing only on the strip-shaped area (spectator seating area) at the top of partial image 215.

[0066] (Processing flow) The flow of the process (moving image generation method) executed by moving image generation device 2 will be described with reference to Fig. 7. Fig. 7 is a flow diagram showing the flow of the moving image generation method executed by moving image generation device 2. The following process may be performed in parallel with the shooting of moving images of racehorses going around the paddock by shooting device 3 (see Fig. 3).

[0067] In S21, the data acquisition unit 201 acquires a predetermined number of time-series still images 211. For example, the data acquisition unit 201 may acquire, from the edge server 4, moving images of racehorses going around a paddock, which are captured by the imaging device 3, and acquire frame images constituting the moving images as the still images 211.

[0068] In S22, the detection unit 202 detects an area in which a horse's body appears from each still image 211 acquired in S21. Specifically, the detection unit 202 detects an area in which a horse's body appears in each still image 211 based on an output value obtained by inputting each still image 211 acquired in S21 to the detection model 212.

[0069] In S23, the partial image generating unit 203 generates a partial image 215 by extracting the area detected in S22 from each still image 211 acquired in S21.

[0070] In S24, the masking unit 204 detects an area in each partial image 215 generated in S23 where a human face appears, and performs masking processing on the detected area. Specifically, the masking unit 204 detects an area in the partial image 215 where a human face appears, based on an output value obtained by inputting the partial image 215 generated in S23 to the face detection model 213, and performs masking processing on the area. Note that the processing of S24 may be performed on the still image 211 acquired in S21. In this case, the processing of S24 is performed after S21 and before S23.

[0071] In S25, the image classification unit 205 classifies the partial images 215 generated in S23 and subjected to the masking process in S24 into each individual racehorse appearing in the partial images 215. Specifically, the image classification unit 205 classifies the partial images 215 based on the output value obtained by inputting the partial images 215 into the individual identification model 214.

[0072] In S26, the moving image generating unit 206 connects the partial images 215 classified into the same category in S25 in chronological order to generate a moving image 216 in which the partial images 215 are used as frame images.

[0073] In S27, the correction unit 207 performs correction to align the positions of the racehorses, which are the detection targets, between frames of the video 216 generated in S26. This completes the video 216 that focuses on each racehorse. The completed video 216 may be made available online so that it can be viewed from terminal devices used by users of the video generation system 7, such as terminal devices 5 and 6 shown in FIG. 3.

[0074] When the above processing is performed in parallel with the shooting of video images of racehorses going around the paddock, new video images (more precisely, frame images constituting the video images) are continuously received from the edge server 4. For this reason, the video generation device 2 may perform the above-mentioned processing of S21 to S27 every time a new video image is received, thereby updating the previously generated video image 216.

[0075] [Modification] The execution entity of each process described in the above embodiment is arbitrary and is not limited to the above example. In other words, the functions of the video generation device 2 can be replaced by multiple devices (which can also be called processors) that can communicate with each other. For example, by distributing each block shown in FIG. 5 among multiple devices, a system having the same functions as the video generation device 2 can be constructed.

[0076] Alternatively, partial images 215 may be classified according to the detection targets captured in the partial images 215, and then the partial images 215 for each classification may be made into a moving image by a separate device. This allows moving images focused on each detection target to be generated by parallel processing using a plurality of devices, making it possible to generate moving images focused on each detection target in a short time.

[0077] [Software implementation example] Some or all of the functions of the video generation devices 1 and 2 may be realized by hardware such as an integrated circuit (IC chip), or by software.

[0078] In the latter case, moving image generation devices 1 and 2 are realized, for example, by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 8. Computer C includes at least one processor C1 and at least one memory C2. Memory C2 stores program P for operating computer C as moving image generation devices 1 and 2. In computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of moving image generation devices 1 and 2.

[0079] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0080] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.

[0081] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0082] [Appendix 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.

[0083] [Appendix 2] Some or all of the above-described embodiments can also be described as follows: However, the present invention is not limited to the following described aspects.

[0084] (Appendix 1) A video generation device comprising: a video generation means for generating a video in which partial images generated by extracting areas in which a specified detection target appears from each of a plurality of still images in a time series are connected in chronological order to generate a video in which the partial images serve as frame images; and a correction means for performing correction to align the positions in which the detection target appears between frames of the video.

[0085] (Appendix 2) 2. The moving image generating device according to claim 1, comprising: a detection means for detecting an area in which the detection target appears from the still image using a detection model constructed by machine learning using an image in which the detection target appears as training data; and a partial image generating means for extracting the area detected by the detection means from the still image and generating the partial image.

[0086] (Appendix 3) 3. The video generation device according to claim 1, further comprising an image classification means for classifying the partial images according to the detection target appearing in the partial images, wherein the video generation means generates a video by connecting the partial images classified into the same classification by the image classification means in chronological order.

[0087] (Appendix 4) The video generation device according to claim 3, wherein the image classification means classifies the partial images by detecting identification information attached to the detection targets in order to identify the plurality of detection targets from the partial images.

[0088] (Appendix 5) 5. The video generating device according to any one of claims 1 to 4, further comprising a masking means for detecting an area in the partial image in which a person's face appears and masking the detected area to make the person unidentifiable.

[0089] (Appendix 6) A method for generating a moving image, the method comprising: at least one processor extracting an area in which a predetermined detection target appears from each of a plurality of still images in a time series, connecting the generated partial images in chronological order to generate a moving image in which the partial images serve as frame images; and performing correction to align the positions in which the detection target appears between frames of the moving image.

[0090] (Appendix 7) A program for causing a computer to operate as the moving image generating device according to any one of Supplementary Notes 1 to 5, the moving image generating program causing the computer to function as each of the means.

[0091] [Appendix 3] Some or all of the above-described embodiments can also be expressed as follows: A moving image generating device including at least one processor that executes a process of generating a moving image having the partial images as frame images by extracting areas in which a predetermined detection target appears from each of a plurality of still images in a time series, and connecting the partial images in chronological order to generate the moving image, and a process of performing a correction process to align the positions in which the detection target appears between frames of the moving image.

[0092] The moving image generating device may further include a memory that stores a program for causing the processor to execute the process of generating the moving image and the process of performing the correction. The program may also be recorded on a computer-readable, non-transitory, tangible recording medium. [Explanation of symbols]

[0093] 1. Video generation device 11 Video generation unit 12 Correction unit 2. Video generation device 202 Detection unit 203 Partial image generation unit 204 Masking Section 205 Image Classification Unit 206 Video Generation Unit 207 Correction Unit

Claims

1. a moving image generating means for generating a moving image using the partial images as frame images by connecting the partial images generated by extracting areas in which a predetermined detection target appears from each of a plurality of still images in time series in time series; a correction means for performing correction to align the positions at which the detection target appears between frames of the moving image; a masking means for detecting an area in the partial image where a person's face is captured and masking the detected area; Equipped with The masking means is A face detection model constructed by machine learning using training data with faces photographed from the front as the correct answer data is used to mask faces photographed from the front and not mask faces photographed from the side, or A moving image generating device that performs masking processing on a face area detected in a band-shaped area at the top end of the partial image, and does not perform masking processing on face areas detected in other areas.

2. a detection means for detecting an area in which the detection target appears from the still image by using a detection model constructed by machine learning using an image in which the detection target appears as training data; 2. The moving image generating device according to claim 1, further comprising: partial image generating means for extracting the region detected by the detecting means from the still image to generate the partial image.

3. an image classification means for classifying the partial images according to the detection target appearing in the partial images; 3. The moving image generating device according to claim 1, wherein the moving image generating means generates a moving image by connecting, in chronological order, the partial images classified into the same classification by the image classifying means.

4. The moving image generating device according to claim 3 , wherein the image classification means classifies the partial images by detecting, from the partial images, identification information attached to the detection targets in order to identify the plurality of detection targets.

5. At least one processor extracting areas in which a predetermined detection target appears from each of a plurality of still images in time series, and connecting the generated partial images in time series order to generate a moving image in which the partial images are used as frame images; performing a correction to align the positions at which the detection target appears between frames of the moving image; detecting an area in the partial image in which a person's face is captured, and performing a masking process on the detected area; Including, The masking treatment is A face detection model constructed by machine learning using training data with faces photographed from the front as the correct answer data is used to mask faces photographed from the front and not mask faces photographed from the side, or A moving image generating method in which a face area detected in a strip-shaped area at the top end of the partial image is masked, and face areas detected in other areas are not masked.

6. Computer, a moving image generating means for generating a moving image in which partial images are connected in chronological order by extracting areas in which a predetermined detection target appears from each of a plurality of still images in time series, and using the partial images as frame images; a correction means for performing correction to align the positions at which the detection target appears between frames of the moving image; a masking means for detecting an area in the partial image where a person's face is captured and masking the detected area; It functions as The masking means is A face detection model constructed by machine learning using training data with faces photographed from the front as the correct answer data is used to mask faces photographed from the front and not mask faces photographed from the side, or A video generation program that performs masking processing on a face area detected in a strip-shaped area at the top end of the partial image, and does not perform masking processing on face areas detected in other areas.

Citation Information

Patent Citations

  • Photographing system and method for supplying image distribution

    JP2004363775A

  • Display device and method, and program

    JP2006099058A

  • Image processing apparatus, image processing method, and program

    JP2006279894A

  • Imaging apparatus, data structure of image file

    JP2009033738A

  • System, method, and program for providing images

    JP2018156453A