Information processing device, information processing method, and program
The information processing device uses image analysis to efficiently extract and compile notable scenes from sports and performance videos, addressing the inefficiency in creating highlight videos by automating scene selection and reducing processing time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-07
- Publication Date
- 2026-04-01
AI Technical Summary
Existing technologies are inefficient in creating highlight videos due to the time-consuming process of selecting notable scenes from videos of athletes and performers.
An information processing device and method that uses image analysis technology to extract and identify scenes of interest from a video based on specified times, outputting information on the location of these scenes for efficient compilation into highlight videos.
Solves the inefficiency in creating highlight videos by automating the scene selection process, reducing processing time and load, and enabling easy compilation of desired scenes.
Smart Images

Figure 0007838590000001 
Figure 0007838590000002 
Figure 0007838590000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Technologies related to the present invention are disclosed in Patent Documents 1 to 3 and Non-Patent Document 1.
[0003] Patent Document 1 discloses a technique for detecting the line of sight of referees, recorders, etc. in sports competitions and calculating the position to be photographed by a camera for competitive shooting based on the detection result.
[0004] Patent Document 2 discloses a technique for generating a free viewpoint video using multi-viewpoint videos taken from different viewpoints of the same scene.
[0005] Patent Document 3 discloses a technique for calculating the feature amount of each of a plurality of key points of a human body included in an image, and searching for an image including a human body with a similar posture or a human body with a similar movement based on the calculated feature amount, or classifying those with similar postures or movements together.
[0006] Patent Document 4 discloses a technique for detecting the attention state of spectators, determining the shooting position based on the detection result, and flying a drone to the determined shooting position for shooting.
[0007] Patent Document 5 discloses a technique for extracting a notable scene from a moving image based on the posture of a player.
[0008] Patent Document 6 discloses a technique for generating data indicating the content and result of a competition based on the movement information of a video taken of a competition with movement.
[0009] Non-Patent Document 1 discloses a technique related to human skeleton estimation.
Prior Art Documents
Patent Documents
[0010] [Patent Document 1] Japanese Patent Publication No. 2008-5208 [Patent Document 2] International Publication No. 2018 / 030206 [Patent Document 3] International Publication No. 2021 / 084677 [Patent Document 4] Japanese Patent Publication No. 2019-193209 [Patent Document 5] Japanese Patent Publication No. 2021-141434 [Patent Document 6] Japanese Patent Publication No. 11-339009 [Non-patent literature]
[0011] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 [Overview of the project] [Problems that the invention aims to solve]
[0012] Highlight videos are created by extracting and compiling notable scenes from videos of athletes and other performers, and then providing them to viewers. However, a problem with creating these highlight videos is that the process of selecting the notable scenes to extract is time-consuming and inefficient.
[0013] As described above, the technologies described in Patent Documents 1 and 4 are for assisting with filming, and do not assist in the process of creating highlight videos. The technology described in Patent Document 2 processes filmed footage to generate new footage, but as described above, it generates free-viewpoint footage from multi-viewpoint footage and does not create highlight videos. The technology described in Patent Document 3 is for searching for images containing human bodies with similar postures or movements, and for classifying those with similar postures or movements together, and does not describe the creation of highlight videos. The technology described in Patent Document 5 is for extracting scenes of interest, but it has the problem that if the amount of data in the video to be processed is large, the time required for computer processing will be large. The technology described in Patent Document 6 is for generating data showing the content and results of competitions, and does not create highlight videos. The technology described in Non-Patent Document 1 is a technology related to the estimation of a person's skeleton, and does not describe the creation of highlight videos.
[0014] The technologies described in Patent Documents 1 to 6 and Non-Patent Document 1 alone had the problem of not being able to solve the aforementioned problem of the workability of creating highlight videos.
[0015] One example of the object of the present invention is to provide an information processing device, an information processing method, and a program that solve the problems of the workability of creating highlight videos, in view of the above-mentioned problems. [Means for solving the problem]
[0016] According to one aspect of the present invention, An extraction means that uses image analysis technology to extract a scene of interest from a portion of a first video of the player that is identified based on a specified time, An output means that outputs information indicating the location of the scene of interest within the first video, An information processing device having the following is provided.
[0017] According to one aspect of the present invention, Computers Extract a highlight scene from a portion specified based on a specified time in a first video capturing a player, using image analysis technology, Output information indicating the position of the highlight scene in the first video, An information processing method is provided.
[0018] According to one aspect of the present invention, A computer is An extraction means for extracting a highlight scene from a portion specified based on a specified time in a first video capturing a player, using image analysis technology, An output means for outputting information indicating the position of the highlight scene in the first video, A program that causes the computer to function as such is provided.
Advantages of the Invention
[0019] According to one aspect of the present invention, the problem of workability in creating highlight videos is solved.
Brief Description of the Drawings
[0020] The above-described objects, as well as other objects, features, and advantages, will become more apparent from the public embodiments described below and the accompanying drawings.
[0021] [Figure 1] It is a diagram showing an example of a functional block diagram of an information processing apparatus. [Figure 2] It is a diagram showing an example of a hardware configuration of an information processing apparatus. [Figure 3] It is a diagram for explaining the processing of a processing unit. [Figure 4] It is a diagram schematically showing an example of information output by an information processing apparatus. [Figure 5] It is a flowchart showing an example of a processing flow of an information processing apparatus. [Figure 6] It is a diagram schematically showing another example of information output by an information processing apparatus. [Figure 7]This diagram schematically illustrates another example of the information output by an information processing device. [Figure 8] This diagram schematically illustrates another example of the information output by an information processing device. [Modes for carrying out the invention]
[0022] Embodiments of the present invention will be described below with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.
[0023] <First Embodiment> Figure 1 is a functional block diagram showing an overview of the information processing device 10 according to the first embodiment. The information processing device 10 comprises an extraction unit 11 and an output unit 12.
[0024] The extraction unit 11 extracts a scene of interest from a portion of a first video, which is a recording of a player in a sport or other performance, based on a specified time, using image analysis technology. The output unit 12 outputs information indicating the location of the scene of interest within the first video.
[0025] According to the information processing device 10 with this configuration, the problem of the workability of creating highlight videos is solved.
[0026] <Second Embodiment> "overview" The information processing device 10 of this embodiment is a more concrete example of the information processing device 10 of the first embodiment.
[0027] The information processing device 10 of this embodiment uses image analysis technology to assist in the process of creating a highlight video by extracting and compiling noteworthy scenes from a first video of a player in sports or other performances. Examples of image analysis technologies used by the information processing device 10 include, but are not limited to, face recognition, human figure recognition, posture recognition, motion recognition, appearance attribute recognition, gradient feature detection of images, color feature detection of images, object recognition, and character recognition.
[0028] "Hardware configuration" Next, an example of the hardware configuration of the information processing device 10 will be described. Each functional unit of the information processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into memory, a storage unit such as a hard disk that stores those programs (which can store programs that are pre-installed at the time of shipment, as well as programs downloaded from storage media such as CDs (Compact Discs) or from servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device.
[0029] Figure 2 is a block diagram illustrating the hardware configuration of the information processing device 10. As shown in Figure 2, the information processing device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The information processing device 10 does not necessarily have peripheral circuitry 4A. The information processing device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0030] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input devices include, for example, keyboards, mice, microphones, physical buttons, touch panels, etc. Output devices include, for example, displays, speakers, printers, mailers, etc. Processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0031] "Functional Configuration" Next, the functional configuration of the information processing device 10 of this embodiment will be described in detail. Figure 1 shows an example of a functional block diagram of the information processing device 10 of this embodiment. As shown in the figure, the information processing device 10 has an extraction unit 11 and an output unit 12.
[0032] The extraction unit 11 uses image analysis technology to extract notable scenes from the first video of the player. The output unit 12 then outputs information indicating the location of the notable scenes (the notable scenes extracted by the extraction unit 11) within the first video.
[0033] A "player" is someone who participates in sports or other performances. Performances include, but are not limited to, singing, music, dancing, drama, theater, talk shows, etc.
[0034] The "first video" is the source video for the highlight video. In other words, the highlight video is created from the first video.
[0035] "Featured scenes" are scenes that are candidates for inclusion in the highlight video. For example, an operator can decide which scenes to include in the highlight video from among the extracted featured scenes. The operator can recognize the extracted featured scenes based on the "information indicating the location of the featured scenes in the first video" output by the information processing device 10. The extraction unit 11 extracts featured scenes from the first video using image analysis technology.
[0036] Next, the process of extracting scenes of interest using image analysis technology will be described. In this embodiment, as shown in Figure 3, an image analysis system 20 is provided that analyzes images and outputs the analysis results. The image analysis system 20 may be part of the information processing device 10, or it may be an external device that is physically and / or logically independent from the information processing device 10. The extraction unit 11 uses the image analysis system 20 to extract scenes of interest from the first video.
[0037] The image analysis system 20 is described below. The image analysis system 20 includes at least one of the following functions: face recognition, human figure recognition, posture recognition, motion recognition, appearance attribute recognition, image gradient feature detection, image color feature detection, object recognition, and character recognition.
[0038] The face recognition function extracts facial features from a person. Furthermore, the similarity between facial features may be compared and calculated (e.g., to determine if they are the same person). Alternatively, the extracted facial features may be compared with the facial features of multiple players pre-registered in the database to identify which player the person in the image belongs to. In addition, the extracted facial features may be compared with the facial features of target players pre-registered in the database to detect the target player from the first video. There may be one or more target players. Note that the comparison of the extracted facial features with the facial features pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0039] The human figure recognition function extracts the physical characteristics of a person (for example, overall characteristics such as body shape, height, and clothing). Furthermore, the similarity between the physical characteristics may be compared and calculated (e.g., to determine if they are the same person). Alternatively, the extracted physical characteristics may be compared with the physical characteristics of multiple human players pre-registered in the database to identify which player the person in the image belongs to. In addition, the extracted physical characteristics may be compared with the physical characteristics of the target player pre-registered in the database to detect the target player from the first video. The target player may be one person or multiple people. Note that the comparison of the extracted physical characteristics with the physical characteristics pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0040] The posture recognition function and motion recognition function detect the joint points of a person and connect the joint points to construct a stick figure model. Then, using the information from this stick figure model, the height of the person is estimated, posture features are extracted, and movement is identified based on changes in posture. Furthermore, the similarity between posture features and movement features may be compared and calculated (e.g., determining whether the posture or movement is the same). In addition, the estimated height may be compared with the heights of multiple players pre-registered in the database to identify which player the person in the image is. Alternatively, the estimated height may be compared with the heights of target players pre-registered in the database to detect the target player from the first video. There may be one or more target players. Note that the comparison of the estimated height with the heights pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0041] The posture recognition function and the motion recognition function may be implemented using the technologies disclosed in Patent Document 3 and Non-Patent Document 1 mentioned above.
[0042] The appearance attribute recognition function recognizes appearance attributes associated with a person (for example, clothing color, shoe color, hairstyle, wearing of hats and ties, etc., there are a total of 100 or more appearance attributes). Furthermore, the similarity of the recognized appearance attributes may be compared and calculated (it is possible to determine if they are the same attribute). Alternatively, the recognized appearance attributes may be compared with the appearance attributes of multiple players pre-registered in the database to identify which player the person in the image is. In addition, the recognized appearance attributes may be compared with the appearance attributes of the target player pre-registered in the database to detect the target player from the first video. There may be one or more target players. Note that the comparison of the recognized appearance attributes with the appearance attributes pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0043] Image gradient feature detection functions include SIFT, SURF, RIFF, ORB, BRISK, CARD, and HOG. These functions detect the gradient features of each frame image. For example, the gradient features of the detected image may be compared with the gradient features of the target image pre-registered in the database to detect the target image (scene) from the first video. The comparison between the gradient features of the detected image and the gradient features of the target image pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0044] The image color feature detection function generates data that indicates the color characteristics of the image, such as a color histogram. According to this function, the color characteristics of each frame image are detected. For example, the detected image color characteristics may be compared with the color characteristics of target images that are pre-registered in the database to detect the target image (scene) from the first video. Note that the comparison of the detected image color characteristics with the color characteristics of target images that are pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead of the image analysis system 20.
[0045] Object recognition functionality is implemented using engines such as YOLO (which can extract general objects [e.g., equipment and facilities used in sports and other performances], as well as people). By using object recognition functionality, objects can be detected from images.
[0046] The character recognition function recognizes numbers, letters, etc. It may also identify which player a person in the image is by comparing the recognized numbers in the area containing a person with the numbers (jersey numbers, etc.) of multiple players pre-registered in a database. Alternatively, it may detect a target player from the first video by comparing the recognized numbers in the area containing a person with the numbers (jersey numbers, etc.) of a target player pre-registered in a database. The target player may be one person or multiple players. The comparison of the recognized numbers in the area containing a person with the numbers (jersey numbers, etc.) of a target player pre-registered in the database may be performed by the image analysis system 20, or by the extraction unit 11 instead.
[0047] As shown in Figure 3, the extraction unit 11 inputs the first video to the image analysis system 20. The extraction unit 11 then acquires the analysis results of the first video output from the image analysis system 20.
[0048] When using the face recognition function, among the analysis results output from the image analysis system 20, • Facial features extracted from the first video, and information indicating the location within the first video of the scene from which each facial feature was extracted. • Information indicating the players detected in the first video, and information indicating the location of each player in the scene within the first video. • Information indicating the location of the player being detected in the first video scene. It includes at least one of the following.
[0049] The position of a scene within the first video is indicated, for example, by the elapsed time from the beginning of the first video. The same applies to subsequent videos.
[0050] When using the human figure recognition function, among the analysis results output from the image analysis system 20, • Human physical characteristics extracted from the first video, and information indicating the location within the first video of the scene from which each human physical characteristic was extracted. • Information indicating the players detected in the first video, and information indicating the location of each player in the scene within the first video. • Information indicating the location of the player being detected in the first video scene. It includes at least one of the following.
[0051] When using the posture recognition function and / or motion recognition function, among the analysis results output from the image analysis system 20, • Information indicating posture and / or movement detected from the first video, and information indicating the position of each posture and / or movement in the first video scene. • Information indicating the players detected in the first video, and information indicating the location of each player in the scene within the first video. • Information indicating the location of the player being detected in the first video scene. It includes at least one of the following.
[0052] When using the appearance attribute recognition function, among the analysis results output from the image analysis system 20, • Information indicating the visual attributes detected from the first video, and information indicating the location within the first video of the scene in which each visual attribute was detected. • Information indicating the players detected in the first video, and information indicating the location of each player in the scene within the first video. • Information indicating the location of the player being detected in the first video scene. It includes at least one of the following.
[0053] When using the gradient feature detection function for images, among the analysis results output from the image analysis system 20, • Gradient features of each frame image, • Information indicating the position within the first video of a scene that has similar gradient features to the image (scene) being detected. It includes at least one of the following.
[0054] When using the image color feature detection function, among the analysis results output from the image analysis system 20, • Color characteristics of each frame image, • Information indicating the location within the first video of a scene that has similar color features to the image (scene) being detected. It includes at least one of the following.
[0055] When the object recognition function is used, the analysis results output from the image analysis system 20 include information indicating the position of the object to be detected in the first video scene.
[0056] When using the character recognition function, among the analysis results output from the image analysis system 20, • Information indicating the players detected in the first video, and information indicating the location of each player in the scene within the first video. • Information indicating the location of the player being detected in the first video scene. It includes at least one of the following.
[0057] The extraction unit 11 extracts noteworthy scenes from the first video based on the analysis results output from the image analysis system 20 as described above.
[0058] The highlight is, • Scenes showing players in a specific posture. • Scenes showing players performing specific movements. • Scenes in which a specific player is shown. • Scenes in which a specific player is shown in a specific posture. • Scenes in which a specific player is shown performing a specific movement. It is at least one of the following.
[0059] A predetermined posture, predetermined movement, and predetermined player are registered in advance. For example, the image analysis system 20 may be configured to detect only the predetermined posture, predetermined movement, and predetermined player. Alternatively, the image analysis system 20 may be configured to detect not only the predetermined posture, predetermined movement, and predetermined player, but also other postures, movements, and players. The extraction unit 11 may then extract the predetermined posture, predetermined movement, and predetermined player from among the postures, movements, and players detected by the image analysis system 20.
[0060] For example, by designating popular or noteworthy players as "specified players," scenes featuring these players can be extracted as "noteworthy scenes." Similarly, by designating victory poses, good plays, and other postures and movements as "specified postures and movements," scenes of victory poses or good plays can be extracted as noteworthy scenes. The noteworthy scenes described above are just examples; other scenes may also be designated as noteworthy scenes.
[0061] The aforementioned "scenes showing a specified subject (posture, movement, player, etc.)" may consist only of frame images that show the specified subject, or it may consist of a frame image showing the specified subject and a predetermined number of frame images before and after it. The "determined number of frame images before and after the frame image showing the specified subject" do not need to show the specified subject. In this case, for example, a play before a victory pose can be included as a scene of interest. One scene consists of at least two consecutive frame images.
[0062] After the above process extracts the scenes of interest from the first video, the output unit 12 outputs information indicating the location of the scenes of interest within the first video. This information is indicated, for example, by the elapsed time from the beginning of the first video. If multiple scenes of interest are extracted, the output unit 12 outputs information indicating the location of each of the multiple scenes of interest within the first video.
[0063] Figure 4 schematically shows an example of the information output by the output unit 12. The information shown in Figure 4 includes the file name, serial number, location of the scene of interest, and reason for extraction. Note that it is sufficient for at least the location of the scene of interest to be shown, and other information does not need to be displayed.
[0064] "File name" is the file name of the first video.
[0065] The "serial number" is a number used to identify multiple selected scenes of interest from one another.
[0066] The "location of the scene of interest" indicates the position of the extracted scene of interest within the first video. In the example shown in Figure 4, the location of the scene of interest is indicated by the elapsed time from the beginning of the first video.
[0067] The "Reason for Selection" indicates the reason why a particular scene was selected as a noteworthy scene. For example, the reason might be a specific player, a specific posture, or a specific movement shown in each noteworthy scene.
[0068] For example, when the system receives user input to select one of several featured scenes listed as shown in Figure 4, the output unit 12 may start playback of the first video from the beginning of the selected featured scene. The output unit 12 can use information indicating the position of each featured scene within the first video to enable playback from the beginning of the selected featured scene.
[0069] Next, an example of the processing flow of the information processing device 10 will be explained using the flowchart in Figure 5.
[0070] In S10, the information processing device 10 acquires a first video of the player. For example, the information processing device 10 acquires a first video input by the worker, or it acquires a video specified by the worker as the first video from among multiple video files stored in an accessible storage device.
[0071] In S11, the information processing device 10 extracts noteworthy scenes from the first video using image analysis technology. For example, after inputting the first video into the image analysis system 20, the information processing device 10 obtains the analysis results of the first video output from the image analysis system 20. Then, based on these analysis results, the information processing device 10 extracts noteworthy scenes from the first video.
[0072] A scene of interest is at least one of the following: a scene showing a player in a specific posture, a scene showing a player in a specific movement, a scene showing a specific player, a scene showing a specific player in a specific posture, or a scene showing a specific player in a specific movement.
[0073] In S12, the information processing device 10 outputs information indicating the position of the scene of interest extracted in S11 within the first video. The information processing device 10 outputs information such as that shown in Figure 4.
[0074] "Effects and Effects" The information processing device 10 of this embodiment uses image analysis technology to extract noteworthy scenes from a first video of the player and outputs information indicating the location of the noteworthy scenes within the first video. The person creating the highlight video can select the scenes to include in the highlight video from among the noteworthy scenes.
[0075] Furthermore, the information processing device 10 of this embodiment can analyze the first video using at least one of the following functions: face recognition, human figure recognition, posture recognition, motion recognition, appearance attribute recognition, image gradient feature detection, image color feature detection, object recognition, and character recognition. Therefore, it is possible to extract scenes of interest from various perspectives.
[0076] For example, according to the information processing device 10 of this embodiment, scenes in which a player is shown in a predetermined posture, scenes in which a player is shown in a predetermined movement, scenes in which a predetermined player is shown, scenes in which a predetermined player is shown in a predetermined posture, scenes in which a predetermined player is shown in a predetermined movement, etc., can be extracted as scenes of interest. As a result, scenes desired by the viewer can be extracted as scenes of interest.
[0077] <Third Embodiment> The information processing device 10 of this embodiment differs from the first and second embodiments in that it can analyze a portion of the first video as described above, while excluding other portions from the analysis. This will be explained in detail below.
[0078] The extraction unit 11 accepts input specifying a time. The extraction unit 11 then extracts the scene of interest from a portion of the first video that is identified based on the specified time, and uses that portion as the subject of the image analysis described above. Other parts of the first video (parts that are not identified based on the specified time) are not subject to the image analysis described above.
[0079] "Specifying the time" is done, for example, by the person creating the highlight video. The person enters the approximate time of scoring plays or moments when the audience was excited.
[0080] The "part identified based on the specified time" refers to frame images taken during a time period identified based on the specified time, for example, from frame images taken a predetermined time before the specified time to frame images taken a predetermined time after the specified time. The predetermined time is a design consideration. For example, the extraction unit 11 can identify frame images taken a predetermined time before the specified time and frame images taken a predetermined time after the specified time, based on the timestamp of the first video (information indicating the time each frame image was taken).
[0081] The extraction unit 11 may, for example, not input the entire first video to the image analysis system 20, but rather extract only the portion specified based on a designated time from the first video and input only the extracted portion to the image analysis system 20. Alternatively, the extraction unit 11 may input the entire first video to the image analysis system 20, as well as information indicating the portion to be analyzed.
[0082] The other configurations of the information processing device 10 in this embodiment are the same as those in the first and second embodiments.
[0083] According to the information processing device 10 of this embodiment, the same effects and advantages as those of the information processing device 10 of the first and second embodiments are achieved.
[0084] Furthermore, the information processing device 10 of this embodiment can analyze only a portion of the first video, rather than the entire video. As a result, the processing load on the image analysis system 20 is reduced, and the time required for image analysis is shortened. For example, if the operator knows the approximate timing of scoring scenes or exciting scenes in advance, the information processing device 10 of this embodiment is useful.
[0085] <Fourth Embodiment> The information processing device 10 of this embodiment differs from the first to third embodiments in that it further has a function to extract noteworthy scenes from the first video based on the results of analyzing a second video of spectators watching the player. This will be explained in detail below.
[0086] The extraction unit 11 extracts noteworthy scenes from the first video based on the results of analyzing the first video as described in the second and third embodiments, as well as the results of analyzing the second video, which shows spectators watching the player. The process of extracting noteworthy scenes from the first video based on the results of analyzing the first video is the same as described in the second and third embodiments.
[0087] As shown in Figure 3, the extraction unit 11 inputs the second video to the image analysis system 20. The extraction unit 11 then acquires the analysis results of the second video output from the image analysis system 20.
[0088] When using the posture recognition function and / or motion recognition function, the analysis results output from the image analysis system 20 include information indicating the posture and / or motion detected in the second video, and information indicating the position of each posture and / or motion in the second video.
[0089] When using the gradient feature detection function for images, among the analysis results output from the image analysis system 20, • Gradient features of each frame image, • Information indicating the position within a second video of a scene that has similar gradient features to the image (scene) being detected. It includes at least one of the following.
[0090] When using the image color feature detection function, among the analysis results output from the image analysis system 20, • Color characteristics of each frame image, • Information indicating the location within a second video of a scene that has similar color features to the detected image (scene). It includes at least one of the following.
[0091] When the object recognition function is used, the analysis results output from the image analysis system 20 include information indicating the position of the detected object in the second video scene.
[0092] Furthermore, in this embodiment, the image analysis system 20 may also have a facial expression detection function. When the facial expression detection function is used, the analysis results output from the image analysis system 20 include information indicating the facial expressions of the audience members detected in the second video, and information indicating the position of the scene in the second video in which each audience member with a particular facial expression is shown.
[0093] The extraction unit 11 detects the target scene in the second video based on the analysis results of the second video output from the image analysis system 20 as described above.
[0094] The scenes to be detected are: • Scenes showing spectators in a specific posture, • Scenes showing spectators performing specific movements. • Scenes showing audience members with a specific facial expression, It is at least one of the following.
[0095] Predetermined postures, movements, and facial expressions are registered in advance. For example, the image analysis system 20 may be configured to detect only predetermined postures, movements, and facial expressions. Alternatively, the image analysis system 20 may be configured to detect not only predetermined postures, movements, and facial expressions, but also other postures, movements, and facial expressions. The extraction unit 11 may then extract predetermined postures, movements, and facial expressions from among those detected by the image analysis system 20.
[0096] For example, by defining certain postures, movements, and facial expressions as predetermined, such as standing, raising both hands in joy, standing up, jumping for joy, happy expressions, or excited expressions, scenes in which the audience is happy or excited can be detected as target scenes. Note that the above-mentioned target scenes are just examples, and other scenes may also be used as target scenes. In addition, target scenes may be detected based on the audio data of the second video. For example, scenes where the audio volume is higher than a certain threshold may be used as target scenes.
[0097] The aforementioned "scenes in which a predetermined subject (posture, movement, and facial expression) is captured" may consist only of frame images in which the predetermined subject is captured, or it may consist of frame images in which the predetermined subject is captured and a predetermined number of frame images before and after it. One scene consists of at least two consecutive frame images.
[0098] After detecting the target scene in the second video using the process described above, the extraction unit 11 extracts the scene of interest from the first video based on the detection result. Specifically, the extraction unit 11 extracts the scene in the first video that was filmed at the same time as the target scene detected in the second video as the scene of interest from the first video. For example, the extraction unit 11 can identify the scene in the first video that was filmed at the same time as the target scene detected in the second video based on the timestamps (information indicating the time each frame image was filmed) of the first video and the second video, respectively.
[0099] The other configurations of the information processing device 10 in this embodiment are the same as those in the first to third embodiments.
[0100] According to the information processing device 10 of this embodiment, the same effects and advantages as those of the information processing device 10 of the first to third embodiments are achieved.
[0101] Furthermore, according to the information processing device 10 of this embodiment, it is possible to extract noteworthy scenes from the first video based on the results of analyzing a second video in which spectators watching the player were filmed. This means that, according to the information processing device 10 of this embodiment, noteworthy scenes from the first video can be extracted from the first video from a different perspective than the information processing device 10 of the first embodiment, which extracts noteworthy scenes from the first video based on the results of analyzing a first video in which the player was filmed.
[0102] Furthermore, according to the information processing device 10 of this embodiment, a scene in the first video that was filmed at the same time as a scene in the second video in which an audience member is shown in a predetermined posture, movement, or facial expression can be extracted as a scene of interest. In this case, for example, a scene in which the audience member is happy and excited can be extracted as a scene of interest.
[0103] <Fifth Embodiment> The information processing device 10 of this embodiment differs from the first to fourth embodiments in that it displays information indicating the location of a scene of interest within the first video on a distinctive UI (user interface) screen. This will be described in detail below.
[0104] The extraction unit 11 groups the scenes of interest extracted from the first video according to their content. The output unit 12 then outputs information indicating the location of each scene of interest within the first video, separated by group. For example, the extraction unit 11 groups the scenes of interest by the players they are filmed with, the postures of the players they are filmed with, the movements of the players they are filmed with, the postures of the spectators in the videos filmed at the same time, the movements of the spectators in the videos filmed at the same time, or the facial expressions of the spectators in the videos filmed at the same time. Note that a single scene may belong to multiple groups.
[0105] Figure 6 schematically shows an example of a UI screen output by the output unit 12. The UI screen shown in Figure 6 displays the file name, player index, and scene index.
[0106] "File name" is the file name of the first video.
[0107] The "Player Index" is a list of the names of the players shown in the first video.
[0108] The "Scene Index" is a list of scenes shown in the first video. For example, scenes of good plays, scenes of players celebrating, scenes where the crowd gets excited, etc.
[0109] In a UI screen like the one shown in Figure 6, when a user selects one of several indices, the output unit 12 may further display scene location information in accordance with that user input, as shown in Figure 7. In the example shown in Figure 7, "Jun Tanaka," enclosed in frame W, has been selected by user input. The scene location column displays information indicating the location of the featured scene (the featured scene for which Jun Tanaka is the reason for extraction) in which the selected Jun Tanaka appears. In the example shown in Figure 7, the starting position of the featured scene is indicated by the elapsed time from the beginning of the first video.
[0110] The other configurations of the information processing device 10 in this embodiment are the same as those in the first to fourth embodiments.
[0111] According to the information processing device 10 of this embodiment, the same effects and advantages as those of the information processing device 10 of the first to fourth embodiments are achieved.
[0112] Furthermore, according to the information processing device 10 of this embodiment, information indicating the location of a scene of interest within the first video can be displayed on a distinctive UI screen as shown in Figures 6 and 7. Specifically, the scene of interest can be grouped according to its content, and information indicating the location of the scene of interest within the first video can be output separately for each group. With this information processing device 10 of this embodiment, a worker creating a highlight video can easily find the desired scene of interest from among multiple scenes of interest. As a result, the problem of the workability of creating a highlight video is solved.
[0113] <Sixth Embodiment> The information processing device 10 of this embodiment differs from the first to fifth embodiments in that it acquires a plurality of first videos and outputs information indicating the position of a scene of interest within each of the plurality of first videos.
[0114] When the playing area where players are playing (baseball field, stadium, concert hall, etc.) is large, or when multiple players are playing simultaneously, multiple cameras may be used to film the area. Multiple first videos are videos generated by filming the same playing area at the same time with multiple cameras in this manner. The multiple cameras may be filming different subjects (players, scoreboard, clock, manager, etc.), different locations (different locations within the same area), or the same subject from different angles.
[0115] The extraction unit 11 performs the image analysis described in the first to fifth embodiments for each of the multiple first videos. The output unit 12 then outputs information indicating the location of the scene of interest within the multiple first videos.
[0116] Figure 8 schematically shows an example of a UI screen output by the output unit 12. The UI screen shown in Figure 8 displays the player index, scene index, and scene position.
[0117] The "Player Index" is a list of the names of the players who appear in any of the multiple first videos.
[0118] A "scene index" is a list of scenes that appear in any of the first videos. For example, scenes of good plays, scenes of players celebrating, scenes of the crowd getting excited, etc.
[0119] "Scene position" is information indicating the location of the extracted scene of interest within the first video. In the example shown in Figure 8, the location of the scene of interest belonging to the group selected by the operator within the first video is shown. In the example shown in Figure 8, the group related to "Jun Tanaka," enclosed in frame W, is selected at that time. Therefore, the scene position column shows the location of the scene of interest in which Jun Tanaka appears. In the example shown in Figure 8, the starting position of each scene of interest is indicated by information linking the file name of the first video with the elapsed time from the beginning of that first video. As illustrated, multiple scenes of interest extracted from multiple videos are displayed together in a list. Also, the scene position may be displayed according to the selection of a single index, as in the example explained using Figures 6 and 7.
[0120] "Variations" Here, a modified version of the sixth embodiment will be described. When the technology of the sixth embodiment is combined with the technology of the third embodiment, which extracts noteworthy scenes from the first video based on the results of analyzing a second video of spectators watching the player, the information processing device 10 can perform the following processing.
[0121] First, the extraction unit 11 identifies where within the play area the audience members included in the detection target scene detected from the second video are looking.
[0122] Specifically, the extraction unit 11 uses image analysis to determine the direction the spectator is facing (direction of gaze, direction of face, or direction of body). Next, based on the map of the play area, the installation position of each of the multiple cameras within the play area, and the background image included in the detection target scene, the extraction unit 11 determines the orientation of each of the multiple cameras at the time the detection target scene was captured. Then, based on the map of the play area, the installation position of each of the multiple cameras within the play area, the orientation of each of the multiple cameras at the time the detection target scene was captured, and the identified direction the spectator is facing, the extraction unit 11 determines where the spectator is looking within the play area. These processes can be implemented using any relevant technology.
[0123] If the scene to be detected includes multiple spectators, the extraction unit 11 may identify the direction each of the multiple spectators is facing, and then statistically calculate the direction (the direction most people are facing, or the average of the directions multiple spectators are facing) and identify that as the direction the spectators are facing.
[0124] Furthermore, at least some of these processes may be performed by the image analysis system 20.
[0125] Next, the extraction unit 11 identifies the camera that captured the position within the play area that the audience is looking at in the detected scene, based on the map of the play area, the installation position of each of the multiple cameras within the play area, the orientation of each of the multiple cameras at the time the detection target scene was captured, and the position within the play area that the audience is looking at in the detected target scene. Then, the extraction unit 11 extracts the scene from the first video captured by the identified camera that was captured at the same time as the detection target scene detected in the second video, as the scene of interest in the first video.
[0126] The other configurations of the information processing device 10 in this embodiment are the same as those in the first to fifth embodiments.
[0127] According to the information processing device 10 of this embodiment, the same effects and advantages as those of the information processing device 10 of the first to fifth embodiments are achieved.
[0128] Furthermore, the information processing device 10 of this embodiment can output the positions of notable scenes extracted from multiple first videos in a single output. In cases where there are multiple players playing simultaneously, such as in baseball, soccer, or concerts, the event may be filmed with multiple cameras. In this case, a more appealing highlight image can be generated by creating a highlight image from multiple first videos produced by filming with multiple cameras. However, the process of watching each of the multiple first videos and selecting the parts to include in the highlight video from each is extremely time-consuming. The information processing device 10 of this embodiment, which outputs the positions of notable scenes extracted from multiple first videos in a single output, improves the efficiency of the process of selecting the parts to include in the highlight video from multiple first videos.
[0129] Furthermore, according to the above-described modified version of the information processing device 10 of this embodiment, for example, a scene from the first video generated by a camera that was filming the position the audience was looking at when the audience was happy and excited can be extracted as a scene of interest.
[0130] <Variation> Here, we will describe some modifications applicable to the first to sixth embodiments.
[0131] -Experimental Variation 1- The extraction unit 11 may, after detecting scenes in which each of the multiple players is shown using the above technology, process each scene using a predetermined method to calculate statistics for each player. The output unit 12 may then output the calculated statistics.
[0132] -Variation 2- The image analysis system 20 may detect multiple postures and / or movements from the first video, then group them together based on similar postures and movements, and output the grouping results. This process can be implemented using the technology described in Patent Document 3. The output unit 12 may then output the grouping results. Based on this output information, the operator can grasp an overview of what postures and movements were detected in the first video. Based on this understanding, the operator can then construct a rough storyline for the highlight video to be created. After constructing the storyline, the operator can find the desired scenes of interest from a UI screen such as those shown in Figures 4, 6, 7, or 8, and create the highlight video.
[0133] -Variation 3- The extraction unit 11 may accept input of highlight videos created in the past. The extraction unit 11 may then extract scenes as noteworthy scenes in which a player is shown with a posture or movement similar to that of a player included in a previously created highlight video.
[0134] In this case, the extraction unit 11 may also create a highlight video by combining the extracted highlight scenes in the same order as previously created highlight videos.
[0135] The embodiments of the present invention have been described above with reference to the drawings, but these are illustrative examples of the present invention, and various other configurations can be adopted. The configurations of the embodiments described above may be combined with each other, or some configurations may be replaced with other configurations. Furthermore, the configurations of the embodiments described above may be modified in various ways without departing from the spirit of the invention. In addition, the configurations and processes disclosed in each of the embodiments and modifications described above may be combined with each other.
[0136] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their contents do not conflict.
[0137] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An extraction means that uses image analysis technology to extract a scene of interest from a portion of a first video of the player that is identified based on a specified time, An output means that outputs information indicating the location of the scene of interest within the first video, An information processing device having 2. The information processing device according to claim 1, wherein the scene of interest is a scene in which a player in a predetermined posture or a player in a predetermined movement is shown. 3. The information processing device according to 1 or 2, wherein the scene of interest is a scene in which a predetermined player is shown. 4. The information processing device according to any one of 1 to 3, wherein the extraction means further extracts the scene of interest from the first video based on the results of analyzing a second video of spectators watching the player. 5. The information processing device according to 4, wherein the scene of interest is a scene in the first video filmed at the same time as a scene in the second video film showing an audience member in a predetermined posture, an audience member in a predetermined movement, or an audience member with a predetermined facial expression. 6. The extraction means groups the extracted scenes of interest according to their content, The output means is an information processing device according to any one of 1 to 5 that outputs information indicating the position of the scene of interest within the first video, separated by group. 7. The extraction means is The information processing device described in 6 groups the aforementioned scenes of interest by each player, each posture of each player, each movement of each player, each posture of spectators in videos filmed at the same time, each movement of spectators in videos filmed at the same time, or each facial expression of spectators in videos filmed at the same time. 8. The computer, Using image analysis technology, we extract notable scenes from a portion of the first video of the player that is identified based on a specified time. The system outputs information indicating the location of the scene of interest within the first video. Information processing methods. 9. Computers, An extraction method that uses image analysis technology to extract a scene of interest from a portion of a first video of the player that is identified based on a specified time. Output means for outputting information indicating the location of the scene of interest within the first video, A program that makes it function as such. 10. An extraction means that extracts scenes of interest from a first video of the player using image analysis technology, and groups the scenes of interest according to their content, An output means that outputs information indicating the location of the scene of interest within the first video, divided into the aforementioned groups, An information processing device having 11. Computers, Using image analysis technology, we extract notable scenes from the first video recording of the player. The aforementioned scenes of interest were grouped according to their content. For each of the aforementioned groups, information indicating the location of the scene of interest within the first video is output. Information processing methods. 12. Computers, An extraction means that extracts scenes of interest from a first video of the player using image analysis technology, and groups the scenes of interest according to their content. Output means that outputs information indicating the location of the scene of interest within the first video, divided into the aforementioned groups. A program that makes it function as such. [Explanation of Symbols]
[0138] 10 Information Processing Devices 11 Extraction part 12 Output section 20 Image Analysis System 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A bus
Claims
1. A means for receiving input specifying a time, An extraction means for extracting a scene of interest from a portion of a first video, based on the results of image analysis of a portion of the first video that was filmed of the player and identified based on a specified time, and the results of image analysis of a second video that was filmed of the audience watching the player. An output means that outputs information indicating the location of the scene of interest within the first video, It has, The extraction means is an information processing device that extracts a portion of the first video that was filmed at the same time as a scene in the second video in which an audience member is shown in a predetermined posture, a scene in which an audience member is shown in a predetermined movement, or a scene in which an audience member is shown with a predetermined facial expression, as the scene of interest.
2. The information processing apparatus according to claim 1, wherein the aforementioned scene of interest is a scene in which a player in a predetermined posture or a player in a predetermined movement is shown.
3. The information processing apparatus according to claim 1 or 2, wherein the aforementioned scene of interest is a scene in which a predetermined player is shown.
4. The extraction means groups the extracted scenes of interest according to their content, The information processing device according to any one of claims 1 to 3, wherein the output means outputs information indicating the position of the scene of interest within the first video, separated by group.
5. The extraction means is The information processing device according to claim 4, which groups the aforementioned scenes of interest by each player, each posture of each player, each movement of each player, each posture of a spectator in a video filmed at the same time, each movement of a spectator in a video filmed at the same time, or each facial expression of a spectator in a video filmed at the same time.
6. Computers Accepts input specifying the time, Based on the results of image analysis of a portion of a first video in which the player was filmed, which is identified based on the specified time, and the results of image analysis of a second video in which the audience watching the player was filmed, a scene of interest is extracted from the portion of the first video. Information indicating the location of the scene of interest within the first video is output. The process for extracting the aforementioned scenes of interest involves extracting a portion of the first video that was filmed at the same time as a scene in the second video in which an audience member is shown in a predetermined posture, a scene in which an audience member is shown in a predetermined movement, or a scene in which an audience member is shown with a predetermined facial expression, as the scenes of interest.
7. Computers A means for receiving input specifying a time, Extraction means for extracting a scene of interest from a portion of a first video, based on the results of image analysis of a portion of the first video that was filmed of the player and identified based on a specified time, and the results of image analysis of a second video that was filmed of the audience watching the player. Output means for outputting information indicating the location of the scene of interest within the first video, To make it function as, The extraction means is a program that extracts a portion of the first video that was filmed at the same time as a scene in the second video in which an audience member is shown in a predetermined posture, a scene in which an audience member is shown in a predetermined movement, or a scene in which an audience member is shown with a predetermined facial expression, as the scene of interest.
Citation Information
Patent Citations
Analytic data generating device
JP1999339009A
Camera automatic control system for athletics, camera automatic control method, camera automatic control unit, and program
JP2008005208A
Image processing device, method and program
JP2008021225A
Control apparatus and photographing method
JP2019193209A
Video information output device, video information output system, video information output program, and video information output method
JP2020170980A