Information processing apparatus, information processing method, program
The information processing apparatus addresses the challenge of generating viewer-focused video content by analyzing scene-related information from events, using social networking data and image analysis to create engaging digest videos with relevant clips from multiple angles.
Patent Information
- Application Number
- JP2023517064
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-26
- Filing Date
- 2022-02-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-02-08
AI Technical Summary
Existing systems struggle to accurately generate video content that reflects the interests and concerns of viewers, as they fail to effectively identify specific scenes of interest from social networking service data.
An information processing apparatus that includes a specifying unit to analyze scene-related information from events like sports games or concerts, using image analysis and metadata to generate clip sets and digest videos based on viewer interests, identified through social networking service data and image analysis of multiple camera angles.
The apparatus effectively generates digest videos that include scenes of high viewer interest, providing a more engaging viewing experience by incorporating clips from various angles and types, reducing processing load and improving content relevance.
Smart Images

Figure 0007704196000001 
Figure 0007704196000002 
Figure 0007704196000003
Abstract
Description
Technical Field
[0001] The present technology relates to the technical field of an information processing apparatus, an information processing method, and a program for generating a digest video.
Background Art
[0002] Video content is preferably created based on the interests and concerns of the viewing user. For example, Patent Document 1 below discloses a system for generating television content so as to include content with a high degree of viewer interest from information posted on a social networking service (SNS).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, it is difficult to grasp which specific scenes are the scenes in which viewers are interested from the information obtained from SNS or the like, and there are cases where appropriate video content cannot be generated.
[0005] The present technology has been made in view of such problems, and an object thereof is to provide video content reflecting the interests and concerns of viewers.
Means for Solving the Problems
[0006] The information processing apparatus according to the present technology includes a specifying unit that specifies auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event. and a clip set generation unit that generates a clip set including one or more clip videos obtained from the imaging device, using the analysis result obtained by image analysis processing on the video obtained from the imaging device that captures the event and the auxiliary information; It is provided with. An event is, for example, an event such as a sports game or a concert. Also, auxiliary information is, for example, information used to generate a digest video and information used to determine which part of the captured video to cut out. For example, in the case of a sports game, specifically, information such as player names, scene types, and play types is regarded as auxiliary information.
Brief Description of Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Embodiments for Carrying Out the Invention
[0008] Hereinafter, with reference to the accompanying drawings, embodiments of the information processing apparatus according to the present technology will be described in the following order. <1. System Configuration> <2. Processing Flow> <2-1. First Processing Flow> <2-2. Second Processing Flow> <2-3. Third Processing Flow> <2-4. Flow of the Generation Process of the Clip Collection> <3. Regarding Scoring> <3-1. Scoring Method> <3-2. Processing Flow in Video Selection Using the Score> <4. Modification Example> <5. Computer Device> <6. Summary> <7. The Present Technology>
[0009] <1. System Configuration> An example of the system configuration of this embodiment will be described with reference to FIG. 1. The information processing apparatus 1 of the present embodiment is an apparatus that generates a digest video DV about events such as sports games, concerts, and stage shows. The generated digest video DV is distributed to viewers.
[0010] In the following description, a sports game is cited as an example of an event. In particular, the information processing apparatus 1 that generates a digest video DV of an American football game will be described.
[0011] The digest video DV is a video that collects important scenes for understanding the flow of the game. Also, the digest video DV can be regarded as a highlight video.
[0012] The information processing apparatus 1 includes a posted data extraction unit 2, a metadata extraction unit 3, a video analysis unit 4, and a video generation unit 5.
[0013] The posted data extraction unit 2 performs a process of extracting keywords from texts, hashtags, videos, etc. posted on an SNS (Social Networking Service). For this purpose, the information processing apparatus 1 is configured to be able to communicate with the SNS server 100 via the communication network NW.
[0014] The keywords extracted by the posted data extraction unit 2 are, for example, the names of players participating in the game, their jersey numbers, or the names of coaches and referees. These pieces of information are information that can identify a person. The player names include not only first names and family names but also nicknames, etc.
[0015] Also, the keywords extracted by the posted data extraction unit 2 may be information of scene types indicating the content of the play. Specifically, it is type information about scoring scenes such as touchdowns and field goals, or type information about various fouls such as offside and holding. Or it may be of types such as super plays and blunders as information indicating better-than-usual plays or failed plays.
[0016] The information extracted by the contribution data extraction unit 2 is information that serves as an indicator for generating the digest video DV. In particular, the information posted on the SNS is information used to generate a digest video DV along with the interests of the viewers.
[0017] The information extracted by the contribution data extraction unit 2 is information about a specific scene in the event, and this is described as "scene-related information".
[0018] The metadata extraction unit 3 performs a process of extracting metadata including information representing the progress of the game, etc. The metadata may be, for example, information independently distributed by a company operating the game, or information input by a recorder (scorer) who records various information such as the progress of the game while watching the game, or data distributed from a company handling information about sports. Alternatively, it may be information about the progress of the game uploaded on the web.
[0019] As an example of metadata, information in which the type information of scenes occurring during the game such as touchdowns, field goals, fouls, player substitutions, and player ejections, the occurrence time of the scene, player information related to the scene, and information such as score changes accompanying the occurrence of the scene are associated is metadata.
[0020] The metadata may be distributed each time a specific scene occurs in the game, or may be distributed collectively after the game ends.
[0021] The metadata extraction unit 3 is information about a specific scene in the event, and this information is also regarded as "scene-related information".
[0022] The information processing apparatus 1 is configured to enable mutual communication with the metadata server 200 via the communication network NW so that the metadata extraction unit 3 can execute the extraction process of the metadata.
[0023] The video analysis unit 4 performs processing to receive video from a plurality of imaging devices CA arranged at the game venue, and also performs image analysis processing on the received video. In addition, the video analysis unit 4 performs processing to acquire the broadcast video VA which is the broadcast video, and performs image analysis processing on the broadcast video VA.
[0024] Note that in FIG. 1, the first imaging device CA1, the second imaging device CA2, and the third imaging device CA3 are illustrated as examples of the imaging device CA, but this is just an example. It is also possible that only one imaging device CA is installed at the game venue, or four or more imaging devices CA are installed at the game venue.
[0025] Also, the video obtained from the first imaging device CA1 is taken as the first video V1, the video obtained from the second imaging device CA2 is taken as the second video V2, and the third video obtained from the third imaging device CA3 is taken as V3.
[0026] Each imaging device CA is synchronized so that it is possible to know the images captured at the same timing by referring to the time code.
[0027] The video analysis unit 4 obtains information on the subject being imaged at each time by image analysis processing. Examples of the subject information include the name of the subject such as the player's name, the back number information, the imaging angle, the posture of the subject, etc. Also, the subject may be specified based on facial features, hairstyle, hair color, expression, etc.
[0028] The video analysis unit 4 obtains information on the scene type for specifying the scene by image analysis processing. Examples of the scene type information include information such as whether the imaged scene is a scoring scene, a foul scene, a player substitution scene, an injury scene, etc. The scene type may be specified by detecting the posture of the subject described above. For example, the judge's content may be estimated by detecting the posture of the referee to specify the scene type, or the scoring scene may be detected by detecting the player's guts pose.
[0029] The video analysis unit 4 identifies in-points and out-points through image analysis processing. The in-points and out-points are information for identifying the cut-out range of the video captured by the imaging device CA. In the following description, the video within a predetermined range cut out by a set of in-points and out-points will be referred to as the "clip video CV".
[0030] The in-points and out-points may be determined, for example, by identifying the moment when the play of the detection target occurs through image analysis processing and using that as a basis point. Also, when detecting in-points and out-points based on the broadcast video VA, it may be done by detecting the timing of video switching. That is, the video analysis unit 4 may perform image analysis processing on the broadcast video VA and identify the in-points and out-points by detecting the switching points of the imaging device CA.
[0031] The video analysis unit 4 attaches information obtained through image analysis processing to the video. For example, it is associated and stored that in a certain time period in the first video V1, players A and B are imaged, and that the said time period is a scene of a touchdown. Thereby, for example, when it is desired to create a digest video DV using the scene where a specific player is imaged, the time period when the specific player is imaged can be easily identified.
[0032] The video analysis unit 4 identifies the game progress by executing image analysis processing on the broadcast video VA. The broadcast video VA is generated by connecting specific partial videos (clip videos CV) using the first video V1, the second video V2, the third video V3, etc. as materials, and superimposing various information such as score information and player name information.
[0033] In the image analysis processing, by recognizing subtitles, three-dimensional images, etc. superimposed on the video, the transition of scores, player substitutions, the player names of the players imaged in the image, the elapsed time in the game, etc. are identified.
[0034] In addition, the video analysis unit 4 may assign a score to each video by performing image analysis processing. The score may be calculated as the likelihood when the imaged subject is identified, or may be calculated as an index indicating whether the video presented to the viewer is appropriate or not.
[0035] Note that in FIG. 1, the configuration in which the video analysis unit 4 acquires a video from the imaging device CA is shown, but the video may be acquired from a storage device in which the video captured by the imaging device CA is stored.
[0036] The video generation unit 5 performs a process of generating a digest video DV using the first video V1, the second video V2, and the third video V3.
[0037] For this purpose, the video generation unit 5 includes an identification unit 10, a clip set generation unit 11, and a digest video generation unit 12 (see FIG. 2).
[0038] The identification unit 10 performs a process of identifying auxiliary information SD for generating the digest video DV. Here, an example of the flow of generating the digest video DV is shown.
[0039] Suppose a scoring scene occurs in a certain sports game. In this case, a clip set CS for the scoring scene is generated. The clip set CS is a combination of a plurality of clip videos CV. For example, a clip video CV obtained by cutting out the time period in which the scoring scene is imaged from the first video V1 captured by the first imaging device CA1, a clip video CV obtained by cutting out the time period in which the scoring scene is imaged from the second video V2 captured by the second imaging device CA2, and a clip video CV obtained by cutting out the time period in which the scoring scene is imaged from the third video V3 captured by the third imaging device CA3 are combined to generate a clip set CS for the scoring scene.
[0040] Such clip sets CS are generated, for example, for the number of scoring scenes, or for the number of foul scenes, or for the number of player substitution scenes.
[0041] The digest video DV is generated by selecting and combining the clip set CS to be presented to the viewer from the plurality of clip sets CS thus generated.
[0042] For example, auxiliary information SD is used to select the clip video CV to be included in the clip set CS. The auxiliary information SD is used as a keyword for selecting the clip set CS included in the digest video DV from the plurality of clip sets CS. If the name of a player who is an SNS is frequently posted, it can be determined that the viewer's interest in the player is high. In that case, a scoring scene or a foul scene in which the player is involved is selected and incorporated into the digest video DV.
[0043] Note that not only the player name and the above-mentioned nickname, but also information that can identify the player may be used. For example, keywords such as a position name or a referee may be used.
[0044] Alternatively, the auxiliary information SD may be used as a keyword as scene type information. For example, if there are many posts about foul scenes on SNS, it can be determined that the viewer's interest in foul scenes is high. In that case, the clip set CS of the foul scene is selected and incorporated into the digest video DV.
[0045] Note that the auxiliary information SD may be type information such as a scoring scene or a foul scene, or may be a keyword indicating type information such as a more detailed field goal scene, a touchdown scene, or a specific foul name.
[0046] Further, the auxiliary information SD may indicate the order of combination of the clip videos CV included in the clip set CS. By generating the clip collection CS based on the auxiliary information SD, for example, as shown in FIG. 3, when the first video V1 is a wide-angle video captured from an overhead view from the side of the field, the second video V2 is a telephoto video captured near the player holding the ball, and the third video V3 is a video captured from the goal post side, it becomes possible to combine the clip videos CV cut out from each video in an appropriate order.
[0047] Note that the auxiliary information SD indicating the combination order may be different according to the scene type. For example, the scoring scene may start from the wide-angle video, and the foul scene may start from the telephoto video.
[0048] Alternatively, the auxiliary information SD may be information indicating whether it is a broadcast video. There is a possibility that the viewer has already watched the broadcast video VA about the game. Since showing the same video to such a viewer will not provide significant information to the viewer, it is conceivable to generate the digest video DV so that it includes videos captured from angles that the viewer has not watched. The auxiliary information SD indicating whether it is a broadcast video is used for the selection of the clip collection CS or the selection of the clip video CV in such a case.
[0049] The clip collection generation unit 11 generates the clip video CV based on the auxiliary information SD. Specifically, by presenting the specified auxiliary information SD such as the player name to the video analysis unit 4, the video analysis unit 4 determines the in-point and out-point of the video in which the player is captured and generates the clip video CV.
[0050] The clip collection generation unit 11 combines the clip videos CV to generate the clip collection CS. The combination order of the clip videos CV may be based on the auxiliary information SD or may be a predetermined order determined in advance.
[0051] That is, the clip collection generation unit 11 generates the clip collection CS using the analysis result of the image analysis process by the video analysis unit 4 and the auxiliary information SD.
[0052] Note that when combining two clip videos CV, the clip collection generation unit 11 may insert an image representing a video switch between the clip videos CV.
[0053] The digest video generation unit 12 combines the clip collection CS generated by the clip collection generation unit 11 to generate a digest video DV. The combination order of the clip collections CS is determined, for example, according to the occurrence time of each scene. An image or the like representing a video switch may be inserted between the clip collections CS.
[0054] The generated digest video DV may be posted on the SNS or uploaded onto a web page.
[0055] <2. Processing flow> Some examples of the processing executed by the information processing apparatus 1 will be described.
[0056] <2-1. First processing flow> Examples of the first processing flow are shown in each of FIGS. 4 to 7. Specifically, an example of the processing flow executed by the posting data extraction unit 2 of the information processing apparatus 1 is shown in FIG. 4, an example of the processing flow executed by the metadata extraction unit 3 is shown in FIG. 5, an example of the processing flow executed by the video analysis unit 4 is shown in FIG. 6, and an example of the processing flow executed by the video generation unit 5 is shown in FIG. 7.
[0057] The posting data extraction unit 2 analyzes the posting data of the SNS in step S101 of FIG. 4. High-frequency keywords and highly noticeable keywords are extracted by this analysis process. These keywords are, for example, the player names and scene types described above.
[0058] Next, in step S102, the post data extraction unit 2 determines whether the extracted keyword is related to the target event. Specifically, it determines whether the extracted person's name exists as a member of the team participating in the game that is the generation target of the digest video DV, or determines whether the extracted keyword is related to the target game.
[0059] If it is determined that the keyword is related to the target event, in step S103, the post data extraction unit 2 performs a process of outputting the extracted keyword to the metadata extraction unit 3.
[0060] On the other hand, if it is determined that the keyword is not related to the target event, the post data extraction unit 2 does not perform the process of step S103 and determines in step S104 whether the event has ended.
[0061] If it is determined that the event has not ended, the post data extraction unit 2 returns to the process of step S101 to continue extracting keywords. On the other hand, if it is determined that the event has ended, the post data extraction unit 2 ends the series of processes shown in FIG. 4.
[0062] Note that in FIG. 4 and each subsequent figure, since it is an example of generating a clip set CS for generating a digest video DV in parallel with the progress of the event, the process of determining whether the event has ended is executed in step S104.
[0063] On the contrary, when generating the clip set CS and the digest video DV after the end of the event, instead of the determination process in step S104, a process of determining whether keyword extraction and the like have been completed for all the post data posted on the SNS during the time period when the event was held may be executed.
[0064] By executing the series of processes shown in FIG. 4, the contribution data extraction unit 2 continuously extracts keywords from the contribution data posted on the SNS from the start to the end of an event such as a sports game, and outputs them to the metadata extraction unit 3 as appropriate.
[0065] In parallel with the execution of the process shown in FIG. 4 by the contribution data extraction unit 2, the metadata extraction unit 3 executes the series of processes shown in FIG. 5. Specifically, in step S201, the metadata extraction unit 3 analyzes the metadata acquired from the metadata server 200 and extracts information for identifying the scene that occurred in the event. For example, in the case of an American football game, the time when a scene corresponding to a touchdown, which is one of the scene types, occurred, the name of the player who scored a touchdown, information on the change in score due to the touchdown, etc. are extracted.
[0066] Subsequently, in step S202, the metadata extraction unit 3 determines whether the keyword extracted from the SNS post was obtained from the contribution data extraction unit 2.
[0067] If the keyword information has not been obtained, the metadata extraction unit 3 returns to the process of step S201.
[0068] If the keyword information has been obtained, the metadata extraction unit 3 identifies the metadata related to the obtained keyword in step S203.
[0069] Subsequently, in step S204, the metadata extraction unit 3 outputs the identified metadata to the video analysis unit 4.
[0070] Then, in step S205, the metadata extraction unit 3 determines whether the event has ended.
[0071] If it is determined that the event has not ended, the metadata extraction unit 3 returns to the process of step S201 to perform the process of analyzing the metadata. If it is determined that the event has ended, the metadata extraction unit 3 ends the series of processes shown in FIG. 5.
[0072] By executing the series of processes shown in FIG. 5 by the metadata extraction unit 3, from the start to the end of an event such as a sports game, the analysis process of the metadata accumulated in the metadata server 200 as an external information processing device is continuously executed, and the information of each scene occurring during the game is extracted.
[0073] In parallel with the execution of the process shown in FIG. 4 by the posting data extraction unit 2 and the execution of the process shown in FIG. 5 by the metadata extraction unit 3, the video analysis unit 4 executes the series of processes shown in FIG. 6.
[0074] In step S301, the video analysis unit 4 performs video analysis by performing image recognition processing on a plurality of videos such as the first video V1, the second video V2, the third video V3, and the broadcast video VA, and identifies the back numbers, the faces of the players, the balls, etc. captured in the video. Further, the video analysis unit 4 may specify the camera angle, or may specify the in-point and out-point for generating the clip video CV.
[0075] In the face recognition process, likelihood information indicating the plausibility of the recognition result may be calculated. The likelihood information is used in the video selection process and the like in the subsequent video generation unit 5.
[0076] The information specified by the image recognition process is associated with time information such as the elapsed time of the game and the elapsed time since the start of recording for each of the plurality of videos and stored.
[0077] In step S302, the video analysis unit 4 determines whether the event has ended. If it is determined that the event has not ended, the video analysis unit 4 returns to the process of step S301 to continue the video analysis process. On the other hand, if it is determined that the event has ended, the video analysis unit 4 ends the series of processes shown in FIG. 6.
[0078] By executing a series of processes shown in FIG. 6 by the video analysis unit 4, various kinds of information are extracted from the video captured from the start to the end of an event such as a sports game.
[0079] The video generation unit 5 generates a digest video DV according to the processing results of the posting data extraction unit 2, the metadata extraction unit 3, and the video analysis unit 4.
[0080] Specifically, in step S401 of FIG. 7, the video generation unit 5 determines whether keywords and metadata have been acquired.
[0081] When the keyword posted to the SNS has been acquired from the posting data extraction unit 2 or when the information about the metadata has been acquired from the metadata extraction unit 3, the video generation unit 5 proceeds to step S402 and performs a process of generating a clip video CV for the target scene based on the keyword or metadata. This process generates the clip video CV based on the in-point and out-point specified by the video analysis unit 4 for the target scene.
[0082] After generating the clip video CV, in step S403, the video generation unit 5 combines the clip videos CV to generate a clip set CS for the target scene. The clip video CV may be generated, for example, by combining the first video V1, the second video V2, and the third video V3 in a predetermined order.
[0083] Alternatively, a template may be prepared so that videos are combined in the order of a predetermined camera angle according to the scene type, and each clip video CV is applied to the template based on the camera angle information for each imaging device CA so that the clip videos CV are combined in an optimal order.
[0084] After generating the clip video CV, the video generation unit 5 returns to the process of step S401.
[0085] In the determination process of step S401, if it is determined that keywords or metadata have not been acquired, the video generation unit 5 proceeds to step S404 and determines whether the event has ended.
[0086] If it is determined that the event has not ended yet, the video generation unit 5 returns to step S401 and continues to generate the clip video CV and the clip collection CS.
[0087] On the other hand, if it is determined that the event has ended, the video generation unit 5 proceeds to step S405 and combines the clip collection CS to generate the digest video DV.
[0088] The digest video DV is basically generated by combining the clip collections CS for each scene that occurred during the game in chronological order.
[0089] If there is a limit on the playback time length of the digest video DV, the digest video DV will be generated while making a selection to include the clip collection CS with a high priority from among the clip collections CS.
[0090] The clip collection CS with a high priority includes the clip collection CS corresponding to the scene where either team scored, the clip collection CS corresponding to the scene where the viewer's interest is presumed to be high from the posted data on SNS, and the like.
[0091] In addition, when selecting the clip collection CS, the posted data posted within a predetermined period (such as 10 minutes or 30 minutes) after the end of the game may be used. For example, it is presumed that the posted data posted within a predetermined period after the end of the game includes posts that summarize the game, posts that mention a scene like watching the game again, and the like.
[0092] By selecting the clip collection CS based on such posted data, it becomes possible to generate a digest video DV that the viewers are highly interested in.
[0093] After generating the digest video DV, the video generation unit 5 performs a process of saving the digest video DV in step S406. The place where the digest video DV is saved may be a storage unit inside the information processing apparatus 1, or may be a storage unit of a server apparatus different from the information processing apparatus 1.
[0094] <2-2. Second processing flow> Examples of the second processing flow are shown in the respective figures from FIG. 8 to FIG. 11. Note that, for processes similar to those described in the first processing flow, the same step numbers are assigned and the description is omitted as appropriate.
[0095] The post data extraction unit 2 analyzes the post data of the SNS in step S101 of FIG. 8. By this analysis process, keywords with high appearance frequencies such as player names and scene types, and keywords with high attention levels are extracted.
[0096] Next, the post data extraction unit 2 determines in step S102 whether the extracted keyword is related to the target event.
[0097] If it is determined that the keyword is related to the target event, the post data extraction unit 2 performs a process of classifying the extracted keyword in step S110.
[0098] For example, the extracted keyword is classified into any one of a keyword related to a person such as a player, a referee, or a coach, a keyword related to a scoring scene such as a field goal or a touchdown, and a keyword related to a foul scene such as offside or holding.
[0099] Note that the three classifications shown here are merely examples, and the keyword may be classified into other categories.
[0100] After classifying the keyword, the post data extraction unit 2 outputs the classification result to the metadata extraction unit 3 in step S111.
[0101] On the other hand, if it is determined that it is not related to the target event, or after executing step S111, the post data extraction unit 2 determines whether the event has ended in step S104 without performing the processes of step S110 and step S111.
[0102] If it is determined that the event has not ended, the post data extraction unit 2 returns to the process of step S101 to continue extracting keywords. On the other hand, if it is determined that the event has ended, the post data extraction unit 2 ends the series of processes shown in FIG. 8.
[0103] In parallel with the execution of the process shown in FIG. 8 by the post data extraction unit 2, the metadata extraction unit 3 executes a series of processes shown in FIG. 9.
[0104] The metadata extraction unit 3 determines whether it has obtained the classification result of the keyword in step S210 of FIG. 9. If it is determined that the classification result has been obtained, the metadata extraction unit 3 performs a branch process according to the classification result in step S211.
[0105] For example, if the extracted keyword is related to a person, the metadata extraction unit 3 identifies the metadata including the person related to the keyword in step S212.
[0106] Or, if the extracted keyword is related to a scoring scene, the metadata extraction unit 3 identifies the metadata about the scoring scene in step S213.
[0107] Also, if the extracted keyword is related to a foul scene, the metadata extraction unit 3 identifies the metadata about the foul scene in step S214.
[0108] After executing any one of steps S212, S213, or S214, the metadata extraction unit 3 proceeds to step S204 and outputs the identified metadata and the above-described classification result to the video analysis unit 4.
[0109] Then, in step S205, the metadata extraction unit 3 determines whether the event has ended.
[0110] If it is determined that the event has not ended, the metadata extraction unit 3 returns to the process of step S210 to perform the acquisition determination of the classification result. On the other hand, if it is determined that the event has ended, the metadata extraction unit 3 ends the series of processes shown in FIG. 9.
[0111] In parallel with the process shown in FIG. 8 by the submission data extraction unit 2 and the process shown in FIG. 9 by the metadata extraction unit 3, the video analysis unit 4 executes the series of processes shown in FIG. 10.
[0112] In step S310, the video analysis unit 4 determines whether it has acquired the metadata and the classification result from the metadata extraction unit 3. If it is determined that the metadata has not been acquired, the video analysis unit 4 executes the process of step S310 again.
[0113] On the other hand, if it is determined that the metadata has been acquired, the video analysis unit 4 proceeds to step S311 and performs branch processing according to the classification result.
[0114] For example, if the metadata is related to a person, in step S312, the video analysis unit 4 performs back number recognition and face recognition by image recognition processing to identify the time zone when the identified person was imaged.
[0115] Alternatively, if the metadata is related to a scoring scene, in step S313, the video analysis unit 4 performs scoreboard recognition by image recognition processing to identify the scoring scene.
[0116] Scoreboard recognition by image recognition processing may, for example, perform processing to detect a location where a scoreboard installed in a venue is imaged and extract the score of the scoreboard, or detect changes in the scores of both teams by recognizing subtitles, graphics, etc. superimposed on the captured image by analyzing the broadcast video VA.
[0117] Note that since the time when a scoring scene occurred is clear from the metadata, instead of performing image recognition processing on the entire captured video, image recognition processing may be performed on the video within a predetermined range before and after the specified time centered on the specified time. Thereby, it is possible to reduce the processing load and shorten the processing time related to the image recognition processing.
[0118] Also, when the extracted keyword is related to a foul scene, the video analysis unit 4 performs detection of a foul display by image recognition processing in step S314 to identify the foul scene.
[0119] The image recognition processing for identifying a foul scene may, for example, identify the timing of the occurrence of a foul scene by recognizing a yellow flag thrown into the field, or identify a foul scene by recognizing subtitles, graphics, etc. superimposed on the captured image by analyzing the broadcast video VA. Also, in the case of soccer, by detecting the posture of the referee, a scene where a yellow card or a red card is shown to the target player may be identified as a foul scene.
[0120] In the image analysis processing of step S314, processing may be performed on the video of a predetermined section based on the metadata in the same manner as in step S313.
[0121] After executing any one of steps S312, S313, or S314, the video analysis unit 4 proceeds to step S315 and identifies the camera angle by image analysis processing. The information on the camera angle identified here is used in the process of generating the subsequent clip collection CS.
[0122] Subsequently, in step S316, the video analysis unit 4 executes image analysis processing for identifying the in-point and out-point. Note that the in-point and out-point may be determined based on the occurrence timing of the scene. For example, the time 15 seconds before the occurrence timing of the scene may be set as the in-point, and the time 20 seconds after the in-point may be set as the out-point.
[0123] In step S302, the video analysis unit 4 determines whether the event has ended. If it is determined that the event has not ended, the video analysis unit 4 returns to the process of step S310. On the other hand, if it is determined that the event has ended, the video analysis unit 4 ends the series of processes shown in FIG. 10.
[0124] The video generation unit 5 generates a digest video DV according to the processing results of the posting data extraction unit 2, the metadata extraction unit 3, and the video analysis unit 4.
[0125] Specifically, in step S410 of FIG. 11, the video generation unit 5 determines whether it has detected that the in-point and out-point have been identified.
[0126] If it has detected that the in-point and out-point have been identified, the video generation unit 5 proceeds to step S411 and performs a process of generating a clip video CV based on the in-point and out-point.
[0127] After generating the clip video CV, in step S403, the video generation unit 5 combines the clip videos CV to generate a clip set CS for the target scene.
[0128] After generating the clip video CV, the video generation unit 5 returns to the process of step S410.
[0129] In the determination process of step S410, when it is determined that it has not been detected that the in-point and out-point have been specified, the video generation unit 5 proceeds to step S404 to determine whether the event has ended.
[0130] If it is determined that the event has not ended yet, the video generation unit 5 returns to step S410.
[0131] On the other hand, if it is determined that the event has ended, the video generation unit 5 proceeds to step S405, combines the clip set CS to generate the digest video DV, and in the subsequent step S406, performs the process of saving the digest video DV.
[0132] <2-3. Third processing flow> The third processing flow is an example of the case where the digest video DV is generated without using metadata.
[0133] Specifically, it will be described with reference to FIGS. 8, 10, and 11.
[0134] The submission data extraction unit 2 extracts and classifies the keywords related to the event by executing the series of processes shown in FIG. 8. The classification result is output to the video analysis unit 4 in step S111.
[0135] Since the metadata extraction unit 3 does not need to perform the analysis of metadata, it does not perform any processing.
[0136] The video analysis unit 4 determines whether it has obtained the classification result of the keywords instead of determining whether it has obtained the metadata in step S310 of FIG. 10.
[0137] Then, according to the classification result of the keywords, the respective processes from S311 to S316 are appropriately executed.
[0138] The video generation unit 5 generates the digest video DV by executing the series of processes shown in FIG. 11.
[0139] In this way, it is possible to generate a digest video DV with an appealing power to viewers using only the post data to the SNS without using metadata.
[0140] <2-4. Flow of generation process of clip set> Regarding the generation process of the clip set CS described in step S403 of FIG. 7 and FIG. 11, a specific processing flow will be described.
[0141] The first example is an example using different templates for each scene type.
[0142] In step S501 of FIG. 12, the video generation unit 5 performs a branch process according to the scene type of the target scene. The type of the target scene may be estimated from keywords or determined based on metadata.
[0143] When the scene type is a touch-down scene, in step S502, the video generation unit 5 selects a template for the touch-down scene.
[0144] As described above, the template is information that determines which video of what camera angle is combined in what order.
[0145] When the scene type is a field goal scene, in step S503, the video generation unit 5 selects a template for the field goal scene.
[0146] When the scene type is a foul scene, in step S504, the video generation unit 5 selects a template for the foul scene.
[0147] After selecting any template in steps S502, S503 or S504, the video generation unit 5 executes a process of generating the clip set CS using the selected template in step S505.
[0148] Also, in step S501, when it is determined that the scene type does not correspond to any of them, the video generation unit 5 adopts the target section in the broadcast video VA as the clip set CS in step S506.
[0149] The target section may be determined, for example, based on the posting time to the SNS or based on the scene occurrence time in the metadata.
[0150] After executing either the process of step S505 or S506, the video generation unit 5 finishes the generation process of the clip set CS.
[0151] Another example is an example in which not only the generation of the clip set CS but also the determination of the in-point and out-point for the generation of the clip video CV are made according to the scene type of the target scene.
[0152] Specifically, it is the process executed instead of steps S402 and S403 in FIG. 7, and the process executed instead of steps S411 and S403 in FIG. 11. This process will be described as step S421 (see FIG. 13).
[0153] In step S501, the video generation unit 5 performs branch processing according to the scene type of the target scene.
[0154] When the scene type is a touch-down scene, the video generation unit 5 determines the in-point and out-point for the touch-down scene and generates the clip video CV in step S510. At this time, the in-point and out-point may be determined, for example, so that the clip video CV has an optimal length.
[0155] Next, in step S502, the video generation unit 5 selects a template for the touch-down scene.
[0156] Also, when the scene type is the field goal scene, in step S511, the video generation unit 5 determines in-points and out-points for the field goal scene and generates the clip video CV.
[0157] Next, in step S503, the video generation unit 5 selects a template for the field goal scene.
[0158] Furthermore, when the scene type is the foul scene, in step S512, the video generation unit 5 determines in-points and out-points for the foul scene and generates the clip video CV.
[0159] Next, in step S504, the video generation unit 5 selects a template for the foul scene.
[0160] After executing any one of steps S502, S503, or S504, in step S505, the video generation unit 5 executes a process of generating the clip set CS using the selected template.
[0161] Also, in step S501, when it is determined that the scene type does not correspond to any, in step S506, the video generation unit 5 adopts the target section in the broadcast video VA as the clip set CS.
[0162] The target section may be determined based on, for example, the posting time to the SNS, or may be determined based on the scene occurrence time in the metadata.
[0163] After executing either of the processes in steps S505 or S506, the video generation unit 5 finishes the generation process of the clip set CS.
[0164] Note that in FIGS. 12 and 13, an example in which one template is prepared for the foul scene is shown, but different templates may be prepared according to the type of foul. Also, templates may be prepared not only for the illustrated cases but also for other scene types such as the injury scene.
[0165] <3. Scoring> <3-1. Scoring Method> When there is a limit on the playback time of the clip collection CS, it may not be possible to combine all the selected clip videos CV. In such a case, a scoring process of attaching scores to each clip video CV may be performed so that the clip video CV with a higher score is preferentially included in the clip collection CS.
[0166] FIG. 14 shows an example of the score given as a result of scoring for the size of the subject and the score given as a result of scoring for the orientation of the subject with respect to the clip video CV for each imaging device CA. Note that each score is a value in the range of 0 to 1, and the larger the value, the better the score.
[0167] The first video V1 is an overhead view video, and since the subject is imaged small, the score for the size of the subject is 0.02. Also, regarding the orientation of the subject, since the subject is small and the orientation is difficult to understand and the parts of the subject's face cannot be clearly distinguished, the score for the orientation of the subject is 0.1.
[0168] The second video V2 is a telephoto video in which a player holding a ball is shown large, and the score for the size of the subject is 0.85. Also, regarding the orientation of the subject, since the subject is facing the front with respect to the imaging device CA and the parts of the subject's face are clearly imaged, the score for the orientation of the subject is 0.9.
[0169] The third video V3 is an overhead view video that images a relatively narrow area, and since the size of the subject is not very large, the score for the size of the subject is 0.1. Also, regarding the orientation of the subject, since the subject is small and the orientation is difficult to understand and the parts of the subject's face cannot be clearly distinguished, the score for the orientation of the subject is 0.1.
[0170] The fourth video V4 is the video captured by the fourth imaging device CA4. The fourth video V4 is a telephoto video in which the subject is captured large, and the score for the size of the subject is set to 0.92. However, since the orientation of the subject is not directly facing the imaging device CA, the score for the orientation of the subject is set to 0.1.
[0171] When prioritizing videos in which the subject appears large, the fourth video V4 is preferentially selected. Also, when prioritizing videos in which the front of the subject is captured, the second video V2 is preferentially selected.
[0172] In this way, by selecting the clip video CV with reference to different scores according to the purpose, it is possible to generate an appealing clip set CS and a digest video DV.
[0173] Note that the scoring may be calculated not only for each clip video CV but also for each clip set CS including a plurality of clip videos CV. And when selecting the clip set CS included in the digest video DV, it may be such that the clip set CS with a high score for each clip set CS given by the scoring is likely to be included.
[0174] Also, in the scoring process of the clip video CV, the clip video CV including the imaging image with the highest score may be selected, or the clip video CV may be selected based on the average score of each imaging image. The average score is, for example, the average of the scores calculated for each imaging image included in the clip video CV.
[0175] <3-2. Processing Flow in Video Selection Using Scores> The specific processing procedure of the clip set CS generation process described in step S403 of FIG. 7 and FIG. 11 will be described. In particular, in this example, an example of generating the clip set CS using scores will be described.
[0176] Regarding the scoring process, it is executed by the video analysis unit 4 after step S301 in FIG. 6 and step S316 in FIG. 10. Therefore, at the stage of executing the series of processes shown in FIG. 15, scores are given to each clip video CV in various states.
[0177] In step S601 of FIG. 15, the video generation unit 5 selects a clip video CV whose score is equal to or higher than the threshold value. Thereby, videos with low scores that are not attractive to viewers can be excluded.
[0178] In step S602, the video generation unit 5 generates a clip collection CS by combining the clip videos CV in the order of scores.
[0179] The score given by the scoring process can be regarded as an indicator indicating that the video is easy for viewers to watch and is appropriate for understanding what happened in the scene.
[0180] By generating the clip collection CS by combining the clip videos CV with high scores in order, viewers who watch the clip collection CS can correctly understand what happened in the scene. In other words, it is possible to prevent a situation where viewers watch clip videos CV with low scores and cannot understand the events that occurred in the scene.
[0181] Another example of the process of generating the clip collection CS using scores will be described with reference to FIG. 16.
[0182] Note that this example is a process of generating the clip video CV and the clip collection CS, and is a process executed instead of step S402 and step S403 in FIG. 7, or a process executed instead of step S411 and step S403 in FIG. 11.
[0183] Similar to the example described above, this process is described with reference to FIG. 16 as step S421. Note that each process shown in FIG. 16 is described as being executed by the video generation unit 5, but some processes may be executed by the video analysis unit 4.
[0184] In step S501, the video generation unit 5 performs a branch process according to the scene type of the target scene.
[0185] When the scene type is a touch-down scene, in step S610, the video generation unit 5 selects a video (imaging device CA) optimal for the touch-down scene. A plurality of videos may be selected. That is, a plurality of imaging devices CA may be selected.
[0186] Next, in step S502, the video generation unit 5 selects a template for the touch-down scene.
[0187] When the scene type is a field goal scene, in step S611, the video generation unit 5 selects a video optimal for the field goal scene.
[0188] Next, in step S503, the video generation unit 5 selects a template for the field goal scene.
[0189] Furthermore, when the scene type is an irregular scene, in step S612, the video generation unit 5 selects a video optimal for the irregular scene.
[0190] Next, in step S504, the video generation unit 5 selects a template for the irregular scene.
[0191] After executing any of steps S502, S503, or S504, in step S613, the video generation unit 5 determines in-points and out-points for sections where the score is equal to or higher than the threshold from the section where the target scene was imaged, and generates a clip video CV. This process is executed for each selected video.
[0192] Next, in step S505, the video generation unit 5 executes a process of generating a clip set CS using the selected template.
[0193] Also, in step S501, when it is determined that the scene type does not correspond to any of them, the video generation unit 5 adopts the target section in the broadcast video VA as the clip set CS in step S506.
[0194] The target section may be determined based on, for example, the posting time to SNS, or may be determined based on the scene occurrence time in the metadata.
[0195] After executing either the process of step S505 or S506, the video generation unit 5 finishes the generation process of the clip set CS.
[0196] Another example of the process of generating the clip set CS using the score will be described with reference to FIG. 17.
[0197] Note that this example is a process of generating the clip video CV and the clip set CS, and is a process executed instead of steps S402 and S403 in FIG. 7, or a process executed instead of steps S411 and S403 in FIG. 11.
[0198] In step S601 of FIG. 17, the video generation unit 5 selects a clip video CV whose score is equal to or higher than the threshold value. Thereby, videos with low scores that are not attractive to viewers can be excluded.
[0199] In step S613, the video generation unit 5 cuts out a section with a score equal to or higher than the threshold value from the selected clip videos CV and newly generates it as the clip video CV. Specifically, the in-point and out-point of the section with a score higher than the threshold value are determined to generate the clip video CV. This process is executed for each selected video.
[0200] In step S602, the video generation unit 5 generates a clip collection CS by combining the clip videos CV in the order of scores.
[0201] As a result, a high-score section is further carefully selected and cut out from the clip videos CV with high scores, so that a digest video DV using only videos that viewers are highly interested in can be generated.
[0202] <4. Modification Example> In the above example, it is shown that posting data is extracted from the SNS. Here, the posting data to be extracted may be from an unspecified number of accounts or a specific account. By extracting the posting data for an unspecified number of accounts, it becomes possible to better grasp the interests of viewers. On the other hand, by extracting the posting data for a specific account used by team members or those who are doing a live broadcast of a game, etc., the possibility of extracting incorrect information can be reduced. That is, a certain amount of noise can be removed.
[0203] Also, the extraction of the posting data may be the posting data itself, or the information obtained after performing statistical processing on the posting data. For example, it may be information extracted by statistical processing, such as keywords that appear frequently in the information posted in a recent predetermined time. These information may be extracted in the SNS server 100 that manages the postings to the SNS, or obtained from another server device that analyzes the postings to the SNS server 100.
[0204] The video analysis unit 4 showed an example of analyzing the broadcast video VA. When analyzing the broadcast video VA, not only image analysis processing but also voice analysis processing for analyzing the voices of the live commentator and the narrator may be performed. As a result, it becomes possible to more specifically and accurately identify the scenes that occurred during the game, and it also becomes easier to identify the players related to the scene. Further, voice analysis processing may be used to determine the in-point and out-point for generating the clip video CV. Also, by analyzing the voices of the audience's cheers, etc., it may be possible to grasp the occurrence timing of the scene and identify the scene type.
[0205] In addition to the scene types described above, rough play scenes, miss play scenes, good play scenes, commemorative play scenes, etc. may be detected and included in the digest video DV. Note that a commemorative play is, for example, a play at the moment when a certain player's cumulative score reaches a predetermined value, or a play when the previous record is rewritten.
[0206] In the example of using a template, when the corresponding video does not exist, the clip collection CS may be generated without combining the video. For example, when there is no enlarged video of the shooting angle specified by the template, the clip collection CS is generated without including that video.
[0207] Depending on the sport, the gestures of the referees may be finely set according to the type of play and the type of foul. In such a case, a dedicated imaging device CA for imaging the referee is arranged in the venue, and by identifying the posture and gestures of the referee through image analysis processing, it becomes possible to identify the content of the play that occurred during the game, that is, the scene type, etc.
[0208] The information on the scene type obtained in this way can be used, for example, instead of metadata. Note that the referees to be subjected to the image analysis processing may include not only the chief referee but also the assistant referees, etc.
[0209] In addition, in the above example, an example of generating a digest video DV by combining a plurality of clip sets CS has been described. However, a digest video DV may be generated from a single clip set CS. Specifically, when there is only one clip set CS to be presented to the viewer, the digest video DV may be generated to include only the single clip set CS.
[0210] <5. Computer Device> With reference to FIG. 18, the configuration of a computer device including an arithmetic processing unit that realizes the above-described information processing apparatus 1 will be described.
[0211] The CPU 71 of the computer device functions as an arithmetic processing unit that performs the various processes described above, and executes various processes according to a program stored in the ROM 72 or a non-volatile memory unit 74 such as an EEP-ROM (Electrically Erasable Programmable Read-Only Memory), or a program loaded from the storage unit 79 into the RAM 73. The RAM 73 also appropriately stores data and the like necessary for the CPU 71 to execute various processes. The CPU 71, ROM 72, RAM 73, and non-volatile memory unit 74 are interconnected via a bus 83. An input / output interface (I / F) 75 is also connected to this bus 83.
[0212] An input unit 76 composed of an operator and an operation device is connected to the input / output interface 75. For example, as the input unit 76, various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, and a remote controller are assumed. An operation of the user is detected by the input unit 76, and a signal corresponding to the input operation is interpreted by the CPU 71.
[0213] In addition, a display unit 77 composed of an LCD or an organic EL panel, etc., and an audio output unit 78 composed of a speaker, etc. are connected to the input / output interface 75 either integrally or separately. The display unit 77 is a display unit that performs various displays, and is configured by, for example, a display device provided on the housing of a computer device, a separate display device connected to the computer device, etc. The display unit 77 executes displays of images for various image processes, moving images of processing targets, etc. on the display screen based on instructions from the CPU 71. Also, based on instructions from the CPU 71, the display unit 77 performs displays such as various operation menus, icons, messages, etc., that is, displays as a GUI (Graphical User Interface).
[0214] In some cases, a storage unit 79 composed of a hard disk, a solid-state memory, etc., and a communication unit 80 composed of a modem, etc. are connected to the input / output interface 75.
[0215] The communication unit 80 performs communication processing via a transmission path such as the Internet, and communication by means of wired / wireless communication with various devices, bus communication, etc.
[0216] In addition, a drive 81 is connected to the input / output interface 75 as necessary, and a removable storage medium 82 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory is appropriately mounted. The drive 81 can read data files such as programs used for each process from the removable storage medium 82. The read data files are stored in the storage unit 79, or images and audio included in the data files are output by the display unit 77 and the audio output unit 78. Also, computer programs, etc. read from the removable storage medium 82 are installed in the storage unit 79 as necessary.
[0217] In this computer device, for example, software for the processing of the present embodiment can be installed via network communication by the communication unit 80 or via a removable storage medium 82. Alternatively, the software may be stored in advance in the ROM 72, the storage unit 79, or the like.
[0218] The CPU 71 performs a processing operation based on various programs, thereby executing the necessary information processing and communication processing as the information processing apparatus 1 provided with the arithmetic processing unit described above. Note that the information processing apparatus 1 is not limited to being configured by a single computer device as shown in FIG. 2, and may be configured by a plurality of computer devices being systematized. The plurality of computer devices may be systematized by a LAN (Local Area Network) or the like, or may be arranged remotely by a VPN (Virtual Private Network) or the like using the Internet or the like. The plurality of computer devices may include a computer device as a server group (cloud) that can be used by a cloud computing service.
[0219] <6. Summary> As described in each of the above examples, the information processing apparatus 1 includes a specifying unit 10 that specifies auxiliary information SD for generating a digest video DV based on scene-related information about a scene that occurred in an event such as a sports game. An event is, for example, an event such as a sports game or a concert. The auxiliary information SD is, for example, information used for generating the digest video DV, and is information used for determining which part of the captured video to cut out. For example, in the case of a sports game, specifically, information such as player names, scene types, and play types is used as the auxiliary information. By specifying the auxiliary information SD, it is possible to specify the time period to be cut out from the captured video, and thus it is possible to generate the digest video DV.
[0220] The scene-related information may be information including metadata distributed from another information processing device (metadata server 200). Metadata is information including the progress of events such as sports. Taking a sports game as an example, it includes information such as the time information when a specific play occurred, the names of the players related to the play, and the score information that changed as a result of the play. By specifying the auxiliary information SD based on such metadata, it is possible to more appropriately specify the time zone to be cut out from the captured video.
[0221] The scene-related information may include information related to posts by users of a social networking service (SNS). Various posts are made on the SNS according to the progress of the event. By analyzing the content of the posts on the SNS, it becomes possible to identify scenes that viewers are highly interested in. By specifying the auxiliary information SD based on the scene-related information, which is the information obtained from such an SNS, it is possible to generate a digest video DV including appropriate scenes that match the viewers' interests. Note that, as described above, the information related to posts by users of the SNS is information related to the information posted on the SNS, and includes, for example, keywords with a high appearance frequency in a recent predetermined time. This information may extract keywords based on the information posted on the SNS, or may obtain keywords presented by a service attached to the SNS, or may obtain keywords presented by a service different from the SNS.
[0222] The auxiliary information SD may be information indicating whether it has been adopted as the broadcast video VA. For example, if it is possible to specify the section adopted as the broadcast video VA in the captured video, it is possible to specify the section not adopted as the broadcast video VA. This makes it possible to generate the digest video DV so as to include the clip video CV that is not adopted as the broadcast video VA. Therefore, it becomes possible to provide the viewer with the digest video DV including new videos.
[0223] The auxiliary information SD may be keyword information. The keyword information is, for example, information such as player name information, scene type information, play type information, and equipment name. By using the keyword information, it is possible to realize the process of specifying the time zone to be cut out from the captured video with a small processing load.
[0224] The keyword information may be scene type information. For example, the clip video CV to be cut out from the captured video is determined based on the scene type information. Therefore, it is possible to generate the digest video DV including the clip video CV corresponding to a predetermined scene type.
[0225] The keyword information may be information for identifying the participants in the event. If the event is a sports game, the scene to be cut out from the captured video is determined based on the keyword information such as the names and back numbers of the players who participated in the game. Therefore, it is possible to generate a digest video DV focusing on a specific player.
[0226] The auxiliary information SD may be information used for generating the clip set CS including one or more clip videos CV obtained from a plurality of imaging devices CA that image the event. For example, when a specific play type is selected as the auxiliary information SD, a section in which the specific play type is imaged is cut out from a plurality of videos (such as the first video V1 and the second video V2) imaged by a plurality of imaging devices CA and combined, thereby generating the clip set CS related to the play type. By generating the digest video DV so as to include the clip collection CS generated in this way, one play can be viewed from different angles, and a digest video DV that is easier for the viewer to understand the play situation can be generated.
[0227] The clip collection CS is assumed to be a combination of clip videos CV that captured specific scenes in an event, and the auxiliary information SD may include information on the predetermined combination order of the clip videos CV. The clip collection CS is assumed to be a combination of a plurality of clip videos CV as partial videos captured from different angles of one play. In the generation of such a clip collection CS, by connecting videos in a predetermined order, it is possible to provide the viewer with a video that can view one play from different angles, and it is possible to reduce the processing burden for determining the connection order.
[0228] The information on the combination order of the clip videos CV may be information according to the scene type for a specific scene. That is, the predetermined order may be an appropriate order different for each type of scene. For example, when generating one clip collection CS for one field goal that occurred in an American football game, in order for the viewer to correctly recognize the situation regarding the field goal, or to enhance the sense of presence, by connecting the clip videos CV in a specific order, an appropriate clip collection CS for the field goal can be generated. The template for the specific order is defined such that videos captured from different angles, such as a video from the side, a video from the back side of the goal, a video from the front side of the goal, an overhead video, etc., are connected in a predetermined order. By fitting the videos of each imaging device CA according to this template, an appropriate clip collection CS can be automatically generated. And the processing burden for determining the connection order of the videos can be reduced. Further, the template may be different depending on the scene type.
[0229] The information processing apparatus 1 may include a clip set generation unit 11 that generates a clip set CS using the auxiliary information SD. Thereby, a series of processes from specifying the auxiliary information SD to generating the clip video CV and generating the clip set CS are executed in the information processing apparatus 1. When the information processing apparatus 1 is a single apparatus, there is no need to transmit the information necessary for generating the clip set CS from specifying the auxiliary information SD to other information processing apparatuses, and the processing load can be reduced. Note that another short video, image, or the like may be sandwiched between the clip videos CV.
[0230] The clip set generation unit 11 of the information processing apparatus 1 may generate a clip set CS by combining the clip videos CV. For example, the clip set CS is generated only by combining the clip videos CV without sandwiching another video therebetween. Thereby, it is possible to reduce the processing load required for generating the clip set CS.
[0231] The clip set CS may be a combination of clip videos CV that capture a specific scene in an event. By combining a plurality of clip videos CV obtained by cutting out videos captured from different angles for a certain scene, a clip set CS that allows the scene to be confirmed from different angles is generated. Thereby, it is possible to generate a digest video DV that is easy for the user to understand the events that occurred in each scene.
[0232] The clip set generation unit 11 of the information processing apparatus 1 may generate a clip set CS using the analysis result obtained by image analysis processing on the video obtained from the imaging apparatus CA that images an event and the auxiliary information SD. It becomes possible to identify information about the subject of the video, the type information of the scene, etc. by performing image analysis processing on the video. As a result, a clip collection CS corresponding to the auxiliary information SD can be generated, and an appropriate digest video DV can be generated.
[0233] The image analysis processing may be processing for identifying a person shown in the video. By appropriately identifying the person shown in the video through the image analysis processing, it becomes possible to identify the clip video CV to be included in the clip collection CS based on keywords such as the player's name. Therefore, the processing burden related to the selection of the clip video CV can be reduced.
[0234] The image analysis processing may be processing for identifying the type of the scene shown in the video. By appropriately identifying the type of the scene shown in the video through the image analysis processing, it becomes possible to identify the clip video CV to be included in the clip collection CS based on keywords such as the scene type. Therefore, the processing burden related to the selection of the clip video CV can be reduced.
[0235] The image analysis processing may be processing for identifying the in-point and the out-point. By identifying the in-point and the out-point through the image analysis processing, it is possible to cut out the video of an appropriate section as the clip video CV. Therefore, it is possible to generate an appropriate clip collection CS and a digest video DV.
[0236] The image analysis processing may include processing for assigning a score to each clip video CV. Depending on the duration of the clip video CV, it may not be possible to include all the clip videos CV that captured the scene in one clip collection CS. Also, there are clip videos CV that should not be included in the clip collection CS. By scoring each clip video CV, a clip collection CS can be generated by combining only appropriate clip video CVs.
[0237] The information processing method of this embodiment is a process in which a computer device identifies auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event.
[0238] The program to be executed by the information processing apparatus 1 described above can be recorded in advance in an HDD (Hard Disk Drive) as a recording medium built into a device such as a computer device, or in a ROM in a microcomputer having a CPU. Alternatively, the program can be temporarily or permanently stored (recorded) in a removable recording medium such as a flexible disk, CD-ROM (Compact Disk Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), Blu-ray Disc (registered trademark), magnetic disk, semiconductor memory, memory card, etc. Such a removable recording medium can be provided as so-called package software. In addition to installing such a program from a removable recording medium to a personal computer or the like, it can also be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet.
[0239] Note that the effects described in this specification are merely examples and are not limited, and there may be other effects.
[0240] Also, the above examples can be combined in any way, and various effects described above can be obtained even when various combinations are used.
[0241] <7. The present technology> The present technology can also adopt the following configuration. (1) An information processing apparatus including a specifying unit that specifies auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event. information processing apparatus. (2) The scene-related information is information including metadata distributed from another information processing apparatus. The information processing apparatus according to (1) above. (3) The scene-related information is assumed to include information related to posts by users of a social networking service. The information processing apparatus according to any one of (1) to (2) above. (4) The auxiliary information is information indicating whether it has been adopted as a broadcast video. The information processing apparatus according to any one of (1) to (3) above. (5) The auxiliary information is keyword information. The information processing apparatus according to any one of (1) to (4) above. (6) The keyword information is scene type information. The information processing apparatus according to (5) above. (7) The keyword information is information for identifying participants in the event. The information processing apparatus according to (5) above. (8) The auxiliary information is information used for generating a clip collection including one or more clip videos obtained from a plurality of imaging devices that image the event. The information processing apparatus according to any one of (1) to (7) above. (9) The clip collection is a combination of clip videos that image a specific scene in the event, The auxiliary information includes information on the order of combining the clip videos determined in advance. The information processing apparatus according to the above (8). (10) The information on the combination order is information corresponding to the scene type for the specific scene The information processing apparatus according to the above (9). (11) Comprising a clip set generation unit that generates the clip set using the auxiliary information The information processing apparatus according to any one of the above (8) to the above (10). (12) The clip set generation unit generates the clip set by combining the clip videos The information processing apparatus according to the above (11). (13) The clip set is obtained by combining clip videos that captured a specific scene in the event The information processing apparatus according to the above (12). (14) The clip set generation unit generates the clip set using the analysis result obtained by image analysis processing on the video obtained from the imaging device that images the event and the auxiliary information The information processing apparatus according to any one of the above (11) to the above (13). (15) The image analysis processing is processing for identifying a person shown in the video The information processing apparatus according to the above (14). (16) The image analysis processing is processing for identifying the type of the scene shown in the video The information processing apparatus according to the above (14). (17) The image analysis processing is processing for identifying in - points and out - points The information processing apparatus according to the above (14). (18) The image analysis processing includes processing for assigning a score to each clip video The information processing apparatus according to the above (14). (19) The computer device executes a process of identifying auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event. Information processing method. (20) Based on scene-related information about a scene that occurred in an event, cause an arithmetic processing device to execute a function of identifying auxiliary information for generating a digest video. Program.
Explanation of Signs
[0242] 1 Information processing device 10 Identification unit 11 Clip set generation unit 200 Metadata server (other information processing device) CA Imaging device DV Digest video SD Auxiliary information CV Clip video CS Clip set VA Broadcast video
Claims
1. A specifying unit that specifies auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event; A clip set generation unit that generates a clip set including one or more clip videos obtained from the imaging device, using the analysis result obtained by image analysis processing on the video obtained from the imaging device that captures the event and the auxiliary information. An information processing apparatus comprising: An information processing apparatus.
2. The scene-related information is information including metadata distributed from another information processing apparatus. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
3. The scene-related information is considered to include information related to posts by users of a social networking service. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
4. The auxiliary information is information indicating whether it has been adopted as a broadcast video. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
5. The auxiliary information is keyword information. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
6. The keyword information is scene type information. The information processing apparatus according to claim 5. The information processing apparatus according to claim 5.
7. The keyword information is information for identifying participants in the event. The information processing apparatus according to claim 5. The information processing apparatus according to claim 5.
8. The clip set is a combination of the clip videos that captured specific scenes in the event. The auxiliary information includes information on the predetermined order of combining the clip videos. The information processing apparatus according to claim 1. The auxiliary information includes information on the predetermined order of combining the clip videos. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
9. The information on the combining order is information corresponding to the scene type of the specific scene. The information processing apparatus according to claim 8. The information processing apparatus according to claim 8.
10. The clip set generation unit generates the clip set by combining the clip videos. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
11. The clip set is a combination of the clip videos that captured specific scenes in the event. The information processing apparatus according to claim 10. The information processing apparatus according to claim 10.
12. The image analysis processing is processing for identifying a person shown in the video. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
13. The image analysis processing is processing for identifying the type of scene shown in the video. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
14. The image analysis processing is processing for identifying in-points and out-points. The information processing apparatus according to claim 1. The information processing apparatus according to claim 1.
15. The image analysis process includes a process of assigning scores to each of the clipped videos. The information processing apparatus according to claim 1.
16. A process of identifying auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event, and A process of generating a clip set including one or more clipped videos obtained from the imaging device, using the analysis result obtained by an image analysis process on the video obtained from the imaging device that captures the event and the auxiliary information, which is executed by a computer device Information processing method.
17. A function of identifying auxiliary information for generating a digest video based on scene-related information about a scene that occurred in an event, and A function of generating a clip set including one or more clipped videos obtained from the imaging device, using the analysis result obtained by an image analysis process on the video obtained from the imaging device that captures the event and the auxiliary information, which is executed by an arithmetic processing unit Program.
Citation Information
Patent Citations
Content generation device
JP2017107404A
JPP6765558B