Information processing device, information processing method, and computer program

The integration of large-scale language models for audio and video analysis allows the extraction of user-defined highlight scenes from video content, addressing the inefficiencies in existing technologies by enhancing scene selection and reducing processing load.

WO2025173423A1PCT designated stage Publication Date: 2025-08-21SONY GROUP CORP

Patent Information

Application Number
PCT/JP2025/000085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-15
Filing Date
2025-01-06
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Users face challenges in efficiently extracting specific scenes from video content based on their interests, as existing technologies lack the ability to integrate audio and video analysis with user instructions effectively.

Method used

An information processing device and method that utilizes a large-scale language model to analyze video data, generate summary text data, and extract specific scenes based on user prompts, integrating audio and video analysis to create a highlight scene list.

Benefits of technology

Enables efficient extraction of user-defined highlight scenes from video content by associating video analysis with audio summaries, reducing processing load and improving scene selection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000085_21082025_PF_FP_ABST
    Figure JP2025000085_21082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an information processing device that extracts scenes based on instructions from a user. The information processing device is provided with: a video analysis unit that analyzes video data and extracts a plurality of specific scenes; and a summary generation unit that performs recognition processing of voice data associated with the video data to generate summary text data summarizing speech. The information processing device: associates the plurality of specific scenes with the summary text data; selects scene entries of highlight scenes indicated by user prompts from a scene list comprising a plurality of scene entries that include scene information and summary text data based on video analysis of each scene; and extracts a highlight scene list.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and computer program

[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing device, an information processing method, and a computer program for processing video information.

[0002] A huge amount of video content is available through various media, including broadcasting and streaming services, across a wide range of genres, including various sports, movies, TV dramas, music, and education. It is difficult for users to watch all of the video content they are interested in, and there are times when they want to watch only the scenes they want to see, or check the scenes they want to see first. For example, in the case of sports videos, these may include scenes showing the goals scored and lost by the team they are rooting for, or scenes of the final stages of a game where action is likely to occur.

[0003] For example, an information processing device has been proposed that includes a setting unit that sets a recognition unit that detects detected metadata, which is metadata related to a predetermined recognition target, by performing recognition processing on the recognition target, based on a sample scene, which is a scene from content specified by a user, and a generation unit that generates an extraction rule for extracting scenes from the content based on the sample scene and the detected metadata detected by the configured recognition unit by performing the recognition processing on the content as a processing target, and that extracts scenes similar to the sample scene (see Patent Document 1).

[0004] Also proposed is an information processing device that includes a control unit that performs a first control process to determine an analysis engine for scene detection from a plurality of analysis engines based on scene detection information for scene detection in an input video, and a second control process to determine an analysis engine from a plurality of analysis engines to obtain second result information related to a scene based on scene-related information about the scene obtained as first result information by the analysis engine determined in the first control process, and that determines in points and out points for the scene identified in the scene detection phase in the scene extraction phase to determine the cut-out range of video data (see Patent Document 2).

[0005] WO2023 / 233998WO2021 / 241430

[0006] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever "Robust Speech Recognition via Large-scale Weak Supervision"(arXiv:2212.04356v1 [eess.AS] 6 Dec 2022)GPT-4 Technical Report (arXiv:2303.08774v4 [cs.CL] 19 Dec 2023)

[0007] An object of the present disclosure is to provide an information processing device, an information processing method, and a computer program that perform processing to extract a specific scene from a video.

[0008] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an information processing device comprising: a video analysis unit that analyzes video data and extracts a plurality of specific scenes; and a summary generation unit that recognizes and processes audio data related to the video data to generate summary text data that summarizes speech, and associates the plurality of specific scenes with the summary text data.

[0009] The information processing device according to the first aspect further includes a prompt creation unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user, and a scene extraction unit that selects, based on the prompt text created by the prompt creation unit, a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries including scene information based on video analysis for each scene and summary text data, and extracts a highlight scene list.

[0010] The prompt creation unit includes an intention interpretation unit that interprets a task intended by the user from a user prompt entered as text by the user, and a text scene data creation unit that creates text scene data by collecting text data of scene entries that match the task intended by the user from the scene list.The scene extraction unit extracts highlight scenes indicated by the user prompt based on prompt text including the user prompt and the text scene data.

[0011] A second aspect of the present disclosure is an information processing method having: a first processing step including the steps of: a video analysis step of analyzing video data and extracting a plurality of specific scenes; and a summary generation step of recognizing and processing audio data related to the video data to generate summary text data that summarizes speech, wherein the first processing step associates the plurality of specific scenes with the summary text data; a prompt creation step of creating prompt text suitable for scene extraction processing from a user prompt entered as text by a user; and a scene extraction step of selecting, based on the prompt text created in the prompt creation step, a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries each including scene information based on video analysis for each scene and summary text data; and a second processing step of creating a highlight scene list consisting of scene entries of highlight scenes, wherein the second processing step includes the steps of: a video analysis step of analyzing video data and extracting a plurality of specific scenes; and a summary generation step of generating summary text data by recognizing and processing audio data related to the video data, wherein the first processing step associates the plurality of specific scenes with the summary text data, wherein the first processing step associates the plurality of specific scenes with the summary text data, wherein the first processing step associates the plurality of specific scenes with the summary text data, wherein the second processing step associates the plurality of specific scenes with the summary text data, wherein the first ...

[0012] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format to cause a computer to function as: a first processing unit that analyzes video data to extract a plurality of specific scenes, recognizes and processes audio data related to the video data to generate summary text data that summarizes the speech, and associates the plurality of specific scenes with the summary text data; and a second processing unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user, and, based on the prompt text, selects a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries each including scene information based on video analysis for each scene and summary text data, and creates a highlight scene list consisting of scene entries of highlight scenes.

[0013] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing device according to the first aspect of the present disclosure can be obtained.

[0014] According to the present disclosure, it is possible to provide an information processing device, an information processing method, and a computer program that integrate audio information and video analysis information related to video and process the information to extract scenes based on user instructions.

[0015] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited to these. Furthermore, the present disclosure may also bring about additional effects in addition to the effects described above.

[0016] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description based on the embodiments and accompanying drawings.

[0017] FIG. 1 is a diagram showing an example configuration of a scene extraction system 100 to which the present disclosure is applied. FIG. 2 is a diagram schematically showing the data structure of a scene list (after video analysis). FIG. 3 is a diagram schematically showing the data structure of a scene list (after speech recognition). FIG. 4 is a diagram schematically showing the data structure of a scene list (after summary generation). FIG. 5 is a diagram showing an example configuration of a highlight scene list screen. FIG. 6 is a diagram generally showing an implementation example of the scene extraction system 100. FIG. 7 is a diagram showing a specific implementation example of the scene extraction system 100. FIG. 8 is a diagram showing a specific implementation example of the scene extraction system 100. FIG. 9 is a diagram showing a specific configuration example of a pre-processing device 601. FIG. 10 is a diagram showing the data structure of a scene entry. FIG. 11 is a diagram showing an overview of processing performed by a video analysis unit 901. FIG. 12 is a diagram showing an example configuration of a video analysis unit 901 implemented using AI. FIG. 13 is a diagram showing an analysis engine 1203 and example parameters used for scene detection. FIG. 14 is a diagram showing an example of the analysis engine 1203 and parameters used for scene extraction. FIG. 15 is a diagram showing an example of the analysis engine 1203 and parameters used for detailed description. FIG. 16 is a flowchart showing the processing procedure executed by the video analysis unit. FIG. 17 is a diagram showing a specific configuration example of the demonstration device 602. FIG. 18 is a diagram showing the data structure of input data to the intention interpretation unit 1701 in accordance with a template instruction. FIG. 19 is a diagram showing another example of the data structure of input data to the intention interpretation unit 1701 in accordance with a template instruction. FIG. 20 is a diagram showing an example of pseudo-programming code for creating text scene data. FIG. 21 is a diagram showing an example of a data format of text scene data. FIG. 22 is a diagram showing an example of a prompt input to the scene extraction unit 1703. FIG. 23 is a diagram showing an example of a prompt input to the scene description creation unit 1704. FIG. 24 is a diagram showing an example of a highlight video generation screen display. FIG. 25 is a diagram showing an example of a highlight video generation screen display. FIG. 26 is a diagram showing an example of a highlight video generation screen display. FIG. 27 is a diagram showing a display example of the highlight video generation screen.Fig. 28 is a flowchart showing the processing operation of the pre-processing device 601. Fig. 29 is a flowchart showing the processing operation of the demonstration device 602. Fig. 30 is a diagram showing an example of the hardware configuration of the information processing device 2000.

[0018] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.

[0019] A. Overview B. Implementation example B-1. Pre-processing device B-2. Video analysis B-2-1. Overview of video analysis B-2-2. AI engine for video analysis B-2-3. AI engine used for each process B-2-4. Processing operation of the video analysis unit B-3. Demonstration device B-4. Generation of highlight video C. Configuration example of information processing device

[0020] A. Overview Technology for extracting specific scenes from video based on user instructions is already known (see, for example, Patent Document 1). The scene extraction technology according to the present disclosure is characterized in that it extracts scenes based on user instructions by integrating audio information and video analysis information using a large-scale language model (LLM). The large-scale language model is an AI (Artificial Intelligence) model trained on a massive amount of text data.

[0021] According to the present disclosure, video data is analyzed to obtain information for each scene, and a large-scale language model is used to generate a summary of the utterances for each scene from text data of all utterances obtained by speech recognition of audio data related to the video data.The video analysis information for each scene obtained from the original video data is then integrated with the summary to extract scenes based on user instructions.A large-scale language model can also be used for scene extraction.

[0022] Furthermore, the scene extraction technology according to the present disclosure generates metadata describing the extracted scenes and associates them with each scene, and is characterized by its ability to generate metadata that reflects user instructions. Large-scale language models are also utilized to generate the metadata for each scene.

[0023] 1 schematically shows an example configuration of a scene extraction system 100 to which the present disclosure is applied. The illustrated scene extraction system 100 includes a video analysis unit 101, a voice recognition unit 102, a summary generation unit 103, a prompt creation unit 104, a scene extraction unit 105, and a user input unit 106. Video data associated with audio data, such as broadcast video or streaming video, is supplied to the scene extraction system 100.

[0024] The video analysis unit 101 performs video analysis on the input video data using AI for video analysis (described below). Specifically, the video analysis unit 101 sequentially performs the following processes: scene detection, which detects the location where a specific scene (event) occurred in the video data; scene extraction, which determines the range of time information related to the detected scene; and detailed description, which identifies scene information describing the detailed content of the extracted scene. The video analysis unit 101 then outputs a scene list listing the time information (in-points and out-points) and scene information for each scene. For example, if the video to be processed is a sports video, the video analysis unit 101 can detect the in-points and out-points of specific scenes, such as scoring scenes, and generate scene information using the technology disclosed in Patent Document 2, for example.

[0025] 2 shows a schematic diagram of the data structure of the scene list output from the video analysis unit 101. The scene list includes an entry for each scene (scene entry), and each scene entry describes various information based on the results of video analysis, such as a scene ID that uniquely identifies the scene, the scene start frame position, the scene time interval, the scene start time and end time, and scene information describing detailed content of the scene (for example, in the case of a sports video, the score of the team's team and the score of the opposing team).

[0026] The speech recognition unit 102 performs speech recognition processing on the speech data related to the video data and converts all spoken content into text data. For example, if the video to be processed is a sports video, the speech recognition unit 102 performs speech recognition on all commentary and outputs text data of all spoken content. The speech recognition unit 102 uses a trained model that converts speech into text (Speech2Text), such as Whisper (see Non-Patent Document 1). The speech recognition unit 102 may perform speech recognition processing using an existing library that is stored, shared, or made public, for example, through a source code management service.

[0027] The output of the video analysis unit 101 and the output of the speech recognition unit 102 are associated on the time axis, and a scene list listing information for each scene is created, and text data of all utterances in each scene is generated by speech recognition. For example, as shown in Figure 3, each scene entry in the scene list may include text data of all utterances in the corresponding scene.

[0028] The summary generation unit 103 generates summary text data by summarizing the text data of all utterances output from the speech recognition unit 102. For example, if the video to be processed is a sports video, the summary generation unit 103 generates a summary commentary by summarizing the text data of all game commentary output from the speech recognition unit 102. Although the text data of the entire game commentary may be used as is in the scene extraction process at the subsequent stage, generating a summary commentary makes it possible to compress the amount of data and reduce the load of processing the text data in the scene extraction process at the subsequent stage.

[0029] The summary generation unit 103 can generate text data of a summary for long text data using a large-scale language model such as GPT (Generative Pre-trained Transformer)-4 (see Non-Patent Document 2). The summary generation unit 103 may perform the process of generating a summary using an existing library that is saved, shared, or made public by, for example, a source code management service.

[0030] When text data of all utterances in a scene corresponding to each scene entry in the scene list is stored as described above (see FIG. 3 ), the summary generation unit 103 generates summary text data that summarizes the text data of all utterances for each scene. Therefore, when the processing of the summary generation unit 103 is completed, summary text data associated with each scene in the scene list has been generated. For example, as shown in FIG. 4 , each scene entry in the scene list may further include summary text data that summarizes the text data of all utterances in the corresponding scene, in addition to the text data of all utterances obtained by speech recognition. It should be fully understood that the present disclosure makes it possible to create a scene list that includes both information derived from video analysis and information derived from speech recognition.

[0031] The user input unit 106 receives a user prompt that indicates what kind of scene the user wants to see from the original video data. The user input unit 106 is composed of a keyboard or the like, and it is assumed that the user prompt is text input. However, the user input unit 106 may also input text data obtained by voice recognition of the user's speech as the user prompt.

[0032] The prompt creation unit 104 takes the user prompt input to the user input unit 106, i.e., the user's instruction as to what scene the user wants from the original video data, and creates prompt text suitable for input to the subsequent scene extraction unit 105. The user may input the instruction as text using a keyboard or by voice. The user may also input the instruction using an input device other than a keyboard (such as a mouse or touch panel).

[0033] The scene extraction unit 105 inputs the prompt text created by the prompt creation unit 104 and extracts scenes that match the user prompt, i.e., highlight scenes, from a scene list (see FIG. 4) containing summary text data. The scene extraction unit 105 also generates metadata for each extracted scene based on the user prompt. The scene metadata here refers to, for example, a description of the scene that matches the user prompt. The summary text data for each scene generated by the summary generation unit 103 is a description of the scene, but does not necessarily match the user's intent. In contrast, the scene extraction unit 105 generates metadata including a description of the scene that matches the user prompt, based on the user prompt. The scene extraction unit 105 can be implemented using a large-scale language model such as GPT-4 (see Non-Patent Document 2). The scene extraction unit 105 may also extract highlight scenes from the scene list using an existing library that is stored, shared, or made public via a source code management service, for example.

[0034] The scene extraction system 100 can obtain, as an output from the scene extraction unit 105, a highlight scene list that is a collection of scene lists of each highlight scene intended by the user from the video to be processed.

[0035] The highlight scene list can be used in a variety of ways. Because the highlight scene list contains time information (in-point and out-point) and metadata for each highlight scene, a highlight scene list screen can be presented to the user using thumbnail images of the highlight scenes (representative frames such as the in-point image frame, for example) and metadata. Furthermore, the highlight video selected by the user from the list screen can be played back and displayed.

[0036] FIG. 5 shows an example of the configuration of a highlight scene list screen. In the example shown in FIG. 5, thumbnails #1 to #9 of a total of nine highlight scenes are displayed together with metadata #1 to #9, respectively. If nine or more highlight scenes are extracted, a scroll bar or the like may be provided so that thumbnails of all the highlight scenes can be viewed by scrolling the screen. Furthermore, the number of highlight scenes displayed at one time on the list screen is arbitrary. The user can confirm whether or not they wish to view each highlight scene based on the thumbnail image and the description content of the metadata. Then, when the user finds a highlight scene they wish to view, they can indicate their intention to select it by, for example, touching the desired thumbnail on the list screen. In response to the selection of a thumbnail, the display may switch to a full-screen display of a playback image of that highlight scene.

[0037] B. Implementation Example The scene extraction system 100 may be configured to perform the processes of creating a scene list from video and extracting highlight scenes from the scene list in real time. However, because the process of creating a scene list from video is a heavy load, it is more realistic to perform the process of creating the scene list as a pre-process and then extract scenes specified by the user using the created scene list.

[0038] 6 shows a schematic diagram of an example implementation of the scene extraction system 100. In the example implementation shown in FIG. 6, the scene extraction system 100 is composed of a pre-processing device 601 and a demonstration processing device 602.

[0039] The pre-processing device 601 is mainly composed of the video analysis unit 101, the voice recognition unit 102, and the summary generation unit 103 from the components shown in Figure 1, and generates, as pre-processing, from the video to be processed, a scene list that lists time information (in-point and out-point) and scene information for each scene, and summary text data corresponding to each scene.

[0040] 1, the demonstration device 602 is mainly composed of the prompt creation unit 104 and the scene extraction unit 105, and extracts scenes that match the user prompt, i.e., highlight scenes, from the scene list and summary text data for each scene based on the user's instructions regarding the type of scene the user wants. The demonstration device 602 can present the user with a list screen of highlight scenes and play back and display the highlight scene the user selects from the list screen.

[0041] The pre-processing device 601 and the demonstration device 602 may be configured on a single information processing device (e.g., a personal computer (PC)), or may be physically separate devices. In the latter case, as shown in FIG. 7 , the pre-processing device 601 may be located on a cloud, and the demonstration device 602 may be an information terminal operated by a user. Also, as shown in FIG. 8 , one pre-processing device 601 may provide services to the information terminals of multiple users (i.e., a scene list for a certain video content created by one pre-processing device 601 may be shared by the information terminals of multiple users who watch the same video content).

[0042] B-1. Pre-processing device Fig. 9 shows a specific example of the configuration of the pre-processing device 601. The pre-processing device 601 includes a video analysis unit 901, a voice recognition unit 902, and a summary generation unit 903. These components 901 to 903 are almost the same as the components of the same names shown in Fig. 1.

[0043] The video analysis unit 901 sequentially executes the following processes: scene detection, which detects the location where a specific scene (event) occurred from the input video data; scene extraction, which determines the range of time information related to the detected scene; and detailed description, which specifies scene information that describes the detailed content of the extracted scene, and outputs a scene list 921 (see, for example, Figure 2) that lists the time information (in point and out point) and scene information for each scene.

[0044] The speech recognition unit 902 performs speech recognition processing on speech data related to video data and converts all spoken content into text data. The speech recognition unit 102 uses a trained model that converts speech into text (Speech2Text), such as Whisper (see Non-Patent Document 1). The text data output from the speech recognition unit 902 is associated with each scene in the scene list created by the video analysis unit 901, and a scene list 922 (see, for example, FIG. 3) is created that stores text data of all spoken utterances, which are the speech recognition results for the scene corresponding to each scene entry.

[0045] Each scene entry in the scene list is formatted into an appropriate prompt template according to template instructions 911 and then input to summary generator 903 .

[0046] The summary generation unit 903 uses a large-scale language model such as GPT-4 (see Non-Patent Document 2) to generate summary text data, which is a summary of the text data of all utterances, which is the speech recognition result of each scene entry in the scene list, and keywords. The summary generation unit 903 also generates a pre-description for each scene. However, the pre-description generated by the summary generation unit 903 is not a description that reflects a user's instructions (i.e., the pre-description generated by the summary generation unit 903 is essentially different from the scene description generated by the scene description creation unit 1704, which will be described later, in response to a user's instructions). Furthermore, since the pre-description is not directly related to the present disclosure, detailed descriptions of the method for generating the pre-description and the content of the pre-description will be omitted.

[0047] In this way, the pre-processing device 601 can create a scene list 923 for the input video data, which consists of scene entries that store scene information based on the video analysis results for each scene, text data of all utterances based on the voice recognition results, summary text data, and pre-explanation.

[0048] Figure 10 shows a schematic diagram of the data structure of a scene entry for one scene. Note that Figure 10 shows an example of a scene extracted from an American football game. Scene entries can be written in any language format. In the example shown in Figure 10, one data item is written per line.

[0049] The line indicated by the reference numeral 1001 indicates that the processing target is a video. The line indicated by the reference numeral 1002 indicates the name of the video analysis AI (AI_Sports_Scene_Extraction) used to extract play scenes from sports video. The line indicated by the reference numeral 1003 indicates the scene ID assigned to the scene entry.

[0050] The line section indicated by the reference numeral 1010 describes scene information based on the analysis results by the video analysis unit 901. For example, the reliability of the video analysis (confidence), the scene start frame position (startFrame), the frame interval of the scene (duration), the scene start time (start) and the scene end time (end), ..., the game clock (gameClock), the teammate's score (score1), the opposing team's score (score2), the teammate's score change (scoreChange1), the opposing team's score change (scoreChange2), ..., and the like are stored as scene information.

[0051] The line section indicated by the reference numeral 1020 describes information based on audio information, i.e., information on text data obtained by speech recognition of audio data related to video data. Of this line section, the "all utterance text data (relevant_comment)" line stores text data of all utterances, which are the speech recognition results of the scene by the speech recognition unit 902. The "summary text data (summarized_comment)" line stores summarized text data generated by the summary generation unit 903 for the text data of all utterances stored in the "all utterance text data (relevant_comment)" line. The "scene context description (pre-description)" line stores a context description of the scene.

[0052] FIG. 28 shows the processing operation of the pre-processing device 601 in the form of a flowchart.

[0053] The video analysis unit 901 sequentially executes the following processes: scene detection, which detects the location where a specific scene occurred from the input video data; scene extraction, which determines the range of time information related to the detected scene; and detailed description, which specifies scene information that describes the detailed content of the extracted scene, and outputs a scene list 921 that lists the time information and scene information for each scene (step S2801).

[0054] Meanwhile, the voice recognition unit 902 performs voice recognition processing on the voice data related to the video data and converts all the spoken content into text data (step S2802). However, the video data analysis processing by the video analysis unit 901 and the voice recognition processing by the voice recognition unit 902 may be performed simultaneously in parallel, or the voice recognition processing by the voice recognition unit 902 may be performed first.

[0055] The text data output from the voice recognition unit 902 is associated with each scene in the scene list created by the video analysis unit 901, and a scene list 922 (see, for example, Figure 3) is created that stores the text data of all utterances, which are the voice recognition results for the scene corresponding to each scene entry.

[0056] Each scene entry in the scene list created by the video analysis unit 901 is formatted into an appropriate prompt template according to template instructions 911 and then input to the summary generation unit 903 .

[0057] The summary generation unit 903 then uses a large-scale language model such as GPT-4 (see Non-Patent Document 2) to summarize the text data of all utterances for each scene, and generates summarized text data and keywords (step S2803). The summary generation unit 903 also generates a pre-description for each scene.

[0058] In this way, the pre-processing device 601 creates and outputs a scene list 923 consisting of scene entries that store scene information for each scene, text data of all utterances based on the speech recognition results, summary text data, and pre-explanation for the input video data (step S2804).

[0059] The created scene list with summary text data is output to the demonstration device 602. The demonstration device 602 can use the scene list supplied from the pre-processing device 601 to extract highlight scenes desired by the user and play back highlight footage. Of course, the pre-processing device 601 may output the scene list to a device other than the demonstration device 602.

[0060] B-2. Video Analysis B-2-1. Overview of Video Analysis The processing in the video analysis unit 901 can be realized using, for example, the technology disclosed in Patent Document 2. Fig. 11 shows an overview of the processing executed by the video analysis unit 901. The video analysis unit 901 receives as input video data 1111 to be processed and parameters 1112 that describe various specifications related to the processing, and sequentially executes the processing of scene detection 1101, scene extraction 1102, and detailed description 1103.

[0061] Scene detection 1101 is a process for identifying a location where a specific scene (event) occurred from video data 1111 and outputting the scene occurrence time as time information. For example, in the case of video of an American football (hereinafter simply referred to as "American football") game, scenes to be detected include "touchdown," "field goal," "long run," and "quarterback (QB) sack." Scene detection may detect one type of scene from these, or multiple types of scenes. For example, when detecting a "touchdown" scene, the time of the moment of the touchdown is detected as the scene occurrence time, and the number of detected touchdowns is output. The scenes to be detected may be set in advance, or may be set, added, or changed by the user at any time.

[0062] Scene extraction 1102 is a process of determining the range of time information related to the identified scene with respect to the scene occurrence time identified in scene detection 1101. Specifically, scene extraction determines the in-point and out-point for the identified scene, thereby determining the range of video data to be cut out.

[0063] The detailed description 1103 is a process for specifying scene information to be extracted from the video data of a scene for which the in-point and out-point have been specified. For example, in the case of a touchdown scene, the detailed description may specify the name of the player who made the touchdown and the score, and may also specify the name of the player who made the pass before that.

[0064] In this way, the video analysis unit 901 identifies the time when a scene occurs in scene detection, identifies the width of the time information related to the scene (the range to be cut out) in scene extraction, performs processing to extract scene information for each scene in detailed description, and outputs a scene list 1113 that lists the scene time information (in point and out point) and scene information for each scene.

[0065] The parameters 1112 input to the video analysis unit 901 simultaneously with the video data may, for example, specify parameters (scene detection information) for detecting a desired scene in scene detection, parameters (scene-related information) for determining the scene cutout range in a desired manner in scene extraction, parameters used in both scene detection and scene extraction (scene detection information and scene-related information), and various setting items for analyzing the video data.

[0066] B-2-2. AI Engine for Video Analysis The video analysis unit 901 uses multiple AI engines for video analysis to perform the above-mentioned processes of scene detection 1101, scene extraction 1102, and detailed description 1103. Fig. 12 shows an example configuration of the video analysis unit 901 realized using AI. In the example shown in Fig. 12, the video analysis unit 901 includes an interface 1201 for inputting and outputting data, an AI process manager 1202, and multiple analysis engines 1203 as AI engines.

[0067] The analysis engine 1203 is configured with, for example, a deep neural network and dictionary data (DIC database), and functions as a recognizer or extractor with improved processing accuracy through AI machine learning (deep learning, etc.), and performs various types of recognition processing, extraction processing, and judgment processing.

[0068] The interface 1201 passes necessary parameters input from the outside so that the analysis engine 1203 can execute various analysis processes, for example, to the AI ​​process manager 1202. The interface 1201 also receives analysis results obtained using each analysis engine 1203 from the AI ​​process manager 1202 and passes them to the outside (for example, the demonstration device 602).

[0069] For each of the processes of scene detection, scene extraction, and detailed description, the AI ​​process manager 1202 selects one or more optimal analysis engines 1203 according to the processing content from among the multiple analysis engines 1203. The analysis engines 1203 selected for each of the processes of scene detection, scene extraction, and detailed description may be implemented in an information processing device that constitutes the pre-processing device 601, or may be provided outside the information processing device that constitutes the pre-processing device 601.

[0070] Parameters may be set for each analysis engine 1203. For example, when processing video data of an American football game, parameters for setting "touchdown," "field goal," and "quarterback sack" as scenes to be detected are set in the analysis engine 1203 used for scene detection. The parameters may be provided in association with the video data, or the user may set the parameters for the pre-processing device 601. In the latter case, the pre-processing device 601 may provide the user with a UI (User Interface) for setting the necessary parameters.

[0071] Below is an example of a plurality of analysis engines that can be used by the video analysis unit 901 as the analysis engine 1203, categorized by purpose.

[0072] Audio subtitling engine: An analysis engine that analyzes audio data and extracts text data, and is assigned parameters such as a language specification. Object recognition engine: An analysis engine that recognizes objects such as people, animals, and objects that appear in video data, and is assigned parameters such as a type of object. Character recognition engine: An analysis engine that detects characters that appear in video data or characters that are superimposed on video data. Parameters assigned to the character recognition engine include, for example, a parameter that specifies the language type (English, Japanese, etc.). Face recognition engine: An analysis engine that recognizes the facial areas of people that appear in video data, and is assigned parameters such as a specific person or a specific gender. Sports data analysis engine: A sports data analysis engine is, for example, an analysis engine that analyzes STATS (statistics) information provided outside the video analysis unit 901, and is assigned parameters such as information that identifies a player or a scene. STATS information may include, for example, text information that describes the progress of a game or numerical information that describes a player's performance over a specific period. Highlight generation engine: An analysis engine that extracts highlight scenes from a match, and is assigned parameters such as parameters that specify the length of a highlight video and parameters that identify a player in order to generate a highlight video of a specific player. The highlight generation engine may generate highlight videos by working with an excitement detection engine (described below), for example. The highlight generation engine may also extract time information for generating highlight videos without actually generating a highlight video. Time-shortened version generation engine: An analysis engine that generates a time-shortened version of video data, and performs processing such as identifying scenes where the game is interrupted in order to eliminate scenes where the game is interrupted. The time-shortened version generation engine is assigned parameters such as parameters that specify the length of the time-shortened version of video data and parameters that identify unnecessary scenes. Emotion recognition engine: An analysis engine that estimates emotions by analyzing the shapes of each facial feature of a person captured in video data, and is assigned parameters such as parameters that identify a player and a team.The emotion recognition engine may analyze emotions by incorporating the functions of a face recognition engine, or may analyze the emotions of a subject by working in conjunction with the face recognition engine. Excitement detection engine: An analysis engine that analyzes whether a scene is exciting, and is assigned a volume parameter as a threshold for determining whether a scene is exciting. Camera angle recognition engine: An analysis engine that identifies camera angles and changes in camera angles, and also detects the size of the subject in the image, i.e., whether a close-up of a player or the entire game is being captured. The camera angle recognition engine, for example, acquires a person's skeletal information through image recognition, and if the acquired skeletal information is of a certain size and only includes the upper body, determines as an analysis result that the image is a bust-up. The camera angle recognition engine is assigned, for example, a parameter specifying a wide shot or a parameter specifying the angle of view. Pan-tilt recognition engine: An analysis engine that identifies the panning and tilting of the camera, and is assigned parameters such as a parameter specifying either pan or tilt, and a parameter specifying a change in angle.

[0073] These various analysis engines may be provided for each type of sport. For example, an object recognition engine for American football that specializes in recognizing American football players, a football ball, goalposts, etc., and an object recognition engine for soccer that specializes in recognizing soccer players, a soccer ball, goals, etc. may be provided.

[0074] B-2-3. AI engines used for each process Next, specific examples of the analysis engines and parameters used for each process of scene detection, scene extraction, and detailed description will be explained.

[0075] 13 shows examples of analysis engines 1203 and parameters used for scene detection. The analysis engines 1203 used for scene detection can be classified according to the processing content, such as an external data analysis engine 1301, an object detection engine 1302, a voice analysis engine 1303, a camerawork analysis engine 1304, and a character recognition engine 1305. However, analysis engines 1203 other than those listed above may also be used for scene detection.

[0076] The external data analysis engine 1301 analyzes information acquired from outside the video analysis unit 901 and performs processing that contributes to identifying the time when a scene occurred, such as the sports data analysis engine (described above) that analyzes STATS information. The external data analysis engine 1301 is assigned parameters such as STATS information for each sport type.

[0077] The object detection engine 1302 performs image analysis to recognize objects such as people, animals, and objects that appear in video data, and is an example of an object recognition engine (described above). Specifically, the object detection engine 1302 identifies balls, people, and sports equipment that appear in the video. The object detection engine 1302 may also identify superimposed images such as subtitles, character images, and 3D images that are superimposed on images. An excitement detection engine (described above) that determines the excitement level in the venue and an emotion recognition engine (described above) that identifies the faces of players and analyzes their expressions can also be considered object detection engines 1302. The object detection engine 1302 is assigned parameters such as dictionaries that differ for each type of sport, such as American football or soccer. In other words, by assigning a different dictionary as a parameter for each type of sport, the object detection engine 1302 functions as an object detection engine for a specific sport.

[0078] The voice analysis engine 1303 performs processing to analyze voice data, and is exemplified by the voice subtitling engine (described above). The excitement detection engine (described above), which detects excitement in a venue by analyzing changes in the volume of the voice, can also be considered the voice analysis engine 1303. The voice analysis engine 1303 is provided with parameters such as a vocabulary list for each type of sport, such as American football or soccer. In other words, by providing a different vocabulary list for each type of sport as a parameter, the voice analysis engine 1303 functions as a voice analysis engine for a specific sport.

[0079] The camerawork analysis engine 1304 performs processing to identify the angle of view and analyze changes, and corresponds to, for example, the camera angle recognition engine (described above) and the pan-tilt recognition engine (described above). It can also be said that the excitement detection engine (described above) utilizes the camerawork analysis engine 1304. The camerawork analysis engine 1304 is provided with parameters such as a dictionary for each type of sport, such as American football or soccer.

[0080] The character recognition engine 1305 performs character recognition processing on subtitles and score displays superimposed on an image. For example, the character recognition engine 1305 detects a touchdown scene by performing character recognition processing on an image on which the decorative word "touchdown" is superimposed. The character recognition engine 1305 can also identify the movement of the score superimposed as subtitles through character recognition and estimate the scene that occurred from the change in the score. The character recognition engine 1305 can also function as an excitement detection engine (described above) that estimates the level of excitement by performing scene detection based on character recognition processing. The character recognition engine 1305 is assigned parameters such as a dictionary for each sport type, such as American football or soccer. Parameters may be assigned based on scoring rules for each sport type and on scene types in addition to the sport type. For example, in American football, different parameters may be assigned for detecting touchdown scenes and field goal scenes.

[0081] 14 shows examples of analysis engines 1203 and parameters used for scene extraction. The analysis engines 1203 used for scene extraction can be classified according to the processing content into a camera switching analysis engine 1401, an exciting section analysis engine 1402, a fixed-second cutout engine 1403, a camerawork analysis engine 1404, and an object detection engine 1405. However, analysis engines 1203 other than those listed above may also be used for scene extraction.

[0082] The camera switching analysis engine 1401 analyzes whether or not the image capture device is being switched by the switcher, the switching timing, etc. The camera switching analysis engine 1401 may function to detect the end of an exciting section by detecting the frequency of switching, or may detect the timing of switching and be used to determine the range of scene extraction by the highlight generation engine (described above). The camera switching analysis engine 1401 is assigned parameters such as thresholds for determining switching, for example.

[0083] The excitement section analysis engine 1402 performs processing to determine the extraction range of an excitement scene and also functions as an excitement detection engine (described above). The excitement section analysis engine 1402 is assigned parameters such as a threshold value for determining whether or not an excitement occurs.

[0084] The fixed-second cutout engine 1403 performs processing to cut out a range before and after a specified time, and is used by the highlight generation engine (described above) that extracts scenes that are highlights of a match, the time-saving version generation engine (described above) that generates time-saving versions of video data (such as processing to identify game interruption scenes in order to eliminate scenes where the match is interrupted), etc. The fixed-second cutout engine 1403 is assigned parameters such as the time when a scene occurred and the number of seconds to cut out.

[0085] The camerawork analysis engine 1404 performs processing to identify the angle of view and analyze changes, similar to those used in scene detection, and corresponds to the camera angle recognition engine (described above) and pan / tilt recognition engine (described above). The camerawork analysis engine 1404 is provided with parameters such as dictionaries for each type of sport, such as American football or soccer. The results of the camerawork analysis are also used by the highlight generation engine (described above) that generates highlight videos and the time-saving generation engine (described above) that generates time-saving videos.

[0086] The object detection engine 1405 used in scene extraction is the same as that used in scene detection, and a description thereof will be omitted here.

[0087] 15 shows examples of analysis engines 1203 and parameters used in detailed description. The analysis engines 1203 used in detailed description can be classified according to the processing content into an external data analysis engine 1501, a uniform number recognition engine 1502, a voice analysis engine 1503, a character recognition engine 1504, etc. However, analysis engines 1203 other than those listed above may also be used in detailed description.

[0088] The external data analysis engine 1501 is the same as that used in scene detection, and a description thereof will be omitted here.

[0089] The uniform number recognition engine 1502 performs image analysis to recognize the uniform numbers (uniform numbers) of players appearing in an image. This process is performed, for example, to identify key players in important plays. The uniform number recognition engine is used in the highlight generation engine (described above) and the time-saving version generation engine (described above). The uniform number recognition engine 1502 is provided with parameters such as a dictionary for each sport, for example, American football or soccer.

[0090] The voice analysis engine 1503 performs processing to analyze voice data and is used, for example, to identify key players in important plays. Therefore, it is used by the highlight generation engine (described above) and the time-saving version generation engine (described above). The voice analysis engine 1503 is assigned parameters such as a vocabulary list for American football or soccer, and may also be assigned parameters for identifying the language.

[0091] The character recognition engine 1504 performs processing to detect character information, including superimposed subtitles, from an image, and is used, for example, to extract detailed information about the content of a play that has been made. For example, in the case of video of an American football game, yardage notations are superimposed on the field at regular intervals from the end line, and the character recognition engine 1504 is used to calculate, from these yardage notations, how far the ball was advanced in a play, i.e., the number of yards gained in a long run. The character recognition engine 1504 is provided with parameters such as dictionaries for each type of sport, such as American football or soccer, and scoring rules for each type of sport.

[0092] B-2-4. Processing Operation of Video Analysis Unit The video analysis unit 901 uses the AI ​​engine described above to sequentially execute the following processes: scene detection, which detects the location where a specific scene occurred from the video data; scene extraction, which determines the range of time information related to the detected scene; and detailed description, which specifies scene information that describes the detailed contents of the extracted scene, to create a scene list made up of scene entries (see FIG. 10, for example) that summarize the time information and scene information for each scene.

[0093] 16 is a flowchart showing the processing procedure executed by the video analysis unit 901. The processing procedure executed by the video analysis unit 901 to create a scene list will be described below with reference to FIG.

[0094] When the AI ​​process manager 1202 acquires STATS information (step S1601), it performs scene detection processing on the acquired STATS information, i.e., detects a target scene from the input video based on the acquired STATS information (step S1602).Then, the AI ​​process manager 1202 determines whether the target scene has been detected from the input video in the scene detection (step S1603).

[0095] If it is determined that the scene to be detected has been detected in the scene detection (Yes in step S1603), the process proceeds to step S1604, where the AI ​​process manager 1202 selects an analysis engine and parameters to be used in scene extraction according to the scene type. That is, the AI ​​process manager 1202 selects an analysis engine suitable for detecting in-points and out-points according to the scene type, and assigns parameters to the selected analysis engine.

[0096] Next, the AI ​​process manager 1202 performs an in-point detection process using the selected analysis engine and parameters to extract time information (step S1205), and then performs an out-point detection process to extract time information (step S1206).

[0097] Next, the AI ​​process manager 1202 outputs the time information of the IN point and OUT point as the scene start position (start) and scene end position (end) of the scene (i.e., writes it to the scene entry of the scene) (step S1607). In step S1607, information such as the scene type and scene occurrence time of the scene acquired in the scene detection executed in the processing of step S1602 is also output as data of the scene entry of the scene.

[0098] On the other hand, if it is determined that the scene to be detected could not be detected in the scene detection (No in step S1603), the AI ​​process manager 1202 determines whether detailed information related to the scene to be detected is included (step S1608). That is, the AI ​​process manager 1202 determines whether detailed information, which is information that supplements the already detected scene, can be acquired from the acquired STATS information. The detailed information is, for example, scene information indicating the uniform number information and player name information of the players who played an active role in the scene, and the content of the play (such as the number of yards gained).

[0099] If it is determined that detailed information can be obtained from the scene (Yes in step S1608), the AI ​​process manager 1202 proceeds to step S1607 and performs a process of outputting the obtained detailed information, scene type, etc. as scene information (i.e., a process of writing them to the scene entry of the scene).

[0100] Also, if it is determined that detailed information cannot be obtained from the scene (No in step S1608), the AI ​​process manager 1202 terminates this processing.

[0101] The AI ​​process manager 1202 appropriately executes the process shown in Fig. 16 each time it acquires STATS information. For example, in sports video, the process shown in Fig. 16 is executed each time it acquires STATS information as information indicating the content of one play. By repeating the process shown in Fig. 16, it is possible to identify the time when the scene to be detected occurred, determine the extraction range, and extract detailed information from video data of a single game, thereby creating a scene list consisting of scene entries for each scene.

[0102] In this mode, for example, when STATS information is updated in parallel with the progress of a match, the AI ​​process manager 1202 acquires the STATS information in step S1601 each time the STATS information is updated and executes the subsequent processes, and as the match progresses, scene entries are created by detecting the occurrence time, in points, and out points of the scenes to be detected and extracting detailed information. For example, in parallel with the progress of the match, or after the match video has been recorded, the video analysis unit 901 creates a scene list made up of scene entries that store scene information necessary for editing highlight scenes.

[0103] B-3. ​​Demonstration Device Fig. 17 shows a specific example of the configuration of the demonstration device 602. The demonstration device 602 includes an intention interpretation unit 1701, a text scene data creation unit 1702, a scene extraction unit 1703, and a scene description creation unit 1704. The intention interpretation unit 1701 and the text scene data creation unit 1702 correspond to the components included in the prompt creation unit 104 shown in Fig. 1. The scene extraction unit 1703 and the scene description creation unit 1704 correspond to the components included in the scene extraction unit 105 shown in Fig. 1.

[0104] The intention interpretation unit 1701 interprets the intention of a user prompt (instruction of a desired scene) 1721 entered as text into the user input unit 106 (not shown in FIG. 17 ) using a keyboard or the like. The text input by the user specifying a desired scene is shaped into the format of an appropriate prompt template according to template instructions 1711, and then input to the intention interpretation unit 1701. The intention interpretation unit 1701 determines the task intended by the user from the user prompt entered as text. The task to be determined differs depending on the type of video to be processed (the type of sport in the case of sports video). For example, when video of an American football game is to be processed, the intention interpretation unit 1701 indicates the presence or absence of an intention, such as "Does the user intend to extract a scoring scene?" or "Does the user intend to specify a phase of the game?", using output labels such as "1st quarter," "2nd quarter," ..., and "No."

[0105] The intention interpretation unit 1701 can be realized using a large-scale language model such as GPT-4 (see Non-Patent Document 2). The intention interpretation unit 1701 may perform user intention interpretation processing using an existing library that is saved, shared, and made public by, for example, a source code management service.

[0106] 18 schematically shows the data structure of input data to the intent interpretation unit 1701 in accordance with the template instructions. The input shown in the figure, in the first half indicated by reference numeral 1801, instructs the unit to interpret the user prompt as "Do you intend to extract a scoring scene?" and to output "YES" if the user prompt intends to extract a scoring scene, and to output "NO" in other cases (for example, if the user prompt intends to specify a phase of the game). The second half indicated by reference numeral 1802 (the part between "===User prompt starts===" and "===User prompt ends===") stores the user prompt entered as text.

[0107] 19 shows another example of the data structure of input data to the intent interpretation unit 1701 according to the template instruction. The input shown in the figure, in the first half indicated by the reference numeral 1901, instructs to interpret the user prompt "Are you intending to specify a phase of the game?" and to specify the output format when the user prompt is intended to specify a phase of the game, i.e., to output "#1Q" or "#1st" if the first quarter is specified, to output "#2Q" or "#2nd" if the second quarter is specified, to output "#1Q" or "#1st" and "#2Q" or "#2nd" if the first and second quarters are specified, and to output "UNKNOWN" if the quarter is not specified or is unknown. The latter half of the text denoted by reference numeral 1902 (the portion sandwiched between "===User prompt starts===" and "===User prompt ends===") stores the user prompt entered as text.

[0108] The text scene data creation unit 1702 receives the result of the interpretation of the user's intention by the intention interpretation unit 1701, filters each scene entry in the scene list with summary text data 1722 output from the pre-processing device 602 (summary generation unit 903), extracts scene entries that match the intention of the user prompt, and creates text scene data 1723. A predetermined analysis engine may be used to filter the scene entries.

[0109] The text scene data creation unit 1702 creates text scene data using rule-based processing. The rule-based processing performed by the text scene data creation unit 1702 can be coded using, for example, Python or any other programming language. FIG. 20 shows a simplified example of pseudo-programming code describing the rule-based processing performed by the text scene data creation unit 1702. The portion indicated by reference numeral 2051 instructs the intention interpretation unit 1701 to extract scene data corresponding to the phase (quarter) specified by the user from the scene list when the intention interpretation unit 1701 interprets an intention to specify a game phase (quarter_prompt). Furthermore, the portion indicated by reference numeral 2052 instructs the intention interpretation unit 1701 to extract data from the corresponding scene list and write text scene data when the intention interpretation unit 1701 interprets an intention to extract a scoring scene (scoring_prompt). The description rules include rules for converting scene entries corresponding to changes in the team's score into text scene data, and rules for converting scene entries corresponding to changes in the opposing team's score into text scene data.

[0110] The text scene data creation unit 1702 converts the structured data (see FIG. 10 ) of each scene entry matching the user prompt into text scene data consisting of a single line of plain text data according to the rule-based processing shown in FIG. 20 . FIG. 21 shows an example of the data format of text scene data 1723 output from the text scene data creation unit 1702. The illustrated text scene data is based on the assumption that the user prompt is interpreted as an attempt to extract scoring scenes. The text scene data includes, from the beginning, a scene ID, the name of the scoring team, the number of points scored, keywords, and summary text data, with each piece of data actually listed on the same line, separated by commas. The text scene data creation unit 1702 creates text scene data consisting of a single line for each scene entry matching the user prompt, and outputs this text scene data to the subsequent scene extraction unit 1703 in a format sorted, for example, by scene ID.

[0111] The text-input user prompt and the text scene data created by the text scene data creation unit 1702 are formatted into an appropriate prompt template according to the template instruction 1712, and then input to the scene extraction unit 1703. FIG. 22 shows an example of a prompt input to the scene extraction unit 1703. The illustrated prompt contains the following instruction at the location indicated by reference numeral 2201: "Select a scene from the provided game commentary data that best suits the user's criteria. Data for each scene is written on a separate line in the format of 'ID, play information.' Avoid selecting scenes based solely on chronological order; follow the user's prompt to determine the relevance of the scenes." The text scene data created by the text scene data creation unit 1702 is stored at the location indicated by reference numeral 2202, {play_data}, and the text-input user prompt is stored as is at the location indicated by reference numeral 2203, {user_prompt}.

[0112] The scene extraction unit 1703 selects the ID of a scene indicated by the user prompt based on prompt text, which includes a user prompt 1721 entered as text by the user and text scene data 1723 created by the text scene data creation unit 1702 and is shaped into an appropriate prompt format based on the template instruction 1712. As a result, the scene extraction unit 1703 can create a highlight scene list 1724 in which scene entries of highlight scenes indicated by the user prompt are extracted from the scene list. The scene extraction unit 1703 can perform a process of selecting the ID of a scene indicated by the user prompt from the text scene data using a large-scale language model, such as GPT-4 (see Non-Patent Document 2). Because the scene entry includes time information (in point and out point) of the scene, it is possible to play back video of the highlight scene based on the highlight scene list (described below).

[0113] The scene description creation unit 1704 creates metadata including a scene description corresponding to the user prompt for each scene entry included in the highlight scene list 1724, based on the summary text data included in the scene entry and the user prompt entered as text by the user. The scene description creation unit 1704 can create the metadata using a large-scale language model. However, in consideration of execution speed, GPT3.5 turbo is used for the scene description creation unit 1704.

[0114] 23 shows an example of a prompt input to the scene description creation unit 1704. The prompt includes instructions such as "Describe the scene briefly," "Describe the score in the format of '{your team name}:{score}' and '{opponent team name}:{score}'," "Describe the score changes for each team in the format of '{your team name} {score change} and {opponent team name} {score change}'," "Please use summary text data to enhance your description," "This description is based on a user prompt," and "Generate a simple, one-line description." Because the prompt includes the instructions "Please use summary text data to enhance your description" and "This description is based on a user prompt," the scene description creation unit 1704 can create metadata including a scene description that corresponds to the user prompt (i.e., in accordance with the user's intention). Then, the scene description creation unit 1704 follows the prompt shown in FIG. 23 and adds the created metadata for each scene to the corresponding scene entry in the highlight scene list 1724 according to the template instructions 1713 .

[0115] In this manner, the demonstration device 602 can extract highlight scenes from input video (e.g., sports video) based on user instructions and create a highlight scene list that stores metadata including time information (in and out points) and scene information (such as scores and information about game phases) for each highlight scene, as well as scene descriptions that reflect the user instructions.

[0116] FIG. 29 shows the processing operation of the demonstration device 602 in the form of a flowchart.

[0117] First, the user input unit 106 accepts a user prompt 1721 that indicates what kind of scene the user wants from the original video data.

[0118] Then, the intention interpretation unit 1701 interprets the intention of the user prompt 1721 input as text (step S2901). Specifically, the intention interpretation unit 1701 determines the task intended by the user from the user prompt input as text.

[0119] When the text scene data creation unit 1702 receives the result of the interpretation of the user's intention by the intention interpretation unit 1701, it filters each scene entry in the scene list with summary text data 1722 supplied from the pre-processing device 602, extracts scene entries that match the intention of the user prompt, and creates text scene data 1723 (step S2902).

[0120] Next, the scene extraction unit 1703 receives the user prompt 1721 and the text scene data 1723, selects the ID of the scene specified by the user prompt, and extracts the highlight scene specified by the user prompt (step S2903).

[0121] Next, for each scene entry included in the highlight scene list 1724, the scene description creation unit 1704 creates metadata including a scene description corresponding to the user prompt based on the summary text data included in the scene entry and the user prompt entered as text by the user (step S2904).

[0122] In this way, the demonstration device 602 uses the scene list supplied from the pre-processing device 601 to extract highlight scenes based on the user's instructions, and generates metadata for each highlight scene that reflects the user's instructions, thereby completing and outputting the highlight scene list 1724 (step S2905).

[0123] B-4. Highlight Video Generation The demonstration device 602 can provide the user with a highlight video generation service that uses the created highlight scene list to generate highlight videos from the original video based on the user's instructions, or to assist in searching for highlight scenes from the entire original video.

[0124] 24 to 27 show configuration examples of highlight video generation screens provided by the demonstration device 602. However, it is assumed that the video to be processed is a sports video of an American football game.

[0125] FIG. 24 shows an example of a highlight video generation screen display in an initial state immediately after the demonstration device 602 receives a scene list of video data from the pre-processing device 601, or before a prompt has been input by the user.

[0126] The highlight video generation (Highlight Generator) screen shown in Fig. 24 includes a first display section 2401 that displays the original video (Original Video) and a second display section 2402 that displays the highlight scene video (Highlight Video). The first display section 2401 and the second display section 2402 may each include UI components such as a progress bar indicating the playback position of the video, a speaker volume icon, a full-screen display button, and a kebab menu (menu icon). However, in the initial state, no highlight scene has been extracted, i.e., no display signal has been generated, so the second display section 2402 is in a no-display state (No Highlight Video is generated).

[0127] Below the first display unit 2401 and the second display unit 2402, time information obtained from the processing results of the video analysis unit 901 and the voice recognition unit 902 of the pre-processing device 601 is displayed using a timeline.

[0128] The timeline indicated by reference numeral 2411 displays "Play Scene," i.e., the time position at which a play scene in the video data occurred. The timeline indicated by reference numeral 2412 displays "Scoring Event," i.e., the time position at which a score occurred in the play scene. The scene list includes scene entries created for each scene extracted from the original video data, and the timelines 2411 and 2412 can be displayed based on scene information stored in each scene entry, which describes the time information and content of the scene. Of course, additional timelines may be provided to display the positions at which scenes (events) other than scoring scenes occurred.

[0129] The timeline indicated by reference numeral 2413 displays "Live Commentary," i.e., the time position at which a game commentary occurred, based on the recognition result of the voice recognition unit 902 for the voice data related to the video data. Reference numeral 2414 displays "Audio," i.e., the waveform signal of the voice data related to the video data. The timeline indicated by reference numeral 2415 can display "Highlight Scene," i.e., the time position of a highlight scene extracted by the scene extraction unit 1703. However, since FIG. 24 assumes an initial state in which no highlight scene has yet been extracted, the timeline 2415 is empty.

[0130] Furthermore, the highlight video generation screen shown in FIG. 24 has an input area on the right side of the second display section 2402 where the user can input instructions such as the scenes they want to view. The input area indicated by reference numeral 2421 is a type designation section for designating the type of sport in the sports video. The input area indicated by reference numeral 2422 is a user prompt input section where the user can input a user prompt indicating a desired scene as text. The user can input the user prompt as text into the user prompt input section 2422 using, for example, a keyboard. The user's speech may be recognized by voice recognition and the user prompt may be input as text into the user prompt input section 2422. The input area indicated by reference numeral 2423 is a display number designation section for designating the number of highlight scenes to be displayed at one time in the highlight scene list (described below). The user can designate the number of highlight scenes to be displayed by sliding a slide bar in the display number designation section 2423, or can designate the number of highlight scenes by entering a numerical value. Reference numeral 2424 is a "Generate" button. When the user presses the "Generate" button, the demonstration apparatus 602 starts generating highlight scenes based on the conditions specified in the type specifying section 2421 and the user prompt input section 2422. Fig. 24 shows an example of the highlight video generation screen in an initial state, with the type specifying section 2421, the user prompt input section 2422, and the number of images to display specifying section 2423 all remaining blank (empty).

[0131] Fig. 25 shows an example in which a user has input into the type designation section 2421, the user prompt input section 2422, and the display count designation section 2423 of the highlight video generation screen in the initial state shown in Fig. 24. Even in the state shown in Fig. 25, the basic configuration of the highlight video generation screen is the same as in the initial state shown in Fig. 24.

[0132] In the example shown in FIG. 25 , the type designation section 2421 designates the type of sport “American football,” the user prompt input section 2422 inputs the user prompt “Please extract touchdown scene.” as text indicating that a touchdown scene is desired as a highlight scene, and the number-to-display designation section 2423 designates “3” as the desired number of highlight scenes.

[0133] When the "Generate" button is pressed in the input state shown in Figure 25, highlight scene generation processing begins in the demonstration device 602. The intention interpretation unit 1701 interprets the intention of the user prompt "Please extract touchdown scene." The text scene data creation unit 1702 creates text scene data from the scene list created by the pre-processing device 601 based on the intention interpretation result of the user prompt by the intention interpretation unit 1701. Then, the scene extraction unit 1703 extracts and creates a highlight scene list. In addition, the scene description creation unit 1704 creates metadata including a scene description corresponding to the user prompt for each highlight scene included in the highlight scene, and stores the metadata in the corresponding scene entry of the highlight scene list.

[0134] 26 shows an example of a highlight video generation screen display that has been updated based on a highlight scene list created by the demonstration device 602 in response to a user prompt. In the state shown in FIG. 26, the basic configuration of the highlight video generation screen is the same as in the initial state shown in FIG.

[0135] As shown in FIG. 26 , a highlight scene list window 2600 pops up in approximately the center of the highlight video generation screen. The highlight scene list 2600 displays thumbnail images (representative frames such as in-point image frames, for example) 2601-2603 of the selected highlight scenes, the specified number of which is three. The "Highlight Scene" timeline 2415 also displays the time positions of the highlight scenes 2601-2603. The thumbnail images 2601-2603 of the highlight scenes each display corresponding metadata 2604-2606. The highlight scene list window 2600 includes a "Selected Scenes" button 2611 for listing the specified number of highlight scenes, and an "All Scenes" button 2612 for listing all highlight scenes. FIG. 26 shows the state in which the "Selected Scenes" button 2611 is selected. When the "All Scenes" button 2612 is pressed, thumbnail images of all highlight scenes are displayed in the highlight scene list window 2600 (not shown).

[0136] When the user selects one of highlight scenes 2601 to 2603 displayed in highlight scene list window 2600, highlight scene list window 2600 closes and a playback video of the selected highlight scene is displayed in second display portion 2402. Figure 27 shows a state in which highlight scene 2601 is selected in highlight scene list window 2600 and a playback video of highlight scene 2601 is displayed in second display portion 2402.

[0137] C. Configuration Examples of Information Processing Devices This section C describes the configuration of information processing devices that can be used for the pre-processing device 601 and the demonstration device 602. The pre-processing device 601 and the demonstration device 602 may each be realized by one information processing device, or the functions of both the pre-processing device 601 and the demonstration device 602 may be realized by one information processing device.

[0138] 30 shows an example of the hardware configuration of an information processing device 2000. This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by an information terminal such as a personal computer, a tablet, or a smartphone.

[0139] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing processes with a high computational load on the information processing device 2000 (for example, processes related to model learning such as a speech recognition model or a summary generation model), it is desirable that the CPU 2001 be a multi-core CPU (for example, Apple M1 Max, etc.), or that the information processing device 2000 be further equipped with a multi-core processor such as a GPU (Graphics Processing Unit) or GPGPU (General-purpose computing on graphics processing units) (for example, NVIDIA's "Quadro A6000"). However, for convenience, these will be collectively referred to as the CPU 2001 below.

[0140] The ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).

[0141] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and other components. The CPU 2001 executes various application programs in an execution environment provided by the OS through the cooperative operation of the ROM 2002 and RAM 2003, thereby enabling various functions and services to be realized. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. Furthermore, examples of application programs executed on the information processing device 2000 include the following: At least some of the application programs may be computer programs provided as libraries. (1) A program that performs video analysis processing using an analysis engine. (2) A program that performs speech recognition processing. (3) A program that processes text data using a large-scale language model (e.g., generating a summary). (4) A program that integrates video analysis information and text data obtained from speech recognition results to extract scenes from video data based on user instructions. (5) A program that generates explanatory text for each scene in response to user instructions using a large-scale language model.

[0142] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).

[0143] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 30 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.

[0144] The input unit 2008 is composed of an input control circuit that generates an input signal based on user input and outputs it to the CPU 2001. If the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, mouse, and touch panel, and may also include a camera and microphone used for remote conferences and face-to-face customer service. The output unit 2009 includes display devices such as a liquid crystal display (LCD) device, an organic electroluminescence (EL) display device, and an LED (light emitting diode), as well as an audio output device such as a speaker. User prompts are issued using the input unit 2008, and GUI screens (see, for example, FIGS. 4 and 24 to 27) are displayed using the output unit 2009.

[0145] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.

[0146] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 113. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.

[0147] The communication unit 2013 is a device that performs wireless communication via Wi-Fi (registered trademark), Bluetooth (registered trademark), or cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, and the like. Programs executed on the information processing device 2000 are installed from an external device, for example, via the communication unit 2013. An acoustic signal that is the subject of the summary generation process according to the present disclosure is captured, for example, via the communication unit 2013.

[0148] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.

[0149] According to the present disclosure, by integrating video analysis information and information based on voice recognition, scenes can be extracted from video data based on user instructions. Furthermore, according to the present disclosure, metadata such as explanatory text based on user instructions can also be generated for each extracted scene. Furthermore, according to the present disclosure, by linking and managing the extracted scene information and metadata, the metadata can be utilized when playing back video of the scene.

[0150] The present disclosure can be suitably realized by utilizing an AI engine for video analysis and a large-scale language model for text data processing, but of course, means other than an AI engine and a large-scale language model may also be used.

[0151] The present disclosure can also be applied to video data of sports competitions to extract scenes of scoring goals or specific phases in the game that a user desires, and generate metadata such as explanatory text for each scene based on the user's instructions. The present disclosure can be applied to processing video data of various sports by selecting the AI ​​engine and AI engine parameters used for video analysis depending on the sport. Of course, the present disclosure is not limited to sports videos, but can also be applied to processing video data of various genres.

[0152] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.

[0153] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.

[0154] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.

[0155] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.

[0156] The present disclosure may also be configured as follows.

[0157] (1) An information processing device comprising: a video analysis unit that analyzes video data and extracts a plurality of specific scenes; and a summary generation unit that recognizes and processes audio data related to the video data and generates summary text data that summarizes speech, and associates the plurality of specific scenes with the summary text data.

[0158] (2) The information processing device described in (1) above further comprises: a prompt creation unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user; and a scene extraction unit that selects a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of multiple scene entries including scene information based on video analysis for each scene and summary text data, based on the prompt text created by the prompt creation unit, and extracts a highlight scene list.

[0159] (3) The information processing device described in (2) above, wherein the prompt creation unit includes: an intention interpretation unit that interprets the task intended by the user from a user prompt entered as text by the user; and a text scene data creation unit that collects text data of scene entries that match the task intended by the user from the scene list and creates text scene data; and the scene extraction unit extracts highlight scenes indicated by the user prompt based on prompt text including the user prompt and the text scene data.

[0160] (4) The information processing device according to any one of (2) or (3), further comprising: a metadata creation unit that creates metadata including a scene description for each scene entry in the highlight scene list in accordance with an instruction from a user.

[0161] (5) The information processing device according to (4), wherein the metadata creation unit creates metadata including a scene description according to an instruction from a user, based on summary text data included in a scene entry and the user prompt.

[0162] (6) The information processing device according to any one of (2) to (5), wherein the summary generation unit generates a summary of text data obtained by speech recognition of audio data related to the video data using a large-scale language model.

[0163] (7) The information processing device according to any one of (2) to (6), wherein the scene extraction unit extracts a scene entry of the highlight scene from the scene list using a large-scale language model.

[0164] (8) The information processing device according to any one of (2) to (7), wherein the prompt creation unit uses a large-scale language model to create a prompt text from a user prompt that is text input by a user.

[0165] (9) The information processing device according to any one of (4) or (5), wherein the metadata creation unit extracts a scene entry for the highlight scene from the scene list using a large-scale language model.

[0166] (10) An information processing method comprising: a first processing step of associating the plurality of specific scenes with the summary text data, the first processing step including the steps of: a video analysis step of analyzing video data and extracting a plurality of specific scenes; and a summary generation step of recognizing and processing audio data related to the video data to generate summary text data summarizing speech, the first processing step associating the plurality of specific scenes with the summary text data; a prompt creation step of creating prompt text suitable for scene extraction processing from a user prompt entered as text by a user; and a scene extraction step of selecting, based on the prompt text created in the prompt creation step, a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries each including scene information based on video analysis for each scene and summary text data, the second processing step including the steps of:

[0167] (11) A computer program written in a computer-readable format to cause a computer to function as: a first processing unit that analyzes video data to extract multiple specific scenes, recognizes and processes audio data related to the video data to generate summary text data that summarizes the speech, and associates the multiple specific scenes with the summary text data; and a second processing unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user, and, based on the prompt text, selects a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of multiple scene entries each containing scene information based on video analysis for each scene and summary text data, and creates a highlight scene list consisting of scene entries of highlight scenes.

[0168] 100...Scene extraction system, 101...Video analysis unit, 102...Speech recognition unit, 103...Summary generation unit, 104...Prompt creation unit, 105...Scene extraction unit, 106...User input unit, 601...Pre-processing device, 602...Demonstration device, 901...Video analysis unit, 902...Speech recognition unit, 903...Summary generation unit, 1201...Interface, 1202...AI process manager, 1203...AI engine, 1701...Intention interpretation unit, 1702...Text scene data creation unit, 1703...Scene extraction unit, 1704...Scene description creation unit, 2000...Information processing device, 2001...CPU, 2002...ROM, 2003...RAM, 2004...Host bus, 2005...Bridge, 2006...Expansion bus, 2007...Interface unit 2008...input unit, 2009...output unit, 2010...storage unit, 2011...drive, 2012...removable recording medium, 2013...communication unit

Claims

1. An information processing device comprising: a video analysis unit that analyzes video data and extracts a plurality of specific scenes; and a summary generation unit that recognizes and processes audio data related to the video data and generates summary text data that summarizes the speech; and associates the plurality of specific scenes with the summary text data.

2. The information processing device of claim 1, further comprising: a prompt creation unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user; and a scene extraction unit that selects a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries including scene information based on video analysis for each scene and summary text data, based on the prompt text created by the prompt creation unit, and extracts a highlight scene list.

3. The information processing device of claim 2, wherein the prompt creation unit includes: an intention interpretation unit that interprets the task intended by the user from a user prompt entered as text by the user; and a text scene data creation unit that collects text data of scene entries from the scene list that match the task intended by the user and creates text scene data; and the scene extraction unit extracts highlight scenes indicated by the user prompt based on prompt text including the user prompt and the text scene data.

4. The information processing device according to claim 2, further comprising a metadata creating unit that creates metadata including a scene description for each scene entry in the highlight scene list in accordance with an instruction from a user.

5. The information processing device according to claim 4, wherein the metadata creation unit creates metadata including a scene description in response to an instruction from a user, based on summary text data included in a scene entry and the user prompt.

6. The information processing device according to claim 2, wherein the summary generation unit generates a summary of text data obtained by speech recognition of audio data related to the video data using a large-scale language model.

7. The information processing device according to claim 2, wherein the scene extraction unit extracts scene entries of the highlight scenes from the scene list using a large-scale language model.

8. The information processing device according to claim 2, wherein the prompt creation unit uses a large-scale language model to create prompt text from a user prompt entered as text by a user.

9. The information processing device according to claim 4, wherein the metadata creation unit extracts scene entries for the highlight scenes from the scene list using a large-scale language model.

10. An information processing method comprising: a first processing step including the steps of: a video analysis step of analyzing video data and extracting a plurality of specific scenes; and a summary generation step of recognizing and processing audio data related to the video data to generate summary text data that summarizes utterances, wherein the first processing step associates the plurality of specific scenes with the summary text data; a prompt creation step of creating prompt text suitable for scene extraction processing from a user prompt entered as text by a user; and a scene extraction step of selecting, based on the prompt text created in the prompt creation step, a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries each containing scene information based on video analysis for each scene and summary text data, wherein the second processing step creates a highlight scene list consisting of scene entries of highlight scenes.

11. A computer program written in a computer-readable format to cause a computer to function as: a first processing unit that analyzes video data to extract a plurality of specific scenes, recognizes and processes audio data related to the video data to generate summary text data that summarizes the speech, and associates the plurality of specific scenes with the summary text data; and a second processing unit that creates prompt text suitable for scene extraction processing from a user prompt entered as text by a user, and based on the prompt text, selects a scene entry of a highlight scene indicated by the user prompt from a scene list consisting of a plurality of scene entries each containing scene information based on video analysis of each scene and summary text data, and creates a highlight scene list consisting of scene entries of highlight scenes.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    WO2021241430A1

  • Information processing device, information processing method, and program

    WO2023233998A1

  • Motion picture summary automatic generation apparatus and method, and computer program

    JP2008148121A

  • Conference system, summarization device, method of controlling conference system, method of controlling summarization device, and program

    JP2019139572A

  • Tagging device for moving images, method, and program

    JP2020079982A

Cited By

  • Bridge monitoring method, medium, device and system based on multi-agent cooperation

    CN122473752A