Information processing device, information processing method, and recording medium

The information processing device addresses the challenge of catering to varied viewer preferences by generating customized edited images and timetables through metadata analysis and instruction-based scene identification, enhancing viewer satisfaction.

WO2025253993A1PCT designated stage Publication Date: 2025-12-11NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/019256
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-07
Filing Date
2025-05-28
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing video editing technologies struggle to generate diverse content patterns that cater to varying viewer preferences, limiting satisfaction among multiple viewers when generating highlight videos, thumbnail images, and video timetables from a single video.

Method used

An information processing device that acquires videos, generates metadata, and identifies scenes based on instruction information to create customized edited images and timetables by analyzing video content using metadata and instruction information, allowing for multiple patterns of content generation.

Benefits of technology

Enables the generation of diverse edited images and timetables tailored to individual viewer preferences, increasing satisfaction by providing customized content even from the same video source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025019256_11122025_PF_FP_ABST
    Figure JP2025019256_11122025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present disclosure comprises a video acquiring unit, a metadata generating unit, an instruction information acquiring unit, and an editing unit. The video acquiring unit acquires a processing target video, which is at least one video. The metadata generating unit generates metadata relating to the content of each of at least one part of the processing target video. The instruction information acquiring unit acquires instruction information specifying the content of a scene to be extracted. The editing unit identifies, on the basis of the metadata, a part of the processing target video having content that is related to the scene to be extracted, as indicated by the instruction information.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and recording medium

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] A technology related to this disclosure is disclosed in Patent Document 1. Patent Document 1 discloses a technology for generating a summary text from a presentation video using a summary model generated by machine learning.

[0003] International Publication No. 2023 / 166746

[0004] There is a demand for video editing technology that can generate highlight videos, thumbnail images, etc. (hereinafter, these may be collectively referred to as "edited images") from videos, and that can generate video timetables.

[0005] Patent Literature 1 discloses a video editing technology that generates summary text from a video. However, the technology disclosed in Patent Literature 1 can only generate one pattern of summary text from a single video. Since preferences vary from person to person, when generating content (highlight video, thumbnail images, summary text, etc.) for multiple viewers, it is difficult to increase the satisfaction of multiple viewers by generating and providing only one pattern of content.

[0006] One example of a purpose of this disclosure is to provide a new video editing technique.

[0007] According to one aspect of this disclosure, there is provided an information processing device having: a video acquisition means for acquiring a target video, which is at least one video; a metadata generation means for generating metadata relating to the content of at least one portion of the target video; an instruction information acquisition means for acquiring instruction information specifying the content of a scene to be extracted; and an editing means for identifying the portion whose content is related to the scene to be extracted, indicated by the instruction information, based on the metadata.

[0008] Furthermore, according to one aspect of this disclosure, an information processing method is provided in which one or more computers acquire a target moving image, which is at least one moving image, generate metadata regarding the content of at least one portion of the target moving image, acquire instruction information specifying the content of a scene to be extracted, and, based on the metadata, identify the portion whose content is related to the scene to be extracted as indicated by the instruction information.

[0009] Furthermore, according to one aspect of this disclosure, a program is provided that causes a computer to function as: a video acquisition means for acquiring at least one video to be processed; a metadata generation means for generating metadata relating to the content of at least one portion of the video to be processed; an instruction information acquisition means for acquiring instruction information specifying the content of a scene to be extracted; and an editing means for identifying, based on the metadata, the portion whose content is related to the scene to be extracted as indicated by the instruction information.

[0010] According to one example of this disclosure, a new video editing technique is provided.

[0011] FIG. 1 is a diagram showing an example of a functional block diagram of an information processing device. FIG. 2 is a flowchart showing an example of a processing flow of the information processing device. FIG. 3 is a diagram showing an example of a hardware configuration of the information processing device. FIG. 4 is a diagram showing another example of a functional block diagram of the information processing device. FIG. 5 is a diagram showing an example of information processed by the information processing device. FIG. 6 is a diagram showing an example of information processed by the information processing device. FIG. 7 is a diagram showing an example of a timetable generated by the information processing device. FIG. 8 is a flowchart showing another example of a processing flow of the information processing device.

[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In this disclosure, the drawings relate to one or more embodiments. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate.

[0013] First Embodiment Fig. 1 is a functional block diagram showing an overview of an information processing device 10. Fig. 2 is a flowchart showing an example of the flow of processing executed by the information processing device 10.

[0014] 1, the information processing device 10 includes a video acquisition unit 11, a metadata generation unit 12, an instruction information acquisition unit 13, and an editing unit 14. These functional units execute the process of the flowchart in FIG.

[0015] At S10, the video acquisition unit 11 acquires at least one video to be processed. At S11, the metadata generation unit 12 generates metadata regarding the content of at least one portion of the video to be processed. At S12, the instruction information acquisition unit 13 acquires instruction information specifying the content of a scene to be extracted. At S13, the editing unit 14, based on the metadata, identifies a portion of at least one portion of the video to be processed whose content is related to the "scene to be extracted" indicated by the instruction information.

[0016] The order of processing is not limited to the order shown in the flowchart of Fig. 2 and can be changed as appropriate. For example, the instruction information may be acquired in S12 before the metadata is generated in S11. The instruction information may also be acquired in S12 before the video to be processed is acquired in S10, or these steps may be performed in parallel.

[0017] In this way, the information processing device 10 acquires the target moving image, which is the object of editing, as well as "instruction information" that specifies the content of the scene to be extracted. Then, the information processing device 10 identifies a part of the target moving image based on the acquired instruction information.

[0018] According to this information processing device 10, it is possible to identify a portion of the moving image to be processed that corresponds to the content of the instruction information. Even when the same moving image is to be processed, by inputting instruction information with different content to the information processing device 10, it is possible to identify a portion of different content from the moving image to be processed. In other words, by inputting instruction information with various content, it is possible to identify portions of various patterns from the moving image to be processed. Furthermore, by inputting multiple patterns of instruction information, it is possible to identify portions of multiple patterns from the moving image to be processed.

[0019] Using a portion of the moving image to be processed thus identified, various edited images (highlight moving images, thumbnail images, etc.) can be generated, or a timetable can be generated.

[0020] Furthermore, the information processing device 10 generates metadata related to the content of the moving image for at least a portion of the moving image to be processed. Then, the information processing device 10 identifies scenes to be extracted from the moving image to be processed based on the metadata and the instruction information. With this information processing device 10, it is possible to efficiently and accurately identify scenes to be extracted that are indicated by the instruction information.

[0021] In this way, the information processing device 10 provides a new video editing technique.

[0022] <<Second Embodiment>> <Overview> An information processing apparatus 10 according to a second embodiment is a specific implementation of the configuration of the information processing apparatus 10 according to the first embodiment. A detailed description will be given below.

[0023] <Hardware Configuration> First, an example of the hardware configuration of the information processing device 10 will be described. Each functional unit of the information processing device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the realization method and device. The software includes programs that are pre-stored in the device before shipping, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.

[0024] FIG. 3 is a block diagram illustrating an example of the hardware configuration of an information processing device 10. As shown in FIG. 3, the information processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The information processing device 10 does not necessarily have to have the peripheral circuit 4A. Note that the information processing device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices may have the above hardware configuration.

[0025] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a central processing unit (CPU) or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, and touch panel. Examples of output devices include a display, projection device, speaker, printer, and mailer. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0026] <Functional Configuration> Next, a detailed description will be given of the functional configuration of the information processing device 10. Fig. 1 is an example of a functional block diagram of the information processing device 10. As shown in the figure, the information processing device 10 has a video acquisition unit 11, a metadata generation unit 12, an instruction information acquisition unit 13, an editing unit 14, and an output unit 15.

[0027] The moving image acquisition unit 11 acquires a processing target moving image, which is at least one moving image.

[0028] The entire video recorded in one video file may be "one video." Alternatively, a continuous portion of a video recorded in one video file may be "one video."

[0029] The moving image acquisition unit 11 can acquire the processing target moving image by performing at least one of the following acquisition examples 1 and 2.

[0030] In Acquisition Example 1, the user specifies or inputs at least one moving image, and the moving image acquisition unit 11 acquires the at least one moving image specified or input by the user as a moving image to be processed.

[0031] A user is a person or organization that generates edited images from a moving image to be processed using the information processing device 10. Various people or organizations can be users.

[0032] For example, a user may be a content provider who generates content such as long videos or edited images and provides them to viewers. Alternatively, a user may be a viewer who watches content provided by a content provider. Alternatively, a user may be a viewer who watches videos that the viewer owns. Alternatively, a user may be a business that generates edited images from videos generated by a content provider based on a request from the content provider. Note that the examples of users here are merely examples and are not limited to these.

[0033] The user may perform an input to designate at least one video file as a video to be processed from among video files stored in a storage device accessible from the information processing device 10. The video acquisition unit 11 may then acquire the at least one video file designated by the input as a video to be processed. Alternatively, the user may perform an operation to transmit (e.g., upload) at least one video file stored in a specified storage device to the information processing device 10. The video acquisition unit 11 may then acquire the at least one video file transmitted by the operation as a video to be processed. In these examples, the entire video recorded in each of the one or more designated or transmitted video files becomes the video to be processed. The specified storage device may be provided within the information processing device 10, or may be provided in an external device accessible from the information processing device 10. The same assumption regarding the specified storage device applies hereinafter.

[0034] In addition to the above-described operations of specifying and transmitting a video file, the user may also specify a portion of the video recorded in the video file. The video acquisition unit 11 may then acquire the portion of the video specified by this operation as the video to be processed. The specification of a portion of the video recorded in the video file can be achieved using widely known technology.

[0035] The user may perform the above-described operations via an input device of the information processing device 10. Alternatively, the information processing device 10 may be a server. The user may perform the above-described operations via a client terminal. Examples of the client terminal include, but are not limited to, a smartphone, a tablet terminal, a personal computer, a television, a mobile phone, a smart watch, smart glasses, a game terminal, etc.

[0036] In the second acquisition example, the moving image acquisition unit 11 acquires at least one moving image selected according to a predetermined rule from among the moving images stored in a predetermined storage device as a moving image to be processed.

[0037] The predetermined rule may define the videos to be selected based on the attribute information of the videos. Examples of the attribute information of the videos include, but are not limited to, the shooting date and time, the video length, the shooting location, the file name, the video title, the video tag, and the camera angle (high angle, low angle, horizontal angle, etc.). Examples of the predetermined rule include, but are not limited to, "selecting videos shot on the same day as videos to be processed" and "selecting videos shot between 9:00 and 12:00 on the same day as videos to be processed."

[0038] In the case of acquisition example 2, the video acquisition unit 11 acquires at least one video as a video to be processed at a predetermined timing. The predetermined timing may be a predetermined time of day, such as X o'clock. Alternatively, the video acquisition unit 11 may acquire at least one video as a video to be processed at other intervals, such as once a week or once a month. Furthermore, the video acquisition unit 11 may acquire at least one video as a video to be processed at a timing when an instruction is input from the user.

[0039] The metadata generation unit 12 generates metadata relating to the content of each of at least one portion of the moving image to be processed. For example, the metadata generation unit 12 can generate metadata relating to the content of each of multiple portions of the moving image to be processed.

[0040] First, the process of dividing a moving image to be processed into a plurality of portions will be described. Hereinafter, a portion of a moving image to be processed may be referred to as a "moving image portion to be processed."

[0041] The lengths of the multiple moving image portions to be processed may be the same or different. Furthermore, the moving image to be processed may be divided so that the entire moving image is included in one of the moving image portions to be processed. Alternatively, any part of the moving image to be processed may not be included in any of the moving image portions to be processed. Furthermore, the moving image may be divided so that there are overlapping parts included in multiple moving image portions to be processed, or so that there are no overlapping parts included in multiple moving image portions to be processed.

[0042] In one example, the metadata generating unit 12 divides the moving image to be processed into a plurality of moving image portions each having a predetermined number of frames (for example, every 1 frame, every 10 frames) according to a predetermined rule.

[0043] In another example, the metadata generation unit 12 divides the target video into multiple segments by using a technique for analyzing video and dividing it into multiple segments. An example of a technique for analyzing video and dividing it into multiple segments is a technique for dividing the video into multiple scenes. For example, the metadata generation unit 12 uses this technique to detect scene changes in the target video. The metadata generation unit 12 then determines a single scene from a scene change to the immediately following scene change as a single target video. Note that the use of a technique for dividing a video into multiple scenes is merely an example, and the metadata generation unit 12 may divide the target video into multiple segments by using other techniques. In this example, the lengths of the multiple target video segments may differ from one another.

[0044] After dividing the video to be processed into multiple parts, the metadata generation unit 12 generates information indicating the display timing within the video file of each of the multiple video parts to be processed based on the elapsed time from the beginning of the video file, and stores the information in a specified storage device.

[0045] Next, a process for generating metadata for each of a plurality of moving image portions to be processed will be described.

[0046] The metadata for each of the plurality of moving image portions to be processed relates to the content of the respective moving image portion to be processed.

[0047] The metadata may include at least one of the following: characters, objects, camera angles (high angle, low angle, horizontal angle, etc.), shooting techniques (close-up, wide-angle), and descriptions. Note that the metadata may also include other information.

[0048] The metadata generation unit 12 generates the above-described metadata by analyzing each of the multiple processing target video portions. For example, the metadata generation unit 12 can identify characters using facial recognition technology, etc. The metadata generation unit 12 can identify characters in each of the multiple processing target video portions using the appearance feature values ​​of each of the multiple pre-registered characters.

[0049] The metadata generation unit 12 can also identify objects that appear using object detection technology, a classifier, etc. Examples of objects include, but are not limited to, cars, bicycles, balls, microphones, goals, back screens, spectator seats, dogs, cats, etc. For example, the metadata generation unit 12 can identify objects that appear in each of multiple portions of the video to be processed using an object detection model, a classifier, etc. that have been generated in advance by machine learning.

[0050] In addition, the metadata generation unit 12 can identify the camera angle (high angle, low angle, horizontal angle, etc.) and shooting technique (close-up, wide-angle) of each of multiple portions of the videos to be processed, for example, using an estimation model generated in advance by machine learning.

[0051] The metadata generation unit 12 can also generate a description indicating the content of each of a plurality of target video segments by using a large-scale visual language model that combines a large-scale language model and video recognition AI (Artificial Intelligence). The description may be written in natural language.

[0052] 5 is a diagram illustrating an example of metadata generated by the metadata generating unit 12. In the illustrated example, the items of partial identification information, range, person, and description are linked to one another.

[0053] The part identification information item indicates information that identifies multiple video parts to be processed from each other. The range item indicates information that identifies each of multiple video parts to be processed. The information in parentheses in the figure is the identification information of the video file. The range of the video part to be processed within the video file is identified by the elapsed time from the beginning of the video file. The person item indicates the person who appears in each video part to be processed. The description item indicates a description of the content of each video part to be processed.

[0054] Returning to FIG. 4, the instruction information acquisition unit 13 acquires instruction information that specifies the content of the scene to be extracted.

[0055] The instruction information acquisition unit 13 can acquire instruction information corresponding to the processing target moving image acquired by the moving image acquisition unit 11. The instruction information acquisition unit 13 may acquire one piece of instruction information corresponding to the processing target moving image acquired by the moving image acquisition unit 11, or may acquire multiple pieces of instruction information.

[0056] In one example, the user determines the content of the instruction information and inputs the determined instruction information to the information processing device 10. The user may input one piece of instruction information to the information processing device 10, or may input multiple pieces of instruction information to the information processing device 10.

[0057] The user may perform an operation to input at least one piece of instruction information via an input device of the information processing device 10. Alternatively, the information processing device 10 may be a server. The user may perform the above-described operation via a client terminal.

[0058] In another example, as shown in Fig. 6, at least one piece of instruction information is stored in advance in a predetermined storage device. The instruction information acquisition unit 13 acquires the at least one piece of instruction information stored in the predetermined storage device. The user can register the at least one piece of instruction information in advance in the predetermined storage device. This at least one piece of instruction information registered in advance is repeatedly used for general purposes in editing multiple moving images to be processed.

[0059] The configuration of the instruction information will now be described. The instruction information may be written in a natural language. Examples of such instruction information include, but are not limited to, "highlight scenes of Taro Tokyo" and "highlight scenes of the game."

[0060] The instruction information may also include an image. For example, the instruction information may be composed of natural language and an image. An example of natural language to be included in the instruction information of this example is "highlight scenes of the players appearing in the input image." Other examples of natural language to be included in the instruction information of this example include "scenes shot at the same camera angle as the input image," "scenes with a background similar to that of the input image," etc. Note that the examples here are merely examples and are not limited to these.

[0061] 4, the editing unit 14 identifies, from at least one portion of a moving image to be processed, a portion of the moving image to be processed whose content is related to the "scene to be extracted" indicated in the instruction information acquired by the instruction information acquisition unit 13, based on the metadata generated by the metadata generation unit 12. When the instruction information acquisition unit 13 acquires multiple pieces of instruction information, the editing unit 14 can perform processing to identify the portion of the moving image to be processed using each of the multiple pieces of instruction information.

[0062] The editing unit 14 may identify the content of the scene to be extracted by processing the instruction information written in natural language using a large-scale language model. By performing this processing, the scene to be extracted can be specified from the instruction content indicated by the instruction information.

[0063] For example, the editing unit 14 may input a prepared prompt, such as "Please provide five specific examples of scenes to be extracted that are specified in this instruction information," along with instruction information written in natural language to the large-scale language model. When instruction information such as "highlight scenes of a basketball game" is input to the large-scale language model along with the prompt, it is expected that "scoring scenes," "blocking scenes," "competitive play scenes," etc. will be shown as specific examples of scenes to be extracted. Depending on the content of the instruction information, it is expected that other scenes to be extracted will be determined, such as "Tokyo Taro scoring scenes," "scenes featuring Team A's mascot character," "high-angle scenes," "close-up scenes," etc.

[0064] The editing unit 14 compares the content of the scene to be extracted thus identified with the content of each of at least one portion of the moving image to be processed indicated by the metadata, and identifies the portion of the moving image to be processed whose content is related to the scene to be extracted.

[0065] For example, if the content of the identified scene to be extracted is a "scoring scene," the editing unit 14 can identify a portion of the video to be processed that is described as a scoring scene in the metadata description. The editing unit 14 may identify a portion of the video to be processed that includes a word such as "scoring scene" or a similar word in the metadata description. Similar words can be identified using a thesaurus stored in advance in a predetermined storage device.

[0066] Also, for example, if the content of the identified scene to be extracted includes a person's name, the editing unit 14 can refer to the metadata and identify a portion of the video to be processed in which the person is included among the characters.

[0067] Furthermore, for example, if the content of the identified scene to be extracted includes the name of a specific object (e.g., Team A's mascot character), the editing department 14 can refer to the metadata and identify a portion of the video to be processed that includes that object among the objects that appear.

[0068] Also, for example, if a camera angle is specified in the content of the identified scene to be extracted, the editing unit 14 can refer to the metadata and identify the portion of the video to be processed that is linked to the specified camera angle.

[0069] Also, for example, if a shooting technique is specified in the content of the identified scene to be extracted, the editing unit 14 can refer to the metadata and identify a portion of the video to be processed that is linked to the specified shooting technique.

[0070] Furthermore, the editing unit 14 can identify a portion of a moving image to be processed by combining two or more of the above-mentioned identification methods. For example, if the identified scene to be extracted is "the scene where Taro Tokyo scores a goal," the editing unit 14 can identify a portion of a moving image to be processed that is described as a scoring scene in the metadata description and includes Taro Tokyo as a character.

[0071] After identifying a portion of the video to be processed whose content is related to the ``scene to be extracted'' based on the metadata and instruction information of the portion of the video to be processed, the editing unit 14 can perform at least one of the following two processes.

[0072] A process of generating an edited image of the video to be processed based on the identified portion of the video to be processed. A process of generating a timetable indicating the display timing of the identified portion of the video to be processed within the video to be processed.

[0073] First, the process of generating an edited image will be described. The edited image is a highlight image or thumbnail image that is a shortened version of the target moving image. For example, the editing unit 14 can generate a highlight image by connecting identified portions of multiple target moving images in chronological order within the target moving image.

[0074] In addition, the editing department 14 may use widely known highlight image generation technology to extract further portions from the identified portions of the multiple videos to be processed, and connect the extracted portions in chronological order within the videos to be processed to generate a highlight image.

[0075] Furthermore, the editing unit 14 can identify at least one frame image from among the identified portions of the moving images to be processed as a thumbnail image. There are various means for identifying at least one frame image from among the identified portions of the moving images to be processed as a thumbnail image. The editing unit 14 can perform this identification in accordance with a predetermined rule. An example of a rule is, but is not limited to, "a person's face is shown at a size equal to or larger than a threshold."

[0076] The editing unit 14 may use a large-scale visual language model to generate at least one of a description, a title, and a caption for the generated edited image. For example, the editing unit 14 can achieve this generation by inputting a prepared prompt, such as "Please decide on a title for this image," together with the generated edited image into the large-scale visual language model.

[0077] Next, the process of generating a timetable will be described. As shown in Figure 7, the timetable indicates the display timing of the identified portions of the moving image to be processed within the moving image to be processed. In the figure, the display timing of each portion of the moving image to be processed is indicated by the elapsed time from the beginning of the moving image file. Note that when the moving image to be processed is made up of multiple moving image files, the timetable can indicate the display timing of each portion of the moving image to be processed for each moving image file by the elapsed time from the beginning of each moving image file.

[0078] As described above, when generating metadata, the metadata generation unit 12 divides the target moving image into multiple target moving image portions. The metadata generation unit 12 can then generate information indicating the display timing within the moving image file of each of the multiple target moving image portions based on the elapsed time from the beginning of the moving image file. The editing unit 14 can use this information to generate the timetable.

[0079] The editing unit 14 may use the large-scale visual language model to generate at least one of the description, title, and caption for each of the identified portions of the video to be processed. For example, the editing unit 14 can achieve this generation by inputting a prepared prompt, such as "Please decide on a title for this video," into the large-scale visual language model along with each of the identified portions of the video to be processed.

[0080] Then, the editing unit 14 may assign at least one of a description, a title, and a caption to each of the identified portions of the moving image to be processed in the timetable, as shown in FIG.

[0081] Note that when the instruction information acquisition unit 13 acquires multiple pieces of instruction information, the editing unit 14 can execute a process of identifying a portion of the moving image to be processed from the moving image to be processed using each piece of instruction information. That is, the editing unit 14 can execute a process of identifying a portion of the moving image to be processed from the moving image to be processed for each piece of instruction information. The editing unit 14 can then generate an edited image or a timetable for each piece of instruction information. When multiple pieces of instruction information are acquired corresponding to one moving image to be processed, the editing unit 14 can generate multiple edited images or multiple timetables corresponding to that one moving image to be processed.

[0082] 4 , the output unit 15 outputs at least one of the edited image and the time table generated by the editing unit 14. When the instruction information acquisition unit 13 acquires a plurality of pieces of instruction information, the output unit 15 outputs at least one of the edited image and the time table generated by the editing unit 14 for each piece of instruction information.

[0083] The output unit 15 can also output, together with the edited image, at least one of the description, title, and heading of the edited image generated by the editing unit 14. The output unit 15 can also output a timetable (see FIG. 7 ) including at least one of the description, title, and heading of a part of the moving image to be processed whose display timing is indicated in the timetable generated by the editing unit 14.

[0084] For example, the output unit 15 can output at least one of the edited image and the timetable via an output device included in the information processing device 10. The output unit 15 can also transmit at least one of the edited image and the timetable to another device. The output unit 15 can also store at least one of the edited image and the timetable in a predetermined storage device.

[0085] In one example, the information processing device 10 accepts input of instruction information via an input device included in the information processing device 10. Then, the information processing device 10 generates at least one of an edited image and a timetable based on the input instruction information. In this case, the output unit 15 can output at least one of the edited image and the timetable via the output device included in the information processing device 10.

[0086] In another example, the information processing device 10 is a server. The information processing device 10 generates at least one of an edited image and a timetable based on instruction information transmitted from a client terminal. In this case, the output unit 15 can transmit at least one of the edited image and the timetable to the client terminal.

[0087] In another example, the information processing device 10 acquires instruction information stored in advance in a predetermined storage device. Then, the information processing device 10 generates at least one of an edited image and a timetable based on the acquired instruction information. In this case, the output unit 15 outputs at least one of the edited image and the timetable to the predetermined storage device, and stores at least one of the edited image and the timetable in the predetermined storage device.

[0088] Next, an example of the flow of processing by the information processing device 10 will be described. Note that the purpose of this description is to explain the flow of processing. Details of each process have been described above, so a description thereof will be omitted here.

[0089] First, the information processing device 10 acquires at least one moving image to be processed (S20).

[0090] Next, the information processing device 10 generates metadata relating to the content of each part of at least one moving image to be processed among the moving images to be processed (S21).

[0091] Next, the information processing device 10 acquires instruction information that specifies the content of the scene to be extracted (S22).

[0092] Next, the information processing device 10 identifies a portion whose content is related to the scene to be extracted, which is indicated by the instruction information acquired in S22, based on the metadata generated in S21 (S23).

[0093] Next, the information processing device 10 generates and outputs at least one of an edited image and a timetable based on the identification result of S23 (S24).

[0094] The order of processing is not limited to the order shown in the flowchart of Fig. 8 and can be changed as appropriate. For example, the instruction information may be acquired in S22 before the metadata is generated in S21. The instruction information may also be acquired in S22 before the video to be processed is acquired in S20, or these steps may be performed in parallel.

[0095] <Usage Scenarios> Next, a description will be given of usage scenarios of the information processing device 10. Note that the usage scenarios illustrated here are merely examples, and the usage scenarios of the information processing device 10 are not limited to these.

[0096] In the first usage scenario, the user of the information processing device 10 is a content provider who generates content such as long videos and edited images and provides it to viewers. The content provider provides the generated content to viewers, for example, using a video distribution platform. That is, the content provider provides the generated content to viewers via a communication network such as the Internet. Note that the content provider may provide the generated content to viewers by mailing a recording medium on which the content is recorded, or by other means, such as providing the content via television broadcast.

[0097] The content provider inputs the generated long moving image to the information processing device 10 as a moving image to be processed, and generates edited images, a timetable, and the like from the long moving image.

[0098] The content provider can input a plurality of pieces of instruction information corresponding to the long video into the information processing device 10, and generate a plurality of edited images, a plurality of timetables, and the like from the long video.

[0099] In addition, the content provider can register at least one (e.g., multiple) pieces of instruction information in advance in a predetermined storage device. Based on the input long video and at least one (e.g., multiple) pieces of instruction information registered in advance, the information processing device 10 can generate at least one edited image and at least one (e.g., multiple) timetables, etc. from the long video.

[0100] The content provider provides the generated edited images and timetables to viewers. The content provider provides the generated edited images and timetables to viewers by linking them to long videos. The content provider provides the generated edited images and timetables to viewers using a method similar to the content provision method described above. For example, the content provider provides the generated edited images and timetables to viewers using a video distribution platform.

[0101] The content provider may request the generation of edited images or timetables from a business that generates edited images or timetables from videos, and the business may then use the information processing device 10 to generate the edited images or timetables.

[0102] In the second usage scenario, the user of the information processing device 10 is a viewer who watches content provided by a content provider via streaming distribution. The content provider provides the generated content to the viewer using, for example, a video distribution platform that distributes videos via streaming distribution. The viewer then searches for a desired video on the video distribution platform and watches the searched video.

[0103] In one example, this video distribution platform has the functions of the information processing device 10. That is, the video distribution platform provides a user interface (UI) screen for viewing videos to a viewer's terminal via a communication network. For example, the UI screen is provided using a web browser or a dedicated application. Examples of client terminals include, but are not limited to, smartphones, tablet terminals, personal computers, televisions, mobile phones, smartwatches, smart glasses, and game terminals.

[0104] The viewer searches for a desired video by performing a predetermined operation on the UI screen. Then, the viewer performs an operation on the UI screen to generate an edited image or a timetable of the searched video. The viewer inputs at least one piece of instruction information along with an instruction input for generating the edited image or the timetable.

[0105] The video distribution platform generates at least one edited image and at least one timetable from the searched video based on the searched video and at least one input instruction information, and then transmits the generated at least one edited image and at least one timetable to a viewer's terminal.

[0106] In the third usage scenario, the user of the information processing device 10 is a viewer who watches videos that the user owns on the user's terminal device. The videos may be videos that the viewer has generated / shot themselves, or videos that the viewer has downloaded or obtained by other means.

[0107] A viewer can install a predetermined program on their own terminal device to realize the functions of the information processing device 10 on their own terminal device. Examples of terminal devices include, but are not limited to, smartphones, tablet devices, personal computers, televisions, mobile phones, smart watches, smart glasses, and game terminals.

[0108] A viewer inputs a predetermined video he or she owns and at least one piece of instruction information into a terminal device, and generates edited images, a timetable, and the like from the video.

[0109] <Effects> According to the information processing device 10 of the second embodiment, the same effects as those of the information processing device 10 of the first embodiment are achieved.

[0110] Furthermore, the information processing device 10 acquires "instruction information" that specifies the content of the scene to be extracted in addition to the target moving image that is the target of editing. Then, the information processing device 10 identifies a portion of the target moving image based on the acquired instruction information. Then, the information processing device 10 can generate an edited image or a timetable based on the identified portion.

[0111] According to such an information processing device 10, edited images and timetables can be generated according to the content of the instruction information. Even when the same video is used as the processing target video, edited images and timetables with different contents can be generated by inputting instruction information with different contents into the information processing device 10. That is, by inputting instruction information with various contents, edited images and timetables with various contents can be generated. Furthermore, by inputting multiple patterns of instruction information, edited images and timetables with multiple patterns can be generated.

[0112] In one example, a viewer of a video can create a desired edited image by inputting desired instruction information into the information processing device 10. In this way, by generating and providing an edited image customized for each viewer, it is possible to increase the satisfaction of multiple viewers with different tastes.

[0113] In another example, a content provider who provides content to viewers can easily generate multiple patterns of edited images and timetables by simply inputting multiple patterns of instruction information into the information processing device 10. By providing multiple patterns of edited images and timetables to multiple viewers, it is possible to increase the satisfaction of multiple viewers with different tastes.

[0114] Furthermore, the information processing device 10 can output at least one of a description, a title, and a caption of the generated edited image together with the edited image. Furthermore, the information processing device 10 can output a timetable including at least one of a description, a title, and a caption of a portion of the moving image to be processed whose display timing is indicated in the generated timetable. Based on the description, title, caption, etc., the viewer can understand the content of the edited image and the content of the portion of the moving image to be processed indicated in the timetable.

[0115] Furthermore, the information processing device 10 acquires instruction information written in a natural language and processes the instruction information written in a natural language using a large-scale language model, thereby identifying the content of the scene to be extracted. Because the instruction information can be input in a natural language, the user can easily input the instruction information.

[0116] <<Third Embodiment>> An information processing apparatus 10 according to a third embodiment generates feedback information for a plurality of edited images generated based on a plurality of pieces of instruction information. This will be described in detail below.

[0117] The instruction information acquisition unit 13 acquires a plurality of pieces of instruction information for the moving image to be processed, and the editing unit 14 generates a plurality of edited images from the moving image to be processed based on each of the plurality of pieces of instruction information.

[0118] For example, a content provider who generates content such as long videos and edited images and provides them to viewers generates a plurality of edited images using the information processing device 10. Then, the content provider uses a video distribution platform to provide the generated plurality of edited images to viewers.

[0119] The editing unit 14 generates feedback information for each of the plurality of edited images. The editing unit 14 generates feedback information for each of the plurality of edited images based on at least one of the viewing history of each of the plurality of edited images and viewer input information. For example, the video distribution platform records the viewing history of each video being distributed. The video distribution platform also accepts and registers user input of ratings for each of the distributed videos. The editing unit 14 can acquire at least one of the viewing history of each of the plurality of edited images and viewer input information from the video distribution platform.

[0120] The viewing history includes the number of times each video has been viewed, the attributes of the viewers who viewed each video (gender, age group, residential area, occupation, nationality, etc.), and statistical values ​​thereof. The statistical values ​​of viewer attributes indicate, for example, the proportion of viewers with each attribute among all viewers who viewed each video (e.g., the proportion of men, etc.). The viewing history may also indicate the number of times each video has been viewed by viewers with each attribute.

[0121] The viewer input information includes statistical values ​​(e.g., average, maximum, minimum, mode, median, etc.) of the rating values ​​(e.g., 5-point rating) of each video entered by each viewer, the number of "GOOD" ratings for each video, the number of "BAD" ratings for each video, etc.

[0122] The editing unit 14 generates feedback information based on the viewing history and the input information of the viewer. The feedback information indicates at least one of the number of times each of the edited images has been viewed, statistics on viewer attributes, and ratings from the viewer.

[0123] The number of times each of the edited images has been viewed and the statistical values ​​of the viewer attributes are indicated in the viewing history.

[0124] The viewer ratings may be, for example, a statistical value of the rating values ​​of each video entered by each viewer indicated in the viewer input information, the number of "GOOD" ratings for each video, or the number of "BAD" ratings for each video. The viewer ratings may also be calculated ratings calculated using a predetermined formula based on these values. For example, the calculated rating may be calculated by subtracting the number of "BAD" ratings for each video from the number of "GOOD" ratings for each video.

[0125] The editing unit 14 can extract low-rated edited images from among the plurality of edited images based on the feedback information, and can notify the user of the instruction information used to generate the extracted low-rated edited images.

[0126] In one example, the editing unit 14 can extract low-rated edited images by relative evaluation of multiple edited images generated from the same processing target video. For example, the editing unit 14 may rank multiple edited images generated from the same processing target video based on the feedback information, and extract a predetermined number of edited images from the lowest ranked ones as low-rated edited images. Examples of ranking include, but are not limited to, descending order of the number of views, the number of views by viewers with a predetermined attribute, the highest viewer rating value, the most number of "GOOD" ratings, the fewest number of "BAD" ratings, and the highest calculated rating value. Alternatively, the editing unit 14 may calculate the standard deviation of these values ​​for each edited image, and extract edited images with a standard deviation value below a threshold as low-rated edited images.

[0127] Alternatively, the editing unit 14 may extract low-rated edited images based on absolute evaluation. For example, the editing unit 14 may extract, as low-rated edited images, edited images whose number of views, the number of views by viewers with a predetermined attribute, the viewer evaluation value, the number of “GOOD” evaluations, the number of “BAD” evaluations, or the calculated evaluation value satisfies predetermined conditions.

[0128] The editing unit 14 can notify the user of the instruction information used to generate the extracted low-rated edited image. The editing unit 14 may generate a list of low-rated instruction information and present it to the user, or may notify the user by other methods.

[0129] Based on the notification, the user can understand the instruction information that will result in the generation of a low-rated edited image.The user can then reflect the content of that information in the generation of future edited images.For example, as described in the second embodiment, if at least one piece of instruction information is registered in advance in a predetermined storage device and is used repeatedly and generally in editing multiple moving images to be processed, the user can update the instruction information registered in the predetermined storage device.Specifically, the user can delete or change the content of the instruction information that will result in the generation of a low-rated edited image.

[0130] Other configurations of the information processing apparatus 10 of the third embodiment are similar to those of the information processing apparatus 10 of the first and second embodiments.

[0131] According to the information processing device 10 of the third embodiment, the same effects as those of the information processing device 10 of the first and second embodiments are realized.

[0132] Furthermore, when the information processing device 10 generates multiple edited images from the target moving image based on each of multiple pieces of instruction information, it can generate feedback information for each of the multiple edited images. Then, based on the generated feedback information, the information processing device 10 can extract low-rated edited images from the multiple edited images and notify the user of the instruction information used to generate the extracted low-rated edited images.

[0133] Based on such notifications, the user can understand what kind of instruction information is used to generate edited images that viewers like and what kind of instruction information is used to generate edited images that viewers dislike.The user can then optimize the instruction information used to generate future edited images based on the information they have learned.As a result, the user can efficiently generate edited images that viewers tend to like.

[0134] <<Modifications>> Modifications applicable to the information processing apparatus 10 of the first to third embodiments will be described below.

[0135] Scenes that can be extracted from the moving image to be processed are scenes that are included in the moving image to be processed. Even if instruction information specifying a scene that is not included in the moving image to be processed as a scene to be extracted is input to the information processing device 10, the specified scene cannot be extracted.

[0136] If the user does not know which scenes are included in the moving image to be processed, a problem may arise in that the user inputs instruction information to the information processing device 10 specifying a scene not included in the moving image to be processed as a scene to be extracted.

[0137] Therefore, the instruction information acquisition unit 13 generates and outputs information about scenes that can be designated as scenes to be extracted based on the metadata. When executing a process for accepting input of instruction information, the instruction information acquisition unit 13 can present information about scenes that can be designated as scenes to be extracted to the user. For example, the instruction information acquisition unit 13 may display information about scenes that can be designated as scenes to be extracted on a UI screen that accepts input of instruction information. Alternatively, information about scenes that can be designated as scenes to be extracted may be called up on the screen by operating a UI component on the UI screen that accepts input of instruction information.

[0138] The information about a scene that can be designated as a scene to be extracted may be the metadata itself. Details of the metadata are as described in the second embodiment. Alternatively, the instruction information acquisition unit 13 may format the metadata (such as rearrange the data) according to a predetermined rule. Then, the instruction information acquisition unit 13 may treat the formatted data as information about a scene that can be designated as a scene to be extracted.

[0139] According to this modification, the same effects as those of the first to third embodiments are achieved. Furthermore, according to this modification, the user can grasp scenes that can be designated as scenes to be extracted when inputting instruction information. As a result, it is possible to prevent the inconvenience of inputting instruction information to the information processing device 10 that designates a scene that is not included in the moving image to be processed as a scene to be extracted.

[0140] Although this disclosure has been described above with reference to the embodiments, this disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of this disclosure within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0141] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order. However, the order of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of the content.

[0142] Some or all of the above embodiments may be described as, but are not limited to, the following notes. 1. An information processing device comprising: video acquisition means for acquiring a target video, which is at least one video; metadata generation means for generating metadata relating to the content of at least one portion of the target video; instruction information acquisition means for acquiring instruction information specifying the content of a scene to be extracted; and editing means for identifying, based on the metadata, the portion whose content is related to the scene to be extracted, indicated by the instruction information. 2. The information processing device described in 1, wherein the editing means performs at least one of the following processes: generating an edited image of the target video based on the identified portion; and generating a timetable indicating the display timing of the identified portion in the target video. 3. The information processing device described in 2, wherein the editing means performs at least one of the following processes: generating at least one of a description, title, and caption of the edited image; and generating at least one of a description, title, and caption of the portion whose display timing is indicated in the timetable. 4. 4. The information processing device of any one of 1 to 3, wherein the instruction information acquisition means acquires the instruction information written in natural language, and the editing means processes the instruction information written in natural language with a large-scale language model to identify the content of the scene to be extracted. 5. The information processing device of any one of 1 to 4, wherein the instruction information acquisition means acquires a plurality of pieces of instruction information for the processing target video, and the editing means generates a plurality of edited images from the processing target video based on each of the plurality of pieces of instruction information, and generates feedback information for each of the plurality of edited images based on at least one of a viewing history of each of the plurality of edited images and viewer input information. 6. The information processing device of 5, wherein the editing means generates the feedback information indicating at least one of the number of views, statistical values ​​of viewer attributes, and viewer evaluations for each of the plurality of edited images.7. The information processing device according to 6, wherein the editing means extracts low-rated edited images from among the plurality of edited images based on the feedback information, and notifies the user of the instruction information used to generate the extracted low-rated edited images. 8. The information processing device according to any of 1 to 7, wherein the instruction information acquisition means generates and outputs information about scenes that can be specified as scenes to be extracted based on the metadata. 9. An information processing method in which one or more computers acquire a processing target moving image that is at least one moving image, generate metadata about the content of at least one portion of the processing target moving image, acquire instruction information that specifies the content of the scene to be extracted, and identify the portion whose content is related to the scene to be extracted that is specified in the instruction information based on the metadata. 10. A program that causes a computer to function as: a video acquisition means that acquires at least one video to be processed; a metadata generation means that generates metadata regarding the content of at least one portion of the video to be processed; an instruction information acquisition means that acquires instruction information that specifies the content of a scene to be extracted; and an editing means that, based on the metadata, identifies the portion whose content is related to the scene to be extracted that is indicated by the instruction information.

[0143] This application claims priority based on Japanese Patent Application No. 2024-092860, filed on June 7, 2024, the disclosure of which is incorporated herein by reference in its entirety.

[0144] REFERENCE SIGNS LIST 10 Information processing device 11 Video acquisition unit 12 Metadata generation unit 13 Instruction information acquisition unit 14 Editing unit 15 Output unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. An information processing device having: a video acquisition means for acquiring at least one video to be processed; a metadata generation means for generating metadata relating to the content of at least one portion of the video to be processed; an instruction information acquisition means for acquiring instruction information specifying the content of a scene to be extracted; and an editing means for identifying the portion whose content is related to the scene to be extracted indicated by the instruction information based on the metadata.

2. The information processing device according to claim 1, wherein the editing means performs at least one of the following processes: generating an edited image of the video to be processed based on the identified portion; and generating a timetable indicating the display timing of the identified portion within the video to be processed.

3. The information processing device according to claim 2, wherein the editing means executes at least one of the following processes: generating at least one of a description, a title, and a heading of the edited image; and generating at least one of a description, a title, and a heading of the portion whose display timing is indicated in the timetable.

4. An information processing device described in any one of claims 1 to 3, wherein the instruction information acquisition means acquires the instruction information written in natural language, and the editing means identifies the content of the scene to be extracted by processing the instruction information written in natural language using a large-scale language model.

5. An information processing device as described in any one of claims 1 to 4, wherein the instruction information acquisition means acquires a plurality of pieces of instruction information for the video to be processed, and the editing means generates a plurality of edited images from the video to be processed based on each of the plurality of pieces of instruction information, and generates feedback information for each of the plurality of edited images based on at least one of the viewing history of each of the plurality of edited images and viewer input information.

6. An information processing device according to claim 5, wherein said editing means generates said feedback information indicating at least one of the number of times each of said plurality of edited images has been viewed, statistics on viewer attributes, and evaluations from viewers.

7. The information processing device according to claim 6, wherein the editing means extracts the edited image with a low rating from among the plurality of edited images based on the feedback information, and notifies the user of the instruction information used to generate the extracted edited image with a low rating.

8. An information processing device according to any one of claims 1 to 7, wherein the instruction information acquisition means generates and outputs information relating to scenes that can be designated as scenes to be extracted based on the metadata.

9. An information processing method in which one or more computers acquire at least one video to be processed, generate metadata regarding the content of at least one portion of the video to be processed, acquire instruction information specifying the content of a scene to be extracted, and, based on the metadata, identify the portion whose content is related to the scene to be extracted as indicated by the instruction information.

10. The information processing method according to claim 9, wherein the one or more computers perform at least one of the following processes: generating an edited image of the video to be processed based on the identified portion; and generating a timetable indicating the display timing of the identified portion within the video to be processed.

11. The information processing method according to claim 10, wherein the one or more computers execute at least one of the following processes: generating at least one of a description, title, and heading of the edited image; and generating at least one of a description, title, and heading of the portion whose display timing is indicated in the timetable.

12. An information processing method described in any one of claims 9 to 11, wherein the one or more computers acquire the instruction information written in natural language and identify the content of the scene to be extracted by processing the instruction information written in natural language using a large-scale language model.

13. An information processing method described in any one of claims 9 to 12, wherein the one or more computers acquire a plurality of pieces of instruction information for the video to be processed, generate a plurality of edited images from the video to be processed based on each of the plurality of pieces of instruction information, and generate feedback information for each of the plurality of edited images based on at least one of the viewing history of each of the plurality of edited images and viewer input information.

14. The information processing method according to claim 13, wherein one or more computers generate the feedback information indicating at least one of the number of views of each of the plurality of edited images, statistics on viewer attributes, and ratings from viewers.

15. A recording medium having recorded thereon a program that causes a computer to function as: a video acquisition means for acquiring at least one video to be processed; a metadata generation means for generating metadata relating to the content of at least one portion of the video to be processed; an instruction information acquisition means for acquiring instruction information specifying the content of a scene to be extracted; and an editing means for identifying the portion whose content is related to the scene to be extracted, as indicated by the instruction information, based on the metadata.

16. The recording medium according to claim 15, wherein the editing means executes at least one of the following processes: generating an edited image of the video to be processed based on the identified portion; and generating a timetable indicating the display timing of the identified portion within the video to be processed.

17. The recording medium according to claim 16, wherein the editing means executes at least one of the following processes: generating at least one of a description, a title, and a caption of the edited image; and generating at least one of a description, a title, and a caption of the portion whose display timing is indicated in the timetable.

18. A recording medium described in any one of claims 15 to 17, wherein the instruction information acquisition means acquires the instruction information written in natural language, and the editing means identifies the content of the scene to be extracted by processing the instruction information written in natural language using a large-scale language model.

19. A recording medium described in any one of claims 15 to 18, wherein the instruction information acquisition means acquires a plurality of pieces of instruction information for the video to be processed, and the editing means generates a plurality of edited images from the video to be processed based on each of the plurality of pieces of instruction information, and generates feedback information for each of the plurality of edited images based on at least one of the viewing history of each of the plurality of edited images and viewer input information.

20. The recording medium according to claim 19, wherein the editing means generates the feedback information indicating at least one of the number of times each of the plurality of edited images has been viewed, statistics on viewer attributes, and evaluations from viewers.

Citation Information

Patent Citations

  • Video generation method and device, computer equipment and medium

    CN118138854A

  • Digest providing system, digest providing server, index preparation terminal, digest providing method, program therefor and storage medium storing the program

    JP2003050812A

  • Information processor, information processing method, and computer program

    JP2004295568A

  • Information processing apparatus, information processing method, and program

    JP2012249156A

  • Method for specifying scene of broadcast program, evaluation method, device for specifying scene of broadcast program, and program

    JP2017091401A