Video processing method and device, electronic equipment, storage medium and program product
By storyboarding and multimodal information analysis of videos, and using a large language model to generate and select video clips, the problem of low video production efficiency in existing technologies is solved, and high-quality video clips can be efficiently generated.
Patent Information
- Application Number
- CN202410318317.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-09-16
AI Technical Summary
In the prior art, producing a short video feature requires a lot of manpower and time, resulting in low production efficiency and an inability to quickly generate high-quality short videos.
By performing storyboard processing on the video to be processed, multimodal information is obtained, and a fragment summary of the storyboard fragment is generated using the trained first language model. The target storyboard fragment is selected through the second language model, and video fragments are fused to generate the target video fragment.
The efficiency of video production is improved, and the content of the generated video clips is more accurate, meeting the actual application needs.
Smart Images

Figure CN120658910A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology and may involve fields such as video processing, artificial intelligence, and cloud technology. Specifically, the present application relates to a video processing method, device, electronic device, storage medium, and program product. Background Art
[0002] With the continuous advancement of film and television technology and equipment, a large number of different types of film and television videos such as TV series, variety shows, and movies have emerged.
[0003] Taking TV series as an example, each episode typically lasts around 45 minutes, and watching a full series often takes dozens of hours. To quickly attract users to watch, short story clips corresponding to the relevant episodes can be created based on the highlights of the story video, such as trailers for the entire series and next episodes. This can increase user interest and guide them to watch the full series.
[0004] Currently, relevant short story films are usually edited and produced by editors after manual plot analysis, which requires huge manpower, time, and software and hardware costs. In addition, the video production cycle is long and the production efficiency is low. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a video processing method, device, electronic device, storage medium, and program product that can effectively improve the efficiency of video production. To achieve this purpose, the technical solutions provided by the embodiments of the present application are as follows:
[0006] In one aspect, an embodiment of the present application provides a video processing method, the method comprising:
[0007] Get the video to be processed;
[0008] Performing storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed;
[0009] Determining multimodal information of each storyboard segment, wherein the multimodal information includes: text information corresponding to the storyboard segment and each video frame in the storyboard segment;
[0010] Based on the multimodal information of each storyboard segment and the first instruction information, a segment summary of each storyboard segment is generated using a trained first language model; the first instruction information is used to instruct the generation of a segment summary of the storyboard segment;
[0011] Acquire second instruction information, where the second instruction information is used to instruct video segment selection and corresponding video segment selection conditions;
[0012] Based on the segment summaries of the storyboard segments and the second instruction information, determining at least one target storyboard segment from the storyboard segments by using the trained second language model;
[0013] The target storyboard segments are fused to generate target video segments corresponding to the video to be processed.
[0014] On the other hand, an embodiment of the present application further provides a video processing device, the device comprising:
[0015] Video acquisition module, used to acquire the video to be processed;
[0016] A storyboard module, configured to perform storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed;
[0017] a multimodal information determination module, configured to determine multimodal information of each storyboard segment, wherein the multimodal information includes text information corresponding to the storyboard segment and each video frame in the storyboard segment;
[0018] a segment summary determination module, configured to generate a segment summary for each storyboard segment using a trained first language model based on the multimodal information of each storyboard segment and first instruction information; wherein the first instruction information is used to instruct the generation of a segment summary for the storyboard segment;
[0019] An instruction acquisition module, configured to acquire second instruction information, wherein the second instruction information is used to instruct video segment selection and corresponding video segment selection conditions;
[0020] a target storyboard determining module, configured to determine at least one target storyboard segment from each of the storyboard segments based on the segment summary of each storyboard segment and the second instruction information and using a trained second language model;
[0021] The target video generation module is used to fuse the target storyboard segments to generate a target video segment corresponding to the video to be processed.
[0022] Optionally, the segment summary determination module is further configured to:
[0023] Obtaining video metadata of the video to be processed;
[0024] For each of the storyboard segments, identifying objects in each video frame in the storyboard segment to obtain object information of the storyboard segment; wherein the multimodal information also includes the object information of the storyboard segment;
[0025] For each of the storyboard segments, a segment summary of the storyboard segment is generated by a trained first language model based on the multimodal information of the storyboard segment, the video metadata, and the first instruction information.
[0026] Optionally, the target mirror determination module may be used to:
[0027] performing content aggregation on the segment summaries of the storyboard segments to obtain a video summary of the video to be processed;
[0028] Based on the segment summaries of the storyboard segments, the video summary of the video to be processed and the second instruction information, at least one target storyboard segment is determined from the storyboard segments through the trained second language model.
[0029] Optionally, for each storyboard segment, the segment summary determination module may be configured to:
[0030] Extracting visual features from at least a portion of the video frames in the storyboard segment to obtain visual features of the storyboard segment;
[0031] Performing text feature extraction on the text information corresponding to the storyboard segment and the first instruction information to obtain text features corresponding to the storyboard segment;
[0032] The visual features of the storyboard segment and the corresponding text features are fused into a multimodal feature, and a segment summary of the storyboard segment is generated based on the fused features.
[0033] Optionally, the segment summary determination module may be configured to:
[0034] Mapping the visual features of the storyboard segment to a feature space corresponding to the text features to obtain mapped features corresponding to the visual features;
[0035] The mapped features and the text features are fused to obtain fused features.
[0036] Optionally, the second instruction information is generated in the following manner:
[0037] In response to a video editing trigger operation for the video to be processed, a video editing setting interface is displayed; the video editing setting interface displays editing operation options, and the editing operation options include a video clip selection operation option;
[0038] In response to a triggering operation on the video clip selection operation option, displaying a video clip selection condition setting interface;
[0039] Receiving at least one of an input operation of selecting a keyword or an instruction template selection operation for the video to be processed through the video segment selection condition setting interface;
[0040] Based on at least one of the received keyword selection operation or instruction template selection operation for the video to be processed, second instruction information corresponding to the video to be processed is obtained.
[0041] Optionally, the target video generation module is further configured to:
[0042] Get the maximum duration limit of the target video segment to be generated;
[0043] The target video generation module can be used to:
[0044] Determining the duration of each target storyboard segment;
[0045] If the sum of the durations of the target storyboard segments is less than or equal to the maximum duration limit, then the target storyboard segments are spliced together to generate a target video segment corresponding to the video to be processed;
[0046] If the sum of the durations of the target storyboard segments is greater than the maximum duration limit, then based on the maximum duration limit and relevant information corresponding to the target storyboard segments, select some video frames from the target storyboard segments, and splice the selected video frames to generate a target video segment corresponding to the video to be processed;
[0047] The relevant information corresponding to each target storyboard segment includes at least one of the following:
[0048] the correlation between the target storyboard segments;
[0049] a correlation between each of the target storyboard segments and a video summary of the video to be processed, wherein the video summary is obtained by performing content aggregation on the segment summaries of each of the storyboard segments;
[0050] Correlation between video frames of different target storyboards.
[0051] Optionally, the first language model is trained in the following manner:
[0052] Acquire a plurality of first training samples and first sample instruction information, each of the first training samples including multimodal information of a sample video storyboard segment and a sample segment summary of the sample video storyboard segment; the sample video storyboard segment is a storyboard segment in the first sample video;
[0053] Based on the plurality of first training samples and the first sample instruction information, a training operation is continuously performed on the pre-trained multimodal large model until a first training end condition is met, and the trained multimodal large model is used as the first large language model, wherein the training operation includes:
[0054] Inputting the first sample instruction information and the multimodal information of each of the sample video storyboard segments into a pre-trained multimodal large model to obtain a predicted segment summary of each of the sample video storyboard segments;
[0055] determining a first training loss based on a difference between the predicted segment summary and the sample segment summary of each of the sample video storyboard segments;
[0056] Model parameters of the multimodal large model are adjusted based on the first training loss.
[0057] Optionally, the second largest language model is trained in the following manner:
[0058] Acquire a plurality of second training samples and second sample instruction information, each of the second training samples including a sample segment summary of each storyboard segment of a second sample video and a label of each storyboard segment, wherein a storyboard segment label indicates whether the storyboard segment is a target storyboard segment in the corresponding second sample video;
[0059] Based on the plurality of second training samples, the second sample instruction information, and the second sample instruction information, a training operation is continuously performed on the large language model to be trained until a second training end condition is met, and the trained large language model is used as a second large language model, wherein the training operation includes:
[0060] Inputting the sample segment summary of each storyboard segment in each second training sample of the second sample instruction information into the second large language model, obtaining a predicted probability of each storyboard segment in each second training sample, wherein the predicted probability represents a probability that the storyboard segment is a target storyboard segment in the corresponding second sample video;
[0061] Determining a second training loss based on the predicted probability of each storyboard segment and the label of each storyboard segment in each of the second training samples;
[0062] Adjust model parameters of the second language model based on the second training loss.
[0063] An embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.
[0064] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.
[0065] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.
[0066] The beneficial effects of the technical solution provided by the embodiments of the present application are as follows:
[0067] The video processing method provided in the embodiment of the present application combines multiple modal information of each storyboard segment in the video, performs semantic understanding and content extraction of each storyboard segment based on the first large language model, and the obtained segment summaries of each storyboard segment are more accurate. The second large language model is used to more accurately locate the target storyboard segment that meets the selection criteria based on the content based on the second instruction information, thereby realizing the selection of video segments based on the content, improving video production efficiency, and better meeting actual application needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0069] Figure 1 A flowchart of a video processing method provided in an embodiment of the present application;
[0070] Figure 2 Schematic diagram of input and output of the multimodal large model provided in the embodiment of the present application;
[0071] Figure 3 A schematic diagram of a process for generating a target video clip provided in an embodiment of the present application;
[0072] Figure 4 A schematic diagram of a process for generating a target video clip provided in an embodiment of the present application;
[0073] Figure 5 A schematic diagram of a TV drama storyboard segment provided in an embodiment of the present application;
[0074] Figure 6 A schematic diagram of a storyboard segment and a segment summary provided in an embodiment of the present application;
[0075] Figure 7 A schematic diagram of the structure of a video processing system provided in an embodiment of the present application;
[0076] Figure 8A schematic diagram of the structure of a video processing device provided in an embodiment of the present application;
[0077] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0078] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0079] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A including A1 or A2 or A3, and can also be implemented as parameter A including at least two of the three items A1, A2, and A3.
[0080] The embodiments of the present application provide a video processing method, device, electronic device, storage medium and program product. The method combines multiple modal information of each storyboard segment in a video, performs semantic understanding and content extraction on each storyboard segment based on a first large language model, and uses a second large language model based on second instruction information to more accurately locate the target storyboard segment that meets the selection criteria in terms of content, thereby improving video production efficiency and better meeting actual application needs.
[0081] The methods provided in the embodiments of the present application may involve artificial intelligence (AI) technology and may be implemented based on AI technology. For example, a summary of a storyboard segment may be generated using a trained multimodal macromodel. The trained multimodal macromodel may be trained using machine learning (ML) based on a training sample set.
[0082] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0083] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0084] Computer vision (CV) is the science of enabling machines to "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye to identify, detect, and measure objects. Further image processing is performed to transform the computer-generated images into images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, three-dimensional object reconstruction, 3D (three-dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0085] Machine learning (ML) is a multidisciplinary field that draws on a variety of disciplines, including probability theory, statistics, approximation paths, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge and skills, reorganize existing knowledge structures, and continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental path to computer intelligence. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0086] Optionally, the solutions provided in the embodiments of the present application may involve various fields of cloud technology, such as cloud computing, cloud storage, and cloud video in cloud technology. For example, the video processing method provided in the embodiments of the present application may be executed by a server, which may be a cloud server. The data processing involved in the method (such as storyboard processing) may be implemented using cloud computing, and the data storage involved in the embodiments of the present application (such as the storage of storyboard clips) may be implemented using cloud storage.
[0087] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool and be used on demand, flexibly and conveniently.
[0088] Among them, cloud computing is the product of the integration of the development of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.
[0089] Cloud storage is a new concept that has been extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0090] Cloud video refers to a video network platform service based on a cloud computing business model. On the cloud platform, all video suppliers, agents, planning service providers, producers, industry associations, regulatory agencies, industry media, and legal structures are centralized and integrated into a resource pool. These resources are displayed and interact with each other, allowing on-demand communication and consensus building, thereby reducing costs and improving efficiency. This concept is the concept of cloud video. For example, the video to be processed in the embodiments of this application can be a video from the cloud platform.
[0091] It should be noted that in the optional embodiments of the present application, the data related to the object information (videos uploaded by users, etc.) involved, when the embodiments in the present application are applied to specific products or technologies, need to obtain the permission or consent of the object, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. In other words, if the embodiments of the present application involve data related to the object, it needs to be obtained through the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained. The embodiments also need to be implemented with the authorization and consent of the object.
[0092] The following describes several embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, learn from, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0093] Figure 1 A flow chart of a video processing method provided in an embodiment of the present application is shown. The method can be executed by any electronic device, such as a user terminal or a server, or can be implemented by multiple electronic devices in cooperation.
[0094] Among them, the above-mentioned server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server (which can be called the cloud) that provides cloud computing services. The terminal (also called a user terminal or user device) can be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a car terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, and this application does not limit this.
[0095] like Figure 1As shown, the video processing method provided in the embodiment of the present application may include the following steps S110 to S170.
[0096] Step S110: Obtain the video to be processed.
[0097] Step S120: performing storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed.
[0098] The solution provided in the embodiments of the present application can be applied to any application scenario that requires generating corresponding target video clips through video processing. The video to be processed can be a video in any application scenario, such as short videos, film and television videos (TV series, variety shows, movies, cartoons, documentaries, character interviews, etc.), live videos, game videos, sports event videos, etc.
[0099] Among them, the storyboard processing is the segmentation of the video in the time dimension, and the resulting storyboard segments contain relatively coherent video image content over a period of time. This application does not limit the division method of the storyboard, and it can be set as needed. For example, the division can be based on the camera movement of the acquisition device, and the picture captured by each camera movement is regarded as a storyboard segment. It can also be divided based on the changes in the video picture / scene, or it can be divided based on preset time intervals.
[0100] Optionally, when storyboarding the video to be processed, scene detection can be performed on the video to obtain scene change boundaries, and based on the scene change boundaries, the video to be processed can be segmented to obtain multiple storyboard segments. This application does not limit the algorithm / tool used for scene detection. As an optional implementation, PySceneDetect (a scene detection tool / shot segmentation tool) can be used to storyboard the video to be processed.
[0101] Step S130: Determine the multimodal information of each storyboard segment.
[0102] Video is a multifunctional medium that can convey information through multiple modalities such as text and vision. In the embodiment of the present application, the multimodal information of the storyboard segment may include text information corresponding to the storyboard segment and each video frame in the storyboard segment.
[0103] The text information corresponding to the storyboard segments may be text information used to represent the content of the storyboard segments. For example, when the video to be processed is a film or television video, the text information corresponding to the storyboard segments may be the lines of the film or television video; when the video to be processed is a sports event video, the text information corresponding to the storyboard segments may be the real-time score of the game.
[0104] Optionally, for each storyboard segment, the text information corresponding to the storyboard segment may be obtained using at least one of the following methods:
[0105] Automatic Speech Recognition (ASR) is performed on the audio information of the storyboard segment to obtain text information corresponding to the storyboard segment. As an optional implementation, the audio information of the storyboard segment can be input into a trained speech recognition model to obtain recognized text corresponding to the audio information as the text information corresponding to the storyboard segment.
[0106] Character recognition is performed on each video frame in the storyboard segment to obtain text information corresponding to the storyboard segment. As an optional implementation, each video frame in the storyboard segment can be input into a trained OCR recognition model to obtain recognized text for each video frame. Based on the recognized text for each video frame, the text information corresponding to the storyboard segment is determined.
[0107] Optionally, the multimodal information of a storyboard segment may also include object information for the storyboard segment. For each storyboard segment, the object information for the storyboard segment is obtained by identifying the objects in each video frame of the storyboard segment. The objects in the video frames may be real or virtual, animate or inanimate, and this application does not impose any restrictions on this. For example, the objects in the video frames may be stars, pandas, Sun Wukong, sunflowers, oil paintings, etc.
[0108] Optionally, when identifying objects in video frames, each video frame can be input into a trained object recognition model to identify the objects in each video frame. For example, when the video to be processed is a movie, each video frame in the movie can be input into a trained face recognition model to identify the actors in each video frame.
[0109] Step S140: Based on the multimodal information of each storyboard segment and the first instruction information, a segment summary of each storyboard segment is generated by using the trained first language model.
[0110] The storyboard segment summary is a summary of the storyboard segment's content, and the first instruction is used to instruct the generation of a storyboard segment summary. For example, the first instruction could be, "Your task is to output a detailed description of this video segment based on the input you received, including the environment, plot, theme, and emotion. The input you received includes the video's images, characters, and lines. The characters in the segment include XX, and the lines are XXX."
[0111] The first large language model can adopt a multimodal large model, such as the LLaVA (Large Language and Vision Assistant) model, the BLIP (Bootstrapping Language-Image Pre-training) model, etc. The embodiment of the present application does not limit the model structure of the multimodal large model, as long as it can combine the multiple modal information of multiple video clips to generate a segment summary corresponding to the video clip.
[0112] Specifically, for each storyboard segment, based on the multimodal information of the storyboard segment and the first instruction information, the following operations can be performed through the trained first large language model to obtain a segment summary of the storyboard segment: visual features of at least part of the video frames in the storyboard segment are extracted to obtain the visual features of the storyboard segment; text features corresponding to the storyboard segment and the first instruction information are extracted to obtain the text features corresponding to the storyboard segment; the visual features of the storyboard segment and the corresponding text features are fused with multimodal features, and based on the fused features, a segment summary of the storyboard segment is generated.
[0113] Optionally, due to differences in scenes and content between different storyboard segments, the duration and number of video frames contained in different storyboard segments are not exactly the same, and the video image generally does not change significantly within a storyboard segment. Therefore, in order to unify the dimensionality of the model input data and reduce the amount of computation, for each storyboard segment, each video frame in the storyboard segment can be sampled to obtain a preset number of video frames (i.e., at least a portion of the video frames) as the video frames corresponding to the storyboard segment. Furthermore, based on the video frames corresponding to each storyboard segment, the multimodal information of each storyboard segment is determined.
[0114] Optionally, when sampling the storyboard segments, uniform sampling can be performed based on the time dimension to obtain a preset number of video frames. For example, assuming the preset number is 8 frames, the duration of storyboard segment 1 is 8 seconds, and the duration of storyboard segment 2 is 12 seconds, then for storyboard segment 1, one frame of video image can be extracted from each second to obtain 8 frames of video image as the video frames corresponding to storyboard segment 1; for storyboard segment 2, one frame of video image can be extracted from each 1.5 seconds to obtain 8 frames of video image as the video frames corresponding to storyboard segment 2.
[0115] Optionally, for each storyboard segment, when performing visual feature extraction, visual feature extraction can be performed on each video frame based on a plurality of sampled video frames using a visual feature (image feature) extraction model to obtain visual features of each video frame, and the visual features of the storyboard segment can be obtained based on the visual features of the plurality of video frames sampled from the storyboard segment. The present application does not impose any restrictions on the network structure of the visual feature extraction model. For example, structures such as Video SwinTransformer (a video classification network structure based on Transformer) and CLIP (Contrastive Language-Image Pre-training) can be used.
[0116] Optionally, the visual feature extraction model can serve as a visual feature extraction layer in the first large language model, with multiple video frames of each storyboard serving as input to the first large language model, and the visual features of each storyboard extracted via the visual feature extraction layer. Alternatively, multiple video frames of each storyboard can be input to the visual feature extraction model, visual features of each storyboard can be extracted, and the visual features of each storyboard can be used as input to the first large language model. This embodiment of the present application does not impose any restrictions on this, and can be configured as needed.
[0117] Optionally, since the visual features extracted from the video frames are features in the visual space and cannot be used as understandable input for the language model, in this embodiment of the application, semantic mapping can also be performed on the extracted visual features, that is, the visual features of the storyboard segment are mapped to the feature space corresponding to the text features to obtain mapped features corresponding to the visual features, and the mapped features are multimodally fused with the text features to obtain fused features.
[0118] Optionally, when mapping visual features to text features corresponding to text features, semantic mapping of visual features can be achieved through a cross-modal matching module. This application does not restrict the network structure of the cross-modal matching module, and a transformer structure, an MLP (Multi Layer Perceptron) structure, etc. can be used.
[0119] Assume that N video frames are sampled from each storyboard segment. Visual feature extraction is performed on each video frame to obtain 1×d-dimensional visual features. Based on the visual features of N video frames, the N×d-dimensional visual features of the storyboard segment are obtained. The cross-modal matching module converts the N×d-dimensional visual features of the storyboard segment into N×c-dimensional features that can be understood by the language model (i.e., mapped into the feature space of text features).
[0120] In this embodiment of the present application, video metadata of the video to be processed can also be obtained. For each storyboard segment of the video to be processed, a segment summary of the storyboard segment can be generated using the trained first language model based on the multimodal information of the storyboard segment, the video metadata, and the first instruction information. The video metadata includes basic video information such as the video type and video name, and the multimodal information of the storyboard segment includes text information corresponding to the storyboard segment, each video frame in the storyboard segment, and object information of the storyboard segment.
[0121] Optionally, video metadata may also include specific information corresponding to the video type. For example, if the video being processed is a film or TV series, the video metadata may also include the cast information (the correspondence between actors and roles). If the video being processed is a sports event, the video metadata may also include information about the participating teams, referees, coaches, commentators, and other characters.
[0122] Optionally, when extracting multimodal features through the first large language model, the text information, object information corresponding to the storyboard clips and the video metadata of the video to be processed can be fused into the first instruction information, and text feature extraction can be performed on the fused first instruction information to obtain the text features of the storyboard clips.
[0123] For example, Figure 2 As shown in the figure, Xv represents the visual features of each storyboard segment, and Xq represents the textual description of each storyboard segment. The textual description of each storyboard segment includes the corresponding textual information, object information, video metadata, and first instruction information. Semantic mapping is performed on the visual features Xv of each storyboard segment to obtain the mapped features Hv. Feature extraction is performed on the textual description of each storyboard segment to obtain the textual features Hq of each storyboard segment. Based on the visual features Xv and the textual features Xv of each storyboard segment, a segment summary Xa for each storyboard segment is obtained using a multimodal large model.
[0124] Optionally, different first instruction information may be set for different videos to be processed, or for different video types. The first instruction information corresponding to the video to be processed is determined based on the video type to be processed. For example, the first instruction information set for a romance drama may be different from that for a suspense drama.
[0125] Exemplarily, the first instruction information may be set as:
[0126] Instruction 1: "You are an expert at understanding scenes based on visual features in videos. Given a sequence of images, identify the scene in the video and generate a detailed, clear, and accurate description of the scene. The characters in the video clip include XX, and the lines are XX."
[0127] Instruction 2: "Your task is to output a detailed description of this video clip, including the setting, plot, theme, and emotion. The input you receive includes each frame of the video, the characters, the lines, and background information. This is an urban female inspirational drama, and the character's line in the clip is XX."
[0128] Instruction 3: "You are an expert in understanding the descriptions of different scenes in the video. Please use the information provided (including the title, script, lines, and images) to generate a complete description of the scene. You may find character names in the title or script. Use character names in the description whenever possible."
[0129] Optionally, the training process of the first large language model can be divided into a pre-training stage and a fine-tuning stage. Among them, the pre-training stage uses a large number of pictures and texts (or videos and corresponding texts) to pre-train the multimodal large model to be trained to achieve cross-modal alignment of vision and text. The fine-tuning stage adaptively adjusts the pre-trained multimodal large model to make it suitable for different task scenarios and better capture the necessary visual information in the task scenario. Among them, the fine-tuning stage can use fine-tuning techniques such as Fine-tuning, Prompt-tuning, and Instruction-tuning to adjust the pre-trained model.
[0130] Optionally, different video types may have different focus points. For example, romantic TV dramas focus more on the plot, while science fiction films focus more on exciting special effects. Therefore, for different video types, a top language model corresponding to the video type can be trained. When processing a video, the top language model corresponding to the video type to be processed is used for processing. For example, if the video to be processed is a film or TV video, a top language model corresponding to the film or TV video can be trained. Furthermore, top language models for different film and TV genres such as suspense dramas, romantic dramas, and palace intrigue dramas can be trained.
[0131] Optionally, during the fine-tuning phase, the pre-trained multimodal large model can be trained using the following methods:
[0132] Step A1: Acquire a plurality of first training samples and first sample instruction information.
[0133] Each first training sample includes multimodal information of a sample video storyboard segment and a sample segment summary of the sample video storyboard segment. The sample video storyboard segment is a storyboard segment in the first sample video. The first sample video and the to-be-processed video use the same storyboard processing method, and the multimodal information of the sample video storyboard segment is also of the same type as the multimodal information contained in the storyboard segment of the to-be-processed video.
[0134] Optionally, when training the first language model corresponding to the type of film and television drama, the first sample video selected can be a film and television drama video (such as a TV series). Taking the TV series as an example, when aligning the storyboard clips and the sample clip summaries of the storyboard clips, the episode plot descriptions (or episode scripts) of the TV series can be obtained. Based on the episode plot descriptions of the TV series and the storyboard clips of the TV series, the sample clip summaries of the storyboard clips of the TV series can be obtained through the pre-trained multimodal large model (or the existing open source multimodal large model).
[0135] Step A2: Based on multiple first training samples and first sample instruction information, continuously perform training operations on the pre-trained multimodal large model until the first training end condition is met, and use the trained multimodal large model as the first large language model.
[0136] For different training sample sets, the same first sample instruction information may be used, or different first sample instruction information may be used. When different first sample instruction information is used, different natural language instruction templates may be selected.
[0137] Training operations include:
[0138] Step A21: inputting the multimodal information of each sample video storyboard segment of the first sample instruction information into the pre-trained multimodal large model to obtain a predicted segment summary of each sample video storyboard segment;
[0139] Step A22: determining a first training loss based on a difference between the predicted segment summary and the sample segment summary of each sample video storyboard segment;
[0140] Step A23: Adjust the model parameters of the multimodal large model based on the first training loss.
[0141] Among them, the above-mentioned first training end condition and the loss function of the model can be configured according to needs. For example, the first training end condition may include but is not limited to the number of training times reaching a preset number, the loss function convergence (such as the training loss of the model is less than a preset value, or the training losses for multiple consecutive times are less than a preset value, etc.), the test indicators of the model meet the preset indicators, etc. The first training loss of the model represents the deviation between the predicted segment summary of each sample video storyboard segment output by the model and the sample segment summary (real segment summary). Through the training loss function of each training sample and the model, the model is supervisedly trained using the gradient descent algorithm, which can make the segment summary of the storyboard segment predicted by the model can continuously approach the real segment summary of the storyboard segment, so that a well-trained first language model that meets the needs of actual applications can be obtained.
[0142] Step S150: Obtain second instruction information.
[0143] Step S160: Based on the segment summaries of the storyboard segments and the second instruction information, at least one target storyboard segment is determined from the storyboard segments by using the trained second language model.
[0144] Among them, the second instruction information is used to indicate the selection of video clips and the corresponding video clip selection conditions. The video clip selection conditions of different videos may be different. The construction of the second instruction information has a very important impact on the reasoning ability and output effect of the large language model. Appropriate instruction information (prompt words) with sufficient information can better exert the semantic understanding and abstraction ability of the large language model, and help the model generate more accurate and valuable output results. Optionally, the second instruction information can be a target instruction construction template selected from a variety of instruction construction templates, and can also include manually input prompt information (prompt words).
[0145] For example, when used to generate a trailer for a TV series, the instruction construction template can be set to: "As a senior TV series editor, please select the clips that are suitable as the series trailer based on the TV series introduction and the shot description of this episode, and output the trailer's plot summary (also a whole paragraph of text, describing the summary of the selected clip) and the corresponding clip timestamp. The trailer content should try to include suspenseful events, such as separation, quarrels, misunderstandings, conflicts, etc.; try to select clips around the main plot line, and do not reveal all the highlights. The introduction of the TV series is: [Introduction], and the shot descriptions are as follows: [00:00:00 segment plot details 1, 00:00:19 segment plot details 2, ...]".
[0146] Among them, the segmented plot details in the shot description are the fragment summaries of each storyboard obtained through the first language model. The time in the shot description is the starting timestamp of each storyboard. The introduction of the film and television series is included in the video metadata.
[0147] For example, the manually input prompt information may be "This is an urban female inspirational drama. Please select a clip centered around the heroine that is suitable as a series trailer."
[0148] Optionally, the second largest language model may also generate a target video summary corresponding to the target video clip. When the target video clip is a trailer for the next episode of a TV series, a trailer summary for the trailer for the next episode may be generated.
[0149] Because individual storyboard segments only contain the content of that segment, they cannot control the entire video. In this embodiment of the present application, the segment summaries of each storyboard segment can be aggregated to obtain a video summary of the video to be processed. The video summary of the video to be processed is a distillation of the overall content of the video to be processed and a description of the content of the complete video. Subsequently, based on the segment summaries of each storyboard segment, the video summary of the video to be processed, and the second instruction information, at least one target storyboard segment is determined from each storyboard segment using the trained second language model.
[0150] Optionally, when performing content aggregation on the segment summaries of each storyboard segment, a video summary of the video to be processed can be obtained based on the segment summaries of each storyboard segment of the video to be processed through a trained third language model.
[0151] Optionally, the second language model can be trained in the following way:
[0152] Step B1: Acquire multiple second training samples and second sample instruction information.
[0153] Each second training sample includes a sample segment summary of each storyboard segment of a second sample video and a label of each storyboard segment. A label of a storyboard segment indicates whether the storyboard segment is a target storyboard segment in the corresponding second sample video.
[0154] Optionally, the sample segment summaries of each storyboard segment of the second sample video can be generated by the trained first language model, or the sample segment summaries of the storyboard segments in the second training sample can be generated by the same method as the sample segment summaries of the above-mentioned first training sample.
[0155] Optionally, when the second sample video is an episode of a TV series, an existing episode trailer for the TV series main episode can be obtained, and the episode trailer can be storyboarded to obtain multiple trailer storyboard segments in the episode trailer, as well as the timestamps of each trailer storyboard segment in the TV series main episode. The TV series main episode is storyboarded to obtain multiple storyboard segments of the TV series main episode. The trailer storyboard segment in the episode trailer is the target storyboard segment in the TV series main episode. If the storyboard segment is a trailer storyboard segment in the episode trailer, then the storyboard segment is the target storyboard segment in the TV series main episode.
[0156] Step B2: Based on the multiple second training samples and the second sample instruction information, the training operation is continuously performed on the large language model to be trained until the second training end condition is met, and the trained large language model is used as the second large language model.
[0157] The training operations include:
[0158] Step B21: Input the second sample instruction information and the sample segment summary of each storyboard segment in each second training sample into the second largest language model to obtain the predicted probability of each storyboard segment in each second training sample.
[0159] The predicted probability represents the probability that the storyboard segment is the target storyboard segment in the corresponding second sample video;
[0160] Step B22: determining a second training loss based on the predicted probability of each storyboard segment and the label of each storyboard segment in each second training sample;
[0161] Step B23: Adjust the model parameters of the second largest language model based on the second training loss.
[0162] Among them, the above-mentioned second training end condition and the loss function of the model can be configured according to needs. Please refer to the above-mentioned specific description of the first training end condition, and this application will not repeat them here.
[0163] Optionally, the second instruction information may be generated in the following manner:
[0164] In response to a video editing trigger operation for the video to be processed, a video editing setting interface is displayed; wherein the video editing setting interface displays editing operation options, and the editing operation options include a video clip selection operation option;
[0165] In response to a trigger operation on a video clip selection operation option, displaying a video clip selection condition setting interface;
[0166] Receiving at least one of an input operation of a selection keyword or an instruction template selection operation for a video to be processed through a video clip selection condition setting interface;
[0167] Based on at least one of the received keyword selection operation or instruction template selection operation for the video to be processed, second instruction information corresponding to the video to be processed is obtained.
[0168] The user may input a selected keyword, or select a target instruction template from a plurality of pre-set instruction templates, and use at least one of the keyword input by the user or the selected target instruction template as the second instruction information.
[0169] Step S170: Fusing the target storyboard segments to generate a target video segment corresponding to the video to be processed.
[0170] The embodiment of the present application does not limit the method of merging the target storyboard segments, and the method can be set as needed. For example, the target storyboard segments can be spliced in chronological order, or transition effects (transition effects) can be added between the target storyboard segments.
[0171] The target video clip is a video clip composed of some storyboard clips (target storyboard clips) in the video to be processed. For example, if the video to be processed is a TV series, the target video clips may include a synopsis of the previous episode, a preview of the next episode (episode previews), highlights, opening credits, and ending credits. If the video to be processed is a movie, the target video clip may be a movie trailer. If the video to be processed is a live broadcast, the target video clip may be a live broadcast segment.
[0172] In an embodiment of the present application, there may be a duration constraint for the generated target video clip. For example, when the target video clip is a trailer, in order to attract users in a short time and guide users to watch the main film, the trailer is usually within the time range of 3-5 minutes.
[0173] The maximum duration limit of the target video segment to be generated can then be obtained, and the duration of each target storyboard segment can be determined. If the sum of the durations of the target storyboard segments is less than or equal to the maximum duration limit, the target storyboard segments are spliced together to generate the target video segment corresponding to the video to be processed. If the sum of the durations of the target storyboard segments is greater than the maximum duration limit, based on the maximum duration limit and the relevant information corresponding to each target storyboard segment, some video frames are selected from each target storyboard segment, and the selected video frames are spliced together to generate the target video segment corresponding to the video to be processed.
[0174] The relevant information corresponding to each target storyboard segment includes at least one of the following:
[0175] The correlation between each target storyboard segment;
[0176] the correlation between each target storyboard segment and the video summary of the video to be processed, wherein the video summary is obtained by content aggregation of the segment summaries of each storyboard segment;
[0177] Correlation between video frames of different target storyboards.
[0178] Optionally, when selecting part of the video frames from each target storyboard segment based on the correlation between the target storyboard segments, the correlation between the target storyboard segments can be determined based on the segment summary of each target storyboard segment. The stronger the content association, the higher the correlation, and the video frames from the strongly correlated part of the target storyboard segments in each storyboard segment are selected.
[0179] Optionally, when selecting some video frames from each target storyboard segment based on the correlation between each target storyboard segment and the video summary of the video to be processed, the correlation between each target storyboard segment and the video summary of the video to be processed can be determined based on the segment summary of each target storyboard segment and the video summary of the video to be processed. The stronger the content association, the higher the correlation, and video frames from some strongly correlated target storyboard segments can be selected from multiple target storyboard segments.
[0180] Optionally, when selecting some video frames from each target storyboard segment based on the correlation between the video frames of different target storyboard segments, the similarity of the video frames in adjacent target storyboard segments can be determined, and the similar (or repeated) video frames can be deleted, and the video frames in the deleted target storyboard segments can be spliced.
[0181] Optionally, when the sum of the durations of the target storyboard segments exceeds the maximum duration limit, multiple target storyboard segments with higher probabilities may be selected based on the ranking of the predicted probabilities of the target storyboard segments output by the second largest language model, such that the sum of the durations of the multiple target storyboard segments with higher probabilities does not exceed the maximum duration limit. A target video segment corresponding to the video to be processed is generated based on the multiple target storyboard segments with higher probabilities.
[0182] Optionally, since the target storyboard segment is too short to clearly display the content of the segment, and the target storyboard segment is too long to be redundant, the embodiment of the present application can further constrain the duration of the target storyboard segment. For each target storyboard segment, if the duration of the target storyboard segment is greater than the maximum segment duration limit, the duration of the target storyboard segment is adjusted according to the first segment adjustment strategy, so that the adjusted duration of the target storyboard segment is less than the maximum segment duration limit. Among them, the embodiment of the present application does not restrict the first segment adjustment strategy. As an optional implementation method, sampling can be used to shorten the duration of the target storyboard segment.
[0183] If the duration of the target storyboard segment is less than the minimum segment duration limit, the target storyboard segment is adjusted according to the second segment adjustment strategy so that the adjusted duration of the target storyboard segment is greater than the minimum segment duration limit. The present embodiment does not limit the second segment adjustment strategy. As an optional implementation, the target storyboard segment can be adaptively extended before and after the target storyboard segment in the video to be processed (the main film).
[0184] based on Figure 1The video processing method shown combines the multi-modal information of each storyboard segment in the video, performs semantic understanding and content extraction on each storyboard segment based on the first large language model, and the obtained segment summaries of each storyboard segment are more accurate. The second large language model is used to more accurately locate the target storyboard segment that meets the selection criteria based on the second instruction information, thereby realizing the selection of video segments based on the content, improving video production efficiency, and better meeting actual application needs.
[0185] It should be noted that, although the current multimodal large model has significantly improved its understanding of images and videos, it is still very difficult to process long videos, and in practical applications it is difficult to take into account both text extraction (OCR text recognition or speech recognition) and object recognition. The accuracy of the multimodal large model for text recognition and object recognition still cannot exceed that of the single model. Therefore, in an embodiment of the present application, the video to be processed can be first subjected to basic storyboard processing to obtain each storyboard segment, and each storyboard segment is pre-processed based on each single model (character recognition model, speech recognition model, object recognition model) to obtain the text information corresponding to each storyboard segment and the objects in the storyboard segment. Based on the text information corresponding to each storyboard segment, the objects in the storyboard segment, and the video frames in the storyboard segment, etc., a segment summary of each storyboard segment is obtained through the multimodal large model.
[0186] In the embodiment of the present application, when storyboarding is performed on the video to be processed, each storyboard segment obtained also includes corresponding segment identification information.
[0187] Accordingly, when generating the segment summaries of each storyboard segment based on the first language model, the text information corresponding to each storyboard segment as input and each video frame in each storyboard segment may respectively carry segment identification information.
[0188] When selecting target storyboard segments based on the second largest language model, the segment summary of each storyboard segment, the segment identification information of each storyboard segment, and the second instruction information may be used to determine at least one target segment identification information through the trained second largest language model. The segment identification information of each storyboard segment corresponds to the segment summary of each storyboard segment. Based on the determined at least one target segment identification information, at least one target storyboard segment is determined from each storyboard segment. The target segment identification information is a segment identification information in the segment identification information of each storyboard segment.
[0189] Optionally, the segment identification information may be a serial number of the storyboard segment in the video to be processed (main film), or may be a timestamp (including a start timestamp and an end timestamp) corresponding to the storyboard segment in the video to be processed.
[0190] When the segment identification information is a timestamp, the storyboard segments and their corresponding timestamps are obtained by storyboard processing. The timestamps corresponding to the storyboard segments include a start timestamp and an end timestamp, the start timestamp is the timestamp corresponding to the first frame of video (image) in the storyboard segment, and the end timestamp is the timestamp corresponding to the last frame of video (image) in the storyboard segment. Each modal information in the multimodal information of each storyboard segment carries a timestamp corresponding to the storyboard segment, based on the segment summary of each storyboard segment output by the first large language model and its corresponding timestamp. The timestamp corresponding to the target storyboard segment is determined based on the segment summary of each storyboard segment and its corresponding timestamp, as well as the second instruction information, based on the second large language model. Finally, based on the timestamp corresponding to each target storyboard segment, the corresponding target storyboard segment is intercepted from the video to be processed, and the target video segment corresponding to the video to be processed is generated by fusing each target storyboard segment.
[0191] Figure 3 A flow chart of generating target video segments provided in an embodiment of the present application. After obtaining the video to be processed, the video to be processed can be storyboarded to obtain each storyboard segment, and each storyboard segment can be subjected to multimodal preprocessing to identify the object information in each storyboard segment, and speech recognition / character recognition can be performed on each storyboard segment to obtain the text information corresponding to each storyboard segment. Afterwards, each video frame, object information, corresponding text information in each storyboard segment, and video metadata of the video to be processed are input into the first large language model to obtain a segment summary of each storyboard segment. Then, the generated segment summary of each storyboard segment and the second instruction information constructed based on the task requirements are input into the second large language model to determine a number of target storyboard segments from each storyboard segment. Finally, based on each target storyboard segment, a target video segment corresponding to the video to be processed is generated.
[0192] Figure 4A schematic diagram of a process for generating target video segments provided by an embodiment of the present application. After obtaining a video to be processed, the video can be storyboarded to obtain individual storyboard segments. Each storyboard segment can then be subjected to multimodal preprocessing to identify object information within each storyboard segment. Furthermore, each storyboard segment can be subjected to speech recognition / character recognition to obtain text information corresponding to each storyboard segment. Subsequently, each video frame, object information, corresponding text information within each storyboard segment, and video metadata of the video to be processed are input into a first large language model to obtain a segment summary for each storyboard segment. The generated segment summary for each storyboard segment is then input into a third large language model to obtain a video summary for the video to be processed. The generated segment summary for each storyboard segment, the video summary for the video to be processed, and second instruction information constructed based on task requirements are then input into a second large language model to determine a number of target storyboard segments from each storyboard segment. Finally, based on each target storyboard segment, a target video segment corresponding to the video to be processed is generated.
[0193] In order to better understand and illustrate the method provided in the embodiment of the present application, the optional implementation methods of the method provided in the present application are introduced below in conjunction with a specific scenario embodiment. In this scenario embodiment, the generation of episode previews of a TV series is used as an example for explanation.
[0194] Taking the production of the episode trailer of the third episode of the TV series as an example, we can first obtain the third episode of the TV series, and use PySceneDetect to process the third episode of the TV series, and get the following Figure 5 The multiple storyboard segments shown are shown. Each storyboard segment has a corresponding timestamp (including a start timestamp and an end timestamp), and the timestamp corresponding to the storyboard segment is its timestamp in the feature film.
[0195] For each storyboard clip, ASR / OCR is used to identify the lines in each storyboard clip; a face recognition model is used to identify the actors in each storyboard clip.
[0196] Get the first instruction information and the video metadata of the third episode. The first instruction information is "Your task is to output a detailed description of this video clip, including characters, actions, environment, theme, and emotion. The input you receive includes each frame of the video, characters, lines, and basic video information. This is a romantic idol drama. The characters in the clip are XX, the characters' lines in the clip are XX, and the basic video information is XX." The video metadata of the third episode includes basic information such as the TV series type (such as suspense drama, romance drama), TV series title, TV series full series synopsis, TV series third episode synopsis, and TV series cast.
[0197] Next, the dialogue and actors in each storyboard segment, as well as the video metadata for the third episode, are fused into the first instruction information. This fused first instruction information and each video frame (with timestamps) in each storyboard segment are input into the first large language model to obtain a segment summary for each storyboard segment. The segment summary for each storyboard segment is then input into the third large language model to obtain a video summary for the third episode.
[0198] For example, Figure 6 As shown, Figure 6 It contains four storyboard clips from the main film of Episode 3. For storyboard clip 1, the clip summary 1 obtained based on the multimodal information of storyboard clip 1 is "The prince invites the princess to dance"; for storyboard clip 2, the clip summary 2 obtained based on the multimodal information of storyboard clip 2 is "The prince and princess dance together, attracting the audience on the dance floor to watch"; for storyboard clip 3, the clip summary 3 obtained based on the multimodal information of storyboard clip 3 is "The prince lifted the princess up, and there was constant applause all around"; for storyboard clip 4, the clip summary 4 obtained based on the multimodal information of storyboard clip 4 is "The prince slipped and the two fell to the ground."
[0199] It should be noted that, in the figure, two frames of images are used to represent one storyboard segment, but in actual application, one storyboard segment contains multiple frames of images.
[0200] Then, obtain the second instruction information, where the second instruction information is "As a senior film and television drama editor, please select clips suitable as episode trailers based on the film and television drama introduction, the shot description of this episode, and the plot description of this episode, and output the corresponding clip timestamps. The trailer content should contain as many highlight clips as possible, such as sweetness, fighting, hugging, or expressions of intense emotions; try to select clips around the main plot line. The introduction of the film and television drama is: [Introduction], the plot description of this episode is [Plot Details], and the shot descriptions are as follows: [00:00:00 segment plot details 1, 00:00:25 segment plot details 2, ...]".
[0201] The fragment summary (with timestamp) of each storyboard segment and the video summary of the third episode (the plot description of the above episode) are integrated into the second instruction information. The integrated second instruction information is input into the second largest language model to select several highlight segments (i.e., target storyboard segments) from each storyboard segment.
[0202] Finally, by splicing the highlight clips together, a trailer for the third episode of the TV series is generated.
[0203] This application solution parses longer TV series videos and, by combining a large multimodal model with various recognition capabilities, generates a summary of each storyboard segment. Based on these summaries and instruction information constructed based on task requirements, it uses a large language model to locate compelling trailer segments. Finally, after processing, the resulting trailer is automatically generated. By analyzing the plot content of TV series videos and locating compelling, high-profile segments to generate trailers, this significantly improves video production efficiency and attracts users to the main series based on the plot content.
[0204] At the same time, it also supports users to specify plot setting command keywords according to their needs, and AI automatically locates relevant clips to produce trailers, thereby improving video operation effects.
[0205] The video processing method provided in the embodiment of the present application can be applied to Figure 7 The video processing system shown includes a terminal 50 and a server 52. When performing video processing, the terminal 50 can send the video to be processed to the server 52 via the network. The server 52 uses the method shown in steps S110 to S170 above to process the video to obtain a target video segment corresponding to the video to be processed, and then sends the processed target video segment to the terminal 50.
[0206] The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal may be a smartphone, tablet computer, laptop computer, desktop computer, smart home appliance (such as a smart TV), AR / VR device, etc., but is not limited thereto.
[0207] Based on the same principle as the video processing method provided in the embodiment of the present application, the embodiment of the present application provides a video processing device, such as Figure 8 As shown, the video processing device 400 may include a video acquisition module 410, a storyboard module 420, a multimodal information determination module 430, a segment summary determination module 440, an instruction acquisition module 450, a target storyboard determination module 460 and a target video generation module 470.
[0208] The video acquisition module 410 is used to acquire the video to be processed;
[0209] A storyboard module 420 is configured to perform storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed;
[0210] A multimodal information determination module 430 is configured to determine multimodal information of each storyboard segment, wherein the multimodal information includes text information corresponding to the storyboard segment and each video frame in the storyboard segment;
[0211] A segment summary determination module 440 is configured to generate a segment summary for each storyboard segment using a trained first language model based on the multimodal information of each storyboard segment and first instruction information; the first instruction information is configured to instruct the generation of a segment summary for each storyboard segment.
[0212] The instruction acquisition module 450 is used to acquire second instruction information, where the second instruction information is used to instruct video segment selection and corresponding video segment selection conditions;
[0213] a target storyboard determining module 460 for determining at least one target storyboard segment from each of the storyboard segments based on the segment summary of each storyboard segment and the second instruction information and using the trained second language model;
[0214] The target video generation module 470 is used to fuse the target storyboard segments to generate a target video segment corresponding to the video to be processed.
[0215] Optionally, the segment summary determination module 440 is further configured to:
[0216] Obtaining video metadata of the video to be processed;
[0217] For each of the storyboard segments, identifying objects in each video frame in the storyboard segment to obtain object information of the storyboard segment; wherein the multimodal information also includes the object information of the storyboard segment;
[0218] For each of the storyboard segments, a segment summary of the storyboard segment is generated by a trained first language model based on the multimodal information of the storyboard segment, the video metadata, and the first instruction information.
[0219] Optionally, the target mirror determination module 460 may be configured to:
[0220] performing content aggregation on the segment summaries of the storyboard segments to obtain a video summary of the video to be processed;
[0221] Based on the segment summaries of the storyboard segments, the video summary of the video to be processed and the second instruction information, at least one target storyboard segment is determined from the storyboard segments through the trained second language model.
[0222] Optionally, for each storyboard segment, the segment summary determination module 440 may be configured to:
[0223] Extracting visual features from at least a portion of the video frames in the storyboard segment to obtain visual features of the storyboard segment;
[0224] Performing text feature extraction on the text information corresponding to the storyboard segment and the first instruction information to obtain text features corresponding to the storyboard segment;
[0225] The visual features of the storyboard segment and the corresponding text features are fused into a multimodal feature, and a segment summary of the storyboard segment is generated based on the fused features.
[0226] Optionally, the segment summary determination module 440 may be configured to:
[0227] Mapping the visual features of the storyboard segment to a feature space corresponding to the text features to obtain mapped features corresponding to the visual features;
[0228] The mapped features and the text features are fused to obtain fused features.
[0229] Optionally, the second instruction information is generated in the following manner:
[0230] In response to a video editing trigger operation for the video to be processed, a video editing setting interface is displayed; the video editing setting interface displays editing operation options, and the editing operation options include a video clip selection operation option;
[0231] In response to a triggering operation on the video clip selection operation option, displaying a video clip selection condition setting interface;
[0232] Receiving at least one of an input operation of selecting a keyword or an instruction template selection operation for the video to be processed through the video segment selection condition setting interface;
[0233] Based on at least one of the received keyword selection operation or instruction template selection operation for the video to be processed, second instruction information corresponding to the video to be processed is obtained.
[0234] Optionally, the target video generation module 470 is further configured to:
[0235] Get the maximum duration limit of the target video segment to be generated;
[0236] The target video generation module can be used to:
[0237] Determining the duration of each target storyboard segment;
[0238] If the sum of the durations of the target storyboard segments is less than or equal to the maximum duration limit, then the target storyboard segments are spliced together to generate a target video segment corresponding to the video to be processed;
[0239] If the sum of the durations of the target storyboard segments is greater than the maximum duration limit, then based on the maximum duration limit and relevant information corresponding to the target storyboard segments, select some video frames from the target storyboard segments, and splice the selected video frames to generate a target video segment corresponding to the video to be processed;
[0240] The relevant information corresponding to each target storyboard segment includes at least one of the following:
[0241] the correlation between the target storyboard segments;
[0242] a correlation between each of the target storyboard segments and a video summary of the video to be processed, wherein the video summary is obtained by performing content aggregation on the segment summaries of each of the storyboard segments;
[0243] Correlation between video frames of different target storyboards.
[0244] Optionally, the first language model is trained in the following manner:
[0245] Acquire a plurality of first training samples and first sample instruction information, each of the first training samples including multimodal information of a sample video storyboard segment and a sample segment summary of the sample video storyboard segment; the sample video storyboard segment is a storyboard segment in the first sample video;
[0246] Based on the plurality of first training samples and the first sample instruction information, a training operation is continuously performed on the pre-trained multimodal large model until a first training end condition is met, and the trained multimodal large model is used as the first large language model, wherein the training operation includes:
[0247] Inputting the first sample instruction information and the multimodal information of each of the sample video storyboard segments into a pre-trained multimodal large model to obtain a predicted segment summary of each of the sample video storyboard segments;
[0248] determining a first training loss based on a difference between the predicted segment summary and the sample segment summary of each of the sample video storyboard segments;
[0249] Model parameters of the multimodal large model are adjusted based on the first training loss.
[0250] Optionally, the second largest language model is trained in the following manner:
[0251] Acquire a plurality of second training samples and second sample instruction information, each of the second training samples including a sample segment summary of each storyboard segment of a second sample video and a label of each storyboard segment, wherein a storyboard segment label indicates whether the storyboard segment is a target storyboard segment in the corresponding second sample video;
[0252] Based on the plurality of second training samples and the second sample instruction information, a training operation is continuously performed on the large language model to be trained until a second training end condition is satisfied, and the trained large language model is used as a second large language model, wherein the training operation includes:
[0253] Inputting the second sample instruction information and the sample segment summary of each storyboard segment in each second training sample into the second large language model, obtaining a predicted probability of each storyboard segment in each second training sample, wherein the predicted probability represents a probability that the storyboard segment is a target storyboard segment in the corresponding second sample video;
[0254] Determining a second training loss based on the predicted probability of each storyboard segment and the label of each storyboard segment in each of the second training samples;
[0255] Adjust model parameters of the second language model based on the second training loss.
[0256] Optionally, the split mirror module 420 is further configured to:
[0257] Determining segment identification information of each of the storyboard segments;
[0258] The target mirror determination module can be used to:
[0259] Determine at least one target segment identification information by using the segment summary of each storyboard segment, the segment identification information of each storyboard segment, and the second instruction information through the trained second language model;
[0260] At least one target storyboard segment is determined from each of the storyboard segments according to the determined at least one target segment identification information; the target segment identification information is a segment identification information in the segment identification information of each of the storyboard segments.
[0261] Optionally, for each storyboard segment, the text information corresponding to the storyboard segment is obtained by at least one of the following methods:
[0262] Performing speech recognition on the audio information of the storyboard segment to obtain text information corresponding to the storyboard segment;
[0263] Character recognition is performed on each video frame in the storyboard segment to obtain text information corresponding to the storyboard segment.
[0264] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application. The implementation principles are similar and can produce the same technical effects. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, and will not be repeated here.
[0265] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0266] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program stored in the memory, the method in any optional embodiment of the present application can be implemented.
[0267] Figure 9 FIG. 1 shows a schematic structural diagram of an electronic device to which an embodiment of the present invention is applicable. Figure 9 As shown, the electronic device may be a server or a user terminal, and the electronic device may be used to implement the method provided in any embodiment of the present invention.
[0268] like Figure 9 As shown in FIG, the electronic device 2000 may mainly include at least one processor 2001 ( Figure 9 1 ), memory 2002, communication module 2003 and input / output interface 2004 and other components, optionally, the components can be connected and communicated through bus 2005. It should be noted that, Figure 9 The structure of the electronic device 2000 shown in the figure is merely illustrative and does not constitute a limitation on the electronic devices to which the method provided in the embodiments of the present application is applicable.
[0269] Memory 2002 can be used to store operating systems and application programs, etc. Application programs can include computer programs that implement the methods described in the embodiments of the present invention when called by processor 2001, and can also include programs for implementing other functions or services. Memory 2002 can be ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices that can store information and computer programs, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0270] The processor 2001 is connected to the memory 2002 via the bus 2005 and implements corresponding functions by calling the application program stored in the memory 2002. The processor 2001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof, which can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. The processor 2001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0271] The electronic device 2000 can be connected to a network via a communication module 2003 (which may include, but is not limited to, components such as a network interface) to communicate with other devices (such as a user terminal or a server) via the network to implement data interaction, such as sending data to or receiving data from other devices. The communication module 2003 may include a wired network interface and / or a wireless network interface, etc., that is, the communication module may include at least one of a wired communication module and a wireless communication module.
[0272] The electronic device 2000 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 2004. The electronic device 2000 itself can have a display device, and can also be connected to other external display devices through the interface 2004. Optionally, a storage device, such as a hard disk, can also be connected through the interface 2004, so that data in the electronic device 2000 can be stored in the storage device, or data in the storage device can be read, and data in the storage device can also be stored in the memory 2002. It can be understood that the input / output interface 2004 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 2004 can be a component of the electronic device 2000, or it can be an external device connected to the electronic device 2000 when needed.
[0273] Bus 2005, used to connect the various components, may include a path for transmitting information between the components. Bus 2005 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Depending on their function, bus 2005 may be categorized as an address bus, a data bus, a control bus, or the like.
[0274] Optionally, for the solution provided in the embodiment of the present invention, the memory 2002 can be used to store a computer program for executing the solution of the present invention, and be run by the processor 2001. When the processor 2001 runs the computer program, the actions of the method or device provided in the embodiment of the present invention are implemented.
[0275] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.
[0276] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.
[0277] It should be noted that the terms "first," "second," "third," "fourth," "1," "2," etc. (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.
[0278] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0279] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. A video processing method, characterized in that: include: Get the video to be processed; Performing storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed; Determining multimodal information of each storyboard segment, wherein the multimodal information includes: text information corresponding to the storyboard segment and each video frame in the storyboard segment; Based on the multimodal information of each storyboard segment and the first instruction information, a segment summary of each storyboard segment is generated using a trained first language model; the first instruction information is used to instruct the generation of a segment summary of the storyboard segment; Acquire second instruction information, where the second instruction information is used to instruct video segment selection and corresponding video segment selection conditions; Based on the segment summaries of the storyboard segments and the second instruction information, determining at least one target storyboard segment from the storyboard segments by using the trained second language model; The target storyboard segments are fused to generate target video segments corresponding to the video to be processed.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining video metadata of the video to be processed; For each of the storyboard segments, identifying objects in each video frame in the storyboard segment to obtain object information of the storyboard segment; wherein the multimodal information also includes the object information of the storyboard segment; For each of the storyboard segments, the segment summary of the storyboard segment is obtained by: Based on the multimodal information of the storyboard segment, the video metadata and the first instruction information, a segment summary of the storyboard segment is generated by a trained first language model.
3. The method according to claim 1, characterized in that The determining, based on the segment summaries of the respective storyboard segments and the second instruction information, at least one target storyboard segment from the respective storyboard segments by using the trained second language model, comprises: performing content aggregation on the segment summaries of the storyboard segments to obtain a video summary of the video to be processed; Based on the segment summary of each storyboard segment, the video summary of the video to be processed and the second instruction information, at least one target storyboard segment is determined from each storyboard segment through the trained second language model.
4. The method according to claim 1, wherein For each storyboard segment, a segment summary of the storyboard segment is obtained by performing the following operations on the trained first language model based on the multimodal information of the storyboard segment and the first instruction information: Extracting visual features from at least a portion of the video frames in the storyboard segment to obtain visual features of the storyboard segment; Performing text feature extraction on the text information corresponding to the storyboard segment and the first instruction information to obtain text features corresponding to the storyboard segment; The visual features of the storyboard segment and the corresponding text features are fused into a multimodal feature, and a segment summary of the storyboard segment is generated based on the fused features.
5. The method according to claim 4, characterized in that The multimodal feature fusion of the visual features of the storyboard segment and the corresponding text features includes: Mapping the visual features of the storyboard segment to a feature space corresponding to the text features to obtain mapped features corresponding to the visual features; The mapped features and the text features are fused to obtain fused features.
6. The method according to claim 1, wherein The second instruction information is generated in the following manner: In response to a video editing trigger operation for the video to be processed, a video editing setting interface is displayed; the video editing setting interface displays editing operation options, and the editing operation options include a video clip selection operation option; In response to a triggering operation on the video clip selection operation option, displaying a video clip selection condition setting interface; Receiving at least one of an input operation of selecting a keyword or an instruction template selection operation for the video to be processed through the video segment selection condition setting interface; Based on at least one of the received keyword selection operation or instruction template selection operation for the video to be processed, second instruction information corresponding to the video to be processed is obtained.
7. The method according to claim 1, characterized in that The method further comprises: Get the maximum duration limit of the target video segment to be generated; Generating a target video segment corresponding to the video to be processed based on each target storyboard segment includes: Determining the duration of each target storyboard segment; If the sum of the durations of the target storyboard segments is less than or equal to the maximum duration limit, then the target storyboard segments are spliced together to generate a target video segment corresponding to the video to be processed; If the sum of the durations of the target storyboard segments is greater than the maximum duration limit, then based on the maximum duration limit and relevant information corresponding to the target storyboard segments, select some video frames from the target storyboard segments, and splice the selected video frames to generate a target video segment corresponding to the video to be processed; The relevant information corresponding to each target storyboard segment includes at least one of the following: The correlation between the target storyboard segments; a correlation between each of the target storyboard segments and a video summary of the video to be processed, wherein the video summary is obtained by performing content aggregation on the segment summaries of each of the storyboard segments; Correlation between video frames of different target storyboards.
8. The method according to claim 1, characterized in that The first language model is trained in the following way: Acquire a plurality of first training samples and first sample instruction information, each of the first training samples including multimodal information of a sample video storyboard segment and a sample segment summary of the sample video storyboard segment; the sample video storyboard segment is a storyboard segment in the first sample video; Based on the plurality of first training samples and the first sample instruction information, a training operation is continuously performed on the pre-trained multimodal large model until a first training end condition is met, and the trained multimodal large model is used as the first large language model, wherein the training operation includes: Inputting the first sample instruction information and the multimodal information of each of the sample video storyboard segments into a pre-trained multimodal large model to obtain a predicted segment summary of each of the sample video storyboard segments; determining a first training loss based on a difference between the predicted segment summary and the sample segment summary of each of the sample video storyboard segments; Model parameters of the multimodal large model are adjusted based on the first training loss.
9. The method according to claim 1, characterized in that The second largest language model was trained using the following method: Acquire a plurality of second training samples and second sample instruction information, each of the second training samples including a sample segment summary of each storyboard segment of a second sample video and a label of each storyboard segment, wherein a storyboard segment label indicates whether the storyboard segment is a target storyboard segment in the corresponding second sample video; Based on the plurality of second training samples and the second sample instruction information, a training operation is continuously performed on the large language model to be trained until a second training end condition is satisfied, and the trained large language model is used as a second large language model, wherein the training operation includes: Inputting the second sample instruction information and the sample segment summary of each storyboard segment in each second training sample into the second large language model, obtaining a predicted probability of each storyboard segment in each second training sample, wherein the predicted probability represents a probability that the storyboard segment is a target storyboard segment in the corresponding second sample video; Determining a second training loss based on the predicted probability of each storyboard segment and the label of each storyboard segment in each of the second training samples; Adjust model parameters of the second language model based on the second training loss.
10. A video processing device, characterized in that: The device comprises: Video acquisition module, used to acquire the video to be processed; A storyboard module, configured to perform storyboard processing on the video to be processed to obtain a plurality of storyboard segments corresponding to the video to be processed; a multimodal information determination module, configured to determine multimodal information of each storyboard segment, wherein the multimodal information includes text information corresponding to the storyboard segment and each video frame in the storyboard segment; a segment summary determination module, configured to generate a segment summary for each storyboard segment using a trained first language model based on the multimodal information of each storyboard segment and first instruction information; wherein the first instruction information is used to instruct the generation of a segment summary for the storyboard segment; An instruction acquisition module, configured to acquire second instruction information, wherein the second instruction information is used to instruct video segment selection and corresponding video segment selection conditions; a target storyboard determining module, configured to determine at least one target storyboard segment from each of the storyboard segments based on the segment summary of each storyboard segment and the second instruction information and using a trained second language model; The target video generation module is used to fuse the target storyboard segments to generate a target video segment corresponding to the video to be processed.
11. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which implements the method according to any one of claims 1 to 9 when executed by a processor.
13. A computer program product, characterized in that The computer product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Advertisement material automatic fission method and system based on artificial intelligence
CN121258610A