Description video generation method and device based on large model, equipment and medium
By combining a large multimodal model with a large language model, the problems of one-sided content and poor coherence in automated video commentary are solved, high-quality generation of commentary videos is achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202511039028.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-14
AI Technical Summary
In the existing technology, the results of automatic video commentary generation are one-sided, lack connection between segments, and provide poor user viewing experience.
By using a large multimodal model to understand the visual content of non-captioned clips, combining the subtitle text and timestamps, and using a large language model to generate commentary, and through semantic analysis and shot segmentation technology, ensure that the commentary matches the time and semantics of the video content.
Improve the overall quality of commentary videos and the user viewing experience by better reflecting the temporal evolution and semantic cohesion of video content, and improve the fit between commentary and video content.
Smart Images

Figure CN120786152A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical fields of multi-modal, natural language processing, computer vision and deep learning, and specifically relates to a large model-based explanation video generation method, a large model-based explanation video generation device, an electronic device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] Artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of humans, which has both hardware and software technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technologies mainly include natural language processing technology, computer vision technology, speech recognition technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.
[0003] The methods described in this section can not necessarily be the methods previously conceived or adopted. Unless otherwise indicated, nothing in this section should be assumed to be prior art merely because it is included in this section. Similarly, unless otherwise indicated, issues raised in this section should not be assumed to have been recognized in any prior art. SUMMARY
[0004] The present disclosure provides a large model-based explanation video generation method, a large model-based explanation video generation device, an electronic device, a computer readable storage medium and a computer program product.
[0005] According to an aspect of the present disclosure, a large model-based explanation video generation method is provided, comprising: obtaining a plurality of subtitle texts and corresponding first time stamps in a to-be-processed video; determining at least one subtitle-free segment and corresponding second time stamps in the to-be-processed video based on the first time stamps of the plurality of subtitle texts; performing visual content understanding on the at least one subtitle-free segment using a first multi-modal large model to obtain at least one subtitle completion text corresponding to the at least one subtitle-free segment; using a large language model, generating an explanation word for the to-be-processed video based on the plurality of subtitle texts and the corresponding first time stamps and the at least one subtitle completion text and the corresponding second time stamps; and generating an explanation video based on the explanation word.
[0006] According to another aspect of the present disclosure, a large model-based commentary video generation apparatus is provided, comprising: a subtitle obtaining unit configured to obtain a plurality of subtitle texts and corresponding first timestamps in a to-be-processed video; a subtitle-free segment determining unit configured to determine at least one subtitle-free segment and corresponding second timestamps in the to-be-processed video based on the first timestamps of the plurality of subtitle texts; a content understanding unit configured to perform visual content understanding on the at least one subtitle-free segment by using a first multi-modal large model to obtain at least one subtitle completion text corresponding to the at least one subtitle-free segment; a commentary word generating unit configured to generate commentary words for the to-be-processed video based on the plurality of subtitle texts and the corresponding first timestamps and the at least one subtitle completion text and the corresponding second timestamps by using a large language model; and a commentary video generating unit configured to generate a commentary video based on the commentary words.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the above method.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the above method.
[0010] According to one or more embodiments of the present disclosure, by using a multi-modal large model to perform visual content understanding on a subtitle-free segment, the present disclosure can effectively supplement important information that is not explicitly expressed in a video, and overcome the problems of one-sided commentary video content and lack of coherence between segments. In addition, by inputting the subtitle texts, the subtitle completion texts and their respective timestamps into a large language model, the generated commentary words can better reflect the time evolution of the video content, improve the fit of the commentary words with the video content, and thus improve the overall quality of the commentary video and the user viewing experience.
[0011] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings illustrate exemplary embodiments and together with the description, function to explain the principles of the embodiments. The illustrated embodiments are intended to accommodate modifications by one of ordinary skill in the art. In all the drawings, like reference numerals refer to like parts throughout the various figures and embodiments.
[0013] Figure 1 A schematic diagram illustrating an exemplary system in which various methods described herein can be implemented according to embodiments of the present disclosure;
[0014] Figure 2 A flowchart illustrating a method of generating an explanation video according to embodiments of the present disclosure;
[0015] Figure 3 A flowchart illustrating a method of generating an explanation video according to embodiments of the present disclosure;
[0016] Figure 4 A flowchart illustrating a method of generating an explanation video based on an explanation word according to embodiments of the present disclosure;
[0017] Figure 5 A flowchart illustrating a method of determining a video segment matching one explanation word segment in a video to be processed according to embodiments of the present disclosure;
[0018] Figure 6 A flowchart illustrating a method of generating an explanation video based on an explanation word according to embodiments of the present disclosure;
[0019] Figure 7 A block diagram illustrating a structure of an explanation video generation apparatus according to embodiments of the present disclosure; and
[0020] Figure 8 A block diagram illustrating an exemplary electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0021] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, in which various details of the embodiments of the present disclosure are set forth to assist in the understanding of the present disclosure. It will be apparent to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Also, the description is made in the order of description of the accompanying drawings, and the description is made for the purpose of explanation and not for the purpose of limitation.
[0022] In the present disclosure, the terms "first", "second", etc. used in the description of various examples are not intended to limit the positional relationship, timing relationship or importance relationship of the elements, and such terms are only used to distinguish one element from another. In some examples, the first element and the second element can refer to the same instance of the element, and in some cases, based on the context of the description, they can also refer to different instances.
[0023] The terms used in the description of various examples in the present disclosure are only for the purpose of describing the specific examples, and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in the present disclosure encompasses any one of the listed items and all possible combinations.
[0024] In the related art, the automatic video commentary generation result content is one-sided and single, the cutting between different segments is strong, and the user viewing experience is poor.
[0025] To solve the above problems, the present disclosure can effectively supplement the important information not explicitly expressed in the video by using a multi-modal large model to understand the visual content of the non-subtitled segment, overcoming the one-sided commentary of the video content and the lack of connection between segments. In addition, by inputting the subtitle text, the subtitle completion text and their respective timestamps into the large language model, the generated commentary can better reflect the time evolution of the video content, improve the fit of the commentary and the video content, and thus improve the overall quality of the commentary video and the user viewing experience.
[0026] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 A schematic diagram of an example system 100 in which various methods and apparatus described herein can be implemented according to embodiments of the present disclosure is shown. Referring to Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more application programs.
[0028] In embodiments of the present disclosure, the server 120 can run one or more services or software applications that enable the methods of the present disclosure to be performed.
[0029] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0030] In Figure 1 In the illustrated configuration, the server 120 can include one or more components that implement the functionality performed by the server 120. These components can include software components that are executable by one or more processors, hardware components, or combinations thereof. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by the components. It should be understood that a wide variety of system configurations are possible, which can differ from system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0031] A user can use a client device 101, 102, 103, 104, 105, and / or 106 to engage in human-machine interactions. The client device can provide an interface that enables a user of the client device to interact with the client device. The client device can also output information to the user via the interface. Although Figure 1 Only six client devices are depicted, but one of skill in the art will understand that the present disclosure can support any number of client devices.
[0032] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computer devices, such as portable handheld devices, general purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service kiosk devices, service robots, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, and the like. These computer devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or including various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular telephones, smartphones, tablet computers, personal digital assistants (PDAs), and the like. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, and the like. Client devices are capable of executing a variety of different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0033] Network 110 can be any type of network familiar to those skilled in the art that can support data communications using any of a variety of available protocols, including without limitation TCP / IP, SNA, IPX, etc. As examples, one or more of networks 110 can be a LAN, an Ethernet network, a Token Ring network, a WAN, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a local area network (LAN), a wide area network (WAN), a wireless network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a Bluetooth network, a WIFI network), and / or any combination of these and / or other networks.
[0034] Server 120 can include one or more general purpose computers, special purpose server computers (e.g., PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other appropriate arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 can run one or more services or software applications that provide the functionality described below.
[0035] The computing units in the server 120 can run one or more operating systems including any of the operating systems described above, as well as any commercially available server operating systems. Server 120 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0036] In some embodiments, the server 120 can include one or more applications to analyze and consolidate data feeds and / or event updates from users of the client devices 101, 102, 103, 104, 105, and / or 106. The server 120 can also include one or more applications to display the data feeds and / or real-time events via one or more display devices of the client devices 101, 102, 103, 104, 105, and / or 106.
[0037] In some embodiments, the server 120 can be a server of a distributed system, or a server combined with a blockchain. The server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. The cloud server is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS, Virtual Private Server) services.
[0038] The system 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can reside in a variety of locations. For example, databases used by the server 120 can reside locally to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network- or application-specific connection. The databases 130 can be of different types. In certain embodiments, databases used by the server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0039] In certain embodiments, one or more of the databases 130 can also be used by applications to store application data. Databases used by applications can be different types of databases, such as key-value stores, object stores, or regular stores backed by file systems.
[0040] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and apparatuses described according to the present disclosure.
[0041] According to one aspect of the present disclosure, a large model-based explanatory video generation method is provided. As shown in Figure 2 The explanatory video generation method 200 includes: step S201, obtaining a plurality of subtitle texts and corresponding first timestamps in a to-be-processed video; step S202, determining at least one subtitle-free segment and corresponding second timestamps in the to-be-processed video based on the first timestamps of the plurality of subtitle texts; step S203, performing visual content understanding on the at least one subtitle-free segment using a first multi-modal large model to obtain at least one subtitle completion text corresponding to the at least one subtitle-free segment; step S204, using a large language model to generate an explanation word for the to-be-processed video based on the plurality of subtitle texts and corresponding first timestamps and the at least one subtitle completion text and corresponding second timestamps; and step S205, generating an explanatory video based on the explanation word.
[0042] Thus, by using a multi-modal large model to perform visual content understanding on the subtitle-free segment, important information not explicitly expressed in the video can be effectively supplemented, overcoming the problem of one-sidedness of the explanatory video content and lack of coherence between segments. In addition, by inputting the subtitle text, the subtitle completion text and their respective timestamps into the large language model, the generated explanation word can better reflect the temporal evolution of the video content, improving the fit of the explanation word with the video content, thereby improving the overall quality of the explanatory video and the user viewing experience.
[0043] In the present disclosure, the large model (i.e., deep learning large model) has an end-to-end characteristic and can directly generate reply data based on the input data of the user without the aid of functional components other than the deep learning large model or other inputs. In other words, the deep learning large model itself has a generation function. The deep learning large model can be a large language model. A large language model generally refers to a deep learning large model with tens of billions or even hundreds of billions of parameters, which is usually trained on large-scale text data or other modal data. A large language model can be used for various natural language processing tasks, such as text generation, language translation, and question-answering systems, etc. The deep learning large model can also be a multi-modal large model, i.e., a neural network model that simultaneously processes text, images, audio, and other types of input and output.
[0044] The deep learning large model may, for example, adopt an N-layer Transformer network structure with an encoder and a decoder, or a unified pre-trained language model (UniLM) network structure. It can be understood that the deep learning large model can also be other neural network models based on the Transformer network structure, which are not limited herein. The input and output of the deep learning large model are both composed of tokens. Each token can correspond to a single character, character, word, special symbol, or an image block, audio frame, or other modal information. The deep learning large model can be trained using pre-training tasks and generation tasks to have the above generation function.
[0045] Before step S201, a video to be processed can be obtained. The video to be processed can be a movie, a TV series, or any other form or content of video. For ease of understanding, the disclosure will mainly take a movie as an example to illustrate the technical solutions and technical details, but this does not intend to limit the protection scope of the disclosure.
[0046] In some embodiments, the download link of the movie can be obtained using automatic retrieval technology according to the movie name input by the user, and the movie video, i.e., the video to be processed, can be downloaded. In addition, the plot introduction of the movie can also be obtained by searching, using a large language model with network search function, or any other way.
[0047] In step S201, a plurality of subtitle texts and corresponding first time stamps of the video to be processed can be obtained by network search, OCR, or any other way. The first time stamp can include a start time stamp and an end time stamp of the subtitle text.
[0048] In step S202, based on the first time stamp of the plurality of subtitle texts, at least one subtitle-free segment in the video to be processed and the corresponding second time stamp can be obtained.
[0049] In some embodiments, the segment without subtitles between adjacent subtitle texts can be determined as a subtitle-free segment, and the second time stamp of the subtitle-free segment can be determined based on the end time stamp of the previous subtitle text and the start time stamp of the next subtitle text. If the visual content understanding is directly performed on the subtitle-free segment, it is often difficult to provide accurate video time stamp, and only a fuzzy description based on semantic understanding can be output. When these contents are mapped back to specific video segments, time misplacement is easy to occur, and the problem of inaccurate insertion time point of the completed content and disconnection with the picture may occur.
[0050] According to some embodiments, the first timestamp can include a start timestamp and an end timestamp of the corresponding subtitle text. Step S202, determining at least one subtitle-free segment in the to-be-processed video and the corresponding second timestamp based on the first timestamp of the plurality of subtitle texts can include: for any two subtitle texts with a time interval greater than a preset duration, cutting an original segment in the to-be-processed video that does not contain subtitle texts based on the end timestamp of the preceding subtitle text and the start timestamp of the subsequent subtitle text; and performing shot processing on at least one original segment in the to-be-processed video based on a shot detection algorithm to obtain at least one subtitle-free segment and determine the second timestamp of each of the at least one subtitle-free segment.
[0051] Thus, through the above shot processing manner, not only can the subtitle-free segments with too short duration be filtered out, but also each subtitle-free segment can fall within the same shot, so that the subtitle completion text obtained in the subsequent step is always inserted within the natural shot boundary, ensuring the coherence of the commentary video and improving the viewing experience.
[0052] In the present disclosure, "segment" is used to refer to a piece of video, and "segmentation" is used to refer to a piece of text.
[0053] In step S202, any two subtitle texts with a time interval greater than a preset duration can be cut to obtain at least one original segment in the to-be-processed video. Further, each original segment can be processed by shot processing to obtain one or more subtitle-free segments in the original segment and the second timestamp thereof. After all the subtitle-free segments in the original segments are summarized, at least one subtitle-free segment can be obtained.
[0054] In an exemplary embodiment, the preset duration can be 5 seconds.
[0055] In step S203, the first multi-modal large model can be any large model with visual content understanding capability, and the output subtitle completion text can be a description text of the visual content in the subtitle-free segment.
[0056] In some embodiments, the first multi-modal large model can be fine-tuned using fine-tuning technology to help the model master the ability to output more detailed picture descriptions.
[0057] In step S204, the plurality of subtitle texts and the first timestamp thereof, and the plurality of subtitle completion texts (i.e., the result obtained in step S203) and the second timestamp thereof are input into the large language model to obtain the video commentary generated by the large language model. In some embodiments, the plot introduction of the to-be-processed video obtained in advance can also be input into the large language model.
[0058] It can be understood that the second timestamp of the subtitle completion text can be the second timestamp of the subtitle-free segment corresponding to the subtitle completion text obtained in step S202.
[0059] According to some embodiments, as shown in Figure 3 The large model-based commentary video generation method can further include: in step S301, determining a video category of the to-be-processed video in a plurality of preset video categories; and in step S302, obtaining a commentary word generation template corresponding to the video category, wherein the commentary word of the to-be-processed video is generated by the large language model based on the commentary word generation template. It can be understood that Figure 3 The operations and effects of steps S304-S308 in Figure 2 The descriptions of steps S201-S205 in
[0060] Therefore, by selecting different commentary word generation templates based on the video category of the to-be-processed video, the commentary word generated by the large language model can be more consistent with the narrative style and audience preference of different types of videos, thereby improving the diversity and personalization of the commentary content, and further enhancing the user viewing experience.
[0061] Before implementing the commentary video generation method proposed in the present disclosure, a plurality of preset video categories can be pre-set, and one or more commentary word generation templates corresponding to each video category can be prepared.
[0062] In some embodiments, in step S307, the commentary word generation template obtained in step S302 can be input into the large language model together with other input content, so that the large language model generates the commentary word according to the commentary word generation template.
[0063] According to some embodiments, as shown in Figure 3 The large model-based commentary video generation method can further include: in step S303, determining background music for the commentary video based on the video category.
[0064] Therefore, by automatically determining the background music of the commentary video according to the video category, the atmosphere of the commentary video can be more coordinated with the content type, further enhancing the user's viewing experience.
[0065] In step S205, the commentary video can be generated based on the commentary word in various ways. In one exemplary embodiment, a plurality of video segments can be selected from the to-be-processed video, and these video segments can be combined with the commentary word to generate the commentary video.
[0066] Return to step S204. According to some embodiments, the generated commentary using the large language model can include multiple commentary segments. When generating the commentary video, a video segment that synchronously plays can be matched for each commentary segment. The commentary content is usually highly abstract and subjective, such as "she fell into a memory", "the war finally ended", and the like, which are difficult to directly correspond to specific shots. If the picture is not matched properly, for example, the mood, character or scene does not match, it will cause a strong sense of dissonance, seriously affecting the viewing experience of the audience. In addition, manually searching for matching shots sentence by sentence not only has a very high cost, but also has complex challenges such as semantic ambiguity, shot granularity mismatch, and emotional misjudgment in automatic implementation.
[0067] As shown in Figure 4 Step S205, generating a commentary video based on the commentary can include: step S401, performing semantic analysis on one of the multiple commentary segments to obtain a representative semantic tag of the commentary segment; step S402, determining a video segment matching the commentary segment in the to-be-processed video based on the representative semantic tag; and step S403, combining the video segments matched by the multiple commentary segments respectively to generate the commentary video.
[0068] Thus, by performing semantic analysis on each commentary segment and using the obtained representative semantic tag to perform picture matching, the workload of manual searching and matching is effectively reduced, the problem of emotional, character or scene mismatch causing a split viewing experience is solved, thereby improving the fit between the content of the commentary video and the picture, and enhancing the user viewing experience.
[0069] According to some embodiments, step S401, performing semantic analysis on one of the multiple commentary segments to obtain a representative semantic tag of the commentary segment includes: based on a plurality of preset semantic unit dimensions, the commentary segment is disassembled to obtain a plurality of semantic units; and using dependency syntax analysis and keyword extraction, the representative semantic tag is extracted from the plurality of semantic units.
[0070] Thus, by disassembling the commentary segment according to multiple semantic unit dimensions and combining dependency syntax analysis and keyword extraction methods, the representative semantic tag reflecting the essence of the commentary content can be obtained, which helps to realize more fine and contextually consistent picture matching in the subsequent process, further improving the relevance and consistency of the commentary segment and the matching picture.
[0071] According to some embodiments, the multiple semantic unit dimensions can include scene description, character behavior and emotional change.
[0072] In an example embodiment, the narration segment can be "nightfall, the main character stands alone on the bridge, tears slide down the cheeks". Based on the above three semantic unit dimensions, three semantic units can be obtained: "scene description: nightfall, on the bridge", "character behavior: the main character stands alone", and "emotional change: tears slide down the cheeks". Then, various natural language processing tools can be used for dependency syntax analysis to obtain the relationship between each sentence component in each semantic unit. For example, "main character (subject) - alone (adverb) - stand (predicate)". Further, keyword extraction can be performed based on the results of the dependency syntax analysis, thereby extracting representative semantic tags such as "night", "bridge", "main character", "alone", "stand", "tears", and "cheeks".
[0073] It can be understood that the plurality of semantic unit dimensions can also include other semantic unit dimensions. In addition, in addition to the above semantic analysis method combining semantic unit disintegration, dependency syntax analysis, and keyword extraction, semantic analysis can also be performed in other ways to obtain representative semantic tags.
[0074] According to some embodiments, as shown in Figure 5 Based on the representative semantic tags, determining a video segment matching a narration segment in the video to be processed can include: step S501, performing shot segmentation and picture semantic vectorization processing on the video to be processed using a second multi-modal large model to obtain a plurality of shot segments and corresponding picture semantic vectors; step S502, generating a text semantic vector of a narration segment based on the representative semantic tags; step S503, calculating the vector similarity between the text semantic vector of the matching prompt word and the picture semantic vectors of the plurality of shot segments; and step S504, determining a video segment matching a narration segment in the plurality of shot segments based on the vector similarity.
[0075] Thus, by combining shot segmentation and picture semantic vectorization processing, generating a matching prompt word based on representative semantic tags, and performing automatic retrieval based on vector similarity, efficient and accurate matching between narration segments and video pictures can be achieved. This approach not only improves the intelligence and automation of matching, but also effectively reduces the matching errors caused by semantic abstraction or subjective expression, thereby significantly improving the picture matching degree and overall viewing experience of the narration video.
[0076] In some embodiments, in step S501, a shot segmentation algorithm can be applied to the complete video to be processed to divide the video into several shot segments with clear start and end times. Then, for each shot segment, a key frame or representative frame can be selected, and its picture semantic feature in the multi-modal semantic space can be obtained through an image encoder (such as a visual-text model based on a Transformer structure), and the feature can be represented as a high-dimensional vector, i.e., a picture semantic vector of the shot segment. In some embodiments, a structured semantic vector database can be constructed to provide support for subsequent semantic retrieval and matching.
[0077] In some embodiments, in step S502, a large language model or a preset template can be called to convert the representative semantic label obtained in step S401 into a matching prompt word that can reflect the main theme of the commentary segment. The matching prompt word is usually a natural language text that briefly describes the scene, action or emotion. Subsequently, the matching prompt word can be input into a text encoder (such as a CLIP text encoder, BERT, etc.) to obtain its text semantic vector in the multi-modal semantic space (in the same multi-modal semantic space as the picture semantic feature in step S501) for vector similarity calculation.
[0078] In some embodiments, in step S503, the vector similarity between the text semantic vector of the commentary segment and the picture semantic vector of each shot segment can be calculated one by one. Cosine similarity or other measurement methods can be used to calculate the vector similarity. Through this similarity measurement, the closeness of the commentary segment and the pictures of each shot segment in the multi-modal semantic space can be quantified, thereby providing support for subsequent screening of matching video segments.
[0079] In some embodiments, in step S504, the shot segment with the highest similarity or greater than a preset threshold can be screened out to determine the video segment or candidate video segment that matches the commentary segment. In this process, the timestamp, subtitle keyword, audio emotion, etc. can be further combined to optimize the matching accuracy, as will be described below.
[0080] According to some embodiments, the large language model can generate a third timestamp corresponding to each commentary segment. Step S502, determining a video segment that matches a commentary segment in the video to be processed based on the representative semantic label can include determining a time interval between the third timestamp of the commentary segment and the fourth timestamp of the plurality of shot segments. The video segment that matches the commentary segment can be determined from the plurality of shot segments based on the vector similarity and the time interval.
[0081] Thus, in this way, time information can be further introduced on the basis of semantic similarity matching, so that the matched video segment is not only highly relevant to the commentary segment in semantic content, but also closer to the actual plot progress corresponding to the commentary in the time axis, thereby improving the matching degree between the commentary and the video picture and enhancing the user viewing experience.
[0082] In some embodiments, the time interval between the third timestamp and the fourth timestamp can be quantified, the similarity and the time interval can be combined in a weighted combination manner, and the video segment with the optimal sorting result can be selected as the video segment matched with the commentary segment. In other embodiments, a plurality of candidate video segments can be first screened based on the similarity, and the video segment with the smallest time interval can be selected as the matched video segment; or a plurality of candidate video segments can be first screened based on the time interval, and the video segment with the highest similarity can be selected as the matched video segment. It can be understood that the vector similarity and the time interval can also be combined in other ways to determine the video segment matched with the commentary segment from the plurality of shot segments.
[0083] According to some embodiments, the step S502 of determining the video segment matched with the commentary segment in the video to be processed based on the representative semantic label can further include: extracting auxiliary information of the plurality of shot segments, the auxiliary information including subtitle keywords and audio emotions; and performing auxiliary information matching between the representative semantic label and the auxiliary information of the plurality of shot segments to obtain an auxiliary information matching result. The video segment matched with the commentary segment can be determined from the plurality of shot segments based on the vector similarity and the auxiliary information matching result.
[0084] Thus, by comprehensively introducing auxiliary information such as subtitle keywords and audio emotions in the matching process of the commentary segment and the video segment, the association between the video picture and the commentary in multiple dimensions such as semantics, text, and emotion can be more comprehensively reflected, the matching accuracy between the commentary and the video picture can be effectively improved, and the user viewing experience can be further enhanced.
[0085] In some embodiments, the video segment matching the commentary segment can be determined by first calculating the similarity scores between the matching prompt text semantic vector of the commentary segment and the picture semantic vector of each shot segment, respectively; then extracting the subtitle keywords and audio emotion labels of each shot segment, and comparing these auxiliary information with the representative semantic labels of the commentary segment to calculate the auxiliary information matching scores. Finally, the similarity scores and the auxiliary information matching scores can be weighted and combined into a comprehensive score according to the preset weights, and the video segment with the highest comprehensive score is determined as the best matching video segment of the commentary segment. In addition, the similarity scores, the auxiliary information matching scores and the time interval quantization scores can also be weighted and combined into a comprehensive score according to the preset weights, so as to realize multi-dimensional matching.
[0086] According to some embodiments, step S205, generating commentary video based on commentary, can further include: in response to not finding a video segment matching a commentary segment in the video to be processed, generating a video segment matching a commentary segment based on at least one of the following: image illustration generation, slow motion interpolation, and emotion special effect mask.
[0087] Thus, by supplementing with image illustration generation, slow motion interpolation or emotion special effect mask when a video segment matching the commentary segment cannot be found, it can ensure that each commentary segment has corresponding visual content, and improve the coherence and integrity of the commentary video.
[0088] According to some embodiments, step S403, combining the video segments matching the plurality of commentary segments respectively to generate the commentary video can include: converting the plurality of commentary segments into commentary voice; and adaptively aligning the duration of the commentary voice with the duration of the corresponding video segment.
[0089] Thus, by converting the commentary segment into commentary voice and realizing adaptive alignment of the duration of the commentary voice with the duration of the corresponding video segment, it can ensure that the commentary content and the picture are played synchronously, making the commentary video more natural and smooth.
[0090] In the commentary video, in addition to playing the video segment synchronously with the commentary, the highlight segments in the video to be processed can also be played with the original sound, so as to improve the emotional resonance of the viewer and better complete the plot summary. Therefore, it is necessary to determine which segments in the video to be processed are played and where in the commentary video they are played.
[0091] According to some embodiments, the commentary can include a plurality of commentary segments. For example, the commentary can include a plurality of commentary segments, and each commentary segment can be matched with a video segment. Figure 6As shown, step S205, generating the explanation video based on the commentary can include: step S601, performing importance scoring on the plurality of subtitle texts, and determining a plurality of candidate clips in the to-be-processed video based on the scoring results; step S602, matching the plurality of candidate clips with the plurality of commentary segments; and step S603, in response to determining that a target commentary segment in the plurality of commentary segments matches a target clip in the plurality of candidate clips, adding the target clip to the target commentary in the original sound manner.
[0092] In this way, by verifying in combination with the multimodal information of the text and the video frame, the highly consistent of the selected original sound clip in the emotional expression and the plot context is ensured, thereby effectively improving the natural integration of the original sound clip and the overall expressiveness and audience empathy of the explanation video.
[0093] According to some embodiments, step S601, performing importance scoring on the plurality of subtitle texts, and determining a plurality of candidate clips in the to-be-processed video based on the scoring results can include: determining emotional intensity scores of the plurality of subtitle texts by using an emotion analysis model; performing keyword extraction on the plurality of subtitle texts, and determining plot importance scores based on the extracted keywords; and obtaining the scoring results based on the emotional intensity scores and the plot importance scores of the plurality of subtitle texts.
[0094] In this way, by using the emotion analysis model and the keyword extraction technology to respectively perform emotional intensity scoring and plot importance scoring on the plurality of subtitle texts, the emotional fluctuations and the plot highlights in the to-be-processed video can be effectively identified, and the more high-quality and more infectious content in the original video can be used as the original sound material for subsequent generation.
[0095] In some embodiments, the operation of step S602 can refer to the description of steps S501-S504 above. That is, the picture semantic vectors of the plurality of candidate clips and the text semantic vectors of the plurality of commentary segments can be obtained, and matching can be performed based on the vector similarity between the two. In addition, matching can also be performed in combination with the time interval between the candidate clip and the commentary segment, the auxiliary information matching result, and the like.
[0096] In some embodiments, the meaning of adding the target clip to the target commentary in the original sound manner is that, in the finally generated explanation video, the target clip is immediately after the target commentary, and no other audio (such as the explanation voice converted from the target commentary or background music) is played when the target clip is played.
[0097] According to another aspect of the present disclosure, a large model-based explanation video generation apparatus is provided. As shown in FIG. 8, the apparatus can include a processor 801, a memory 802, and a communication interface 803. Figure 7As shown, the commentary video generation apparatus 700 includes a subtitle obtaining unit 710 configured to obtain a plurality of subtitle texts and corresponding first timestamps in a to-be-processed video; a subtitle-free segment determining unit 720 configured to determine at least one subtitle-free segment and corresponding second timestamps in the to-be-processed video based on the first timestamps of the plurality of subtitle texts; a content understanding unit 730 configured to perform visual content understanding on the at least one subtitle-free segment by using a first multi-modal large model to obtain at least one subtitle completion text corresponding to the at least one subtitle-free segment; a commentary generation unit 740 configured to generate a commentary for the to-be-processed video based on the plurality of subtitle texts and the corresponding first timestamps and the at least one subtitle completion text and the corresponding second timestamps by using a large language model; and a commentary video generation unit 750 configured to generate a commentary video based on the commentary.
[0098] It can be understood that the operations of the units 710-750 can be respectively referred to the descriptions of steps S201-S205 above, which will not be repeated here.
[0099] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.
[0100] According to embodiments of the present disclosure, an electronic device, a readable storage medium and a computer program product are also provided.
[0101] Reference Figure 8 A block diagram of an electronic device 800, which can be used for the server or client of the present disclosure, will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computing devices, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computing devices. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections, and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present disclosure described and / or claimed in this document.
[0102] As Figure 8As shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 802 or a computer program loaded into a random access memory (RAM) 803 from a storage unit 808. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0103] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information to the electronic device 800, can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0104] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The computing unit 801 performs various methods, procedures, and / or processes described above. For example, in some embodiments, these methods, procedures, and / or processes can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the methods, procedures, and / or processes described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform these methods, procedures, and / or processes by any other suitable means, such as by means of firmware.
[0105] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0106] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0107] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0108] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0109] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0110] The computer system can include clients and servers. This relationship can be remote, such that the servers are distributed across many clients. The relationship can also be such that the server is remote from the client. The servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The servers can also be servers of a distributed system, or servers combined with a blockchain.
[0111] It should be understood that the various forms of flow shown above can be reordered, steps added or removed. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.
[0112] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-described methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but only by the granted claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, various elements in the embodiments or examples can be combined in various ways. It is important that many of the elements described herein can be replaced by equivalent elements that appear after the present disclosure as technology evolves.
Claims
1. A method for generating an explanation video based on a large model, comprising: Obtain multiple subtitle texts and corresponding first timestamps in the video to be processed; Determining at least one non-caption segment in the to-be-processed video and a corresponding second timestamp based on the first timestamps of the multiple subtitle texts; Performing visual content understanding on the at least one non-captioned segment using the first multimodal large model to obtain at least one captioned completion text corresponding to the at least one non-captioned segment; Generate a commentary for the video to be processed using a large language model based on the multiple subtitle texts and the corresponding first timestamps and the at least one subtitle completion text and the corresponding second timestamp; as well as Based on the commentary, a commentary video is generated.
2. The method according to claim 1, wherein The first timestamp includes a start timestamp and an end timestamp of a corresponding subtitle text, wherein, based on the first timestamps of the multiple subtitle texts, determining at least one subtitle-free segment in the to-be-processed video and the corresponding second timestamp includes: For two subtitle texts whose time interval is greater than a preset time length, based on the end timestamp of the previous subtitle text and the start timestamp of the subsequent subtitle text, intercepting an original segment that does not contain the subtitle text in the to-be-processed video; and Based on a shot detection algorithm, the original segment is subjected to shot-by-shot processing to obtain the at least one non-caption segment, and a second timestamp of each of the at least one non-caption segment is determined.
3. The method according to claim 1 or 2, wherein: The commentary includes a plurality of commentary segments, wherein generating a commentary video based on the commentary includes: Performing semantic analysis on one of the multiple commentary segments to obtain a representative semantic label of the one commentary segment; Determining, based on the representative semantic tag, a video segment in the video to be processed that matches the one commentary segment; and The video clips that match the multiple commentary segments are combined to generate the commentary video.
4. The method according to claim 3, wherein: Performing semantic analysis on one of the multiple commentary segments to obtain a representative semantic label of the commentary segment includes: Decomposing the one commentary segment based on a plurality of preset semantic unit dimensions to obtain a plurality of semantic units; and Dependency parsing and keyword extraction are used to extract the representative semantic tags from the multiple semantic units.
5. The method according to claim 4, wherein The multiple semantic unit dimensions include scene description, character behavior and emotional changes.
6. The method according to claim 3, wherein: Determining, based on the representative semantic tag, a video segment in the video to be processed that matches the one commentary segment includes: Using the second multimodal large model, the video to be processed is subjected to shot segmentation and picture semantic vectorization processing to obtain a plurality of shot segments and corresponding picture semantic vectors; Generating a text semantic vector of the commentary segment based on the representative semantic tag; Calculating vector similarities between the text semantic vector of the one commentary segment and the picture semantic vectors of the plurality of shot segments; and Based on the vector similarity, a video segment matching the one commentary segment is determined from the multiple shot segments.
7. The method according to claim 6, wherein: The commentary includes third timestamps of the multiple commentary segments, wherein, based on the representative semantic tag, determining a video segment in the to-be-processed video that matches the one commentary segment includes: determining a time interval between a third timestamp of the one commentary segment and fourth timestamps of the plurality of shot segments, Wherein, based on the vector similarity and the time interval, a video segment matching the one commentary segment is determined from the multiple shot segments.
8. The method according to claim 6, wherein: Determining, based on the representative semantic tag, a video segment in the video to be processed that matches the one commentary segment further includes: extracting auxiliary information of the plurality of shot segments, the auxiliary information comprising subtitle keywords and audio emotions; and Perform auxiliary information matching on the representative semantic tag and the auxiliary information of the plurality of shot segments to obtain an auxiliary information matching result, Wherein, based on the vector similarity and the auxiliary information matching result, a video segment matching the one commentary segment is determined from the multiple shot segments.
9. The method according to claim 3, wherein: Generating the commentary video based on the commentary further includes: In response to not finding a video segment matching the one commentary segment in the video to be processed, generating a video segment matching the one commentary segment based on at least one of the following: image illustration generation, slow motion interpolation, and emotional effect masking.
10. The method according to claim 3, wherein: Combining the video clips that match the multiple commentary segments to generate the commentary video includes: Converting the plurality of commentary segments into commentary speech; and The duration of the commentary voice is adaptively aligned with the duration of the corresponding video clip and then combined to obtain the commentary video.
11. The method according to claim 1 or 2, wherein: The commentary includes a plurality of commentary segments, wherein generating a commentary video based on the commentary includes: Performing importance scoring on the multiple subtitle texts, and determining multiple candidate segments in the video to be processed based on the scoring results; Matching the plurality of candidate segments with the plurality of commentary segments; and In response to determining that a target commentary segment among the plurality of commentary segments matches a target segment among the plurality of candidate segments, the target segment is added after the target commentary in an original manner.
12. The method according to claim 11, wherein Scoring the importance of the multiple subtitle texts and determining multiple candidate segments in the video to be processed based on the scoring results includes: Determining sentiment intensity scores of the plurality of subtitle texts using a sentiment analysis model; Extracting keywords from the multiple subtitle texts, and determining plot importance scores based on the extracted keywords; and The scoring result is obtained based on the emotion intensity scores and plot importance scores of the multiple subtitle texts.
13. The method according to claim 1 or 2, further comprising: Determining the video category of the video to be processed among multiple preset video categories; as well as A commentary generation template corresponding to the video category is obtained, wherein the commentary of the video to be processed is generated based on the commentary generation template using the large language model.
14. The method according to claim 13, further comprising: Based on the video category, background music for the commentary video is determined.
15. A large-scale model-based explanatory video generation device, comprising: A subtitle acquisition unit, configured to acquire a plurality of subtitle texts and corresponding first timestamps in the video to be processed; a non-caption segment determining unit configured to determine at least one non-caption segment in a to-be-processed video and a corresponding second timestamp based on the first timestamps of the plurality of caption texts; a content understanding unit configured to perform visual content understanding on the at least one non-captioned segment using the first multimodal large model to obtain at least one caption completion text corresponding to the at least one non-captioned segment; a commentary generating unit configured to generate a commentary for the video to be processed based on the multiple subtitle texts and the corresponding first timestamps and the at least one subtitle completion text and the corresponding second timestamp using a large language model; as well as The commentary video generating unit is configured to generate a commentary video based on the commentary words.
16. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; wherein The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.
17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.
18. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.
Citation Information
Cited By
Video generation method and device, storage medium and computer program product
CN121531189A
Method and device for training expression recognition model and electronic equipment
CN121686543A