Method and apparatus for processing videos, and device, medium and program product
By analyzing video attribute data through machine learning models, videos including narration are automatically generated, solving the problems of low efficiency and unstable quality in traditional audio production and achieving efficient and stable video processing.
Patent Information
- Application Number
- PCT/CN2025/111344
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-07-29
- Publication Date
- 2026-02-05
AI Technical Summary
Traditional audio recording production processes are inefficient and of inconsistent quality, rely on manual operation, and are time-consuming.
By analyzing video attribute data using machine learning models, generating text data, and adding audio data, videos including narration are automatically generated.
It improves the efficiency and quality of oral video production, reduces human intervention, and ensures the accuracy and coherence of video content.
Smart Images

Figure CN2025111344_05022026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, media, and programs for processing video.
[0001] This application claims priority to Chinese Patent Application No. 202411038008.7, filed on July 30, 2024, entitled "Method, Apparatus, Device, Medium and Program Product for Processing Video", the entire contents of which are incorporated herein by reference. Technical Field
[0002] Exemplary implementations of this disclosure generally relate to video management, and more particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for adding audio data (e.g., narration) to a video. Background Technology
[0003] Audiovisual presentations are a media form that uses language to describe and convert some video information into audio information for the user. Therefore, audiovisual presentations can help people understand the detailed content of the video through language, such as describing scenes and facial expressions. In traditional methods, the video processing in the production of audiovisual presentations is time-consuming and labor-intensive, and the quality of the content is heavily dependent on the skill level of the production staff, making it difficult to guarantee consistent quality. Therefore, there is a desire to improve the efficiency and quality of the video processing in the production of audiovisual presentations. Summary of the Invention
[0004] In a first aspect of this disclosure, a method for processing video is provided. In this method, in response to receiving a generation request to generate a second video based on a first video, attribute data of the first video is obtained. Based on the attribute data, text data describing the first video is determined. Based on the first video and audio data corresponding to the text data, the second video is generated.
[0005] In a second aspect of this disclosure, an apparatus for processing video is provided. The apparatus includes: an acquisition module configured to acquire attribute data of the first video in response to receiving a generation request to generate a second video based on a first video; a determination module configured to determine text data describing the first video based on the attribute data; and a generation module configured to generate the second video based on the first video and audio data corresponding to the text data.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processor.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, which, when executed by a processor, cause the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] Figures 2A-2C show block diagrams for processing video according to some implementations of this disclosure;
[0013] Figure 3 shows a flowchart of a process for processing video according to some implementations of this disclosure;
[0014] Figure 4 shows a flowchart of a method for processing video according to some implementations of this disclosure;
[0015] Figure 5 shows a block diagram of an apparatus for processing video according to some implementations of the present disclosure; and
[0016] Figure 6 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0017] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0019] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0021] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0022] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0023] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0024] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such an event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0025] Example Environment
[0026] Prolonged video viewing can cause visual fatigue, and for some users with poor eyesight, it may be difficult to clearly see the specific content on the video screen. These factors create a need for verbal narration. When a user wants to understand the content of a video, they can describe the content on the screen verbally, allowing them to grasp the content without watching the video.
[0027] Referring to Figure 1, which describes an application environment according to an example implementation of the present disclosure, Figure 1 shows a block diagram 100 of an application environment according to an exemplary implementation of the present disclosure. As shown in Figure 1, application 110 may be, for example, an application capable of providing audio and video. It should be understood that application 110 can be used to play videos, short video streams, image galleries, and other images. It can also process ordinary videos and add audio data to the video to narrate video scenes, thereby generating narrated images.
[0028] In the context of this disclosure, a specific application environment for a video processing scheme can be described using a video related to a particular object as an example. For instance, a video of a cat can be used as the video to be processed 115, which may include dynamic footage of the cat. When the application plays the video to be processed 115, the user can only hear the original audio content of the video. The video can be processed, and narration can be added to explain the video content, thereby generating spoken video. For example, spoken video can be generated manually. Then, application 110 can play the generated spoken video, thereby outputting the corresponding spoken content 120.
[0029] However, as briefly mentioned earlier, traditional methods for producing narrated videos suffer from low efficiency and poor quality of output content. Taking film as an example, the production of narrated videos in traditional methods may include the following steps: 1) Film viewing: The narrator needs to watch the film several times to understand its core content. 2) Script writing: The narrator needs to write the narration based on the visuals, interval lengths, and the core content of the film. 3) Narration recording: The narrator needs to convert the written narration into an audio version based on the script. 4) Audio track synthesis: The narrator needs to use video editing tools to synthesize the audio track with the video and output the final narrated content.
[0030] Therefore, this method of producing spoken-word videos requires a significant amount of time and manpower, and the quality of the resulting spoken-word content varies considerably. Thus, there is a desire to improve the efficiency and quality of the video processing steps in the spoken-word video production process.
[0031] Video processing overview
[0032] To at least partially address the shortcomings of the prior art, a method for processing video is proposed according to an exemplary implementation of this disclosure. According to one example implementation of this disclosure, the method can be implemented in a processing application. According to another example implementation of this disclosure, the method for processing video can be implemented on any suitable computing device, for example, on a conventional computing device, on a server device, or in combination of both.
[0033] In the following description, an exemplary implementation of this disclosure may be illustrated using a video of a specific object (such as a cat) (e.g., video 115 to be processed in Figure 1) as an example. Referring to Figures 2A-2C, which illustrate block diagrams 200A, 200B, and 200C for processing video according to some implementations of this disclosure, please refer to the summary of an exemplary implementation of this disclosure.
[0034] As shown in Figures 2A-2C, a second video including narration can be generated from a first video (excluding narration) based on a request. Specifically, a second video (e.g., video 115 to be processed) can be generated from a first video (e.g., video 115 to be processed). The second video (e.g., including at least images 210, 220, and 230, and their corresponding narrations, i.e., audio 215, audio 225, and audio 235) can be generated. Specifically, attribute data of the first video can be obtained. In one implementation, the narration can be pre-generated by the processing application. Alternatively and / or additionally, with increased processing speed, the narrated video can be generated and played in real-time while the application 110 is playing the video 115 to be processed. The attribute data of the first video may include, but is not limited to, data related to the visuals and audio of the first video. The data related to the visuals and audio of the first video will be described in detail below.
[0035] Furthermore, based on the attribute data, text data used to describe the first video can be determined. Then, based on the first video and the audio data corresponding to the text data, a second video can be generated.
[0036] Here, the first video may include an audio portion and an image portion, for example, referred to as first audio data and first image data, respectively. The text data can be used to update the first audio data. The text data may, for example, include the text "A cat is watching a person fish" corresponding to audio 215, the text "The person is going to give the fish they caught to the cat" corresponding to audio 225, the text "The cat is eating the fish it just caught" corresponding to audio 235, and so on. It should be understood that the text data can be data used to describe any scene (including images and / or text) of the first video. Subsequently, the audio data corresponding to the text data can be determined, i.e., audio 215, 225, and 235 can be determined respectively.
[0037] In the process of generating the second video, for example, the first image data in the first video can be kept unchanged, and narration can be added to the first audio data in the first video. Specifically, audio data corresponding to the text data is added to the first audio data of the first video to form first updated audio data. The second video is generated based on the first updated audio data and the first video (i.e., the first image data in the first video).
[0038] For example, the second video may include image 210 in block diagram 200A (which may be a frame from video 115 to be processed) and accompanying audio 215. During playback of the second video, the user can hear the voice message "A cat is watching a person fish" from application 110. The second video may also include image 220 in block diagram 200B (which may be another frame from video 115 to be processed) and accompanying audio 225. The user can then hear the voice message "This person is going to give the fish they caught to this cat" from application 110. The second video may also include image 230 in block diagram 200C (which may be yet another frame from video 115 to be processed) and accompanying audio 235. The user can then hear the voice message "This cat is eating the fish it just caught" from application 110.
[0039] Using the exemplary implementation of this disclosure, a second video can be automatically generated by analyzing a first video without human intervention. This effectively improves the efficiency of narration production, ensures the quality of the narrated content output with the video, and makes it easier for users to understand the video content.
[0040] Detailed process of video processing
[0041] The foregoing has described an example implementation of this disclosure; further details regarding video processing will be described below. Figure 3 shows a flowchart of a process 300 for video processing according to some implementations of this disclosure. An example implementation of user-processed video according to this disclosure will be described below with reference to Figure 3, in conjunction with Figures 2A-2C and Figure 1.
[0042] According to an example implementation of this disclosure, when determining text data, the first video can be divided into a group of video segments based on attribute data. Taking Figure 1 as an example, the video 115 to be processed in Figure 1 can be divided into a group of video segments, which may include one or more video segments. Referring to Figures 2A-2C, a group of video segments may include a single video segment having images 210, 220, and 230, or it may include multiple video segments having images 210, 220, and 230 respectively, or it may include a single video segment having image 210 and a single video segment having images 220 and 230. Here, each video segment may include multiple images.
[0043] The method for dividing the first video into a group of video segments can be determined in several ways. For example, a pre-set model (e.g., a trained and fine-tuned machine learning model) can be used. Cue words regarding the division criteria can be input into the model, allowing it to divide the first video according to the content of the cue words. Cue words for division criteria may include, for example, division labels, such as dividing when a scene changes; segment length thresholds, such as the duration of each video segment in a group of video segments. In this case, for example, the first video can be divided into multiple video segments based on the cue words, with each segment no longer than 2 minutes (or other lengths), and so on. This method fully utilizes the powerful processing capabilities of machine learning models to divide the first video into multiple segments that are easy to describe using language.
[0044] Alternatively, image recognition techniques can be used to divide the first video into a group of video segments. For example, segmentation can be based on whether the target object appears in the current frame, or segmentation can be based on scene changes in the first video. It should be understood that the method of dividing the first video into a group of video segments is not limited here and can be set according to different needs.
[0045] The attribute data may include at least one of the following: speech data extracted from the first video, context data of the first video, and keyframe data of the first video, etc. As shown in Figure 3, in process 300, when processing the source video 310 (i.e., the first video), video preprocessing (310) can be performed first. The preprocessing process mainly includes the extraction of attribute data. Then, the next step of video content understanding (330) can be performed based on the attribute data. Specifically, the preprocessing process may include at least the extraction of source audio (i.e., speech data, such as dialogue, monologue, original narration, etc.) (321), context data extraction (322), and keyframe extraction (323).
[0046] For audio data extraction, speech recognition technology can be used to extract the audio portion from the first video. For context data extraction, internet searches can be used to extract the context from the first video. For example, the video title (e.g., movie title), the names of people and locations in the video, etc., can be used as keywords to perform a search. For keyframe data extraction, image recognition technology or computer vision algorithms can be used to determine new keyframes from scene changes, or by identifying the appearance of target objects. Additionally, the first frame of a scene depicting important activities in the first video can be identified as a new keyframe, and so on.
[0047] Furthermore, for a set of video clips, a set of text clips can be generated from the text data. Continuing with Figures 2A-2C as an example, for a set of video clips containing images 210, 220, and 230, a set of text clips can be generated, for example: "A cat is watching a person fish," "The person is going to give the fish they caught to the cat," and "The cat is eating the fish it just caught." Using the exemplary implementation of this disclosure, machine learning models or various image and audio processing algorithms can be used to create spoken audio in a standardized manner, effectively improving the efficiency of video processing during spoken audio production and reducing the instability caused by manual production.
[0048] Referring again to Figure 3, for the video content understanding (330) part, after the video segmentation (331) discussed above, the machine learning model can sequentially implement the steps of understanding keyframes of the video frame (332) and generating text segments (333) using the implementation method of this example. For the cue words instructing the machine learning model to process attribute data, it can at least include understanding the keyframe data in the attribute data.
[0049] According to an example implementation of this disclosure, when generating a set of text snippets, prompt words can be determined for video snippets within a set of video snippets. The prompt words can instruct a machine learning model to process attribute data to generate text snippets describing the video snippets. Then, based on the machine learning model's response to the prompt words, the text snippets describing the video snippets can be determined. For example, the prompt words could specify: "Give a summary description based on audio data, context data, and the content of keyframes," etc. Thus, a machine learning model can be used to summarize the content of the video snippets, and each video snippet can have a corresponding descriptive text.
[0050] Alternatively and / or additionally, the prompts may include text style, such as artistic or documentary style, and may also include word count thresholds, keyword requirements, etc. For example, the prompt could be: "Please generate a summary description of the input video clip. The summary description should be in an artistic style and no more than 200 characters." Here, the word count threshold can be determined based on the amount of input data. In another example implementation, if problems arise such as too much or too little content in the generated set of text clips, the prompts can be optimized and fine-tuned to overcome these issues. Utilizing the exemplary implementation of this disclosure, and employing machine learning models to assist in generating text clips, can improve the efficiency and quality of generating a second video.
[0051] According to an example implementation of this disclosure, contextual transitions can also be considered when determining text segments. For example, the beginning of a text segment can be generated based on a previous video segment preceding the current video segment. In other words, when generating the text segment for the current video segment, a brief recap of the previous video segment can be generated at the beginning. Taking a scenario where multiple video segments are divided based on scene changes, for example, the previous scene can be reviewed first to generate an overview of that scene, and then this overview can be used as the beginning of the text segment for the current video segment.
[0052] By utilizing the exemplary implementation of this disclosure, content connection and transition can be achieved when describing adjacent video segments, resulting in a complete and in-depth content description. This makes it easier for users to distinguish changes in video footage when listening to spoken content, and makes it easier to understand the video content.
[0053] According to an example implementation of this disclosure, each text fragment can be converted into corresponding audio data, which can then be inserted into an appropriate location in the first video. To ensure that the inserted audio data does not affect the viewing experience, when generating the second video, for each video fragment in a set of video fragments, a set of sub-fragments that do not contain audio data can be identified. Then, an audio fragment corresponding to a text fragment can be added to at least one sub-fragment in the set of sub-fragments to serve as narration in the second video.
[0054] Therefore, after understanding the video content (330) in Figure 3, the video content synthesis (340) part is performed, and the text-to-speech (341) step is completed. In this example implementation, a movie can be used as the first video. The first video may include various sounds such as dialogue, monologue, and background music. In this case, a set of sub-segments in the video segment that do not contain audio data can be identified first, and then the audio segment corresponding to the text segment describing the video segment can be inserted into the sub-segments to form the narration in the second video.
[0055] In another example implementation, a set of sub-segments that do not include human voices (e.g., it might include silent sections or background music) can be identified, and the volume of the audio portion (e.g., background music) in this set of sub-segments can be lowered. Then, the audio segment corresponding to the text segment describing the video segment can be inserted into this set of sub-segments, thus creating narration with background music. This increases the likelihood of narration being inserted into the video segment. It should be understood that the timing of the narration insertion can be set according to actual needs.
[0056] Using the exemplary implementation of this disclosure, the narration voice does not overlap with any sound in the video clip, or the narration voice does not overlap with the voice of a character in the video clip, thereby improving the quality of the generated second video and ensuring that the inserted narration voice interferes less with the sound effects of the original video.
[0057] According to an example implementation of this disclosure, when adding an audio segment to at least one sub-segment, a sub-segment matching the audio segment can be selected from a set of sub-segments, and the audio segment can be added to the sub-segment. Alternatively and / or additionally, a sub-segment matching the duration of the audio segment can be selected from a set of sub-segments (e.g., sub-segments excluding sound data, or sub-segments excluding character voice data) based on duration matching. In this case, the duration of the set of sub-segments can be greater than the duration of the audio segment. This ensures that the main plot segments of the film, excluding these sub-segments, are not interfered with by the narration.
[0058] Alternatively and / or additionally, based on content matching, text snippets matching the audio snippets can be selected from a set of sub-snippets. This ensures that the visual content of sub-snippets excluding audio data or excluding voice data matches the narration as closely as possible.
[0059] By utilizing the exemplary implementation of this disclosure, the narration in the generated second video can be made to not affect the original sound effects of the first video, while improving the matching degree between the narration content and the video segments, thereby further improving the quality of the generated second video.
[0060] According to an example implementation of this disclosure, an excessively long audio segment can be divided into multiple parts. Specifically, when adding an audio segment to at least one sub-segment, in response to determining that the duration of the audio segment is greater than the duration of each sub-segment in a set of sub-segments, the audio segment can be divided into a first audio segment and a second audio segment, and a first sub-segment matching a first duration and a second sub-segment matching a second duration are selected from the plurality of sub-segments, respectively. The first audio segment is then added to the first sub-segment, and the second audio segment is added to the second sub-segment.
[0061] In this example implementation, because the audio segment is too long, exceeding the duration of each sub-segment in a set, some content in the generated narration overlaps with parts of the video segment, such as dialogue, thus interfering with the video segment. Therefore, the audio segment can be divided. The division method can be based on the duration of the sub-segments, or on a more granular recognition of the visual content of the sub-segments, etc., and there are no restrictions here.
[0062] By utilizing the exemplary implementation of this disclosure, the narration in the generated second video can be further made to not affect the original sound effects of the first video, thus ensuring the quality of the generated second video.
[0063] According to an example implementation of this disclosure, when adding an audio segment to a sub-segment, an audio track can be created within the sub-segment, and the audio segment can be added to the audio track. Thus, referring to Figure 3, process 300 proceeds to the audio track alignment (342) step. In this example implementation, the first video may include multiple audio tracks (which can be simply referred to as audio tracks). These multiple audio tracks may include, for example, a voice track, a background music track, an ambient sound track, etc. For audio segments of narration, they can be added to a new audio track different from the original audio track of the source video (i.e., the first video).
[0064] By utilizing the exemplary implementation of this disclosure, the narration portion can be decoupled from all the original audio of the first video, thus avoiding the influence of the narration portion on all the original audio of the first video.
[0065] According to an example implementation of this disclosure, the timbre of the audio data can be different from the timbre of the character in the first video. Still taking the first video as an example, for instance, the narration of the present-day scenes in the video can be in a male voice, while the narration of the flashback scenes can be in a female voice. Alternatively, the timbre of a pre-recorded audio clip can be extracted and used as the narration timbre. Alternatively and / or additionally, with the user's authorization, the timbre of a video published by the user can be used as the narration timbre. It should be understood that the method of setting the narration timbre is not limited here. Referring again to Figure 3, after setting the narration timbre, the narration audio can be synchronously merged with the video footage to achieve image synthesis (343), and then the narrated video can be output (350).
[0066] By utilizing the exemplary implementation of this disclosure, and by making the narration's tone different from the tone of the character in the first video, it is possible for users to quickly understand that the currently playing content is narration content without affecting their understanding of the first video content, thereby further improving the quality of the spoken video.
[0067] According to some implementations of this disclosure, a first video is played in response to receiving a first request to disable the narration function; or a second video is played in response to receiving a second request to enable the narration function. Specifically, controls for enabling / disabling the narration function can be provided at application 110. When a user request to disable the narration function is received, the first video, i.e., the original video without narration, can be played directly. When a user request to enable the narration function is received, the second video, i.e., the audio recording including narration, can be played. Using some implementations of this disclosure, different users can be supported in selecting and watching appropriate videos according to their own needs.
[0068] According to an example implementation of this disclosure, when generating a second video, the second video can be generated based on at least a portion of the first video. Specifically, at least a portion of key video segments from the first video can be extracted from a set of video clips. For example, the position, duration, etc., of the key video segments can be pre-specified to facilitate extraction. Alternatively and / or additionally, a pre-trained machine learning model can be used to extract the key video segments. Furthermore, the second video can be generated based on at least a portion of the key video segments from the first video.
[0069] In this example implementation, generating the second video based on at least a portion of the first video can depend on different needs. For example, visually impaired users, users with eye strain, or users who cannot easily watch the entire video might want to listen to the entire first video. In such cases, the second video can be generated based on the entire content of the first video. For example, the need could also be that the user wants an explanatory video of the first video (e.g., a movie synopsis video). In such cases, the second video can be generated based on a portion of the first video. For example, the need could also be that the user wants a simple overview of the first video. In such cases, the second video can be generated based on a small portion of key content from the first video. In another example implementation, the narration rate can also be adjusted according to different needs.
[0070] By utilizing the exemplary implementation of this disclosure, spoken images can be applied to various scenarios to meet diverse needs, thus expanding the scope of application of the video processing solution disclosed herein.
[0071] Example process
[0072] Figure 4 shows a flowchart of a method 400 for processing video according to some implementations of the present disclosure. At block 410, in response to receiving a generation request to generate a second video based on a first video, attribute data of the first video is obtained. At block 420, based on the attribute data, text data describing the first video is determined. At block 430, the second video is generated based on the first video and audio data corresponding to the text data.
[0073] According to one example implementation of this disclosure, generating a second video based on a first video and audio data corresponding to text data includes: adding audio data corresponding to text data to the first audio data of the first video to form first updated audio data; and generating the second video based on the first image data of the first video and the first updated audio data.
[0074] According to an example implementation of this disclosure, determining the text data includes: dividing a first video into a set of video segments based on attribute data, wherein the attribute data includes at least one of the following: audio data extracted from the first video, context data of the first video, and keyframe data of the first video; and generating a set of text segments from the text data for each set of video segments.
[0075] According to an example implementation of this disclosure, generating a set of text fragments includes: for a video fragment in a set of video fragments, determining a prompt word, the prompt word instructing a preset model to process attribute data to generate a text fragment describing the video fragment; and determining the text fragment describing the video fragment based on the preset model's response to the prompt word.
[0076] According to one example implementation of this disclosure, determining a text segment further includes: generating the beginning portion of the text segment based on a previous video segment preceding the video segment.
[0077] According to an example implementation of this disclosure, generating a second video includes: for a set of video segments, determining a set of sub-segments in the video segments that do not include audio data; and adding an audio segment corresponding to a text segment to at least one of the sub-segments in the set of sub-segments to use the audio segment as narration in the second video.
[0078] According to one example implementation of this disclosure, adding an audio segment to at least one sub-segment includes: selecting a sub-segment that matches an audio segment from a set of sub-segments; and adding the audio segment to the sub-segment.
[0079] According to one example implementation of this disclosure, adding an audio segment to a sub-segment includes: creating an audio track within the sub-segment; and adding the audio segment to the audio track.
[0080] According to an example implementation of this disclosure, adding an audio segment to at least one sub-segment includes: in response to determining that the duration of the audio segment is greater than the duration of each sub-segment in a set of sub-segments, dividing the audio segment into a first audio segment and a second audio segment; selecting, respectively, a first sub-segment matching a first duration and a second sub-segment matching a second duration from a plurality of sub-segments; and adding the first audio segment to the first sub-segment and adding the second audio segment to the second sub-segment.
[0081] According to one example implementation of this disclosure, the timbre of the audio data is different from the timbre of the character in the first video.
[0082] According to an example implementation of this disclosure, generating a second video further includes: extracting at least a portion of key video segments from a set of video segments of a first video; and generating a second video based on at least a portion of the key video segments of the first video.
[0083] According to one example implementation of this disclosure, a first video is played in response to receiving a first request to disable the narration function; or a second video is played in response to receiving a second request to enable the narration function.
[0084] Example devices and equipment
[0085] Figure 5 shows a block diagram of an apparatus 500 for processing video according to some implementations of the present disclosure. The apparatus 500 includes: an acquisition module 510 configured to acquire attribute data of the first video in response to receiving a generation request to generate a second video based on a first video; a determination module 520 configured to determine text data describing the first video based on the attribute data; and a generation module 530 configured to generate the second video based on the first video and audio data corresponding to the text data.
[0086] According to some implementations of this disclosure, the generation module 530 is further configured to: add audio data corresponding to the text data to the first audio data of the first video to form first updated audio data; and generate the second video based on the first image data of the first video and the first updated audio data.
[0087] According to an example implementation of this disclosure, the determining module 520 is further configured to divide the first video into a set of video segments based on attribute data, the attribute data including at least one of the following: audio data extracted from the first video, context data of the first video, and keyframe data of the first video; and to generate a set of text segments from the text data for each set of video segments.
[0088] According to one example implementation of this disclosure, the determining module 520 is further configured to determine, for a video segment in a set of video segments, a cue word that instructs a machine learning model to process attribute data to generate a text segment describing the video segment; and to determine the text segment describing the video segment based on the machine learning model's response to the cue word.
[0089] According to one example implementation of this disclosure, the determining module 520 is also configured to generate the beginning portion of a text segment based on a previous video segment preceding the video segment.
[0090] According to an example implementation of this disclosure, the adding module 530 is further configured to, for a set of video segments in a set of video segments, determine a set of sub-segments in the video segments that do not include audio data; and add an audio segment corresponding to a text segment to at least one of the sub-segments in the set of sub-segments to use the audio segment as narration in a second video.
[0091] According to an example implementation of this disclosure, the adding module 530 is also configured to select a sub-segment that matches an audio segment from a set of sub-segments; and to add an audio segment to the sub-segment.
[0092] According to one example implementation of this disclosure, the adding module 530 is also configured to create audio tracks in sub-segments; and to add audio segments in audio tracks.
[0093] According to an example implementation of this disclosure, the adding module 530 is further configured to, in response to determining that the duration of an audio segment is greater than the duration of each of the sub-segments in a set of sub-segments, divide the audio segment into a first audio segment and a second audio segment; select, respectively, a first sub-segment matching the first duration and a second sub-segment matching the second duration of the second audio segment from a plurality of sub-segments; and add the first audio segment to the first sub-segment and add the second audio segment to the second sub-segment.
[0094] According to one example implementation of this disclosure, the timbre of the audio data is different from the timbre of the character in the first video.
[0095] According to an example implementation of this disclosure, the adding module 530 is further configured to: extract at least a portion of key video segments from a set of video segments of a first video; and generate a second video based on at least a portion of the key video segments of the first video.
[0096] According to one example implementation of this disclosure, a first video is played in response to receiving a first request to disable the narration function; or a second video is played in response to receiving a second request to enable the narration function.
[0097] Figure 6 shows a block diagram of a device 600 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 600 shown in Figure 6 can be used to implement the methods described above.
[0098] As shown in Figure 6, the computing device 600 is in the form of a general-purpose computing device. Components of the computing device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 600.
[0099] Computing device 600 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 600.
[0100] The computing device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0101] The communication unit 640 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 600 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0102] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 600 can also communicate as needed with one or more external devices (not shown) via communication unit 640. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 600, or with any device (e.g., network card, modem, etc.) that enables computing device 600 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0103] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0104] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0105] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0106] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0108] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for processing a video, comprising: in response to receiving a generation request for generating a second video based on a first video, obtaining attribute data of the first video; based on the attribute data, determining text data for describing the first video; and generating the second video based on the first video and audio data corresponding to the text data. 2.The method of claim 1, wherein the generating the second video based on the first video and the audio data corresponding to the text data comprises: adding the audio data corresponding to the text data into first audio data of the first video to form first updated audio data; and generating the second video based on first image data of the first video and the first updated audio data. 3.The method of claim 1, wherein the determining the text data comprises: dividing the first video into a set of video segments based on the attribute data, the attribute data comprising at least any one of: speech data extracted from the first video, context data of the first video, and key frame data of the first video; generating a set of text segments in the text data respectively for the set of video segments. for a video segment in the set of video segments, determining a prompt word, the prompt word instructing a preset model to process the attribute data to generate a text segment describing the video segment; and 4. The method of claim 3, wherein generating the set of text snippets comprises: determining the text segment describing the video segment based on a response of the preset model to the prompt word. generating a beginning part of the text segment based on a previous video segment before the video segment. for a video segment in the set of video segments, 5. The method of claim 4, wherein determining the text segment further comprises: determining a set of sub-segments in the video segment that do not include sound data; and 6. The method of claim 3, wherein generating the second video comprises: adding an audio segment corresponding to the text segment to at least one sub-segment in the set of sub-segments to include the audio segment as a voiceover in the second video. 7.The method of claim 6, wherein the adding the audio segment to the at least one sub-segment comprises: selecting a sub-segment from the set of sub-segments that matches the audio segment; and adding the audio segment to the sub-segment. 8.The method of claim 7, wherein the adding the audio segment to the sub-segment comprises: creating an audio track in the sub-segment; and adding the audio segment in the audio track. 9.The method of claim 6, wherein the adding the audio segment to the at least one sub-segment comprises: in response to determining that a time length of an audio segment is greater than a time length of each sub-segment in the set of sub-segments, dividing the audio segment into a first audio segment and a second audio segment; selecting a first sub-segment from the plurality of sub-segments that matches a first time length of the first audio segment and a second sub-segment that matches a second time length of the second audio segment, respectively; and adding the first audio segment to the first sub-segment and adding the second audio segment to the second sub-segment. 10. The method of claim 1, wherein a timbre of the audio data is different from a timbre of a character in the first video.
11. The method of claim 3, wherein generating a second video further comprises: extracting at least a portion of key video segments of the first video from the set of video segments; and generating the second video based on the at least a portion of key video segments of the first video.
12. The method of claim 1, wherein: the first video is played in response to receiving a first request to disable a narration function; or the second video is played in response to receiving a second request to enable a narration function.
13. An apparatus for processing a video, comprising: an obtaining module configured to obtain attribute data of a first video in response to receiving a generation request to generate a second video based on the first video; a determining module configured to determine text data for describing the first video based on the attribute data; and a generating module configured to generate the second video based on the first video and audio data corresponding to the text data.
14. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method of any of claims 1-12.
15. A computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions, when executed by a processor, cause the processor to implement the method of any of claims 1-12.
16. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method of any of claims 1-12.
Citation Information
Patent Citations
Video abstract generation method and device, equipment and storage medium
CN114143479A
Video generation method and device, electronic equipment and storage medium
CN115955585A
Video-text retrieval method and system based on image-text pre-training model
CN117556083A
Metadata generation for video indexing
US20220004574A1