Video generation method and device, storage medium and computer program product

By combining user-perspective descriptions and historical narration language features, and using a target model to extract video segments, personalized videos are generated, solving the problem of lack of distinctive features in existing technologies and achieving diversity and personalization in video generation.

CN121531189APending Publication Date: 2026-02-13MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511500070.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

The videos generated by existing technologies lack user-personalized features, resulting in videos that lack distinctiveness.

Method used

By acquiring the video to be processed and the user's viewpoint description information, the target original audio video extraction model and the target video extraction model are used to extract the original audio video segments and content video segments that match the viewpoint, respectively, and the target video is generated by combining the user's historical video narration language features.

Benefits of technology

The generated videos retain the original audio content that reflects the viewpoint, and incorporate the user's narration language features during the video selection process, thereby enhancing the diversity and personalization of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121531189A_ABST
    Figure CN121531189A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video generation method. The method comprises the following steps: acquiring a to-be-processed video and viewpoint description information of a user for the to-be-processed video; processing the to-be-processed video and the viewpoint description information by adopting a target original sound video extraction model to obtain a first video set; determining a to-be-processed video set based on the first video set and the to-be-processed video; processing the to-be-processed video set and the viewpoint description information by adopting a target video extraction model to obtain a second video set; and generating a target video based on the first video set, the second video set, the viewpoint description information and the historical video commentary language features of the user, thereby solving the problem that the generated video has no features. The embodiment of the invention further discloses video generation equipment, a storage medium and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multimedia processing, and particularly relates to a video generation method, device, storage medium and computer program product. BACKGROUND

[0002] Video generation technology is one of important applications of artificial intelligence in the multimedia field, mainly through algorithm model processing and reconstruction of input text, image or video content to generate new videos meeting user demand. In recent years, with the development of artificial intelligence generated content (AIGC), the technology of automatically generating videos based on text description has been increasingly concerned and widely applied to short video creation, content recommendation and other scenes. At present, the related technology usually adopts a large language model combined with a video generation model to directly generate complete video content according to user text input. This method relies on the understanding and generation ability of the model to the text, resulting in no characteristics of the generated video and lack of personalized features of the user. SUMMARY

[0003] To solve the above technical problems, the embodiments of the present application provide a video generation method, device, storage medium and computer program product, which solve the problem of no characteristics of the generated video in the related technology and realize the personalized features of the user in the generated video.

[0004] To achieve the above object, the technical scheme of the embodiments of the present application is as follows: A video generation method, the method comprising: obtaining a to-be-processed video and viewpoint description information of a user for the to-be-processed video; processing the to-be-processed video and the viewpoint description information by using a target original sound video extraction model to obtain a first video set; determining a to-be-processed video set based on the first video set and the to-be-processed video; processing the to-be-processed video set and the viewpoint description information by using a target video extraction model to obtain a second video set; generating a target video based on the first video set, the second video set, the viewpoint description information and historical video commentary language features of the user.

[0005] In the above scheme, the determination of the to-be-processed video set based on the first video set and the to-be-processed video comprises: eliminating the first video set from the to-be-processed video to obtain a residual video set; Each video in the remaining video set is segmented based on the target segmentation duration to obtain the video set to be processed.

[0006] In the above scheme, generating the target video based on the first video set, the second video set, the viewpoint description information, and the user's historical video narration language features includes: A first target video set is determined based on the first video set; A second target video set is determined based on the second video set; Based on the timing information of the videos to be processed, the first target video set and the second target video set are processed to obtain the video set to be synthesized; The target video is generated based on the set of videos to be synthesized, the viewpoint description information, and the linguistic features of the historical video narration.

[0007] In the above scheme, determining the first target video set based on the first video set includes: Determine the relationship between the first total duration of the videos in the first video set and the first target duration; If the first total duration is greater than the first target duration, the first target video set is determined from the first video set based on the weight of each video in the first video set; If the first total duration is less than or equal to the first target duration, the first video set is determined to be the first target video set. Accordingly, determining the second target video set based on the second video set includes: Calculate the duration difference between the first target duration and the second target duration, and determine the relationship between the second total duration of the videos in the second video set and the duration difference; If the second total duration is greater than the duration difference, the second target video set is determined from the second video set based on the weight of each video in the second video set; If the second total duration is less than or equal to the duration difference, the second video set is determined to be the second target video set.

[0008] In the above scheme, generating the target video based on the set of videos to be synthesized, the viewpoint description information, and the linguistic features of the historical video narration includes: Determine the synthesis order of each video in the set of videos to be synthesized; The synthesis order, the viewpoint description information, and the linguistic features of the historical video narration are processed using a target text model to generate narration text for each video in the video set to be synthesized. The narration of each video is processed to obtain multiple narration audios; The target video is generated based on the multiple audio narrations and the set of videos to be synthesized.

[0009] The method in the above scheme further includes: Acquire a sample video and determine the first audio data of the sample video; The first audio data is segmented to obtain multiple sample sub-audio data; Determine the first sample viewpoint description information for each sample sub-audio data, and extract features from the first sample viewpoint description information to obtain the first sample feature information; Feature extraction is performed on the multiple sample sub-audio data to obtain the second sample feature information; Based on the first sample feature information and the second sample feature information, the initial original audio video extraction model is trained to obtain the target original audio video extraction model.

[0010] The method in the above scheme further includes: Acquire sample videos and segment the sample videos to obtain multiple sample sub-videos; Determine the keyframe images and the second audio data of each sample sub-video; Determine the second sample viewpoint description information for each of the second audio data, and perform feature extraction on multiple second sample viewpoint description information to obtain third sample feature information; Feature extraction is performed on multiple keyframe images to obtain fourth sample feature information; The initial video extraction model is trained based on the feature information of the third sample and the feature information of the fourth sample to obtain the target video extraction model.

[0011] A video generation device, the device comprising: a processor and a memory for storing a computer program capable of running on the processor; The processor is used to run the computer program to implement the steps of the above-described method.

[0012] A computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the above method.

[0013] A computer program product comprising a computer program that, when executed by a processor, implements the above-described method.

[0014] The video generation method, device, storage medium, and computer program product provided in the embodiments of this application can acquire the video to be processed and the user's viewpoint description information regarding the video to be processed. A first video set is obtained by processing the video to be processed and the viewpoint description information using a target original audio video extraction model. A second video set is obtained by determining the video set to be processed based on the first video set and the video to be processed, and then processing the video set to be processed and the viewpoint description information using a target video extraction model. A target video is generated based on the first video set, the second video set, the viewpoint description information, and the user's historical video narration language features. In this way, by combining the user's viewpoint description information, the target original audio video extraction model and the target video extraction model respectively extract original audio video segments and content video segments that conform to the viewpoint, and generate the final video based on the user's historical video narration language features. This ensures that the generated video not only retains the original audio content reflecting the viewpoint but also incorporates the user's narration language features during the video selection process, solving the problem of unoriginal videos in related technologies, thereby significantly improving the diversity and personalization of video generation. Attached Figure Description

[0015] Figure 1 A flowchart illustrating a video generation method provided for an embodiment of this application; Figure 2 A schematic diagram illustrating the determination of a target video extraction model and a target original audio video extraction model in a video generation method provided for an embodiment of this application; Figure 3 A schematic diagram of the structure of a video generation device provided for an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a video generation device provided for an embodiment of this application. Detailed Implementation

[0016] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0017] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0018] Unless otherwise specified, any step in the embodiments of this application performed by the electronic device may be executed by the processor of the electronic device. It is also worth noting that the embodiments of this application do not limit the order in which the electronic device performs the following steps. Furthermore, the methods used to process data in different embodiments may be the same or different methods. It should also be noted that any step in the embodiments of this application can be executed independently by the electronic device; that is, when the electronic device performs any step in the following embodiments, it may not depend on the execution of other steps.

[0019] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0020] This application provides a video generation method, which can be applied to a video generation device, as described above. Figure 1 As shown, the method may include the following steps: Step 101: Obtain the video to be processed and the user's opinion description of the video to be processed.

[0021] The video to be processed can be any length of original video content; for example, it could be a complete recording of a football match. The opinion description information consists of the user's subjective evaluation of the video, highlighting key points or providing commentary. For example, the user might want to showcase exciting goals or emphasize player performances. This opinion description information guides the extraction of video segments from the long video that align with the user's intent and serves as crucial input for generating the commentary.

[0022] Specifically, the system can retrieve videos to be processed from user-uploaded video sources, and simultaneously receive opinion descriptions submitted by users via text input boxes, voice recognition, and other methods. The videos to be processed and the opinion descriptions together constitute the foundational data for subsequent video generation.

[0023] Step 102: Use the target original audio video extraction model to process the video to be processed and the opinion description information to obtain the first video set.

[0024] The target original audio video extraction model is a deep learning-based model used to extract original audio video segments from raw videos that match user viewpoint descriptions. By analyzing the audio data and semantic features of the video to be processed, the model identifies video segments that match the user's viewpoint (i.e., viewpoint description information) and retains the original audio content extracted by the model, thus reflecting the subjective expression contained in the user's viewpoint description information. Furthermore, the target original audio video extraction model can be used with M... s The first video set can be represented as {V}. 原声视频1 =(V[t1到t11] Q1), V 原声视频2 =(V [t2到t22] Q2), ..., V 原声视频n =(V [tn到tnn] Q n )}.

[0025] Specifically, the video to be processed and the opinion description information can be used as input information and fed into the target original audio video extraction model. The target original audio video extraction model then extracts audio data from the video to be processed and performs sentence segmentation, cleaning, and encoding on the audio data. Then, the semantic features of the user's opinion description information are combined with the above-processed audio data for further processing, and finally, original audio video segments that match the opinion description information are selected.

[0026] It should be noted that the first video set refers to a collection of video clips containing original audio extracted based on the target original audio video extraction model. The video clips in the first video set not only retain the original audio, but also express the content that the user is interested in.

[0027] Step 103: Determine the video set to be processed based on the first video set and the video to be processed.

[0028] The video set to be processed is obtained by processing the remaining videos after removing the first video set from the video set to be processed; specifically, the remaining videos can be segmented to obtain the video set to be processed.

[0029] It should be noted that the non-original audio video content in the video set to be processed is usually used to further extract video content so that it can be spliced ​​and combined with the original audio clips to form the final short video (i.e., the target video).

[0030] Step 104: Use the target video extraction model to process the video set to be processed and the opinion description information to obtain the second video set.

[0031] Among them, the target video extraction model is another deep learning-based model. This model extracts video content from the original video (i.e., the set of videos to be processed) that matches the user's viewpoint, without including the original audio. The target video extraction model primarily relies on the visual features and temporal series information of video frames. It can extract video segments that reflect the user's viewpoint, which can then be used for subsequent video synthesis. The target video extraction model can be used with M... v express.

[0032] Specifically, the video set to be processed and the opinion description information can be used as input information and fed into the target video extraction model. The target video extraction model then segments the video set to be processed, extracts keyframes and audio data, and combines this data with the cue words in the opinion description information for further processing. Finally, it outputs a set of video clips that match the user's opinion (i.e., the second video set). It should be noted that the second video set is a collection of video clips extracted from the video set to be processed, which does not contain the original audio but matches the user's opinion description information; the second video set is mainly used to supplement the original audio video clips in the first video set to enrich the video content.

[0033] Step 105: Generate the target video based on the first video set, the second video set, the viewpoint description information, and the user's historical video narration language features.

[0034] Among them, historical video narration language features refer to the narration language style, speaking speed, tone, and other characteristics used by users in their previous short video productions. The system can use natural language processing technology to model based on historical video narration language features to generate personalized narration text consistent with user habits.

[0035] Specifically, in the process of generating the target video, the content of videos in the first and second video sets is comprehensively considered, arranged in chronological order, and combined with the user's historical video narration language characteristics to generate the target video. Furthermore, suitable background music or sound effects can be selected based on the user's preferences, enhancing the video's expressiveness.

[0036] It should be noted that a video set to be synthesized can be generated first based on the first and second video sets, combined with the temporal information of the videos to be processed. Then, the target video can be generated based on the video set to be synthesized, the viewpoint description information, and the language features of the historical video narration.

[0037] In other embodiments of this application, step 103 described above can be implemented in the following ways: A1. Remove the first video set from the videos to be processed to obtain the remaining video set.

[0038] The remaining video set is a collection of multiple video segments obtained by deleting the first video set from the videos to be processed. It's important to note that removing the first video set ensures that subsequent processing only targets the unselected video segments, thereby improving video processing efficiency and guaranteeing the diversity of output content.

[0039] Specifically, the remaining video set {V} can be obtained by removing videos in the time intervals {[t1 to t11], [t2 to t22], ..., [tn to tnn]} from the first video set. [0到t1] V [t11到t2] V[t22到t3] , ..., V [tnn到v总时长]}

[0040] A2. Based on the target segmentation duration, segment each video in the remaining video set to obtain the video set to be processed.

[0041] The target segmentation duration refers to the preset time length for video segmentation, such as 3 seconds. The target segmentation duration can be set based on the user's historical video style preferences to ensure the rhythm and viewing experience of the output video. The video set to be processed can be formed by dividing each video in the remaining video set according to the target segmentation duration, creating multiple sets of temporally consecutive video segments. If the length of a video segment is less than one target segmentation duration, the video segment will be adjusted according to its actual length to ensure that all video segments meet the minimum unit requirement.

[0042] Specifically, regarding V [0到t1] t1 / 3 = X1, t1%3 = Y1. If Y1 equals 0, the video is divided into X1 segments; otherwise, it is divided into X1+1 viewpoint segments. (For V...) [t11到t2] (t2-t11) / 3=X2, (t2-t11)%3=Y2. If Y2 equals 0, then the video is divided into X2 segments; if Y2 does not equal 0, then it is divided into X2+1 viewpoint segments. Similarly, for {V... [t22到t3] ,…,V [tnn到v总时长] The video set {V} is processed to obtain the final video set to be processed. [0到3] V [4到6] V [x+(t-y]*3+1到x+(t-y+1)*3-z] …,V [x+t*3+1到t总]}; where x takes the value of an integer in {t1,t2,t3,…,tn}, t takes the value of an integer greater than 1, y takes the value of an integer less than t, and z takes the value of an integer between 0 and 3.

[0043] It should be noted that if the target segmentation duration is 3 seconds, the video set to be processed can include the following videos: (For V) [0到t1] In the case where t1 / 3 = X1 and t1%3 = Y1, there is a video set {V}. [0到3] V [4到6] , ..., V [(x1-1)*3+1到x1*3]}(Y1=0),{V [0到3] V [4到6] , ..., V [x1*3+1到t1]} (Y1) 0); for V [t11到t2] In the case of (t2-t11) / 3=X2 and (t2-t11)%3=Y2, the video set {V} is given. [t11到t11+3] V [t11+4到t11+6] , ..., V [t11+(x2-1)*3+1到t11+x2*3]}(Y2=0),{V[t11到t11+3] V [t11+4到t11+6] , ..., V [t11+x2*3+1到t12]} (Y2) 0); for V [tnn到V总时长] In the case of (t_total - t_nn) / 3 = X_n and (t_total - t_nn)%3 = Y_n, {V [tnn到tnn+3] V [tnn+4到tnn+6] , ..., V [tnn+(xn-1)*3+1到tnn+xn*3]}(Yn=0),{V [tnn到tnn+3] V [tnn+4到tnn+6] , ..., V [tnn+xn*3+1到t总]}(Yn 0).

[0044] In other embodiments of this application, step 105 described above can be implemented in the following ways: B1. Determine the first target video set based on the first video set.

[0045] Specifically, the first target video set can be determined from the first video set based on the first total duration and the first target duration of the videos in the first video set.

[0046] The first target duration refers to the length of the original audio video to be generated, preset by the user. This first target duration is typically derived through modeling based on historical data (i.e., the user's past habits of publishing narration videos) to match the user's personalized style requirements. Furthermore, the first target duration can be represented by T1.

[0047] B2. Determine the second target video set based on the second video set.

[0048] Specifically, the second target video set can be determined from the second video set based on the second total duration, the first target duration, and the second target duration of the videos in the second video set.

[0049] The second target duration refers to the pre-set video length that the user wants to generate. This second target duration is typically derived through modeling based on historical data (i.e., the user's past habits of publishing narration videos) to match the user's personalized style requirements. Furthermore, the second target duration can be represented by T.

[0050] B3. Based on the temporal information of the videos to be processed, the first target video set and the second target video set are processed to obtain the video set to be synthesized.

[0051] The temporal information of the video to be processed refers to the timeline structure of the original long video (i.e., the video to be processed), that is, the temporal order in which the various video segments are arranged. Specifically, by temporally sorting and splicing the first and second target video sets to be processed, the first and second target video sets can be merged into a coherent video sequence according to the playback logic of the video to be processed, thus forming the video set to be synthesized. This operation process ensures that the final generated video maintains consistency in timeline and logic and can naturally connect video segments from different sources.

[0052] In this embodiment of the application, by integrating multiple video segments based on the timing information of the video to be processed, seamless connection of video content can be achieved, thereby improving the overall smoothness and viewing comfort of the video, and thus enhancing the user's understanding and acceptance of the video content.

[0053] B4. Generate the target video based on the video set to be synthesized, the viewpoint description information, and the language features of historical video narration.

[0054] Specifically, the video set to be synthesized can be combined with opinion description information, and narration that matches the user's style can be automatically generated based on the user's historical language characteristics. Then, the final target video can be generated based on the generated narration and the video set to be synthesized. This process achieves the unification of video content, voice expression and user's personalized style.

[0055] It should be noted that the first and second target video sets are complementary; the first set focuses on the matching degree of content and viewpoint description, while the second set emphasizes the consistency between the original audio and the user's style. The combination of the first and second sets jointly supports the expression of viewpoint description at both visual and auditory levels. Furthermore, the generation process of the target videos needs to incorporate the user's language characteristics, ensuring that the target videos conform to the user's habits and preferences in terms of language expression, thereby achieving multi-dimensional personalized expression.

[0056] In other embodiments of this application, the above-described determination of the first target video set based on the first video set includes: Determine the relationship between the first total duration of the videos in the first video set and the first target duration.

[0057] The first total duration refers to the total playback time of all video segments in the first video set. The first target duration is a desired original audio video duration set by the user based on their past short video style.

[0058] In this embodiment, by calculating the relationship between the first total duration and the first target duration, it can be determined whether the videos in the first video set need to be filtered, thereby reasonably allocating video content and ensuring that the final generated video meets the user's expected duration requirements. Specifically, the relationship between the first total duration and the first target duration determines whether a weighted filtering process is needed; if the first total duration is greater than the first target duration, some videos are filtered from all videos according to weights to meet the first target duration limit; if the first total duration is less than or equal to the first target duration, no weighted filtering is needed, and all videos can be used directly.

[0059] If the first total duration is greater than the first target duration, the first target video set is determined from the first video set based on the weight of each video in the first video set; If the first total duration is less than or equal to the first target duration, the first video set is determined as the first target video set.

[0060] In the first video set, the weight of each video is a numerical indicator derived from a comprehensive evaluation of factors such as the relevance of the video content to the viewpoint described and the video quality. Videos with higher weights indicate that they better meet users' needs in expressing their opinions and are closer to users' historical style preferences. Specifically, deep learning models can be used to extract features from the videos, and weight values ​​can be calculated based on these features.

[0061] In this embodiment, if the total duration is greater than the target duration, the top N videos by weight in the first video set are selected as the first target video set based on the weight of each video in the first video set. If the total duration is less than or equal to the target duration, it indicates that the current first video set already meets the user's duration requirement, and no further selection is needed; the current first video set is directly used as the first target video set. It should be noted that by selecting videos based on weight, video content that best matches the user's viewpoint and style can be prioritized, thereby improving the quality of the final video and enhancing its expressiveness and appeal.

[0062] In other embodiments of this application, the above-described determination of the second target video set based on the second video set includes: Calculate the duration difference between the first target duration and the second target duration, and determine the relationship between the second total duration of the videos in the second video set and the duration difference.

[0063] The second target duration is a desired total video length set by the user based on their past short video style, used to control the length range of the final output video. The duration difference is the second target duration minus the first target duration, used to measure the remaining time space available for non-original audio videos. The second total duration is the total playback time of all videos in the second video set.

[0064] Specifically, if the second total duration exceeds the duration difference, it means that a portion of the videos in the second video set needs to be selected to meet the time limit, and this selection is based on the weight of these videos. In practice, the relationship between the second total duration and the duration difference determines whether the second video set needs to be selected. If the second total duration is greater than the duration difference, some content needs to be selected based on weight to meet the time limit; if it is less than or equal to, all content can be used directly.

[0065] If the second total duration is greater than the duration difference, the second target video set is determined from the second video set based on the weight of each video in the second video set; If the second total duration is less than or equal to the duration difference, the second video set is determined as the second target video set.

[0066] In this embodiment, if the second total duration is greater than the duration difference, then based on the weight of each video in the second video set, the top N videos by weight are selected as the second target video set. If the second total duration is less than or equal to the duration difference, it indicates that the second video set does not require filtering to meet the set time requirement, and therefore the second video set can be directly used as the second target video set.

[0067] It's worth noting that using a weighted approach to filter the second video set prioritizes high-quality videos that align with the chosen viewpoints, ensuring the overall quality of the final videos and enhancing the user viewing experience. Furthermore, by separately determining the duration and weighting the first and second video sets, the total length and content quality of the generated videos can be precisely controlled, enabling personalized video generation solutions and ultimately improving video diversity and user satisfaction.

[0068] In other embodiments of this application, the above-mentioned generation of target videos based on the set of videos to be synthesized, viewpoint description information, and linguistic features of historical video narration includes: Determine the synthesis order of each video in the video set to be synthesized.

[0069] The compositing order refers to the sequence in which multiple video segments are arranged according to a specific logic or timeline when generating the target video. The compositing order is the sequential order in which the original video segments in the set to be composited are joined.

[0070] The target text model is used to process the synthesis order, viewpoint description information, and narration language features of historical videos to generate narration text for each video in the video set to be synthesized.

[0071] The target text model can be a trained natural language processing model used to generate narration that matches the user's style and context. The target text model can be personalized based on the user's historical narration language characteristics (such as speaking speed, tone, frequently used vocabulary, sentence structure, etc.) to ensure that the generated narration is consistent with the user's previous video style. Specifically, the target text model can refer to a large text model, such as Kimi.

[0072] Specifically, the synthesis order, viewpoint description information, and linguistic features of historical video narration can be used as input information and fed into the target text model. The target text model then analyzes and processes the synthesis order, viewpoint description information, and linguistic features of historical video narration to output narration text for each video in the video set to be synthesized.

[0073] In this embodiment of the application, by combining the user's historical language features and current viewpoint description information to generate narration, the rhythm of the newly generated narration is closer to the user's expression habits, which can realize highly personalized video creation, thereby meeting the user's unique expression needs and enhancing the uniqueness and attractiveness of the video.

[0074] The narration for each video is processed to obtain multiple audio narrations.

[0075] The narration audio is the output result after converting the narration text into a speech signal. The processing typically includes text-to-speech (TTS), timbre adjustment, rhythm control, and background music integration. To make the audio more closely resemble the user's previous video style, open-source voice replication technologies (such as OpenVoice and CosyVoice) can be used to convert the narration text into narration audio. It should be noted that the system can learn from the user's past narration audio, extracting voiceprint features from them, and then loading these features into an open-source voice replication model for further learning, thereby achieving consistency in voice style.

[0076] The target video is generated based on multiple audio narrations and a set of videos to be synthesized.

[0077] The process involves synchronously assembling the narration audio with corresponding video segments from the set of videos to be synthesized, forming a complete short video (i.e., the target video). During this process, the system must ensure strict time-series matching between the audio and video, and consider factors such as volume balance, transition effects, and subtitle overlay to improve the quality of the final target video.

[0078] In this embodiment of the application, by precisely synchronizing and integrating the narration audio and video clips into a complete target video, it is possible to ensure that the target video content is clear, smooth, and expressive, thereby effectively improving the viewing experience of the target video and enhancing users' recognition of the target video content and their willingness to share it.

[0079] In other embodiments of this application, the method may further include: Acquire sample videos and determine the first audio data of the sample videos.

[0080] Sample videos refer to the original video footage used for model training. The content of sample videos includes short videos, long video clips, or other video content related to the user's style that the user has previously created. For example... Figure 2 The first audio data shown is the audio portion extracted from the sample video, typically the raw, unprocessed sound signal, such as narration, background noise, and music. It should be noted that by collecting diverse sample videos, different scenarios, tones, and rhythms of speech can be covered, thereby improving the model's ability to understand users' personalized styles.

[0081] The first audio data is segmented to obtain multiple sample sub-audio data.

[0082] Among them, such as Figure 2 The diagram illustrates how a continuous audio stream can be divided into several shorter audio segments, with the audio content within each segment constituting a sample sub-audio data. Performing this segmentation operation helps the model better understand and recognize semantic information in the audio. The segmentation of sample sub-audio data can be dynamically adjusted based on a fixed time interval (e.g., 3 seconds) or based on speech activity detection results to ensure that each sample sub-audio data contains meaningful speech content. This method of segmenting sample sub-audio data not only improves the processability of audio data but also enhances the model's ability to perceive local semantics.

[0083] It should be noted that, as Figure 2 The sample sub-audio data shown can be obtained by segmenting the first audio data and then filtering and cleaning the segmented audio data.

[0084] Determine the first sample viewpoint description information for each sample sub-audio data, and extract features from the first sample viewpoint description information to obtain the first sample feature information.

[0085] The first sample viewpoint description information refers to the text description generated for each sample sub-audio data, summarizing the main viewpoint or theme of the audio content. Specifically, it can be obtained by converting the audio content into text using Automatic Speech Recognition (ASR) technology, followed by processing using Natural Language Processing (NLP) technology. Furthermore, the first sample feature information can be a first sample feature vector, and it can be obtained by encoding the first sample viewpoint description information.

[0086] Feature extraction is performed on multiple sample sub-audio data to obtain the feature information of the second sample.

[0087] The second sample feature information is audio features extracted directly from multiple sample sub-audio data. Furthermore, the second sample feature information can be obtained by encoding the multiple sample sub-audio data.

[0088] Based on the feature information of the first sample and the feature information of the second sample, the initial original audio video extraction model is trained to obtain the target original audio video extraction model.

[0089] Specifically, such as Figure 2 As shown, the feature information of the first sample and the feature information of the second sample can be used as input information to feed into the initial original audio video extraction model for training, thereby obtaining the target original audio video extraction model. Specifically, the initial original audio video extraction model can be a Long Short-Term Memory (LSTM) network model. Furthermore, the first and second feature information can be dimensionality-reduced using Word2Vec (Word to Vector) before being used for model training.

[0090] It should be noted that during the training of the initial original audio video extraction model, previous user commentary videos can be obtained to acquire the original videos. The trained intermediate original audio video extraction model is then used to extract video frames that match the user's input opinion prompts from the original videos. The original audio video frames are separated from the user's commentary videos. A loss function is used to determine whether the original audio extracted by the intermediate original audio video extraction model is infinitely close to the original audio of the user's commentary videos. The model is then optimized and adjusted based on the difference between the two to finally obtain the target original audio video extraction model.

[0091] In other embodiments of this application, the method may further include: Obtain sample videos and segment them to obtain multiple sample sub-videos.

[0092] Among them, such asFigure 2 Therefore, a fixed duration can be used to segment the sample video. This fixed duration can be a time unit set according to task requirements, such as 3 seconds. It should be noted that segmentation can improve the granularity of video processing, enabling the model to more accurately identify sample sub-videos related to the viewpoint, thereby improving training efficiency and model accuracy.

[0093] Determine the keyframe images and the second audio data for each sample sub-video.

[0094] Among them, such as Figure 2 The keyframe images shown can be keyframe images extracted from each sample sub-video, typically the images of the frames in the sample sub-video that best reflect changes in content or structure; and, as... Figure 2 The second audio data shown refers to the audio information extracted from the corresponding sample sub-video, which usually includes speech, background sound effects, etc., and is used for subsequent sound feature analysis.

[0095] Determine the second sample viewpoint description information for each second audio data, and extract features from multiple second sample viewpoint description information to obtain the third sample feature information.

[0096] The second sample opinion description information is a textual description of the opinion content expressed in each sample sub-video, typically generated by human annotators or a large language model. This second sample opinion description information reflects the core opinion or theme of the sample sub-video content. The feature extraction process can use natural language processing techniques (such as Word2Vec) to encode the second sample opinion description information, generating a numerical vector representation that can be used by the model. This representation is called a numeric vector representation. Figure 2 The third sample feature information is shown.

[0097] Feature extraction was performed on multiple keyframe images to obtain the feature information of the fourth sample.

[0098] Feature extraction from keyframe images typically employs image processing techniques such as convolutional neural networks to convert each keyframe into a high-dimensional feature vector, forming a structure like... Figure 2 The fourth sample feature information is shown.

[0099] The initial video extraction model is trained based on the feature information of the third and fourth samples to obtain the target video extraction model.

[0100] Specifically, such as Figure 2As shown, the feature information of the third and fourth samples can be used as input information to feed into the initial video extraction model, so as to train the initial video extraction model and then obtain the target video extraction model. The initial video extraction model is a model architecture that has not undergone specific training, and is usually designed based on structures such as UNet and Transformer.

[0101] It should be noted that during the training of the initial video extraction model, previous user commentary videos can be obtained to acquire the original videos. The trained intermediate video extraction model then extracts video frames matching the user's input opinion prompts from the original videos, separating the video frames excluding the original audio from the user commentary videos. Finally, a loss function is used to extract a set I of video frames from the intermediate video extraction model. x Explaining the video frame set I to the user y By comparing the two models, the intermediate video extraction model is optimized and adjusted based on the difference between them to obtain the target video extraction model.

[0102] In other embodiments of this application, this application can extract content from the video that matches the author's historical short video production style (narration, voice-over, etc.) and combine it with the author's understanding of the long video content to generate narration. At the same time, the original audio of the video that matches the original audio of the viewpoint is preserved, as the original audio is used to express the viewpoint.

[0103] The video generation method provided in the embodiments of this application can combine user viewpoint description information, use a target original audio video extraction model and a target video extraction model to extract original audio video segments and content video segments that conform to the viewpoint, and generate the final video based on the user's historical video narration language features. This allows the generated video to not only retain the original audio content that reflects the viewpoint, but also to introduce the user's narration language features during the video selection process, solving the problem of uncharacteristic videos generated in related technologies, thereby significantly improving the diversity and personalization of video generation.

[0104] Based on the foregoing embodiments, embodiments of this application provide a video generation apparatus that can be applied to... Figure 1 In the corresponding embodiment of the video generation method, refer to Figure 3 As shown, the video generation device 2 may include: an acquisition unit 21, a first processing unit 22, a determination unit 23, a second processing unit 24, and a generation unit 25, wherein: Acquisition unit 21 is used to acquire the video to be processed and user opinion description information about the video to be processed; The first processing unit 22 is used to process the video to be processed and the viewpoint description information using the target original sound video extraction model to obtain the first video set; Determining unit 23 is used to determine the video set to be processed based on the first video set and the video to be processed; The second processing unit 24 is used to process the video set to be processed and the viewpoint description information using a target video extraction model to obtain a second video set; The generation unit 25 is used to generate target videos based on the first video set, the second video set, viewpoint description information, and the user's historical video narration language features.

[0105] In other embodiments of this application, the determining unit 23 is further configured to perform the following steps: Remove the first video set from the videos to be processed to obtain the remaining video set; Each video in the remaining video set is segmented based on the target segmentation duration to obtain the video set to be processed.

[0106] In other embodiments of this application, the generation unit 25 is further configured to perform the following steps: The first target video set is determined based on the first video set; Determine the second target video set based on the second video set; Based on the temporal information of the videos to be processed, the first target video set and the second target video set are processed to obtain the video set to be synthesized. The target video is generated based on the set of videos to be synthesized, the viewpoint description information, and the language features of historical video narration.

[0107] In other embodiments of this application, the generation unit 25 is further configured to perform the following steps: Determine the relationship between the first total duration of the videos in the first video set and the first target duration; If the first total duration is greater than the first target duration, the first target video set is determined from the first video set based on the weight of each video in the first video set; If the first total duration is less than or equal to the first target duration, the first video set is determined as the first target video set.

[0108] In other embodiments of this application, the generation unit 25 is further configured to perform the following steps: Calculate the duration difference between the first target duration and the second target duration, and determine the relationship between the second total duration of the videos in the second video set and the duration difference; If the second total duration is greater than the duration difference, the second target video set is determined from the second video set based on the weight of each video in the second video set; If the second total duration is less than or equal to the duration difference, the second video set is determined as the second target video set.

[0109] In other embodiments of this application, the generation unit 25 is further configured to perform the following steps: Determine the synthesis order of each video in the video set to be synthesized; The target text model is used to process the synthesis order, viewpoint description information and historical video narration language features to generate narration text for each video in the video set to be synthesized; The narration for each video is processed to obtain multiple audio narrations; The target video is generated based on multiple audio narrations and a set of videos to be synthesized.

[0110] In other embodiments of this application, the determining unit 23 is further configured to perform the following steps: Acquire sample videos and determine the first audio data of the sample videos; The first audio data is segmented to obtain multiple sample sub-audio data; Determine the first sample viewpoint description information for each sample sub-audio data, and extract features from the first sample viewpoint description information to obtain the first sample feature information; Feature extraction is performed on multiple sample sub-audio data to obtain the feature information of the second sample; Based on the feature information of the first sample and the feature information of the second sample, the initial original audio video extraction model is trained to obtain the target original audio video extraction model.

[0111] In other embodiments of this application, the determining unit 23 is further configured to perform the following steps: Acquire sample videos and segment them to obtain multiple sample sub-videos; Determine the keyframe images and the second audio data of each sample sub-video; Determine the second sample viewpoint description information for each second audio data, and extract features from multiple second sample viewpoint description information to obtain the third sample feature information; Feature extraction is performed on multiple keyframe images to obtain the feature information of the fourth sample; The initial video extraction model is trained based on the feature information of the third and fourth samples to obtain the target video extraction model.

[0112] It should be noted that the specific implementation process of the steps performed by each unit in the embodiments of this application can be referred to Figure 1 The implementation process of the video generation method provided in the corresponding embodiment will not be described in detail here.

[0113] The video generation apparatus provided in the embodiments of this application can extract original audio video segments and content video segments that conform to the viewpoint by combining the user's viewpoint description information and using the target original audio video extraction model and the target video extraction model, respectively. The final video is generated based on the user's historical video narration language features. This makes the generated video not only retain the original audio content that reflects the viewpoint, but also introduce the user's narration language features during the video selection process. This solves the problem of the lack of distinctive features in the generated videos in related technologies, thereby significantly improving the diversity and personalization of video generation.

[0114] Based on the foregoing embodiments, embodiments of this application provide a video generation device that can be applied to... Figure 1 In the corresponding embodiment of the video generation method, refer to Figure 4 As shown, the video generation device 3 may include: a processor 31, a memory 32, and a communication bus 33, wherein: Communication bus 33 is used to realize the communication connection between processor 31 and memory 32; The memory 32 is used to store computer programs that can run on the processor 31; Processor 31 is used to run computer programs to perform the following steps: Obtain the video to be processed and user comments and descriptions of the video to be processed; A target original audio video extraction model is used to process the video to be processed and the opinion description information to obtain the first video set; Based on the first video set and the videos to be processed, determine the video set to be processed; A target video extraction model is used to process the video set to be processed and the opinion description information to obtain a second video set; The target video is generated based on the first video set, the second video set, the viewpoint description information, and the user's historical video narration language features.

[0115] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Remove the first video set from the videos to be processed to obtain the remaining video set; Each video in the remaining video set is segmented based on the target segmentation duration to obtain the video set to be processed.

[0116] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: The first target video set is determined based on the first video set; Determine the second target video set based on the second video set; Based on the temporal information of the videos to be processed, the first target video set and the second target video set are processed to obtain the video set to be synthesized. The target video is generated based on the set of videos to be synthesized, the viewpoint description information, and the language features of historical video narration.

[0117] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Determine the relationship between the first total duration of the videos in the first video set and the first target duration; If the first total duration is greater than the first target duration, the first target video set is determined from the first video set based on the weight of each video in the first video set; If the first total duration is less than or equal to the first target duration, the first video set is determined as the first target video set.

[0118] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Calculate the duration difference between the first target duration and the second target duration, and determine the relationship between the second total duration of the videos in the second video set and the duration difference; If the second total duration is greater than the duration difference, the second target video set is determined from the second video set based on the weight of each video in the second video set; If the second total duration is less than or equal to the duration difference, the second video set is determined as the second target video set.

[0119] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Determine the synthesis order of each video in the video set to be synthesized; The target text model is used to process the synthesis order, viewpoint description information and historical video narration language features to generate narration text for each video in the video set to be synthesized; The narration for each video is processed to obtain multiple audio narrations; The target video is generated based on multiple audio narrations and a set of videos to be synthesized.

[0120] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Acquire sample videos and determine the first audio data of the sample videos; The first audio data is segmented to obtain multiple sample sub-audio data; Determine the first sample viewpoint description information for each sample sub-audio data, and extract features from the first sample viewpoint description information to obtain the first sample feature information; Feature extraction is performed on multiple sample sub-audio data to obtain the feature information of the second sample; Based on the feature information of the first sample and the feature information of the second sample, the initial original audio video extraction model is trained to obtain the target original audio video extraction model.

[0121] In other embodiments of this application, the processor 31 is used to run a computer program and may also perform the following steps: Acquire sample videos and segment them to obtain multiple sample sub-videos; Determine the keyframe images and the second audio data of each sample sub-video; Determine the second sample viewpoint description information for each second audio data, and extract features from multiple second sample viewpoint description information to obtain the third sample feature information; Feature extraction is performed on multiple keyframe images to obtain the feature information of the fourth sample; The initial video extraction model is trained based on the feature information of the third and fourth samples to obtain the target video extraction model.

[0122] It should be noted that a detailed description of the steps performed by the processor can be found in [reference needed]. Figure 1 The video generation method provided in the corresponding embodiments will not be described in detail here.

[0123] The video generation device provided in the embodiments of this application can extract original audio video segments and content video segments that conform to the viewpoint by combining the user's viewpoint description information and using the target original audio video extraction model and the target video extraction model, respectively. The final video is generated based on the user's historical video narration language features. This makes the generated video not only retain the original audio content that reflects the viewpoint, but also introduce the user's narration language features during the video selection process. This solves the problem of the lack of distinctive features in the generated videos in related technologies, thereby significantly improving the diversity and personalization of video generation.

[0124] Based on the foregoing embodiments, embodiments of this application provide a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement... Figure 1 The corresponding embodiment provides the steps of the video generation method.

[0125] Based on the foregoing embodiments, embodiments of this application provide a computer program product, including a computer program that can be executed by a processor 31 to perform... Figure 1 The corresponding embodiment provides the steps of the video generation method.

[0126] In some embodiments, the computer-readable storage medium may be a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or an optical disk-read-only memory (CD-ROM); or it may be a device that includes one or any combination of the above-mentioned memories.

[0127] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0128] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system, may be in one or more scripts in a (yperText Markup Language, HTML) file, stored in a single file dedicated to the program in question, or stored in multiple co-located files (e.g., files storing one or more modules, subroutines, or code sections).

[0129] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0130] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes ​ The steps of the function specified in one or more boxes.

[0133] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video generation method, characterized in that, The method includes: Obtain the video to be processed and user opinions / descriptions regarding the video to be processed; The target original audio video extraction model is used to process the video to be processed and the opinion description information to obtain the first video set; Based on the first video set and the video to be processed, determine the video set to be processed; The target video extraction model is used to process the video set to be processed and the opinion description information to obtain a second video set; The target video is generated based on the first video set, the second video set, the opinion description information, and the user's historical video narration language features.

2. The method according to claim 1, characterized in that, The step of determining the video set to be processed based on the first video set and the video to be processed includes: Remove the first video set from the videos to be processed to obtain the remaining video set; Each video in the remaining video set is segmented based on the target segmentation duration to obtain the video set to be processed.

3. The method according to claim 1, characterized in that, The step of generating a target video based on the first video set, the second video set, the viewpoint description information, and the user's historical video narration language features includes: A first target video set is determined based on the first video set; A second target video set is determined based on the second video set; Based on the timing information of the videos to be processed, the first target video set and the second target video set are processed to obtain the video set to be synthesized; The target video is generated based on the set of videos to be synthesized, the viewpoint description information, and the linguistic features of the historical video narration.

4. The method according to claim 3, characterized in that, The step of determining the first target video set based on the first video set includes: Determine the relationship between the first total duration of the videos in the first video set and the first target duration; If the first total duration is greater than the first target duration, the first target video set is determined from the first video set based on the weight of each video in the first video set; If the first total duration is less than or equal to the first target duration, the first video set is determined to be the first target video set. Accordingly, determining the second target video set based on the second video set includes: Calculate the duration difference between the first target duration and the second target duration, and determine the relationship between the second total duration of the videos in the second video set and the duration difference; If the second total duration is greater than the duration difference, the second target video set is determined from the second video set based on the weight of each video in the second video set; If the second total duration is less than or equal to the duration difference, the second video set is determined to be the second target video set.

5. The method according to claim 3, characterized in that, The process of generating the target video based on the set of videos to be synthesized, the viewpoint description information, and the linguistic features of the historical video narration includes: Determine the synthesis order of each video in the set of videos to be synthesized; The synthesis order, the viewpoint description information, and the linguistic features of the historical video narration are processed using a target text model to generate narration text for each video in the video set to be synthesized. The narration of each video is processed to obtain multiple narration audios; The target video is generated based on the multiple audio narrations and the set of videos to be synthesized.

6. The method according to claim 1, characterized in that, The method further includes: Acquire a sample video and determine the first audio data of the sample video; The first audio data is segmented to obtain multiple sample sub-audio data; Determine the first sample viewpoint description information for each sample sub-audio data, and extract features from the first sample viewpoint description information to obtain the first sample feature information; Feature extraction is performed on the multiple sample sub-audio data to obtain the second sample feature information; Based on the first sample feature information and the second sample feature information, the initial original audio video extraction model is trained to obtain the target original audio video extraction model.

7. The method according to claim 1, characterized in that, The method further includes: Acquire sample videos and segment the sample videos to obtain multiple sample sub-videos; Determine the keyframe images and the second audio data of each sample sub-video; Determine the second sample viewpoint description information for each of the second audio data, and perform feature extraction on multiple second sample viewpoint description information to obtain third sample feature information; Feature extraction is performed on multiple keyframe images to obtain fourth sample feature information; The initial video extraction model is trained based on the feature information of the third sample and the feature information of the fourth sample to obtain the target video extraction model.

8. A video generation device, characterized in that, The device includes: a processor and a memory for storing computer programs capable of running on the processor; The processor is used to run the computer program to implement the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs that can be executed by one or more processors to implement the steps of the method as described in any one of claims 1 to 7.

10. A computer program product, the computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.