Short video automatic generation method and device based on new media content creation model

By adopting a short video automatic generation method based on the new media content creation model on the video production platform, and using pre-tuned large language model and video generation model, the existing templates lack personalization and single emotional expression are solved, and high-quality, personalized and flexible short video generation is achieved.

CN119865672BActive Publication Date: 2025-06-13HANGZHOU KNOWLEDGE MATRIX INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510330931.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-13
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The templates provided by the existing video production platform lack personalization and artistic quality, and it is difficult to meet the personalized needs of users' actual scenes. At the same time, the video is relatively single in emotional expression, which makes it difficult for the generated short videos to show rich sense of layering and emotional tension, reducing the quality of short videos.

Method used

The short video automatic generation method based on the new media content creation model is adopted. Through the pre-adjusted large language model and video generation model, a more accurate and professional video production plan is generated, and the temporary video materials and system materials uploaded by users are combined to adaptively generate target short videos that meet the actual user's scenarios.

Benefits of technology

The quality of short videos generated is improved, making them more visually and auditoryly more attractive, meeting users' personalized needs, and improving the flexibility of generating short videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119865672B_ABST
    Figure CN119865672B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for automatically generating short videos based on a new media content creation model. The server-side method includes: receiving a short video generation request sent by a client, where the short video generation request carries a video production description text and a user identifier; inputting the video production description text into a pre-fine-tuned large language model to output video production plan information; searching for temporary video materials uploaded by the user and system materials from the BI system through the user identifier; in the case where the temporary video materials are not empty, inputting the video production plan information, the temporary video materials, and the system materials into a video generation model; or, in the case where the temporary video materials are empty, inputting the video production plan information and the system materials into the video generation model; outputting a target short video and sending the target short video to the client. Therefore, by adopting the embodiments of the present application, the generated short videos can show rich levels and emotional tension, thereby improving the quality of short videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of new media content creation, and particularly to a method and device for automatically generating short videos based on a new media content creation model. Background Art

[0002] With the rapid development of self-media platforms, short videos have become an important way of information dissemination. Users share their lives, convey knowledge, and conduct marketing promotions through short videos, and the demand for short video creation is increasing day by day. However, most users lack professional photography and video production knowledge, resulting in obvious deficiencies in the picture composition, color matching, and shot switching of the generated videos, affecting the visual effect and viewing experience of the videos.

[0003] In related technologies, some video production platforms provide simple templates to help users quickly generate videos. The templates provided by existing video production platforms are relatively general, lacking personalization and artistry, and it is difficult to meet the personalized needs of users in actual scenarios. At the same time, the emotional expression of videos is often relatively single, making it difficult for the generated short videos to show rich levels and emotional tension, thus reducing the quality of short videos. Summary of the Invention

[0004] Embodiments of this application provide a method and device for automatically generating short videos based on a new media content creation model. To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0005] In a first aspect, embodiments of this application provide a method for automatically generating short videos based on a new media content creation model, which is applied to a server. The method includes:

[0006] When receiving a short video generation request sent by a client, call a preset new media content creation model; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model;

[0007] Input the video production description text into the pre-fine-tuned large language model, and output video production scheme information corresponding to the video production description text. The video production scheme information includes video script framework information, artistic expression method information of the video, and element supplement description information;

[0008] Search for system materials related to the video production description text in the user-uploaded temporary video materials from the BI system through the user identifier;

[0009] When the temporary video material is not empty, input the video production plan information, the temporary video material, and the system material into the video generation model to generate the target short video; or, when the temporary video material is empty, input the video production plan information and the system material into the video generation model to generate the target short video;

[0010] Output the target short video corresponding to the video production description text, and send the target short video to the client.

[0011] In a second aspect, an embodiment of the present application provides a short video automatic generation device based on a new media content creation model. The device includes:

[0012] A model calling module, configured to call a preset new media content creation model when receiving a short video generation request sent by the client; the short video generation request carries video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model;

[0013] A text processing module, configured to input the video production description text into the pre-fine-tuned large language model, and output the video production plan information corresponding to the video production description text. The video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information;

[0014] A material calling module, configured to search for the temporary video material uploaded by the user and the system material related to the video production description text from the BI system through the user identifier;

[0015] A short video generation module, configured to input the video production plan information, the temporary video material, and the system material into the video generation model to generate the target short video when the temporary video material is not empty; or, input the video production plan information and the system material into the video generation model to generate the target short video when the temporary video material is empty;

[0016] A video sending module, configured to output the target short video corresponding to the video production description text, and send the target short video to the client.

[0017] The technical solution provided by the embodiment of the present application may include the following beneficial effects:

[0018] In an embodiment of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. The large language model can better understand the semantics of the video production description text through fine-tuning, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and relevant materials, and uses advanced generation technology to generate high-quality video content, making the generated short video show rich layers and emotional tension, so that the generated short video is more attractive visually and auditorily, thus improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection when the temporary video materials are empty or not empty, and then can adaptively generate the target short video that meets the actual scenario of the user, meeting the personalized needs of the user's actual scenario, thus improving the flexibility of the generated short video.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0021] Figure 1 is a schematic flowchart of a method for automatically generating short videos based on a new media content creation model provided by an embodiment of the present application;

[0022] Figure 2 is a user interface diagram shown on a client provided by an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of the model architecture of a video generation model provided by an embodiment of the present application;

[0024] Figure 4 is an interaction schematic diagram between a server and a client provided by an embodiment of the present application;

[0025] Figure 5 is a schematic flowchart of a model generation method of a pre-fine-tuned large language model provided by an embodiment of the present application;

[0026] Figure 6 is a schematic structural diagram of a short video automatic generation device based on a new media content creation model provided by an embodiment of the present application;

[0027] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following description and the accompanying drawings fully disclose specific embodiments of the present application, enabling those skilled in the art to practice them.

[0029] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0030] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0031] In the description of the present application, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0032] Currently, the creation of video content mainly relies on some video production platforms that provide simple templates to help users quickly generate videos.

[0033] The inventors have realized that the templates provided by existing video production platforms are relatively general, lacking personalization and artistry, and it is difficult to meet the personalized needs of users in actual scenarios. At the same time, the emotional expression of videos is often relatively single, making it difficult for the generated short videos to show rich levels and emotional tension, thus reducing the quality of short videos.

[0034] To solve the above problems, the present application provides a method and device for automatically generating short videos based on a new media content creation model to solve the problems existing in the above related technical problems. In an embodiment of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. The large language model can better understand the semantics of the video production description text through fine-tuning, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and related materials, and uses advanced generation technology to generate high-quality video content, making the generated short video show rich layers and emotional tension, so that the generated short video is more attractive visually and auditorily, thereby improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection when the temporary video material is empty or not empty, and then can adaptively generate a target short video that meets the actual scenario of the user, satisfying the personalized needs of the user's actual scenario, thereby improving the flexibility of the generated short video. The following uses exemplary embodiments for detailed description.

[0035] The following will combine with the attached Figure 1 - attached Figure 5 , and introduce in detail the method for automatically generating short videos based on the new media content creation model provided by the embodiments of the present application. This method can be implemented depending on a computer program and can run on a short video automatic generation device based on the von Neumann architecture and based on the new media content creation model. This computer program can be integrated in an application or run as an independent tool class application.

[0036] Please refer to Figure 1 , which is a schematic flowchart of a method for automatically generating short videos based on the new media content creation model provided by the embodiments of the present application, applied to the server. As Figure 1 shown, the method of the embodiments of the present application may include the following steps:

[0037] S101, when receiving a short video generation request sent by a client, call a preset new media content creation model; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model;

[0038] Among them, the client is the device or software used by the user to send requests and receive the generated videos. For example, a mobile application, a web interface, etc. The short video generation request is a request initiated by the user to generate a short video. The request contains necessary information such as the video production description text and the user identification. The preset new media content creation model is a pre-configured model system for generating short videos, including two main parts: a pre-fine-tuned large language model and a video generation model. The user identification is the information used to identify the user's identity, such as the user ID or account. The pre-fine-tuned large language model is a fine-tuned large language model used to understand the video production description text and generate video production plan information. The video generation model is an algorithm model used to generate the final video content based on the video production plan information and materials.

[0039] In some embodiments of the present application, first, the user uploads temporary video materials in the BI system, then the user inputs the video production description text in the client (such as a mobile application) and clicks submit. The server receives the short video generation request sent by the client, the server parses the request content, extracts the video production description text and the user identification, and finally the server calls the preset new media content creation model, which includes a pre-fine-tuned large language model and a video generation model.

[0040] For example Figure 2 As shown, the user hopes to send a short video generation request through the client to generate a short video with the theme of "Li Shangyin's 'A Reply to My Wife in the Rainy Night'". The user provides detailed video production description text and uploads some temporary video materials (such as rainy scenes, candlelight, etc.) in the BI system. The video production description text is, for example: "Please make a video of about 30 seconds based on the poem 'A Reply to My Wife in the Rainy Night' by Li Shangyin."

[0041] S102, input the video production description text into the pre-fine-tuned large language model, and output the video production plan information corresponding to the video production description text. The video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information;

[0042] Among them, the video production plan information is the detailed video production plan and guiding information generated by the pre-fine-tuned large language model according to the video production description text. The video script framework information is the specific structure and content arrangement of the video, describing the specific scenes and contents of each part such as the beginning, middle, and end of the video. The video artistic expression method information is various techniques and methods used to enhance the artistic effect and emotional expression of the video. The element supplementary description information is the detailed description and design requirements for specific elements in the video, used to further refine the visual and auditory effects of the video.

[0043] In some embodiments of the present application, after receiving the video production description text, the server inputs it into a pre-fine-tuned large language model, and the pre-fine-tuned large language model generates detailed video production plan information based on the input video production description text.

[0044] For example, the video production description text is: "Please make a 30-second video based on Li Shangyin's poem "Night Rain Sent to the North". After processing by the pre-fine-tuned large language model, the video production plan information can be viewed in the server's log information, as shown below:

[0045] The log information in JSON format is as follows:

[0046] script framework:{

[0047] opening: The camera opens with a close-up of raindrops sliding down bamboo leaves in the dark night, showing the panoramic view of Bashan Mountain in autumn night, with rain curtains covering the mountains and a silhouette of a single lamp swaying in front of the window. The subtitles gradually appear, "You ask when I will return, but there is no date."

[0048] middle: Dynamic transition: Raindrops fall into the pond, ripples spread, and then turn into a close-up of "the water in the autumn pond rises". Contrast between reality and illusion: The camera switches to the poet's profile staring at the candlelight, the picture gradually blurs, and an imaginary picture emerges - under the west window, the candlelight is warm, and the poet and his relatives and friends are talking by candlelight. "

[0049] ending: The camera pulls back to reality, the sound of rain fades, the candlelight goes out, the morning light is faintly visible outside the window, and the subtitles freeze "Let's talk about the night rain in Bashan Mountain".}

[0050] artistic techniques:

[0051] audio:"The background sound is a zither and flute ensemble, with the sound of rain gradually weakening, reinforcing the sense of loneliness and longing. The reading adopts the rhythm of "meaning division" (such as "You ask / when to return / there is no date"), with a slow tone and a slightly rising last sentence, reflecting the tension of time and space.

[0052] Visual symbols: Rain: symbolizes continuous sadness, and conveys emotions through details such as close-ups of raindrops and water ripples. Candlelight: The lonely lamp in reality and the imagined "cut candle" form a cold and warm contrast, metaphorically implying the conflict between hope and reality. Letter: Close-up of the word "return date" written by a brush, with ink smudged, implying an unfulfilled promise.}

[0053] element supplements:

[0054] subtitle_design: "Using running script font, the poem lines appear one by one, matching the rhythm of the picture, and the last line stays for 3 seconds."

[0055] color_contrast: "The real part is mainly cyan-gray, and the imagined scene turns into warm yellow to highlight the emotional contrast."

[0056] Specifically, the video script framework information included in the video production plan information includes: At the beginning (0 - 5 seconds): The shot starts with a close-up of raindrops sliding down bamboo leaves in the dark night, showing the panoramic view of the autumn night in Bashan. The rain curtain shrouds the mountains, and there is a silhouette of a solitary lamp swaying in front of the window. The subtitle gradually appears: "You ask when I'll return, but there's no fixed date." In the middle section (6 - 20 seconds): Dynamic transition: The raindrops fall into the pond, and the ripples spread, dissolving into a close-up of "Autumn pond water rising". Virtual-real contrast: The camera switches to a side view of the poet staring at the candlelight, and the picture gradually fades. An imagined picture emerges - under the west window, the candlelight is warm, and the poet and relatives and friends are chatting while trimming the candle. At the end (21 - 30 seconds): The camera pulls back to reality, the sound of the rain gradually weakens, the candle goes out, and the morning light faintly appears outside the window. The subtitle freezes: "Then talk about the night rain in Bashan."

[0057] The artistic expression techniques information of the video included in the video production plan information includes: Sound effects and music: The background sound selects the ensemble of guzheng and xiao. The sound of the rain gradually weakens from strong to weak, strengthening the layering of loneliness and yearning. The reading adopts the "meaning division" rhythm (such as "You ask / when I'll return / but there's no fixed date"), with a slow and deep tone, and the end sentence slightly rising, reflecting the tension of the interlacing of time and space. Visual symbols: Rain, symbolizing continuous melancholy, conveys emotions through details such as close-ups of raindrops and water ripples. Candlelight, the contrast between the solitary lamp in reality and the "trimming the candle" in imagination forms a contrast between cold and warm, metaphorically suggesting the conflict between hope and reality. Letter, a close-up of writing the two characters "return date" with a writing brush, and the ink smudges, hinting at a promise that cannot be fulfilled.

[0058] The supplementary element description information included in the video production plan information includes: Subtitle design: The running script font is adopted, and the verses appear one by one, matching the rhythm of the picture. The last sentence stays for 3 seconds. Color contrast: The real part is mainly cyan-gray, and the imagined scene turns into warm yellow to highlight the emotional contrast.

[0059] In some embodiments of the present application, the specific process of generating a pre-fine-tuned large language model includes: obtaining the video creation requirements of the short video generation task and a preset video production plan knowledge base. The video creation requirements are used to represent the basic requirements for creating a short video. The preset video production plan knowledge base is constructed based on a standard data set related to short video creation. The standard data set related to short video creation includes video scripts, scene description data, and emotional expression data; using the pre-trained preset large language model as the base model; modifying the base model according to the video creation requirements of the short video generation task to determine the target parameters suitable for the video creation requirements; fine-tuning and optimizing the base model through the sample data and target parameters in the video production plan knowledge base to obtain a pre-fine-tuned large language model.

[0060] In the embodiments of the present application, by obtaining the video creation requirements of the short video generation task and the preset video production plan knowledge base, and making targeted modifications and fine-tuning to the pre-trained large language model based on this data, the performance of the model in the short video creation task can be significantly improved. This fine-tuning process enables the model to better understand the specific requirements of video creation, generate video scripts and production plans that better meet user needs, thereby improving the creation quality and efficiency of short videos.

[0061] Specifically, the specific process of fine-tuning and optimizing the basic model through the sample data and target parameters in the video production plan knowledge base to obtain the pre-fine-tuned large language model includes: dividing the sample data in the video production plan knowledge base according to a preset ratio to obtain a first training set and a second training set; decomposing the target parameters to obtain a first parameter decomposition matrix and a second parameter decomposition matrix, where the first parameter decomposition matrix includes parameters related to the columns of the original model parameter matrix in the basic model, and the second parameter decomposition matrix includes parameters related to the rows of the original model parameter matrix in the basic model; calculating the backpropagation result of the basic model according to each training data, the first parameter decomposition matrix, and the second parameter decomposition matrix in the first training set; calculating the model loss value of the basic model according to the backpropagation result and the preset target loss function; obtaining the fine-tuned basic model when the model loss value reaches the minimum; and performing secondary optimization on the fine-tuned basic model using the second training set to obtain the pre-fine-tuned large language model.

[0062] Specifically, the calculation formula for the backpropagation result is:

[0063]

[0064] Where is the backpropagation result of the th training data, is the original parameter matrix of the basic model, is the bias, is the first parameter decomposition matrix, is the second parameter decomposition matrix, satisfying , is the th training data;

[0065] The preset target loss function is:

[0066]

[0067] Where is the model loss value of the basic model, is the number of samples of the training data, is the number of categories, is the The true label of a training data in a category The true label is 0 or 1, is the th training data's predicted probability in a category is the weight of the regularization term, is the result of backpropagation, is the square of the norm of the gradient of the model parameters after the basic model incorporates the result of backpropagation Specifically, adopting the second training set, the specific process of secondary optimization of the fine-tuned basic model includes: inputting each training data in the second training set into the fine-tuned basic model, and outputting the video production plan information corresponding to each training data; analyzing the performance quantization value of the fine-tuned basic model according to the video production plan information; determining the evaluation value according to the performance quantization value in combination with a preset evaluation function; in the case that the evaluation value does not reach the preset expected value, updating the parameters of the fine-tuned basic model to maximize the expectation of the evaluation value determined by the preset evaluation function.

[0068] Among them, the update can use the policy gradient method in reinforcement learning, and the evaluation function is:

[0069] where is the evaluation value, is a constant used to adjust the range of the evaluation value. The performance quantization value can be determined by analyzing the difference between the video production plan information corresponding to each training data and the label of each training data. S103, search for the system materials related to the temporary video materials uploaded by the user and the video production description text from the BI system through the user identifier;

[0070] Among them, the BI system is the Business Intelligence system, which is used to store and manage the video materials and system materials uploaded by the user. For example, an enterprise-level multimedia material management system, where users can upload and manage their own video materials. The temporary video materials are the temporary video materials uploaded by the user, usually used for specific video generation tasks and can be deleted after the task is completed. The system materials are a preset video material library, which contains various general video clips, background music, fonts, etc., and are used to assist video generation.

[0071]

[0072] ​In some embodiments of the present application, after obtaining the video production plan information output by the pre-fine-tuned large language model, the server searches for the temporary video materials uploaded by the user from the BI system through the user identifier (user identifier user123), then obtains the semantics of the video production description text, and searches for system materials in the preset material library of the BI system whose semantic similarity to the video production description text is greater than the threshold, so as to obtain system materials related to the video production description text.

[0073] S104, in the case where the temporary video material is not empty, input the video production plan information, the temporary video material, and the system material into the video generation model to generate the target short video; or, in the case where the temporary video material is empty, input the video production plan information and the system material into the video generation model to generate the target short video;

[0074] Among them, the video script framework information includes the description of the video start scene, the description of the video middle scene, and the description of the video end scene.

[0075] For example Figure 3 As shown, the video generation model includes a semantic vector extraction module, a material feature extraction module, a material matching module, a video segment generation module, a video splicing module, and a video optimization module.

[0076] Among them, the fact that the temporary video material is not empty indicates that the current user has shot relevant content for generating the short video, and the shot content may be shot by a professional photographer or an ordinary user.

[0077] In some embodiments of the present application, the specific process of generating the target short video includes: the semantic vector extraction module uses natural language processing (NLP) technology to parse the video start scene description, the video middle scene description, and the video end scene description to extract semantic vectors for key scene descriptions; the material feature extraction module, when the temporary video material is not empty, classifies and labels the image frames of the temporary video material and the system material to extract the first material features of the temporary video material and the second material features of the system material; the first material features and the second material features of the system material are composed of visual features and audio features; the material matching module performs material matching on the first material features and the second material features respectively according to the semantic vectors to obtain a first matching result and a second matching result; the first matching result has first target features corresponding to some or all of the semantic vectors, and the second matching result has second target features corresponding to all of the semantic vectors; the video segment generation module generates multiple video segments suitable for the video start scene description, the video middle scene description, and the video end scene description based on the first matching result and the second matching result; the video splicing module splices the multiple video segments based on the artistic expression method information to obtain the initial short video; the artistic expression method information includes sound effect and music information, visual symbol information, and editing technique information; the video optimization module designs subtitles and performs color tone conversion on the initial short video based on the element supplementary description information to obtain the target short video; the element supplementary description information includes subtitle design and color tone contrast information.

[0078] It should be noted that the first matching result having first target features corresponding to some of the semantic vectors indicates that the current user is an ordinary user shooting, that is, some contents in the temporary video material meet the requirements and some do not. The first matching result having first target features corresponding to all of the semantic vectors indicates that the current user is a professional photographer, that is, all contents in the temporary video material meet the requirements. The reason why the second matching result has second target features corresponding to all of the semantic vectors is that the system materials are all standardized and thus all meet the requirements.

[0079] Among them, the semantic vectors include the video start scene semantic vector, the video middle scene semantic vector, and the video end scene semantic vector.

[0080] In some embodiments of the present application, the specific process of generating multiple video segments applicable to the video start scene description, the video middle segment scene description, and the video end scene description based on the first matching result and the second matching result includes: when there is a first target feature corresponding to all semantic vectors in the first matching result, obtaining the material content corresponding to the first target feature corresponding to all semantic vectors from the temporary video material; or, when there is a first target feature corresponding to some of the semantic vectors in the first matching result, obtaining the first material content to be analyzed corresponding to the first target feature corresponding to some of the semantic vectors from the temporary video material; determining the semantic vectors to be analyzed for which the first material features are not matched; determining, from the second target features corresponding to all semantic vectors existing in the second matching result, the second target feature corresponding to the semantic vectors to be analyzed; obtaining the second material content to be analyzed corresponding to the second target feature corresponding to the semantic vectors to be analyzed from the system material; constructing the material content corresponding to all semantic vectors according to the first material content to be analyzed and the second material content to be analyzed; classifying the above material content to obtain the material content corresponding to the video start scene semantic vector, the material content corresponding to the video middle segment scene semantic vector, and the material content corresponding to the video end scene semantic vector; using the material content corresponding to the video start scene semantic vector as the video segment applicable to the video start scene description; using the material content corresponding to the video middle segment scene semantic vector as the video segment applicable to the video middle segment scene description; and using the material content corresponding to the video end scene semantic vector as the video segment applicable to the video end scene description.

[0081] In the embodiments of the present application, when the temporary video material is not empty, short video generation needs to be mainly based on the style of the temporary video material uploaded by the user. Therefore, when there is a first target feature corresponding to all semantic vectors in the first matching result, all the content in the temporary video material meets the requirements, and the temporary video material needs to be used to produce the short video. When there is a first target feature corresponding to some of the semantic vectors in the first matching result, some of the content in the temporary video material meets the requirements. At this time, it is necessary to obtain the system material corresponding to the temporary material that does not meet the requirements from the system material, and then transform the system material corresponding to the temporary material that does not meet the requirements based on the style of the temporary material content that meets the requirements, so that the system material that meets the requirements becomes the material applicable to the user style.

[0082] Specifically, the specific process of constructing the material content corresponding to all semantic vectors according to the content of the first material to be analyzed and the content of the second material to be analyzed includes: analyzing the image features of the first material to be analyzed; using the image features as the algorithm parameters of a preset generation algorithm, and adjusting the image style of the second material to be analyzed through the preset generation algorithm to obtain the optimized second material to be analyzed; the generation algorithm is a diffusion model; combining the first material to be analyzed and the optimized second material to be analyzed into the material content corresponding to all semantic vectors.

[0083] In the embodiment of the present application, when the temporary video material is not empty, the system material is transformed based on the style of the temporary video material uploaded by the user to make it conform to the current user, and a target short video that conforms to the actual scenario of the user is generated, meeting the personalized needs of the actual scenario of the user.

[0084] In some other embodiments of the present application, the specific process of generating the target short video includes: when the temporary video material is empty, classifying and labeling the image frames of the system material to extract the second material features of the system material; performing material matching on the second material features according to the semantic vectors to obtain multiple video segments applicable to the description of the video start scene, the description of the middle video scene, and the description of the video end scene; performing video splicing on the multiple video segments based on the art expression method information to obtain the initial short video; performing subtitle design and color tone conversion on the initial short video based on the element supplementary description information to obtain the target short video.

[0085] It should be noted that when the temporary video material is empty, it means that the user has not uploaded any material, and the short video is generated only based on text. At this time, there is no need to rely on the user's video style, and there is no need to transform the system material at all. Just directly match multiple video segments applicable to the description of the video start scene, the description of the middle video scene, and the description of the video end scene based on the image frames of the system material.

[0086] In the embodiment of the present application, the specific process of performing video splicing on multiple video segments based on the art expression method information to obtain the initial short video is: processing each video segment based on the sound effect and music information, visual symbol information, and editing technique information, and splicing the processed video segments to obtain the initial short video. For example, in the video generation based on "A Letter to My Wife in the Rainy Night", the sound effect and music information: the background music selects the ensemble of guzheng and xiao, and the sound of rain gradually weakens from strong. Visual symbol information: rain symbolizes continuous melancholy, and the candlelight forms a contrast between warm and cold. Editing technique information: uses "dissolve" to connect the real and imaginary scenes, and alternates between fast and slow shots.

[0087] In the embodiments of the present application, during the process of performing subtitle design and color tone conversion on the initial short video based on the element supplementary description information to obtain the target short video, the element supplementary description information includes subtitle design and color tone contrast information. For example, in the generation of the video of "A Letter to My Wife in the Rainy Night", subtitle design: The running script font is adopted, and the lines of the poem appear one by one, and the last line stays for 3 seconds. Color tone contrast: The real part is mainly cyan-gray, and the imagined scene turns into warm yellow.

[0088] S105, output the target short video corresponding to the video production description text, and send the target short video to the client.

[0089] In the embodiments of the present application, after obtaining the target short video, the generated video can be sent back to the client through the backend service, and the user can view and download the generated short video on the front-end interface. The front-end page will display a video player, and the user can play the generated short video. At this time, the video content conforms to the theme of "A Letter to My Wife in the Rainy Night" and has rich layers and emotional tension.

[0090] For example Figure 4 As shown, the server 110 can be a server, which can specifically be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. For example, it is a server device running a new media content creation model and a BI system. When short video automatic generation is required, when the server 110 receives a short video generation request sent by the client 120, it calls a preset new media content creation model; the short video generation request carries the video production description text and the user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model; input the video production description text into the pre-fine-tuned large language model, and output the video production plan information corresponding to the video production description text, where the video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information; through the user identifier, search the system materials related to the video production description text from the user-uploaded temporary video materials in the BI system; in the case where the temporary video materials are not empty, input the video production plan information, the temporary video materials, and the system materials into the video generation model to generate the target short video; or, in the case where the temporary video materials are empty, input the video production plan information and the system materials into the video generation model to generate the target short video; output the target short video corresponding to the video production description text, and send the target short video to the client 120.

[0091] It should be noted that the client 120 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server 110 and the client 120 can be connected through Bluetooth, USB (Universal Serial Bus), or other communication connection methods, and the present invention does not make any restrictions here.

[0092] In an embodiment of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. The large language model can better understand the semantics of the video production description text through fine-tuning, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and relevant materials, and uses advanced generation technology to generate high-quality video content, making the generated short video show rich levels and emotional tension, so that the generated short video is more attractive visually and auditorily, thereby improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection when the temporary video materials are empty or not empty, and then can adaptively generate a target short video that meets the actual scenario of the user, meeting the personalized needs of the user's actual scenario, thereby improving the flexibility of the generated short video.

[0093] Please refer to Figure 5 , which is a schematic flowchart of a model generation method for a pre-fine-tuned large language model provided by an embodiment of the present application. As Figure 5 shown, the method of the embodiment of the present application may include the following steps:

[0094] S201, obtain the video creation requirements of the short video generation task and the preset video production plan knowledge base. The video creation requirements are used to represent the basic requirements for creating a short video. The preset video production plan knowledge base is constructed based on a standard data set related to short video creation. The standard data set related to short video creation includes video scripts, scene description data, and emotional expression data;

[0095] S202, use the pre-trained preset large language model as the basic model;

[0096] S203, modify the basic model according to the video creation requirements of the short video generation task to determine the target parameters applicable to the video creation requirements;

[0097] S204, fine-tune and optimize the basic model through the sample data and target parameters in the video production plan knowledge base to obtain a pre-fine-tuned large language model.

[0098] In the embodiments of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. Through fine-tuning, the large language model can better understand the semantics of the video production description text, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and relevant materials, and uses advanced generation technologies to generate high-quality video content, making the generated short video show rich levels and emotional tension, so that the generated short video is more attractive visually and aurally, thereby improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection when the temporary video materials are empty or not empty, and then can adaptively generate a target short video that meets the actual scenario of the user, meeting the personalized needs of the user's actual scenario, thereby improving the flexibility of the generated short video.

[0099] The following are the device embodiments of the present application, which can be used to execute the method embodiments of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0100] Please refer to Figure 6 , which shows a schematic structural diagram of a short video automatic generation device based on a new media content creation model provided by an exemplary embodiment of the present application. The short video automatic generation device based on the new media content creation model can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a model calling module 10, a text processing module 20, a material calling module 30, a short video generation module 40, and a video sending module 50.

[0101] The model calling module 10 is configured to call a preset new media content creation model when receiving a short video generation request sent by a client; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model;

[0102] The text processing module 20 is configured to input the video production description text into the pre-fine-tuned large language model, and output video production plan information corresponding to the video production description text. The video production plan information includes video script framework information, artistic expression method information of the video, and element supplementary description information;

[0103] The material calling module 30 is configured to search for system materials related to the video production description text of the temporary video materials uploaded by the user from the BI system through the user identifier;

[0104] The short video generation module 40 is used to input video production plan information, temporary video materials, and system materials into a video generation model to generate a target short video when the temporary video materials are not empty; or, when the temporary video materials are empty, input video production plan information and system materials into the video generation model to generate a target short video.

[0105] The video sending module 50 is used to output the target short video corresponding to the video production description text and send the target short video to the client.

[0106] It should be noted that when the short video automatic generation device based on the new media content creation model provided in the above embodiments executes the short video automatic generation method based on the new media content creation model, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the short video automatic generation device based on the new media content creation model provided in the above embodiments and the embodiments of the short video automatic generation method based on the new media content creation model belong to the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.

[0107] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0108] In the embodiments of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. The large language model can better understand the semantics of the video production description text through fine-tuning, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and related materials, and uses advanced generation technologies to generate high-quality video content, making the generated short video show rich levels and emotional tension, so that the generated short video is more attractive visually and auditorily, thus improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection when the temporary video materials are empty or not empty, and then can adaptively generate a target short video that meets the actual scenario of the user, meeting the personalized needs of the user's actual scenario, thus improving the flexibility of the generated short video.

[0109] The present application also provides a computer-readable medium, on which program instructions are stored, and when the program instructions are executed by a processor, the short video automatic generation method based on the new media content creation model provided in each of the above method embodiments is implemented.

[0110] The present application also provides a computer program product containing instructions, which, when running on a computer, enables the computer to execute the short video automatic generation method based on the new media content creation model in each of the above method embodiments.

[0111] Please refer to Figure 7 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 7 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.

[0112] Among them, the communication bus 1002 is used to realize the connection and communication between these components.

[0113] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.

[0114] Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0115] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire electronic device 1000 through various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005, it executes various functions of the electronic device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 1001 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.

[0116] Among them, the memory 1005 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 1005 may also be at least one storage system located far from the aforementioned processor 1001. As Figure 7 shown, in the memory 1005 as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and a short video automatic generation application program based on a new media content creation model.

[0117] In Figure 7 the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user to obtain the data input by the user; and the processor 1001 can be used to call the short video automatic generation application program stored in the memory 1005 and specifically perform the following operations:

[0118] When receiving a short video generation request sent by the client, call a preset new media content creation model; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model;

[0119] Input the video production description text into the pre-fine-tuned large language model, and output the video production plan information corresponding to the video production description text. The video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information;

[0120] Through the user identifier, search in the BI system for the temporary video materials uploaded by the user and the system materials related to the video production description text;

[0121] In the case where the temporary video materials are not empty, input the video production plan information, the temporary video materials, and the system materials into the video generation model to generate a target short video; or, in the case where the temporary video materials are empty, input the video production plan information and the system materials into the video generation model to generate a target short video;

[0122] Output the target short video corresponding to the video production description text, and send the target short video to the client.

[0123] In one embodiment, when the processor 1001 executes to generate the target short video, the following operations are specifically performed:

[0124] The semantic vector extraction module uses natural language processing (NLP) technology to parse the video start scene description, the video middle scene description, and the video end scene description to extract semantic vectors for key scene descriptions;

[0125] When the temporary video material is not empty, the material feature extraction module classifies and labels the image frames of the temporary video material and the system material to extract the first material features of the temporary video material and the second material features of the system material; the first material features and the second material features of the system material are composed of visual features and audio features;

[0126] The material matching module performs material matching on the first material features and the second material features respectively according to the semantic vectors to obtain a first matching result and a second matching result; the first matching result has first target features corresponding to some or all of the semantic vectors, and the second matching result has second target features corresponding to all of the semantic vectors;

[0127] The video segment generation module generates multiple video segments applicable to the video start scene description, the video middle scene description, and the video end scene description based on the first matching result and the second matching result;

[0128] The video splicing module splices the multiple video segments based on the artistic expression method information to obtain an initial short video; the artistic expression method information includes sound effect and music information, visual symbol information, and editing technique information;

[0129] The video optimization module designs subtitles and performs color tone conversion on the initial short video based on the element supplementary description information to obtain the target short video; the element supplementary description information includes subtitle design and color tone contrast information.

[0130] In one embodiment, when the processor 1001 executes to generate multiple video segments applicable to the video start scene description, the video middle scene description, and the video end scene description based on the first matching result and the second matching result, the following operations are specifically performed:

[0131] When the first matching result has first target features corresponding to all of the semantic vectors, obtain the material content corresponding to the first target features corresponding to all of the semantic vectors from the temporary video material;

[0132] Alternatively, when the first matching result has a first target feature corresponding to the semantic vector of the part, obtain the first material content to be analyzed corresponding to the first target feature corresponding to the semantic vector of the part from the temporary video material; determine the semantic vector to be analyzed for which the first material feature is not matched; from the second target features corresponding to all the semantic vectors existing in the second matching result, determine the second target feature corresponding to the semantic vector to be analyzed; obtain the second material content to be analyzed corresponding to the second target feature corresponding to the semantic vector to be analyzed from the system material; construct the material content corresponding to all the semantic vectors according to the first material content to be analyzed and the second material content to be analyzed;

[0133] Classify the above material content to obtain the material content corresponding to the semantic vector of the video start scene, the material content corresponding to the semantic vector of the middle section of the video, and the material content corresponding to the semantic vector of the end scene of the video;

[0134] Use the material content corresponding to the semantic vector of the video start scene as the video segment applicable to the description of the video start scene; use the material content corresponding to the semantic vector of the middle section of the video as the video segment applicable to the description of the middle section of the video; use the material content corresponding to the semantic vector of the end scene of the video as the video segment applicable to the description of the end scene of the video.

[0135] In one embodiment, when the processor 1001 executes to construct the material content corresponding to all the semantic vectors according to the first material content to be analyzed and the second material content to be analyzed, the following operations are specifically executed:

[0136] Analyze the image features of the first material content to be analyzed;

[0137] Use the image features as the algorithm parameters of the preset generation algorithm, and perform image style adjustment on the second material content to be analyzed through the preset generation algorithm to obtain the optimized second material content to be analyzed; the generation algorithm is a diffusion model;

[0138] Merge the first material content to be analyzed and the optimized second material content to be analyzed into the material content corresponding to all the semantic vectors.

[0139] In one embodiment, when the processor 1001 executes to generate the target short video, the following operations are specifically executed:

[0140] In the case where the temporary video material is empty, classify and label the image frames of the system material to extract the second material features of the system material;

[0141] According to the semantic vectors, perform material matching on the second material features to obtain multiple video segments applicable to the description of the video start scene, the middle section of the video, and the end scene of the video;

[0142] Based on the artistic expression information, splice multiple video clips to obtain an initial short video;

[0143] Based on the element supplementary description information, perform subtitle design and color tone conversion on the initial short video to obtain the target short video.

[0144] In one embodiment, when the processor 1001 executes to generate a pre-fine-tuned large language model, the following operations are specifically performed:

[0145] Obtain the video creation requirements of the short video generation task and the preset video production plan knowledge base. The video creation requirements are used to represent the basic requirements for creating a short video. The preset video production plan knowledge base is constructed based on a standard data set related to short video creation. The standard data set related to short video creation includes video scripts, scene description data, and emotional expression data;

[0146] Use the pre-trained preset large language model as the base model;

[0147] Modify the base model according to the video creation requirements of the short video generation task to determine the target parameters applicable to the video creation requirements;

[0148] Fine-tune and optimize the base model through the sample data and target parameters in the video production plan knowledge base to obtain a pre-fine-tuned large language model.

[0149] In one embodiment, when the processor 1001 executes to fine-tune and optimize the base model through the sample data and target parameters in the video production plan knowledge base to obtain a pre-fine-tuned large language model, the following operations are specifically performed:

[0150] Divide the sample data in the video production plan knowledge base according to a preset ratio to obtain a first training set and a second training set;

[0151] Decompose the target parameters to obtain a first parameter decomposition matrix and a second parameter decomposition matrix. The first parameter decomposition matrix includes parameters related to the columns of the original model parameter matrix in the base model, and the second parameter decomposition matrix includes parameters related to the rows of the original model parameter matrix in the base model;

[0152] Calculate the backpropagation result of the base model according to each training data, the first parameter decomposition matrix, and the second parameter decomposition matrix in the first training set;

[0153] Calculate the model loss value of the base model according to the backpropagation result and the preset target loss function;

[0154] Obtain the fine-tuned base model when the model loss value reaches the minimum;

[0155] Using the second training set, the fine-tuned base model is further optimized to obtain a pre-fine-tuned large language model.

[0156] In one embodiment, when the processor 1001 executes the operation of further optimizing the fine-tuned base model using the second training set, the following specific operations are performed:

[0157] Input each training set in the second training set into the fine-tuned base model, and output the video production plan information corresponding to each training set;

[0158] According to the video production plan information, analyze the performance quantization value of the fine-tuned base model;

[0159] According to the performance quantization value, combine a preset evaluation function to determine an evaluation value;

[0160] In the case where the evaluation value does not reach the preset expected value, update the parameters of the fine-tuned base model to maximize the expectation of the evaluation value determined by the preset evaluation function.

[0161] In the embodiments of the present application, on the one hand, in the short video generation process, a pre-fine-tuned large language model and a video generation model are used. The large language model can better understand the semantics of the video production description text through fine-tuning, so as to generate a more accurate and professional video production plan. The video generation model is based on the video production plan and related materials, and uses advanced generation technologies to generate high-quality video content, making the generated short video show rich layers and emotional tension, so that the generated short video is more attractive visually and auditorily, thus improving the quality of the short video. On the other hand, the video generation model can make an adaptive selection in the case where the temporary video material is empty or not empty, and can thus adaptively generate a target short video that meets the actual scenario of the user, meeting the personalized needs of the user's actual scenario, thereby improving the flexibility of the generated short video.

[0162] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above embodiment methods can be completed by instructing relevant hardware through a computer program. The program for automatically generating short videos based on the new media content creation model can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium of the program for automatically generating short videos based on the new media content creation model can be a magnetic disk, an optical disc, a read-only memory, or a random access memory, etc.

[0163] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A short video automatic generation method based on a new media content creation model, characterized in that: Applied to the server, the method includes: Upon receiving a short video generation request sent by a client, a preset new media content creation model is called; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model; the video generation model includes a semantic vector extraction module, a material feature extraction module, a material matching module, a video clip generation module, a video splicing module, and a video optimization module; Input the video production description text into a pre-fine-tuned large language model, and output the video production plan information corresponding to the video production description text, wherein the video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information; the video script framework information is used to describe the specific scenes and contents of each part of the beginning, middle, and end of the video; the artistic expression method information includes sound effect and music information, visual symbol information, and editing technique information; the element supplementary description information includes subtitle design and color contrast information; By using the user ID, searching the BI system for system materials related to the temporary video materials uploaded by the user and the video production description text; When the temporary video material is not empty, the video production scheme information, the temporary video material and the system material are input into the video generation model to generate a target short video; or, when the temporary video material is empty, the video production scheme information and the system material are input into the video generation model to generate a target short video; generating the target short video includes: The semantic vector extraction module uses natural language processing (NLP) technology to parse the scene descriptions at the beginning, middle and end of the video to extract semantic vectors for key scene descriptions. The material feature extraction module classifies and annotates the image frames of the temporary video material and the system material when the temporary video material is not empty, so as to extract a first material feature of the temporary video material and a second material feature of the system material; the first material feature and the second material feature of the system material are composed of visual features and audio features; The material matching module performs material matching on the first material feature and the second material feature respectively according to the semantic vector to obtain a first matching result and a second matching result; the first matching result has a first target feature corresponding to part or all of the semantic vectors, and the second matching result has a second target feature corresponding to all of the semantic vectors; The video segment generation module generates a plurality of video segments suitable for the description of the beginning scene of the video, the description of the middle scene of the video, and the description of the ending scene of the video based on the first matching result and the second matching result; The video splicing module splices the multiple video clips based on the artistic expression technique information to obtain an initial short video; The video optimization module performs subtitle design and tone conversion on the initial short video based on the element supplementary description information to obtain a target short video; Output the target short video corresponding to the video production description text, and send the target short video to the client.

2. The method according to claim 1, characterized in that The semantic vectors include a semantic vector of a scene at the beginning of the video, a semantic vector of a scene in the middle of the video, and a semantic vector of a scene at the end of the video; The generating, based on the first matching result and the second matching result, a plurality of video clips applicable to the description of the beginning scene of the video, the description of the middle scene of the video, and the description of the ending scene of the video comprises: When the first matching result has a first target feature corresponding to all of the semantic vectors, acquiring material content corresponding to the first target feature corresponding to all of the semantic vectors from the temporary video material; Alternatively, when the first matching result has a first target feature corresponding to part of the semantic vector, obtain the first material content to be analyzed corresponding to the first target feature corresponding to part of the semantic vector from the temporary video material; determine the semantic vector to be analyzed that is not matched to the first material feature; determine the second target feature corresponding to the semantic vector to be analyzed from the second target features corresponding to all of the semantic vectors that exist in the second matching result; obtain the second material content to be analyzed corresponding to the second target feature corresponding to the semantic vector to be analyzed from the system material; and construct the material content corresponding to all of the semantic vectors based on the first material content to be analyzed and the second material content to be analyzed; Classify the above-mentioned material contents to obtain the material contents corresponding to the semantic vector of the scene at the beginning of the video, the material contents corresponding to the semantic vector of the scene in the middle of the video, and the material contents corresponding to the semantic vector of the scene at the end of the video; The material content corresponding to the semantic vector of the scene at the beginning of the video is used as a video clip suitable for describing the scene at the beginning of the video; the material content corresponding to the semantic vector of the scene in the middle of the video is used as a video clip suitable for describing the scene in the middle of the video; the material content corresponding to the semantic vector of the scene at the end of the video is used as a video clip suitable for describing the scene at the end of the video.

3. The method according to claim 2, characterized in that The constructing material content corresponding to all the semantic vectors according to the first material content to be analyzed and the second material content to be analyzed includes: Analyzing image features of the first material to be analyzed; The image feature is used as an algorithm parameter of a preset generation algorithm, and the image style of the second material content to be analyzed is adjusted by the preset generation algorithm to obtain an optimized second material content to be analyzed; the generation algorithm is a diffusion model; The first material content to be analyzed and the optimized second material content to be analyzed are merged into material content corresponding to all the semantic vectors.

4. The method according to claim 1, characterized in that The generating of the target short video comprises: When the temporary video material is empty, classifying and marking the image frames of the system material to extract the second material feature of the system material; According to the semantic vector, material matching is performed on the second material feature to obtain a plurality of video clips suitable for describing the scene at the beginning of the video, describing the scene in the middle of the video, and describing the scene at the end of the video; Based on the artistic expression technique information, the multiple video clips are spliced ​​to obtain an initial short video; Based on the element supplementary description information, subtitles are designed and the color tone is converted for the initial short video to obtain a target short video.

5. The method according to claim 1, generating a pre-fine-tuned large language model according to the following steps, comprising: Obtaining video creation requirements and a preset video production solution knowledge base for a short video generation task, wherein the video creation requirements are used to characterize the basic requirements for creating a short video, and the preset video production solution knowledge base is constructed based on a standard data set related to short video creation, and the standard data set related to short video creation includes video scripts, scene description data, and emotion expression data; Use the pre-trained preset large language model as the basic model; According to the video creation requirements of the short video generation task, the basic model is modified to determine target parameters suitable for the video creation requirements; The basic model is fine-tuned and optimized through the sample data in the video production solution knowledge base and the target parameters to obtain a pre-fine-tuned large language model.

6. The method according to claim 5, characterized in that The method of fine-tuning and optimizing the basic model by using the sample data in the video production solution knowledge base and the target parameters to obtain a pre-fine-tuned large language model includes: Dividing the sample data in the video production solution knowledge base according to a preset ratio to obtain a first training set and a second training set; Decomposing the target parameters to obtain a first parameter decomposition matrix and a second parameter decomposition matrix, wherein the first parameter decomposition matrix includes parameters related to columns of the original model parameter matrix in the base model, and the second parameter decomposition matrix includes parameters related to rows of the original model parameter matrix in the base model; Calculating a back propagation result of the basic model according to each training data in the first training set, the first parameter decomposition matrix and the second parameter decomposition matrix; Calculating the model loss value of the basic model according to the back propagation result and the preset target loss function; When the loss value of the model reaches a minimum, a fine-tuned basic model is obtained; The second training set is used to perform secondary optimization on the fine-tuned basic model to obtain a pre-fine-tuned large language model.

7. The method according to claim 6, characterized in that The second training set is used to perform secondary optimization on the fine-tuned basic model, including: Inputting each training set in the second training set into the fine-tuned basic model, and outputting video production solution information corresponding to each training set; Analyze the performance quantification value of the fine-tuned basic model according to the video production solution information; Determine an evaluation value based on the performance quantification value and in combination with a preset evaluation function; In the case where the evaluation value does not reach the preset expected value, the parameters of the fine-tuned basic model are updated to maximize the expectation of the evaluation value determined by the preset evaluation function.

8. A short video automatic generation device based on a new media content creation model implemented using the method described in any one of claims 1 to 7, characterized in that: The device comprises: A model calling module is used to call a preset new media content creation model upon receiving a short video generation request sent by a client; the short video generation request carries a video production description text and a user identifier, and the preset new media content creation model includes a pre-fine-tuned large language model and a video generation model; the video generation model includes a semantic vector extraction module, a material feature extraction module, a material matching module, a video clip generation module, a video splicing module, and a video optimization module; A text processing module is used to input the video production description text into a pre-fine-tuned large language model, and output video production plan information corresponding to the video production description text, wherein the video production plan information includes video script framework information, video artistic expression method information, and element supplementary description information; the video script framework information is used to describe the specific scenes and contents of each part of the beginning, middle, and end of the video; the artistic expression method information includes sound effect and music information, visual symbol information, and editing technique information; the element supplementary description information includes subtitle design and color contrast information; A material calling module, used to search the temporary video material uploaded by the user and the system material related to the video production description text from the BI system through the user identifier; The short video generation module is used to input the video production scheme information, the temporary video material and the system material into the video generation model to generate a target short video when the temporary video material is not empty; or, when the temporary video material is empty, input the video production scheme information and the system material into the video generation model to generate a target short video; generating the target short video includes: The semantic vector extraction module uses natural language processing (NLP) technology to parse the scene descriptions at the beginning, middle and end of the video to extract semantic vectors for key scene descriptions. The material feature extraction module classifies and annotates the image frames of the temporary video material and the system material when the temporary video material is not empty, so as to extract a first material feature of the temporary video material and a second material feature of the system material; the first material feature and the second material feature of the system material are composed of visual features and audio features; The material matching module performs material matching on the first material feature and the second material feature respectively according to the semantic vector to obtain a first matching result and a second matching result; the first matching result has a first target feature corresponding to part or all of the semantic vectors, and the second matching result has a second target feature corresponding to all of the semantic vectors; The video segment generation module generates a plurality of video segments suitable for the description of the beginning scene of the video, the description of the middle scene of the video, and the description of the ending scene of the video based on the first matching result and the second matching result; The video splicing module splices the multiple video clips based on the artistic expression technique information to obtain an initial short video; The video optimization module performs subtitle design and tone conversion on the initial short video based on the element supplementary description information to obtain a target short video; The video sending module is used to output the target short video corresponding to the video production description text and send the target short video to the client.

Citation Information

Patent Citations

  • Short video generation system and method and storage medium

    CN117528143A

  • Video generation method and device, electronic equipment and readable storage medium

    CN118214921A

  • Video generation method, computing device, computer storage medium and computer program product

    CN118972670A