Video generation method, device, electronic device and storage medium
By collecting user information to generate video copy and performing multi-dimensional classification and annotation, the problem of monotony in video generation is solved, and high-quality video generation with the content close to user's intentions and vivid pictures is achieved.
Patent Information
- Application Number
- CN202411425182.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The video images generated in the prior art are monotonous, and the content is not close to the user's real intentions, resulting in low video generation quality.
By collecting video association information input by the user, generating video copy information, and performing multi-dimensional classification and annotation, combining the video materials selected by the user, the target video is generated.
The generated video content is closer to the user's real intentions, and the pictures are vivid and rich, which improves the quality of video generation.
Smart Images

Figure CN119474454B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a video generation method, device, electronic device and storage medium. Background Art
[0002] We-media is a way for the general public to provide and share their own facts and news, connected to the global knowledge system through digital technology. It is a general term for new media that uses modern, electronic means to deliver both normative and non-normative information to a broad audience or specific individuals. With the rise of online we-media, a variety of short videos have emerged.
[0003] Existing video generation technologies include direct text-based video generation. This involves leveraging natural language pre-trained models and computer vision technology to automatically generate short videos based on given language input. However, these videos often lack visual interest and are often monotonous, lacking in relevance to the user's intended purpose. Summary of the Invention
[0004] The present application provides a video generation method, device, electronic device and storage medium to solve the problem that the video generated by the prior art has a monotonous picture and the content is not close enough to the user's real intention, resulting in low video generation quality.
[0005] In a first aspect, the present application provides a video generation method, comprising:
[0006] Based on the user's first operation, collecting video-related information input by the user;
[0007] Determining video text information based on a second operation of the user, wherein the video text information is converted from the video-related information;
[0008] Performing preset title processing on the video copy information to generate at least one video title;
[0009] Perform multi-dimensional classification on the video copy information to generate video annotation information;
[0010] determining a target video title based on a third operation of the user;
[0011] A target video is generated according to the target video title and video tag information, as well as the video material selected by the user.
[0012] Furthermore, the multi-dimensional classification of the video text information to generate video annotation information includes:
[0013] Marking the sentence breaks of the video text information as a first classification label;
[0014] Marking the modal particles in the video text information as a second classification label;
[0015] The keywords of the video text information are identified as a third classification label, wherein the keywords include time, place, characters, professional terms, hot search words, emotional words and appellation words in the video text information.
[0016] Furthermore, generating a target video according to the target video title and video annotation information, and the video material selected by the user, includes:
[0017] Respectively obtain video feature information associated with the first classification label, the second classification label, and the third classification label;
[0018] The video feature information, the target video title and the video material selected by the user are integrated to generate a video played by the digital human voice, wherein the video screen includes the digital human voice explanation display, video material display and subtitle display.
[0019] Furthermore, respectively obtaining the video feature information associated with the first classification label, the second classification label, and the third classification label includes:
[0020] Obtaining the digital human voice playback pause feature associated with the first classification label;
[0021] Obtaining the digital human voice playback expression feature associated with the second classification label;
[0022] The video picture effect features associated with the third classification label are obtained, wherein the video picture effect features include sound effect features, interesting animation features, eye-catching subtitle features and video material features.
[0023] Furthermore, the collecting of the video-related information input by the user based on the first operation of the user includes:
[0024] Receive user-triggered demand instructions;
[0025] In response to the demand instruction, collecting the user's answer to the video-related question through voice interaction in the form of asking questions to generate the video-related information; or
[0026] In response to the demand instruction, the user's answer to the video-related question is collected through interface text interaction in the form of questioning to generate the video-related information.
[0027] Furthermore, before determining the video copy information based on the user's second operation, the method further includes:
[0028] Integrating the video-related information to generate video copy information to be confirmed;
[0029] The video copy information to be confirmed is sent or visually displayed for the user to confirm whether it is used to generate the video.
[0030] Furthermore, determining the video copy information based on the user's second operation includes:
[0031] Upon receiving a confirmation instruction triggered by the user, determining the video copy information to be confirmed as the video copy information;
[0032] When a modification instruction triggered by a user is received, the to-be-confirmed video copy information after the user modification is obtained and determined as the video copy information.
[0033] In a second aspect, the present application provides a video generation device, comprising:
[0034] A collection module, configured to collect video-related information input by the user based on the user's first operation;
[0035] a text confirmation module, configured to determine video text information based on a second operation of the user, wherein the video text information is converted from the video-related information;
[0036] A title generation module, configured to perform preset title processing on the video copy information to generate at least one video title;
[0037] A classification and annotation module is used to perform multi-dimensional classification on the video copy information and generate video annotation information;
[0038] a title confirmation module, which determines the target video title based on the user's third operation;
[0039] The video generation module is used to generate a target video based on the target video title and video annotation information, as well as the video material selected by the user.
[0040] In a third aspect, the present application provides an electronic device comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to execute the video generation method described in the present application.
[0041] In a fourth aspect, the present application further provides a computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the video generation method described in the present application.
[0042] The above technical solution provided by the embodiment of the present application has the following advantages over the prior art: it can generate high-quality videos with content close to the user's true intention and vivid and rich pictures. The method provided by the embodiment of the present application, first, based on the user's first operation, collects the video related information input by the user, and based on the user's second operation, determines the video copy information, wherein the video copy information is converted from the video related information, then performs preset title processing on the video copy information to generate at least one video title, performs multi-dimensional classification on the video copy information to generate video annotation information, and based on the user's third operation, determines the target video title, and finally, generates the target video according to the target video title and video annotation information, as well as the video material selected by the user. By collecting the user's true intention for the video and integrating it, the generated video text content is closer to the user's needs, and by multi-dimensional classification and annotation of the video text content and associating multiple video feature information, the generated video pictures are vivid, eye-catching and impressive, thus ensuring the high quality of video generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0045] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0046] Figure 1 A flowchart of a video generation method provided in an embodiment of the present application;
[0047] Figure 2 A flowchart of another video generation method provided in an embodiment of the present application;
[0048] Figure 3 A module flow chart of a video generation device provided in an embodiment of the present application;
[0049] Figure 4 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] The disclosure below provides many different embodiments or examples for implementing different configurations of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0052] In order to solve the problems in the prior art, the present application provides a video generation method that can generate high-quality videos with content close to the user's true intentions and vivid and rich images.
[0053] Figure 1 A video generation method provided in an embodiment of the present application includes:
[0054] S101. Based on a first operation of a user, collecting video-related information input by the user;
[0055] In the embodiment of the present application, step S101 can be further divided into the following two sub-steps:
[0056] S1011, receiving a demand instruction triggered by a user;
[0057] S1012, responding to the demand instruction, collecting the user's answer to the video-related question by voice interaction in the form of asking questions to generate the video-related information; or
[0058] In response to the demand instruction, the user's answer to the video-related question is collected through interface text interaction in the form of questioning to generate the video-related information.
[0059] The video generation method of the present application can be used to make an application for providing users with video production. Specifically, a control button is set in the software interface of the application. If the user needs to make a video, click the control button to form the user's demand instruction. When the application backend receives the demand instruction, it responds in a timely manner and collects what kind of video the user wants to make and what it is used for by asking the user questions. In one embodiment, questions can be asked to the user in voice through a voice assistant, and the user answers by voice or text. Text questions can also be presented in the software interface, and the user enters the text answer in the interface text box or selects the answer in the option box. It should be noted that in the embodiment of the present application, a large model is used for intelligent matching to generate the next question, that is, the next question will be generated based on the answer to the user's previous question. In this way, the full set of questions corresponding to different users (video-related questions) is equivalent to customized, rather than a set of universal question templates, so that the collected video-related information is closer to the user's true intention, and more accurate and detailed.
[0060] In one embodiment, the video-related questions of this application can be essay questions, fill-in-the-blank questions, or multiple-choice questions. Furthermore, if the user's answer does not meet the requirements of intelligent matching, such as being short, ambiguous, or unclear, a corresponding prompt message will be sent to encourage the user to refine their answer. Furthermore, if the user's answer is not recognized by the large model, such as if the voice is noisy or the input text is garbled, a prompt message will be sent to encourage the user to answer again.
[0061] When the collected video-related information meets the preset rules, the questioning ends. The preset rules here can be understood as the large model has clearly defined the detailed information of the video to be generated by integrating the user's answers, such as: clearly presenting the video content that the user wants to output (collecting video content by asking questions, suitable for users to actively output video content); clearly matching the video content that the user wants (collecting video direction points by asking questions, the large model matches the knowledge base corresponding to the direction point and outputs the video content, suitable for video content that is output based on the user's answers using the large model intelligent matching); and confirming the query (asking the user if there is any additional information).
[0062] It can be seen from this that the first operation in the embodiment of the present application includes the control button clicked when making a video (for example, the "Start" button), the voice input operation, text input operation or option box selection operation of human-computer interaction.
[0063] S102: Determine video text information based on a second operation of the user, wherein the video text information is converted from the video-related information;
[0064] S103, performing preset title processing on the video copy information to generate at least one video title;
[0065] In the embodiments of the present application, the large model integrates and intelligently processes the collected video-related information to obtain video text information, which is the video text content of the target video that is ultimately generated. In one embodiment, the video text information can be displayed on the software interface or sent to the mobile phone number used by the user to log in to the application. The user can then confirm the information by clicking a "Confirm" button on the interface or by replying to the message.
[0066] In the embodiment of the present application, one or more video titles are generated for the user to select. Specifically, the text of the video title is extracted from the video copy information and can reflect the main theme and central meaning of the video copy information. The sentence structure of the video title is the hot search sentence structure, high-like sentence structure, rhetorical question sentence structure, etc. in the knowledge base matched by the large model, aiming to make the video title concise, summarized, vivid and able to quickly attract people's attention. Therefore, the preset title processing in this application includes the selection and extraction of the text and sentence structure of the video title by the large model based on the massive knowledge base.
[0067] S104: Perform multi-dimensional classification on the video text information to generate video annotation information;
[0068] In order to make the final target video vivid and eye-catching, in an embodiment of the present application, the video copy information is multi-dimensionally classified and labeled, specifically: marking the pauses between sentences in the video copy information, the modal particles in the video copy information, and multiple categories of keywords in the video copy information (such as: impressive expressions, key content of the video, eye-catching expressions, interesting and vivid expressions, introduced allusions, important dates and places, etc.).
[0069] In one embodiment, step S104 includes the following sub-steps:
[0070] S1041: Mark the sentence breaks of the video text information as first classification labels;
[0071] S1042: Mark the modal particle in the video text information as a second classification label;
[0072] S1043. Mark the keywords of the video text information as a third classification label, wherein the keywords include time, place, characters, professional terms, hot search words, emotional words and appellation words in the video text information.
[0073] Specifically, the classification labeling method in data labeling is adopted. The first classification label indicates the punctuation points between sentences in the video copy information, that is, the places where the punctuation and pause processing are performed when the voice is played. The second classification label marks the modal particles in the video copy information. The third classification label marks the keywords in the video copy information. In this way, the final target video voice playback content has pauses and intonations, which is close to the state of real human emotional speech, rather than mechanical word-by-word output. At the same time, the key point information in the video copy information is equipped with sound effects, subtitles, animation effects processing and video material insertion, making the video screen more vivid and impressive.
[0074] S105: Determine the target video title based on the user's third operation;
[0075] In an embodiment of the present application, multiple video titles are displayed for the user to choose from. The user can select the video title that he or she is interested in as the target video title, and the target video title is displayed in the final target video.
[0076] S106: Generate a target video based on the target video title and video annotation information, as well as the video material selected by the user.
[0077] In one embodiment, based on the detailed sub-steps of step S104, step S106 includes the following sub-steps:
[0078] S1061. Obtaining pause features of the digital human voice playback associated with the first classification label;
[0079] S1062: Acquire the digital human voice playback expression feature associated with the second classification label;
[0080] S1063: Obtain video picture effect features associated with the third classification tag, wherein the video picture effect features include sound effect features, interesting animation features, eye-catching subtitle features, and video material features;
[0081] S1064: Integrate the pause features of the digital human voice playback, the expression features of the digital human voice playback, the video screen effect features, the target video title, and the video material selected by the user to generate a video played by the digital human voice, wherein the screen of the video includes the digital human voice explanation display, the video material display, and the subtitle display.
[0082] In the embodiments of this application, different classification labels are associated with different video feature information. For example, this video feature information includes pause characteristics of the digital human voice playback, expression characteristics of the digital human voice playback, and video image effect characteristics. This can be understood as having the generated digital human voice play the video, making sentence pauses based on the content of the video copy, displaying different expressions based on different modal particles, and inserting video sound effects, video materials, special subtitles, and fun animated emoticons based on various keywords. Specifically, this can be achieved using large-scale model machine learning and image processing technologies.
[0083] For example, when the video content talks about "patent infringement", a video material of an infringement case can pop up, as well as special subtitles (for example: infringement affects the company's reputation and requires huge compensation), and funny sound effects and interesting animated emoticons that are appropriate to the situation can also pop up.
[0084] It should be noted that the video material in this application can be selected by the user from the videos stored locally on his terminal device, or from the default video material library of the application.
[0085] In an embodiment of the present application, first, based on the user's first operation, the video-related information input by the user is collected, and based on the user's second operation, the video text information is determined, wherein the video text information is converted from the video-related information. Then, the video text information is processed with preset titles to generate at least one video title, and the video text information is multi-dimensionally classified to generate video annotation information. Based on the user's third operation, the target video title is determined. Finally, the target video is generated based on the target video title and video annotation information, as well as the video material selected by the user. By collecting the user's true intention for the video and integrating it, the generated video text content is closer to the user's needs. By multi-dimensionally classifying and annotating the video text content and associating multiple video feature information, the generated video picture is vivid, eye-catching and impressive, thus ensuring the high quality of video generation.
[0086] Furthermore, in order to make the content of the generated video closer to the user's true intention, the embodiment of the present application also provides a method for confirming and modifying the video copy information to ensure that the final video copy information is confirmed or modified by the user to meet the user's video needs.
[0087] For details, see Figure 2 Before determining the video copy information based on the user's second operation, the video generation method in this application further includes:
[0088] S107: Integrate the video-related information to generate video text information to be confirmed;
[0089] S108: Send or visually display the video copy information to be confirmed for the user to confirm whether to use it for generating the video.
[0090] In one embodiment, a large model is first used to integrate and process the collected video-related information to generate the pending video copy information. This is then visually displayed on the software interface, with corresponding "Confirm" and "Edit" buttons provided for user operation. Furthermore, the pending video copy information can be sent to the mobile phone number used by the user to register the application, with corresponding confirmation or modification methods provided within this message. For example, replying with a numeric string indicates confirmation, while replying with the user's modified and adjusted content serves as the final modification confirmation message.
[0091] Specifically, after the pending video copy information is sent or visually displayed to the user, if a confirmation instruction is received from the user, the pending video copy information is determined as the video copy information; if a modification instruction is received from the user, the pending video copy information after the user's modification is obtained and determined as the video copy information. Of course, the pending video copy information after the user's modification can also be set to reconfirm to ensure the accuracy of the video copy information.
[0092] It should be noted that the video copy information in this application is the text content of the generated video, and is also the content played by the digital human voice.
[0093] like Figure 3 As shown, an embodiment of the present application provides a video generation device, including:
[0094] The collection module 301 is used to collect video-related information input by the user based on the user's first operation;
[0095] The text confirmation module 302 is configured to determine video text information based on the user's second operation, wherein the video text information is converted from the video-related information;
[0096] The title generation module 303 is used to perform preset title processing on the video copy information to generate at least one video title;
[0097] The classification and annotation module 304 is used to perform multi-dimensional classification on the video text information and generate video annotation information;
[0098] The title confirmation module 305 determines the target video title based on the user's third operation;
[0099] The video generation module 306 is configured to generate a target video according to the target video title and video annotation information, as well as the video material selected by the user.
[0100] It should be noted that a video generation device in an embodiment of the present application can be a software framework module in a terminal device (such as a mobile phone), and the various modules interact and communicate with each other to execute the video generation method provided by any of the aforementioned method embodiments.
[0101] like Figure 4 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0102] Memory 113, for storing computer programs;
[0103] In one embodiment of the present application, the processor 111 is configured to implement the video generation method provided by any one of the aforementioned method embodiments when executing a program stored in the memory 113 .
[0104] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the video generation method provided in any of the aforementioned method embodiments are implemented.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0107] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0108] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A video generation method, characterized in that: The method comprises: Based on the user's first operation, collecting video-related information input by the user; Determining video text information based on a second operation of the user, wherein the video text information is converted from the video-related information; Performing preset title processing on the video copy information to generate at least one video title; Perform multi-dimensional classification on the video copy information to generate video annotation information; determining a target video title based on a third operation of the user; Generate a target video based on the target video title and video annotation information, and the video material selected by the user; The multi-dimensional classification of the video text information to generate video annotation information includes: Marking the sentence breaks of the video text information as a first classification label; Marking the modal particles in the video text information as a second classification label; Marking the keywords of the video text information as a third classification tag, wherein the keywords include time, place, characters, professional terms, hot search words, emotional words and appellation words in the video text information; The collecting of the video-related information input by the user based on the first operation of the user includes: Receive user-triggered demand instructions; In response to the demand instruction, collecting the user's answer to the video-related question through voice interaction in the form of asking questions to generate the video-related information; or In response to the demand instruction, the user's answer to the video-related question is collected through interface text interaction in the form of questioning to generate the video-related information.
2. The method according to claim 1, characterized in that Generating a target video according to the target video title and video tag information, and the video material selected by the user, includes: Respectively obtain video feature information associated with the first classification label, the second classification label, and the third classification label; The video feature information, the target video title and the video material selected by the user are integrated to generate a video played by the digital human voice, wherein the video screen includes the digital human voice explanation display, video material display and subtitle display.
3. The method according to claim 2, characterized in that The obtaining of video feature information associated with the first classification label, the second classification label, and the third classification label respectively includes: Acquire the digital human voice playback pause feature associated with the first classification label; Obtaining the digital human voice playback expression feature associated with the second classification label; The video picture effect features associated with the third classification label are obtained, wherein the video picture effect features include sound effect features, interesting animation features, eye-catching subtitle features and video material features.
4. The method according to claim 1, wherein Before determining the video copy information based on the second operation of the user, the method further includes: Integrating the video-related information to generate video copy information to be confirmed; The video copy information to be confirmed is sent or visually displayed for the user to confirm whether it is used to generate the video.
5. The method according to claim 4, characterized in that The determining of the video copy information based on the second operation of the user includes: Upon receiving a confirmation instruction triggered by the user, determining the video copy information to be confirmed as the video copy information; When a modification instruction triggered by a user is received, the to-be-confirmed video copy information after the user modification is obtained and determined as the video copy information.
6. A video generating device, characterized in that: The device comprises: A collection module, configured to collect video-related information input by the user based on the user's first operation; a text confirmation module, configured to determine video text information based on a second operation of the user, wherein the video text information is converted from the video-related information; A title generation module, configured to perform preset title processing on the video copy information to generate at least one video title; A classification and annotation module is used to perform multi-dimensional classification on the video copy information and generate video annotation information; a title confirmation module, which determines the target video title based on the user's third operation; A video generation module, configured to generate a target video based on the target video title and video annotation information, as well as the video material selected by the user; The multi-dimensional classification of the video text information to generate video annotation information includes: Marking the sentence breaks of the video text information as a first classification label; Marking the modal particles in the video text information as a second classification label; Marking the keywords of the video text information as a third classification tag, wherein the keywords include time, place, characters, professional terms, hot search words, emotional words and appellation words in the video text information; The collecting of the video-related information input by the user based on the first operation of the user includes: Receive user-triggered demand instructions; In response to the demand instruction, collecting the user's answer to the video-related question through voice interaction in the form of asking questions to generate the video-related information; or In response to the demand instruction, the user's answer to the video-related question is collected through interface text interaction in the form of questioning to generate the video-related information.
7. An electronic device, characterized in that: include: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to execute the method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: Computer-executable instructions are stored, and the computer-executable instructions are used to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video processing method and device, equipment and medium
CN113672765A
Video processing method and electronic device
WO2024051760A1