Video generation method and apparatus, and device, storage medium and product
By obtaining text information from content publishing platforms and generating video themes, and automating the processing of video materials, the problem of long video generation time has been solved, enabling efficient and timely short video production.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2026-03-19
AI Technical Summary
The existing technology for video generation is time-consuming, resulting in low video generation efficiency. This is especially true when producing short videos with a strong time sensitivity, where the timeliness of the theme is easily lost due to insufficient time.
The system retrieves specified types of text information from a content publishing platform using electronic devices, extracts keywords to generate video themes, and extracts keyframes from associated video materials to automatically generate the target video.
It improves the automation level of video generation, saves manual time, enhances video generation efficiency and timeliness, and adapts to rapidly changing video theme requirements.
Smart Images

Figure CN2025113253_19032026_PF_FP_ABST
Abstract
Description
Video generation method, device, equipment, storage medium and product
[0001] The present application claims priority to the Chinese patent application No. 2024112969999, filed on September 14, 2024, and entitled "Video generation method, device, equipment, storage medium and product". TECHNICAL FIELD
[0002] The present application relates to the technical field of computer, in particular to a video generation method, device, equipment, storage medium and product.
[0003] BACKGROUND
[0004] With the rapid development of multimedia technology, video can quickly convey information and ideas, and more intuitively show the true face of things, and gradually become an indispensable part of people's life. With the popularity of various video playing platforms and the rapid growth of video demand, how to quickly and efficiently generate video has become one of the most urgent needs in the current video production field.
[0005] At present, the generation of video needs to consume a large amount of human resources. Video producers collect related video materials for editing according to the hot information on the network, and then publish the produced video. In this process, the video producer needs to consume a lot of time to complete the production of the video, which leads to a long time of video generation, a slow speed of video generation and a low efficiency of video generation. SUMMARY
[0006] The embodiments of the present application provide a video generation method, device, equipment, storage medium and product, which can improve the automation degree of video generation and improve the efficiency of video generation and the timeliness of the generated video.
[0007] In one aspect, the embodiments of the present application provide a video generation method, comprising:
[0008] obtaining text information of a specified type from a first content publishing platform;
[0009] extracting keywords from the text information, and generating a video theme based on the keywords;
[0010] obtaining candidate video materials associated with the video theme from a second content publishing platform based on the video theme; and
[0011] extracting key frames from the candidate video materials based on a preset frame rate, and generating a target video corresponding to the video theme according to the key frames.
[0012] In another aspect, the embodiments of the present application provide a video generation device, comprising:
[0013] an acquisition unit, configured to acquire text information of a specified type from a first content publishing platform;
[0014] a processing unit, configured to extract a keyword from the text information, and generate a video theme based on the keyword;
[0015] The acquisition unit is further configured to acquire candidate video clips associated with the video theme from a second content publishing platform based on the video theme; and
[0016] The processing unit is further configured to extract a key frame from the candidate video clips based on a preset frame rate, and generate a target video corresponding to the video theme according to the key frame.
[0017] In another aspect, an electronic device is provided, which includes at least one processor, and a memory configured to store at least one computer program, which, when executed by the at least one processor, causes the electronic device to implement the video generation method of the first aspect.
[0018] In another aspect, a computer readable storage medium is provided, which stores instructions, which, when executed on a computer, causes the computer to execute the video generation method of the first aspect.
[0019] In another aspect, a computer program product is provided, which includes computer programs or computer instructions, which, when executed by a processor, implement the video generation method of the first aspect.
[0020] Brief Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] FIG. 1 is an architecture schematic diagram of a video generation system provided by an embodiment of the present application;
[0023] FIG. 2 is a flow schematic diagram of a video generation method provided by an embodiment of the present application
[0024] FIG. 3 is a schematic diagram of a method for acquiring text information of a specified type provided by an embodiment of the present application;
[0025] FIG. 4 is a schematic diagram of a method for generating a video theme by a video generation device according to an embodiment of the present application;
[0026] FIG. 5 is a schematic diagram of a user interface for video generation according to an embodiment of the present application;
[0027] FIG. 6 is a schematic diagram of a method for generating a target video based on candidate video materials according to an embodiment of the present application;
[0028] FIG. 7 is another schematic diagram of a method for video generation according to an embodiment of the present application;
[0029] FIG. 8 is a schematic diagram of a user interface for displaying user feedback data according to an embodiment of the present application;
[0030] FIG. 9 is a schematic diagram of a user interface for displaying analysis results according to an embodiment of the present application;
[0031] FIG. 10 is a schematic diagram of an architecture of a method for video generation according to an embodiment of the present application;
[0032] FIG. 11 is a schematic diagram of a structure of a video generation device according to an embodiment of the present application;
[0033] FIG. 12 is a schematic diagram of a structure of an electronic device according to an embodiment of the present application.
[0034] Embodiments
[0035] It should be noted that, in order to make the person skilled in the art better understand the technical solutions provided by the embodiments of the present application, the embodiments of the present application will be described clearly and completely in combination with at least one drawing to describe the implementation manners of the technical solutions provided by the embodiments of the present application. Moreover, the drawings shown by the embodiments of the present application are only exemplary descriptions, for example, the execution order of each step in the drawings can be adaptively adjusted according to the actual application scene. In addition, in the embodiments of the present application, the block diagram shown in each drawing is only a functional entity, which does not necessarily correspond to a physically independent entity. That is, the functional entities can be realized in the form of software, or realized in at least one hardware module or integrated circuit, or realized in different network and / or processor devices and / or microcontroller devices.
[0036] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement at least one module or unit. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0037] It should be noted that "multiple" referred to herein means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship.
[0038] At present, due to the popularity of various video playing platforms, various different videos have been derived, especially videos with relatively short time lengths (i.e., short videos). A short video refers to a video with a time length of less than 5 minutes, which is suitable for people to watch in a mobile state or only with a short leisure time. Since a short video has a fast rhythm, people can obtain information in a short time, and therefore the demand for short videos is rapidly increasing, and the demand for video production is also correspondingly increasing. In the process of producing a short video, a video producer usually needs to manually collect and screen topics on the network first, so as to determine a video theme, and then obtain related video materials for editing, and finally publish the video on a video playing platform. In the process of producing a short video, the screening of a video theme and the editing of a video usually need to consume a large amount of time of the video producer, resulting in a slow video production speed and a low video generation efficiency. Moreover, some video themes have a certain timeliness, and if the video production consumes too long time, the timeliness of the video theme may be lost in the process of video production, resulting in a feedback of the produced video that does not meet the expectation of the video producer.
[0039] Based on this, the embodiments of the present application provide a video generation scheme, which obtains text information of a specified type from a first content publishing platform by an electronic device, and generates a video theme based on the obtained text information. Then the electronic device can obtain associated video materials from a second content publishing platform based on the video theme, the electronic device can extract key frames in the video materials, and generate a target video corresponding to the video theme according to the key frames. It can be seen that the entire video production process can be automatically completed by the electronic device, which improves the automation degree of the video production process, saves the time consumed for manually producing a video, reduces the labor cost, is conducive to improving the efficiency of video generation, and can improve the timeliness of the published video to a certain extent.
[0040] Based on the above description, refer to FIG. 1, which is an architecture diagram of a video generation system provided by an embodiment of the present application. As shown in FIG. 1, the video generation system includes a video generation device 101, a first content publishing platform 102, and a second content publishing platform 103. The video generation device 101 can be directly or indirectly connected to the first content publishing platform 102 through wired or wireless means, and can also be directly or indirectly connected to the second content publishing platform 102 through wired or wireless means. The first content publishing platform 102 and the second content publishing platform 103 can include at least one server, such as a business server or a background server. It should be noted that the number and form of devices shown in FIG. 1 are used for example and do not limit the embodiments of the present application.
[0041] The video generation device 101 is an electronic device, which can be a terminal device, including but not limited to a smart phone (such as an Android phone, an IOS phone, etc.), a tablet computer, a portable personal computer, a mobile Internet device (MID), a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable device, etc. The embodiments of the present application do not limit this.
[0042] In the embodiments of the present application, the first content publishing platform 102 and the second content publishing platform 103 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0043] The video generation device 101 can deploy an artificial intelligence (AI) driven software system, which can call machine learning models for processing different tasks to generate videos. In some embodiments, the video generation device 101 can deploy machine learning models for processing different tasks to generate videos through the deployed machine learning models.
[0044] In some embodiments, the video generation device 101 can be a node device in a distributed system, capable of handling a task in the video generation process, such as acquiring text information or generating video themes. Other node devices in the distributed system can be used to handle other tasks in the video generation process. Thus, by employing a distributed computing approach, the speed of data processing and video generation can be improved to some extent.
[0045] In some embodiments, the video generation device 101 can be an electronic device equipped with an output device, which can be used to output a user interface for setting video generation strategies in the video generation process, and can also output user feedback data received after the generated video is published to a content publishing platform, so that the video creator can monitor the user feedback data, etc. Optionally, the video generation device 101 can also receive instructions input by the video creator based on the user interface.
[0046] In some embodiments, the video generation device 101 may be a cloud device in a cloud computing platform, which can be used for larger-scale data processing and storage. The video generation device 101 may also be an edge device or other devices, which are not limited in this application.
[0047] In this embodiment, the types of content published by the first content publishing platform 102 and the second content publishing platform 103 can be different. For example, the first content publishing platform 102 can be a content publishing platform for publishing text and image content, including image and text information, such as a news sharing platform, a Q&A platform, a discussion platform, etc. The second content publishing platform 103 can be a content publishing platform for publishing video content, such as a social media platform, a short video publishing platform, a long video publishing platform, etc.
[0048] In this embodiment, the first content publishing platform 102 and the second content publishing platform 103 can publish the same type of content. For example, they can be used to publish any type of content, such as image content, text content, and video content.
[0049] The video generation system in this embodiment also includes a database for storing data generated during the video generation process. This database can be a local database on the video generation device 101 or a cloud-based database; this application does not limit the scope of the database. The database can be a NoSQL (Not Only Structured Query Language) database, such as MongoDB or an Elastic Search database, or a relational database, such as a MySQL (My Structured Query Language) database.
[0050] The processing method of the response message provided in the application has the following general process:
[0051] The video generation device 101 can obtain text information of a specified type from the first content publishing platform 102. Specifically, the video generation device 101 can send a content obtaining request to the first content publishing platform 102, and receive a response message returned by the first content publishing platform 102 in response to the content obtaining request. The response message carries a source code document of the first content publishing platform 102, and the text information of the specified type can be extracted from the source code document. Further, the video generation device 101 extracts keywords from the text information, and generates a video theme based on the extracted keywords. After generating the video theme, the video generation device 101 can obtain candidate video materials associated with the video theme from the second content publishing platform 103 based on the video theme, extract key frames from the candidate video materials based on a preset frame rate, and generate a target video corresponding to the video theme according to the key frames.
[0052] In an embodiment, the video generation device 101 can also publish the generated target video in a content publishing platform (such as the second content publishing platform 103). Before publishing the target video, the video generation device 101 can obtain the publishing time corresponding to each of a plurality of videos published in the second type of content publishing platform 103 and the interaction data corresponding to each of the plurality of videos, to determine the optimal publishing time corresponding to the second type of content publishing platform 103. Further, the video generation device 101 can call a video publishing interface of the second type of content publishing platform 103 based on the optimal publishing time, and publish the target video.
[0053] In one implementation, the above-mentioned text information of a specified type, keywords, video themes, key frames, etc. can be saved in a blockchain, which can prevent these information from being tampered with. The blockchain is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. computer technology, and its essence is a decentralized database, which is a series of data blocks associated using cryptography. Each data block contains information of a batch of network transactions, which is used to verify the validity (anti-fake) of the information and generate the next block.
[0054] In the embodiments of the application, the video generation system described in the embodiments of the application is used to more clearly illustrate the technical solutions of the embodiments of the application, and does not constitute a limitation on the technical solutions provided by the embodiments of the application. Those skilled in the art can know that, with the evolution of system architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the application are also applicable to similar technical problems.
[0055] Based on the video generation system described above, the embodiment of the present application provides a video generation method. The video generation method described in the embodiment of the present application can be executed by an electronic device. The electronic device can be the video generation device 101 in the video generation system shown in FIG. 1. Please refer to FIG. 2, which is a flowchart of a video generation method according to an embodiment of the present application. The video generation method includes the following steps S201-S204:
[0056] S201, obtaining text information of a specified type from a first content publishing platform.
[0057] In all embodiments of the present application, the content publishing platform refers to a platform that allows users to create, edit, manage and publish various contents. The content published by the user can be in various forms, such as image content, text content, video content, etc. The content publishing platform can be a social media platform, a news sharing platform, a Q&A platform, a discussion platform, etc. The present application does not limit the content publishing platform.
[0058] The specified type refers to the type of the subject of the video to be generated. For example, the specified type can include education type, entertainment type, film review type, etc. Alternatively, the specified type can also refer to the timeliness type, for example, the specified type can be a hot type (such as current affairs hotspots) and the like.
[0059] The text information refers to the text content contained in the content publishing platform. For example, taking the specified type as a hot type, if the content publishing platform is a news sharing platform, the text information of the specified type can include the text content of the latest news and the text content of the headline news.
[0060] In one possible implementation, the video generation device can simulate the behavior of network users through network crawler technology, automatically access and capture the text information of the specified type in the content publishing platform. Specifically, the video generation device can first send a content acquisition request to the server deployed on at least one content publishing platform. The content acquisition request is used to request to acquire the multimedia content on the content publishing platform. The number of content publishing platforms can be one or more, such as at least one social media platform, at least one news sharing platform, etc. The server on the content publishing platform can be a business server or a background server, which is used to execute the business logic of the content publishing platform, such as querying the database, processing user input, generating response, etc. The acquisition request can be a hypertext transfer protocol (HTTP) request sent by the video generation device, which can be used to acquire the platform content of the content publishing platform.
[0061] In the embodiments of the present application, the video generation device can simulate the behavior of a user accessing the content access platform through a browser, and access the content publishing platform. For example, the video generation device can imitate the user accessing the content publishing platform through a uniform resource locator (URL) based on the content publishing platform, sending an HTTP request to the server on the content publishing platform, and based on the response message returned by the server on the content publishing platform, obtaining the content in the content publishing platform.
[0062] Further, the video generation device can receive a response message returned by the server on each content publishing platform in response to the content acquisition request, and the response message includes the source code document of the content publishing platform. For example, the response message can be an HTTP response, and the source code document can be a Hyper Text Markup Language Document (HTML) document, which can include the content in the corresponding content publishing platform. Wherein, the HTML document can include code tags, i.e. HTML tags, for defining the structure and content of the HTML document, and different contents in the content publishing platform can be contained in different code tags, thereby constituting the entire HTML document.
[0063] Further, the video generation device can extract the text information of a specified type from the source code document based on the code tag associated with the text information of the specified type. Wherein, the code tag associated with the text information of the specified type can be a specific type of HTML tag. For example, if the text information of the specified type is a hot topic on a social media platform, the code tag associated with the text information of the specified type can be a region tag containing the topic title and content, such as Tag. After the video generation device acquires the HTML document, the HTML document can be parsed, and the required information can be extracted by locating the specific HTML tag, that is, the text information of the specified type is obtained.
[0064] In the embodiment of the application, the video generation device can be an electronic device deployed with a crawler system, which can be developed based on a Scrapy framework. The Scrapy framework is an open source framework written in python, which can provide powerful functions such as request scheduling, data parsing, data storage, etc., and is widely used in data collection, information retrieval and other fields. The Scrapy framework can quickly realize data crawling and network data processing by writing crawler scripts, such as defining crawler classes and writing parsing functions. The crawler system includes a scheduler (Scheduler), a downloader (Downloader), a parser (Parser) and a data storage (Data Storage). The scheduler can be used to schedule crawling tasks, set crawling frequency and priority of content publishing platforms required to be crawled, that is, priority of URLs of each content publishing platform. The downloader can be used to download web page content and pass the downloaded content to the parser. The parser can be used to parse the content downloaded by the downloader to parse the content of the content publishing platform (web page), so as to extract the required information, that is, the text information of the specified type. The data storage can be used to store the extracted text information in the database for subsequent processing.
[0065] Specifically, please refer to FIG. 3, which is a method for obtaining text information of a specified type provided by the embodiment of the application. As shown in FIG. 3, the scheduler 302 can be configured with the crawling address of the data source, which is the URL of the data to be acquired, that is, the URL of at least one content publishing platform, and the crawling frequency and the priority of each URL. The server pointed by the URL includes the server 3041 on platform 1, the server 3042 on platform 2, …, and the server 304n on platform n.
[0066] Further, the scheduler 302 triggers the downloader 303 to send an HTTP request to the server corresponding to the URL with the highest priority based on the pre-set priority, for example, sends an HTTP request to the server 3041 on platform 1, and receives the response message returned by the server. Further, the downloader 303 can pass the source code document (HTML document) included in the response message to the parser 305.
[0067] The parser 305 can be configured to parse the source code document to obtain content on the content publishing platform (e.g., a webpage), and specifically, can use a code parsing library 306, such as BeautifulSoup, lxml, or other HTML parsing libraries, to parse the source code document, and based on the parsed results, extract the required information to obtain the specified type of text information. That is, the parser 305 can extract the specified type of text information from the HTML document based on the code tags associated with the specified type of text information. Finally, the extracted text information is stored in the database 307 for subsequent processing.
[0068] In the embodiments of the present application, based on the priority of the pre-set URL, HTTP requests can be sequentially sent to the servers indicated by the subsequent second priority, third priority, and the like, and the specified type of text information can be obtained and stored in the database 307.
[0069] For example, if the content publishing platform is a social media platform, the text information obtained by the video generation device can include popular topics, user comments, and the like. If the content publishing platform is a Q&A platform and a discussion platform, the text information obtained by the video generation device can include questions and answers with high discussion degrees, and posts and user replies with high discussion degrees. If the content publishing platform can also be a news sharing platform, the text information obtained by the video generation device can include the latest news, headline news, and the like.
[0070] Therefore, the video generation device can automatically capture the latest hot information on the content publishing platform. Furthermore, the video generation device can store the captured text information in a NoSQL database (such as MongoDB, Elasticsearch) to facilitate storage and management of the obtained text information, and support fast query and retrieval.
[0071] In the embodiments of the present application, the video generation device can periodically capture the specified type of text information (such as hot information) on the content publishing platform, which can ensure the timeliness and accuracy of the data and provide basic data support for subsequent video generation, such as video editing and publishing.
[0072] In some embodiments, a list of user-agent strings can be maintained in the downloader 303, and when generating an HTTP request, a user-agent string can be selected (e.g., randomly selected) from the list and added to the request header of the HTTP request to be sent to the corresponding server. In addition, a list of proxy Internet Protocol (IP) addresses (e.g., an IP address pool) can also be maintained in the downloader 303, and the proxy IP addresses in the list can be from different geographic locations and Internet service providers, and are proxy IP addresses of at least one proxy server. After the downloader 303 specifies a proxy IP address in the parameter of the HTTP request, the video generation device can send the HTTP request to the proxy server, and the proxy server forwards the HTTP request to the server on the content publishing platform, and then after receiving the response message returned by the server, the response message can be returned to the video generation device. In this way, the behavior of a user accessing the content publishing platform can be simulated through the anti-crawling mechanism to avoid being banned by the content publishing platform.
[0073] In some embodiments, the downloader 303 can automatically retry if an exception occurs during data download, such as a URL error or a server response timeout. The downloader 303 is also configured with an error handling strategy, for example, after the number of retries for the same URL reaches a preset threshold, the URL can be skipped and the next data of the content publishing platform can be automatically crawled. In this way, the stability of data crawling can be ensured through the automatic retry and error handling mechanism.
[0074] In some embodiments, the video generation device can also obtain the text data of the specified type through an Application Programming Interface (API) provided by each content publishing platform.
[0075] S202, extracting keywords from the text information, and generating a video theme based on the keywords.
[0076] In the embodiments of the present application, the keywords can refer to representative and summary words or phrases extracted from the specified type of text information. The specific extraction method can be to analyze and process the text information by using Natural Language Processing (NLP) technology, and extract valuable keywords.
[0077] In addition, the keywords can be combined to generate a video theme. The video theme refers to the idea around which the video content is centered, that is, the core information that the video content conveys. For example, the video theme can be a video theme associated with a specified category, such as hot information, career education, etc.
[0078] In a possible implementation, the specific extraction manner includes:
[0079] generating a text document according to the content in the first content publishing platform;
[0080] parsing the text information to obtain a set of words contained in the text information, the set of words including a plurality of words;
[0081] determining a weight corresponding to each word based on a number of times each word appears in the text document;
[0082] selecting a preset number of words as the plurality of keywords in a descending order of the weights.
[0083] The text document contains various types of text information published and stored in the content publishing platform. Each text document can be divided or identified according to the content publishing platform. The video generation device can extract the text information obtained from each content publishing platform in the database and store the text information in different text documents. In some embodiments, after the video generation device obtains the text information of each content publishing platform, the video generation device can generate a text document corresponding to each content publishing platform based on the obtained text information, and store the text document corresponding to each content publishing platform in the database.
[0084] The parsing can include Tokenization processing, Stop Words Removal processing, and Stemming processing. The Tokenization processing refers to dividing the content in the text information into words or phrases.
[0085] The Tokenization processing refers to dividing the content in the text information into words or phrases.
[0086] The Stop Words Removal processing refers to removing irrelevant words or phrases, such as removing words such as "of", "is", "in", and words such as "the", "and", and "in" in English.
[0087] The Stemming processing refers to restoring words to their stem forms to reduce the diversity of words, for example, restoring the word "running" to "run" and restoring the word "programming" to "program".
[0088] In the embodiments of the present application, the video generation device can call the encapsulated library or framework to parse the text information, such as calling the Natural Language Toolkit (NLTK) library of Python to parse the text information.
[0089] Exemplarily, taking the text information "This is a hot topic on social media: #NLP and its applications. Users are commenting on its potential." as an example, the process of analysis includes:
[0090] Firstly, the text information is processed by word segmentation to obtain a first processing result: "This", "is", "a", "hot", "topic", "on", "social", "media", ":", "#NLP", "and", "its", "applications", ".", "Users", "are", "commenting", "on", "its", "potential", ".".
[0091] Further, the first processing result is processed by stop word removal to obtain a second processing result: "social media", "hot", "topic", "#NLP", "applications", "users", "commenting", "potential".
[0092] Finally, the second processing result is processed by stem extraction to obtain a word set. Since the second processing result is all stems, the word set is determined to be {"social media", "hot", "topic", "#NLP", "applications", "users", "commenting", "potential"}.
[0093] Among them, the number of occurrences of each word in the text document in the word set can be understood as the term frequency, that is, the frequency of the occurrence of the word in each text document. In an implementation manner, if the number of content publishing platforms is one, the video generation device can determine the weight corresponding to each word based on the number of occurrences of each word in the text document in the word set, and the weight can be the value of the number of occurrences, and then select a preset number of words as keywords in the order of weight from large to small.
[0094] In an embodiment, if the first content publishing platform generates a plurality of text documents, the video generation device can determine the weight corresponding to each word based on the sum of the number of occurrences of each word in each text document in the word set, and the weight can be the value of the sum of the number of occurrences, and then select a preset number of words as keywords in the order of weight from large to small.
[0095] In another implementation, the video generation device can extract keywords based on a Term Frequency-Inverse Document Frequency (TF-IDF) algorithm, which can be used to evaluate the importance of words in a document set (including multiple text documents). Among them, the term frequency TF refers to the number of times each word in the word set appears in each text document. The inverse document frequency IDF refers to the discriminability of each word in all documents, which is calculated based on the total number of text documents and the number of text documents containing the word. Then, the value of the term frequency and the inverse document frequency are multiplied to obtain the value of TF-IDF. The IDF in TF-IDF can be used to measure the universality of the word in the entire document set. In the embodiments of the present application, if a word appears frequently in a certain document but less frequently in other documents, it is considered that this word has strong discriminability, so the importance of the word can be calculated based on the value of TF-IDF, and the high-frequency and discriminative keywords can be extracted. That is, the value of TF-IDF corresponding to each word can be determined as the weight corresponding to each word, and then the keywords of a preset number of words can be selected in descending order of weight.
[0096] In another possible implementation, the video generation device can also extract keywords based on the TextRank algorithm. TextRank is a graph-based ranking algorithm that can be used for keyword extraction and text summarization in NLP. Based on the PageRank algorithm, a word co-occurrence graph or a sentence similarity graph is constructed in the text information, and the nodes (representing words or sentences) are sorted to extract important keywords or generate summaries.
[0097] Specifically, the video generation device can first parse the text information to obtain a word set contained in the text information, and the word set includes multiple words. Then, the video generation device can determine the association relationship between each word in the word set based on a preset word window.
[0098] Among them, the preset word window, also known as the co-occurrence window, is composed of a preset number of words, which can be used to judge the association degree between two words, and specifically refers to whether two words appear together in the word window or the relative position relationship between them. For example, if the length (i.e., the number of words) of the preset word window is K, and word 1 and word 2 appear together in the window with a length of K, it is determined that there is an association relationship between word 1 and word 2.
[0099] Thus, each word can be taken as a node, and the association relationship can be taken as an edge to construct an undirected graph. In an embodiment of the present application, if there is an association relationship between word 1 and word 2, there is an edge connecting the nodes corresponding to the two words between word 1 and word 3, and the undirected graph is a weighted undirected graph, that is, the undirected graph includes the initial weights corresponding to each node. The initial weight corresponding to each node can be a set value, for example, 1.
[0100] Further, the video generation device can update the initial weight corresponding to each node based on the association relationship, and select a preset number of words as the extracted keywords in descending order of the updated weights. Specifically, the video generation device can update the initial weight corresponding to each node based on the connection relationship between the nodes in the undirected graph (i.e., the association relationship between the words), such as the out-degree number, the in-degree number of each node, and the iteration formula of the TextRank algorithm, until the weight corresponding to each node meets the iteration end condition (or convergence condition), such as the maximum change of the weight corresponding to each node being less than a preset threshold, or the iteration number reaching a preset maximum value, etc. After the iteration end condition is met, the weight corresponding to each node is the maximum weight of each node. Finally, the video generation device can select the words corresponding to the preset number of nodes as the extracted keywords in descending order of the maximum weight of each node.
[0101] After the video generation device extracts the keywords, the video theme can be generated based on the extracted keywords in a topic modeling manner. Specifically, the video generation device can first construct a document-word vector based on the number of times each keyword appears in the text document, and then the video generation device can analyze the document-word vector based on a set number of themes to obtain a word distribution corresponding to each theme of the multiple themes, and finally the video generation device can generate a video theme based on the word distribution corresponding to each theme of the multiple themes.
[0102] In an embodiment, when the text document includes multiple sub-documents, a document-word matrix is constructed based on the number of times each keyword appears in each sub-document; the document-word matrix is analyzed to obtain a word distribution corresponding to each theme of the multiple themes; and the video theme is generated based on the multiple word distributions.
[0103] The document-word matrix refers to a matrix composed of the relationship between the document and the word, which can be a document-word matrix, i.e., a matrix in which the rows represent the documents and the columns represent the words, or a word-document matrix, i.e., a matrix in which the rows represent the words and the columns represent the documents. The elements in the matrix correspond to each word in the word set, and the value of the element is the number of times or the weight (such as the TF-IDF value) of the word appearing in the document.
[0104] In a possible implementation, the video generation device can extract the video theme based on keywords based on a latent Dirichlet allocation (LDA) model. The LDA is a theme model that can discover (extract) the latent theme from the text document through the probabilistic generation model. The principle is to assume that each sub-document is generated by mixing several themes, and each theme is composed of several words. Through the LDA model, the theme distribution of multiple sub-documents can be modeled to help understand the main content and structure of each sub-document, thereby extracting the latent theme.
[0105] Specifically, after constructing the document-word matrix, the video generation device can parse the document-word matrix based on the set number of themes to obtain the word distribution corresponding to each of the multiple themes under the number of themes. The parsing can be a sampling process on the document-word matrix through Gibbs sampling to obtain the document-theme distribution and the theme-word distribution. The theme-word distribution is the word distribution corresponding to each of the multiple themes under the number of themes, for example, the word distribution corresponding to each of the multiple themes under the number of themes is: theme 1: word 1 (0.7), word 2 (0.3); theme 2: word 3 (0.6), word 2 (0.2), word 4 (0.1), word 5 (0.1), wherein the values in the brackets represent the probability of the corresponding word.
[0106] In another possible implementation, the video generation device can extract the video theme based on keywords based on a non-negative matrix factorization (NMF) algorithm. The NMF is a matrix decomposition technique that can be used to decompose a non-negative matrix into the product of two non-negative matrices. In NLP, NMF can be used for theme extraction and dimension reduction by decomposing the word-document matrix into a document-theme matrix and a theme-word matrix to extract the theme distribution in the keywords.
[0107] Specifically, the parsing can refer to the decomposition of the matrix. After initializing the two decomposition matrices (i.e., the document-theme matrix and the theme-word matrix), the two decomposition matrices are updated alternately through an iterative optimization process to minimize the reconstruction error. After the iterative optimization ends, the word distribution corresponding to each of the multiple themes under the number of themes can be obtained, such as theme 1: word 1 (0.8), word 2 (0.2), theme 2: word 3 (0.7), word 4 (0.3).
[0108] Finally, the video generation device can generate a video theme based on the word distribution corresponding to each of the plurality of themes, i.e., generate a video theme according to the probability of each word under the theme. Specifically, the video generation device can input the word distribution corresponding to each of the plurality of themes, such as word 1 (0.7), word 2 (0.3), into the text generation model to obtain the video theme output by the text generation model. The video theme can be generated based on word 1 and word 2, and the video theme tends to indicate the content of word 1.
[0109] In a possible implementation, in the process of generating a video theme based on the word distribution corresponding to each of the plurality of themes, the video generation device can first input the keywords into a machine learning model for identifying emotional tendencies to obtain an emotional tendency category output by the machine learning model. The machine learning model for identifying emotional tendencies may, for example, be a Bidirectional Encoder Representations from Transformers (BERT), a Generative Pre-trained Transformer (GPT), etc., which are not limited in the present application.
[0110] BERT is a bidirectional encoder that can consider context information at the same time, which helps to understand sentence structure and semantics. GPT can be used for natural language generation tasks and can generate fluent and context-related text. Both BERT and GPT are Large Language Models (LLM), and in the embodiments of the present application, LLM can be used as a machine learning model for identifying emotional tendencies to classify the emotions of keywords.
[0111] Specifically, the video generation device can input the extracted keywords into the machine learning model for identifying emotional tendencies to obtain an emotional tendency category output by the machine learning model. The emotional tendency category may, for example, be divided based on anger, calm, joy, etc., or may be divided based on positive, negative, and neutral.
[0112] Further, the video generation device can input the emotional tendency category and the word distribution corresponding to each of the plurality of themes into the text generation model to obtain a video theme output by the text generation model. For example, the text generation model is GPT or the like.
[0113] Further, after the video generation device obtains the video topics, the video topics can be stored in a NoSQL database (such as MongoDB, Elasticsearch) to store and manage the video topics and support fast query and retrieval. In an embodiment of the present application, the video generation device can store the video topics generated based on the text information periodically crawled, such as the hot information described above, in the database.
[0114] Please refer to FIG. 4, which is a schematic diagram of a method for generating video topics by a video generation device according to an embodiment of the present application. As shown in FIG. 4, the text information 401 is parsed to obtain a word set 402, and then the weights of the words in the word set are calculated, and based on the weights, some words are selected as keywords 403.
[0115] Further, the number of times each keyword 403 appears in the text document is used to construct a document-word matrix 404 and a word-document matrix 405, and based on the two matrices respectively, the word distribution corresponding to each of the multiple topics under the set number of topics is determined 406.
[0116] In an embodiment, the video generation device can directly generate the video topics based on the word distribution corresponding to each of the multiple topics under the set number of topics. The video generation device also inputs the word distribution corresponding to each of the multiple topics under the set number of topics into a text generation model 409 to obtain the video topics output by the text generation model 409.
[0117] In another embodiment, the video generation device can also input the keywords into a machine learning model 407 for identifying sentiment orientation to obtain a sentiment orientation category 408, and then input the sentiment orientation category 408 and the word distribution corresponding to each of the multiple topics under the set number of topics into the text generation model 409 to obtain the video topics 410 output by the text generation model 409.
[0118] In some embodiments, the video generation device can output a user interface for generating a video, and the user can interact with the video generation device based on the user interface. Please refer to FIG. 5, which is a schematic diagram of a user interface 500 for generating a video according to an embodiment of the present application. As shown in FIG. 5, the user interface 500 can include an input control 501 for the user to input a video topic. The user can input a video topic filtered and selected by himself based on the input control 501, and then the video generation device can generate a target video based on the video topic input by the user.
[0119] Optionally, other selection controls can also be included in the user interface, such as a selection control 502 for selecting the video duration, a selection control 503 for selecting the content publishing platform, and a prompt information 504 for prompting the recommended video duration for the content publishing platform.
[0120] Further, based on the triggering operation of the user on the control 505 for automatically generating a video, a target video corresponding to the video theme input by the user is generated.
[0121] In the above embodiment, the way of inputting the video theme judged and screened by the user based on the input control is a semi-automatic scheme with manual assistance. Although the video generation efficiency is not as good as that of the video generation device automatically generating a target video, in some specific cases, a target video with higher content quality can be generated.
[0122] S203, based on the video theme, obtaining candidate video materials associated with the video theme from a second content publishing platform.
[0123] In the embodiment of the present application, the second content publishing platform refers to a content publishing platform including candidate video materials, which can be the same as or different from the second content publishing platform for obtaining text information of a specified type. For example, the content publishing platform for obtaining text information of a specified type is a social media platform, and the second content publishing platform is a video material platform.
[0124] The candidate video material refers to a video content that can be used for editing and does not involve copyright issues, and is mostly long video data. The video generation device can edit the candidate video material to generate a target video corresponding to the video theme. The video generation device can select at least one content publishing platform as the second content publishing platform based on the video theme from a plurality of content publishing platforms from which candidate video materials can be obtained, and then obtain candidate video materials associated with the video from the second content publishing platform.
[0125] In a possible implementation, the video generation device can obtain video related information of each content publishing platform from at least one content publishing platform, and generate a video tag corresponding to each content publishing platform based on the video related information. The video related information can be introduction information for the video, user comments on the video data, or other text information. The video generation device can obtain the video related information of each content publishing platform from at least one content publishing platform based on a crawler technology.
[0126] In an embodiment, video keywords are extracted from the crawled video related information. The extraction principle can be the same as that of extracting keywords from text information of a specified type, which can be referred to the description above, and will not be repeated here.
[0127] Further, the video generation device can generate video tags corresponding to each content publishing platform based on the extracted video keywords. The video generation device can take the video keywords as the video tags. In some embodiments, the video generation device can also filter the video keywords, for example, only keeping the video keywords related to the video content, or only keeping the video keywords related to the video content and user comments, and taking the kept video keywords as the video tags.
[0128] Thus, the video generation device can select, from the content publishing platforms corresponding to the video tags, a content publishing platform corresponding to a video tag matching the video theme as the second content publishing platform. The video generation device can calculate the word similarity between the video tags corresponding to each content publishing platform and the video theme, and then select at least one content publishing platform with the highest word similarity as the second content publishing platform.
[0129] After selecting the second content publishing platform, the video generation device can obtain candidate video materials associated with the video theme from the second content publishing platform. Specifically, the video generation device can obtain video materials from the second content publishing platform based on a crawler technology as the candidate video materials associated with the video theme.
[0130] In some embodiments, in order to reduce the computational load of video editing, the video generation device can only capture a preset number of video materials (such as video files) as the candidate video materials associated with the video theme. The video generation device can obtain the video materials from the second content publishing platform based on the crawler technology, which has been described above and will not be repeated here.
[0131] Thus, after two rounds of data crawling, the video generation device can obtain the candidate video materials, and then the video generation device can store the candidate video materials in a database for subsequent editing.
[0132] S204, extracting key frames from the candidate video materials based on a preset frame rate, and generating a target video corresponding to the video theme according to the key frames.
[0133] In the embodiments of the present application, the preset frame rate is used to extract key frames from the candidate video materials, which is usually expressed in frames per second (FPS), for example, the preset frame rate can be 5 frames per second. The key frame is a video frame extracted from the candidate video materials based on the preset frame rate.
[0134] The video generation device can extract key frames from the candidate video materials based on the preset frame rate through an API for extracting key frames or a preset framework for extracting key frames.
[0135] In the embodiments of the present application, the video generation device can perform extraction processing on each candidate video material to obtain key frames extracted from the candidate video materials. For example, if the video generation device obtains 100 candidate video materials, key frames are extracted from the 100 candidate video materials. The target video is a video generated by the video generation device, which is generated based on the candidate video materials obtained based on the video theme. Therefore, the target video is a video corresponding to the video theme.
[0136] In a possible implementation, in the process of extracting key frames from the candidate video materials based on the preset frame rate, the video generation device can assist in extracting the key frames based on a machine learning model to ensure that the generated video has high content quality and is attractive.
[0137] Specifically, the video generation device can extract an initial video frame set including a plurality of initial video frames from the candidate video materials based on a preset frame rate. For example, the video generation device can configure a preset frame rate, i.e., select a frame rate, such as 5 frames per second. Then, the initial video frame set is obtained by performing extraction processing on the video material at a frame rate of 5 frames per second, and the initial video frame set is a preliminary frame sequence.
[0138] Further, the video generation device can input each initial video frame in the initial video frame set into a machine learning model for scoring image quality to obtain an image quality score corresponding to each video frame output by the machine learning model.
[0139] The machine learning model for scoring image quality can be a convolutional neural network (CNN), such as a deep residual network (ResNet), a VGG network (Visual Geometry Group Network, VGG), etc. The machine learning model for scoring image quality can be obtained by training based on training video frames and corresponding score labels. Users can manually score the image quality of the training video frames to obtain the score labels. The scoring criteria for scoring image quality can include the definition, information amount, and emotional expression of the video frame, which are not limited in the present application.
[0140] Specifically, the video generation model can input each video frame in the initial video frame set into a pre-trained machine learning model, through which feature extraction processing can be performed on each video frame to generate a feature vector (i.e., feature representation information) of each video frame. The feature vector of each video frame can contain visual information in each video frame, such as color, texture, shape, etc. Further, the machine learning model can score each video frame based on the feature vector of each video frame to obtain an image quality score corresponding to each video frame output by the machine learning model. Thus, the video generation device can select a preset number of video frames from the initial video frame set as the plurality of key frames in order of image quality scores from high to low.
[0141] In some embodiments, the machine learning model for scoring image quality can also include an attention mechanism module, which can score more accurately in terms of clarity, information amount, emotional expression, etc.
[0142] In one possible implementation, a preset number of initial video frames are selected from the initial video frame set in order of image quality scores from high to low as a plurality of candidate key frames to form a candidate key frame set.
[0143] Further, video frames with excessively high similarity among the plurality of candidate key frames can be filtered and the filtered video frames can be determined as the plurality of key frames. For example, video frames with excessively high similarity are filtered to obtain filtered candidate key frames, thereby reducing redundant detection.
[0144] Specifically, the video generation device can perform target detection processing on each candidate key frame to obtain a detected target bounding box in each candidate key frame, and further filter the plurality of candidate key frames based on the detected target bounding boxes in each candidate key frame to select the plurality of key frames.
[0145] The target detection processing refers to processing for identifying and locating a specific target (or object) in each candidate key frame, specifically including identifying objects (targets) of a specific category in each candidate key frame, and then accurately marking the positions of these objects (targets), i.e., obtaining a detected target bounding box. The target bounding box can also be referred to as a bounding box, which is used to mark the position information of the detected target in each candidate key frame.
[0146] In some embodiments, the video generation device can input each candidate key frame into a target detection model, which can be, for example, a You Only Look Once (YOLO) model, a Single Shot MultiBox Detector (SSD) model, or the like, without limitation. After inputting each candidate key frame into the target detection model, a target detection box detected by the target detection model in each candidate key frame can be obtained. In an embodiment of the present application, a plurality of targets can be included in one candidate key frame, and the implementation manners of the plurality of targets are the same. For the convenience of description, an embodiment of the present application takes one target as an example for description, that is, the target detection box detected in each candidate key frame is a bounding box including the target.
[0147] Further, the video generation device can perform the following processing:
[0148] For each target detection box,
[0149] From the plurality of target detection boxes, a neighboring target detection box closest to the position of the target detection box is determined;
[0150] The intersection area and the union area between the target detection box and the neighboring target detection box are calculated;
[0151] Based on the intersection area and the union area, a similarity between the candidate key frame corresponding to the target detection box and the candidate key frame corresponding to the neighboring target detection box is determined;
[0152] Based on the similarity, a plurality of key frames are screened from the plurality of candidate key frames.
[0153] In an embodiment of the present application, each candidate key frame has a neighboring candidate key frame. If the similarity of two candidate key frames is high, there is a high probability that the same object (target) exists in the two candidate key frames, and the positions of the object are close. The distance between the target detection boxes in the neighboring candidate key frames can be determined based on the Euclidean distance between the center positions of the target detection boxes in the neighboring candidate key frames, and then the neighboring target detection box closest to the position of the target detection box in the neighboring candidate key frames can be determined. The neighboring target detection box is located in the neighboring candidate key frame. The neighboring candidate key frame is the N candidate key frames adjacent to the candidate key frame to be compared in the candidate key frame sequence included in the candidate key frame set.
[0154] After determining the adjacent target bounding boxes, the video generation device can determine intersection areas and union areas between the target bounding box and each adjacent target bounding box based on position information of the target bounding box in the corresponding candidate key frame, such as position information of four corner points of the target bounding box. The intersection area refers to an area of an overlapping part between the target bounding box and each adjacent target bounding box, and the union area refers to a total area between the target bounding box and each adjacent target bounding box minus the intersection area.
[0155] Further, the video generation device can determine the similarity between adjacent candidate key frames based on the intersection area and the union area. Specifically, the video generation device can calculate an Intersection over Union (IoU), that is, a ratio of the intersection area to the union area, based on the intersection area and the union area between the target bounding box and each adjacent target bounding box, and use the IoU to measure the overlapping degree of two bounding boxes, which can also be used to determine the similarity between two candidate key frames. The video generation device can specifically use the calculated value of the IoU as the similarity between adjacent candidate key frames, or can further process the calculated value of the IoU, such as rounding, and use the value obtained by further processing (such as the rounded IoU value) as the similarity between adjacent candidate key frames.
[0156] Finally, the video generation device can select a plurality of key frames from the plurality of candidate key frames based on the similarity. Since a larger value of the IoU indicates a higher overlapping degree of two bounding boxes, that is, a higher similarity, the video generation device can select a preset number of candidate key frames from adjacent candidate key frames in order of the calculated values of the IoU corresponding to each adjacent target bounding box from small to large, as the plurality of key frames. In this way, adjacent frames with a higher similarity can be avoided from being repeatedly selected.
[0157] In some embodiments, the video generation device can also calculate the similarity, such as the cosine similarity, between each pair of adjacent candidate key frames, and filter out a preset number of candidate key frames in order of the similarity from large to small to obtain the key frames.
[0158] In the embodiments of the present application, since each candidate key frame in the candidate key frame set is extracted from the candidate video material, adjacent video frames in the video have a relatively large similarity. Therefore, by determining the similarity between adjacent candidate key frames for filtering, candidate key frames with a relatively large similarity can be deleted, and key frames can be retained.
[0159] After obtaining the key frame, the video generation device can generate a target video corresponding to the video theme according to the key frame. Specifically, the video generation device can input the key frame into a target generation model to generate a target video corresponding to the video theme based on the target generation model. The target generation model can be a machine learning model for generating a sequence of images based on an image. For example, the target generation model can be a generative adversarial network (GAN), and can also be other machine learning models, which are not limited in the present application. That is, the video generation device can input the key frame into the target generation model to obtain a sequence of images output by the target generation model, and then generate a target video based on the sequence of images.
[0160] For the convenience of description, the embodiments of the present application take the target generation model as an example of GAN. In order to make the generated sequence of images (video frames) have high temporal consistency, the target generation model can be a temporal GAN architecture. The temporal consistency refers to the coherence and consistency of the sequence of images (i.e. video frames) in the time axis in the video.
[0161] In the process of video generation and processing, temporal consistency can be used to ensure the natural transition and logical coherence between video frames, that is, the target generation model can consider the feature consistency of adjacent video frames when generating each video frame, to ensure the coherence of the video segment, thereby avoiding abrupt changes or discordant visual effects, and improving the visual experience of the generated target video.
[0162] In some embodiments, the target generation model can be trained based on high-quality images (such as high-definition images, etc.), so that the target generation model can be used to output a sequence of high-quality images based on the key frame, which is beneficial to improve the quality of the generated target video in the video clip. It can be understood that the target generation model is used to repair and enhance low-quality key frames.
[0163] In some embodiments, after obtaining the video segment generated by the target generation model, the video generation device can enhance the video segment, and take the enhanced video segment as the target video corresponding to the video theme. The enhancement of the video segment refers to further processing and processing of the video segment to improve the video quality, enhance the visual effect or meet the specific playback and display requirements. Specifically, the video generation device can call the pre-packaged library in python or call the corresponding API to perform post-processing such as color correction, denoising, super-resolution, etc. on the generated video segment to improve the overall quality of the video segment, thereby generating the target video.
[0164] In some embodiments, the video generation device can input the key frames into the target generation model in batches to obtain video clips output by the target generation model multiple times, and then combine the video clips output multiple times to obtain the target video output finally.
[0165] Please refer to FIG. 6, which is a schematic diagram of a method for generating a target video based on candidate video materials according to an embodiment of the present application. As shown in FIG. 6, the video generation device first extracts an initial video frame set 602 from the candidate video materials 601 based on a preset frame rate. Then, the video generation device inputs each initial video frame in the initial video frame set 602 into a machine learning model 603 for scoring image quality to obtain an image quality score corresponding to each initial video frame, and selects a candidate key frame set 604 with higher scores.
[0166] Then, the video generation device can perform target detection on each candidate key frame in the candidate key frame set 604, and calculate the similarity 605 between adjacent candidate key frames based on the results of target detection (such as the position information of the target detection frame). The video generation device can filter 606 the candidate key frames with higher similarity in the candidate key frame set 606 based on the similarity 605 between adjacent candidate key frames to obtain a plurality of key frames 607. Finally, the video generation device can input the plurality of key frames 607 into a target generation model 608 to generate a target video 609 corresponding to the video theme based on the image sequence output by the target generation model 608.
[0167] Thus, the video generation device can automatically edit video materials using deep learning and video processing technology, efficiently extract and generate high-quality target videos from video materials, thereby saving the time of video editors for video editing.
[0168] In some embodiments, the video generation device can obtain the background music associated with the video theme based on the video theme, or select the most popular background music from the background music associated with the video theme, and generate the target video based on the image sequence output by the target generation model and the background music. The video generation device can also obtain the corresponding background music based on the sentiment tendency category output by the machine learning model for identifying sentiment tendency, and generate the target video based on the image sequence output by the target generation model and the background music corresponding to the sentiment tendency category.
[0169] In a possible implementation, taking a target generation model as an example, the target generation model can include a generation network and a discrimination network, the generation network, also referred to as a generator, can be used to input a key frame, and generate an image sequence, that is, a high-quality video clip, through a series of convolution and deconvolution operations. The generation network can be a generation network based on a U-net architecture, which can enhance the detail recovery capability of the generation network, thereby improving the quality of the generated video clip. The discrimination network, also referred to as a discriminator, can be used to input a real video clip and a generated pseudo video clip, and distinguish the real and fake videos through convolution operations. The discrimination network can be a discrimination network based on a PatchGAN architecture, which can improve the discrimination capability for local details, thereby improving the accuracy of discrimination.
[0170] During the training of the generation model, the generation network and the discrimination network are first initialized. Then, the training noise can be input into the generation network to output pseudo video data, and the real video data and the pseudo video data can be input into the discrimination network to output the probability that each video data is real video, and finally, based on the probability, the loss data is determined, and the parameters of the generation network and the discrimination network are updated or adjusted based on the loss data. After multiple iterations, until the training end condition is met, the target generation model is obtained. Thus, the quality of the image sequence generated by the generation network can be improved through adversarial training.
[0171] The training noise can be random noise, the pseudo video data output by the generation network is a generated fake video, and the real video data is a real video. The discrimination network can determine the probability that each video data is a real video based on the input fake video and real video, and then calculate the loss data based on the discrimination result of the discrimination network for the real video and the fake video. The loss data can be one or more of a perceptual loss data, a content loss data, and an adversarial loss data. The loss data can be used to represent the ability of the discrimination network to distinguish real videos and fake videos.
[0172] In an embodiment of the present application, the video generation device can first fix the model parameters of the generation network, adjust the parameters of the discrimination network, and then fix the model parameters of the discrimination network and adjust the parameters of the generation network. That is, in multiple iterations, in the current iteration, the parameters of the discrimination network are adjusted based on the loss data, and in the next iteration, the parameters of the generation network are adjusted based on the loss data.
[0173] In another embodiment of the present application, the video generation device can adjust the parameters of the generation network and the discrimination network at each iteration based on the loss data.
[0174] In some embodiments, the video generation device can obtain real video data, i.e., a large amount of high-quality video data. After obtaining a large amount of high-quality video data, the video generation device can also perform data augmentation on it, such as by rotating, cropping, color transforming, etc. to enhance the diversity of the video data contained in the real video data.
[0175] In some embodiments of the present application, the electronic device obtains text information of a specified type from a content publishing platform, and then performs keyword extraction processing on the text information and generates a video theme based on the extracted keywords. After the electronic device generates the video theme, the electronic device can obtain candidate video materials associated with the video theme from a second content publishing platform based on the video theme, extract key frames from the candidate video materials based on a preset frame rate, and generate a target video corresponding to the video theme according to the key frames. As can be seen, the electronic device generates a video theme and obtains video materials based on the video theme to generate a target video corresponding to the video theme. The entire video production can be completed by the electronic device, which improves the automation degree of the video production process and improves the intelligent level of content creation. It can save the time spent by the video producer in video theme screening and video editing, which is conducive to improving the efficiency of video generation and improving the timeliness of the published video to a certain extent.
[0176] The present application also provides a video generation method. The video generation method described in the embodiments of the present application can be executed by an electronic device, which can be a video generation device 101 in the video generation system shown in FIG. 1. Please refer to FIG. 7, which is another flowchart of a video generation method provided by an embodiment of the present application. The video generation method includes the following steps S701-S703:
[0177] S701, obtaining the publication time and interaction data of each video published on a content publishing platform.
[0178] In the embodiments of the present application, the content publishing platform that publishes the video can be the same as or different from the content publishing platform that obtains the text information of a specified type. For example, the content publishing platform can be used only for publishing target videos, and the number can be one or more. The present application does not limit this, and it can be part of the content publishing platforms included in the content publishing platform that obtains the text information of a specified type.
[0179] The publishing time refers to the time when the video is published in 24 hours, such as 9:00. Among them, the video producer can divide 24 hours into multiple publishing time periods, for example, 0-4 time period, 4-8 time period, 8-12 time period, 12-16 time period, 16-20 time period and 20-0 time period. The publishing time corresponding to the multiple videos can refer to the time period to which the publishing time of the multiple videos belongs.
[0180] The interaction data refers to the operation data of the user for the video in a preset time period after the video is published, such as the number of views, the number of likes, the number of comments, and the number of shares.
[0181] In some embodiments, in order to reduce the amount of data for subsequent analysis, the video generation device can only obtain the publishing time and the interaction data of the multiple videos published by the content publishing platform in a preset time period, such as 1 day, 1 week, etc.
[0182] Among them, the specific way of obtaining can be realized based on the crawler technology, and the specific implementation manner can be referred to the foregoing, which will not be described here.
[0183] S702, determining the optimal publishing time based on the publishing time and the interaction data.
[0184] In the embodiments of the present application, the video generation device can analyze the publishing time and the interaction data of the obtained multiple videos to determine the best publishing time of the target video.
[0185] In a possible implementation manner, the video generation device can count the average interaction data of the video in each publishing time period based on the divided multiple publishing time periods, such as the average value of the number of views, the average value of the number of likes, the average value of the number of shares, etc. Then, the average interaction data in each publishing time period is sorted, and the publishing time period with the maximum average interaction data is selected as the optimal publishing time.
[0186] In another possible implementation manner, the video generation device can also count the average interaction data of different types of video themes in each publishing time period, so as to determine the optimal publishing time corresponding to each type of video theme. Then, the video generation device can determine the optimal publishing time according to the type of the video theme corresponding to the target video and the optimal publishing time corresponding to each type of video theme.
[0187] In the embodiments of the present application, the video generation device can determine the optimal publishing time corresponding to each content publishing platform only once, and after determining the optimal publishing time corresponding to the content publishing platform, the optimal publishing time can be stored. In this way, the video generation device can manage the publishing plan to ensure that the target video is published at the best time.
[0188] S703, based on the optimal publishing time, calling a video publishing interface of the first content publishing platform to publish the target video.
[0189] In the embodiments of the present application, the video generation device can call the video publishing interface (API) of the content publishing platform based on the optimal publishing time corresponding to the content publishing platform to publish the generated target video. If the number of first content publishing platforms is more than one, the video generation device can call the video publishing interface (API) of each first content publishing platform respectively to publish the target video. In this way, the video generation device can automatically send the generated video to each content publishing platform, improve the automation degree of video publishing, and improve the efficiency of video publishing.
[0190] Further, the video generation device can obtain the operation of the user after publishing the target video, generate user feedback data, for example, the user feedback data can include the behavior statistics data of the user after publishing the video, such as the number of plays, likes, comments, and shares. The comment data can refer to the comment content and quantity of the user on the video. The user feedback data can also include the average viewing time of the user for the target video, the skip behavior, etc. The skip behavior refers to the behavior of the user dragging the progress bar to watch, which can represent the part of the target video that the user has watched and the part that the user has skipped. The video generation device can periodically obtain the user feedback data to ensure the real-time and accuracy of the data. The video generation device can store the obtained user feedback data in the database and update it regularly.
[0191] In some embodiments, the video generation device can store structured data in a relational database (such as MySQL) and unstructured data in a NoSQL database.
[0192] In some embodiments, the video generation device can periodically obtain the user feedback data corresponding to the target video based on the API interface provided by the content publishing platform. The video generation device can also obtain the user feedback data based on the crawler technology, which will be described in detail in the foregoing implementation manner, and will not be described here. In the embodiments of the present application, the video generation device can capture the user feedback data in the page corresponding to the target video in the content publishing platform based on the crawler technology.
[0193] Specifically, the video generation device can input the user feedback data into a machine learning model (such as a BERT or GPT model) for identifying sentiment trends, to obtain a feedback sentiment category output by the machine learning model. The user feedback data can refer to the comment data described above. The feedback sentiment category can be divided based on positive, negative, and neutral, or can be divided based on specific sentiment categories, such as happy, calm, and angry. The application does not limit this. The video generation device can periodically input updated comment data into the machine learning model for identifying sentiment trends, to obtain updated feedback sentiment categories, so that the sentiment trend of the comment data in the target video can be analyzed based on the feedback sentiment categories corresponding to different time periods, to identify the sentiment change of the user.
[0194] In some embodiments, the video generation device can also output the feedback sentiment category, so that the video producer can intuitively view the feedback sentiment category. Please refer to FIG. 8, which is a schematic diagram of a user interface 800 for displaying user feedback data according to an embodiment of the application. It should be noted that FIG. 8 is only an example, and the user interface for displaying user feedback data is not limited. As shown in FIG. 8, the user interface 800 can display the user feedback data corresponding to the video A, which can be stored in a database. The user interface 800 can include the basic information 801 of the video A, and the total feedback data 802 published in the content publishing platform, such as the total number of plays, the total number of likes, the total number of shares, and the like. The user interface can also display the user feedback data corresponding to each content publishing platform, such as the user feedback data 803 of platform A and the user feedback data 804 of platform B shown in FIG. 8, which include the comment area feedback sentiment category and the new play count statistics, respectively. The comment area feedback sentiment category can be the feedback sentiment category output by the machine learning model for identifying sentiment trends, and a statistical chart generated based on the feedback sentiment category, which can display the proportion of comments corresponding to each feedback sentiment category. The new play count statistics can be generated by the video generation device based on the updated feedback data in the database.
[0195] In some embodiments, the video generation device can construct a dashboard including the user feedback data of the video within a specified time period (such as real-time) through a visualization tool Grafana, to obtain the user interface, based on which the video producer can monitor the feedback of the user, to determine the performance of the video.
[0196] In a possible implementation, the video generation device can further acquire historical user feedback data corresponding to other videos published in the first content publishing platform, and generate a video optimization strategy corresponding to the content publishing platform based on the user feedback data, the historical user feedback data, and the feedback emotion category. The historical user feedback data can be historical feedback information of other videos published in a certain content publishing platform, such as information including a play count, a like count, a comment, a share count, and other user data (such as the average viewing time and the skip behavior described above). In some embodiments, in order to better analyze, the user feedback data can be user feedback information received within a preset time period after the video is published. For example, the user feedback data is user feedback information one week after the target video is published, and the historical user feedback data is also user feedback information one week after other videos are published.
[0197] In a possible implementation, the video generation device can use a regression analysis method to analyze the user feedback data. Regression analysis is a method of determining the quantitative relationship between two or more variables by using statistical principles. Specifically, the video generation device can perform linear regression analysis and multiple regression analysis on the user feedback data. Linear regression analysis can be used to analyze the relationship between a single variable (such as a play count) and a video feature (such as a video duration, a video theme, etc.), and multiple regression analysis can be used to analyze the relationship between multiple variables (such as a video theme, a video duration, etc.) and the overall performance of a video (such as a like count, a comment count, a play count, etc.). The video generation device can train a regression model (i.e., generate a corresponding quantitative relationship) based on the user feedback data and the historical user feedback data. In some embodiments, the video generation device can evaluate the effect of the regression model based on the user feedback data corresponding to the subsequently generated video, and continuously update the regression model (update the quantitative relationship) to optimize.
[0198] In another possible implementation, the video generation device can employ a clustering analysis method to analyze the user feedback data. Clustering analysis refers to grouping a collection of physical or abstract objects into multiple classes composed of similar objects, and analyzing each class. Among them, the video generation device can employ a k-means clustering algorithm (K-Means) clustering, a hierarchical clustering algorithm to analyze the user feedback data. K-Means clustering can be used to divide videos into different categories according to their features, such as tags, topic categories, and length, and identify similar groups for analysis. Hierarchical clustering can be used to build a hierarchy between video tags and discover hierarchical relationships between video tags for further analysis. For example, the video generation device can cluster videos based on their topic types and determine the most popular video topics in the content publishing platform. In another example, the video generation device can determine the relationship between video tags and overall video performance based on the hierarchical relationship between video tags, such as whether secondary tagging has an impact on the number of plays.
[0199] Referring to FIG. 9, FIG. 9 is a schematic diagram of a user interface 900 for displaying data analysis results according to an embodiment of the present application. As shown in FIG. 9, the user interface 900 for displaying data analysis results can include a chart based on pairwise data, such as a chart 901 showing the relationship between video length and the number of plays on the left, where the number of plays (such as the average value) corresponding to different video lengths is obtained based on linear regression analysis. The user interface 900 can also include a chart based on statistical distributions, such as a chart 902 showing the most popular video topics on the right, where the most popular video topics on the content publishing platform (Platform A) are determined based on the relationship between the overall performance of videos of each video topic (type) on the content publishing platform (Platform A). In the present application, FIG. 9 is only an example of displaying analysis results, and is not limited in this regard. The video generation device can display the analysis results as shown in FIG. 9 through data visualization tools such as Matplotlib and Tableau.
[0200] In a possible implementation, the video generation device can determine a video theme category that a user likes (feedback emotion category is positive, such as happy, joyful, and the like) based on the user feedback data, the historical user feedback data, and the feedback emotion category, and sort the video theme categories in descending order of overall performance (i.e., the number of plays, the number of comments, and the number of likes) of the videos to generate a video optimization strategy corresponding to the content publishing platform. The video generation device can assign a corresponding weight to each video category in the sorting result to generate an optimization strategy for the video theme type. The weight obtained in this analysis can also be used to adjust the weight obtained in the last analysis, that is, to adjust the content generation strategy, and then the video generation device can increase the acquisition of text information associated with a video category with a larger weight and reduce the acquisition of text information associated with a video category with a smaller weight when acquiring specified type of text information from the content publishing platform next time. In this way, the video generation device can extract valuable information from the collected user feedback data, and optimize and adjust itself to guide the optimization of subsequent content generation and publishing strategies, which is conducive to improving the relevance and popularity of the subsequently generated videos.
[0201] Please refer to FIG. 10, which is an architecture diagram of a video generation method provided by an embodiment of the present application. As shown in FIG. 10, the video generation method can include four parts: specified type of information capturing 1010, automatic clipping 1020, publishing 1030, and feedback learning 1040.
[0202] The specified type of information capturing 1010 means that the video generation device can use web crawler technology to capture data (i.e., text information) 1011 in the content publishing platform, and then use NLP technology to analyze and extract keywords and generate video themes 1012.
[0203] The automatic clipping 1020 means that the video generation device can extract video frames 1021 in the candidate video material based on a preset frame rate, and then score the extracted video frames by using CNN to select key frames 1022. Then, the video generation device can input the key frames into the GAN to generate high-quality image sequences 1023, and then select background music 1024 according to the keywords to generate a target video based on the image sequences and the background music.
[0204] The publishing 1030 means that the video generation device can call the API interface of the content publishing platform to publish the target video 1031, and the video generation device can count the best publishing time of the content publishing platform before or after publishing the target video to manage / update the publishing plan 1032.
[0205] The feedback learning 1040 refers to that after the video generation device publishes the target video, the video generation device can collect user feedback data 1041 of the published target video, and then can perform data analysis 1042 on the user feedback data by using regression analysis, cluster analysis or the like. Finally, the video generation device can generate a video optimization strategy 1043 based on the data analysis result.
[0206] In the technical solutions provided in some embodiments of the present application, the electronic device obtains text information of a specified type from a content publishing platform, and then performs keyword extraction processing on the text information and generates a video theme based on the extracted keywords. After the electronic device generates the video theme, the electronic device can obtain candidate video materials associated with the video theme from a second content publishing platform based on the video theme, extract key frames from the candidate video materials based on a preset frame rate, and generate a target video corresponding to the video theme according to the key frames. As can be seen, the electronic device generates a video theme and obtains video materials based on the video theme to generate a target video corresponding to the video theme. The entire video production can be completed by the electronic device, which improves the automation degree of the video production process and improves the intelligent level of content creation compared with manual video production. The time spent by the video producer in video theme screening and video editing can be saved, which is conducive to improving the efficiency of video generation and improving the timeliness of the published video to a certain extent.
[0207] Please refer to FIG. 11, which is a structural schematic diagram of a video generation apparatus provided in an embodiment of the present application. The video generation apparatus can be specifically the video generation device 101 shown in FIG. 1. The video generation apparatus shown in FIG. 11 can be used to perform part or all of the functions in the method embodiments described above with reference to FIG. 2 and FIG. 7. The detailed description of each unit is as follows:
[0208] The obtaining unit 1101 is configured to obtain text information of a specified type from a first content publishing platform.
[0209] The processing unit 1102 is configured to extract keywords from the text information and generate a video theme based on the keywords.
[0210] The obtaining unit 1101 is further configured to obtain candidate video materials associated with the video theme from a second content publishing platform based on the video theme.
[0211] The processing unit 1102 is further configured to extract key frames from the candidate video materials based on a preset frame rate, and generate a target video corresponding to the video theme according to the key frames.
[0212] In a possible implementation, the obtaining unit 1101 is specifically configured to:
[0213] sending a content acquisition request to a server on the first content publishing platform;
[0214] receiving a response message returned by the server in response to the content acquisition request, the response message carrying a source code document of the first content publishing platform;
[0215] extracting the text information from the source code document based on a code tag associated with the text information.
[0216] In a possible implementation, the processing unit 1102 is specifically configured to:
[0217] generating a text document according to content in the first content publishing platform;
[0218] parsing the text information to obtain a set of words contained in the text information, the set of words including a plurality of words;
[0219] determining a weight corresponding to each word based on a number of times that each word appears in the text document;
[0220] selecting a preset number of words as a plurality of the keywords in a descending order of the weights.
[0221] In a possible implementation, the processing unit 1102 is specifically configured to:
[0222] constructing a document-word matrix based on a number of times that each keyword appears in each sub-document;
[0223] parsing the document-word matrix to obtain a word distribution corresponding to each topic under a plurality of topics;
[0224] generating the video topic based on a plurality of the word distributions.
[0225] In a possible implementation, the processing unit 1102 is specifically configured to:
[0226] inputting a plurality of the keywords into a machine learning model for identifying sentiment orientation, and outputting a sentiment orientation category;
[0227] inputting the sentiment orientation category and a plurality of the word distributions into a text generation model, and outputting the video topic.
[0228] In a possible implementation, the processing unit 1102 is specifically configured to:
[0229] parsing the text information to obtain a set of words contained in the text information, the set of words including a plurality of words;
[0230] determine a correlation between each of the plurality of words based on a preset word window;
[0231] construct an undirected graph with the plurality of words as nodes and the correlation as edges, the undirected graph including an initial weight corresponding to each node;
[0232] update the initial weight corresponding to each node based on the correlation;
[0233] select a preset number of words in descending order of the updated weight as the plurality of keywords.
[0234] In a possible implementation, the obtaining unit 1101 is specifically configured to:
[0235] obtain video-related information of each content publishing platform from a plurality of content publishing platforms;
[0236] generate a video tag of each content publishing platform based on the video-related information;
[0237] select a content publishing platform corresponding to a video tag matching the video theme from a plurality of the video tags as the second content publishing platform;
[0238] obtain the candidate video material from the second content publishing platform.
[0239] In a possible implementation, the processing unit 1102 is specifically configured to:
[0240] extract a plurality of initial video frames from the candidate video material based on the preset frame rate;
[0241] input the plurality of initial video frames into a machine learning model for scoring image quality, and output an image quality score of each initial video frame;
[0242] select a preset number of video frames from the plurality of initial video frames in descending order of the image quality score as the plurality of key frames.
[0243] In a possible implementation, the processing unit 1102 is specifically configured to:
[0244] select a preset number of initial video frames from the plurality of initial video frames in the order as a plurality of candidate key frames;
[0245] perform target detection processing on each candidate key frame to obtain a target detection frame;
[0246] Based on the plurality of target detection boxes, a plurality of the key frames are screened from the plurality of candidate key frames.
[0247] In a possible implementation, the processing unit 1102 is specifically configured to:
[0248] For each target detection box,
[0249] From the plurality of target detection boxes, a nearest neighbor target detection box to the position of the target detection box is determined;
[0250] The intersection area and the union area between the target detection box and the neighbor target detection box are calculated;
[0251] Based on the intersection area and the union area, a similarity between the candidate key frame corresponding to the target detection box and the candidate key frame corresponding to the neighbor target detection box is determined;
[0252] Based on the similarity, a plurality of the key frames are screened from the plurality of candidate key frames.
[0253] In a possible implementation, the target video is generated by a target generation model according to the key frames, the target generation model includes a generation network and a discrimination network; the video generation apparatus 110 further includes:
[0254] The input unit 1103 is configured to input training noise to the generation network to output pseudo video data; and input real video data and the pseudo video data into the discrimination network to output a probability that each video data in the pseudo video data and the real video data is real video;
[0255] The determination unit 1104 is configured to calculate loss data based on the probability; and train the discrimination network and the generation network based on the loss data to obtain the target generation model.
[0256] In a possible implementation, the determination unit 1104 is specifically configured to:
[0257] In a plurality of iterations,
[0258] In the current iteration, the parameters of the discrimination network are adjusted based on the loss data;
[0259] In the next iteration, the parameters of the generation network are adjusted based on the loss data.
[0260] In a possible implementation, the acquisition unit 1101 is further configured to acquire a plurality of videos each corresponding to a publishing time and interaction data on the first content publishing platform;
[0261] The determination unit 1104 is further configured to determine an optimal publishing time based on the publishing time and the interaction data.
[0262] The publishing unit 1105 is configured to call a video publishing interface of the first content publishing platform based on the optimal publishing time, and publish the target video.
[0263] In a possible implementation, the obtaining unit 1101 is further configured to obtain user feedback data after the target video is published; and input the user feedback data into a machine learning model for identifying sentiment tendency, and output a feedback sentiment category.
[0264] The processing unit 1102 is further configured to generate a video optimization strategy based on the feedback sentiment category, the user feedback data, and historical user feedback data of other videos.
[0265] The units in the video generation apparatus shown in FIG. 11 can be respectively or wholly combined into one or more other units to constitute, or some of the units can be further split into a plurality of units with smaller functions to constitute, which can realize the same operation without affecting the realization of the technical effects of the embodiments of the present application. The units are divided based on logical functions, and in actual application, the function of one unit can also be realized by a plurality of units, or the functions of a plurality of units are realized by one unit. In other embodiments of the present application, the video generation apparatus can also include other units, and in actual application, these functions can also be realized by other units, and can be realized by a plurality of units in cooperation.
[0266] According to another embodiment of the present application, the video generation apparatus shown in FIG. 11 and the video generation method of the embodiments of the present application can be constructed and realized by running a computer program (including program codes) capable of executing each step involved in the corresponding method shown in FIG. 2 and FIG. 7 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM).
[0267] Based on the same inventive concept, the video generation apparatus provided in the embodiments of the present application has similar principles and beneficial effects to those of the video generation apparatus in the method embodiments for solving problems, and the principles and beneficial effects of the method embodiments can be referred to. For brevity, the description is not repeated here.
[0268] Please refer to FIG. 12, which is a structural schematic diagram of an electronic device provided in an embodiment of the present application. The electronic device 120 at least includes a processor 1201, an input device 1202, an output device 1203 and a memory 1204. The processor 1201, the input device 1202, the output device 1203 and the memory 1204 in the electronic device 120 are connected through a bus or other means.
[0269] The memory 1204 is a memory device in the electronic device 120, and is used to store programs and data. In the embodiment of the present application, the memory 1204 can include an internal storage medium of the webpage data encryption processing device, and of course can also include an extended storage medium supported by the electronic device 120. The memory 1204 provides a storage space, and the storage space stores an operating system of the electronic device 120. In addition, the storage space also stores a computer program (including program codes). It should be noted that the computer storage medium can be a high-speed RAM memory, and optionally can be at least one computer storage medium away from the processor. The processor can be referred to as a central processing unit (CPU), which is the core and control center of the webpage data encryption processing device, and is used to run the computer program stored in the memory 1204.
[0270] In an embodiment, the computer program stored in the memory 1204 can be loaded and executed by the processor 1201 to implement the corresponding steps of the method in the above method embodiments.
[0271] It should be understood that, in the embodiment of the present application, the processor 1001 can be a central processing unit (CPU), and the processor 1001 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0272] The embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program includes program instructions, and the program instructions are executed by a processor to perform the steps performed in all the above embodiments.
[0273] The embodiments of the present application further provide a computer program product or computer program, which comprises computer instructions stored in a computer readable storage medium, and the computer instructions are executed by a processor of a computer device to perform the method in any of the above embodiments.
[0274] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the above-mentioned program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.
[0275] The above only discloses a preferred embodiment of the present application, and of course cannot limit the scope of the present application. Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be implemented, and equivalent changes made according to the claims of the present application still fall within the scope of the present application.
[0276] In addition, it is particularly necessary to point out that when the above embodiments of the present application are applied to specific products or technologies, if it is necessary to obtain the data of a user, the permission or consent of the user is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in relevant countries and regions.
Claims
1. A video generation method, executed by an electronic device, comprising: obtaining text information of a specified type from a first content publishing platform; extracting keywords from the text information, and generating a video theme based on the keywords; obtaining candidate video materials associated with the video theme from a second content publishing platform based on the video theme; and extracting key frames from the candidate video materials based on a preset frame rate, and generating a target video corresponding to the video theme according to the key frames. The obtaining of the text information of the specified type from the first content publishing platform comprises:
2. The method of claim 1, wherein, sending a content obtaining request to a server on the first content publishing platform; receiving a response message returned by the server in response to the content obtaining request, the response message carrying a source code document of the first content publishing platform; extracting the text information from the source code document based on a code tag associated with the text information. The extraction of the keywords from the text information comprises:
3. The method of claim 1 or 2, wherein, generating a text document according to content in the first content publishing platform; parsing the text information to obtain a set of words contained in the text information, the set of words including a plurality of words; determining a weight corresponding to each word in the plurality of words based on a number of times each word appears in the text document; selecting a preset number of words as the plurality of keywords in descending order of the weights. The text document includes a plurality of sub-documents, and the generation of the video theme based on the keywords comprises:
4. The method of claim 3, wherein, constructing a document-word matrix based on a number of times each keyword appears in each sub-document; parsing the document-word matrix to obtain a word distribution corresponding to each theme under a plurality of themes; generating the video theme based on the plurality of word distributions. The generation of the video theme based on the plurality of word distributions comprises:
5. The method of claim 4, wherein, inputting the plurality of keywords into a machine learning model for identifying sentiment orientation, and outputting a sentiment orientation category; inputting the sentiment orientation category and the plurality of word distributions into a text generation model, and outputting the video theme. The extraction of the keywords from the text information comprises:
6. The method of claim 1 or 2, wherein, parsing the text information to obtain a set of words contained in the text information, the set of words including a plurality of words; determining an association relationship between each word in the plurality of words based on a preset word window; constructing an undirected graph by taking the plurality of words as nodes and the association relationship as edges, the undirected graph including an initial weight corresponding to each node; updating the initial weight corresponding to each node based on the association relationship; selecting a preset number of words as the plurality of keywords in descending order of the updated weights. The obtaining of the candidate video materials associated with the video theme from the second content publishing platform based on the video theme comprises:
7. The method of any one of claims 1-6, wherein, obtaining video-related information of each content publishing platform from a plurality of content publishing platforms; generating a video tag of each content publishing platform based on the video-related information; and From the multiple video tags, a content publishing platform corresponding to a video tag matching the video theme is selected as the second content publishing platform; From the second content publishing platform, the candidate video material is obtained.
8. The method of any one of claims 1-7, wherein, The extracting key frames from the candidate video material based on the preset frame rate comprises: Based on the preset frame rate, multiple initial video frames are extracted from the candidate video material; The multiple initial video frames are input into a machine learning model for scoring image quality, and an image quality score of each initial video frame is output; In the order from high to low of the image quality scores, a preset number of video frames are selected from the multiple initial video frames as the multiple key frames.
9. The method of claim 8, wherein, The selecting, in the order from high to low of the image quality scores, a preset number of video frames from the multiple initial video frames as the multiple key frames comprises: In the order, a preset number of initial video frames are selected from the multiple initial video frames as multiple candidate key frames; Each candidate key frame is subjected to target detection processing to obtain a target detection box; Based on the multiple target detection boxes, the multiple candidate key frames are filtered to obtain the multiple key frames.
10. The method of claim 9, wherein, The filtering, based on the multiple target detection boxes, the multiple candidate key frames to obtain the multiple key frames comprises: For each target detection box, From the multiple target detection boxes, a nearest neighbor target detection box to the position of the target detection box is determined; An intersection area and a union area between the target detection box and the neighbor target detection box are calculated; Based on the intersection area and the union area, a similarity between a candidate key frame corresponding to the target detection box and a candidate key frame corresponding to the neighbor target detection box is determined; Based on the similarity, the multiple candidate key frames are filtered to obtain the multiple key frames.
11. The method of any one of claims 1-10, wherein, The target video is generated by a target generation model based on the key frames, and the target generation model comprises a generation network and a discrimination network. The method further comprises: A training noise is input into the generation network to output pseudo video data; Real video data and the pseudo video data are input into the discrimination network to output a probability that each video data in the pseudo video data and the real video data is real video; Based on the probability, loss data is calculated; Based on the loss data, the discrimination network and the generation network are trained to obtain the target generation model.
12. The method of claim 11, wherein, The training, based on the loss data, the discrimination network and the generation network to obtain the target generation model comprises: In multiple iterations, In the current iteration, based on the loss data, parameters of the discrimination network are adjusted; In the next iteration, based on the loss data, parameters of the generation network are adjusted.
13. The method of any one of claims 1-12, further comprising: Obtaining a respective publishing time and interaction data of multiple videos that have been published on the first content publishing platform; Based on the publishing time and the interaction data, determining an optimal publishing time; Based on the optimal publishing time, a video publishing interface of the first content publishing platform is called to publish the target video.
14. The method of claim 13, further comprising: obtaining user feedback data after publishing the target video; inputting the user feedback data into a machine learning model for identifying sentiment tendency to output a feedback sentiment category; generating a video optimization strategy based on the feedback sentiment category, the user feedback data, and historical user feedback data of other videos.
15. A video generation apparatus, comprising: an obtaining unit configured to obtain text information of a specified type from a first content publishing platform; a processing unit configured to extract a keyword from the text information and generate a video theme based on the keyword; the obtaining unit is further configured to obtain candidate video materials associated with the video theme from a second content publishing platform based on the video theme; and the processing unit is further configured to extract key frames from the candidate video materials based on a preset frame rate and generate a target video corresponding to the video theme according to the key frames.
16. An electronic device, comprising: at least one processor; a memory configured to store at least one computer program, which, when executed by the at least one processor, causes the electronic device to implement the video generation method of any one of claims 1-14.
17. A computer readable medium having stored thereon a computer program, which, when executed by a processor, implements the video generation method of any one of claims 1-14.
18. A computer program product comprising a computer program stored in a computer readable storage medium, which, when read and executed by a processor of an electronic device, causes the electronic device to perform the video generation method of any one of claims 1-14.
Citation Information
Patent Citations
Information processing method and device, electronic equipment, storage medium and program product
CN114491149A
Video generation method and device, and medium
CN114501076A
Video generation method and device, electronic equipment and storage medium
CN117676277A
Theme extraction method and device, related equipment and computer program product
CN118332103A
Generating theme-based videos
US20180068019A1