Content generation method and device
By mapping the labels of the input content with the multimodal knowledge graph, the timeline for video generation is constructed, which solves the problem of low efficiency of traditional video generation methods and realizes efficient and automated video content generation.
Patent Information
- Application Number
- CN202510273934.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
Smart Images

Figure CN120201143A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of multimedia technologies, and in particular, to a content generation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] Currently, the generation of video content generally involves manual collection and editing of materials, and finally forms the video content desired by users. Since the acquisition and editing of materials both take a lot of time, this method of generating video content has low efficiency.
[0003] With the rapid development of digital media technologies, the demand for content creation is increasing, and it is becoming increasingly difficult to meet the needs of content creation using traditional video content generation methods.
[0004] It should be noted that the above content is not necessarily prior art and does not limit the patent protection scope of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a content generation method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above-mentioned technical problems.
[0006] One aspect of the embodiments of the present application provides a content generation method, and the method includes: Obtain input content, where the input content includes at least one of text, audio, image, or video; Extract content tags based on the input content; Map the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes; Construct a timeline for content generation based on the mapped nodes; Generate target content based on the mapped nodes and the timeline.
[0007] Optionally, the constructing a timeline for content generation based on the mapped nodes includes: Obtain associated nodes connected to the mapped nodes and entity relationships with the associated nodes; Determine time information and / or causal relationships corresponding to the mapped nodes based on the associated nodes and the entity relationships; Construct a timeline for content generation based on the time information and the causal relationships.
[0008] Optionally, the target content includes voiceover audio; Correspondingly, the generating target content based on the mapped nodes and the timeline includes: Obtain relevant semantic relationships from the multimodal knowledge graph based on the mapping nodes; Generate a voice commentary text based on the semantic relationships, the content tags, and the timeline; Convert the voice commentary text into a voice commentary audio by using a text-to-speech conversion method.
[0009] Optionally, the generation of the target content based on the mapping nodes and the timeline includes: Obtain multimedia materials matching the content tags from the multimodal knowledge graph based on the mapping nodes, where the multimedia materials include video clips, pictures, and audio; Insert and edit the multimedia materials based on the timeline to generate the target content.
[0010] Optionally, the extraction of the content tags based on the input content includes: When the input content includes input text, use natural language processing methods to extract the first text tags of the input text, where the text tags include at least partial information of events, people, locations, and emotions; When the input content includes input audio, convert the input audio into text by using a speech recognition method, use natural language processing methods to extract the second text tags of the converted text, and use a speech emotion analysis method to obtain the emotion tags of the input audio; When the input content includes input images and / or input videos, use computer vision methods to extract the video image tags of the input images and / or the input videos.
[0011] Optionally, the method further includes: Obtain the target format parameters of the target platform, where the target format parameters include resolution, frame rate, and duration; Process the target content by using a video processing tool to obtain a video that conforms to the target format parameters.
[0012] Optionally, the method further includes: When authorized by the current user, obtain the target knowledge graph information corresponding to the current user; Correspondingly, the generation of the target content based on the mapping nodes and the timeline further includes: Generate the target content based on the target knowledge graph information, the mapping nodes, and the timeline.
[0013] Another aspect of the embodiments of the present application provides a content generation device, and the device includes: An acquisition module for acquiring input content, where the input content includes at least one of text, audio, image, or video; An extraction module for extracting content tags based on the input content; A mapping module for mapping the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes; A construction module for constructing a timeline for content generation based on the mapped nodes; A generation module for generating target content based on the mapped nodes and the timeline.
[0014] Another aspect of the embodiments of the present application provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method described above is implemented.
[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: By acquiring input content, extracting content tags based on the input content, mapping the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes, constructing a timeline for content generation based on the mapped nodes, and generating target content based on the mapped nodes and the timeline, it is possible to quickly obtain matching materials through the mapping of the content tags of the input content to the multimodal knowledge graph for video generation, thereby improving the efficiency of video generation and meeting the increasing demand for content creation; at the same time, by constructing a timeline for content generation through the mapped nodes of the knowledge graph and generating videos according to the constructed timeline, the dynamic construction ability of the timeline can be realized, and the degree of automation of video generation can be improved. Description of the Drawings
[0018] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0019] Figure 1 Schematically shows the operating environment diagram of the content generation method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the content generation method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 the sub-step flowchart of step S202 in; Figure 4 Schematically shows Figure 2 the sub-step flowchart of step S206 in; Figure 5 Schematically shows Figure 2 the sub-step flowchart of step S208 in; Figure 6 Schematically shows Figure 2 another sub-step flowchart of step S208 in; Figure 7 Schematically shows the new flowchart of the content generation method according to Embodiment 1 of the present application; Figure 8 Schematically shows the block diagram of the content generation device according to Embodiment 2 of the present application; and Figure 9 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the protection scope of the present application.
[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0022] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of execution of the steps. They are only used to conveniently describe this application and distinguish each step, and therefore should not be construed as a limitation to this application.
[0023] First, provide the following explanations of the terms involved in this application: Knowledge graph: It is a data model based on a graph structure, used to represent entities (such as people, things, events, etc.) and their complex relationships in the real world. Its core feature is that entities are represented by nodes, and the logical associations between entities are described by edges, thus forming a semantic network.
[0024] Multimodal: It refers to multiple forms of representation of data or information, that is, an information can exist in multiple different forms such as text, image, audio, video, etc. A modality, simply speaking, is a form of representation of data or information. Multimodal means that an information or data can have multiple different forms of representation, and these forms can exist alone or in combination.
[0025] Natural Language Processing (NLP): It is an interdisciplinary field that combines linguistics, computer science, and artificial intelligence. Its core goal is to enable computers to understand, interpret, and generate human natural language.
[0026] Dynamic programming algorithm: It is an algorithm design technique widely used in multiple fields such as mathematics, computer science, and economics. It solves complex problems by decomposing the original problem into relatively simple sub-problems and uses the solutions of these sub-problems to construct the solution of the original problem.
[0027] Topological sorting algorithm: It is an algorithm for directed acyclic graphs. It arranges all vertices in a directed acyclic graph into a linear sequence such that for each directed edge (u, v) in the graph, vertex u appears before vertex v in the linear sequence. Such a linear sequence is called a sequence that satisfies the topological order, simply referred to as a topological sequence. The purpose of topological sorting is to obtain a linear sequence that satisfies all dependency relationships.
[0028] Speech emotion analysis: It refers to identifying the emotional state of the speaker, such as happiness, anger, sorrow, etc., by analyzing the speech signal.
[0029] Text-to-speech conversion: It is a technology that converts written text information into natural and fluent speech output.
[0030] Secondly, to facilitate the understanding of the technical solutions provided by the embodiments of this application by those skilled in the art, the related technologies are described below: With the rapid development of digital media technology, the demand for content creation is increasing. However, traditional video content generation methods require a large amount of time in the processes of material acquisition, material editing, etc., with low efficiency, and it is increasingly difficult to meet the growing demand for content creation. Moreover, when generating videos, related video content generation methods all rely on manual setting of the video content generation order and lack the ability of dynamic construction.
[0031] Therefore, the embodiments of this application provide a content generation technical solution. In this technical solution, by extracting the content tags of the user input content and obtaining the corresponding content for video generation according to the mapping of the content tags and a pre-constructed multi-modal knowledge graph, the efficiency of video generation can be improved, thus meeting the growing demand for content creation. At the same time, by constructing the timeline of content generation through the mapping nodes of the knowledge graph and generating videos according to the constructed timeline, the dynamic construction ability of the timeline can be realized. See the following for details.
[0032] Finally, for ease of understanding, an exemplary operating environment is provided below.
[0033] As Figure 1 shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where: The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated application programs, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.
[0034] The service platform 2 can be configured to communicate with the client 6, etc. through the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and combinations thereof, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0035] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing services for writing video materials to the client and generating videos based on the video materials.
[0036] The client 6 can be an electronic device running operating systems such as Windows, Android™, or iOS, such as a smartphone, tablet device, laptop computer, virtual reality device, gaming device, set-top box, in-vehicle terminal, or smart TV. Based on the above operating systems, various applications can be run, such as an application for running and uploading video materials to generate videos.
[0037] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading objects, etc.
[0038] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.
[0039] The technical solutions of the present application will be introduced through multiple embodiments below. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.
[0040] Embodiment 1 Figure 2 A flowchart of a content generation method according to Embodiment 1 of the present application is schematically shown. It should be noted that the execution subject of the content generation method in the embodiments of the present application can be the service platform.
[0041] As Figure 2 shown, the content generation method may include steps S200 to S208, where: Step S200: Obtain input content, where the input content includes at least one of text, audio, image, or video.
[0042] Step S202: Extract content tags based on the input content.
[0043] Step S204: Map the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes.
[0044] Step S206: Construct a timeline for content generation based on the mapped nodes.
[0045] Step S208: Generate target content based on the mapped nodes and the timeline.
[0046] The content generation method provided in this embodiment obtains input content, extracts content tags based on the input content, maps the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes, constructs a timeline for content generation based on the mapped nodes, and generates target content based on the mapped nodes and the timeline. It can quickly obtain matching materials through the mapping of the content tags of the input content to the multimodal knowledge graph to generate videos, thereby improving the efficiency of video generation and meeting the increasing demand for content creation. At the same time, by constructing a timeline for content generation based on the mapped nodes of the knowledge graph and generating videos according to the constructed timeline, the dynamic construction ability of the timeline can be realized, and the automation degree of video generation can be improved.
[0047] The following combines Figure 2 , and elaborates on each step in steps S200 to S208 and other optional steps in detail.
[0048] Step S200 , obtain input content, where the input content includes at least one of text, audio, image, or video.
[0049] For input content in text form, it can specifically be descriptive text of the generated content. For example, it can be a descriptive text about a certain tourist destination. Optionally, the input content can also be requirement description text of the generated content, such as "Please generate a promotional video for XX tourist destination". For input content in audio form, it can be descriptive audio of the generated content or requirement description audio of the generated content. For input content in image form, it can be an image description that is expected to be included in the generated content. For example, it can be images corresponding to some key scenes. For input content in video form, it can be a video description that is expected to be included in the generated content, such as videos of some core content.
[0050] Step S202 , extract content tags based on the input content.
[0051] Specifically, different processing methods can be adopted according to the different forms of the input content to extract content tags related to the input content, so as to facilitate subsequent mapping to the nodes of the multimodal knowledge graph.
[0052] In an optional embodiment, as Figure 3 shown, step S202 can specifically include: Step S600, when the input content includes input text, use natural language processing methods to extract the first text tags of the input text, and the text tags include at least some information of events, people, locations, and emotions.
[0053] Step S602, when the input content includes input audio, convert the input audio into text by using a speech recognition method, extract the second text tags of the converted text by using a natural language processing method, and obtain the emotion tags of the input audio by using a speech emotion analysis method.
[0054] Step S604, when the input content includes input images and / or input videos, extract the video image tags of the input images and / or input videos by using a computer vision method.
[0055] For example, if the input content includes the input text "Paris is the capital of romance, and the Eiffel Tower and the banks of the Seine are must-visit places.", then the natural language processing method can be used to extract the locations "Paris, Eiffel Tower, Seine River" and the emotion "romance" in the input text as the first text tags.
[0056] For another example, if the input content is input audio, and the corresponding audio content is: "Paris is the capital of romance, and the Eiffel Tower and the banks of the Seine are must-visit places.", then the input audio can be first converted into text by using a speech recognition method, and then the natural language processing method can be used to extract the locations "Paris, Eiffel Tower, Seine River" and the emotion "romance" as the second text tags; at the same time, the speech emotion analysis method can be used to obtain the emotion tags of the input audio, such as emotion tags like "pleased" and "cheerful".
[0057] For another example, if the input content is an input image of the Eiffel Tower, the computer vision method can be used to extract the video image tag "Eiffel Tower"; or, if the input content is an input video of Paris, the computer vision method can also be used to extract the video image tag "Paris".
[0058] Among them, the computer vision method can include technologies such as object detection, scene classification, and person recognition. It can be understood that the input content can include both input text and input video at the same time. When the input content includes both input text and input video at the same time, when extracting the video image tags of the video, elements related to the input text description can be extracted as the video image tags, so that the content extracted from the video is consistent with the text description; of course, the video image tags of the video can also be extracted independently.
[0059] In this embodiment, by processing the input content through methods such as natural language processing, speech recognition, speech emotion analysis, and computer vision, corresponding processing can be performed on different forms of input content, and content tags related to the input can be effectively extracted.
[0060] Step S204 , map the content tags to the nodes in the pre-constructed multimodal knowledge graph to obtain the mapped nodes.
[0061] The construction of a multi-modal knowledge graph can be as follows: First, collect multi-modal data such as text, images, audio, and video from multiple sources. For example, obtain multi-modal data from news websites, social media, academic databases, etc. Then, the collected data can be preprocessed. For example, the text can be preprocessed by word segmentation and part-of-speech tagging, the images can be preprocessed by resizing and noise reduction, and the audio can be preprocessed by format conversion. After preprocessing the collected data, entities are identified from each modal data, such as names of people and places in the text, and people in the images, and then associations between entities are established. Analyze the data of different modalities to determine the relationships between entities, such as causal relationships, temporal relationships, and spatial relationships between entities. Finally, a unified knowledge representation is formed. In a multi-modal knowledge graph, the nodes of the knowledge graph represent entities or concepts, while the edges represent the semantic relationships between entities. In practical applications, an open-source graph database can be used to store and query the multi-modal knowledge graph. Among them, the open-source graph database can be GraphDB, for example.
[0062] Specifically, after extracting the content tags of the input content, the similarity between the content tags and the nodes in the multi-modal knowledge graph can be calculated. When the similarity reaches a preset threshold, the corresponding node is determined as the mapping node. For example, if the content tag is "Word A", "Word A" can be converted into a vector, and then the cosine similarity between "Word A" and the nodes in the multi-modal knowledge graph is calculated. If the calculated result shows that the cosine similarity between "Word A" and "Node B" is greater than the preset threshold, then "Node B" is used as the mapping node.
[0063] Step S206 , construct a timeline for content generation based on the mapping node.
[0064] After determining the mapping nodes of the content tags, the causal relationships and temporal information of the events can be obtained according to the connected nodes of the mapping nodes and the relationships with the connected nodes, and a timeline for content generation can be constructed based on these causal relationships and temporal information.
[0065] In an alternative embodiment, as Figure 4 shown, step S206 may include: Step S300, obtain the associated nodes connected to the mapping node and the entity relationships with the associated nodes.
[0066] Step S302, determine the temporal information and / or causal relationships corresponding to the mapping node based on the associated nodes and entity relationships.
[0067] Step S304, construct a timeline for content generation based on the temporal information and causal relationships.
[0068] Among them, the associated node can be a node directly or indirectly connected to the mapping node.
[0069] Specifically, after determining the mapping node, the associated nodes can be obtained according to the connection relationship between the mapping node and other nodes, and the entity relationship with the associated nodes is determined through the edges between the nodes; the time information and / or causal relationship corresponding to the mapping node are obtained through the relevant information of the associated nodes and their entity relationships; finally, the timeline for content generation is constructed based on the obtained time information and causal relationship.
[0070] For example, if the mapping node "Eiffel Tower" is mapped according to the content label, the node connected to the "Eiffel Tower" node is "1889", and their entity relationship is "construction time", then the time information of the "Eiffel Tower" can be determined as the construction time is 1889 based on the associated node and the entity relationship. If the "Tokyo Tower" is also mapped according to the content label, then the construction time of the "Tokyo Tower" can also be determined as 1958 based on the associated node and the entity relationship. In this way, if the finally generated video is a video including famous buildings around the world, the timeline for the corresponding part can be determined as "Eiffel Tower" → "Tokyo Tower".
[0071] Another example, if the causal relationship obtained according to the entity relationship is that "Internet" is a prerequisite for "e-commerce", then the timeline can be determined as: Internet → e-commerce.
[0072] When constructing the timeline for content generation, algorithms such as dynamic programming algorithm and topological sorting can be used to automatically sort and generate the timeline. Among them, the dynamic programming algorithm decomposes the complex timeline construction problem into a series of interrelated sub-problems, and uses the optimal solutions of the sub-problems to construct the global optimal timeline. When processing time information, fully consider the order and dependency of the content elements corresponding to different time nodes, and through the analysis and solution of each sub-problem, find the best time arrangement plan. The topological sorting algorithm focuses on processing causal relationships, and can sort the content elements according to these causal relationships to determine their order on the timeline, ensuring that the dependency and order are not disrupted.
[0073] In this embodiment, by obtaining the associated nodes connected to the mapping node and the entity relationships with the associated nodes, determining the time information and / or causal relationship corresponding to the mapping node based on the associated nodes and entity relationships, and constructing the timeline for content generation based on the time information and causal relationship, the timeline included in the generated content can be automatically determined, realizing the ability of dynamic construction of the timeline for content generation.
[0074] Step S208 , generate the target content based on the mapping node and the timeline.
[0075] Specifically, relevant materials or elements generated from the multimodal knowledge graph can be obtained according to the mapping nodes, and then the target content can be generated according to the constructed timeline. Among them, the target content can include videos, images, audio, and subtitles.
[0076] In an alternative embodiment, as Figure 5 shown, the target content includes voice-over audio, and step S208 may include: Step S400, obtaining relevant semantic relationships from the multimodal knowledge graph based on the mapping nodes.
[0077] Step S402, generating voice-over text based on the semantic relationships, content tags, and timeline.
[0078] Step S404, converting the voice-over text into voice-over audio using a text-to-speech conversion method.
[0079] Specifically, the mapping node can be used as the core to obtain semantic relationships related to the mapping node from the multimodal knowledge graph. Then, based on the obtained semantic relationships, content tags, and timeline, generative artificial intelligence is used to generate voice-over text. Finally, the voice-over text is converted into voice-over audio using text-to-speech technology (TTS).
[0080] For example, if the mapping node is "Battle of Chibi", the relevant semantic relationships obtained from the multimodal knowledge graph may be: "Cao Cao was the initiator of the Battle of Chibi", "Liu Bei and Sun Quan joined forces to fight against Cao Cao in the Battle of Chibi", "The Battle of Chibi took place in the area of Chibi on the Yangtze River", etc. If the content tags are "War background", "War process", and "War result", then based on these semantic relationships, content tags, and timeline, generative artificial intelligence can be used to generate voice-over text about the Battle of Chibi. Finally, conversion using TTS can obtain voice-over audio about the Battle of Chibi.
[0081] In the case where the target content includes voice-over audio, the voice-over text can be directly used as the subtitle in the target content. In practical applications, the target content may not include voice-over audio. In this case, the target content can also include independent subtitles. In this situation, the subtitles can be generated using generative artificial intelligence based on the keywords, emotional information of the input content, and the constructed timeline. Generating subtitles based on emotional information (such as tension) can enhance the fit with the content.
[0082] In this embodiment, by obtaining relevant semantic relationships from the multimodal knowledge graph based on mapping nodes, generating speech commentary text based on the semantic relationships, content tags, and timeline, and using a text-to-speech conversion method to convert the speech commentary text into speech commentary audio, it is possible to automatically and accurately generate speech commentary audio related to the video, improving the intelligence and accuracy of content generation.
[0083] In an alternative embodiment, as Figure 6 shown, step S208 may further include: Step S500, obtaining multimedia materials that match the content tags from the multimodal knowledge graph based on mapping nodes, where the multimedia materials include video clips, pictures, and audio.
[0084] Step S502, inserting and editing the multimedia materials based on the timeline to generate the target content.
[0085] Among them, the audio can be the audio corresponding to background music, voice, or special sound effects.
[0086] For example, if the content tags are "Eiffel Tower" and "Tokyo Tower", then based on the mapping nodes, video clips, pictures, and related audio of the "Eiffel Tower" and "Tokyo Tower" can be obtained from the multimodal knowledge graph. Then, according to the previously constructed timeline "Eiffel Tower → Tokyo Tower", the multimedia materials such as video clips, pictures, and related audio of the "Eiffel Tower" and "Tokyo Tower" are inserted and edited successively, and finally the target content is generated.
[0087] In this embodiment, by obtaining multimedia materials that match the content tags from the multimodal knowledge graph based on mapping nodes, inserting and editing the multimedia materials based on the timeline to generate the target content, it is possible to automatically obtain relevant materials from the multimodal knowledge graph for content generation, enabling the target content to include various modal information and improving the efficiency of content generation.
[0088] In an alternative embodiment, as Figure 7 shown, the content generation method of the embodiment of the present application may further include: Step S700, obtaining target format parameters of the target platform, where the target format parameters include resolution, frame rate, and duration.
[0089] Step S702, processing the target content using a video processing tool to obtain a video that conforms to the target format parameters.
[0090] Specifically, the target format parameters of the target platform can be obtained from the official documentation or platform description of the target platform, and then a video processing tool such as FFmpeg is used to process the target content to obtain a video that conforms to the target format parameters.
[0091] In this embodiment, by obtaining the target format parameters of the target platform and using a video processing tool to process the target content to obtain a video that conforms to the target format parameters, videos with different format parameters can be generated according to the format parameter requirements of different platforms, so as to ensure that the generated videos can be played smoothly on each platform.
[0092] In an alternative embodiment, the content generation method of the embodiments of the present application may further include: obtaining target knowledge graph information corresponding to the current user under the authorization of the current user; in step S208, when generating the target content based on the mapping nodes and the timeline, it may further include: generating the target content based on the target knowledge graph information, the mapping nodes and the timeline.
[0093] The current user is also the user corresponding to the input content. The target knowledge graph information may be information in a multimodal knowledge graph or knowledge graph information constructed based on the user. It should be noted that the knowledge graph information constructed based on the user is constructed under the authorization of the user.
[0094] Specifically, under the authorization of the current user, relevant information of the user can be obtained from the knowledge graph as the target knowledge graph information according to the relevant information of the user, such as the user's video style, preferred video category, hobbies, education level, etc.; when generating the target content, the obtained target knowledge graph information is combined to generate the target content. For example, if the user's video style is an anime style, a method of style transfer can be used to generate target content with an anime style.
[0095] In this embodiment, by obtaining the target knowledge graph information corresponding to the current user under the authorization of the current user and generating the target content based on the target knowledge graph information, the mapping nodes and the timeline, the generated target content can conform to the characteristics of the current user, and the personalization of content generation can be improved.
[0096] To make the present application easier to understand, the following provides an exemplary application for automatically generating a travel promotional video.
[0097] 1. Input: A descriptive article (such as text) about a tourist destination.
[0098] For example: "Paris is the capital of romance, and the Eiffel Tower and the banks of the Seine are must-visit places." 2. Step 1: Input data processing and label extraction: Analyze the input text and use natural language processing techniques (such as named entity recognition, sentiment analysis) to extract key labels: (1) Location tags: Paris, Eiffel Tower, Seine River.
[0099] (2) Emotion tag: Romantic.
[0100] (3) Time information: Historical background (e.g., 19th century, modernization).
[0101] 3. Step 2: Mapping tags to knowledge graph nodes: (1) "Paris" is mapped to the location node, associated with historical and cultural background, romantic tag, etc.
[0102] (2) "Eiffel Tower" is mapped to the landmark node, associated with the construction year (1889) and the modernization process of Paris.
[0103] (3) "Seine River" is mapped to the scenic spot node, associated with tags such as tourism and romance.
[0104] (4) "Romantic" is mapped to the emotion tag node, associated with the cultural background of Paris and the emotional tendencies of tourists.
[0105] Through these mappings, ensure that each tag is correctly corresponding to the entities and emotion information in the knowledge graph.
[0106] 4. Step 3: Constructing a timeline based on the knowledge graph: When constructing the timeline, infer the order and logic of events according to the mapped tags and node relationships. For example: (1) Paris becomes the capital of romance: Through the association between Paris and the romantic tag in the knowledge graph, it is determined that this is the cultural image gradually established by Paris in the 19th century.
[0107] (2) The Eiffel Tower is completed (in 1889): As a symbol of Paris's romantic image, the completion time of the Eiffel Tower is a key node at this time.
[0108] (3) The Seine River becomes a famous scenic spot: After the completion of the Eiffel Tower, the Seine River becomes a popular romantic scenic spot for tourists.
[0109] The constructed timeline is: ① Paris becomes the capital of romance (early 19th century) ② The Eiffel Tower is completed (in 1889) ③ The Seine River becomes a famous scenic spot (from the late 19th century to the early 20th century) 5. Step 4: Content generation and matching: Based on the timeline, automatically select and generate multimedia content related to each time node: (1) Video generation and editing: The system searches for matching video materials through the knowledge graph. For example, when "Paris" is entered, the system retrieves videos of scenic spots related to Paris; when "Eiffel Tower" is entered, relevant landmark videos are retrieved. All video materials will be inserted and edited in the order of the timeline to ensure that each video clip is displayed according to the time node.
[0110] (2) Narrative audio generation: The system uses TTS technology to generate narration audio based on the structured information in the input text. For example: “Paris has become a romantic capital and a symbol of the global cultural and artistic center…” "In 1889, the Eiffel Tower was built and became one of the symbols of Paris..." “The banks of the Seine have become a romantic destination for tourists.” (3) Subtitle generation: Based on the emotional information and event descriptions in the video, subtitles that are synchronized with the video content are automatically generated. The subtitle content is displayed synchronously according to the timeline order and emotional tags (such as romance, history).
[0111] 6. Step 5: Cross-platform adaptation and output: (1) Video format adjustment: Automatically adjust the resolution, frame rate, and format according to the requirements of the target platform to ensure smooth playback on various devices.
[0112] (2) Personalized recommendation: Based on the user’s historical preferences and interests, the system will generate a personalized version of the video. For example, if the user prefers historical and cultural content, the system will enhance the historical background information in the video.
[0113] 7. Output results: Automatically generate a 1- to 2-minute tourism promotional video showing romantic attractions in Paris, such as the Eiffel Tower and the banks of the Seine, with audio narration, background music, and synchronized subtitles. The video is presented in timeline order to ensure that the content is consistent with the description entered by the user and that the emotion fits the romantic theme.
[0114] In this exemplary application, video generation is performed by mapping content tags with a multimodal knowledge graph, which can improve the efficiency of video generation and meet the growing demand for content creation. At the same time, a timeline for content generation is constructed through the mapping nodes of the knowledge graph, and video generation is performed based on the constructed timeline, which can realize the dynamic construction capability of the timeline in video generation.
[0115] Embodiment 2 Figure 8A block diagram of a content generation device according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 8 shown, the device 800 may include: an acquisition module 810, an extraction module 820, a mapping module 830, a construction module 840, and a generation module 850, where: The acquisition module 810 is configured to acquire input content, where the input content includes at least one of text, audio, image, or video; The extraction module 820 is configured to extract content tags based on the input content; The mapping module 830 is configured to map the content tags to nodes in a pre-constructed multimodal knowledge graph to obtain mapped nodes; The construction module 840 is configured to construct a timeline for content generation based on the mapped nodes; The generation module 850 is configured to generate target content based on the mapped nodes and the timeline.
[0116] In an alternative embodiment, the construction module 840 is further configured to: acquire associated nodes connected to the mapped nodes and the entity relationships with the associated nodes; determine time information and / or causal relationships corresponding to the mapped nodes based on the associated nodes and the entity relationships; construct a timeline for content generation based on the time information and the causal relationships.
[0117] In an alternative embodiment, the target content includes voice commentary audio; correspondingly, the generation module 850 is further configured to: acquire relevant semantic relationships from the multimodal knowledge graph based on the mapped nodes; generate voice commentary text based on the semantic relationships, the content tags, and the timeline; convert the voice commentary text into voice commentary audio using a text-to-speech conversion method.
[0118] In an alternative embodiment, the generation module 850 is further configured to: acquire multimedia materials matching the content tags from the multimodal knowledge graph based on the mapped nodes, where the multimedia materials include video clips, pictures, and audio; insert and edit the multimedia materials based on the timeline to generate the target content.
[0119] In an alternative embodiment, the extraction module 820 is further configured to: When the input content includes input text, use natural language processing methods to extract a first text label of the input text, where the text label includes at least partial information of events, people, locations, and emotions; When the input content includes input audio, use speech recognition methods to convert the input audio into text, use natural language processing methods to extract a second text label of the converted text, and use speech emotion analysis methods to obtain an emotion label of the input audio; When the input content includes input images and / or input videos, use computer vision methods to extract video image labels of the input images and / or the input videos.
[0120] In an alternative embodiment, the apparatus 800 is further configured to: Obtain target format parameters of a target platform, where the target format parameters include resolution, frame rate, and duration; Use video processing tools to process the target content to obtain a video that conforms to the target format parameters.
[0121] In an alternative embodiment, the apparatus 800 is further configured to: When authorized by the current user, obtain target knowledge graph information corresponding to the current user; Correspondingly, the generation module 850 is further configured to: When authorized by the current user, obtain target knowledge graph information corresponding to the current user; Correspondingly, the generation of the target content based on the mapping node and the timeline further includes: Generate target content based on the target knowledge graph information, the mapping node, and the timeline.
[0122] Embodiment III Figure 9 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing a content generation method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 9As shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 can also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the content generation method. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.
[0123] In some embodiments, the processor 10020 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0124] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0125] It should be noted that Figure 9 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.
[0126] In this embodiment, the content generation method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0127] Embodiment 4 The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the content generation method in the embodiments are implemented.
[0128] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the content generation method in the embodiment. In addition, the computer-readable storage medium can also be used to temporarily store various data that have been output or will be output.
[0129] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0130] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0131] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.
Claims
1. A content generation method, characterized in that: The method comprises: Acquire input content, where the input content includes at least one of text, audio, image or video; extracting content tags based on the input content; Mapping the content tag with a node in a pre-constructed multimodal knowledge graph to obtain a mapping node; Constructing a timeline of content generation based on the mapping nodes; Target content is generated based on the mapping nodes and the timeline.
2. The method according to claim 1, characterized in that: The constructing a timeline of content generation based on the mapping node includes: Acquire an associated node connected to the mapping node and an entity relationship with the associated node; Determine the time information and / or causal relationship corresponding to the mapping node based on the associated node and the entity relationship; A timeline of content generation is constructed based on the time information and the causal relationship.
3. The method according to claim 2, characterized in that The target content includes voice commentary audio; Correspondingly, the generating of target content based on the mapping node and the timeline includes: Acquire relevant semantic relations from the multimodal knowledge graph based on the mapping nodes; Generate a voice commentary text based on the semantic relationship, the content tag and the timeline; The voice commentary text is converted into voice commentary audio using a text-to-speech conversion method.
4. The method according to claim 2, characterized in that: The generating of target content based on the mapping node and the timeline includes: Acquire multimedia materials matching the content tag from the multimodal knowledge graph based on the mapping node, the multimedia materials including video clips, pictures and audio; The multimedia material is inserted and edited based on the timeline to generate the target content.
5. The method according to claim 1, characterized in that The extracting content tags based on the input content includes: In the case where the input content includes input text, a natural language processing method is used to extract a first text tag of the input text, wherein the text tag includes at least part of information of events, persons, places and emotions; In the case where the input content includes input audio, converting the input audio into text using a speech recognition method, extracting a second text label of the converted text using a natural language processing method, and acquiring an emotional label of the input audio using a speech emotion analysis method; In the case where the input content includes an input image and / or an input video, a computer vision method is used to extract video image tags of the input image and / or the input video.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtain target format parameters of the target platform, wherein the target format parameters include resolution, frame rate, and duration; The target content is processed using a video processing tool to obtain a video that meets the target format parameters.
7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: With the authorization of the current user, obtain the target knowledge graph information corresponding to the current user; Correspondingly, the generating of the target content based on the mapping node and the timeline further includes: Target content is generated based on the target knowledge graph information, the mapping nodes and the timeline.
8. A content generating device, characterized in that: The device comprises: An acquisition module, used to acquire input content, wherein the input content includes at least one of text, audio, image or video; An extraction block is used to extract content tags based on the input content; A mapping module, used to map the content tags with nodes in a pre-built multimodal knowledge graph to obtain mapping nodes; A construction module, used for constructing a timeline of content generation based on the mapping nodes; A generation module is used to generate target content based on the mapping node and the timeline.
9. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.