Video content generation method and device, electronic device, and storage medium

By acquiring topic information and shooting elements and matching them in the video material library, the problem of similar short video content was solved, personalized video content was generated, and video quality was improved.

CN115580758BActive Publication Date: 2026-02-03CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211249471.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-12
Publication Date
2026-02-03
Estimated Expiration
2042-10-12

AI Technical Summary

Technical Problem

The existing short video production templates are relatively simple, resulting in most creators producing similar video content, which affects the normal release of short videos.

Method used

By acquiring the first topic information, filtering out the target semantic type, identifying shooting elements, and matching the content in the video material library, the content to be combined is obtained by combining interactive information, and finally combined into personalized video content.

Benefits of technology

It enhances the personalization of video content, improves video quality, and makes the generated video content more diverse and expressive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115580758B_ABST
    Figure CN115580758B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a video content generation method and device, electronic equipment and storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring first topic information, and screening at least two target semantic types from preset video semantic types according to the first topic information. Acquire shooting information and identify shooting elements from the shooting information. According to each target semantic type, the first topic information and the shooting element, content matching is performed in the video material library to obtain the material content corresponding to the target semantic type. Detect the interaction information, and according to the interaction information, obtain the to-be-combined content from the material content corresponding to each target semantic type, so as to combine all the to-be-combined content corresponding to the target semantic type, and obtain the video content. It can be seen that the embodiment of the application can generate personalized video content and improve the video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a video content generation method and apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, short video creation is popular on various online platforms, and many self-media creators use short videos to share knowledge and attract customers. However, existing short video production templates are still relatively simple, resulting in most creators producing similar video content, which in turn affects the normal release of short videos. Summary of the Invention

[0003] The main objective of this application is to provide a video content generation method, apparatus, electronic device, and storage medium, which aim to generate personalized video content.

[0004] To achieve the above objectives, a first aspect of this application proposes a video content generation method, the method comprising:

[0005] Obtain information about the first topic;

[0006] Based on the first topic information, at least two target semantic types are selected from the preset video semantic types, and each target semantic type is used to determine different video semantic features;

[0007] Acquire shooting information and identify shooting elements from the shooting information;

[0008] Based on each of the target semantic types, the first topic information, and the shooting elements, content matching is performed in the video material library to obtain the material content corresponding to the target semantic type;

[0009] Detect interaction information, and based on the interaction information, obtain the content to be combined from the material content corresponding to each of the target semantic types;

[0010] The content to be combined corresponding to all the target semantic types is combined to obtain the video content.

[0011] In some embodiments, before filtering out the target semantic type from preset video semantic types based on the first topic information, the method further includes:

[0012] Collect video samples;

[0013] The video sample is subjected to content analysis to obtain the first video content, and the second topic information is extracted from the first video content.

[0014] Obtain the semantic type related to the first video content from the preset video semantic types, and use it as the first semantic type;

[0015] Establish a correspondence between the second topic information and the first semantic type;

[0016] The step of filtering the target semantic type from the preset video semantic types based on the first topic information includes:

[0017] Based on the first topic information and the corresponding relationship, the target semantic type is selected from the video semantic types.

[0018] In some embodiments, obtaining the semantic type related to the first video content from a preset video semantic type as the first semantic type includes:

[0019] Based on a preset video semantic type, a material classification standard is obtained, which is used to determine different semantic types within the video semantic type;

[0020] Based on the material classification criteria, the first video content is grouped to obtain multiple sets of material data;

[0021] Based on the video semantic type, label the semantic type corresponding to each group of material data to add it to the first semantic type;

[0022] The method further includes:

[0023] Add the aforementioned sets of material data to the video material library.

[0024] In some embodiments, after extracting the second topic information from the first video content, the method further includes:

[0025] Based on the second topic information, video capture and processing are performed to obtain the target video;

[0026] The target video is analyzed to obtain the second video content;

[0027] The first video content is grouped according to the material classification criteria to obtain multiple sets of material data, including:

[0028] Based on the material classification criteria, the first video content and the second video content are grouped to obtain multiple sets of material data.

[0029] In some embodiments, the step of performing content matching in the video material library based on each of the target semantic types, the first topic information, and the shooting elements to obtain the material content corresponding to the target semantic type includes:

[0030] Obtain the text material corresponding to each of the target semantic types from the video material library;

[0031] Based on the first topic information and the shooting elements, tag recognition is performed to obtain text tags and multimedia tags;

[0032] Based on the text tags, content matching is performed on the text material to obtain the text content;

[0033] Obtain the multimedia materials corresponding to the text content from the video material library;

[0034] Based on the multimedia tags, content matching is performed on the multimedia materials to obtain multimedia content;

[0035] The text content and the multimedia content are used as the material content corresponding to the target semantic type.

[0036] In some embodiments, detecting interaction information and obtaining the content to be combined from the material content corresponding to each target semantic type based on the interaction information includes:

[0037] Detect voice information and perform information matching in the material content based on the voice information to obtain reference content;

[0038] The reference content is then displayed.

[0039] The editing information of the reference content is detected, and the reference content is updated according to the editing information to obtain the content to be combined.

[0040] In some embodiments, the step of combining the content to be combined corresponding to all the target semantic types to obtain video content includes:

[0041] Obtain the sorting information corresponding to each of the target semantic types;

[0042] Obtain a video timeline, which includes multiple playback time segments strung together in chronological order;

[0043] According to the sorting information, the playback time segment corresponding to each of the target semantic types is obtained from the video timeline, and used as the target time segment for each target semantic type.

[0044] The content to be combined corresponding to each of the target semantic types is imported into the target time period of the target semantic type to obtain the video content.

[0045] To achieve the above objectives, a second aspect of this application provides a video content generation apparatus, the apparatus comprising:

[0046] The acquisition module is used to obtain information about the first topic.

[0047] The filtering module is used to filter at least two target semantic types from preset video semantic types based on the first topic information, wherein each target semantic type is used to determine different video semantic features;

[0048] The identification module is used to acquire shooting information and identify shooting elements from the shooting information;

[0049] The matching module is used to perform content matching in the video material library according to each of the target semantic types, the first topic information and the shooting elements, to obtain the material content corresponding to the target semantic type;

[0050] An interaction module is used to detect interaction information and, based on the interaction information, obtain the content to be combined from the material content corresponding to each of the target semantic types.

[0051] The combination module is used to combine the content to be combined corresponding to all the target semantic types to obtain video content.

[0052] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0053] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0054] The video content generation method, apparatus, electronic device, and storage medium proposed in this application can obtain first topic information and then filter out suitable target semantic types from preset video semantic types to form a semantic framework for the video content. Based on this, shooting elements are identified to determine the shooting scene, which is then combined with the first topic information and the target semantic type to match material content that matches the shooting scene and topic in the video material library, providing a more accurate material reference. Furthermore, based on detected interaction information, content to be combined to meet user needs is obtained from the material content to form the final video content, which can improve the personalization of the generated video content and improve video quality. Attached Figure Description

[0055] Figure 1 This is a schematic flowchart of a video content generation method provided in an embodiment of this application;

[0056] Figure 2 This is a schematic diagram of a process for constructing a correspondence in an embodiment of this application;

[0057] Figure 3 yes Figure 1 A specific flowchart of step S130;

[0058] Figure 4 yes Figure 1 A specific flowchart of step S140;

[0059] Figure 5 This is a schematic diagram of the structure of the video content generation device provided in the embodiments of this application;

[0060] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0062] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0064] First, let's analyze some of the terms used in this application:

[0065] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0066] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0067] Information extraction is a text processing technique that extracts factual information such as entities, relationships, and events from natural language text and outputs it as structured data. Information extraction is a technique for extracting specific information from text data. Text data is composed of specific units, such as sentences, paragraphs, and chapters. Text information is composed of smaller, specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of information extraction. Of course, text information extraction techniques can extract information of various types.

[0068] Image captioning generates natural language descriptions for images, helping applications understand the semantics expressed in the image's visual scene. For example, image captioning can convert image retrieval into text retrieval, classify images, and improve retrieval results. While people can often describe the details of an image's visual scene with a quick glance, automatically adding descriptions to images is a comprehensive and challenging computer vision task, requiring the conversion of complex information contained within the image into natural language descriptions. Compared to ordinary computer vision tasks, image captioning not only requires identifying objects in an image but also associating the identified objects with natural semantics and describing them in natural language. Therefore, image captioning requires extracting deep features from the image, associating them with semantic features, and converting them to generate descriptions.

[0069] Currently, short video creation is popular on various online platforms, and many self-media creators use short videos to share knowledge and attract customers. However, existing short video production templates are still relatively simple, resulting in most creators producing similar video content, which in turn affects the normal release of short videos.

[0070] Based on this, embodiments of this application provide a video content generation method and apparatus, electronic device, and storage medium, aimed at generating personalized video content.

[0071] The video content generation method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the video content generation method in this application is described.

[0072] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0073] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0074] The video content generation method provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video content generation method, but is not limited to the above forms. The following description uses a terminal as an example.

[0075] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0076] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0077] Figure 1 This is a flowchart illustrating a video content generation method provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S100 to S150.

[0078] Step S100: Obtain information about the first topic.

[0079] In this embodiment, the first topic information may include at least one of topic keywords, phrases, or sentences, without specific limitations. The methods for obtaining the first topic information include, but are not limited to: the terminal launching an application and generating an input box in the application interface to receive the first topic information entered by the user in the input box, such as the user manually entering the title of a photograph; the terminal generating multiple topic tags in the application interface and obtaining the topic tag selected by the user from the multiple topic tags as the first topic information.

[0080] The acquisition of multiple topic tags includes, but is not limited to: collecting trending topic data (such as trending search lists, blog post titles, blog post tags, video titles, video descriptions, video tags, and high-frequency bullet comments) from social networking or news websites; obtaining multiple trending tags from the trending topic data through word segmentation or clustering; then, statistically processing each trending tag based on at least one statistical indicator, such as topic popularity (e.g., number of likes or scores) and frequency of occurrence, to obtain the corresponding statistical value for each trending tag; and finally, selecting trending tags whose statistical values ​​meet preset conditions as topic tags. These preset conditions are specified and adjusted manually; for example, a preset condition is that the statistical value is greater than a specified value, without limitation.

[0081] Step S110: Based on the first topic information, select at least two target semantic types from the preset video semantic types.

[0082] In this embodiment, the preset video semantic type may include various semantic types related to video content, such as title, opening, main content, summary, and dialogue, without specific limitations. Each semantic type is used to determine different video semantic features; therefore, the video semantic type is used to determine the semantic structure and relationships of the video content. Specifically, in step S110, the terminal can obtain the correspondence between topic information and semantic types, and then obtain the target semantic type corresponding to the first topic information based on the correspondence. The aforementioned correspondence can be manually specified or obtained through data analysis, without limitation.

[0083] Step S120: Acquire shooting information and identify shooting elements from the shooting information.

[0084] In this embodiment, the captured information can be a captured image taken by the terminal's camera device, or a screen recording obtained by recording the application interface of the terminal, etc., without specific limitations. Correspondingly, the captured elements can be the constituent objects contained in the captured image or screen recording, and the captured elements can be obtained by performing image description processing on the captured information (e.g., using an image description model based on CNN+LSTM+attention) or image object detection (e.g., using an object detection model based on R-CNN), also without specific limitations. For example, if the captured information is a street scene image, the captured elements can include pedestrians, buildings, trees, and streetlights, etc. If the captured information is a screen recording, the captured elements can include the software name, software icon, AI character image, and other user interface (UI) design elements in the screen recording.

[0085] Step S130: Based on each target semantic type, first topic information, and shooting elements, perform content matching in the video material library to obtain the material content corresponding to the target semantic type.

[0086] In this embodiment, the video material library is used to store material data corresponding to different semantic types. This material data includes, but is not limited to, text, images, emoticons, charts, music, audio, video, animation, subtitles, screen filters, post-production effects, and video project files. In practical applications, multiple matching texts can be pre-annotated to the material data. By performing content matching (such as text matching algorithms based on BM25 or deep learning) between the first topic information and shooting elements and the matching texts of each semantic type, matching texts that meet the matching conditions are obtained as target texts. The data of the annotated target texts is then extracted from the material data as the material content.

[0087] In some optional implementations, the matching text may include style tags. Style tags determine the presentation style of the video content. Style tags can be general, story-based, expert-oriented, humorous, or educational, and can also be determined based on the current user of the terminal, thus achieving personalized classification of the material content according to the current user's video production style, without specific limitations. By matching with the primary topic information and shooting elements, it ensures that the presentation style of the video content aligns with the actual topic and shooting scene.

[0088] Optionally, the terminal can determine the current user based on the logged-in account on the terminal, or it can identify the current user from the shooting elements, and there is no limitation.

[0089] Step S140: Detect interaction information and, based on the interaction information, obtain the content to be combined from the material content corresponding to each target semantic type.

[0090] In this embodiment, the interactive information may include, but is not limited to, user input information in the application interface of the terminal, user voice information, and image information captured by the user. Accordingly, the terminal can collect voice information through an audio recording device and image information through a shooting device.

[0091] In one alternative implementation, the terminal can display the content corresponding to each target semantic type, such as displaying text, playing videos, and playing music in the application interface. When a user selects a piece of content from the content through the application interface, the terminal can retrieve the selected content as the content to be combined.

[0092] Step S150: Combine all the content to be combined corresponding to the target semantic types to obtain the video content.

[0093] It is understandable that all the content to be combined is processed by combining them. Specifically, this can be done by importing all the content to be combined into the same video timeline, and adjusting the order in which the various content materials appear through the video timeline to finally obtain the video content.

[0094] In some optional implementations, the terminal can obtain sorting information corresponding to each target semantic type and obtain a video timeline, wherein the video timeline includes multiple playback time segments strung together in chronological order, and the duration of each playback time segment can be preset. Based on this, the terminal obtains the playback time segments corresponding to each target semantic type from the video timeline according to the sorting information as the target time segments for the target semantic types. The content to be combined corresponding to each target semantic type is then imported into the target time segment of the target semantic type to obtain the video content. In practical applications, the terminal can respond to adjustment commands to adjust the duration or sorting of each time segment in the video timeline to obtain an adjusted video timeline. Optionally, the adjustment command can be a command generated by the terminal based on a manual adjustment operation detected from the application interface, without specific limitations.

[0095] In another optional implementation, the terminal can display all content to be combined, obtain confirmation instructions for each piece of content to be combined, and then sort all content according to the order in which the confirmation instructions are obtained, obtaining a sorting result. Finally, the terminal can combine each piece of content sequentially according to the sorting result. For example, in a practical application, when a user selects any piece of content to be combined through the application interface, the terminal detects a confirmation instruction for that piece of content.

[0096] In another optional implementation, the terminal can also obtain sorting instructions for all content to be combined, identify the sorting result from the sorting instructions, and then sequentially combine each content to be combined according to the sorting result. The sorting instructions can be generated in ways including, but not limited to: the user dragging the operation controls corresponding to each content to be combined through the application interface to adjust the arrangement order of the operation controls, and then the terminal generates sorting instructions based on the adjusted arrangement order of the operation controls.

[0097] In another optional implementation, the terminal can also sequentially combine the content to be combined according to the semantic order set for each target semantic type. For example, assuming the preset video semantic types include title, opening, main content, summary, and dialogue, the corresponding semantic order can be: title, opening, main content, dialogue, summary. That is, for content A to be combined corresponding to the title, content B to be combined corresponding to the opening, content C to be combined corresponding to the main content, content D to be combined corresponding to the summary, and content E to be combined corresponding to the dialogue, the video content generated by the terminal is ABCED.

[0098] As can be seen, the video content generation method provided in this application, by obtaining first topic information, can filter out suitable target semantic types from preset video semantic types to form the semantic framework of the video content. Based on this, shooting elements are identified to determine the shooting scene, which is then combined with the first topic information and the target semantic type to match material content that conforms to the shooting scene and topic in the video material library, providing a more accurate material reference. Furthermore, based on the detected interaction information, content to be combined to meet user needs is obtained from the material content to form the final video content, which can improve the personalization of the generated video content and improve video quality.

[0099] For some alternative implementation methods, please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating a process for constructing a correspondence in an embodiment of this application. For example... Figure 2 As shown, before step S110, the following steps S200 to S230 may also be included.

[0100] Step S200: Collect video samples.

[0101] In this application embodiment, the video sample can be a popular video, a highly praised video, or a highly collected video collected through channels such as social networking sites, news websites, and video websites, without limitation.

[0102] Step S210: Perform content analysis on the video sample to obtain the first video content, and extract the second topic information from the first video content.

[0103] For example, content analysis of video samples may include, but is not limited to, version analysis, material extraction, subtitle extraction, deduplication and error correction, and punctuation recognition.

[0104] In some optional implementations, step S210 can also be: extracting keyframes from the video sample as the first video content. The frame content is obtained by performing image description processing or image object detection on the keyframes, and then the frame content of all keyframes is taken as the second topic information. Keyframes are the video frames where the key actions of characters or objects in the video sample are located. Algorithms for extracting keyframes include, but are not limited to: motion analysis-based keyframe extraction algorithms, such as analyzing the optical flow of object motion in the video frames contained in the video sample and selecting the video frame with the fewest optical flow movements as the keyframe; and video clustering-based keyframe extraction algorithms, such as dividing the video frames contained in the video sample into several clusters and selecting the image frame closest to the cluster center in each cluster as the keyframe.

[0105] Step S220: Obtain the semantic type related to the first video content from the preset video semantic types, and use it as the first semantic type.

[0106] In this embodiment, the terminal can obtain the semantic type labeled by a person for the first video content, or it can identify the semantic type related to the first video content from the video semantic type through data analysis, without limitation.

[0107] Furthermore, in some optional implementations, step S220 may include, but is not limited to, the following steps S221 to S223.

[0108] Step S221: Based on the preset video semantic type, obtain the material classification standard. The material classification standard is used to determine the different semantic types in the video semantic type.

[0109] In the embodiments of this application, the material classification criteria may include classification criteria set for different semantic types, and the classification criteria include, but are not limited to, at least one of video time period, video content keywords, and video frame category.

[0110] Step S222: Based on the material classification criteria, the first video content is grouped to obtain multiple sets of material data.

[0111] In some optional implementations, after step S210, the terminal can further perform video capture processing (such as web crawling or video search) based on the second topic information to obtain the target video, and then perform content analysis on the target video to obtain the second video content. The content analysis of the target video can be referred to the explanation of the video sample content analysis in step S210, and will not be repeated here. Based on this, in step S222, the terminal can perform content grouping processing on the first video content and the second video content according to the material classification criteria to obtain multiple sets of material data. It is evident that further searching numerous video resources on the Internet based on topic information can improve the sample reliability of content analysis and material grouping.

[0112] Step S223: According to the semantic type of the video, label the semantic type corresponding to each group of material data to add it to the first semantic type.

[0113] In some optional implementations, the terminal can acquire reference videos, obtain grouped content based on material classification standards, and semantic types labeled for each group, using this as training data to train a recognition model. Specifically, in steps S222 and S223, the first video content is input into the recognition module for recognition processing, resulting in multiple sets of material data and the corresponding semantic types for each set. The recognition module can employ models based on classification algorithms such as logistic regression, Naive Bayes, decision trees, support vector machines, random forests, or gradient boosting trees; no specific limitations are imposed.

[0114] Correspondingly, after step S223, the terminal can also add multiple sets of material data to the video material library. Therefore, by grouping, labeling, and storing the material data, it is only necessary to search for different semantic types to obtain material content matching the semantic type from the video material library.

[0115] As can be seen, through the above steps S221 to S223, video content is segmented and semantically labeled according to the specified classification criteria, thereby achieving more intelligent semantic classification.

[0116] Step S230: Establish a correspondence between the second topic information and the first semantic type.

[0117] Accordingly, step S110 can specifically be: based on the first topic information and the corresponding relationship, select the target semantic type from the video semantic types.

[0118] As can be seen, through the above steps S200 to S230, a correspondence between topic information and semantic type is established, which facilitates the quick filtering of matching semantic type based on topic information in practical applications.

[0119] For some alternative implementation methods, please refer to Figure 3 , Figure 3 yes Figure 1 A schematic diagram of a specific process for step S130. For example... Figure 3 As shown, step S130 includes, but is not limited to, the following steps S131 to S136.

[0120] Step S131: Obtain the text material corresponding to each target semantic type from the video material library.

[0121] Step S132: Based on the first topic information and shooting elements, perform tag recognition to obtain text tags and multimedia tags.

[0122] In this embodiment of the application, text tags are used to label text data, and multimedia tags are used to label multimedia data. Multimedia data refers to media data other than text, including but not limited to images, music, audio, video, and animation.

[0123] Step S133: Based on the text tags, perform content matching in the text material to obtain the text content.

[0124] Step S134: Obtain the multimedia materials corresponding to the text content from the video material library.

[0125] In this embodiment, the video material library can use a hierarchical storage structure, a tree-like storage structure, or a linked storage structure to store different text content and its corresponding multimedia materials, without specific limitations. That is, a parent node can be constructed for the text content, and then child nodes of the parent node can be constructed using the corresponding multimedia materials. Specifically, when the multimedia materials include at least two different media types, associated nodes can also be constructed for the different media types according to a hierarchy. For example, the parent node corresponding to the text content is connected to different image modal child nodes, and each image modal child node is connected to the corresponding animation modal child node and sound modal child node. Therefore, by traversing the node paths, a relatively comprehensive material chain can be obtained.

[0126] Step S135: Based on the multimedia tags, perform content matching in the multimedia materials to obtain the multimedia content.

[0127] The multimedia materials include multimedia data with multiple tags, and the multimedia content can be quickly extracted by tag matching.

[0128] Step S136: Use the text content and multimedia content as the material content corresponding to the target semantic type.

[0129] As can be seen, by combining steps S131 to S136, at least two interactive communication media related to the theme information and shooting scene elements can be selected, which can enrich the diversity of material content and help improve the intuitiveness and vividness of video content display.

[0130] For some alternative implementation methods, please refer to Figure 4 , Figure 4 yes Figure 1 A schematic diagram of a specific process for step S140. For example... Figure 4 As shown, step S140 includes, but is not limited to, the following steps S141 to S143.

[0131] Step S141: Detect speech information and perform information matching in the source material based on the speech information to obtain reference content.

[0132] In this embodiment of the application, the reference content may include, but is not limited to, text, images, music, voice, emoticons, charts, videos, animations, subtitles, screen filters, post-production effects, and video project files. In step S141, specifically, the terminal can collect voice data through a recording device and identify voice information from the voice data through automatic speech recognition (ASR) technology.

[0133] Step S142: Display the reference content.

[0134] In step S142, specifically, the terminal can display text, images, subtitles, screen filters, post-production effects, and video project files in the application interface, as well as play videos, animations, music, and audio for user reference.

[0135] Step S143: Detect the editing information of the reference content, and update the reference content according to the editing information to obtain the content to be combined.

[0136] In this embodiment, the edit information is generated based on editing operations on the reference content. Editing operations can represent user actions performed on the content displayed in the application interface. These operations include, but are not limited to, content replacement, deletion, addition, or modification. For example, a user can edit the specific content of a text in the application interface, delete an image, add a newly uploaded video, or modify a subtitle.

[0137] As can be seen, by using steps S141 to S143, reference content can be quickly filtered out from the source material through voice recognition, and users can easily update the reference content according to their own needs, thus further improving the personalization and diversity of video content.

[0138] The following examples, Tables 1 and 2, illustrate this process. Table 1 shows a sample of the source material, and Table 2 shows a sample of the content to be combined. It should be understood that this is merely an example and does not constitute a specific limitation on the source material or the content to be combined. As shown in Tables 1 and 2, users can use the style tag "General" and select the title (h3), opening (o3), main content (c1), main content (c3), summary (s3), and dialogue (a3) ​​from the source material to create entirely new video content. The process is convenient.

[0139] Table 1. Material Content Illustration

[0140]

[0141] Table 2. Schematic diagram of contents to be combined

[0142]

[0143] Please see Figure 5 This application also provides a video content generation apparatus that can implement the above-described video content generation method. The video content generation apparatus includes an acquisition module 510, a filtering module 520, a recognition module 530, a matching module 540, an interaction module 550, and a combination module 560, wherein:

[0144] Module 510 is used to obtain information about the first topic.

[0145] The filtering module 520 is used to filter at least two target semantic types from the preset video semantic types based on the first topic information, and each target semantic type is used to determine different video semantic features.

[0146] The recognition module 530 is used to acquire shooting information and identify shooting elements from the shooting information.

[0147] The matching module 540 is used to perform content matching in the video material library based on each target semantic type, first topic information and shooting elements to obtain the material content corresponding to the target semantic type.

[0148] The interaction module 550 is used to detect interaction information and, based on the interaction information, obtain the content to be combined from the material content corresponding to each target semantic type.

[0149] The combination module 560 is used to combine the content to be combined corresponding to all target semantic types to obtain video content.

[0150] The specific implementation of the video content generation device is basically the same as the specific embodiment of the video content generation method described above, and will not be repeated here.

[0151] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described video content generation method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0152] Please see Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0153] The processor 601 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0154] The memory 602 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 602 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called and executed by the processor 601 using the video content generation method of the embodiments of this application.

[0155] The input / output interface 603 is used to implement information input and output;

[0156] The communication interface 604 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0157] Bus 605 transmits information between various components of the device (e.g., processor 601, memory 602, input / output interface 603, and communication interface 604);

[0158] The processor 601, memory 602, input / output interface 603, and communication interface 604 are connected to each other within the device via bus 605.

[0159] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video content generation method.

[0160] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0161] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0162] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0164] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0165] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0166] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0167] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0168] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0169] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0170] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0171] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for generating video content, characterized in that, The method includes: Obtain information about the first topic; Based on the first topic information, at least two target semantic types are selected from the preset video semantic types, and each target semantic type is used to determine different video semantic features; Acquire shooting information and identify shooting elements from the shooting information; Based on each of the target semantic types, the first topic information, and the shooting elements, content matching is performed in the video material library to obtain the material content corresponding to the target semantic type; Detect interaction information, and based on the interaction information, obtain the content to be combined from the material content corresponding to each of the target semantic types; The content to be combined corresponding to all the target semantic types is combined to obtain the video content; The step of performing content matching in the video material library based on each of the target semantic types, the first topic information, and the shooting elements to obtain the material content corresponding to the target semantic type includes: Obtain the text material corresponding to each of the target semantic types from the video material library; Based on the first topic information and the shooting elements, tag recognition is performed to obtain text tags and multimedia tags; Based on the text tags, content matching is performed on the text material to obtain the text content; Obtain the multimedia materials corresponding to the text content from the video material library; Based on the multimedia tags, content matching is performed on the multimedia materials to obtain multimedia content; The text content and the multimedia content are used as the material content corresponding to the target semantic type; The step of detecting interaction information and obtaining the content to be combined from the material content corresponding to each target semantic type based on the interaction information includes: Detect voice information and perform information matching in the material content based on the voice information to obtain reference content; The reference content is then displayed. The editing information of the reference content is detected, and the reference content is updated according to the editing information to obtain the content to be combined.

2. The method according to claim 1, characterized in that, Before filtering the target semantic type from the preset video semantic types based on the first topic information, the method further includes: Collect video samples; The video sample is subjected to content analysis to obtain the first video content, and the second topic information is extracted from the first video content. Obtain the semantic type related to the first video content from the preset video semantic types, and use it as the first semantic type; Establish a correspondence between the second topic information and the first semantic type; The step of filtering the target semantic type from the preset video semantic types based on the first topic information includes: Based on the first topic information and the corresponding relationship, the target semantic type is selected from the video semantic types.

3. The method according to claim 2, characterized in that, The step of obtaining a semantic type related to the first video content from a preset video semantic type, as the first semantic type, includes: Based on a preset video semantic type, a material classification standard is obtained, which is used to determine different semantic types within the video semantic type; Based on the material classification criteria, the first video content is grouped to obtain multiple sets of material data; Based on the video semantic type, label the semantic type corresponding to each group of material data to add it to the first semantic type; The method further includes: Add the aforementioned sets of material data to the video material library.

4. The method according to claim 3, characterized in that, After extracting the second topic information from the first video content, the method further includes: Based on the second topic information, video capture and processing are performed to obtain the target video; The target video is analyzed to obtain the second video content; The first video content is grouped according to the material classification criteria to obtain multiple sets of material data, including: Based on the material classification criteria, the first video content and the second video content are grouped to obtain multiple sets of material data.

5. The method according to any one of claims 1 to 4, characterized in that, The step of combining the content to be combined corresponding to all the target semantic types to obtain video content includes: Obtain the sorting information corresponding to each of the target semantic types; Obtain a video timeline, which includes multiple playback time segments strung together in chronological order; According to the sorting information, the playback time segment corresponding to each of the target semantic types is obtained from the video timeline, and used as the target time segment for each target semantic type. The content to be combined corresponding to each of the target semantic types is imported into the target time period of the target semantic type to obtain the video content.

6. A video content generation device, characterized in that, The apparatus is used to implement the video content generation method according to any one of claims 1 to 5, the apparatus comprising: The acquisition module is used to obtain information about the first topic. The filtering module is used to filter at least two target semantic types from preset video semantic types based on the first topic information, and each target semantic type is used to determine different video semantic features; The identification module is used to acquire shooting information and identify shooting elements from the shooting information; The matching module is used to perform content matching in the video material library according to each of the target semantic types, the first topic information and the shooting elements, to obtain the material content corresponding to the target semantic type; An interaction module is used to detect interaction information and, based on the interaction information, obtain the content to be combined from the material content corresponding to each of the target semantic types. The combination module is used to combine the content to be combined corresponding to all the target semantic types to obtain video content.

7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the video content generation method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video content generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video generation method and device

    CN112528073A

  • Video display and generation method and device

    CN112689189A