A method, system, device and storage medium for automatically generating videos

By identifying key elements in video frames and inserting popular elements, the cumbersome problem of traditional video production process is solved, and video data containing popular elements is efficiently and automatically generated, improving video quality and dissemination efficiency.

CN119583900BActive Publication Date: 2025-06-27BEIJING JING PARTNER TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510127656.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-06-27
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

The production process of traditional videos is cumbersome, time-consuming and difficult to quickly capture and utilize popular elements, affecting the influence and dissemination efficiency of videos.

Method used

By identifying the key elements in the video frame, calling the corresponding popular elements, and inserting them into the video frame, generating the target video frame, and finally automatically generating video data containing the popular elements.

Benefits of technology

Automatically organize and automatically generate video data containing popular elements, improving the quality and dissemination efficiency of video data, and shortening the video production cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119583900B_ABST
    Figure CN119583900B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system, device and storage medium for automatically generating videos, which belong to the technical field of video processing. The method includes: importing video data to be processed, where the video data includes at least one initial video frame; identifying key elements of the initial video frame, where the key elements are the elements to be monitored in the initial video frame; obtaining popular elements corresponding to the key elements according to the key elements; inserting the popular elements into the initial video frame where the key elements are located to obtain a target video frame; generating and exporting final video data according to the target video frame. The present application can call popular elements based on the key elements in the initial video frame, then insert the popular elements into the initial video frame to obtain a target video frame, and finally generate final video data with the target video frame, so as to realize the automatic sorting and automatic generation of video data containing popular elements and improve the overall quality of the video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of video processing. Specifically, it relates to a method, system, device, and storage medium for automatically generating videos. Background Art

[0002] With the rise of social media and video sharing platforms, the creation and dissemination of video content have increasingly become important ways of information dissemination. However, the traditional video production process is cumbersome and requires a large amount of time for material sorting, feature extraction, creative concept design, and post-production. Especially in the current context of pursuing content novelty and attractiveness, how to quickly capture and utilize popular elements in video creation has become the key to enhancing the influence and dissemination efficiency of videos.

[0003] Therefore, it is particularly important to develop a video generation system that can automatically analyze material features and intelligently integrate popular elements. Summary of the Invention

[0004] The technical problem to be solved by this application is to overcome the deficiencies of the prior art and provide a method, system, device, and storage medium for automatically generating videos.

[0005] To solve the above technical problem, the basic concept of the technical solution adopted in this application is:

[0006] In the first aspect of this application, a method for automatically generating videos is provided. The method includes:

[0007] Import video data to be processed, where the video data includes at least one initial video frame;

[0008] Identify the key elements of the initial video frame, where the key elements are the elements to be monitored in the initial video frame;

[0009] Obtain the popular elements corresponding to the key elements according to the key elements. The popular elements include at least text, music, special effects, and filters;

[0010] Insert the popular elements into the initial video frame where the key elements are located to obtain a target video frame;

[0011] Generate and export the final video data according to the target video frame.

[0012] By adopting the above technical solution, first, identify the key elements in the initial video frame, call the popular elements based on the key elements, then insert the popular elements into the initial video frame to obtain the target video frame, and finally generate the final video data with the target video frame, realizing the automatic sorting and automatic generation of video data containing popular elements, and improving the efficiency and video quality of the automatically generated video data.

[0013] In a possible implementation, identifying the key elements of the initial video frame includes:

[0014] Extracting the elements contained in the initial video frame to obtain an element set;

[0015] Taking the elements that are the same in the element set and the standard set as the key elements, where the standard set stores multiple elements and the popular elements corresponding to each element.

[0016] In a possible implementation, extracting the elements contained in the initial video frame to obtain an element set includes:

[0017] Extracting the text, objects, people, human postures, and backgrounds in the initial video frame to obtain the element set.

[0018] By adopting the above technical solution, the present application extracts elements such as text, objects, people, human postures, and backgrounds in the initial video frame to obtain an element set, improving the breadth and accuracy of the obtained element set. Further, by comparing the element set containing multiple elements with the elements in the standard set, and taking the elements that are the same in the element set and the standard set as the key elements, the accuracy of the obtained key elements is correspondingly improved when the accuracy and breadth of the element set are both guaranteed.

[0019] In a possible implementation, inserting the popular element into the initial video frame where the key element is located to obtain the target video frame includes:

[0020] Calculating the similarity between the popular element and the key element;

[0021] Determining whether the similarity is greater than a preset value;

[0022] If so, using the popular element to replace the key element to obtain the target video frame;

[0023] If not, inserting the popular element into the initial video frame and establishing a guiding identifier for the key element and the popular element to obtain the target video frame.

[0024] By adopting the above technical solution, after obtaining the popular elements, calculate the similarity between the popular elements and the key elements and compare the similarity with a preset value. If the similarity is greater than the preset value, it indicates that the popular elements are similar or identical to the key elements. At this time, directly replace the key elements with the popular elements to achieve the purpose of removing duplicate elements, reduce the element redundancy in the obtained target video frame, and improve the quality of the target video frame. If the similarity is less than or equal to the preset value, it indicates that the popular elements are not the same as the key elements. At this time, retain the key elements and the popular elements and establish a guiding identifier for the key elements and the popular elements, so that when playing the target video frame, after locking the key elements, the popular elements can be quickly located according to the guiding identifier, and the popular elements can attract more user attention to achieve the purpose of drainage.

[0025] In a possible implementation manner, establishing the guiding identifier for the key elements and the popular elements includes:

[0026] Connect the key elements and the popular elements with a line segment; or

[0027] Identify the key elements and the popular elements with the same color.

[0028] In a possible implementation manner, the method further includes:

[0029] Obtain a popular video frame corresponding to the key element according to the key element;

[0030] Insert the popular video frame after the initial video frame where the key element is located to obtain a target video frame.

[0031] By adopting the above technical solution, in addition to being able to insert popular elements into the initial video frame, the present application can also insert popular video frames between the initial video frames, that is, accurately locate the insertion position of the popular video frame on the video timeline at the time point when the key element appears, so as to integrate the popular video frame. While improving the quality of the video data, it also shortens the video production cycle and improves the efficiency of the automatically generated video data.

[0032] In a possible implementation manner, the method further includes:

[0033] When the initial video frame corresponds to multiple popular video frames, determine whether there is a user-set video frame among the multiple popular video frames;

[0034] If so, retain the user-set video frame as the popular video frame and insert it after the initial video frame;

[0035] If not, retain the popular video frame with the highest current popularity and insert it after the initial video frame.

[0036] By adopting the above technical solution, when the initial video frame corresponds to multiple popular video frames, the video frames set by the user are preferentially inserted after the initial video frame, so as to realize personalized generation of video data; otherwise, the video frame with the highest current popularity is retained and inserted after the initial video frame, avoiding the problem that multiple popular video frames are inserted after the initial video frame, resulting in inconsistent content between the two initial video frames before and after, ensuring natural transition and consistent content, and improving the overall quality of the video.

[0037] In the second aspect of the present application, a video automatic generation system is provided. The system includes:

[0038] A data import module for importing video data to be processed, where the video data includes at least one initial video frame;

[0039] A data recognition module for recognizing key elements of the initial video frame, where the key elements are elements to be monitored in the initial video frame;

[0040] A data matching module for obtaining popular elements corresponding to the key elements according to the key elements;

[0041] A data processing module for inserting the popular elements into the initial video frame where the key elements are located to obtain a target video frame;

[0042] A data generation module for generating and exporting final video data according to the target video frame.

[0043] In the third aspect of the present application, a video automatic generation device is provided. The device includes: a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, any of the above video automatic generation methods is implemented.

[0044] In the fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, any of the above video automatic generation methods is implemented.

[0045] In summary, the present application includes at least one of the following beneficial technical effects:

[0046] First, this application extracts elements such as text, objects, people, people's postures, and backgrounds from the initial video frame to obtain an element set. By comparing the elements in the element set containing multiple elements with the elements in the standard set, the same elements in the element set and the standard set are used as key elements to ensure the accuracy of the obtained key elements. Then, popular elements are called based on the key elements, and the popular elements are inserted into the initial video frame to obtain the target video frame. Finally, the target video frame is used to generate the final video data, realizing the automatic sorting and automatic generation of video data containing popular elements, and improving the efficiency and video quality of the automatically generated video data;

[0047] This application can also call popular video frames based on the key elements, and accurately locate the insertion position of the popular video frames on the video timeline at the time points when the key elements appear, realizing the integration of popular video frames. While improving the quality of the video data, it also shortens the video production cycle and improves the efficiency of the automatically generated video data. Description of the Drawings

[0048] The drawings, as a part of this application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application, but do not constitute an improper limitation to this application. Obviously, the drawings in the following description are only some embodiments, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0049] Figure 1 is a flowchart of a video automatic generation method in an embodiment of this application;

[0050] Figure 2 is a block diagram of a video automatic generation system in an embodiment of this application.

[0051] Description of the reference numerals: 1. Data import module; 2. Data recognition module; 3. Data matching module; 4. Data processing module; 5. Data generation module.

[0052] It should be noted that these drawings and text descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed Embodiments

[0053] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will combine the attached drawings in the embodiments of this application Figure 1-2 to clearly and completely describe the technical solutions in the embodiments.

[0054] It should be noted that the terms "including" and "having" and any variations thereof in the description, claims, and the above-mentioned drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0055] As Figure 1 shown, the present application provides a method for automatically generating a video. This method is mainly used for automatically analyzing multi-dimensional elements such as language, text, images, and human actions in unorganized video materials, and performing intelligent matching and insertion according to current popular elements. Finally, the video materials are improved and supplemented to generate the required video data, which not only improves the quality of the obtained video data but also shortens the cycle of producing video data. Specifically, the steps of the video automatic generation method are as follows.

[0056] Step S1: Import the video data to be processed, and the video data includes at least one initial video frame.

[0057] The video data to be processed is unorganized video materials uploaded by the user, and the video data to be processed includes at least one initial video frame. The present application supports the user to import the video data to be processed in different video formats, and the supported video formats include but are not limited to AVI, MP4, MOV, and WMV.

[0058] Step S2: Identify the key elements of the initial video frame, and the key elements are the elements that need to be monitored in the initial video frame.

[0059] First, extract the elements contained in the initial video frame to obtain an element set. A multi-feature analysis model based on deep learning is used to extract the elements contained in the initial video frame. The elements in the initial video frame include, but are not limited to, speech, text, objects, human postures, backgrounds, etc. The multi-feature analysis model is pre-trained. The multi-feature analysis model at least includes a text extraction sub-model, an object extraction sub-model, and a human action extraction sub-model. Among them, the text extraction sub-model is mainly used to convert the speech in the initial video frame into text, and then extract the converted text and the text that originally existed in the initial image frame to obtain text-type elements. The object extraction sub-model is mainly used to extract the objects and people in the initial video frame. Specifically, it identifies the edges of the objects or people, and then sets a marking frame around the objects or people. The marking frame selects the local area of the initial video frame where the objects or people are located. Then, the objects or people within the marking frame are used as object-type elements. Here, the marking frame is a virtual frame used to mark objects or people. The human action extraction sub-model is coupled with the object extraction sub-model. When the object extraction sub-model extracts a person, the marking frame and the person within the marking frame are sent to the human action extraction sub-model together. The human action extraction sub-model uses a pose estimation algorithm to detect and classify the actions, expressions, or other interaction behaviors of the person within the marking frame. Further, in the initial image frame, except for text, objects, and people, the remaining area is used as the background, and the background will also be used as one of the elements. The background includes information such as background size, color, and brightness.

[0060] Based on the multiple elements obtained from the multi-feature analysis model, multiple elements belonging to the same initial video frame are used as an element set. Therefore, when the video data to be processed contains multiple initial video frames, an element set corresponding to each initial video frame will be obtained, that is, multiple element sets will be obtained correspondingly.

[0061] Then, in the order of the initial video frames in the video data to be processed, check whether there are elements in the standard set that are the same as the elements in the element set corresponding to the initial video frame. If there are the same elements, then use the same elements as the key elements in the initial video frame.

[0062] It should be noted that the standard set stores multiple elements and popular elements corresponding to each element. The popular elements can be elements involved in hot topics, popular trends, popular music, and filter effects on current social media and video platforms based on big data analysis, real-time monitoring, and analysis. For example, elements such as the text of widely discussed titles and topics, the background music used in popular trends and popular music, and the special effects and filters in filter effects. That is, the popular elements include but are not limited to text, music, special effects, and filters. At the same time, the elements that trigger these popular elements or are related to these popular elements are set as the elements corresponding to the popular elements. The key to determining whether an element is a key element that triggers a popular element can be determined by the frequency of simultaneous appearance of the element and the popular element. For example, if the number of times element A and popular element B appear simultaneously in a hot topic is greater than 1000 times, it indicates that element A has a corresponding relationship with popular element B, and both element A and popular element B will be stored in the standard set. In practical applications, other judgment methods can also be used to determine whether an element is a key element that triggers a popular element, and this application does not make any restrictions.

[0063] It should also be noted that the elements in the standard set can also correspond to video frames. The judgment condition is also that when the number of times an element and a video frame appear in a hot topic is greater than a preset number, the element and the video frame are bound, and both of them and their binding relationship are stored in the standard set. For the sake of easy distinction, the video frames stored in the standard set are called popular video frames.

[0064] In practical applications, the popular elements and popular video frames in the standard set can also be video frames that are pre-organized and cropped by users, but the specific insertion positions are not determined by users. Therefore, they are preferentially stored in the standard set. When an element corresponding to the popular element and the popular video frame appears, the popular element and the popular video frame are called and the insertion operation is executed.

[0065] Therefore, according to the elements in the initial video frame, popular elements and / or popular video frames are matched from the standard set.

[0066] In addition, the elements in the standard set, the popular elements corresponding to the elements, and the popular video frames are updated periodically. That is, after a period of time, the elements involved in hot topics, popular trends, popular music, and filter effects on current social media and video platforms are monitored and analyzed again, and the elements in the standard set, the popular elements corresponding to the elements, and the popular video frames are dynamically updated according to the monitoring results.

[0067] Step S3: Obtain the popular elements and / or popular video frames corresponding to the key elements according to the key elements.

[0068] Since the standard set stores elements, popular elements corresponding to the elements, and popular video frames, and the key elements refer to the elements in the element set that are the same as the elements in the standard set, the key elements also have a corresponding relationship with the popular elements and popular video frames. Therefore, after obtaining the key elements, according to the corresponding relationship between the elements, popular elements, and popular video frames in the standard set, the corresponding popular elements and / or popular video frames of the key elements are obtained.

[0069] Step S4: Insert the popular element into the initial video frame where the key element is located and / or insert the popular video frame after the initial video frame where the key element is located to obtain the target video frame.

[0070] The process of inserting the popular element is as follows:

[0071] First, calculate the similarity between the popular element and the key element. This application uses a similarity calculation model based on deep learning, which is obtained after being pre-trained and optimized. The similarity calculation model can analyze and calculate the similarity between the key element and the popular element. The closer the key element and the popular element are to being the same, the higher the similarity. For example, if the key element is text and the popular element is also text, then the two are related. If the meanings represented by the two texts are the same or the texts are exactly the same, the similarity is 100%.

[0072] Then, determine whether the similarity between the popular element and the key element is greater than a preset value. The preset value is set in advance and is used to evaluate whether the key element and the popular element are similar or the same. The preset value set in this application is 80%. That is, when the similarity between the popular element and the key element is greater than 80%, it means that the two are similar or the same and can be substituted for each other.

[0073] If the comparison result shows that the similarity is greater than the preset value, it means that the popular element and the key element are similar or the same. At this time, directly replace the key element with the popular element, insert the popular element into the position where the key element is located, and use the initial video frame with the popular element inserted as the target video frame, so that there are no duplicate elements in the target video frame, reducing the element redundancy in the obtained target video frame and improving the quality of the target video frame;

[0074] If the comparison result shows that the similarity is less than or equal to the preset value, it indicates that the popular element and the key element are different. In this case, both the key element and the popular element are retained, and a guiding identifier for the key element and the popular element is established. The establishment of this guiding identifier includes, but is not limited to: connecting the key element and the popular element with a line segment between the key element and the popular element, or using the same color to identify the key element and the popular element. Then, the initial video frame with the guiding identifier for the key element and the popular element established is used as the target video frame. In this way, when playing this target video frame subsequently, after locking the key element, the popular element can be quickly located according to this guiding identifier, attracting more user attention with the popular element, enhancing the influence and dissemination efficiency of video data, and achieving the purpose of drainage.

[0075] In practical applications, the popular element can also be set to float above the key element, or the popular element is connected to the key element. For example, when the human gesture is "OK", the popular element is placed in the area adjacent to the gesture, so that the popular element can be quickly locked based on the human gesture.

[0076] The insertion process of the popular video frame is as follows:

[0077] After obtaining the popular video frame, the popular video frame is inserted after the initial video frame where the key element is located, and the inserted popular video frame is used as the target video frame. Therefore, this application can accurately locate the insertion position of the popular video frame on the video timeline at the time point when the key element appears, shortening the video production cycle.

[0078] It should be noted that when there are multiple key elements on the same initial video frame, and each of the multiple key elements corresponds to a different popular video frame, it is necessary to determine whether the popular video frame is the video frame set by the user. If so, the user-defined video frame is preferentially inserted after the initial video frame. Otherwise, the video frame with the highest current popularity of the popular video frame is retained and inserted after the initial video frame, avoiding the problem of inconsistent content between the two consecutive initial video frames caused by inserting multiple popular video frames after the initial video frame, ensuring natural transition and coherent content, and improving the overall quality of the video.

[0079] Step S5: Generate and export the final video data according to the target video frame.

[0080] After obtaining the target video frame, the original position where the initial video frame is located is replaced with the target video frame, or the target video frame is inserted after the initial video frame, and then output in sequence according to the order of the video frames, so as to obtain the final video data.

[0081] In summary, the implementation principle of a video automatic generation method in this application is as follows: First, elements such as text, objects, people, people's postures, and backgrounds in the initial video frame are extracted to obtain an element set. The elements in the element set containing multiple elements are compared with the elements in the standard set, and the same elements in the element set and the standard set are used as key elements to ensure the accuracy of the obtained key elements. Then, popular elements and / or popular video frames are called based on the key elements, and the popular elements are inserted into the initial video frame where the key elements are located and / or the popular video frames are inserted after the initial video frame where the key elements are located to obtain target video frames. Finally, the target video frames are used as the basis for generating the final video data, realizing the automatic sorting and automatic generation of video data containing popular elements and popular video frames, and improving the efficiency and video quality of the automatically generated video data.

[0082] As Figure 2 shown, this application provides a video automatic generation system, which includes a data import module 1, a data recognition module 2, a data matching module 3, a data processing module 4, and a data generation module 5.

[0083] The data import module 1 is used to import video data to be processed, and the video data includes at least one initial video frame.

[0084] The data recognition module 2 is used to recognize the key elements of the initial video frame, and the key elements are the elements to be monitored in the initial video frame.

[0085] The data matching module 3 is used to obtain the popular elements corresponding to the key elements according to the key elements.

[0086] The data processing module 4 is used to insert the popular elements into the initial video frame where the key elements are located to obtain target video frames.

[0087] The data generation module 5 is used to generate and export the final video data according to the target video frames.

[0088] The modules involved in the embodiments described in this application can be implemented in software or in hardware. The described modules can also be set in a processor. For example, it can be described as: A processor includes a data import module 1, a data recognition module 2, a data matching module 3, a data processing module 4, and a data generation module 5. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the data import module 1 can also be described as "a module for importing video data to be processed".

[0089] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0090] To better execute the procedures of the above method, the present application also provides a video automatic generation device, which includes a memory and a processor.

[0091] Among them, the memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory can include a program storage area and a data storage area. The program storage area can store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the above video automatic generation method, etc.; the data storage area can store data involved in the above video automatic generation method, etc.

[0092] The processor can include one or more processing cores. By running or executing instructions, programs, code sets or instruction sets stored in the memory, the processor calls the data stored in the memory and executes various functions of the present application and processes data. The processor can be at least one of an application specific integrated circuit, a digital signal processor, a digital signal processing device, a programmable logic device, a field programmable gate array, a central processing unit, a controller, a microcontroller and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above processor functions can also be others, and the embodiments of the present application do not make specific limitations.

[0093] The present application also provides a computer-readable storage medium, such as including: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs, etc. The computer-readable storage medium stores a computer program that can be loaded and executed by the processor to implement the above video automatic generation method.

[0094] The above are only the preferred embodiments of the present application, and do not impose any form of limitation on the present application. Although the present application has been disclosed above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications into equivalent embodiments by using the technical content prompted above within the scope of the technical solution of the present application. The implementation schemes in the above embodiments can also be further combined or replaced. However, as long as it does not deviate from the technical solution of the present application, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application still fall within the scope of the present application.

Claims

1. A video automatic generation method, the method is used for unorganized video materials, characterized in that: include: Importing video data to be processed, wherein the video data includes at least one initial video frame; Identifying key elements of the initial video frame, where the key elements are elements that need to be monitored in the initial video frame, including: extracting elements contained in the initial video frame to obtain an element set; using the same elements in the element set and the standard set as key elements, where the standard set stores multiple elements and popular elements and / or popular video frames corresponding to each element, monitoring and analyzing the elements involved in hot topics, popular trends, popular music, and filter effects on current social media and video platforms, and dynamically updating the elements in the standard set and the popular elements and / or popular video frames corresponding to the elements according to the monitoring results, where the popular elements at least include text, music, special effects, and filters; Obtaining, according to the key element, a popular element and / or a popular video frame corresponding to the key element; Inserting the popular element into the initial video frame where the key element is located and / or inserting the popular video frame into the initial video frame where the key element is located to obtain a target video frame; wherein inserting the popular element into the initial video frame where the key element is located to obtain the target video frame includes: calculating the similarity between the popular element and the key element; judging whether the similarity is greater than a preset value; if so, replacing the key element with the popular element to obtain the target video frame; if not, inserting the popular element into the initial video frame and establishing a guide mark of the key element and the popular element, including: connecting the key element and the popular element with a line segment; or marking the key element and the popular element with the same color to obtain the target video frame; wherein inserting the popular video frame into the initial video frame where the key element is located to obtain the target video frame includes: when multiple key elements appear on the same initial video frame, and the multiple key elements correspond to different popular video frames respectively, judging whether the popular video frame is a video frame set by a user, and when the popular video frame is not a video frame set by the user, retaining the video frame with the highest current popularity among the popular video frames; Final video data is generated and exported according to the target video frame.

2. The video automatic generation method according to claim 1, characterized in that: The step of extracting the elements contained in the initial video frame to obtain an element set includes: The element set is obtained by extracting text, objects, characters, character postures, and background from the initial video frame.

3. A video automatic generation system, characterized in that: include: A data import module (1), used for importing video data to be processed, wherein the video data comprises at least one initial video frame; A data identification module (2) is used to identify key elements of the initial video frame, wherein the key elements are elements that need to be monitored in the initial video frame, including: extracting elements contained in the initial video frame to obtain an element set; using the same elements in the element set and the standard set as key elements, wherein the standard set stores multiple elements and popular elements and / or popular video frames corresponding to each element, monitoring and analyzing elements involved in hot topics, popular trends, popular music, and filter effects on current social media and video platforms, and dynamically updating the elements in the standard set and the popular elements and / or popular video frames corresponding to the elements according to the monitoring results, wherein the popular elements at least include text, music, special effects, and filters; A data matching module (3), configured to obtain, according to the key element, a popular element and / or a popular video frame corresponding to the key element; The data processing module (4) is used to insert the popular element into the initial video frame where the key element is located and / or insert the popular video frame after the initial video frame where the key element is located to obtain a target video frame; wherein inserting the popular element into the initial video frame where the key element is located to obtain the target video frame comprises: calculating the similarity between the popular element and the key element; judging whether the similarity is greater than a preset value; if so, replacing the key element with the popular element to obtain the target video frame; if not, inserting the popular element into the initial video frame and establishing a guide mark of the key element and the popular element, comprising: connecting the key element and the popular element with a line segment; or marking the key element and the popular element with the same color to obtain the target video frame; wherein inserting the popular video frame after the initial video frame where the key element is located to obtain the target video frame comprises: when multiple key elements appear on the same initial video frame and the multiple key elements correspond to different popular video frames respectively, judging whether the popular video frame is a video frame set by a user, and when the popular video frame is not a video frame set by the user, retaining the video frame with the highest current popularity among the popular video frames; A data generation module (5) is used to generate and export final video data according to the target video frame.

4. A video automatic generation device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method according to any one of claims 1 to 2 is implemented.

5. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.

Citation Information

Patent Citations

  • Target object tracking method based on color-structure features

    CN104240266A

  • Video file generation method and terminal

    CN106803909A

  • Video synthesis method and device, electronic equipment and computer readable storage medium

    CN110958386A