A method for making a video special effect cover based on artificial intelligence

By collecting behavioral data to generate popularity curves, constructing dynamic graph networks, and using conditional generative adversarial networks to optimize cover effects, this approach solves the problems of low efficiency and homogenization in the manual production of online education video special effects covers. It achieves intelligent and automated cover generation, improving the attractiveness of the covers and the learning conversion rate.

CN120512573BActive Publication Date: 2025-12-05BEIJING BLUE DIAMOND CULTURE MEDIA CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510859721.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-12-05
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the field of online education videos, the manual production of video effects covers is inefficient, suffers from severe homogenization, lacks self-learning and continuous optimization, resulting in insufficient cover appeal, low click-through rates, and inadequate learning conversion rates.

Method used

By collecting real-time behavioral data to generate time-series heat maps, performing multimodal alignment processing, constructing a dynamic graph network, generating candidate covers, collecting user interaction feedback, strengthening the cover learning iterative model, and using conditional generative adversarial networks to optimize cover generation.

Benefits of technology

It has enabled the intelligent and automated creation of video special effects covers, improving production efficiency, quickly responding to user preferences and content changes, and increasing the click-through rate and learning conversion rate of covers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512573B_ABST
    Figure CN120512573B_ABST
Patent Text Reader

Abstract

The application belongs to the field of video production and provides a video special effect cover production method based on artificial intelligence, which collects user real-time behavior data by setting a buried point between a video page and a player, generates a video heat curve graph, filters peak periods of the heat curve, extracts feature vectors of different modal data, maps them to a unified dimension space to form aligned vectors input into a classifier, extracts video semantic labels, establishes a node and edge set, integrates training data, calculates edge weights, dynamically adjusts edge weights, constructs a dynamic graph network, inputs a conditional vector into a conditional generative adversarial network, generates multiple candidate covers, maintains candidate cover sequences of each video, uploads them to an online education video page, polls covers at regular intervals, collects user interaction feedback and iteratively optimizes the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of video production, and particularly relates to a video special effect cover production method based on artificial intelligence. BACKGROUND

[0002] The video special effect cover refers to a video thumbnail that is beautified and attractive by adding dynamic visual effects before the video is published or displayed.

[0003] In China, the video special effect cover production method, device, electronic equipment and storage medium are disclosed in CN117714774B. The configuration method comprises the following steps: acquiring an original video, displaying the original video and a corresponding time axis through a play control of a client to sequentially represent the play time of a picture frame; selecting a picture frame corresponding to a target play time to generate an initial dynamic picture; performing a special effect selection operation on a special effect, rendering the initial dynamic picture according to the selected target special effect to obtain a preview dynamic picture; displaying the preview dynamic picture through the play control, and responding to an export operation of the client to export the preview dynamic picture as a target file; and solving the problem that multiple operation interfaces are required to edit and process a video when a video cover is produced.

[0004] In the field of online education videos, there is a diversified and intelligent design demand for video special effect covers. Although there are methods to solve the problem of complicated production operation process, the existing production methods are still mainly manual, and in the face of new courses online or new segments added to original courses every day, due to the complicated and time-consuming links, the labor cost is high, and only standard templates can be used for production, resulting in the problem of cover style homogeneity. On the other hand, the produced video special effect cover is not attractive enough, and it is difficult to catch the eyes of learners, resulting in low click rate. Covers that are too general or do not match the core content of the course cannot effectively remind the importance of the section, so the learning conversion rate is insufficient. After the cover is produced, there is a lack of a closed-loop feedback method for self-learning and continuous optimization. SUMMARY

[0005] The application provides a video special effect cover production method based on artificial intelligence, which aims to solve the problems of low efficiency, serious homogeneity and lack of self-learning and continuous optimization of the cover generation model in manual production of video special effect covers.

[0006] To solve the above technical problems, the application provides a video special effect cover production method based on artificial intelligence: collecting real-time behavior data to generate a time series heat curve; performing multi-modal alignment processing to extract video semantic tags; constructing a dynamic graph network to generate corresponding candidate covers; collecting user interaction feedback to strengthen the cover learning iteration model.

[0007] As a preferred embodiment, the specific steps of collecting real-time behavior data are: assigning a unique identifier resourceId to each educational video resource in the online education background, setting JS buries in the online education page and video player in the front end, uploading behavior events in real time, using integrated player SDK, returning SDK logs to the background server, the background server uses the default window size of 2000 to pull logs, checks the mandatory fields such as resource_id in the logs, marks the logs missing mandatory fields, defines a legal event type table {play, pause, finish, click} in advance, wherein play represents playing, pause represents pausing, finish represents closing, and click represents clicking the video cover, verifies the predefined enumeration value for each log, processes illegal logs whose event type is not in the type table, maps the event type to an integer value {play: 1, pause: 2, finish: 3, click: 4}, unifies the timestamp to seconds, uses the timestamp as the primary key, removes high-frequency redundant events in the logs, normalizes the text content to UTF-8, removes invisible characters, and encapsulates the processed log data into a structured object.

[0008] As a preferred embodiment, the specific steps of generating a timing heat curve are: assigning weights to event types, the playing type is the most important, the weight is set to 0.6, the pause type is secondary, the weight is set to 0.3, the closing type is special, when the closing event occurs, it has a negative impact on the weight, and the weight is set to -0.1, dividing the video into different time windows with a fixed length window of 1s, dividing events into corresponding time windows, the window identifier is windowId, for each windowId event, calculating the heat value according to the number of event types and the type, the heat value calculation formula is:

[0009] heatScore(windowId)=total(1)×0.6+total(2)×0.3+

[0010] total(3)×(-0.1), wherein heatScore is the heat value, total is used to count the number of event types in the time window, the window identifier and the weight are combined into a structured object, the heat sequence is arranged in ascending order of time, the sliding average method is used to eliminate the occasional peaks of the heat sequence, and the heat sequence is input into the drawing engine to render the curve.

[0011] As a preferred embodiment, the specific steps of performing the multi-modal alignment processing are: calculating the mean mu and standard deviation sigma in the heat curve, setting the heat value greater than mu+sigma in the heat sequence as the peak point, merging adjacent peak points into peak periods to form a list of non-overlapping peak periods, unifying the timestamps of the peak periods in the list, extracting the feature vectors of the different modal data corresponding to the timestamps, including visual feature vectors, audio feature vectors, and text feature vectors, introducing linear projection layers to the feature vectors respectively, mapping the feature vectors of different modal data to a unified dimensional space, performing layer normalization and linear activation on the mapping results to enhance feature separability, and outputting multi-modal alignment vectors.

[0012] As a preferred embodiment, the specific steps of extracting video semantic labels are: pre-constructing a semantic ontology of online education scenarios, defining a label set, such as formula derivation, experimental demonstration, difficult point summary, knowledge review, etc., sorting the multi-modal alignment vectors by timestamp to form a fusion vector sequence where t represents the timestamp, alignVector represents the alignment vector, and for each time t i , merging the alignment vectors of the previous and subsequent W frames to form context input As input to the multi-modal semantic classifier, it is mapped to the semantic space through the feedforward layer and activation function, and outputs the prediction distribution for the pre-defined label set. A threshold of 0.6 is set for the pre-defined labels as a compromise point suitable for online education scenarios that consider accuracy and completeness, without missing important labels while effectively suppressing low-confidence labels. The results of each output of the classifier are compared with the threshold, and the labels greater than the set threshold are selected. Then the labels are merged with synonyms and hierarchical, and the unified labels are merged according to the pre-constructed semantic of online education scenarios. The labels adjacent in time are merged into continuous intervals, and the core labels in the continuous intervals are obtained by counting the frequency of the labels.

[0013] As a preferred embodiment, the specific steps of constructing a dynamic graph network are: assigning a unique identifier to each online education video and establishing a video node v for the video. The video node is corresponding to the video content through the unique identifier. The semantic labels extracted are established as label nodes l, and the name of the label node is corresponding to the video scenario. At the same time, event nodes e are established for event types, and the numerical value of the event type is corresponding to the event node. The three types of nodes are integrated into training data; for each pair of video nodes and label nodes, video nodes and label nodes with co-occurrence relationship are established as edges E v,l , the co-occurrence relationship is the corresponding multi-modal feature vector and video time period, and the weight of edge E v,l is calculated, and the formula is where w v,l represents edge Ev,l weight of edge E v,l represents the confidence of the corresponding segment of video v being predicted as label l, count v,l represents the number of times label l appears in video v, ∑ l count v,l represents the sum of all label appearances in video v, for each pair of event node and label node, an edge E l,e exists between the label node and the event node if there is frequent touch, that is, after a user performs a certain event, a certain label has a high probability of being extracted in the current time period, the edge E l,e is calculated, and the weight of the edge E where w l,e represents the weight of the edge E l,e , m l,e represents the number of times label l appears after event e is triggered, ∑ e m l,e represents the sum of the number of times label l appears after all events are triggered, for each pair of video node and event node, an edge E v,e is established between the video node and the event node if there is a triggering behavior, that is, the video node has performed a certain event when playing or using, the edge E v,e is calculated, and the weight of the edge E where w v,e represents the weight of the edge E v,e , n v,e represents the number of times video v triggers event e, ∑ e n v,e represents the sum of all triggering event times in video v; traverse the event node, the video node, the label node, the edge set, and the edge weight set to construct a graph model, and dynamically adjust w v,l according to the heat, periodically count the specific execution events of a user in a certain video to adjust w v,e , count the number of times the label is associated with the event to adjust w l,e , update the adjusted edge weight and timestamp to the graph model, retain the update record for training, and form a dynamic network graph network by continuously evolving.

[0014] As a preferred embodiment, the specific step of generating the corresponding candidate cover is: finding the label node and event node with the strongest association from the target video node, selecting the popular label node according to the edge weight, and forming a condition vector with the video frame corresponding to the timestamp and the heat value, and inputting the condition vector into the conditional generative adversarial network. The conditional generative adversarial network continuously adjusts its parameters through the generator, improves the discrimination ability of the judge to true and false samples, and finally generates the cover candidate by the generator. Automatically add special effect elements to the video frame, and combine edge weight, event node and label node for decision-making of text title, layout mode, etc.

[0015] As a preferred embodiment, the specific step of collecting user interaction feedback is: storing the generated candidate cover and the condition vector corresponding to the video resource unique identifier into the cover library, maintaining a sequence of current candidate covers for each online education video, rotating the online education video cover once every period, burying points in the cover display interface, listening to the user's interaction action on the current candidate cover, and recording the event type executed by the user after playing the video. After receiving the data in the background, write it to the message queue, process it using the stream processing framework, calculate the click rate and complete play rate, and return the result to the conditional generative adversarial network.

[0016] As a preferred embodiment, the specific step of strengthening the cover learning iteration model is: combining the condition vector of the candidate cover in the dynamic graph and the information of the user interaction feedback, setting a reward function, taking the user interaction feedback as a reward and punishment signal, calculating the reward value, and the reward value formula is: award = 0.7 x R click + 0.3 x average dwell , wherein award represents the reward value, R click is the click rate, average dwell is the average dwell time. The calculated reward value is fed back to the conditional generative adversarial network as a reward signal to update the parameters in the generator. After each display, feedback and update, the candidate cover is put into actual testing to enable the cover generation model to continuously learn and optimize itself.

[0017] The beneficial effects of the present application are:

[0018] 1. The present application realizes the intelligence and automation of video special effect cover production, generates a heat curve by collecting user behavior data, constructs a dynamic graph, avoids subjective bias in manual production, and improves the efficiency of dealing with different online education video cover production.

[0019] 2. Through the combination of conditional generative adversarial network, multiple sets of candidate covers are generated, a reinforcement learning strategy model is constructed, user interaction behavior is taken as a reward signal, the model is continuously iterated and optimized, user preferences and online education content changes can be quickly responded to, and the click rate and learning conversion rate are effectively improved.

[0020] Legend

[0021] Figure 1 It is a flow chart of a video special effect cover production method based on artificial intelligence.

[0022] Figure 2 It is a thermal curve diagram of a video special effect cover production method based on artificial intelligence. DETAILED DESCRIPTION

[0023] In order to make the technical means, creative features, purposes and effects realized by the present application easy to understand, the present application will be further described below in combination with specific embodiments, but the following embodiments are only preferred embodiments of the present application, not all. Based on the embodiments in the embodiments, other embodiments obtained by those skilled in the art without creative labor also belong to the protection scope of the present application.

[0024] Embodiment 1, as Figure 1 It is a video special effect cover production method based on artificial intelligence. Real-time behavior data is collected between the background and online education video to generate a heat curve diagram of the video, multi-modal data with high heat value is aligned and processed, and video semantic labels are extracted. The specific implementation steps are as follows:

[0025] Step one, collect real-time behavior data and generate a time series heat curve diagram;

[0026] Step two, perform multi-modal alignment processing and extract video semantic labels;

[0027] Step three, construct a dynamic graph network and generate corresponding candidate covers;

[0028] Step four, collect user interaction feedback and strengthen cover learning iteration model;

[0029] The application discloses a video special effect cover production method based on artificial intelligence, collects real-time behavior data, and generates a time sequence heat curve diagram, wherein the specific steps of collecting real-time behavior data are as follows: a unique identifier resourceID is allocated to each education video resource in an online education background, JS burying points are set in a front-end online education page and a video player, behavior events are uploaded in real time, a set player SDK is utilized, SDK logs are returned to a background server, the background server adopts a default 2000-window size to pull logs, compulsory fields such as resource_id in the logs are checked, logs lacking compulsory fields are marked, a legal event type table {play, pause, finish, click} is defined in advance, wherein play represents playing, pause represents pausing, finish represents closing, and click represents clicking a video cover, the predefined enumeration values are verified for each log, illegal logs whose event types are not in the type table are processed, the event types are mapped to integer values {play: 1, pause: 2, finish: 3, click: 4}, timestamps are unified to seconds, the logs are removed from high-frequency redundant events, text contents are subjected to UTF-8 normalization processing, invisible characters are removed, and the processed log data is encapsulated into a structured object; the specific steps of generating the time sequence heat curve diagram are as follows: weights are allocated to event types, the playing type is the most important, the weight is set to 0.6, the pausing type is secondary, the weight is set to 0.3, the closing type is relatively special, the closing event has a negative influence on the weight, and the weight is set to-0.1, a fixed length window 1s is taken as a unit, the video is divided into different time windows, events are divided into corresponding time windows, the window identifier is windowId, the events occurring in each windowId are calculated according to the event type occurrence frequency and the type, the heat value is calculated according to the following formula:

[0030] heatScore(windowId)=total(1)×0.6+total(2)×0.3+

[0031] total(3)×(-0.1), wherein heatScore is the heat value, total is used to count the event type occurrence frequency in the time window, the window identifier and the weight are combined into a structured object, the heat sequence is arranged in ascending order according to time, the sliding average method is utilized to eliminate the accidental peaks of the heat sequence, the heat sequence is input into a drawing engine, a curve diagram is rendered, and the heat curve diagram data is shown in Table 1.

[0032]

[0033]

[0034] Table 1

[0035] A heat curve graph drawn according to Table 1 is shown in FIG. 1, in which the horizontal axis represents windowId and the vertical axis represents the heat score. Figure 2

[0036] Based on the above steps, the specific steps of performing multi-modal alignment processing are as follows: calculating the mean μ and standard deviation σ in the heat curve, setting the heat value greater than μ+σ in the heat sequence as a peak point, merging adjacent peak points into a peak period to form a list of non-overlapping peak periods, unifying the timestamps of the peak periods in the list, extracting the feature vectors of the different modal data corresponding to the timestamps, including visual feature vectors, audio feature vectors, and text feature vectors, introducing linear projection layers to the feature vectors respectively, mapping the feature vectors of different modal data to a unified dimensional space, performing layer normalization and linear activation on the mapping results to enhance feature separability, and outputting a multi-modal alignment vector;

[0037] The specific steps of extracting video semantic tags are as follows: constructing a semantic ontology of online education scenarios in advance, defining a tag set, such as formula derivation, experimental demonstration, difficulty summary, knowledge review, etc., sorting the multi-modal alignment vectors according to the timestamps to form a fusion vector sequence where t represents the timestamp, alignVector represents the alignment vector, and for each time t i , the alignment vectors of the previous and subsequent W frames are merged to form a context input as input to the multi-modal semantic classifier, which is mapped to a semantic space through a feedforward layer and an activation function, and outputs a prediction distribution for the predefined tag set. A threshold of 0.6 is set for the predefined tags as a compromise point suitable for online education scenarios that consider accuracy and completeness, without missing important tags while effectively suppressing low-confidence tags. The results of each output of the classifier are compared with the threshold, and the tags greater than the set threshold are selected. The tags are then merged with synonyms and hierarchical, and the timestamps of adjacent tags are merged into continuous intervals. The frequency of the tags is counted to obtain the core tags in the continuous intervals.

[0038] Based on the above steps, the specific steps of constructing a dynamic graph network are as follows: assigning a unique identifier to each online education video and establishing a video node v for the video, the video node corresponding to the video content through the unique identifier, establishing a label node l for the extracted semantic tags, the name of the label node corresponding to the video scene, establishing an event node e for the event type, the value of the event type corresponding to the event node, and integrating the three types of nodes into training data; for each pair of video nodes and label nodes, establishing an edge E v,l between the video nodes and label nodes that have a co-occurrence relationship, i.e., the aligned multi-modal feature vectors correspond to the video time period, calculating the weight of the edge E v,l , and the formula is​ where w v,l is the weight of edge E v,l , p v,l is the confidence of the corresponding segment of video v being predicted as label l, count v,l is the number of times label l appears in video v, ∑ l count v,l is the sum of all labels appearing in video v, between each pair of event node and label node, there is an edge E l,e established if there is frequent touch between the label node and the event node, which means that after a user performs a certain event, a certain label has a high probability of being extracted in the current time period, the weight of edge E l,e is calculated, and the formula is where w l,e is the weight of edge E l,e , m l,e is the number of times label l appears after event e is triggered, ∑ e m l,e is the sum of the number of times label l appears after all events are triggered, between each pair of video node and event node, there is an edge E v,e established if there is a triggering behavior between the video node and the event node, which means that a certain event has been performed when the video node is playing or being used, the weight of edge E v,e is calculated, and the formula is where w v,e is the weight of edge E v,e , n v,e is the number of times video v triggers event e, ∑ e n v,e is the sum of the number of times all triggering events in video v, traverse the event nodes, video nodes, label nodes, edge set, and weight set of each edge, construct the graph model, and dynamically adjust w v,l according to the heat, periodically count the specific execution events of a user in a certain video to adjust w v,e , count the number of times labels are associated with events to adjust w l,e , update the adjusted edge weight and timestamp to the graph model, keep the update record for training, and form a dynamic graph network by continuously evolving.

[0039] The specific steps for generating the corresponding candidate cover are: finding the label node and event node most closely associated with the target video node, selecting the popular label node according to the edge weight, and forming a condition vector with the video frame corresponding to the timestamp and the heat value, inputting the condition vector into the conditional generative adversarial network, the conditional generative adversarial network continuously adjusts its parameters through the generator, the discriminator improves the ability to distinguish between true and false samples, and finally the generator generates the cover candidate, automatically adds special effect elements to the video frame, and combines the edge weight with the event node and the label node to make decisions on the text title and layout mode.

[0040] Based on the above steps, the specific steps for collecting user interaction feedback are: storing the generated candidate cover and the condition vector corresponding to the video resource unique identifier in the cover library, maintaining a sequence of current candidate covers for each online education video, rotating the online education video cover once regularly, burying points in the cover display interface, listening to user interaction actions on the current candidate cover, recording the event types executed by the user after playing, writing messages to the message queue after receiving data in the background, processing using the stream processing framework, calculating key indicators such as click rate and complete play rate, and returning the results to the conditional generative adversarial network; the specific steps for strengthening the cover learning iteration model are: combining the condition vector of the candidate cover in the dynamic graph and the user interaction feedback information, setting a reward function, using user interaction feedback as a reward and punishment signal, calculating the reward value, and the reward value formula is: award = 0.7 x R click + 0.3 x average dwell , where award represents the reward value, R click is the click rate, average dwell is the average dwell time, the calculated reward value is used as a reward signal to feed back to the conditional generative adversarial network, and the parameters in the generator are updated, and after each display, feedback and update, the candidate cover is put into actual testing to enable the cover generation model to continuously learn and optimize itself.

[0041] The above describes embodiments of the present application, and without departing from the embodiments of the present application and its broader aspects, those skilled in the art can make data modifications and method changes based on this place, the appended claims are for all such data modifications and method changes that do not deviate from the embodiments of the present application.

Claims

1. A method for creating video special effects covers based on artificial intelligence, characterized in that: Collecting real-time behavioral data and generating time-series heat curves involves cleaning and deduplicating the collected real-time user behavior data, calculating the heat value within a time window, and generating a time-series heat curve. Performing multimodal alignment processing and extracting video semantic tags refers to selecting the peak time period in the popularity curve, extracting data from different modalities, performing alignment processing, mapping to a unified dimension, and then extracting semantic tags. Constructing a dynamic graph network to generate corresponding candidate covers involves assigning unique identifiers to online education videos and establishing video nodes (v) for each video. These video nodes correspond to the video content through their unique identifiers. Extracted semantic tags are used to create tag nodes (l), whose names correspond to video scenes. Event nodes (e) are created for each event type, with their numerical values ​​corresponding to the event nodes. These three types of nodes are integrated into training data, and edges (E) are established between video nodes and tag nodes that co-occur. v,l Calculate edge E v,l The weights are given by the formula: Where w v,l Representing edge E v,l The weights, p v,l This represents the confidence that the segment corresponding to video v is predicted to be label l, count. v,l ∑ represents the number of times label l appears in video v. l count v,l E represents the sum of all occurrences of tags in video v, and establishes edges E for tag nodes and event nodes that frequently interact. l,e Calculate edge E l,e The weights are given by the formula: Where w l,e Representing edge E l,e The weight, m l,e ∑ represents the number of times label l appears after event e is triggered. e m l,e The sum of the number of times label l appears after all events are triggered is used to establish an edge E for video nodes and event nodes that have triggering behaviors. v,e Calculate edge E v,e The weights are given by the formula: Where w v,e Representing edge E v,e The weights, n v,e ∑ represents the number of times video v triggers event e. e n v,e The sum of the number of times all events are triggered in video v is represented by the graph model constructed by traversing event nodes, video nodes, tag nodes, and collecting edges and their weights, and dynamically adjusting w based on the popularity value. v,l Adjust w by counting the number of events executed in online education videos. v,e Adjust the number of times the statistical tags are associated with the event. l,e The adjusted edge weights and timestamps are updated to the graph model to form a dynamic graph network. Nodes are selected based on the edge weights to form conditional vectors, which are then input into a conditional generative adversarial network to generate multiple candidate video covers. Collecting user interaction feedback to strengthen the cover learning iterative model involves uploading the generated candidate video covers to the online education video page and collecting user interaction feedback to iteratively optimize the model.

2. The method for creating a video special effects cover based on artificial intelligence according to claim 1, characterized in that: The specific steps for collecting real-time behavioral data are as follows: assign a unique identifier to the online education video resource, set up JS tracking points between the page and the video player, upload behavioral events in real time, use the integrated player SDK to send logs back to the backend server, check the unique identifiers in the logs, mark logs that lack unique identifiers, predefine a table of valid event types, verify the predefined event type values ​​for each log, process logs whose event types are not in the type table, map the event types to integer values ​​{play: 1, pause: 2, finish: 3, click: 4}, remove high-frequency redundant events from the logs, perform UTF-8 normalization on the text content, remove invisible characters, and encapsulate the processed log data into a structured object. The specific steps for generating the time-series heatmap are as follows: Assign weights to event types, setting the weight for playback type to 0.6, pause type to 0.3, and close type to -0.1; divide the video into 1-second units; generate time windows; assign event types to their corresponding time windows, with each time window identified by a windowId; and calculate the heat value for each event occurring within a windowId based on the frequency and type of the event. The heat value calculation formula is: heatScore(windowId)=total(1)×0.6+total(2)×0.3+ total(3)×(-0.1), where heatScore is the heat value, and total is used to count the number of times this type of event occurs within the time window. The time window identifier and heat value are combined into a structured object, arranged in ascending order of time to form a heat sequence. The occasional spikes in the heat sequence are eliminated by using a moving average method, and then passed into the drawing engine to render the curve.

3. The method for creating a video special effects cover based on artificial intelligence according to claim 1, characterized in that: The specific steps for performing multimodal alignment processing are as follows: calculate the mean μ and standard deviation σ in the heat curve, set the heat values ​​greater than μ+σ in the heat sequence as peak points, merge adjacent peak points into peak time periods, form a peak time period list, unify the timestamps of the peak time period list, extract feature vectors of different modal data within the corresponding timestamps, including visual feature vectors, audio feature vectors, and text feature vectors, introduce linear projection layers into the feature vectors respectively, map the feature vectors of different modal data to a unified dimensional space, perform layer normalization and linear activation on the mapping results, and output multimodal alignment vectors; The specific steps for extracting semantic tags from videos are as follows: First, a semantic ontology for the online education scenario is pre-constructed; then, a tag set is defined; and finally, multimodal alignment vectors are sorted by timestamp to form a fused vector sequence. Where t represents the timestamp, and alignVector represents the alignment vector. For each time t... i The alignment vectors of the W frames before and after merging are used to form the context input. The input is a multimodal semantic classifier, which is mapped to the semantic space through a feedforward layer and an activation function, and outputs the predicted distribution of a predefined set of labels.

4. The method for creating a video special effects cover based on artificial intelligence according to claim 3, characterized in that: The specific steps for extracting video semantic tags also include: setting a threshold of 0.6 for predefined tags, comparing the results of the classifier output with the threshold, selecting tags that are greater than the set threshold, merging the tags into synonyms and levels, merging them into unified tags based on the semantics of the pre-constructed online education scenario, merging tags with adjacent timestamps into continuous intervals, counting the frequency of tag occurrences, and obtaining the core tags within the continuous intervals.

5. The method for creating a video special effects cover based on artificial intelligence according to claim 1, characterized in that: The specific steps for generating corresponding candidate covers are as follows: Filter tag nodes and event nodes from video nodes, select popular tag nodes based on edge weights, form a conditional vector with the corresponding timestamp video frame and popularity value, input it into a conditional generative adversarial network, the conditional generative adversarial network adjusts its own parameters through the generator, the judge identifies real and fake samples, generates multiple cover candidates, and automatically adds special effects elements to the video frames, combining edge weights with event nodes and tag nodes to provide text layout decisions.

6. The method for creating a video special effects cover based on artificial intelligence according to claim 1, characterized in that: The specific steps for collecting user interaction feedback are as follows: the generated candidate covers and condition vectors, along with the unique identifiers of video resources, are stored in the cover library; a sequence of candidate covers is maintained for each online education video; the online education video covers are rotated periodically; data points are embedded in the cover display interface to listen for user interaction actions with the current candidate cover and record the types of events executed by the user after playback; the data is received in the background and written into a message queue; it is processed using a stream processing framework; key indicators are calculated; and the results are sent back to the conditional generative adversarial network.

7. The method for creating a video special effects cover based on artificial intelligence according to claim 1, characterized in that: The specific steps of the enhanced cover learning iterative model are as follows: Combining the conditional vectors of candidate covers in the dynamic graph and information from user interaction feedback, a reward function is set, using user interaction feedback as the reward / penalty signal, and the reward value is calculated. The reward value formula is: award = 0.7 × R click +0.3×average dwell Where award represents the reward value, R click For click-through rate, average dwell The calculated reward value is used as the average dwell time and fed back to the conditional generative adversarial network to update the parameters in the generator. After subsequent display, feedback and update, the candidate cover is put into actual testing so that the cover generation model can continuously learn and optimize itself.

Citation Information

Patent Citations

  • Method, device, electronic device and storage medium for making video special effects cover

    CN117714774B

  • Video generation method, video display method and device

    CN115357755A

  • Personalized recommendation system and method for plasticized products in combination with user portraits

    CN119046537A