A Short Video Advertising Frame Interpolation Method Based on Generative Adversarial Networks
By combining the optical flow method and audio information to detect the occlusion area, using the ant colony algorithm to verify the motion hypothesis, and using the adversarial network model to compensate for details, the problem of difficult to obtain the motion information in the occlusion area in the short video advertising insert frame is solved, and the video quality and fluency are improved.
Patent Information
- Application Number
- CN202510662890.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-22
AI Technical Summary
In the prior art, when short video advertisements are inserted, the motion information in the occlusion area is difficult to accurately obtain, resulting in the loss of picture details, high computational complexity, and high hardware performance requirements.
The method based on the generation of adversarial network is adopted, combined with the optical flow method to detect the occlusion area, fused audio information to evaluate the intensity of the movement, used the ant colony algorithm to verify the motion hypothesis, and detailed compensation recovery is performed through the adversarial network model.
It improves the accuracy and reliability of motion analysis in the occlusion area, enhances the practicality and quality of video processing, and improves video fluency and audience experience.
Smart Images

Figure CN120186382B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular to a short video advertisement frame interpolation method based on a generative adversarial network. Background Art
[0002] Short video advertisement frame interpolation is a technology that inserts new frames between adjacent frames of an original short video advertisement to improve the smoothness and viewing experience of the advertisement video. In the prior art, a motion compensation algorithm is used to perform frame interpolation processing on the advertisement in the short video. Motion estimation is performed on two adjacent frames, and according to the motion trajectory and speed of the object, the frames that do not exist in the original video are compensated to achieve the purpose of increasing the video frame rate. The generated intermediate frames conform to the smooth motion relationship of the original video and can make the video smoother. However, in such a technology, when there is an object occlusion in the video, it is difficult to accurately obtain the motion information of the occluded part. For example, in a sports competition video, players intersperse and block each other. The MEMC algorithm may misestimate the motion trajectory of the occluded player, resulting in picture dislocation and ghosting after frame interpolation. Precise motion estimation and compensation require a large amount of calculation and have high requirements for hardware performance. During the motion compensation process, in order to reduce the computational complexity, the algorithm simplifies the motion model, resulting in the loss of some picture detail information, making the interpolated advertisement video less detailed than the original video and reducing the video quality. In response to this, we propose a short video advertisement frame interpolation method based on a generative adversarial network. Summary of the Invention
[0003] To solve the above technical problems, a short video advertisement frame interpolation method based on a generative adversarial network is provided. This technical solution solves the problems of difficult accurate acquisition of motion information and loss of picture details during occlusion.
[0004] To achieve the above object, the technical solution adopted by the present invention is: a short video advertisement frame interpolation method based on a generative adversarial network, and the advertisement frame interpolation steps are as follows:
[0005] S1. Obtain the short video advertisement to be inserted, and detect the image area with occlusion in the advertisement video based on the optical flow method;
[0006] S2. Establish different motion hypotheses for the occluded area, obtain the audio information to judge the current intensity of motion, fuse the audio information into the cost function calculation, calculate the rationality of the motion hypothesis, sort the calculation results, and select the motion hypothesis with the smallest rationality as the result;
[0007] S3. Obtain the surrounding images of the video occlusion, map the surrounding image pixel points into ant colonies, assign the initial moving direction and speed under the motion hypothesis to the pixel points, perform pheromone update, and statistically verify the motion hypothesis results;
[0008] S4. Obtain different short video data, where the video data includes various motion scene data. After preprocessing and integrating the data, construct an adversarial network model, use the obtained data for training and optimization, and compensate and restore the details of the verified motion hypotheses;
[0009] S5. Perform frame interpolation on the compensated and restored video advertisement.
[0010] Preferably, in step S1, the optical flow method detection steps are as follows:
[0011] Read the short video advertisement to be inserted, decompose the video into image frames, and preprocess the image frames;
[0012] Based on the optical flow algorithm, calculate the optical flow vector of each pixel point in the image, calculate frame by frame to obtain a data set;
[0013] Traverse each optical flow vector in the optical flow field to perform occlusion detection. Set a threshold. When the optical flow vector is greater than the threshold, it is judged as an abnormal pixel point, and the image is marked to mark the possible occlusion area.
[0014] Preferably, the optical flow algorithm calculates the optical flow vector of each pixel point in the image for two adjacent preprocessed images. The optical flow vector includes the velocity components of the pixel point in the horizontal and vertical directions. Calculate the optical flow vector frame by frame in sequence to obtain the data of the optical flow vector;
[0015] The specific occlusion detection steps are as follows:
[0016] Set a threshold for judging the abnormality of the optical flow vector. For each group of optical flow vectors, calculate the norm. Compare the norm with the threshold. When the calculated norm of the optical flow vector is greater than the threshold, judge the pixel point corresponding to the current optical flow vector as an abnormal pixel point, and perform marking processing on the abnormal pixel point. The marked area represents the possible occlusion area. By defining the marking rules, mark the abnormal pixel point as 1 according to the rules to indicate the possible occlusion situation; when the norm of the optical flow vector is less than the threshold, mark it as 0 to indicate that there is no occlusion situation, and obtain the marked image of the occlusion area in each frame of the image.
[0017] Preferably, in step S2, the steps for calculating the rationality of the motion hypothesis are as follows:
[0018] Propose the existing motion hypotheses for the judged occlusion image area;
[0019] Extract the audio data from the short video advertisement, perform format conversion processing, extract the features of the audio, calculate the features, and judge the intensity of the motion;
[0020] Define a cost function, comprehensively consider the optical flow information and the audio information for comprehensive calculation to obtain the cost function value;
[0021] The rationality is directly measured by the cost function value, that is, the smaller the cost function value, the higher the rationality of the motion hypothesis;
[0022] Sort the cost function values of all motion hypotheses, and select the motion hypothesis at the end of the sorting as the final result.
[0023] Preferably, the calculation of judging the intensity of motion is carried out by analyzing the audio features to obtain a quantization index. Multiply the root mean square of the volume in the audio features by the weight coefficient, and add the product of the frequency and the sum of the weight coefficients as the measurement value of judging the intensity of motion;
[0024] Define a cost function. For each motion hypothesis, calculate the cost related to the optical flow information and the cost related to the audio information respectively;
[0025] Calculate the cost related to the optical flow information. Based on the motion hypothesis, predict the optical flow situation, compare the predicted optical flow with the actual optical flow, calculate the difference to obtain the cost related to the optical flow information, and add up the gaps between the actual optical flow vector and the optical flow vector predicted based on the motion hypothesis to obtain the cost related to the optical flow;
[0026] Calculate the cost related to the audio information. Based on the motion hypothesis, estimate the intensity of motion index, compare the predicted index with the intensity of motion index calculated from the actual audio, calculate the difference between them to obtain the cost related to the audio information;
[0027] Set weights for the cost related to the optical flow information and the cost related to the audio information respectively. Multiply the cost related to the optical flow information and the cost related to the audio information by the corresponding weight values respectively, and add the products to obtain the cost function value of the current motion hypothesis. Calculate the cost function values of all motion hypotheses in turn.
[0028] Preferably, the motion hypothesis result verification step in step S3 is as follows:
[0029] Based on the detected occlusion area in the video, determine the surrounding unoccluded area image and extract it;
[0030] Map the pixel points in the surrounding area image into an ant colony, and record the position coordinate information of each pixel point in the image;
[0031] For each pixel point, based on the established different motion hypotheses, assign corresponding initial moving directions and speeds;
[0032] Define the initial value of pheromone, set the initial pheromone concentration for the position of each pixel point, and initialize the pheromone concentration of all positions to a fixed value;
[0033] Let the pixel points move in the image according to the given initial moving direction and speed. During the movement, update the pheromone. When the pixel point moves to a position point, increase the pheromone concentration at that position point;
[0034] Based on the passage of time, the pheromone concentration at all position points naturally decays according to the decay rate;
[0035] During the movement of the pixel points, record information, set evaluation indicators to quantify the degree of compliance, calculate the average error between the predicted pixel point position and the actual object position and the entropy value of the pheromone concentration distribution; Based on the calculated indicators, score each motion hypothesis. The higher the score, the more consistent the motion hypothesis is with the actual situation and the better the verification effect;
[0036] Based on the scoring results, judge whether the current motion hypothesis results are consistent.
[0037] Preferably, the average error calculation steps are as follows:
[0038] For each established motion hypothesis, determine the number of pixel points participating in the verification. For each pixel point, find the predicted coordinate position and the corresponding coordinate position of the actual object, calculate the distance difference between the predicted and actual coordinate positions. By taking the square root of the sum of the squares of the horizontal and vertical coordinate differences, add up the distance values of all pixel points and divide by the total number of pixel points participating in the verification to obtain the average error between the predicted positions of all pixel points and the actual object position;
[0039] The entropy value calculation steps are as follows:
[0040] Divide the image into a certain number of small regions, calculate the total sum of the pheromone concentrations in all small regions. For each small region, divide the pheromone concentration by the total sum of the pheromone concentrations to obtain the proportion of the pheromone concentration in the small region in the overall. Take the logarithm of the overall proportion and multiply it by its own proportion, then take the opposite number. Add up the values obtained by calculating all small regions to get the entropy value of the pheromone concentration distribution.
[0041] Preferably, the specific steps for scoring each motion hypothesis are as follows:
[0042] Set a weight for the average error and the entropy value respectively. For each motion hypothesis, divide 1 by the value of the average error, multiply by the weight corresponding to the average error, add the product of the entropy value and its corresponding weight. The obtained value is the score of the current motion hypothesis. Obtain the verification score threshold, compare the calculated value with the threshold. When it is greater than the threshold, the motion hypothesis result is verified accurately.
[0043] Preferably, in step S4, the adversarial network model consists of a generator and a discriminator. The generator is used to generate the details of the motion hypothesis, and the discriminator is used to judge the authenticity.
[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0045] By using the optical flow method to detect the occluded area, the present invention lays a foundation for subsequent analysis, establishes a motion hypothesis and fuses audio information to evaluate the rationality, comprehensively considers visual and auditory information, can analyze the motion state of the occluded area more comprehensively and accurately, avoids the limitations of a single information source, dynamically verifies the hypothesis from the pixel level by simulating the behavior of ant colonies, makes the verification process more in line with the actual motion law, enhances the reliability of the result, and with the help of the adversarial network model, compensates and restores the details of the motion hypothesis, improves the accuracy of motion analysis, can meet the requirements of various application scenarios such as precise advertising placement and video content understanding, and enhances the practicality and value of video processing. Brief Description of the Drawings
[0046] Figure 1 It is a flowchart of the advertisement frame interpolation steps of the present invention. Detailed Embodiments
[0047] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and other obvious variations can be thought of by those skilled in the art.
[0048] Referring to Figure 1 as shown, a short video advertisement frame interpolation method based on a generative adversarial network, and the advertisement frame interpolation steps are as follows:
[0049] S1. Obtain the short video advertisement to be inserted, and detect the image area with occlusion in the advertisement video based on the optical flow method;
[0050] S2. Establish different motion hypotheses for the occluded area, obtain audio information to judge the current intensity of motion, fuse the audio information into the cost function calculation, calculate the rationality of the motion hypothesis, sort the calculation results, and select the motion hypothesis with the smallest rationality as the result;
[0051] S3. Obtain the images around the video occlusion, map the surrounding image pixel points into an ant colony, assign the initial moving direction and speed under the motion hypothesis to the pixel points, perform pheromone update, and statistically verify the motion hypothesis result;
[0052] S4. Obtain different short video data, the video data includes various motion scene data, preprocess and integrate the data after preprocessing, construct an adversarial network model, use the obtained data for training and optimization, and compensate and restore the details of the verified motion hypothesis;
[0053] S5. Perform frame interpolation processing on the compensated and restored video advertisement.
[0054] This application detects occluded regions based on the optical flow method, which can accurately identify parts in the advertising video that may affect the visual effect and information transmission, providing a precise target area for subsequent processing. By combining audio information to judge the intensity of motion and incorporating it into the cost function calculation, it makes full use of the multi-modal information of the video, making the evaluation of the rationality of motion assumptions more comprehensive and accurate. Selecting the hypothesis with the minimum rationality as the result helps to uncover more unique and potentially valuable motion patterns; using the idea of the ant colony algorithm to verify motion assumptions, mapping pixel points to ant colonies and updating pheromones, this bionic method can more meticulously simulate the motion behavior of pixel points, verifying the correctness of motion assumptions from a microscopic level and improving the reliability of the results. Building an adversarial network model and training and optimizing it with a large amount of short video data to compensate and restore the details of motion assumptions can effectively repair and improve the information loss or incompleteness problems caused by occlusion in the video, enhancing the quality and readability of the video; frame interpolation processing can increase the frame rate of the video without changing the original content of the video, making the video playback smoother and enhancing the viewing experience of the audience. Especially for videos that have undergone complex processing before, frame interpolation processing helps to eliminate possible visual stuttering, making the advertising video more natural and attractive.
[0055] The optical flow method detection steps in step S1 are as follows:
[0056] Read the short video advertisement to be inserted, decompose the video into image frames, and preprocess the image frames;
[0057] Based on the optical flow algorithm, calculate the optical flow vector of each pixel point in the image, calculate frame by frame to obtain a data set;
[0058] Traverse each optical flow vector in the optical flow field for occlusion detection. Set a threshold. When the optical flow vector is greater than the threshold, it is judged as an abnormal pixel point, and the image is marked as a possible occluded area.
[0059] This application decomposes the video into image frames and preprocesses them, providing a clear and standardized data basis for subsequent precise analysis, effectively removing noise interference factors, and improving the accuracy of analysis; calculating the optical flow vector of each pixel point based on the optical flow algorithm and processing frame by frame can meticulously capture the motion information of objects in the image, accurate to the motion state of each pixel point, thus comprehensively and precisely describing the motion situation in the video; by traversing the optical flow vectors for occlusion detection and using the set threshold to judge abnormal pixel points, it can quickly and accurately identify possible occluded areas. This method has high efficiency and accuracy, can quickly locate problem areas in a large amount of image data, provide a precise target for subsequent processing, helps to improve the efficiency and quality of the entire video processing process, and better guarantees the clear display of advertising content and the viewing experience of the audience.
[0060] The optical flow algorithm calculates the optical flow vector of each pixel in two adjacent pre-processed frames. The optical flow vector includes the velocity components of the pixel in the horizontal and vertical directions. The optical flow vector is calculated frame by frame to obtain the optical flow vector data.
[0061] The occlusion detection steps are as follows:
[0062] A threshold is set for judging the abnormality of the optical flow vector. For each group of optical flow vectors, the modulus is calculated and compared with the threshold. When the modulus of the calculated optical flow vector is greater than the threshold, the pixel corresponding to the current optical flow vector is judged to be an abnormal pixel, and the abnormal pixel is marked. The marked area indicates the area where occlusion may exist. By defining the marking rules, the abnormal pixel is marked as 1 according to the rules to indicate possible occlusion; when the modulus of the optical flow vector is less than the threshold, it is marked as 0 to indicate that there is no occlusion, and a marked image of the occlusion area in each frame of the image is obtained.
[0063] The specific calculation steps are as follows:
[0064] Optical flow vector calculation, the t-th frame image after preprocessing is I t (x,y), the t+1th frame image is I t+1 (x,y), where x represents the horizontal coordinate of the image and y represents the vertical coordinate of the image;
[0065] For the pixel point with coordinates (x, y) in the image, the optical flow vector between the tth frame and the t+1th frame is calculated based on the optical flow algorithm. The calculation formula is:
[0066]
[0067] where u t (x,y) is the velocity component of the pixel in the horizontal direction, v t (x,y) is the velocity component of the pixel in the vertical direction;
[0068] Calculate each two adjacent frames of images in turn to obtain the data set of optical flow vectors Where N is the total number of image frames after video decomposition;
[0069] Occlusion detection is achieved by setting a threshold T for judging the abnormality of optical flow vectors. Calculate its modulus The calculation formula is:
[0070] Define a labeling function L t (x, y) is used to mark whether a pixel is an abnormal pixel. The rules are as follows:
[0071]
[0072] Through the above marking function, the marking image of the occluded area in each frame image is obtained, where L t The area where (x,y)=1 indicates the area where there may be occlusion, L t The area where (x, y)=0 indicates an area where there is no occlusion.
[0073] The steps for calculating the rationality of the motion hypothesis in step S2 are:
[0074] Propose a motion hypothesis for the judged occluded image area;
[0075] Extract audio data from short video ads, convert the format, extract audio features, calculate the features, and determine the intensity of the movement;
[0076] Define the cost function, comprehensively consider the optical flow information and audio information for comprehensive calculation, and obtain the cost function value;
[0077] Directly use the cost function value to measure rationality, that is, the smaller the cost function value, the higher the rationality of the motion hypothesis;
[0078] Sort the cost function values of all motion hypotheses and select the motion hypothesis at the end of the sort as the final result.
[0079] This application not only considers the optical flow information in the video image, but also incorporates audio data and its characteristics, analyzes the motion of the occluded area from multiple dimensions, and makes the judgment of the rationality of the motion hypothesis more comprehensive and accurate; by calculating the audio features to judge the intensity of the motion, it can more carefully reflect the motion conditions in the video. Intense action scenes may be accompanied by large audio amplitudes and complex frequency changes. Combined with the optical flow information, this can more accurately depict the motion state of the occluded area; define the cost function and calculate the optical flow and audio information comprehensively, convert the rationality of the motion hypothesis into a specific numerical value, facilitate the quantitative evaluation and comparison of different motion hypotheses, and intuitively reflect the rationality of each hypothesis; sort all motion hypotheses according to the cost function value, and select the motion hypothesis at the end of the sort as the final result, which helps to screen out the hypothesis that best matches the actual motion of the video, improve the accuracy of the motion analysis of the occluded area, and provide a more reliable basis for subsequent video processing.
[0080] The intensity of exercise is determined by analyzing audio features to obtain a quantitative index. The root mean square of the volume in the audio features is multiplied by the weight coefficient, and the sum of the frequency and the weight coefficient is added to form a measure of the intensity of exercise.
[0081] Define the cost function. For each motion hypothesis, calculate the cost related to the optical flow information and the cost related to the audio information separately;
[0082] Calculate the cost related to the optical flow information. Based on the motion hypothesis, predict the optical flow situation, compare the predicted optical flow with the actual optical flow, calculate the difference to obtain the cost related to the optical flow information, and add up the gaps between the actual optical flow vectors and the optical flow vectors predicted based on the motion hypothesis to get the cost related to the optical flow;
[0083] Calculate the cost related to the audio information. Based on the motion hypothesis, estimate the motion intensity index, compare the predicted index with the motion intensity index calculated from the actual audio, calculate the difference between them to obtain the cost related to the audio information;
[0084] Set weights for the cost related to the optical flow information and the cost related to the audio information respectively. Multiply the cost related to the optical flow information and the cost related to the audio information by their corresponding weight values respectively, and add up the products to obtain the value of the cost function for the current motion hypothesis. Calculate the values of the cost functions for all motion hypotheses in turn.
[0085] The specific calculation process is as follows:
[0086] Judge the calculation of the motion intensity metric. Let the root mean square of the volume in the audio feature be RMS, and its weight coefficient be w rms ; the frequency-related value be f, and its weight coefficient be w f ;
[0087] The calculation formula for the motion intensity metric I is: I = w rms × RMS + w f × f
[0088] Define the relevant calculations of the cost function. Suppose there are n motion hypotheses in total. For the i-th motion hypothesis (i = 1, 2,..., n);
[0089] The cost related to the optical flow information Calculate: Let the set of actual optical flow vectors be The set of optical flow vectors predicted based on the i-th motion hypothesis be The relevant cost calculation formula is:
[0090]
[0091] where ‖·‖2 represents the Euclidean norm, and j represents the element index in the optical flow vector set;
[0092] The cost related to the audio information Calculate. Let the motion intensity index predicted based on the i-th motion hypothesis be The motion intensity index calculated from the actual audio be I;
[0093]
[0094] Cost function value J i Calculation: Let the weight of the cost related to the optical flow information be w optical , and the weight of the cost related to the audio information be w audio ;
[0095] And w optical +w audio =1
[0096] 0 ≤ w optical , w audio ≤ 1
[0097]
[0098] Perform the above calculations for i = 1, 2,..., n in sequence, and the cost function values J 1 , J 2 , …, J n of all n motion hypotheses can be obtained.
[0099] The motion hypothesis result verification step in step S3 is as follows:
[0100] Based on the detected occluded regions in the video, determine the surrounding non-occluded region images and extract them;
[0101] Map the pixel points in the surrounding region images to ant colonies, and record the position coordinate information of each pixel point in the image;
[0102] For each pixel point, based on the established different motion hypotheses, assign corresponding initial moving directions and speeds;
[0103] Define the initial value of pheromone, set the initial pheromone concentration for the position of each pixel point, and initialize the pheromone concentrations of all positions to a fixed value;
[0104] Let the pixel points move in the image according to the assigned initial moving directions and speeds. During the movement, update the pheromone. When the pixel point moves to a position point, increase the pheromone concentration at that position point;
[0105] Based on the passage of time, the pheromone concentrations of all position points decay naturally according to the decay rate;
[0106] During the movement of the pixel points, record information, set evaluation indicators to quantify the degree of conformity, calculate the average error between the predicted pixel point positions and the actual object positions and the entropy value of the pheromone concentration distribution; Based on the calculated indicators, score each motion hypothesis. The higher the score, the more consistent the motion hypothesis is with the actual situation and the better the verification effect;
[0107] Based on the scoring results, determine whether the current motion hypothesis results are consistent.
[0108] In this application, by determining and extracting the image of the unoccluded area around the occluded area in the video, the surrounding information can be used to assist in verifying the motion hypothesis. The pixel points in the surrounding area image are mapped into an ant colony, and each pixel point is given an initial moving direction and speed. This simulation method based on the ant colony algorithm can effectively simulate the motion behavior of pixel points under different motion hypotheses. An evaluation index is set to quantify the degree of compliance, and each motion hypothesis is scored by calculating the average error between the predicted pixel point position and the actual object position and the entropy value of the pheromone concentration distribution.
[0109] The steps for calculating the average error are as follows:
[0110] For each established motion hypothesis, determine the number of pixel points participating in the verification. For each pixel point, find the predicted coordinate position and the corresponding coordinate position of the actual object, calculate the distance difference between the predicted and actual coordinate positions, take the square root of the sum of the squares of the horizontal and vertical coordinate differences, add up the distance values of all pixel points, and divide by the total number of pixel points participating in the verification to obtain the average error between the predicted positions of all pixel points and the actual object position.
[0111] The steps for calculating the entropy value are as follows:
[0112] Divide the image into a certain number of small regions, calculate the total sum of the pheromone concentrations in all small regions. For each small region, divide the pheromone concentration by the total sum of the pheromone concentrations to obtain the proportion of the pheromone concentration in the small region in the overall. Take the logarithm of the overall proportion and multiply it by its own proportion, then take the opposite number. Add up the values obtained by calculating all small regions to get the entropy value of the pheromone concentration distribution.
[0113] The specific steps for scoring each motion hypothesis are as follows:
[0114] Set a weight for the average error and the entropy value respectively. For each motion hypothesis, divide 1 by the value of the average error, multiply it by the weight corresponding to the average error, add the product of the entropy value and its corresponding weight. The obtained value is the score of the current motion hypothesis. Obtain the verification score threshold, compare the calculated value with the threshold. When it is greater than the threshold, the motion hypothesis result is verified accurately.
[0115] By determining and extracting the images of the unoccluded areas around the occluded area, this application can obtain additional valid information. These peripheral information complements the information of the occluded area itself, so that when verifying the motion hypothesis, it no longer solely relies on the limited data of the occluded area. Mapping the pixel points in the surrounding area into an ant colony and endowing them with an initial moving direction and speed, and simulating the motion behavior of pixel points by virtue of the characteristics of the ant colony algorithm. Updating and decaying pheromones during the movement of pixel points can make the verification process adapt to different scenarios and time changes.
[0116] In step S4, the adversarial network model consists of a generator and a discriminator. The generator is used to generate the details of the motion hypothesis, and the discriminator is used to judge the authenticity.
[0117] The generator of this application can generate the details of the motion hypothesis. By learning the motion patterns and characteristics in a large number of short video data, it supplements the missing or blurred information for the motion hypothesis of the occluded area. The occlusion detection module can accurately detect the occluded image areas in the advertising video, providing a clear target for subsequent processing. The hypothesis verification module verifies the results of the established motion hypothesis. By setting reasonable verification methods and indicators, it ensures the reliability of the motion hypothesis.
[0118] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A short video advertisement frame interpolation method based on a generative adversarial network, characterized in that The steps of inserting advertisement frames are as follows: S1. Obtain the short video advertisement to be inserted, and detect the occluded image areas in the advertisement video based on the optical flow method; S2. Establish different motion hypotheses for the occluded areas, obtain the audio information to judge the current intensity of motion, integrate the audio information into the calculation of the cost function, calculate the rationality of the motion hypotheses, sort the calculation results, and select the motion hypothesis with the lowest rationality as the result; S3. Obtain the images around the video occlusion, map the pixel points of the surrounding images into ant colonies, assign the initial moving directions and speeds under the motion hypothesis to the pixel points, perform pheromone update, and statistically verify the motion hypothesis results; S4. Obtain different short video data, where the video data includes various motion scene data, preprocess and integrate the data after preprocessing, construct an adversarial network model, use the obtained data for training and optimization, and compensate and restore the details of the verified motion hypothesis; S5. Perform frame insertion processing on the compensated and restored video advertisement.
2. The short video advertisement frame interpolation method based on a generative adversarial network according to claim 1, wherein The optical flow method detection steps in step S1 are as follows: Read the short video advertisement to be inserted, decompose the video into image frames, and preprocess the image frames; Based on the optical flow algorithm, calculate the optical flow vectors of each pixel point in the image, calculate frame by frame to obtain a data set; Traverse each optical flow vector in the optical flow field, perform occlusion detection, set a threshold, and when the optical flow vector is greater than the threshold, judge it as an abnormal pixel point, mark the image, and mark it as a possible occlusion area.
3. The short video advertisement frame interpolation method based on a generative adversarial network according to claim 2, wherein, The optical flow algorithm calculates the optical flow vectors of each pixel point in the image for two adjacent preprocessed images. The optical flow vectors include the velocity components of the pixel point in the horizontal and vertical directions. Calculate the optical flow vectors frame by frame in sequence to obtain the data of the optical flow vectors; The specific steps of occlusion detection are as follows: Set a threshold for judging the abnormality of the optical flow vector. For each group of optical flow vectors, calculate the norm, compare the norm with the threshold. When the calculated norm of the optical flow vector is greater than the threshold, judge the pixel point corresponding to the current optical flow vector as an abnormal pixel point, and perform marking processing on the abnormal pixel point. The marked area indicates a possible occlusion area. By defining the marking rules, mark the abnormal pixel point as 1 according to the rules to indicate a possible occlusion situation; when the norm of the optical flow vector is less than the threshold, mark it as 0 to indicate that there is no occlusion situation, and obtain the marked image of the occlusion area in each frame of the image.
4. A short video advertisement frame interpolation method based on a generative adversarial network according to claim 1, characterized in that, The steps of calculating the rationality of the motion hypothesis in step S2 are as follows: Propose the existing motion hypotheses for the judged occluded image areas; Extract the audio data from the short video advertisement, perform format conversion processing, extract the features of the audio, calculate the features, and judge the intensity of motion; Define a cost function, comprehensively consider the optical flow information and audio information for comprehensive calculation to obtain the cost function value; Directly use the cost function value to measure the rationality, that is, the smaller the cost function value, the higher the rationality of the motion hypothesis; Sort the cost function values of all motion hypotheses, and select the motion hypothesis at the end of the sorting as the final result.
5. The short video advertisement frame interpolation method based on a generative adversarial network according to claim 4, wherein The calculation of the intensity of exercise is judged by analyzing audio features to obtain a quantization index. Multiply the root mean square of the volume in the audio features by the weight coefficient, and add the product of the frequency and the sum of the weight coefficients as the measurement value for judging the intensity of exercise. Define the cost function. For each motion hypothesis, calculate the cost related to the optical flow information and the cost related to the audio information respectively. Calculate the cost related to the optical flow information. Based on the motion hypothesis, predict the optical flow situation, compare the predicted optical flow with the actual optical flow, calculate the difference to obtain the cost related to the optical flow information, and add up the gaps between the actual optical flow vectors and the optical flow vectors predicted based on the motion hypothesis to get the cost related to the optical flow. Calculate the cost related to the audio information. Based on the motion hypothesis, predict the intensity index of exercise, compare the predicted index with the intensity index of exercise calculated from the actual audio, calculate the difference between them to obtain the cost related to the audio information. Set weights for the cost related to the optical flow information and the cost related to the audio information respectively. Multiply the cost related to the optical flow information and the cost related to the audio information by the corresponding weight values respectively, and add the products to obtain the value of the cost function for the current motion hypothesis. Calculate the values of the cost functions for all motion hypotheses in turn.
6. A method for short video advertisement frame interpolation based on a generative adversarial network according to claim 1, characterized in that, The verification step of the motion hypothesis result in step S3 is as follows: Based on the detected occlusion area in the video, determine the image of the surrounding unoccluded area and extract it. Map the pixel points in the surrounding area image into an ant colony, and record the position coordinate information of each pixel point in the image. For each pixel point, based on the established different motion hypotheses, assign corresponding initial moving directions and speeds. Define the initial value of pheromone. Set the initial pheromone concentration for the position of each pixel point, and initialize the pheromone concentrations of all positions to a fixed value. Let the pixel points move in the image according to the assigned initial moving directions and speeds. During the movement, update the pheromone. When the pixel point moves to a position point, increase the pheromone concentration at that position point. Based on the passage of time, the pheromone concentrations of all position points decay naturally according to the decay rate. During the movement of the pixel points, record information, set the evaluation index to quantify the degree of conformity, calculate the average error between the predicted pixel point position and the actual object position and the entropy value of the pheromone concentration distribution; based on the calculated index, score each motion hypothesis. The higher the score, the more consistent the motion hypothesis is with the actual situation and the better the verification effect. Based on the scoring results, judge whether the current motion hypothesis results are consistent.
7. A short video advertisement frame interpolation method based on a generative adversarial network according to claim 6, characterized in that The calculation steps of the average error are as follows: For each established motion hypothesis, determine the number of pixel points participating in the verification. Find the predicted coordinate position and the corresponding coordinate position of the actual object for each pixel point, calculate the distance difference between the predicted and actual coordinate positions, take the square root of the sum of the squares of the horizontal and vertical coordinate differences, add up the distance values of all pixel points, and divide by the total number of pixel points participating in the verification to obtain the average error between the predicted positions of all pixel points and the actual object position. The calculation steps of the entropy value are as follows: Divide the image into a number of small regions, calculate the sum of the pheromone concentrations of all small regions. For each small region, divide the pheromone concentration by the sum of the pheromone concentrations to obtain the proportion of the pheromone concentration of the small region in the overall. Take the logarithm of the overall proportion and multiply it by its own proportion, then take the opposite. Add up the values obtained by calculation for all small regions, which is the entropy value of the pheromone concentration distribution.
8. A short video advertisement frame interpolation method based on a generative adversarial network according to claim 6, characterized in that The specific steps for scoring each motion hypothesis are as follows: Set a weight for the average error and the entropy value respectively. For each motion hypothesis, divide 1 by the value of the average error, multiply it by the weight corresponding to the average error, and add the product of the entropy value and its corresponding weight. The obtained value is the score of the current motion hypothesis. Obtain the verification score threshold, and compare the calculated value with the threshold. When it is greater than the threshold, the motion hypothesis result is verified accurately.
9. A short video advertisement frame interpolation method based on a generative adversarial network according to claim 1, characterized in that In step S4, the adversarial network model consists of a generator and a discriminator. The generator is used to generate the details of the motion hypothesis, and the discriminator is used to judge the authenticity.
Citation Information
Patent Citations
Multi-reference frame rapid movement estimation method based on effective coverage
CN1633184A
Occlusion processing for frame rate conversion using deep learning
WO2022075688A1