Video management method based on artificial intelligence

By combining facial recognition and expression analysis with a cross-modal matching model, the problem of insufficient ad placement evaluation in existing technologies is solved, achieving accurate matching between ads and videos, improving user experience, and optimizing ad delivery effectiveness.

CN120915978AActive Publication Date: 2025-11-07YINMEI (SHANDONG) CULTURE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511109352.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-07
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing video management methods fail to effectively assess the impact of ad placement on ad performance, leading to an underestimation of the potential value of ads and a lack of in-depth understanding of user viewing behavior, which affects user experience and ad delivery effectiveness.

Method used

By locating the center of a user's pupil through facial recognition, calculating the user's interest value by combining facial expressions, extracting image text and scene features, evaluating the matching degree between advertisements and videos using a cross-modal matching model, and optimizing advertising strategies based on user interaction data.

Benefits of technology

It enables intelligent determination of personalized ad insertion positions, improves the contextual consistency between ads and videos, enhances user acceptance, and optimizes ad performance by adjusting ad strategies through self-learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915978A_ABST
    Figure CN120915978A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of video management, and particularly relates to a video management method based on artificial intelligence. The method comprises the following steps: firstly, extracting an interest value of a user in a video watching process through facial recognition and expression analysis, and judging an attention concentration time frame; and then text features and scene features of the time frame and the advertisement video key frame are extracted, semantic alignment and matching scoring are performed on the video and the advertisement by using a cross-modal matching model, and the advertisement with high matching degree is screened and inserted into the video. The system further collects interaction behavior data generated after the user watches the advertisement, calculates the advertisement adhesion score, dynamically adjusts the advertisement putting strategy, and achieves the continuous optimization of the advertisement effect. The method improves the correlation between the advertisement and the video content and the user experience, and has the advantages of high precision, high adaptability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of video management, and particularly relates to a video management method based on artificial intelligence. BACKGROUND

[0002] With the acceleration of the commercialization process of video platforms, advertising has become the core issue of balancing platform revenue and user experience. Not only does it require the ad to reach the user, but it also needs to accurately measure the quality of the user's deep interaction with the ad (such as continuous attention after clicking and interaction continuity); and the increasing sensitivity of users to ad interference forces platforms to build more sophisticated ad effectiveness evaluation systems to optimize ad placement strategies.

[0003] In the prior art, the evaluation of the effect of the advertisement in the video is often only processed by relying on the click rate and the stay time, without considering that the position of the advertisement has a huge influence on the evaluation effect of the advertisement, which is easy to underestimate the potential value of the advertisement. SUMMARY

[0004] The present application proposes a video management method based on artificial intelligence to solve the technical problems in the above background art.

[0005] In order to achieve the above purpose, the technical solution adopted by the present application is as follows:

[0006] First, the pure video data is obtained, the user's pupil center is located by facial recognition, the user's interest value is calculated combined with facial expressions, and the user's attention concentration time frame is determined;

[0007] The image of the user's attention concentration time frame in the video data is selected, the image is feature-extracted, and the image text features and image scene features are extracted; the content features of the advertisement video to be inserted are extracted, including the advertisement text features and the advertisement scene features;

[0008] Through cross-modal matching, the image of the user's attention concentration time frame and the features of the advertisement to be inserted are aligned, and the matching degree score of the video and the advertisement is output by the cross-modal matching model; a first determination threshold is set, and if the matching degree is greater than the first determination threshold, the advertisement is included in the candidate advertisement pool;

[0009] The advertisements in the candidate advertisement pool are sorted in descending order of matching degree score, the maximum number of insertions is determined combined with the total video duration, and the top N advertisements in the sorting are selected as the advertisements to be inserted;

[0010] The advertisement to be inserted is inserted into the position adjacent to the attention concentration time frame of the video, and the video containing the advertisement is generated and pushed to the user;

[0011] Obtaining interaction data of a user when watching a video containing an advertisement, calculating a stickiness score of the user to the advertisement based on the interaction data, if the stickiness score is greater than a preset second determination threshold, determining that the advertisement is effectively matched with the video, retaining the advertisement insertion strategy, if the stickiness score is less than or equal to the preset second determination threshold, selecting a next advertisement from a candidate advertisement pool to re-perform the insertion and determination process;

[0012] Before extracting the content features of the advertisement video to be inserted, the method further comprises extracting key feature frames from the advertisement video, and the key feature frames are extracted from the beginning, middle and end of the advertisement video respectively.

[0013] As preferred, the specific implementation of using face recognition to locate the pupil center of the user and combining facial expression to calculate the interest value of the user comprises:

[0014] First, the eye ROI is extracted, the pupil edge points are extracted, and the least square method is used to fit an ellipse for the pupil edge points, and the center of the ellipse is the pupil center;

[0015] An initial feature map of the face ROI is extracted by using a lightweight CNN Wherein is the size of the feature map, is the number of channels, the distance weight between the feature map position and the expression key node is calculated to generate a spatial weight map , Wherein, is the feature map position corresponding to the Euclidean distance of the nearest expression node in the original image, is the diagonal length of the face ROI;

[0016] The channel dimension of the feature map is globally averaged pooled, the importance of the channel is learned through a fully connected layer, and a channel weight map is generated , Wherein, is the c-th channel feature map, is the global average pooling, is the fully connected layer, is the Sigmoid function;

[0017] The weight map and the basic feature map are fused, and the expression feature vector is obtained through global pooling , Wherein is element-wise multiplication, is outer product expansion;

[0018] The moving standard deviation of the pupil center in the sliding window is calculated, and the stability is obtained after normalization Wherein is the number of frames in the window, the mean value of the pupil center in the window, the horizontal and vertical coordinates of the pupil center in the video frame image coordinate system at frame t, and then the pupil stability coefficient is calculated where D is the diagonal length of the eye ROI region;

[0019] the expression feature vector is input into the pre-trained SVM to score and output the expression positivity ;

[0020] The interest value is obtained by weighted fusion of the pupil stability coefficient and the expression positivity.

[0021] As a preferred embodiment, after obtaining the interest value, the specific implementation of determining the user attention concentration time frame is as follows:

[0022] The average interest value is calculated using a sliding window wherein is the average interest value of the t-th frame, is the size of the sliding window, is the original interest value of the (t-w)th frame;

[0023] When is greater than a set threshold value, it is marked as a candidate frame, and the duration of consecutive candidate frames is counted If is greater than a set duration threshold value, all frames in the interval are determined as user attention concentration time frames.

[0024] As a preferred embodiment, the image of the user attention concentration time frame in the video data is selected, and image features are extracted, including image text features and image scene features. The content features of the advertisement video to be inserted include advertisement text features and advertisement scene features. The extraction of the features of the two is implemented as text feature extraction and scene feature extraction:

[0025] The text feature extraction is as follows:

[0026] A text area detector is used to detect the text area in the image of the attention concentration time frame and the key frame in the advertisement video, and the text area in the image is located.

[0027] A neural network is used to recognize the detected text area, and a text sequence is output.

[0028] A pre-trained BERT model is used to encode the word vectors of the recognized text sequence, and a basic text feature vector is obtained.

[0029] A GPT small sample learning model is used to expand the semantic of the basic text feature vector, and a related expanded text vector is generated.

[0030] Fusing the basic text feature vector and the extended text feature vector to obtain a final text feature;

[0031] The scene feature extraction is:

[0032] The image scene visual feature vectors of the two are extracted by using a Transformer model;

[0033] The image features are mapped to a semantic space based on a pre-trained CLIP model, matched with preset scene category texts, and scene semantic feature vectors are generated;

[0034] The image scene visual feature vectors and the scene semantic feature vectors are spliced and fused to obtain scene features.

[0035] As a preferred, the extraction of the text feature and the scene feature is that the feature extraction operation is respectively performed on the user attention concentrated time frames extracted in all videos and the three key feature frames extracted from the front, middle and rear of the advertisement video, and then a weighting operation is performed to obtain the final text feature and scene feature of the two.

[0036] As a preferred, the implementation of aligning the image of the user attention concentrated time frame and the features of the advertisement to be inserted by cross-modal matching is that:

[0037] The image text feature vector and the image scene feature vector of the user attention concentrated time frame are mapped to the same high-dimensional semantic space through a cross-modal mapping network together with the advertisement text feature vector and the advertisement scene feature vector of the advertisement to be inserted;

[0038] The cosine similarity of the image text mapping vector and the advertisement text mapping vector is calculated to obtain a text matching degree ;

[0039] The cosine similarity of the image scene mapping vector and the advertisement scene mapping vector is calculated to obtain a scene matching degree

[0040] The matching degree score is obtained by weighted sum of the text matching degree and the scene matching degree , wherein, is a weight coefficient.

[0041] As a preferred, the interaction data of the video containing the advertisement includes an advertisement click rate and an advertisement stay duration proportion.

[0042] As a preferred, the calculation method of the adhesion score of the user to the advertisement based on the interaction data is:

[0043] First, the synergy index is calculated , wherein For the advertisement click rate, For the advertisement stay time length proportion;

[0044] Calculate the adhesion score Wherein The adhesion score.

[0045] Compared with the prior art, the advantages and positive effects of the present application are:

[0046] 1. By combining facial recognition with pupil tracking, combining expression recognition technology, real-time calculation of user interest value, accurate positioning of user attention concentration time frame, realization of intelligent judgment of personalized advertisement insertion position.

[0047] 2. At the same time, the text features and scene features in the image are extracted, and the video content and the advertisement content are deep semantic modeling and fusion through BERT, GPT, Transformer, CLIP and other pre-training models, to improve the consistency of the advertisement and the video context.

[0048] 3. Through the cross-modal mapping network, the video frame and the advertisement content are mapped to a unified high-dimensional semantic space, and the semantic matching degree is used to measure the fit degree between the two, so as to filter out the most relevant advertisement for insertion, effectively improving the user acceptance.

[0049] 4. Based on the real interaction data of users (such as click rate, stay time, etc.), the adhesion score is calculated, the advertisement effect is dynamically evaluated, the feedback mechanism is formed, the advertisement strategy is automatically adjusted, and the self-learning and continuous optimization of the advertisement effect are realized. BRIEF DESCRIPTION OF DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0051] Figure 1 It is a structural flow diagram of a video management method based on artificial intelligence. DETAILED DESCRIPTION

[0052] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the following will further illustrate the present application with the help of the drawings and embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0053] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure aspects of the present application.

[0054] With the rapid growth of video content in social media, streaming platforms, and online education, among other fields, advertisement insertion has become one of the main ways for video platforms to monetize. However, existing methods of advertisement insertion in video management often use fixed time nodes or simple content keyword matching for pushing, lacking a deep understanding of user viewing behavior, which can easily cause disconnection between advertisements and video content, thereby affecting user experience and advertisement placement effectiveness. To this end, a video management method based on artificial intelligence is proposed, which matches advertisements and videos by personnel, thereby improving user experience and advertisement placement effectiveness. The specific implementation process is as shown in Figure 1

[0055] To realize real-time recognition of user interest, first, pure video data is obtained, the user's pupil center is located using facial recognition, the user's interest value is calculated in combination with facial expressions, and the user's attention concentration time frame is determined. Specifically, the specific implementation of the use of facial recognition to locate the user's pupil center and the calculation of the user's interest value in combination with facial expressions includes: first, the eye ROI is extracted, the pupil edge points are extracted, and the pupil edge points are fitted with an ellipse using the least squares method. The center of the ellipse is the pupil center. This algorithm minimizes the error of all edge points to the boundary of the fitted ellipse to obtain the optimal ellipse parameters, including the center coordinates, the length of the major and minor axes, and the rotation angle, etc. Finally, the geometric center of the fitted ellipse is regarded as the center position of the user's pupil in the current frame, which is used as the core reference data for subsequent calculation of pupil stability and user interest value. This method has the advantages of high calculation accuracy and strong robustness, and can effectively cope with slight occlusion, changes in light, and other actual scene disturbances. A lightweight CNN is used to extract the initial feature map of the facial ROI wherein is the feature map size, is the number of channels, the distance weight between the feature map position and the expression key node is calculated to generate a spatial weight map , wherein, is the feature map position corresponding to the Euclidean distance of the nearest expression node in the original image, is the diagonal length of the facial ROI; the channel dimension of the feature map is globally averaged pooled, the importance of the channel is learned through a fully connected layer, and a channel weight map is generated , wherein, is the c-th channel feature map, is the global average pooling, is the fully connected layer,​ is a sigmoid function; the weight map is fused with the base feature map, and an expression feature vector is obtained through global pooling , wherein is an element-wise multiplication, is an outer product expansion; the moving standard deviation of the pupil center in the sliding window is calculated, and the stability is obtained after normalization wherein is the number of frames in the window, is the mean value of the pupil center in the window, is the horizontal and vertical coordinates of the pupil center in the video frame image coordinate system at t frame, and the pupil stability coefficient is calculated wherein D is the diagonal length of the eye ROI region; the expression feature vector is scored by a pre-trained SVM , and the expression positivity is output ; the pupil stability coefficient and the expression positivity are weighted and fused to obtain the interest value.

[0056] After obtaining the interest value, the user attention concentration time frame is determined, and the average interest value is calculated using a sliding window wherein is the average interest value of the t frame, is the size of the sliding window, is the original interest value of the (t-w) frame; when is greater than a set threshold, it is marked as a candidate frame, and the duration of consecutive candidate frames is counted If is greater than a set duration threshold, all frames in the interval are determined as user attention concentration time frames.

[0057] In order to ensure the semantic consistency and scene coherence of the advertisement and the video content, the application further extracts multi-modal features from the video key frame and the advertisement video content. The images of the user attention concentration time frames in the video data are selected, and the image features are extracted, including image text features and image scene features; the content features of the inserted advertisement video are extracted, including advertisement text features and advertisement scene features. Specifically, the inserted advertisement video is processed first, and before extracting the content features of the inserted advertisement video, key feature frames need to be extracted from the advertisement video, and one frame is extracted as a key feature frame at the beginning, middle and end of the advertisement video. The purpose of this is to ensure that the key frames are processed.

[0058] Then, text feature extraction and scene feature extraction are performed on each user attention concentration time frame and the three key feature frames extracted from the beginning, middle and end of the advertisement video, respectively.

[0059] The text feature extraction is to detect the text area of the image of the attention concentration time frame and the key frame in the advertisement video by using a text detector, locate the text area in the image; the detected text area is subjected to text recognition by using a neural network, and a text sequence is output; the recognized text sequence is subjected to word vector coding based on a pre-trained BERT model, and a basic text feature vector is obtained; the basic text feature vector is subjected to semantic expansion by using a GPT small sample learning model, and a related expanded text vector is generated; the basic text feature vector and the expanded text feature vector are fused, and a final text feature is obtained. Specifically, in order to realize semantic understanding of the image text in the video key frame and the advertisement content, the text feature extraction step includes multiple stages, aiming to convert the text information in the original image into deep vector features that can be used for semantic matching. The process specifically includes the following operations: first, an advanced text detector (such as an algorithm based on EAST or DBNet) is used to detect the text area of the image of the user attention concentration time frame and the key feature frame in the advertisement video. This step can accurately locate the potential text area in the image and output the position information of the text box for subsequent processing. Next, a deep learning-based neural network text recognition model (such as CRNN, Rosetta, etc.) is applied to the detected text area, and the character information in the image is converted into a machine-readable text sequence. The recognition process can cope with the rotation, blur or complex background in the image, ensuring that the text information is as accurately restored as possible. The recognized text sequence is then input into a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model for word vector coding processing. The BERT model can capture the semantic relationship between texts through bidirectional modeling, thereby outputting high-quality basic text feature vectors. To further improve the semantic representation capability, the present application introduces a GPT (Generative Pre-trained Transformer) small sample learning model to perform semantic expansion on the basic text feature vector. The GPT model generates expanded text content related to the original semantics, excavates potential meanings and context supplements, and forms a set of expanded text feature vectors. This stage not only improves the understanding of short texts or abstract sentences, but also enhances the recognition ability of hidden associations in the advertisement and video context. Finally, the basic text feature vector and the expanded text feature vector are fused by using weighted splicing, attention mechanism or feature superposition, etc., to generate the final text feature representation. The final text feature has richer semantic dimensions and can accurately express the text connotation in the original image, providing stronger text semantic support for subsequent cross-modal matching.

[0060] The scene feature extraction is to extract two image scene visual feature vectors by using a Transformer model; the image features are mapped to a semantic space based on a pre-trained CLIP model, matched with preset scene category texts to generate scene semantic feature vectors; and the image scene visual feature vectors and the scene semantic feature vectors are spliced and fused to obtain scene features. Specifically, in order to realize accurate understanding and semantic modeling of the scene presented by the user attention concentrated frame image and the advertisement key frame image, the scene feature extraction method includes three core steps of visual feature extraction, semantic feature mapping and multi-modal feature fusion, and the specific process is as follows: first, an image coding model based on a Transformer structure is used to extract visual features of the image. This model can fully utilize the self-attention mechanism to model the context relationship of different regions in the image, so as to extract the global scene features of the image. Compared with the traditional convolutional neural network (CNN), the Transformer structure has stronger modeling ability for scene elements in complex backgrounds, and can identify high-level semantic categories such as "urban street scene", "office space" and "natural scenery". Then, the image visual features extracted above are input into a pre-trained CLIP (Contrastive Language-Image Pretraining) model to further realize the mapping of the image to the semantic space. The CLIP model has been trained on a large number of image-text pairs and has strong cross-modal understanding ability, which can convert image content into semantic vectors consistent with language description. In the present application, a plurality of text descriptions of typical scene categories (such as "home kitchen", "business conference room", "outdoor sports scene", etc.) are preset, and the image features are compared and matched with these text categories to obtain the scene semantic feature vectors corresponding to the image. Finally, in order to take into account the original visual information of the image and its semantic representation, a feature splicing and fusion strategy is used to jointly encode the image visual feature vectors and the scene semantic feature vectors. The fusion operation can use vector-level splicing, weighted superposition or attention guiding mechanism, aiming to form a composite scene feature vector that has both low-level visual features and high-level semantic labels. This scene feature not only has strong semantic expression ability, but also can improve the scene consistency judgment between the video and the advertisement in the subsequent cross-modal matching stage, thereby enhancing the relevance of the advertisement placement and the user acceptance.

[0061] The extraction of the text features and the scene features is respectively performed on the user attention concentrated time frames extracted from all the videos and the three key feature frames extracted from the beginning, the middle and the end of the advertisement video, and then a weighting operation is performed to obtain the final text features and scene features, that is, the text features and the scene features are actually extracted for two parts: one is the user attention concentrated time frames selected from the video watched by the user for analysis; and the other is that one key frame is extracted from the beginning, the middle and the end of each advertisement video to be inserted to reflect the overall content of the advertisement. After the text features and the scene features are extracted from the two sources of image frames, the more representative frames are highlighted through a weighting calculation method, and finally more accurate and comprehensive video and advertisement features are obtained for subsequent matching to judge the matching degree of the advertisement and the video.

[0062] Then, in order to improve the accuracy of the advertisement recommendation, the image of the user attention concentrated time frame and the features of the advertisement to be inserted are aligned through cross-modal matching, and a matching degree score of the video and the advertisement is output through a cross-modal matching model; a first determination threshold is set, and if the matching degree is greater than the first determination threshold, the advertisement is included in the candidate advertisement pool. Specifically, the image text feature vector and the image scene feature vector of the user attention concentrated time frame are mapped to the same high-dimensional semantic space through a cross-modal mapping network together with the advertisement text feature vector and the advertisement scene feature vector of the advertisement to be inserted; the cosine similarity of the image text mapping vector and the advertisement text mapping vector is calculated to obtain a text matching degree ; the cosine similarity of the image scene mapping vector and the advertisement scene mapping vector is calculated to obtain a scene matching degree ; and the matching degree score is obtained by weighting and summing the text matching degree and the scene matching degree , wherein, is a weight coefficient.

[0063] The advertisements in the candidate advertisement pool are sorted in descending order according to the matching degree score, the maximum insertion number is determined in combination with the total duration of the video, the top N advertisements in the sorting are selected as the advertisements to be inserted, and the advertisements to be inserted are inserted into the position adjacent to the attention concentrated time frame of the video, thereby generating a video containing the advertisements and pushing the video to the user.

[0064] Finally, the interactive data of the user watching the video containing the advertisements is obtained, the adhesion score of the user to the advertisement is calculated based on the interactive data, if the adhesion score is greater than a preset second determination threshold, it is determined that the advertisement and the video are matched and effective, the advertisement insertion strategy is retained, and if the adhesion score is less than or equal to the preset second determination threshold, the next advertisement is selected from the candidate advertisement pool to re-execute the insertion and judgment process. The interactive data of the video containing the advertisements includes the advertisement click rate and the advertisement stay duration proportion. The calculation method of the adhesion score of the user to the advertisement based on the interactive data is: first, the synergy index wherein is the advertisement click rate, is the advertisement dwell time proportion; calculate the adhesion score wherein is the adhesion score. When the score is higher than the threshold value, it indicates that the advertisement matches the video well, and the current advertisement insertion strategy is considered effective, and the insertion record of the advertisement in the current video will be retained; if the adhesion score is lower than or equal to the threshold value, it indicates that the advertisement effect is poor, and the system will select the next best advertisement from the candidate advertisement pool and re-execute the insertion and judgment process until the effectiveness condition is met or the candidate pool is exhausted.

[0065] The above only describes the preferred embodiments of the present application and is not intended to limit the present application in other forms. Any person skilled in the art can use the disclosed technical content to make changes or modifications to equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application without departing from the technical solution content of the present application still belongs to the protection scope of the technical solution of the present application.

Claims

1. A method for video management based on artificial intelligence, characterized in that, The method comprises the following steps: First, pure video data is obtained, the pupil center of a user is located by facial recognition, the interest value of the user is calculated by combining facial expressions, and the time frame of user attention concentration is determined; An image of the time frame of user attention concentration in the video data is selected, and image feature extraction is performed to extract image text features and image scene features; Content features of the advertisement video to be inserted are extracted, including advertisement text features and advertisement scene features; Through cross-modal matching, the features of the image of the time frame of user attention concentration and the advertisement to be inserted are aligned, and a matching score of the video and the advertisement is output by the cross-modal matching model; a first determination threshold is set, and if the matching score is greater than the first determination threshold, the advertisement is included in a candidate advertisement pool; The advertisements in the candidate advertisement pool are sorted in descending order of matching score, the maximum number of insertions is determined in combination with the total duration of the video, and the top N advertisements are selected as the advertisements to be inserted; The advertisements to be inserted are inserted into the position adjacent to the time frame of attention concentration in the video, an advertisement-containing video is generated, and the advertisement-containing video is pushed to the user; Interaction data of the user when watching the advertisement-containing video is obtained, the adhesion score of the user to the advertisement is calculated based on the interaction data, if the adhesion score is greater than a preset second determination threshold, it is determined that the advertisement and the video are matched and effective, the insertion strategy of the advertisement is retained, and if the adhesion score is less than or equal to the preset second determination threshold, the next advertisement is selected from the candidate advertisement pool to re-execute the insertion and judgment process; Before the content features of the advertisement video to be inserted are extracted, key feature frames are extracted from the advertisement video, and one frame is extracted as a key feature frame at the beginning, middle and end of the advertisement video. 2.The video management method based on artificial intelligence of claim 1, wherein, The specific implementation of locating the pupil center of the user by facial recognition and calculating the interest value of the user by combining facial expressions comprises: First, the eye ROI is extracted, the pupil edge points are extracted, and the ellipse is fitted by the least square method for the pupil edge points, and the center of the ellipse is the pupil center; Adopt lightweight CNN to extract the initial feature map of face ROI wherein is the feature map size, is the channel number, calculate the distance weight between the feature map position and the expression key node, and generate a spatial weight map , wherein, is the feature map position corresponding to the Euclidean distance of the nearest expression node in the original image, is the diagonal line length of the face ROI; Global average pooling is performed on the feature map channel dimension, and the importance of the channel is learned through a fully connected layer to generate a channel weight map , wherein, is the cth channel feature map, is the global average pooling, is the fully connected layer, is the Sigmoid function; The weight graph is fused with the basic feature graph, and an expression feature vector is obtained through global pooling , wherein is an element-wise multiplication, is an outer product expansion; The moving standard deviation of the pupil center in the sliding window is calculated, and the stability is obtained after normalization wherein is the number of frames in the window, is the mean value of the pupil center in the window, is the horizontal and vertical coordinates of the pupil center in the video frame image coordinate system at t frame, and then the pupil stability coefficient is calculated wherein D is the diagonal length of the eye ROI region; by the pre-trained SVM on the expression feature vector score, outputting an expression positivity ; The interest value is obtained by weighted fusion of the pupil stability coefficient and the expression activity. 3.The video management method based on artificial intelligence of claim 2, wherein, After the interest value is obtained, the specific implementation of determining the time frame of user attention concentration comprises: Calculating average interest value using sliding window wherein is the average interest value for the t-th frame, is the size of the sliding window, is the original interest value for the (t-w)-th frame; When greater than a set threshold, marked as a candidate frame, statistics of the duration of consecutive candidate frames , if greater than a set duration threshold, all frames in the interval are also determined as user attention concentration time frames. 4.The video management method based on artificial intelligence of claim 1, wherein, The image of the time frame of user attention concentration in the video data is selected, and image feature extraction is performed to extract image text features and image scene features; the content features of the advertisement video to be inserted are extracted, including advertisement text features and advertisement scene features, and the extraction of the features of the two comprises text feature extraction and scene feature extraction: The text feature extraction comprises: A text region in the image is located by using a text detector to detect the text region in the image and the key frame in the advertisement video; Text recognition is performed on the detected text region by using a neural network to output a text sequence; A word vector of the recognized text sequence is encoded based on a pre-trained BERT model to obtain a basic text feature vector; A GPT small sample learning model is used to perform semantic expansion on the basic text feature vector to generate a related expanded text vector; The basic text feature vector and the expanded text feature vector are fused to obtain a final text feature; The scene feature extraction comprises: The image scene visual feature vector is extracted by using a Transformer model; The image feature is mapped to a semantic space based on a pre-trained CLIP model, matched with a preset scene category text, and a scene semantic feature vector is generated; The image scene visual feature vector and the scene semantic feature vector are spliced and fused to obtain a scene feature.

5. The video management method based on artificial intelligence according to claim 4, characterized in that, The extraction of the text feature and the scene feature is a feature extraction operation performed on the user attention concentrated time frames extracted from all videos and three key feature frames extracted from the front, middle and rear of the advertisement video, and then a weighting operation is performed to obtain the final text feature and scene feature. 6.The video management method based on artificial intelligence of claim 1, wherein, The implementation of aligning the image of the user attention concentrated time frame and the feature of the advertisement to be inserted by cross-modal matching, and outputting the matching score of the video and the advertisement by the cross-modal matching model is as follows: The image text feature vector and the image scene feature vector of the user attention concentrated time frame are mapped to the same high-dimensional semantic space through a cross-modal mapping network together with the advertisement text feature vector and the advertisement scene feature vector of the advertisement to be inserted. A cosine similarity between the image text mapping vector and the advertisement text mapping vector is calculated to obtain a text matching degree ; Calculate a cosine similarity of the image scene mapping vector and the advertisement scene mapping vector to obtain a scene matching degree ; The text matching degree and the scene matching degree are weighted and summed to obtain a matching degree score wherein is a weight coefficient. 7.The video management method based on artificial intelligence of claim 1, wherein, The interaction data of the video containing the advertisement includes the advertisement click rate and the proportion of the advertisement stay time. 8.The video management method based on artificial intelligence of claim 7, wherein, The calculation method of the user's adhesion score based on the interaction data is as follows: First, the synergy index is calculated wherein is the advertisement click rate, is the advertisement dwell time proportion; calculating a sticking score wherein is the sticking score.

Citation Information

Patent Citations

  • Video advertisement broadcasting method, device and system

    CN103503463A

  • Advertisement implanting method and advertisement implanting system

    CN105141987A

  • Intelligent advertisement putting processing method and system based on user data

    CN119417538A

  • Short video advertisement putting optimization evaluation method based on deep learning

    CN119477425A

  • Dynamic advertisement content intelligent delivery method based on user emotion recognition

    CN119887307A

Cited By

  • Marketing data real-time processing method and system based on cloud edge collaboration

    CN121919263A