Advertisement video collection generation method and device, equipment and storage medium

By obtaining the scene type of advertising video and evaluating multimodal information of video clips using text matching algorithms, and generating video highlights that are highly related to advertising goals, the problem of inaccurate generation of advertising video highlights in the prior art is solved and the advertising effect is improved.

CN120343356AActive Publication Date: 2025-07-18BEIJING LIANSHI LEGEND NETWORK TECH CO LTD

Patent Information

Application Number
CN202510580237.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-07-18
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The prior art cannot accurately filter and generate video highlights that are consistent with the advertising goals based on different advertising scenario types, resulting in poor advertising results.

Method used

By obtaining the scene type of the ad, the preset text matching algorithm is used to evaluate the similarity of the text, color, audio and other information of the video clip, and a video highlight is generated based on the similarity and weight correction factors.

Benefits of technology

It improves the accuracy and pertinence of advertising video highlights, ensures the high correlation between video clips and advertising goals, and improves advertising effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343356A_ABST
    Figure CN120343356A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, solves the problem that in the prior art, a video collection matched with an advertisement target cannot be accurately screened and generated according to different scene types, and provides an advertisement video collection generation method and device, equipment and a storage medium. The method comprises the steps that the scene type of an advertisement to be put is acquired, and the scene type comprises an introduction scene, a product display scene, a user experience scene and a problem solving scene; according to the scene type, utilizing a preset text matching algorithm to carry out similarity evaluation on character information related to the video clip content and a preset text template, and determining a target similarity corresponding to each video clip; and according to the target similarity and a preset similarity threshold value, synthesizing the video clips of which the target similarity is greater than the similarity threshold value into a video collection. According to the method, the video collection which better conforms to the advertisement target and can accurately convey the advertisement information is generated, and the advertisement effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of the invention patent application with the application date of September 12, 2024, the invention name of "Video Content Analysis Method, Device and Equipment Based on Multimodal Data Processing", and the application number of 202411279794.X. Technical Field

[0002] The present invention relates to the field of image processing technology, and particularly relates to a method, device, equipment and storage medium for generating an advertising video collection. Background Art

[0003] In the current field of advertising video production and marketing, the generation of advertising video collections is a key means to improve advertising effects and user experience. With the increasingly fierce competition in the advertising market, brands and advertisers not only need to produce high-quality advertising content, but also need to improve the conversion rate and user engagement of advertisements through accurate recommendations and precise advertising placements. Traditional methods for generating advertising video collections usually rely on manual screening and basic time segment editing, or rely on rough keyword matching algorithms, resulting in the generated video collections may not be precise enough and cannot effectively resonate with the target audience, especially there are limitations in the precise matching of advertising scenarios and content.

[0004] The existing Chinese patent CN112004111A discloses a news video information extraction method for global deep learning, including: at the video decoding layer, the shot label module marks each dynamic shot through the TSM spatio-temporal model to generate the label of each dynamic shot; the similarity calculation module calculates the similarity of all labels through the BM25 algorithm, and the shot splicing module splices the dynamic shots with similar labels into a theme video; the image processing module obtains the theme video, processes each frame of the image in the theme video by using the optical flow method, the gray histogram method, the Lucas–Kanade algorithm and the image entropy calculation method to obtain key frames, and sends them to the key frame cache module for caching; at the image parsing layer, the well-known person detection module retrieves the key frames, uses the YOLOv3 model for target object detection and occupation detection, and uses the Facenet model to identify well-known persons; the key target detection module uses the Facenet model to identify the target objects in the key frames. Although the above-mentioned patent's news video information extraction method for global deep learning provides certain support for the content extraction and analysis of news videos through means such as spatio-temporal models, similarity calculations, image processing, and target detection, there are still some technical problems in the generation of advertising video collections. First of all, this patent focuses on the scene understanding and target object recognition of news videos, and mainly generates video segments through shot labels, similarity calculations, and person and target object detections. This technology faces the problem of insufficient accuracy in advertising videos because advertising videos often have higher creativity, emotional appeals, and diversity of scene types (such as introduction, display, experience, problem-solving, etc.), and the existing technical solutions cannot accurately identify and match these complex advertising scenes and themes. Secondly, using traditional image processing methods such as the optical flow method and the gray histogram method to extract key frames, although it helps to capture important content in the video, for the rapidly changing, dynamic scenes, and subtle emotional expressions that may appear in advertising videos, the effects of these methods may not be ideal, resulting in some key advertising elements or emotional turns being missed in the generated video collection, thereby affecting the conveyance of advertising effects. Finally, the processing method of this patent relies more on static images and target object recognition, lacking a comprehensive understanding and processing of the potential text information, audio information, and emotional information in advertising videos, and unable to comprehensively improve the accuracy and personalized matching degree of advertising video collections. Therefore, the above-mentioned patent has certain limitations in the accuracy and effect of generating advertising video collections, especially in the precise matching with advertising scene types, audience interests, and advertising goals, and still needs to be further optimized.

[0005] Therefore, how to accurately screen and generate video collections that match the advertising goals according to different scene types is an urgent problem to be solved. Summary of the Invention

[0006] In view of this, the present invention provides a method, device, equipment and storage medium for generating an advertising video collection, so as to solve the problem in the prior art that it is impossible to accurately screen and generate a video collection that matches the advertising target according to different scene types.

[0007] The technical solution adopted by the present invention is:

[0008] In the first aspect, the present invention provides a method for generating an advertising video collection, and the method includes:

[0009] Obtain the scene type of the advertisement to be put on the market, where the scene type includes: introduction scene, product display scene, user experience scene and problem-solving scene;

[0010] According to the scene type, use a preset text matching algorithm to evaluate the similarity between the text information related to the video clip content and the preset text template, and determine the target similarity corresponding to each video clip;

[0011] According to the target similarity and the preset similarity threshold, synthesize the video clips with the target similarity greater than the similarity threshold into a video collection.

[0012] Preferably, the step of "According to the scene type, use a preset text matching algorithm to evaluate the similarity between the text information related to the video clip content and the preset text template, and determine the target similarity corresponding to each video clip" includes:

[0013] Use a preset text matching algorithm to evaluate the similarity between the text information related to the video clip content and the preset text template, and determine the first similarity, second similarity, third similarity and fourth similarity corresponding to the personnel information, color information, item information and audio feature information respectively;

[0014] Obtain a preset weight correction factor, where the weight correction factor is greater than 1;

[0015] Obtain the initial weights corresponding to the first similarity, second similarity, third similarity and fourth similarity respectively, where the sum of the initial weights is equal to 1;

[0016] If the scene type is an introduction scene, use the weight correction factor to correct the initial weights corresponding to the second similarity and the fourth similarity;

[0017] If the scene type is a product display scene, use the weight correction factor to correct the initial weights corresponding to the second similarity and the third similarity;

[0018] If the scene type is a user experience scene, the initial weights corresponding to the first similarity and the initial weights corresponding to the third similarity are corrected using the weight correction factor;

[0019] If the scene type is a problem-solving scene, the initial weights corresponding to the first similarity, the initial weights corresponding to the third similarity, and the initial weights corresponding to the fourth similarity are corrected using the weight correction factor;

[0020] Based on the corrected initial weights, the first similarity, the second similarity, the third similarity, and the fourth similarity are weighted and averaged to determine the target similarity.

[0021] Preferably, before obtaining the scene type of the advertisement to be delivered, it further includes:

[0022] Obtain the original video material to be parsed;

[0023] Transcode the original video material to determine the standard video data in a preset encoding format;

[0024] Use the shot segmentation technology to decompose the standard video data into multiple video segments;

[0025] Use a preset feature processing algorithm to embed the blind watermark information containing video material information into each frame image in each video segment;

[0026] Compress each video segment after embedding the blind watermark information to determine the compressed video data;

[0027] Identify the blind watermark information for each frame of the compressed image and output the identification result;

[0028] If the blind watermark information is identified, obtain the video material information embedded in the blind watermark information as the target video material information;

[0029] If the blind watermark information is not identified, perform feature extraction and matching on each frame of the compressed image, and determine the target video material information based on the matching result;

[0030] Input each frame of the compressed image into a pre-trained feature extraction model to output key feature information;

[0031] Input the key feature information into a multi-modal large language model to output the text information.

[0032] Preferably, the step of using the shot segmentation technology to decompose the standard video data into multiple video segments includes:

[0033] Perform grayscale conversion on each frame image in the standard video data to obtain grayscale images for each frame;

[0034] Use an edge detection algorithm to perform edge detection on each frame of the grayscale image and output edge feature information;

[0035] Determine the edge feature change difference based on the edge feature information corresponding to adjacent frame grayscale images;

[0036] Determine video boundary information based on the edge feature change difference and a preset change threshold;

[0037] Decompose the standard video data based on the video boundary information to determine each video segment.

[0038] Preferably, the embedding of blind watermark information including video material information into each frame image in each video segment by using a preset feature processing algorithm includes:

[0039] Convert the video material information to be embedded into a binary format to determine the encoded blind watermark information;

[0040] Perform decomposition processing on each frame image in each video segment to obtain regional images;

[0041] Perform discrete cosine transform on the regional images to convert the regional images from spatial domain images to frequency domain images;

[0042] Adjust the high-frequency components in the frequency domain images and embed the blind watermark information into the frequency domain images;

[0043] Perform inverse discrete cosine transform on the frequency domain images and recombine them to complete the embedding of the blind watermark information in each video segment.

[0044] Preferably, if the blind watermark information is not recognized, the feature extraction and matching of each frame of the compressed image are performed, and based on the matching result, the determination of the target video material information includes:

[0045] Input each frame of the compressed image into a pre-trained self-supervised vision transformation model to output encoded feature information;

[0046] Use the approximate nearest neighbor algorithm to perform feature matching between the encoded feature information and the feature template information of each video material template in the video material database, and output the matching result;

[0047] Based on the matching result, output the video material template corresponding to the feature template information that matches the encoded feature information as the target video material information.

[0048] Preferably, when inputting each frame of compressed image into a pre-trained feature extraction model, the output key feature information includes:

[0049] Performing decoding processing on the compressed video data to obtain audio data;

[0050] Inputting the compressed image into a pre-trained face recognition and classification model, classifying and labeling the face features in the recognized compressed image to determine personnel information;

[0051] Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting main color information, where the main color information at least includes the advertising brand color tone detected in the compressed image or the main color tone extracted from the advertising landscape painting;

[0052] Inputting the compressed image into an object detection model, locating and classifying the objects in the compressed image to determine item information, where the item information at least includes the item category and the item location;

[0053] Inputting the audio data into a pre-trained speech transcription model to output the audio feature information in the audio data;

[0054] Inputting the personnel information, color information, item information, and audio feature information into a multi-modal large language model respectively to output the key feature information.

[0055] In a second aspect, the present invention provides an advertising video collection generation device, and the device includes:

[0056] A scene type acquisition module, configured to acquire the scene type of the advertisement to be delivered, where the scene type includes: an introduction scene, a product display scene, a user experience scene, and a problem-solving scene;

[0057] A similarity evaluation module, configured to, according to the scene type, use a preset text matching algorithm to evaluate the similarity between the text information related to the video segment content and a preset text template, and determine the target similarity corresponding to each video segment;

[0058] A video segment synthesis module, configured to synthesize the video segments with a target similarity greater than the similarity threshold into a video collection according to the target similarity and a preset similarity threshold.

[0059] In a third aspect, an embodiment of the present invention further provides an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, the method in the first aspect of the above implementation manner is implemented.

[0060] Fourthly, an embodiment of the present invention further provides a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method in the first aspect of the above-mentioned implementation manner is implemented.

[0061] In summary, the beneficial effects of the present invention are as follows:

[0062] The method, device, equipment and storage medium for generating an advertising video collection provided by the present invention, the method includes: obtaining the scene type of the advertisement to be placed, where the scene type includes: introduction scene, product display scene, user experience scene and problem-solving scene; according to the scene type, using a preset text matching algorithm, evaluating the similarity between the text information related to the video segment content and the preset text template, and determining the target similarity corresponding to each video segment; according to the target similarity and the preset similarity threshold, synthesizing the video segments with the target similarity greater than the similarity threshold into a video collection. The present invention first classifies the scenes of the advertisement video to be placed, clearly divides them into different types such as introduction scene, product display scene, user experience scene and problem-solving scene, thereby providing a more accurate framework for the generation of the advertising collection; on this basis, using a preset text matching algorithm, evaluating the similarity between the text information in the video segment and the preset scene template, and ensuring that the selected segment is highly relevant to a specific scene type by calculating the similarity between each video segment and the preset template; further, a similarity threshold is set, and when the similarity of a certain video segment is greater than this threshold, it is merged into the corresponding advertising collection. This screening method based on scene type and text similarity solves the problem that in the generation of traditional advertising video collections, it is impossible to accurately screen and match advertising content according to specific scene types, thereby generating a video collection that better meets the advertising objectives and can accurately convey advertising information, and improving the pertinence and accuracy of the advertising effect. Description of the Drawings

[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings, and these are all within the protection scope of the present invention.

[0064] Figure 1 It is a schematic flowchart of the overall work of the method for generating an advertising video collection in Embodiment 1 of the present invention;

[0065] Figure 2 It is a schematic flowchart of decomposing the standard video data into multiple video segments in Embodiment 1 of the present invention;

[0066] Figure 3Schematic flowchart of embedding blind watermark information containing video material information into each frame image of each video segment in Embodiment 1 of the present invention;

[0067] Figure 4 Schematic flowchart of identifying and analyzing the compressed video data in Embodiment 1 of the present invention;

[0068] Figure 5 Schematic flowchart of identifying blind watermark information in each frame of compressed image in the compressed video data in Embodiment 1 of the present invention;

[0069] Figure 6 Schematic flowchart of feature extraction and matching for each frame of the compressed image in Embodiment 1 of the present invention;

[0070] Figure 7 Schematic flowchart of inputting each frame of compressed image into a pre-trained feature extraction model and outputting key feature information in Embodiment 1 of the present invention;

[0071] Figure 8 Structural block diagram of an advertising video collection generation device in Embodiment 3 of the present invention;

[0072] Figure 9 Schematic diagram of the structure of an electronic device in Embodiment 4 of the present invention. Detailed implementation manners

[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device comprising the said elements. If there is no conflict, the embodiments of the present invention and the various features in the embodiments may be combined with each other, and all are within the protection scope of the present invention.

[0074] Embodiment 1

[0075] Please refer to Figure 1 , Embodiment 1 of the present invention discloses a method for generating an advertising video collection, and the method includes:

[0076] S1: Obtain the original video material to be parsed;

[0077] Specifically, obtaining the original video material to be parsed involves collecting and importing all the original video files that need to be analyzed and processed. These materials are from different camera devices, recording environments, or external resources. This process ensures that all the original video materials to be processed are centralized in a manageable storage location, facilitating subsequent transcoding, segmentation, and analysis operations. By collecting and managing these original video materials, high-quality input data can be provided for each subsequent processing step, laying the foundation for the entire video processing and analysis process.

[0078] S2: Perform transcoding processing on the original video material to determine the standard video data in a preset encoding format;

[0079] Specifically, transcoding is performed on the original video material to ensure video file format consistency and data security. When a user uploads the original video material, unified transcoding will be performed on the original video material to convert videos in various different formats into a preset standard encoding format. For example: First, identify the original format and encoding information of the uploaded video, decode the video file into uncompressed original data, and then re-encode the decoded data into a preset standard format, such as H.264 or H.265 encoding, set a unified resolution and frame rate, and add encryption and error detection mechanisms during the encoding process to improve the security of video data. This not only realizes the unified management of videos in different formats, simplifies subsequent processing steps, but also reduces the storage space requirement by optimizing the encoding format and compression parameters, while ensuring that the video quality is not significantly affected. The transcoding process also adds encryption and error detection mechanisms to further improve the security of video data and prevent data from being damaged or accessed without authorization during transmission and storage. Through this process, video materials can be managed more effectively, ensuring the compatibility and security of all data in subsequent processing.

[0080] S3: Using the shot segmentation technique, decompose the standard video data into multiple video segments;

[0081] Specifically, using the shot segmentation technique, decompose the standard video data into multiple video segments. First, identify the shot transition points, which are usually determined by significant changes in the picture content, scene transitions, or significant changes in the audio. To ensure the accuracy of segmentation, advanced algorithms such as edge detection and motion analysis are used. When the change amplitude of the picture edge exceeds the preset threshold, it is identified as a shot transition point; subsequently, the standard video data is segmented at these transition points to generate multiple consecutive and content-coherent video segments, and each segment retains the important scene information in the original video, facilitating subsequent processing and analysis.

[0082] In one embodiment, please refer to Figure 2 , S3 includes:

[0083] S31: Perform grayscale conversion on each frame image in the standard video data to obtain each frame of grayscale image;

[0084] Specifically, first, perform grayscale conversion on each frame image in the standard video data, that is, convert the color image into a grayscale image. Grayscale conversion is to convert the RGB (red, green, blue) color information of the image into a grayscale value, and calculate the grayscale value by weighted averaging the color values of each pixel. Specifically, the commonly used conversion formula is: grayscale value = 0.299 * R + 0.587 * G + 0.114 * B. Such processing can simplify the image data, reduce the computational complexity, and highlight the structural and shape features in the image, facilitating subsequent edge detection.

[0085] S32: Use an edge detection algorithm to perform edge detection on each frame of the grayscale image and output edge feature information;

[0086] Specifically, next, apply an edge detection algorithm such as Sobel, Canny, Laplacian, etc. to process each frame of the grayscale image to detect the edge information in the image. The purpose of edge detection is to find the regions in the image where the pixel values change significantly, and these regions usually correspond to the boundaries and contours of objects. The edge detection algorithm identifies the pixels with drastic changes in grayscale values in the image by calculating the image gradient or using a differential operator, and marks these pixels as edge points, thereby generating a binary image containing edge feature information.

[0087] S33: Determine the edge feature change difference based on the edge feature information corresponding to adjacent frames of grayscale images;

[0088] Specifically, after the edge detection is completed, compare the edge feature information of adjacent frames to determine the change of the edge features. The specific operation is to calculate the edge image difference between adjacent frames, and identify the regions with edge changes in the image through pixel-level differential operations. The edge feature change difference is obtained by calculating the difference of the edge pixel values at the corresponding positions of adjacent frames, and these differences can be quantified as the change amplitude, reflecting the dynamic change of the scene content over time.

[0089] S34: Determine the video boundary information based on the edge feature change difference and a preset change threshold;

[0090] Specifically, according to the calculated edge feature change difference, compare it with the preset change threshold. If the edge feature change difference between a certain frame and its adjacent frame exceeds the preset threshold, it is considered that there is a significant scene switch between this frame and the previous frame, that is, it is determined as a video boundary. The preset change threshold is used to filter out minor and insignificant changes to ensure that only obvious content changes (such as scene switches) are recognized as video boundaries. In this way, the boundary information of each shot in the video is accurately determined.

[0091] S35: Decompose the standard video data according to the video boundary information to determine each video segment.

[0092] Specifically, finally, according to the determined video boundary information, the standard video data is segmented at these boundary points to generate multiple independent video segments. Each video segment represents a continuous scene or shot, preserving the coherence and integrity of the content. The decomposed video segments facilitate subsequent processing and analysis, such as embedding blind watermarks, feature extraction, content annotation, etc. Through this method, video data can be effectively managed and processed, improving the parsing and processing efficiency of video content.

[0093] Specifically, starting from grayscale conversion, to edge detection, then to the calculation of edge feature changes and the determination of video boundaries, and finally the video is decomposed into multiple segments. This process realizes the refined processing and management of video content, laying a solid foundation for subsequent in-depth analysis and applications.

[0094] S4: Using a preset feature processing algorithm, embed the blind watermark information containing video material information into each frame image of each video segment;

[0095] Specifically, using a preset feature processing algorithm, embed the blind watermark information containing video material information into each frame image of each video segment. This is to ensure that the original video material can be accurately identified and associated during subsequent video editing. The specific operations include: First, select a suitable feature processing algorithm, such as the Discrete Cosine Transform (DCT) or the Discrete Wavelet Transform (DWT), to transform each frame image into the frequency domain. Then, select specific frequency coefficients (usually low-frequency or medium-frequency coefficients) in the frequency domain for minor adjustments and embed the blind watermark information into them. These embedded blind watermark information can contain key data such as unique material identifiers, ensuring that while not affecting the visual quality of the video, minor adjustments are made to the image so that the watermark information is invisible to the human eye but can be extracted through a specific algorithm later. Finally, use the inverse transform to restore the image to the spatial domain, generating frame images containing blind watermarks. In this way, each frame of each video segment carries unique identification information, realizing accurate material tracking and identification, and effectively solving the problem of locating video materials in the final video.

[0096] In one embodiment, please refer to Figure 3 , where S4 includes:

[0097] S41: Convert the video material information to be embedded into a binary format to determine the encoded blind watermark information;

[0098] Specifically, first, convert the video material information to be embedded (such as unique identifiers, copyright information, etc.) into binary format. Convert text or other forms of information into binary data streams through encoding algorithms, such as ASCII encoding or UTF-8 encoding. After converting this information into binary format, perform specific encryption or hashing processing to generate the final blind watermark information. The purpose of this processing is to ensure the stability and security of the information during the embedding process, facilitating subsequent embedding and extraction in the frequency-domain image.

[0099] S42: Decompose each frame image in each video segment respectively to obtain regional images;

[0100] Specifically, next, decompose each frame image in each video segment. Usually, the image is divided into several small blocks (such as 8x8 or 16x16 macroblocks) to facilitate more fine-grained frequency-domain processing. The decomposition processing enables each small block to be processed independently, ensuring that the embedding process has the least impact on the overall image. These decomposed regional images can better analyze and process local features, which is beneficial to improving the embedding efficiency and concealment of the blind watermark.

[0101] S43: Perform a discrete cosine transform on the regional images to convert the regional images from spatial-domain images to frequency-domain images;

[0102] Specifically, after obtaining the regional images, apply the discrete cosine transform (DCT) to each regional image to convert it from the spatial domain to the frequency domain. The DCT decomposes the image data into different frequency components, making the energy of the image concentrated in a few low-frequency components. These low-frequency components often correspond to the main features of the image. Through this transformation, blind watermark information can be more effectively embedded in the frequency domain without significantly affecting the visual quality of the original image.

[0103] S44: Adjust the high-frequency components in the frequency-domain images and embed the blind watermark information into the frequency-domain images;

[0104] Specifically, select the high-frequency components in the frequency-domain images for adjustment and embed the blind watermark information into them. The specific method is to make slight adjustments to the selected DCT coefficients according to the binary blind watermark information, such as increasing or decreasing the values of specific coefficients to represent binary 0 or 1. The low-frequency components contain the main information of the image, and slight adjustments will not significantly affect the image quality. The middle-frequency components balance the concealment and robustness of information embedding. In this way, the blind watermark information is securely and covertly embedded into the frequency-domain features of the image.

[0105] S45: Perform an inverse discrete cosine transform on the frequency-domain images and then recombine them to complete the embedding of the blind watermark information in each video segment.

[0106] Specifically, finally, the adjusted frequency domain image is subjected to inverse discrete cosine transform (IDCT) to convert it back to the spatial domain image and restore the visual effect of the original image. After IDCT processing, each regional image is reassembled back into the original frame image to ensure that each frame retains the blind watermark information completely. In this way, the blind watermark information is seamlessly embedded into each frame image in the video clip, and the stable embedding and subsequent extractability of the information are ensured without affecting the overall visual quality. Each video clip processed in this way carries unique identification information, realizing accurate tracking and protection of the material.

[0107] S5: compressing each video segment after the blind watermark information is embedded to determine compressed video data;

[0108] Specifically, each video clip after the blind watermark information is embedded is compressed to determine the compressed video data, the original film is stored in the company's intranet environment, and a low-resolution version of the video is compressed and stored in the cloud for labeling and analysis. Video compression technology significantly reduces the size of video files by optimizing encoding and reducing resolution, thereby reducing storage and transmission costs. The original film is stored in the intranet to ensure the high quality of the video and data security, and prevent unauthorized access. The low-resolution version of the video is sufficient to meet the needs of content recognition and annotation, and its smaller file size makes cloud storage and access more efficient. Storing low-resolution videos in the cloud enables remote access and sharing, facilitates collaboration and data processing among team members in different locations, improves work efficiency and flexibility, and ensures data security and consistency. This solution takes into account the quality, security and ease of use of video data through efficient compression and distributed storage, providing the team with a flexible working environment and strong technical support.

[0109] S6: using a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information.

[0110] Specifically, using a preset feature extraction algorithm, the compressed video data is identified and analyzed to output text information related to the video content and target video material information. This process first analyzes each frame of the compressed video through advanced feature extraction algorithms, such as multi-modal large language models, speech transcription models, face recognition models, object detection models, and color analysis models. These algorithms can identify various contents in the video, such as text, speech, faces, objects, and colors. Then, the identified information is integrated and analyzed to generate detailed text descriptions and annotations, including scene descriptions, dialogue contents, character information, and object positions. These information can not only be used for the understanding and retrieval of video content, but also help generate accurate identifications and descriptions related to video materials, ensuring that these materials can be accurately associated and utilized in subsequent use. This process greatly improves the efficiency and accuracy of video content analysis and provides strong technical support for content management and applications.

[0111] In one embodiment, please refer to Figure 4 , S6 includes:

[0112] S61: Identify the blind watermark information for each frame of the compressed image and output the identification result;

[0113] Specifically, use the extraction algorithm corresponding to the preset blind watermark embedding algorithm to process each frame of the image in the compressed video data, extract the embedded blind watermark information, and determine the target video material information based on the extraction result. This step ensures that in the case of compressed video, the source and specific content of the video segment can still be accurately located and identified, thus ensuring the traceability of the material and the effectiveness of content management.

[0114] In one embodiment, please refer to Figure 5 , S61 includes:

[0115] S611: Input each frame of the compressed image into a pre-trained self-supervised vision transformation model and output encoded feature information;

[0116] Specifically, first, input each frame of the compressed image into a pre-trained self-supervised vision transformation model (such as DINO-V2). This model learns rich visual features from a large amount of unlabeled data through self-supervised learning. The model processes the input image, extracts the key features of the image, and encodes these features into high-dimensional feature vectors. These feature vectors contain various information such as the texture, edges, shape, and color of the image, and can accurately describe the image content. The output encoded feature information will be used as the basis for subsequent feature matching.

[0117] S612: Using the approximate nearest neighbor algorithm, perform feature matching between the encoded feature information and the feature template information of each video material template in the video material database, and output the matching result;

[0118] Specifically, then, perform feature matching on the encoded feature information using the approximate nearest neighbor algorithm (ANN). The approximate nearest neighbor algorithm can quickly find the feature template most similar to the target feature in the high-dimensional feature space. Compare the encoded feature information of each frame of image with the feature template information of each video material template pre-stored in the video material database to calculate the similarity. Through the ANN algorithm, find the feature template closest to the encoded feature information and output the matching result. The matching result includes the matching degree and relevant information of each frame of image with the most similar material template in the database.

[0119] S613: According to the matching result, output the video material template corresponding to the feature template information that matches the encoded feature information as the target video material information.

[0120] Specifically, finally, determine the target video material information according to the feature matching result. For each frame of image, select the feature template information with the highest matching degree and output its corresponding video material template as the target video material information to ensure that even in the absence of blind watermark information, the source and specific content of the video segment can be accurately identified through feature matching. In this way, each frame of image can be efficiently and accurately associated with the original video material, realizing the precise identification and management of video content.

[0121] S62: If the blind watermark information is recognized, obtain the video material information embedded in the blind watermark information as the target video material information;

[0122] Specifically, after successfully recognizing the blind watermark information, extract the embedded video material information from the recognized watermark data. The blind watermark usually contains a unique material identifier, copyright information, or other relevant metadata. Through these information, the source and specific content of the video material can be accurately determined and used as the target video material information to ensure that each frame of image can be accurately associated with the corresponding original material.

[0123] S63: If the blind watermark information is not recognized, perform feature extraction and matching on each frame of the compressed image, and determine the target video material information according to the matching result;

[0124] Specifically, if the blind watermark information cannot be recognized, a feature extraction and matching method is used to further analyze the compressed image. The extracted features include information such as texture, edge, shape, and color, forming a unified image feature encoding. Then, the extracted image features are matched with the material features in the database to determine the most similar material, ensuring that even in the absence of a blind watermark, the source and content of the video clip can be accurately identified. Based on the matching results, the target video material information is determined and output, thus realizing the precise identification and management of materials.

[0125] S64: Input each frame of the compressed image into a pre-trained feature extraction model to output key feature information;

[0126] Specifically, inputting each frame of the compressed image into a pre-trained feature extraction model to output key feature information involves processing each frame of the image through a pre-trained feature extraction model in the field of computer vision, such as a convolutional neural network (CNN) or a Transformer model. This model has been trained on a large amount of image data and can automatically extract important visual features in the image, such as edges, textures, color distributions, and shapes. After inputting the image, the model will perform a series of convolutional layers, activation functions, and pooling operations to gradually extract high-level features in the image. These features will be encoded into a high-dimensional feature vector, which contains the core information of the image. These encoded key feature information can not only accurately describe the image content but also be used for subsequent feature matching, classification, or retrieval tasks, ensuring the effective utilization and analysis of image data.

[0127] In one embodiment, please refer to Figure 6 , where S64 includes:

[0128] S641: Decode the compressed video data to obtain audio data;

[0129] Specifically, first, decode the compressed video data to extract the audio data. This process includes restoring the compressed video file to its original uncompressed state and extracting the audio part from the video stream through a decoder. The decoding process usually involves converting the compressed encoded data stream (such as MP4 or H.264 format) into an uncompressed audio format that can be further processed (such as WAV or PCM). The extracted audio data contains all the sound information in the video, including dialogue, background music, and ambient sound, providing a basis for subsequent audio analysis and processing.

[0130] S642: Input the compressed image into a pre-trained face recognition classification model to classify and label the face features in the recognized compressed image and determine the personnel information;

[0131] S643: Input the compressed image into a pre-trained color analysis model to analyze the color distribution in the compressed image and extract the main color information, where the main color information at least includes the advertising brand hue detected in the compressed image or the main hue extracted from the advertising landscape painting;

[0132] S644: Input the compressed image into an object detection model to locate and classify the objects in the compressed image and determine the item information, where the item information at least includes the item category and the item location;

[0133] Specifically, input the compressed image into a pre-trained face recognition classification model, a color analysis model, and an object detection model respectively to output the personnel information, color information, and item information. Specifically, this step involves passing each frame of the compressed image into multiple specially trained deep learning models to extract and analyze different types of information. First, input the image into the face recognition classification model, which is trained using a large amount of pre-annotated facial data and can accurately identify the facial features in the image and perform classification and annotation. For example, it can recognize famous people, actors, or other known individuals appearing in the image, such as recognizing the facial features of a reporter in a news report. Second, input the image into the color analysis model, which analyzes the color distribution in the image and extracts the main color information, such as the brand hue detected in an advertising image or the main hue extracted from a landscape painting. Finally, input the image into the object detection model, which locates and classifies the objects in the image, such as recognizing the products in a shopping advertisement or traffic signs in a video, and determines their categories and locations. These models each process different aspects of the image. By integrating this information, a comprehensive description of the image content can be obtained. For example, in a movie clip, the actors, background colors, and main objects in the scene can be recognized, providing key details for in-depth understanding and analysis of the video.

[0134] S645: Input the audio data into a pre-trained speech transcription model to output the audio feature information in the audio data;

[0135] Specifically, input the extracted audio data into a pre-trained speech transcription model to convert the speech content in the audio into text information. The speech transcription model uses natural language processing technology and speech recognition algorithms to recognize the language, tone, and syllables in the audio and transcribe them into text. This process can recognize and record conversations, speeches, or other speech information and generate corresponding text descriptions. The audio feature information includes the text form of the speech content, enabling the information in the audio to be further analyzed and processed.

[0136] S646: Input the personnel information, color information, item information, and audio feature information into a multi-modal large language model respectively to output the key feature information.

[0137] Specifically, the personnel information, color information, item information extracted from the image, and the audio feature information extracted from the audio are input into the multi-modal large language model. The multi-modal large language model can integrate data from different sources, understand and analyze the relationships between various types of information. The model will comprehensively analyze these multi-modal data and generate key feature information describing the video content. This includes integrating the visual features in the image and the speech content in the audio to provide a comprehensive understanding and summary of the content. This fusion analysis method enables the effective integration and interpretation of various information extracted from the video, generating profound insights into the video content.

[0138] S65: Input the key feature information into the multi-modal large language model and output the text information.

[0139] Specifically, inputting the key feature information into the multi-modal large language model to output the text information first involves integrating the key feature information extracted from the image (such as personnel information, color information, item information) and the audio (such as speech transcription text) into a unified data format. These feature information will be input into the multi-modal large language model, which uses deep learning technology to integrate data of different modalities (vision, audio, text, etc.) and understand the relationships between these information. By analyzing and contextually modeling the input multi-modal data, the model can generate a coherent text description that accurately expresses the key points of the video content. The output text information usually includes a summary of the video scene, dialogue content, detailed descriptions of people and objects, and other relevant information. This comprehensive analysis ability enables the multi-modal large language model to provide a comprehensive and accurate summary of the video content, enhancing the understanding and utilization of video data.

[0140] Embodiment 2

[0141] In the above Embodiment 1, analysis results related to the content of each video segment are obtained through content analysis, including personnel information, color information, item information, and audio feature information. In the actual advertising video editing scenario, the relevance of the video segments is evaluated based on the content analysis results, and a video collection is edited accordingly, which can significantly improve the quality and viewing experience of the video. For this reason, in one implementation, please refer to Figure 7 , after S6, it further includes:

[0142] Obtain the scene type of the advertisement to be launched, where the scene type includes: introduction scene, product display scene, user experience scene, and problem-solving scene;

[0143] Specifically, obtain the scene type of the advertisement to be launched. Advertisements are usually divided into multiple scene types, such as introduction scenes, product display scenes, user experience scenes, and problem-solving scenes, etc. Introduction scenes are usually used to attract the audience's attention and present the core information or brand of the advertisement; product display scenes focus on showing the functions, appearance, and usage methods of products to help the audience intuitively understand the value of the products; user experience scenes show how users use the products and the experiences and feedback obtained during the usage process; problem-solving scenes are used to show how the products help users solve practical problems and emphasize the effects and practicality of the products. The steps to obtain the scene type can be completed through the analysis of advertisement planning content, brand requirements, and market demands. The advertisement editing team usually pre-sets these scene types to ensure that the advertisement can be displayed according to the set plot or marketing strategy. At the same time, video analysis technology is used to extract key features (such as products, people, scene elements) in the video and match these features with the predefined scene types to quickly determine the scene category to which the video segment belongs.

[0144] According to the scene type, use a preset text matching algorithm to evaluate the similarity between the text information related to the video segment content and the preset text template, and determine the target similarity corresponding to each video segment;

[0145] Specifically, after confirming the scene type, further analyze the video segment according to these scene types. Specifically, use a text matching algorithm to determine the similarity between the segment and the target scene. First, the content of the video segment is usually converted into text form through a speech transcription model or subtitle extraction technology. These text information reflects the core content of the video segment, such as product names, function descriptions, user feelings, etc. Use text matching algorithms, such as TF-IDF, BERT, etc., to calculate the similarity between the text information in the video segment and the preset scene text template. Each scene type has a set of preset text templates, and the text templates contain keywords or sentence patterns related to the scene. The core of the similarity evaluation is to calculate the semantic similarity between the text in the video segment and the template text, and then evaluate the matching degree between the video segment and the scene type.

[0146] In one embodiment, the step of using a preset text matching algorithm to evaluate the similarity between the text information related to the video segment content and the preset text template according to the scene type and determine the target similarity corresponding to each video segment includes:

[0147] Use a preset text matching algorithm to evaluate the similarity between the text information related to the video segment content and the preset text template, and determine the first similarity, second similarity, third similarity, and fourth similarity corresponding to personnel information, color information, item information, and audio feature information respectively;

[0148] Specifically, first, a text matching algorithm is used to process the text information related to video clips. The text information includes personnel information, color information, item information, and audio feature information. The similarity between this text information and a preset text template is evaluated. The text template usually represents a specific scene type, such as an introduction scene, a product display scene, etc., and contains keywords and related sentences. Through algorithms such as TF-IDF or BERT, the semantic similarity between the video text information and the template is calculated, and the first similarity, the second similarity, the third similarity, and the fourth similarity corresponding to the personnel information, color information, item information, and audio feature information are respectively determined. These similarities reflect the matching degree of each video clip with the relevant features in the scene.

[0149] Obtain a preset weight correction factor, where the weight correction factor is greater than 1;

[0150] Specifically, in order to adjust the evaluation of different similarities, a preset weight correction factor is obtained. This correction factor is usually greater than 1, indicating that some similarities need to be amplified. The setting of the weight correction factor is based on specific scene requirements. For example, in some scenes, color information may be more important than other features, so the weight of color similarity needs to be increased. The weight correction factor is used to dynamically adjust the importance of different similarities, so that the final result can better meet the scene requirements and advertising planning goals.

[0151] Obtain the initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity respectively, where the sum of the initial weights is equal to 1;

[0152] Specifically, after calculating the similarity, an initial weight is assigned to each similarity. The initial weight determines the influence of each similarity in the total similarity calculation, and the sum of all initial weights is equal to 1. According to the advertising planning goal, the assignment of the initial weight can be preset to different values. For example, in a product display scene, a higher weight may be given to item information, while in a user experience scene, a higher weight may be given to personnel information.

[0153] If the scene type is an introduction scene, use the weight correction factor to correct the initial weights corresponding to the second similarity and the fourth similarity;

[0154] Specifically, in the introduction scenario, the color information and audio feature information of the video are usually the key factors attracting the audience's attention. Therefore, the similarity weights related to these features, namely the second similarity and the fourth similarity, are corrected using the weight correction factor. The correction process includes multiplying the weight correction factor by the second similarity to obtain a new second similarity, and multiplying the weight correction factor by the fourth similarity to obtain a new fourth similarity. This means that the initial weights of color and audio are amplified, making these features play a greater role in similarity evaluation and ensuring that the visual and auditory effects of the introduction scenario are more attractive.

[0155] If the scenario type is a product demonstration scenario, the weight correction factor is used to correct the initial weights corresponding to the second similarity and the initial weight corresponding to the third similarity.

[0156] Specifically, in the product demonstration scenario, the color information and item information in the video are particularly important. Therefore, the weight correction factor is used to correct the second similarity and the third similarity. This adjustment ensures that the visual effect and item details during product demonstration are more prominent, attracting the audience's attention and prompting them to focus on the product's features.

[0157] If the scenario type is a user experience scenario, the weight correction factor is used to correct the initial weights corresponding to the first similarity and the initial weight corresponding to the third similarity.

[0158] Specifically, the user experience scenario focuses on demonstrating the interaction experience between people and products. Therefore, the weight correction factor is used to correct the initial weights corresponding to the first similarity and the initial weight corresponding to the third similarity. In this type of scenario, the weights of these two types of information are amplified to ensure that the interaction process between the user and the product and the display of the product itself are more prominent, enhancing the user's sense of immersion.

[0159] If the scenario type is a problem-solving scenario, the weight correction factor is used to correct the initial weights corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity.

[0160] Specifically, in the problem-solving scenario, it is usually necessary to highlight the personnel information, item information, and audio feature information to show how the product solves the user's actual problems. Therefore, the weight correction factor is used to correct the first similarity, the third similarity, and the fourth similarity. This can better present how the product solves the user's problems through interaction with the user, the problem-solving process, and relevant audio prompts, enhancing the persuasiveness and authenticity of the scenario.

[0161] According to the corrected initial weights, perform weighted average processing on the first similarity, second similarity, third similarity, and fourth similarity to determine the target similarity;

[0162] Specifically, process the previously obtained first similarity, second similarity, third similarity, and fourth similarity. Each similarity represents the matching degree of different types of information. In the previous steps, each similarity was corrected for weights according to the scenario requirements to ensure that the content emphasized by the scenario is highlighted. The weighted average processing means multiplying each corrected initial weight by the corresponding similarity value, then adding these results and taking the average to calculate the overall target similarity. The target similarity indicates the matching degree of the video segment with a specific scenario after integrating multi-modal information. Through this process, the system can comprehensively evaluate the fit degree of each video segment with the preset scenario under multi-modal information.

[0163] According to the target similarity and the preset similarity threshold, synthesize the video segments with the target similarity greater than the similarity threshold into a video collection.

[0164] Specifically, after obtaining the target similarity, compare this value with a preset similarity threshold. The similarity threshold is a standard preset by the system to determine whether a video segment meets the minimum requirements of the scenario. This threshold can be adjusted according to the requirements of different scenarios. For example, the product display scenario may require a higher similarity threshold, while the introduction scenario may allow a lower threshold. The similarity threshold is usually a standard set based on user expectations or advertising goals. If the target similarity is greater than or equal to this threshold, it means that the video segment has a high correlation and can well match the requirements of the scenario and is suitable for the next step of processing. Those segments with a similarity lower than the threshold will be considered not to match the current scenario and thus be excluded from the selection of the video collection. After determining which video segments meet the similarity requirements, the system will perform synthesis processing on these segments. The synthesis process includes combining multiple video segments with high similarity according to a certain logic or strategy, such as chronological order, plot development, visual consistency, etc., to form a complete advertising collection. This step can be adjusted according to the design of the marketing strategy. For example, combine product display segments with user experience segments, or arrange problem-solving segments and product display segments in an orderly manner to ensure that the collection has a coherent narrative structure.

[0165] Embodiment 3

[0166] Please refer to Figure 8 , Embodiment 3 of the present invention also provides an advertising video collection generation device, and the device includes:

[0167] A video material acquisition module, configured to acquire the original video material to be parsed;

[0168] A transcoding processing module, used for performing transcoding processing on the original video material to determine standard video data in a preset coding format;

[0169] A shot segmentation module, used to decompose the standard video data into multiple video segments by using a shot segmentation technology;

[0170] A blind watermark embedding module is used to embed blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm;

[0171] A compression processing module, used to compress each video segment after the blind watermark information is embedded, and determine the compressed video data;

[0172] The feature extraction module is used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information.

[0173] Specifically, the advertising video highlights generation device provided by the embodiment of the present invention is adopted, and the device includes: a video material acquisition module, which is used to acquire the original video material to be parsed; a transcoding processing module, which is used to transcode the original video material and determine the standard video data of a preset encoding format; a shot segmentation module, which is used to decompose the standard video data into multiple video segments using the shot segmentation technology; a blind watermark embedding module, which is used to embed the blind watermark information containing the video material information into each frame image in each video segment using a preset feature processing algorithm; a compression processing module, which is used to compress each video segment after the blind watermark information is embedded, and determine the compressed video data; a feature extraction module, which is used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and the target video material information. This device achieves accurate association between finished films and materials, and accurately identifies and analyzes video content through a series of orderly processing steps. First, by acquiring and transcoding the original video material, all data formats are ensured to be unified. Then, the standard video data is decomposed into multiple easy-to-manage segments using shot segmentation technology, and blind watermark information is embedded, so that each material segment can be accurately identified in the finished film, even after editing and compression. Finally, the compressed video is identified and analyzed through a feature extraction algorithm, and detailed text information related to the video content and target video material information are output. This method not only improves the matching accuracy of video materials and finished films, ensuring that each piece of material can be accurately located and identified, but also can efficiently extract and summarize video content, provide more detailed analysis and annotation, and solve the efficiency and accuracy problems in traditional video processing and analysis.

[0174] Example 4

[0175] In addition, in combination with Figure 1 the method for generating an advertising video collection according to Embodiment 1 of the present invention described above can be implemented by an electronic device. Figure 9 FIG. shows a schematic hardware structure diagram of the electronic device provided in Embodiment 4 of the present invention.

[0176] The electronic device may include a processor and a memory storing computer program instructions.

[0177] Specifically, the above-mentioned processor may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0178] The memory may include a mass storage for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In a suitable case, the memory may include a removable or non-removable (or fixed) medium. In a suitable case, the memory may be internal or external to the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory or a combination of two or more of these.

[0179] The processor reads and executes the computer program instructions stored in the memory to implement any one of the advertising video collection generation methods in the above embodiments.

[0180] In one example, the electronic device may further include a communication interface and a bus. Among them, as Figure 9 shown, the processor, the memory, and the communication interface are connected by the bus and complete communication with each other.

[0181] The communication interface is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present invention.

[0182] The bus includes hardware, software, or both, and couples components of the device together. By way of example and not limitation, the bus can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus can include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0183] Embodiment 5

[0184] In addition, in combination with the advertisement video collection generation method in the above Embodiment 1, Embodiment 5 of the present invention can also be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the advertisement video collection generation methods in the above embodiments is implemented.

[0185] In summary, embodiments of the present invention provide an advertisement video collection generation method, apparatus, device, and storage medium.

[0186] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0187] The functional blocks shown in the above-described structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0188] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant location, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0189] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps. That is to say, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.

[0190] As described above, the above is only the specific implementation manner of the present invention. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for generating an advertising video collection, characterized in that, The method includes: Obtaining the scene type of the advertisement to be delivered, where the scene type includes: introduction scene, product display scene, user experience scene, and problem-solving scene; According to the scene type, using a preset text matching algorithm, evaluating the similarity between the text information related to the video clip content and a preset text template, and determining the target similarity corresponding to each video clip; According to the target similarity and a preset similarity threshold, synthesizing the video clips with a target similarity greater than the similarity threshold into a video collection.

2. The method for generating an advertising video collection according to claim 1, characterized in that The step of according to the scene type, using a preset text matching algorithm, evaluating the similarity between the text information related to the video clip content and a preset text template, and determining the target similarity corresponding to each video clip includes: Using a preset text matching algorithm, evaluating the similarity between the text information related to the video clip content and a preset text template, and determining the first similarity, second similarity, third similarity, and fourth similarity corresponding to the personnel information, color information, item information, and audio feature information respectively; Obtaining a preset weight correction factor, where the weight correction factor is greater than 1; Obtaining the initial weights corresponding to the first similarity, second similarity, third similarity, and fourth similarity respectively, where the sum of the initial weights is equal to 1; If the scene type is an introduction scene, using the weight correction factor to correct the initial weights corresponding to the second similarity and the fourth similarity; If the scene type is a product display scene, using the weight correction factor to correct the initial weights corresponding to the second similarity and the third similarity; If the scene type is a user experience scene, using the weight correction factor to correct the initial weights corresponding to the first similarity and the third similarity; If the scene type is a problem-solving scene, using the weight correction factor to correct the initial weights corresponding to the first similarity, the third similarity, and the fourth similarity; According to the corrected initial weights, performing a weighted average process on the first similarity, second similarity, third similarity, and fourth similarity to determine the target similarity.

3. The method for generating an advertising video collection according to claim 1, characterized in that, Before the step of obtaining the scene type of the advertisement to be delivered, it further includes: Obtaining the original video material to be analyzed; Performing transcoding processing on the original video material to determine standard video data in a preset coding format; Using a shot segmentation technique to decompose the standard video data into multiple video clips; Using a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image of each video clip; Performing compression processing on each video clip after embedding the blind watermark information to determine compressed video data; Performing blind watermark information recognition on each frame of the compressed image and outputting the recognition result; If the blind watermark information is recognized, obtaining the video material information embedded in the blind watermark information as the target video material information; If the blind watermark information is not recognized, feature extraction and matching are performed on each frame of the compressed image, and based on the matching result, the target video material information is determined; Input each frame of the compressed image into a pre-trained feature extraction model to output key feature information; Input the key feature information into a multi-modal large language model to output the text information.

4. The method for generating an advertising video collection according to claim 3, wherein The use of shot segmentation technology to decompose the standard video data into multiple video segments includes: Perform grayscale conversion on each frame image in the standard video data to obtain each frame of grayscale image; Use an edge detection algorithm to perform edge detection on each frame of the grayscale image to output edge feature information; Determine the edge feature change difference based on the edge feature information corresponding to adjacent frame grayscale images; Determine the video boundary information based on the edge feature change difference and a preset change threshold; Decompose the standard video data according to the video boundary information to determine each video segment.

5. The method for generating an advertising video collection according to claim 3, wherein, The use of a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image in each video segment includes: Convert the video material information to be embedded into a binary format to determine the encoded blind watermark information; Perform decomposition processing on each frame image in each video segment to obtain regional images; Perform discrete cosine transform on the regional images to convert the regional images from spatial domain images to frequency domain images; Adjust the high-frequency components in the frequency domain images to embed the blind watermark information into the frequency domain images; Perform inverse discrete cosine transform on the frequency domain images and then recombine them to complete the embedding of the blind watermark information in each video segment.

6. The method for generating an advertising video collection according to claim 4, characterized in that, If the blind watermark information is not recognized, feature extraction and matching are performed on each frame of the compressed image, and based on the matching result, the target video material information is determined, including: Input each frame of the compressed image into a pre-trained self-supervised vision transformation model to output encoded feature information; Use the approximate nearest neighbor algorithm to perform feature matching between the encoded feature information and the feature template information of each video material template in the video material database, and output the matching result; Based on the matching result, output the video material template corresponding to the feature template information that matches the encoded feature information as the target video material information.

7. The method for generating an advertising video collection according to claim 3, characterized in that Inputting each frame of the compressed image into a pre-trained feature extraction model to output key feature information includes: Perform decoding processing on the compressed video data to obtain audio data; Input the compressed image into a pre-trained face recognition classification model to classify and label the face features in the recognized compressed image to determine personnel information; Input the compressed image into a pre-trained color analysis model to analyze the color distribution in the compressed image and extract the main color information, where the main color information at least includes the advertisement brand color tone detected in the compressed image or the main color tone extracted from the advertisement landscape painting; Input the compressed image into an object detection model to locate and classify the objects in the compressed image to determine item information, where the item information at least includes the item category and the item location; Input the audio data into a pre-trained speech transcription model to output the audio feature information in the audio data; Input the personnel information, color information, item information, and audio feature information into a multi-modal large language model respectively to output the key feature information.

8. An advertising video collection generation device, characterized in that, The device includes: A scene type acquisition module, configured to acquire the scene type of the advertisement to be delivered, where the scene type includes: an introduction scene, a product display scene, a user experience scene, and a problem-solving scene; A similarity evaluation module, configured to, according to the scene type, use a preset text matching algorithm to evaluate the similarity between the text information related to the video segment content and a preset text template, and determine the target similarity corresponding to each video segment; A video segment synthesis module, configured to synthesize the video segments with a target similarity greater than the similarity threshold into a video collection according to the target similarity and a preset similarity threshold.

9. An electronic device, characterized in that, Including: At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1-7 when the computer program instructions are executed by the processor.

10. A storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • News video information extraction method for global deep learning

    CN112004111A

  • Method and system for monitoring advertisement broadcast

    CN102799605A

  • Method and device for detecting advertisements

    CN103235956A

  • Video analysis platform, matching method, accurate advertisement delivery method and system

    CN106686404A

  • Method and system for converting text into video

    CN108986186A

Cited By

  • Feature-enhanced advertising service scene image generation method

    CN120852570A

  • A feature-enhanced advertisement service scene image generation method

    CN120852570B