Advertisement video highlight generation method and device, equipment and storage medium
By acquiring the scene type of the advertising video and using text matching algorithms to evaluate the multimodal information of the video segments, the problem of insufficient accuracy in the generation of advertising video compilations was solved, resulting in video compilations that are more in line with advertising objectives and improving advertising effectiveness.
Patent Information
- Application Number
- CN202510580237.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Existing technologies are unable to accurately filter and generate video compilations that match advertising objectives based on different scenario types, especially lacking accurate identification and matching in terms of the complexity of creativity and emotional expression in advertising videos.
By obtaining the scene type of the advertisement to be placed, a preset text matching algorithm is used to evaluate the similarity of text, audio, color and object information of video clips, and combined with a weight correction factor, a video compilation with a target similarity higher than the threshold is generated.
It enables precise filtering and matching based on the type of advertising video scene, generating video compilations that better match advertising objectives, thereby improving the targeting and accuracy of advertising effectiveness.
Smart Images

Figure CN120343356B_ABST
Abstract
Description
[0001] The present application is a divisional application of the invention patent application with the application number 202411279794.X, the title of which is "Video content analysis method, device and equipment based on multi-modal data processing", and which was filed on September 12, 2024. TECHNICAL FIELD
[0002] The present application relates to the technical field of image processing, and in particular to an advertisement video highlight generation method, device, equipment and storage medium. BACKGROUND
[0003] In the current field of advertisement video production and marketing, the generation of advertisement video highlights is a key means to improve advertisement effectiveness and user experience. With the increasing competition in the advertisement market, brands and advertisers not only need to produce high-quality advertisement content, but also need to improve the conversion rate of advertisements and user engagement through precise recommendation and precise advertisement placement. Traditional advertisement video highlight generation methods usually rely on manual screening and basic time segment editing, or rely on rough keyword matching algorithms, resulting in generated video highlights that may not be accurate enough to effectively resonate with target audiences, especially in the precise matching of advertisement scenarios and content.
[0004] The existing Chinese patent CN112004111A discloses a global deep learning news video information extraction method, which includes: in the video decoding layer, each dynamic shot is labeled by the TSM spatio-temporal model through the lens label module, generating the label of each dynamic shot; the similarity calculation module calculates the similarity of all labels through the BM25 algorithm, and the lens splicing module splices the dynamic shots with similar labels into a theme video; the image processing module obtains the theme video, processes each frame of image in the theme video using the optical flow method, the gray histogram method, the Lucas-Kanade algorithm and the image entropy calculation method, obtains the key frame, and sends it to the key frame cache module for caching; in the image analysis layer, the famous person detection module calls the key frame, uses the YOLOv3 model for target object detection and occupation detection, and uses the Facenet model to identify famous people; the key target detection module identifies the target object in the key frame using the Facenet model; the global deep learning news video information extraction method proposed in the above patent provides certain support for content extraction and analysis of news videos through spatio-temporal model, similarity calculation, image processing and target detection, etc., but there are still some technical problems in the generation of advertisement video highlights. First of all, this patent focuses on the scene understanding and target object recognition of news videos, mainly generating video clips through lens label, similarity calculation and person and target object detection. This technology faces the problem of insufficient accuracy in advertisement videos, because advertisement videos often have higher creativity, emotional appeal and diversity of scene types (such as introduction, display, experience, problem solving, etc.), and the existing technical solutions cannot accurately identify and match these complex advertisement scenes and themes; secondly, the key frame is extracted based on traditional image processing methods such as optical flow method and gray histogram method, which helps to capture important content in the video, but for the rapid changes, dynamic scenes and subtle emotional expressions that may appear in advertisement videos, the effect of these methods may not be ideal, resulting in the omission of some key advertisement elements or emotional turning points in the generated video highlights, and thus affecting the conveyance of advertisement effect. Finally, the processing method of this patent relies more on static images and target object recognition, lacks comprehensive understanding and processing of potential text information, audio information and emotional information in advertisement videos, and cannot fully improve the accuracy and personalized matching degree of advertisement video highlights. Therefore, the above patent has certain limitations in the accuracy and effect of advertisement video highlight generation, especially in the accurate matching of advertisement scene types, audience interests and advertisement targets, which still needs to be further optimized.
[0005] Therefore, how to accurately screen and generate video highlights that match the advertisement target according to different scene types is a problem to be solved. SUMMARY
[0006] Therefore, the application provides an advertisement video highlight generation method and device, equipment and a storage medium to solve the problem that the prior art cannot accurately screen and generate video highlights that match the advertisement target according to different scene types.
[0007] The technical scheme adopted by the application is:
[0008] In a first aspect, the application provides an advertisement video highlight generation method, which comprises:
[0009] Obtaining a scene type of an advertisement to be launched, wherein the scene type comprises an introduction scene, a product display scene, a user experience scene and a problem solving scene;
[0010] According to the scene type, using a preset text matching algorithm to perform similarity evaluation on the text information related to the video segment content and the preset text template to determine the target similarity corresponding to each video segment;
[0011] According to the target similarity and a preset similarity threshold, video segments with a target similarity greater than the similarity threshold are combined into a video highlight.
[0012] Preferably, the similarity evaluation on the text information related to the video segment content and the preset text template according to the scene type using the preset text matching algorithm to determine the target similarity corresponding to each video segment comprises:
[0013] Using a preset text matching algorithm to perform similarity evaluation on the text information related to the video segment content and the preset text template to determine the first similarity, the second similarity, the third similarity and the fourth similarity corresponding to the personnel information, the color information, the article information and the audio feature information respectively;
[0014] Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1;
[0015] Obtaining the initial weight corresponding to each of the first similarity, the second similarity, the third similarity and the fourth similarity, wherein the sum of the initial weights is equal to 1;
[0016] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;
[0017] If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor;
[0018] If the scene type is a user experience scene, the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity are corrected by using the weight correction factor;
[0019] If the scene type is a problem solving scene, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected by using the weight correction factor;
[0020] The first similarity, the second similarity, the third similarity, and the fourth similarity are weighted and averaged according to the corrected initial weights to determine a target similarity.
[0021] Preferably, before the scene type of the to-be-launched advertisement is obtained, the method further comprises:
[0022] Obtaining original video material to be parsed;
[0023] Converting the original video material to determine standard video data of a preset encoding format;
[0024] Using a shot cutting technique to decompose the standard video data into a plurality of video segments;
[0025] Using a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image in each video segment;
[0026] Compressing each video segment after embedding the blind watermark information to determine compressed video data;
[0027] Identifying the blind watermark information for each frame of the compressed image and outputting an identification result;
[0028] If the blind watermark information is identified, the video material information embedded in the blind watermark information is obtained as target video material information;
[0029] If the blind watermark information is not identified, feature extraction and matching are performed on each frame of the compressed image, and the target video material information is determined according to the matching result;
[0030] Inputting each frame of the compressed image into a pre-trained feature extraction model to output key feature information;
[0031] Inputting the key feature information into a multi-modal large language model to output the text information.
[0032] Preferably, the standard video data is decomposed into a plurality of video segments using a shot cutting technique, which comprises:
[0033] Converting each frame image in the standard video data into a gray scale image to obtain a gray scale image of each frame;
[0034] Performing edge detection on each frame of the gray scale image using an edge detection algorithm to output edge feature information;
[0035] Determining an edge feature change difference according to the edge feature information corresponding to adjacent frame gray scale images;
[0036] Determining video boundary information according to the edge feature change difference and a preset change threshold;
[0037] Decomposing the standard video data according to the video boundary information to determine each video segment.
[0038] Preferably, the embedding of the blind watermark information containing video material information into each frame image in each video segment using a preset feature processing algorithm comprises:
[0039] Converting the video material information to be embedded into a binary format to determine encoded blind watermark information;
[0040] Respectively decomposing each frame image in each video segment to obtain a region image;
[0041] Performing discrete cosine transform on the region image to convert the region image from a spatial domain image to a frequency domain image;
[0042] Adjusting the high-frequency component in the frequency domain image to embed the blind watermark information in the frequency domain image;
[0043] Recombining the frequency domain image after inverse discrete cosine transform to complete the embedding of the blind watermark information in each video segment.
[0044] Preferably, if the blind watermark information is not identified, feature extraction and matching are performed on each frame of the compressed image, and the target video material information is determined according to the matching result, which comprises:
[0045] Inputting each frame of the compressed image into a pre-trained self-supervised visual transformation model to output encoded feature information;
[0046] Using an approximate nearest neighbor algorithm to perform feature matching between the encoded feature information and feature template information of each video material template in a video material database to output a matching result;
[0047] According to the matching result, outputting the video material template corresponding to the feature template information matching the encoded feature information as the target video material information.
[0048] Preferably, the inputting each frame of compressed image into the pre-trained feature extraction model and outputting key feature information comprises:
[0049] Decoding the compressed video data to obtain audio data;
[0050] Inputting the compressed image into the pre-trained face recognition classification model to classify and label the face features in the recognized compressed image, and determining personnel information;
[0051] Inputting the compressed image into the pre-trained color analysis model to analyze the color distribution in the compressed image and extract main color information, wherein the main color information at least includes an advertisement brand tone detected in the compressed image or a main tone extracted in an advertisement landscape painting;
[0052] Inputting the compressed image into a target detection model to locate and classify objects in the compressed image and determine article information, wherein the article information at least includes article categories and article locations;
[0053] Inputting the audio data into a pre-trained speech transcription model to output audio feature information in the audio data;
[0054] Inputting the personnel information, color information, article information and audio feature information into a multi-modal large language model respectively to output the key feature information.
[0055] In a second aspect, the present application provides an advertisement video highlight generation device, the device comprising:
[0056] A scene type acquisition module is configured to acquire a scene type of an advertisement to be played, wherein the scene type includes an introduction scene, a product display scene, a user experience scene and a problem solving scene;
[0057] A similarity evaluation module is configured to perform similarity evaluation on text information related to the content of the video segment and a preset text template according to the scene type by using a preset text matching algorithm, to determine a target similarity corresponding to each video segment;
[0058] A video segment synthesis module is configured to synthesize video segments with a target similarity greater than a similarity threshold into a video highlight according to the target similarity and the preset similarity threshold.
[0059] In a third aspect, the present application further provides an electronic device, comprising at least one processor, at least one memory and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, the method of the first aspect in the above-mentioned embodiments is implemented.
[0060] In a fourth aspect, the embodiments of the present application further provide a storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method of the first aspect in the above-described embodiments.
[0061] In summary, the beneficial effects of the present application are as follows:
[0062] The advertisement video highlight generation method, device and equipment and storage medium provided by the present application, the method comprises: obtaining the scene type of the advertisement to be put, wherein the scene type comprises: introduction scene, product display scene, user experience scene and problem solving scene; according to the scene type, the text matching algorithm is used to evaluate the similarity between the text information related to the video segment content and the preset text template, and the target similarity corresponding to each video segment is determined; according to the target similarity and the preset similarity threshold, the video segment with the target similarity greater than the similarity threshold is combined into a video highlight. The present application firstly classifies the video of the advertisement to be put into scene types, and clearly divides it into different types such as introduction scene, product display scene, user experience scene and problem solving scene, so as to provide a more accurate framework for the generation of advertisement highlights; on this basis, the preset text matching algorithm is used to evaluate the similarity between the text information in the video segment and the preset scene template, and the similarity between each video segment and the preset template is calculated to ensure that the selected segment is highly related to a specific scene type; further, a similarity threshold is set, and when the similarity of a certain video segment is greater than the threshold, it is combined into the corresponding advertisement highlight. This screening method based on scene type and text similarity solves the problem that in the traditional advertisement video highlight generation, the advertisement content cannot be accurately screened and matched according to the specific scene type, so as to generate a video highlight that is more in line with the advertisement target and can accurately convey the advertisement information, and improve the pertinence and accuracy of the advertisement effect. BRIEF DESCRIPTION OF DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. For those skilled in the art, other drawings can also be obtained without creative labor on the premise of not paying the drawings, and these are within the protection scope of the present application.
[0064] Figure 1 The flowchart of the whole work of the advertisement video highlight generation method in embodiment 1 of the present application is shown in the figure.
[0065] Figure 2 The flowchart of the whole work of the advertisement video highlight generation method in embodiment 1 of the present application is shown in the figure.
[0066] Figure 3A flowchart of embedding blind watermark information containing video material information into each frame image of each video segment in embodiment 1 of the present application;
[0067] Figure 4 A flowchart of identifying and analyzing the compressed video data in embodiment 1 of the present application;
[0068] Figure 5 A flowchart of blind watermark information identification of each frame compressed image in the compressed video data in embodiment 1 of the present application;
[0069] Figure 6 A flowchart of feature extraction and matching of each frame compressed image in embodiment 1 of the present application;
[0070] Figure 7 A flowchart of inputting each frame compressed image into a pre-trained feature extraction model and outputting key feature information in embodiment 1 of the present application;
[0071] Figure 8 A structure block diagram of the advertisement video highlight generation device in embodiment 3 of the present application;
[0072] Figure 9 A structure diagram of the electronic device in embodiment 4 of the present application. DETAILED DESCRIPTION
[0073] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be noted that, in this document, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or sequence between these entities or operations. In the description of the present application, it should be understood that the orientations or positional relationships indicated by terms such as center, upper, lower, front, rear, left, right, vertical, horizontal, top, bottom, inner, outer and the like are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. Moreover, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the elements defined by the statement “include” do not exclude the presence of additional identical elements in the process, method, article or device including the elements. If there is no conflict, the embodiments of the present application and the various features in the embodiments can be combined with each other, and all within the scope of protection of the present application.
[0074] Embodiment 1
[0075] Please refer to Figure 1 Embodiment 1 of the present application discloses a method for generating an advertisement video highlight, the method comprising:
[0076] S1: obtaining original video materials to be analyzed;
[0077] Specifically, obtaining original video materials to be analyzed involves collecting and importing all original video files that need to be analyzed and processed. These materials come from different camera equipment, recording environments or external resources. This process ensures that all original video materials to be processed are concentrated in a manageable storage location, facilitating subsequent transcoding, segmentation and analysis operations. By collecting and managing these original video materials, high-quality input data can be provided for each subsequent processing step, laying the foundation for the entire video processing and analysis process.
[0078] S2: transcoding the original video materials to determine standard video data in a preset encoding format;
[0079] Specifically, the original video material is transcoded to ensure video file format consistency and data security. When a user uploads the original video material, a uniform transcoding process is performed on the original video material to convert various formats of video into a preset standard encoding format, such as H.264 or H.265 encoding. The uniform resolution and frame rate are set, and encryption and error detection mechanisms are added during the encoding process to improve the security of video data. This not only realizes the uniform management of different formats of video and simplifies the subsequent processing steps, but also reduces the storage space requirement by optimizing the encoding format and compression parameters, while ensuring that the video quality is not significantly affected. The transcoding process also adds encryption and error detection mechanisms to further improve the security of video data, preventing data from being damaged or unauthorized access during transmission and storage. Through this process, video material can be more effectively managed to ensure compatibility and security of all data in subsequent processing.
[0080] S3: using shot segmentation technology to divide the standard video data into multiple video segments;
[0081] Specifically, the standard video data is divided into multiple video segments using shot segmentation technology. First, the shot switching points are identified, which are usually determined by significant changes in picture content, scene changes or significant changes in audio. To ensure the accuracy of the segmentation, advanced algorithms such as edge detection and motion analysis are used, and when the edge change amplitude of the picture exceeds the preset threshold, it is identified as a shot switching point. Then, the standard video data is segmented at these switching points to generate multiple continuous and coherent video segments, each of which retains important scene information in the original video, facilitating subsequent processing and analysis.
[0082] In an embodiment, referring to Figure 2 , the S3 comprises:
[0083] S31: performing gray scale conversion on each frame of image in the standard video data to obtain a gray scale image of each frame;
[0084] Specifically, first, each frame of image in the standard video data is converted to a gray scale image, i.e. a color image is converted to a gray scale image. Gray scale conversion is to convert the RGB (red, green, blue) color information of the image to a gray scale value, which is calculated by weighting the color value of each pixel. Specifically, the commonly used conversion formula is: gray scale value = 0.299 * R + 0.587 * G + 0.114 * B. This processing can simplify the image data, reduce the computational complexity, and highlight the structure and shape features in the image, which is convenient for subsequent edge detection.
[0085] S32: performing edge detection on each frame of the grayscale images using an edge detection algorithm to output edge feature information;
[0086] Specifically, next, an edge detection algorithm such as Sobel, Canny, Laplacian, etc. is applied to process each frame of grayscale images to detect the edge information in the images. The purpose of edge detection is to find regions in the image where the pixel values change significantly, which usually correspond to the boundaries and contours of objects. The edge detection algorithm identifies pixels with dramatic changes in grayscale values in the image by calculating image gradients or using difference operators, and marks these pixels as edge points, thereby generating a binary image containing edge feature information.
[0087] S33: determining an edge feature change difference value based on the edge feature information corresponding to adjacent frames of grayscale images;
[0088] Specifically, after edge detection is completed, the edge feature information of adjacent frames is compared to determine the change of edge features. The specific operation is to calculate the edge image difference value between adjacent frames, and identify the edge change area in the image through pixel-level difference operation. The edge feature change difference value is obtained by calculating the difference of edge pixel values at corresponding positions of adjacent frames, and these differences can be quantified as change amplitudes, reflecting the dynamic change of scene content over time.
[0089] S34: determining video boundary information based on the edge feature change difference value and a preset change threshold;
[0090] Specifically, according to the calculated edge feature change difference value, it is compared with the preset change threshold. If the edge feature change difference value between a frame and its adjacent frame exceeds the preset threshold, it is considered that there is a significant scene cut between the frame and the previous frame, i.e. it is determined as a video boundary. The preset change threshold is used to filter out subtle and insignificant changes, ensuring that only obvious content changes (such as scene cuts) are identified as video boundaries. In this way, the boundary information of each shot in the video is accurately determined.
[0091] S35: decomposing the standard video data based on the video boundary information to determine each video segment.
[0092] Specifically, finally, according to the determined video boundary information, the standard video data is segmented at these boundary points to generate multiple independent video segments, each of which represents a continuous scene or shot, preserving the coherence and integrity of the content. The decomposed video segments facilitate subsequent processing and analysis, such as embedding blind watermarks, feature extraction, content labeling, etc. Through this method, video data can be effectively managed and processed, improving the efficiency of video content analysis and processing.
[0093] Specifically, from grayscale conversion to edge detection, to the calculation of edge feature changes and the determination of video boundaries, the video is finally decomposed into multiple segments. This process realizes the fine processing and management of video content, laying a solid foundation for subsequent deep analysis and application.
[0094] S4: using a preset feature processing algorithm, embedding blind watermark information containing video material information into each frame image in each video segment;
[0095] Specifically, using a preset feature processing algorithm, embedding blind watermark information containing video material information into each frame image in each video segment, in order to ensure that the original video material can be accurately identified and associated in subsequent film editing. The specific operation includes: first, select a suitable feature processing algorithm, such as discrete cosine transform (DCT) or discrete wavelet transform (DWT), to convert each frame image into the frequency domain. Then, select specific frequency coefficients (usually low or medium frequency coefficients) in the frequency domain for slight adjustment to embed blind watermark information. These embedded blind watermark information can contain unique material identifiers and other key data, ensuring that the image is slightly adjusted without affecting the video visual quality, so that the watermark information is invisible to the human eye, but can be extracted through a specific algorithm in the future. Finally, use the inverse transform to restore the image to the spatial domain to generate a frame image containing blind watermark. In this way, each frame of each video segment carries unique identification information, achieving accurate material tracking and identification, effectively solving the problem of positioning video materials in the film.
[0096] In an embodiment, referring to Figure 3 , the S4 includes:
[0097] S41: converting the video material information to be embedded into a binary format to determine the encoded blind watermark information;
[0098] Specifically, first, the video material information to be embedded (such as unique identifier, copyright information, etc.) is converted into binary format, and text or other forms of information are converted into binary data stream through encoding algorithms, such as ASCII encoding or UTF-8 encoding. After the information is converted into binary format, specific encryption or hash processing is performed to generate the final blind watermark information. The purpose of this processing is to ensure the stability and security of the information during the embedding process, facilitating subsequent embedding and extraction in the frequency domain image.
[0099] S42: Each frame image in each video segment is processed for decomposition to obtain regional images;
[0100] Specifically, next, each frame image in each video segment is processed for decomposition, which typically involves dividing the image into small blocks (such as 8x8 or 16x16 macroblocks) to facilitate more granular frequency domain processing. Decomposition allows each small block to be processed independently, ensuring that the embedding process has minimal impact on the overall image. These decomposed regional images can better analyze and process local features, which is beneficial to improve the embedding efficiency and concealment of blind watermark.
[0101] S43: Discrete cosine transform is applied to the regional images to convert them from spatial domain images to frequency domain images;
[0102] Specifically, after obtaining the regional images, discrete cosine transform (DCT) is applied to each regional image to convert it from the spatial domain to the frequency domain. DCT decomposes image data into different frequency components, allowing the energy of the image to concentrate in a few low-frequency components, which often correspond to the main features of the image. Through this transformation, blind watermark information can be more effectively embedded in the frequency domain without significantly affecting the visual quality of the original image.
[0103] S44: High-frequency components in the frequency domain image are adjusted to embed blind watermark information in the frequency domain image;
[0104] Specifically, high-frequency components are selected in the frequency domain image for adjustment to embed blind watermark information. The specific method is to adjust the selected DCT coefficients slightly according to the binary blind watermark information, such as increasing or decreasing the value of a specific coefficient to represent binary 0 or 1. Low-frequency components contain the main information of the image, and slight adjustments will not significantly affect the image quality, while medium-frequency components balance the concealment and robustness of information embedding. In this way, blind watermark information is securely and covertly embedded in the frequency domain features of the image.
[0105] S45: After inverse discrete cosine transform is applied to the frequency domain image, it is recombined to complete the embedding of blind watermark information in each video segment.
[0106] Specifically, finally, the adjusted frequency domain image is subjected to inverse discrete cosine transform (IDCT) to convert it back to a spatial domain image, restoring the visual effect of the original image. After IDCT processing, each region image is recombined back into the original frame image, ensuring that each frame retains the blind watermark information completely. In this way, the blind watermark information is seamlessly embedded in each frame image in the video segment, and under the premise of not affecting the overall visual quality, ensures the stable embedding and subsequent extractability of the information. Each video segment processed in this way carries unique identification information, achieving precise tracking and protection of the material.
[0107] S5: compressing each video segment after embedding the blind watermark information to determine compressed video data;
[0108] Specifically, each video segment after embedding the blind watermark information is compressed to determine compressed video data, the original piece is stored in the company's intranet environment, and a low-resolution version of the video is compressed and stored in the cloud for tagging and analysis. Video compression technology significantly reduces the size of video files by optimizing encoding and reducing resolution, thereby reducing storage and transmission costs. Storing the original piece in the intranet ensures the high quality and data security of the video, preventing unauthorized access. The low-resolution version of the video is sufficient to meet the needs of content recognition and labeling, and its smaller file size makes cloud storage and access more efficient. Storing low-resolution videos in the cloud enables remote access and sharing, making it easier for team members to collaborate and process data from different locations, improving work efficiency and flexibility while ensuring data security and consistency. This scheme takes into account the quality, security and ease of use of video data through efficient compression and distributed storage, providing a flexible working environment and strong technical support for the team.
[0109] S6: using a pre-set feature extraction algorithm to identify and analyze the compressed video data, outputting text information related to the video content and target video material information.
[0110] Specifically, the compressed video data is identified and analyzed using a pre-set feature extraction algorithm, outputting text information and target video material information related to the video content. This process first analyzes each frame of the compressed video using advanced feature extraction algorithms such as multi-modal large language models, speech transcription models, face recognition models, object detection models, and color analysis models. These algorithms can identify various content in the video, such as text, speech, faces, objects, and colors. Then, the identified information is integrated and analyzed to generate detailed text descriptions and annotations, including scene descriptions, dialogue content, character information, and object locations. These information can not only be used for video content understanding and retrieval, but also help generate accurate identification and description of video material, ensuring accurate association and utilization of these materials in subsequent use. This process greatly improves the efficiency and accuracy of video content analysis, providing strong technical support for content management and application.
[0111] In an embodiment, referring to Figure 4 , the S6 comprises:
[0112] S61: Blind watermark information identification is performed on each frame of the compressed image, and an identification result is output.
[0113] Specifically, the pre-set blind watermark embedding algorithm corresponding to the extraction algorithm is used to process each frame of image in the compressed video data, and the embedded blind watermark information is extracted. According to the extraction result, the target video material information is determined. This step ensures that the source and specific content of the video segment can be accurately located and identified even in the case of compressed video, thereby ensuring the traceability of the material and the effectiveness of the content management.
[0114] In an embodiment, referring to Figure 5 , the S61 comprises:
[0115] S611: Input each frame of the compressed image into a pre-trained self-supervised visual transformation model, and output encoded feature information.
[0116] Specifically, first, each frame of compressed image is input into a pre-trained self-supervised visual transformation model (such as DINO-V2). The model learns rich visual features from a large amount of unlabeled data through self-supervised learning. The model processes the input image, extracts the key features of the image, and encodes these features into high-dimensional feature vectors. These feature vectors contain various information such as texture, edge, shape, and color of the image, which can accurately describe the image content. The output encoded feature information will serve as the basis for subsequent feature matching.
[0117] S612: Perform feature matching on the encoded feature information with the feature template information of each video material template in the video material database using the approximate nearest neighbor algorithm, and output the matching result;
[0118] Specifically, then, the encoded feature information is matched using the approximate nearest neighbor algorithm (ANN). The approximate nearest neighbor algorithm can quickly find the most similar feature template to the target feature in a high-dimensional feature space. The encoded feature information of each frame of image is compared with the feature template information of each video material template pre-stored in the video material database, and the similarity is calculated. Through the ANN algorithm, the feature template closest to the encoded feature information is found, and the matching result is output, which includes the matching degree and related information of each frame of image with the most similar material template in the database.
[0119] S613: According to the matching result, output the video material template corresponding to the feature template information matched with the encoded feature information as the target video material information.
[0120] Specifically, finally, according to the feature matching result, the target video material information is determined. For each frame of image, the feature template information with the highest matching degree is selected, and the corresponding video material template is output as the target video material information, ensuring that even without blind watermark information, the source and specific content of the video segment can be accurately identified through feature matching. In this way, each frame of image can be efficiently and accurately associated with the original video material, realizing accurate identification and management of video content.
[0121] S62: If the blind watermark information is identified, the video material information embedded in the blind watermark information is obtained as the target video material information;
[0122] Specifically, after successfully identifying the blind watermark information, the embedded video material information is extracted from the identified watermark data. The blind watermark usually contains a unique material identifier, copyright information or other related metadata. Through these information, the source and specific content of the video material can be accurately determined, and it is used as the target video material information, ensuring that each frame of image can be accurately associated with the corresponding original material.
[0123] S63: If the blind watermark information is not identified, perform feature extraction and matching on each frame of the compressed image, and determine the target video material information according to the matching result;
[0124] Specifically, if the blind watermark information fails to be identified, a feature extraction and matching method is used to further analyze the compressed image. The extracted features include texture, edge, shape, and color information, forming a unified image feature code. Then, the extracted image features are matched with the material features in the database to determine the most similar material, ensuring that even without blind watermarking, the source and content of the video segment can be accurately identified. According to the matching result, the target video material information is determined and output, thereby realizing accurate identification and management of materials.
[0125] S64: input each frame of compressed image into the pre-trained feature extraction model, and output key feature information;
[0126] Specifically, each frame of compressed image is input into the pre-trained feature extraction model, and key feature information is output. This process involves passing each frame of image through a pre-trained feature extraction model in the field of computer vision, such as a convolutional neural network (CNN) or a Transformer model. This model has been trained on a large amount of image data and can automatically extract important visual features such as edges, textures, color distributions, and shapes from images. After inputting the image, the model will extract high-level features from the image layer by layer through a series of convolutional layers, activation functions, and pooling operations. These features will be encoded into a high-dimensional feature vector, which contains the core information of the image. These encoded key feature information not only accurately describes the image content, but also can be used for subsequent feature matching, classification, or retrieval tasks, ensuring effective utilization and analysis of image data.
[0127] In an embodiment, please refer to Figure 6 , the S64 comprises:
[0128] S641: decoding the compressed video data to obtain audio data;
[0129] Specifically, first, the compressed video data is decoded to extract the audio data. This process includes restoring the compressed video file to its original uncompressed state and extracting the audio portion of the video stream through a decoder. The decoding process usually involves converting the compressed encoded data stream (such as MP4 or H.264 format) into an uncompressed audio format (such as WAV or PCM) that can be further processed. The extracted audio data contains all the sound information in the video, including dialogues, background music, and environmental sounds, providing a basis for subsequent audio analysis and processing.
[0130] S642: input the compressed image into the pre-trained face recognition classification model, classify and label the recognized face features in the compressed image, and determine the personnel information;
[0131] S643: input the compressed image into a pre-trained color analysis model to analyze the color distribution in the compressed image and extract primary color information, wherein the primary color information at least includes an advertising brand hue detected in the compressed image or a dominant hue extracted in an advertising landscape painting;
[0132] S644: input the compressed image into a target detection model to locate and classify objects in the compressed image and determine item information, wherein the item information at least includes an item category and an item location;
[0133] Specifically, the compressed image is input into a pre-trained face recognition classification model, a color analysis model, and a target detection model respectively to output personnel information, color information, and item information. Specifically, this step involves passing each frame of compressed image into multiple specially trained deep learning models to extract and analyze different types of information. First, the input image is passed to a face recognition classification model, which is trained using a large amount of pre-labeled facial data, capable of accurately recognizing facial features in the image and classifying and labeling, for example, identifying celebrities, actors, or other known figures appearing in the image, such as identifying the facial features of a reporter in a news report. Second, the image is input into a color analysis model, which analyzes the color distribution in the image and extracts primary color information, such as brand hues detected in advertising images or dominant hues extracted in landscape paintings. Finally, the image is input into a target detection model, which locates and classifies objects in the image, such as identifying goods in a shopping advertisement or traffic signs in a video, and determines their category and location. These models each process different aspects of the image, and by integrating this information, a comprehensive description of the image content can be obtained, such as identifying actors, background colors, and main objects in the scene in a movie clip, providing key details for in-depth understanding and analysis of the video.
[0134] S645: input the audio data into a pre-trained speech transcription model to output audio feature information in the audio data;
[0135] Specifically, the extracted audio data is input into a pre-trained speech transcription model to convert the speech content in the audio into text information. The speech transcription model uses natural language processing techniques and speech recognition algorithms to recognize language, intonation, and syllables in the audio and transcribe them into text. This process can identify and record conversations, speeches, or other voice information and generate corresponding text descriptions. Audio feature information includes the text form of speech content, so that the information in the audio can be further analyzed and processed.
[0136] S646: input the personnel information, color information, item information, and audio feature information into a multi-modal large language model respectively to output the key feature information.
[0137] Specifically, the personnel information, color information, item information extracted from the image, and the audio feature information extracted from the audio are input into the multi-modal large language model. The multi-modal large language model can fuse data from different sources, understand and analyze the relationship between various types of information. The model will comprehensively analyze these multi-modal data and generate key feature information describing the video content. This includes integrating visual features in images and speech content in audio to provide a comprehensive content understanding and summary. This fusion analysis method enables the effective integration and interpretation of various information extracted from the video, generating deep insights into the video content.
[0138] S65: inputting the key feature information into the multi-modal large language model and outputting the text information.
[0139] Specifically, inputting the key feature information into the multi-modal large language model and outputting the text information first involves integrating the key feature information extracted from images (such as personnel information, color information, item information) and audio (such as speech transcription text) into a unified data format. These feature information will be input into the multi-modal large language model, which uses deep learning technology to fuse data from different modalities (vision, audio, text, etc.), understand the relationship between these information. Through analysis and context modeling of the input multi-modal data, the model can generate a coherent text description that accurately expresses the key points of the video content. The output text information usually includes a summary of the video scene, detailed descriptions of the dialogue content, characters and objects, and other related information. This comprehensive analysis capability enables the multi-modal large language model to provide a comprehensive and accurate summary of the video content, enhancing the understanding and utilization of video data.
[0140] Embodiment 2
[0141] In the above embodiment 1, the analysis results related to the content of each video segment are obtained through content analysis, including personnel information, color information, item information and audio feature information. In the real advertising video editing scene, the relevance of the video segments is evaluated according to the content analysis results, and the video highlights are edited accordingly, which can greatly improve the quality and viewing experience of the video. For this purpose, in an implementation, please refer to Figure 7 , after the S6, it further includes:
[0142] Obtaining the scene type of the to-be-played advertisement, wherein the scene type includes an introduction scene, a product display scene, a user experience scene, and a problem solving scene.
[0143] Specifically, the scene type of the advertisement to be launched is obtained. Advertisements are generally divided into multiple scene types, such as introduction scenes, product display scenes, user experience scenes, and problem solving scenes. The introduction scene is usually used to attract the attention of the audience and present the core information or brand of the advertisement. The product display scene focuses on displaying the function, appearance, and use method of the product to help the audience intuitively understand the value of the product. The user experience scene shows how the user uses the product and the experience and feedback obtained in the use process. The problem solving scene shows how the product helps the user solve practical problems and emphasizes the effectiveness and practicality of the product. The step of obtaining the scene type can be completed by analyzing the advertising planning content, brand requirements, and market demand. The advertisement editing team usually presets these scene types to ensure that the advertisement can be displayed according to the set plot or marketing strategy. At the same time, video analysis technology is used to extract key features (such as products, characters, and scene elements) in the video, and these features are matched with the predefined scene types to quickly determine the scene category to which the video segment belongs.
[0144] According to the scene type, the text matching algorithm is used to evaluate the similarity between the text information related to the content of the video segment and the preset text template to determine the target similarity corresponding to each video segment.
[0145] Specifically, after confirming the scene type, the video segment is further analyzed according to the scene type. The similarity between the segment and the target scene is determined by a text matching algorithm. First, the content of the video segment is usually converted into text form through a speech transcription model or subtitle extraction technology. These text information reflects the core content of the video segment, such as product name, function description, and user experience. The text matching algorithm, such as TF-IDF or BERT, is used to calculate the similarity between the text information in the video segment and the preset scene text template. Each scene type has a set of preset text templates, and the text template contains keywords or sentence patterns related to the scene. The core of the similarity evaluation is to calculate the semantic similarity between the text in the video segment and the template text, and then evaluate the matching degree of the video segment and the scene type.
[0146] In an embodiment, the step of evaluating the similarity between the text information related to the content of the video segment and the preset text template according to the scene type using the preset text matching algorithm to determine the target similarity corresponding to each video segment includes:
[0147] The preset text matching algorithm is used to evaluate the similarity between the text information related to the content of the video segment and the preset text template to determine the first similarity, the second similarity, the third similarity, and the fourth similarity corresponding to the personnel information, the color information, the article information, and the audio feature information, respectively.
[0148] Specifically, first, the text matching algorithm is used to process the text information related to the video segment, including personnel information, color information, item information and audio feature information. These text information is evaluated for similarity with preset text templates, which usually represent specific scene types, such as introduction scenes, product display scenes, etc., containing keywords and related sentences. The semantic similarity between the video text information and the template is calculated by algorithms such as TF-IDF or BERT, and the first similarity, the second similarity, the third similarity and the fourth similarity corresponding to the personnel information, the color information, the item information and the audio feature information are determined respectively. These similarities reflect the matching degree of each video segment with the relevant features in the scene.
[0149] A preset weight correction factor is obtained, wherein the weight correction factor is greater than 1;
[0150] Specifically, in order to adjust the evaluation of different similarities, a preset weight correction factor is obtained, which is usually greater than 1, indicating that some similarities need to be amplified. The setting of the weight correction factor depends on the specific scene requirements. For example, in some scenes, color information may be more important than other features, so the weight of color similarity needs to be increased. The weight correction factor is used to dynamically adjust the importance of different similarities, so that the final result can better meet the scene requirements and advertising planning goals.
[0151] The first similarity, the second similarity, the third similarity and the fourth similarity respectively correspond to each initial weight, wherein the sum of each initial weight is equal to 1;
[0152] Specifically, after similarity calculation, each similarity is assigned an initial weight, which determines the influence of each similarity in the total similarity calculation, and the sum of all initial weights is equal to 1. According to the advertising planning goal, the initial weight can be preset to different values, for example, the product display scene may give higher weight to item information, while the user experience scene may give higher weight to personnel information.
[0153] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are modified using the weight correction factor;
[0154] Specifically, in the introduction scene, the color information and audio feature information of the video are usually the key factors to attract the audience's attention, so the weight correction factor is used to correct the similarities related to these features, i.e., the second similarity and the fourth similarity. The correction process includes multiplying the weight correction factor by the second similarity to obtain a new second similarity, and multiplying the weight correction factor by the fourth similarity to obtain a new fourth similarity. This means that the initial weights of color and audio are amplified, so that these features occupy a larger proportion in the similarity evaluation, ensuring that the visual and auditory effects of the introduction scene are more attractive.
[0155] If the scene type is a product display scene, the weight correction factor is used to correct the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity.
[0156] Specifically, in the product display scene, the color information and item information in the video are particularly important, so the weight correction factor is used to correct the second similarity and the third similarity. This adjustment ensures that the visual effects and item details of the product during display are more prominent, which can attract the audience's attention and make them focus on the characteristics of the product.
[0157] If the scene type is a user experience scene, the weight correction factor is used to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity.
[0158] Specifically, the user experience scene focuses on showing the interaction between people and products, so the weight correction factor is used to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity. In this type of scene, the weights of these two types of information are amplified to ensure that the interaction process between the user and the product and the display of the product itself are more prominent, enhancing the user's sense of identification.
[0159] If the scene type is a problem solving scene, the weight correction factor is used to correct the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity.
[0160] Specifically, in the problem solving scene, it is usually necessary to highlight personnel information, item information, and audio feature information to show how the product solves the user's actual problem. Therefore, the weight correction factor is used to correct the first similarity, the third similarity, and the fourth similarity, which can better present how the product solves the user's troubles through interaction with the user, the process of solving the problem, and related audio prompts, enhancing the persuasiveness and authenticity of the scene.
[0161] According to the modified initial weights, the first similarity, the second similarity, the third similarity and the fourth similarity are weighted and averaged to determine a target similarity;
[0162] Specifically, the first similarity, the second similarity, the third similarity and the fourth similarity obtained before are processed, each of which represents the matching degree of different types of information. In the previous steps, each similarity is modified according to the scene requirements to ensure that the content valued by the scene is highlighted. The weighted average processing means that the modified initial weights are multiplied by the corresponding similarity values respectively, and then the results are added and averaged to calculate the overall target similarity. The target similarity represents the matching degree of the video clip after integrating multi-modal information with a specific scene. Through this process, the system can comprehensively evaluate the fit degree of each video clip with the preset scene under multi-modal information.
[0163] According to the target similarity and a preset similarity threshold, video clips with a target similarity greater than the similarity threshold are combined into a video highlight.
[0164] Specifically, after obtaining the target similarity, the value is compared with a preset similarity threshold. The similarity threshold is a standard set by the system in advance to determine whether a video clip meets the minimum requirements of the scene. This threshold can be adjusted according to the requirements of different scenes. For example, a product demonstration scene may require a higher similarity threshold, while an introduction scene may allow a lower threshold. The similarity threshold is usually a standard set based on user expectations or advertising goals. If the target similarity is greater than or equal to the threshold, it indicates that the video clip has high relevance and can well match the requirements of the scene, making it suitable for the next step. Those below the threshold are considered not to match the current scene and are excluded from the selection of the video highlight. After determining which video clips meet the similarity requirements, the system will perform synthesis processing on these clips. The synthesis process includes combining multiple high-similarity video clips according to certain logic or strategies, such as time sequence, plot development, visual consistency, etc., to form a complete advertising highlight. This step can be adjusted according to the design of the marketing strategy, such as combining product demonstration clips with user experience clips or arranging problem-solving clips and product demonstration clips in order to ensure that the highlight has a coherent narrative structure.
[0165] Embodiment 3
[0166] See Figure 8 Embodiment 3 of the present application also provides an advertising video highlight generation device, which comprises:
[0167] A video material acquisition module is configured to acquire original video material to be analyzed.
[0168] a transcoding processing module, configured to perform transcoding processing on the original video material to determine standard video data in a preset encoding format;
[0169] a shot segmentation module, configured to use a shot segmentation technique to decompose the standard video data into a plurality of video clips;
[0170] a blind watermark embedding module, configured to use a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image in each video clip;
[0171] a compression processing module, configured to perform compression processing on each video clip after embedding the blind watermark information to determine compressed video data;
[0172] a feature extraction module, configured to use a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information.
[0173] Specifically, the advertisement video highlight generation device provided by the embodiment of the application comprises: a video material acquisition module configured to acquire original video material to be analyzed; a transcoding processing module configured to perform transcoding processing on the original video material to determine standard video data in a preset encoding format; a shot segmentation module configured to use a shot segmentation technique to decompose the standard video data into a plurality of video clips; a blind watermark embedding module configured to use a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image in each video clip; a compression processing module configured to perform compression processing on each video clip after embedding the blind watermark information to determine compressed video data; and a feature extraction module configured to use a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information. Through a series of ordered processing steps, the device realizes accurate association of clips and material, and accurate identification and analysis of video content. First, the original video material is acquired and transcoded to ensure that all data formats are uniform. Then, the standard video data is decomposed into a plurality of easily managed clips using a shot segmentation technique, and blind watermark information is embedded to enable accurate identification of each material clip in the finished product in the subsequent process, even after editing and compression processing. Finally, the compressed video is identified and analyzed through a feature extraction algorithm to output detailed text information related to the video content and target video material information. This method not only improves the matching accuracy of video material and finished product, ensures that each piece of material can be accurately located and identified, but also efficiently extracts and summarizes video content, provides more detailed analysis and annotation, and solves the problems of efficiency and accuracy in traditional video processing and analysis.
[0174] Embodiment 4
[0175] In addition, in combination with Figure 1 The advertisement video highlight generation method of the embodiment 1 of the present application can be implemented by an electronic device. Figure 9 A hardware structure schematic diagram of the electronic device provided by the embodiment 4 of the present application is shown.
[0176] The electronic device can include a processor and a memory having computer program instructions stored therein.
[0177] Specifically, the processor can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits of the embodiments of the present application.
[0178] The memory can include a mass storage for data or instructions. By way of example and not limitation, the memory can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory can include removable or non-removable (or fixed) media. Where appropriate, the memory can be internal or external to the data processing apparatus. In certain embodiments, the memory is non-volatile solid-state memory. In certain embodiments, the memory includes read-only memory (ROM). Where appropriate, this ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0179] The processor reads and executes the computer program instructions stored in the memory to implement any one of the advertisement video highlight generation methods in the above embodiments.
[0180] In one example, the electronic device can further include a communication interface and a bus. Wherein, as Figure 9 The processor, the memory, the communication interface are connected through the bus and complete the communication among each other.
[0181] The communication interface is mainly used to realize the communication between each module, device, unit and / or equipment in the embodiments of the present application.
[0182] The bus includes hardware, software, or both, to couple components of the device to each other and to couple components to other devices. While the application is not limited to a particular bus structure, examples of busses include an Accelerated Graphics Port (AGP) or other graphics bus, a Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, the bus can include one or more buses. Although the application is described and shown with respect to a particular bus, the application contemplates any suitable bus or interconnect.
[0183] Embodiment 5
[0184] In addition, in combination with the advertisement video highlight generation method in Embodiment 1 described above, Embodiment 5 of the application can also provide a computer readable storage medium for implementation. The computer readable storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any one of the advertisement video highlight generation methods in the above embodiments.
[0185] In summary, the embodiments of the application provide an advertisement video highlight generation method, device, equipment and storage medium.
[0186] It needs to be clear that the application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the application.
[0187] The functional blocks shown in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, functional cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium that can store or transfer information. Examples of the machine-readable medium include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, and the like. The code segments can be downloaded via a computer network such as the Internet, an intranet, and the like.
[0188] The user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant place, and provide corresponding operation portal for the user to choose authorization or rejection.
[0189] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps are executed simultaneously.
[0190] The above description is only a specific implementation of the present application, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, module and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here. It should be understood that the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. An advertisement video highlight generation method characterized by comprising: The method comprises: acquiring a scene type of an advertisement to be launched, wherein the scene type comprises an introduction scene, a product display scene, a user experience scene, and a problem solving scene; according to the scene type, using a preset text matching algorithm to perform similarity evaluation on text information related to the content of the video segment and a preset text template to determine a target similarity corresponding to each video segment; according to the target similarity and a preset similarity threshold, combining video segments with a target similarity greater than the similarity threshold into a video highlight; wherein the similarity evaluation according to the scene type and using the preset text matching algorithm on the text information related to the content of the video segment and the preset text template to determine the target similarity corresponding to each video segment comprises: the content of the video segment is converted into a text form through a speech transcription model or a subtitle extraction technology, and the text information reflects the core content of the video segment, including product names, function descriptions, and user experiences; using a text matching algorithm, the text matching algorithm comprises TF-IDF and BERT, to perform similarity calculation on the text information in the video segment and a preset scene text template, wherein each scene type has a set of preset text templates, and the text templates contain keywords or sentence patterns related to the scene; the core of the similarity evaluation is to calculate the semantic similarity of the text in the video segment and the template text, and then evaluate the matching degree of the video segment and the scene type; the similarity evaluation according to the scene type and using the preset text matching algorithm on the text information related to the content of the video segment and the preset text template to determine the target similarity corresponding to each video segment comprises: using a preset text matching algorithm to perform similarity evaluation on text information related to the content of the video segment and a preset text template to determine a target similarity corresponding to each video segment; acquiring a preset weight correction factor, wherein the weight correction factor is greater than 1; acquiring initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1; if the scene type is an introduction scene, using the weight correction factor to correct the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity; if the scene type is a product display scene, using the weight correction factor to correct the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity; if the scene type is a user experience scene, using the weight correction factor to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity; if the scene type is a problem solving scene, using the weight correction factor to correct the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity; The first similarity, the second similarity, the third similarity and the fourth similarity are weighted and averaged according to the modified initial weights to determine the target similarity.
2. The advertisement video highlight generation method of claim 1, wherein, The method further comprises, before the acquiring of the scene type of the advertisement to be launched, the following steps: Acquiring original video material to be parsed; Performing transcoding processing on the original video material to determine standard video data of a preset encoding format; Using a shot segmentation technique to decompose the standard video data into a plurality of video segments; Embedding blind watermark information containing video material information into each frame image in each video segment using a preset feature processing algorithm; Performing compression processing on each video segment after the blind watermark information is embedded to determine compressed video data; Performing blind watermark information identification on each frame compressed image in the compressed video data to output an identification result; If the blind watermark information is identified, video material information embedded in the blind watermark information is acquired as target video material information; If the blind watermark information is not identified, feature extraction and matching are performed on each frame compressed image, and the target video material information is determined according to a matching result; Inputting each frame compressed image into a pre-trained feature extraction model to output key feature information; Inputting the key feature information into a multi-modal large language model to output the text information.
3. The advertisement video highlight generation method of claim 2, wherein, The method further comprises, before the acquiring of the scene type of the advertisement to be launched, the following steps: Performing grayscale conversion on each frame image in the standard video data to acquire each frame grayscale image; Using an edge detection algorithm to perform edge detection on each frame grayscale image to output edge feature information; Determining an edge feature change difference value according to edge feature information corresponding to adjacent frame grayscale images; Determining video boundary information according to the edge feature change difference value and a preset change threshold; According to the video boundary information, the standard video data is decomposed to determine each video segment.
4. The advertisement video highlight generation method of claim 2, wherein, The method further comprises, before the acquiring of the scene type of the advertisement to be launched, the following steps: Converting the video material information to be embedded into a binary format to determine encoded blind watermark information; Performing decomposition processing on each frame image in each video segment to acquire a region image; Performing discrete cosine transformation on the region image to convert the region image from a spatial domain image to a frequency domain image; Adjusting a high-frequency component in the frequency domain image to embed the blind watermark information in the frequency domain image; After inverse discrete cosine transformation is performed on the frequency domain image, the frequency domain image is recombined to complete embedding of the blind watermark information in each video segment.
5. The advertisement video highlight generation method of claim 3, wherein, The method further comprises, before the acquiring of the scene type of the advertisement to be launched, the following steps: Inputting each frame compressed image into a pre-trained self-supervised visual transformation model to output encoded feature information; Using an approximate nearest neighbor algorithm to perform feature matching on the encoded feature information and feature template information of each video material template in a video material database to output a matching result; According to the matching result, the video material template corresponding to the feature template information matched with the coded feature information is output as the target video material information.
6. The advertisement video highlight generation method of claim 2, wherein, The inputting the compressed image into the pre-trained feature extraction model and outputting key feature information includes: Decoding the compressed video data to obtain audio data; Inputting the compressed image into a pre-trained face recognition classification model to classify and label the face features in the recognized compressed image to determine personnel information; Inputting the compressed image into a pre-trained color analysis model to analyze the color distribution in the compressed image and extract main color information, wherein the main color information at least includes an advertisement brand tone detected in the compressed image or a main tone extracted in an advertisement landscape painting; Inputting the compressed image into a target detection model to locate and classify objects in the compressed image to determine item information, wherein the item information at least includes an item category and an item location; Inputting the audio data into a pre-trained speech transcription model to output audio feature information in the audio data; Inputting the personnel information, color information, item information, and audio feature information into a multi-modal large language model respectively to output the key feature information.
7. An advertisement video highlight generation apparatus characterized by comprising: The device comprises: A scene type acquisition module configured to acquire a scene type of an advertisement to be launched, wherein the scene type includes an introduction scene, a product display scene, a user experience scene, and a problem solving scene; A similarity evaluation module configured to, according to the scene type, utilize a preset text matching algorithm to evaluate the similarity between text information related to the content of a video segment and a preset text template, and determine a target similarity corresponding to each video segment; A video segment synthesis module configured to, according to the target similarity and a preset similarity threshold, synthesize video segments with a target similarity greater than the similarity threshold into a video highlight; According to the scene type, utilizing a preset text matching algorithm to evaluate the similarity between text information related to the content of a video segment and a preset text template, and determining a target similarity corresponding to each video segment includes: the content of a video segment is converted into a text form through a speech transcription model or a subtitle extraction technology, and these text information reflects the core content of the video segment, including product names, function descriptions, and user experiences; utilizing a text matching algorithm, the text matching algorithm includes TF-IDF and BERT, to calculate the similarity between the text information in the video segment and a preset scene text template, wherein each scene type has a set of preset text templates, and the text templates contain keywords or sentence patterns related to the scene; the core of the similarity evaluation is to calculate the semantic similarity between the text in the video segment and the template text, and then to evaluate the matching degree of the video segment and the scene type; According to the scene type, utilizing a preset text matching algorithm to evaluate the similarity between text information related to the content of a video segment and a preset text template, and determining a target similarity corresponding to each video segment includes: The preset text matching algorithm is used to perform similarity evaluation on the text information related to the video segment content and the preset text template, to determine a first similarity, a second similarity, a third similarity and a fourth similarity corresponding to the personnel information, the color information, the article information and the audio feature information respectively. A preset weight correction factor is obtained, wherein the weight correction factor is greater than 1; Each initial weight corresponding to the first similarity, the second similarity, the third similarity and the fourth similarity is obtained, wherein the sum of the initial weights is equal to 1; If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected by using the weight correction factor; If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected by using the weight correction factor; If the scene type is a user experience scene, the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity are corrected by using the weight correction factor; If the scene type is a problem solving scene, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity and the initial weight corresponding to the fourth similarity are corrected by using the weight correction factor; The first similarity, the second similarity, the third similarity and the fourth similarity are weighted and averaged according to the corrected initial weights, to determine a target similarity.
8. An electronic device, comprising: The computer program instructions, when executed by the processor, implement the method of any one of claims 1-6. The computer program instructions, when executed by the processor, implement the method of any one of claims 1-6.
9. A storage medium having stored thereon computer program instructions, characterized in that,
Citation Information
Patent Citations
News video information extraction method for global deep learning
CN112004111A
Method and system for monitoring advertisement broadcast
CN102799605A
Method and system for converting text into video
CN108986186A