Target video material identification and extraction method, device, equipment and medium

By embedding blind watermark information in video clips and combining it with a self-supervised visual transformation model and feature matching algorithm, the problem of video material tracking difficulty in existing technologies is solved, and efficient and accurate video material management and copyright protection are achieved.

CN120186353BActive Publication Date: 2025-09-16BEIJING LIANSHI LEGEND NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510172324.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-09-16
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

Existing technologies are unable to track and manage video materials efficiently and accurately, especially in large-scale video data. Traditional methods rely on cumbersome computational steps and cannot cope with the management and retrieval of unlabeled video materials, making copyright protection and precise matching of video materials difficult.

Method used

By embedding blind watermark information in video clips and combining it with a self-supervised visual transformation model and feature matching algorithm, video material tracking is achieved. The blind watermark information is used to directly extract video material information, while the self-supervised visual transformation model is used to supplement feature matching, using an approximate nearest neighbor algorithm to perform feature matching in a video material database.

Benefits of technology

It achieves efficient, robust and accurate tracking of target video materials in complex video environments, is suitable for real-time processing of large-scale video data, and ensures accurate management and copyright protection of video materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186353B_ABST
    Figure CN120186353B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, solves the problem in the prior art that the target video material cannot be tracked, and provides a method, device, equipment and medium for identifying and extracting target video material. The method includes: compressing each video segment after embedding blind watermark information to determine the compressed video data; performing blind watermark information identification on each frame of compressed image in the compressed video data, and if the blind watermark information is identified, obtaining the video material information in the blind watermark information as the target video material information; if the blind watermark information is not identified, inputting each frame of compressed image into a self-supervised visual transformation model, and outputting encoded feature information; using an approximate nearest neighbor algorithm, performing feature matching between the encoded feature information and the feature template information in the video material database, and outputting the video material template corresponding to the matched feature template information as the target video material information. The present invention realizes the tracking of the target video material.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of the invention patent application filed on September 12, 2024, with the invention name “Video content analysis method, device and equipment based on multimodal data processing” and application number 202411279794.X. Technical Field

[0002] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment and medium for identifying and extracting target video material. Background Art

[0003] With the increasing richness and diversity of video content, the demand for video material management and video content retrieval has become increasingly important. Copyright protection and material extraction have become key links in video editing and creation. In practical applications, the acquisition of video materials often relies on manual search and annotation, which is inefficient and prone to errors. In addition, the copyright issue of video materials is becoming increasingly serious. Protecting the legitimate rights and interests of video content has become an urgent issue. To this end, blind watermarking technology has been widely used as a hidden identification technology. It embeds invisible watermark information in videos for copyright protection or traceability. However, existing blind watermark information extraction and video material matching technologies still face certain challenges.

[0004] Traditional video material extraction and matching methods usually rely on specific video tags or metadata, but these methods are often unable to cope with the management and retrieval of large amounts of untagged video materials. At the same time, current technical solutions often rely on cumbersome calculation steps in the process of video material retrieval, and the overall processing time is long. They are not suitable for real-time processing of large-scale video data, which limits their widespread promotion in practical applications. Therefore, existing technologies have major technical bottlenecks in efficiently extracting blind watermark information from compressed videos, accurately matching video materials, and improving recognition accuracy.

[0005] Existing Chinese patent CN112004111A discloses a news video information extraction method based on global deep learning, including: at the video decoding layer, a shot labeling module labels each dynamic shot using the TSM spatiotemporal model to generate a label for each dynamic shot; a similarity calculation module calculates the similarity of all labels using the BM25 algorithm, and a shot stitching module stitches dynamic shots with similar labels into a theme video; an image processing module obtains the theme video and processes each frame of the theme video using the optical flow method, grayscale histogram method, Lucas–Kanade algorithm, and image entropy calculation method. , obtain key frames, and send them to the key frame cache module for caching; in the image parsing layer, the well-known person detection module retrieves key frames, uses the YOLOv3 model to perform target object detection and occupation detection, and uses the Facenet model to identify well-known people; the key target detection module uses the Facenet model to identify the target object in the key frame; although the above patent also discloses color distribution recognition and human subject recognition, as well as logically connecting the visual content of multiple shots, it still cannot accurately associate the information of the film and the material: the correlation between the film and the original material is not strong, resulting in the inability to accurately track the video material.

[0006] Therefore, how to track the target video material is an urgent problem to be solved. Summary of the Invention

[0007] In view of this, the present invention provides a method, device, equipment and medium for identifying and extracting target video material, so as to solve the problem that the target video material cannot be tracked in the prior art.

[0008] The technical solution adopted in the present invention is:

[0009] In a first aspect, the present invention provides a method for identifying and extracting target video material, the method comprising:

[0010] Compressing each video segment after embedding the blind watermark information to determine compressed video data, wherein the blind watermark information includes video material information;

[0011] Performing blind watermark information recognition on each compressed image frame in the compressed video data, and if the blind watermark information is recognized, obtaining the video material information embedded in the blind watermark information as target video material information related to the video content;

[0012] If the blind watermark information is not recognized, inputting each frame of the compressed image into a pre-trained self-supervised visual transformation model to output encoding feature information;

[0013] Using an approximate nearest neighbor algorithm, the encoding feature information is matched with feature template information of each video material template in the video material database, and a matching result is output;

[0014] According to the matching result, the video material template corresponding to the feature template information matched with the encoding feature information is output as the target video material information.

[0015] Preferably, before compressing each video segment after embedding the blind watermark information, the method further includes:

[0016] Transcoding the original video material to be parsed to determine standard video data in a preset encoding format;

[0017] Performing grayscale conversion on each frame image in the standard video data to obtain a grayscale image of each frame;

[0018] Using an edge detection algorithm, edge detection is performed on the grayscale image of each frame, and edge feature information is output;

[0019] Determine the edge feature change difference based on the edge feature information corresponding to the grayscale images of adjacent frames;

[0020] Determining video boundary information based on the edge feature change difference and a preset change threshold;

[0021] Decomposing the standard video data according to the video boundary information to determine each of the video segments;

[0022] Using a preset feature processing algorithm, blind watermark information containing video material information is embedded into each frame image in each video clip.

[0023] Preferably, the step of embedding blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm includes:

[0024] Convert the video material information to be embedded into binary format and determine the encoded blind watermark information;

[0025] Decompose each frame image in each video clip to obtain a regional image;

[0026] Performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image;

[0027] Adjusting the high-frequency components in the frequency domain image and embedding blind watermark information into the frequency domain image;

[0028] The frequency domain images are subjected to inverse discrete cosine transformation and then reassembled to complete the embedding of blind watermark information in each video segment.

[0029] Preferably, after outputting the video material template corresponding to the feature template information matched with the encoding feature information as the target video material information based on the matching result, the method further includes:

[0030] Input each frame of compressed image into the pre-trained feature extraction model and output key feature information;

[0031] The key feature information is input into a multimodal large language model, and text information related to the video content is output.

[0032] Preferably, the step of inputting each frame of compressed image into a pre-trained feature extraction model and outputting key feature information comprises:

[0033] Decoding the compressed video data to obtain audio data;

[0034] The compressed image is input into a pre-trained face recognition classification model, and the facial features in the recognized compressed image are classified and labeled to determine the person information;

[0035] Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting primary color information, wherein the primary color information at least includes the advertising brand hue detected in the compressed image or the main hue extracted from the advertising landscape;

[0036] Inputting the compressed image into the object detection model, locating and classifying objects in the compressed image, and determining object information, wherein the object information includes at least the object category and the object location;

[0037] Inputting the audio data into a pre-trained speech transcription model and outputting audio feature information in the audio data;

[0038] The personnel information, color information, object information and audio feature information are respectively input into a multimodal large language model, and the key feature information is output.

[0039] Preferably, after inputting the key feature information into the multimodal large language model and outputting the text information, the method further includes:

[0040] Obtaining the scene type of the advertisement to be delivered, wherein the scene types include: introduction scene, product display scene, user experience scene, and problem-solving scene;

[0041] Based on the scene type, using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity to determine the target similarity corresponding to each video clip;

[0042] According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlights.

[0043] Preferably, the method of evaluating the similarity between text information related to the video clip content and a preset text template using a preset text matching algorithm based on the scene type to determine the target similarity corresponding to each video clip includes:

[0044] Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined, respectively;

[0045] Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1;

[0046] Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1;

[0047] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0048] If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor;

[0049] If the scenario type is a user experience scenario, modifying the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity using the weight modification factor;

[0050] If the scenario type is a problem-solving scenario, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0051] According to the corrected initial weights, a weighted average process is performed on the first similarity, the second similarity, the third similarity, and the fourth similarity to determine a target similarity.

[0052] In a second aspect, the present invention provides a device for identifying and extracting target video material, the device comprising:

[0053] A compression module, configured to compress each video segment after embedding the blind watermark information to determine compressed video data, wherein the blind watermark information includes video material information;

[0054] a first target video material determination module, configured to perform blind watermark information recognition on each compressed image frame in the compressed video data, and if the blind watermark information is recognized, obtain the video material information embedded in the blind watermark information as target video material information related to the video content;

[0055] A feature extraction module is configured to input each frame of the compressed image into a pre-trained self-supervised visual transformation model and output encoding feature information if the blind watermark information is not recognized;

[0056] A feature matching module is used to perform feature matching between the encoded feature information and feature template information of each video material template in the video material database using an approximate nearest neighbor algorithm, and output a matching result;

[0057] The second target video material determination module is configured to output the video material template corresponding to the feature template information matched with the encoding feature information as the target video material information based on the matching result.

[0058] In a third aspect, an embodiment of the present invention further provides an electronic device comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect of the above-mentioned embodiment.

[0059] In a fourth aspect, an embodiment of the present invention further provides a storage medium having computer program instructions stored thereon, which implements the method of the first aspect of the above-mentioned embodiment when the computer program instructions are executed by a processor.

[0060] In summary, the beneficial effects of the present invention are as follows:

[0061] The present invention provides a method, device, equipment and medium for identifying and extracting target video materials. The method includes: compressing each video clip after embedding blind watermark information to determine compressed video data, wherein the blind watermark information includes video material information; performing blind watermark information identification on each frame of compressed image in the compressed video data, and if the blind watermark information is identified, obtaining the video material information embedded in the blind watermark information as target video material information related to the video content; if the blind watermark information is not identified, inputting each frame of the compressed image into a pre-trained self-supervised visual transformation model to output encoding feature information; using an approximate nearest neighbor algorithm, performing feature matching between the encoding feature information and feature template information of each video material template in a video material database, and outputting a matching result; based on the matching result, outputting the video material template corresponding to the feature template information matched with the encoding feature information as the target video material information. The present invention combines the feature extraction of blind watermark information and self-supervised visual transformation model to realize the tracking of target video material, ensuring the accurate extraction and identification of target video material information in different video processing links. First, when compressing the video clip, a blind watermark containing video material information is embedded. The blind watermark information is used to track the target material in the later processing. When analyzing the compressed video data, the blind watermark information is first tried to be identified. If the watermark information is successfully identified, the embedded video material information is directly extracted from the watermark, thereby clarifying the target video material and completing the tracking task. However, if the blind watermark information cannot be identified due to compression loss or other reasons, this process is supplemented by inputting each frame of compressed image into a pre-trained self-supervised visual transformation model. The self-supervised model learns the video The visual features of the frame are used to output the encoded feature information, which can represent the key information of the video content. Next, the approximate nearest neighbor algorithm is used to match the encoded feature information with the feature templates of each video material in the video material database, and the video material template most relevant to the current frame image is obtained through feature matching. Through the above two methods - blind watermark information recognition and feature matching of the self-supervised visual transformation model, the target video material can be tracked in the compressed video data. It can not only cope with the situation where the blind watermark information is damaged, but also ensure the accurate tracking and extraction of the target material through deep learning matching of visual features. This technical solution ensures the efficiency, robustness and accuracy of video material tracking, and is especially suitable for accurate tracking and management of material content in complex video environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work, and these are all within the scope of protection of the present invention.

[0063] Figure 1 Schematic diagram of the overall working process of the target video material identification and extraction method in Example 1 of the present invention;

[0064] Figure 2 Schematic diagram of the process of decomposing the standard video data into multiple video segments in embodiment 1 of the present invention;

[0065] Figure 3 Schematic diagram of the process of embedding blind watermark information containing video material information into each frame image in each video clip in embodiment 1 of the present invention;

[0066] Figure 4 Schematic diagram of the process of identifying and analyzing the compressed video data in Example 1 of the present invention;

[0067] Figure 5 Schematic diagram of the process of blind watermark information recognition for each frame of compressed image in compressed video data in embodiment 1 of the present invention;

[0068] Figure 6 Schematic diagram of the process of extracting and matching features of the compressed image of each frame in embodiment 1 of the present invention;

[0069] Figure 7 This is a flow chart of inputting each frame of compressed image into a pre-trained feature extraction model and outputting key feature information in Example 1 of the present invention;

[0070] Figure 8 This is a structural block diagram of a target video material identification and extraction device in Example 3 of the present invention;

[0071] Figure 9 This is a schematic diagram of the structure of an electronic device in Example 4 of the present invention. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further limitations, elements defined by the phrase "comprising..." do not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the elements. The embodiments of the present invention and the features thereof may be combined with each other if there is no conflict, and all are within the scope of protection of the present invention.

[0073] Example 1

[0074] See Figure 1 Embodiment 1 of the present invention discloses a method for identifying and extracting target video material, the method comprising:

[0075] S1: Get the original video material to be parsed;

[0076] Specifically, acquiring the raw video material to be parsed involves collecting and importing all the raw video files that need to be analyzed and processed, originating from various cameras, recording environments, or external sources. This process ensures that all raw video material to be processed is centralized in a manageable storage location, facilitating subsequent transcoding, segmentation, and analysis operations. By collecting and managing this raw video material, high-quality input data is provided for each subsequent processing step, laying the foundation for the entire video processing and analysis process.

[0077] S2: performing transcoding processing on the original video material to determine standard video data in a preset encoding format;

[0078] Specifically, the original video material is transcoded to ensure video file format consistency and data security. After a user uploads the original video material, it undergoes a unified transcoding process, converting various video formats into a preset standard encoding format. For example, the original format and encoding information of the uploaded video are first identified, the video file is decoded into uncompressed raw data, and the decoded data is re-encoded into a preset standard format, such as H.264 or H.265 encoding, with a uniform resolution and frame rate. Encryption and error detection mechanisms are incorporated into the encoding process to enhance the security of the video data. This not only enables unified management of videos of different formats and simplifies subsequent processing steps, but also reduces storage space requirements by optimizing the encoding format and compression parameters, while ensuring that video quality is not significantly affected. The transcoding process also incorporates encryption and error detection mechanisms to further enhance the security of the video data, preventing data corruption or unauthorized access during transmission and storage. This process enables more efficient management of video material, ensuring the compatibility and security of all data during subsequent processing.

[0079] S3: Decomposing the standard video data into multiple video segments using a shot segmentation technique;

[0080] Specifically, the standard video data is decomposed into multiple video segments using shot segmentation technology. First, shot switching points are identified. These switching points are usually determined by significant changes in image content, scene changes, or significant changes in audio. To ensure the accuracy of the segmentation, advanced algorithms such as edge detection and motion analysis are used. When the amplitude of the detected image edge change exceeds a preset threshold, it is identified as a shot switching point. Subsequently, the standard video data is segmented at these switching points to generate multiple continuous and coherent video segments. Each segment retains important scene information from the original video, facilitating subsequent processing and analysis.

[0081] In one embodiment, see Figure 2 , said S3 includes:

[0082] S31: performing grayscale conversion on each frame image in the standard video data to obtain a grayscale image of each frame;

[0083] Specifically, each frame of the standard video data is first grayscale converted, converting the color image into a grayscale image. Grayscale conversion converts the image's RGB (red, green, and blue) color information into grayscale values, calculating the grayscale value by taking a weighted average of the color values ​​of each pixel. Specifically, the commonly used conversion formula is: Grayscale value = 0.299*R + 0.587*G + 0.114*B. This process simplifies the image data, reduces computational complexity, and highlights the structural and shape features in the image, facilitating subsequent edge detection.

[0084] S32: Using an edge detection algorithm, perform edge detection on the grayscale image of each frame and output edge feature information;

[0085] Specifically, edge detection algorithms, such as Sobel, Canny, and Laplacian, are applied to each grayscale image frame to detect edge information. The goal of edge detection is to identify areas in the image where pixel values ​​vary significantly. These areas typically correspond to the boundaries and outlines of objects. Edge detection algorithms calculate image gradients or use differential operators to identify pixels in the image where grayscale values ​​vary dramatically. These pixels are marked as edge points, thereby generating a binary image containing edge feature information.

[0086] S33: determining edge feature change differences based on edge feature information corresponding to grayscale images of adjacent frames;

[0087] Specifically, after edge detection is complete, edge feature information from adjacent frames is compared to determine changes in edge features. This involves calculating the edge image differences between adjacent frames and identifying areas of image edge change through pixel-level differential operations. Edge feature change differences are calculated by calculating the differences in edge pixel values ​​at corresponding locations in adjacent frames. These differences can be quantified as magnitudes of change, reflecting the dynamic changes in scene content over time.

[0088] S34: Determine video boundary information based on the edge feature change difference and a preset change threshold;

[0089] Specifically, the calculated edge feature change difference is compared with a preset change threshold. If the edge feature change difference between a frame and its adjacent frame exceeds the preset threshold, it is considered that there is a significant scene switch between the frame and the previous frame, that is, it is determined to be a video boundary. The preset change threshold is used to filter out subtle and insignificant changes, ensuring that only obvious content changes (such as scene switches) are identified as video boundaries. In this way, the boundary information of each shot in the video can be accurately determined.

[0090] S35: Decomposing the standard video data according to the video boundary information to determine each of the video segments.

[0091] Specifically, based on the determined video boundary information, the standard video data is segmented at these boundary points to generate multiple independent video segments. Each segment represents a continuous scene or shot, preserving the coherence and integrity of the content. The decomposed video segments facilitate subsequent processing and analysis, such as blind watermark embedding, feature extraction, and content tagging. This approach effectively manages and processes video data, improving the efficiency of video content analysis and processing.

[0092] Specifically, the process involves grayscale conversion, edge detection, calculation of edge feature changes, and determination of video boundaries, ultimately breaking the video into multiple segments. This process enables refined processing and management of video content, laying a solid foundation for subsequent in-depth analysis and applications.

[0093] S4: Using a preset feature processing algorithm, embed the blind watermark information containing the video material information into each frame image in each video clip;

[0094] Specifically, a blind watermark containing video source information is embedded into each frame of each video clip using a preset feature processing algorithm. This ensures accurate identification and association of the original video source material during subsequent editing. The specific steps include: First, a suitable feature processing algorithm, such as discrete cosine transform (DCT) or discrete wavelet transform (DWT), is selected to convert each frame into the frequency domain. Next, specific frequency coefficients (typically low- or mid-frequency coefficients) are selected within the frequency domain and slightly adjusted to embed the blind watermark information. This embedded blind watermark information can contain key data such as a unique source identifier, ensuring that subtle adjustments to the image remain invisible to the human eye without compromising visual quality, yet can be subsequently extracted using a specific algorithm. Finally, an inverse transform is used to restore the image to the spatial domain, generating a frame containing the blind watermark. In this way, each frame of each video clip carries unique identifying information, enabling accurate source material tracking and identification, effectively solving the problem of locating video source material in the final film.

[0095] In one embodiment, see Figure 3 , said S4 includes:

[0096] S41: Convert the video material information to be embedded into a binary format and determine the encoded blind watermark information;

[0097] Specifically, first, the video material information to be embedded (such as unique identifiers, copyright information, etc.) is converted into binary format. The text or other forms of information are converted into a binary data stream through an encoding algorithm, such as ASCII encoding or UTF-8 encoding. After converting this information into binary format, a specific encryption or hashing process is performed to generate the final blind watermark information. The purpose of this process is to ensure that the information remains stable and secure during the embedding process, facilitating subsequent embedding and extraction in the frequency domain image.

[0098] S42: Decompose each frame image in each video clip to obtain a regional image;

[0099] Specifically, each frame in each video clip is then decomposed, typically into several small blocks (such as 8x8 or 16x16 macroblocks) to facilitate finer-grained frequency domain processing. The decomposition process allows each small block to be processed independently, ensuring that the embedding process has minimal impact on the overall image. These decomposed regional images can better analyze and process local features, which helps improve the embedding efficiency and concealment of the blind watermark.

[0100] S43: performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image;

[0101] Specifically, after obtaining the regional image, discrete cosine transform (DCT) is applied to each regional image to convert it from the spatial domain to the frequency domain. DCT decomposes the image data into different frequency components so that the energy of the image is concentrated on a few low-frequency components. These low-frequency components often correspond to the main features of the image. Through this transformation, blind watermark information can be more effectively embedded in the frequency domain without significantly affecting the visual quality of the original image.

[0102] S44: adjusting the high-frequency components in the frequency domain image, and embedding the blind watermark information into the frequency domain image;

[0103] Specifically, high-frequency components are selected in the frequency domain image and adjusted to embed the blind watermark information therein. This is accomplished by making small adjustments to selected DCT coefficients based on the binary blind watermark information, such as increasing or decreasing the value of a specific coefficient to represent a binary 0 or 1. Low-frequency components contain the primary information of the image, and minor adjustments will not significantly affect image quality. Mid-frequency components, on the other hand, balance the stealth and robustness of information embedding. In this way, the blind watermark information is securely and covertly embedded within the frequency domain features of the image.

[0104] S45: performing inverse discrete cosine transform on the frequency domain images and then recombining them to complete embedding of blind watermark information in each video segment.

[0105] Specifically, the adjusted frequency domain image is finally subjected to an inverse discrete cosine transform (IDCT) to convert it back into a spatial domain image, restoring the visual effect of the original image. After IDCT processing, each regional image is reassembled back into the original frame image, ensuring that the blind watermark information is fully retained in each frame. In this way, the blind watermark information is seamlessly embedded into each frame image in the video clip, and the information is firmly embedded and subsequently retrievable without affecting the overall visual quality. Each video clip processed in this way carries unique identification information, achieving accurate tracking and protection of the material.

[0106] S5: compressing each video segment after embedding the blind watermark information to determine compressed video data;

[0107] Specifically, each video clip after the blind watermark information is embedded is compressed to determine the compressed video data. The original video is stored in the company's intranet environment, and a low-resolution version of the video is compressed and stored in the cloud for labeling and analysis. Video compression technology significantly reduces the size of video files by optimizing encoding and reducing resolution, thereby reducing storage and transmission costs. Storing the original video in the intranet ensures high video quality and data security, preventing unauthorized access. The low-resolution version of the video is sufficient to meet the needs of content recognition and labeling, and its smaller file size makes cloud storage and access more efficient. Storing low-resolution videos in the cloud enables remote access and sharing, facilitating collaboration and data processing among team members in different locations, improving work efficiency and flexibility, and ensuring data security and consistency. This solution takes into account the quality, security, and ease of use of video data through efficient compression and distributed storage, providing the team with a flexible working environment and strong technical support.

[0108] S6: Using a preset feature extraction algorithm, the compressed video data is identified and analyzed, and text information related to the video content and target video material information are output.

[0109] Specifically, a preset feature extraction algorithm is used to identify and analyze the compressed video data, outputting text information related to the video content and information about the target video material. This process first analyzes the compressed video frame by frame using advanced feature extraction algorithms, such as multimodal large language models, speech transcription models, face recognition models, object detection models, and color analysis models. These algorithms are capable of identifying various video content, such as text, speech, faces, objects, and colors. The identified information is then integrated and analyzed to generate detailed text descriptions and annotations, including scene descriptions, dialogue content, character information, and object locations. This information can not only be used to understand and retrieve video content, but also helps generate precise identifiers and descriptions related to the video material, ensuring accurate association and utilization of this material in subsequent use. This process significantly improves the efficiency and accuracy of video content analysis, providing strong technical support for content management and application.

[0110] In one embodiment, see Figure 4 , the S6 includes:

[0111] S61: performing blind watermark information recognition on the compressed image of each frame and outputting the recognition result;

[0112] Specifically, the extraction algorithm corresponding to the preset blind watermark embedding algorithm is used to process each frame of the compressed video data, extract the embedded blind watermark information, and determine the target video material information based on the extraction results. This step ensures that even in the case of compressed video, the source of the video clip and its specific content can still be accurately located and identified, thereby ensuring the traceability of the material and the effectiveness of content management.

[0113] In one embodiment, see Figure 5 , the S61 includes:

[0114] S611: Inputting the compressed image of each frame into a pre-trained self-supervised visual transformation model, and outputting encoding feature information;

[0115] Specifically, each compressed image frame is first fed into a pre-trained self-supervised visual transformation model (such as DINO-V2). This model uses self-supervised learning to learn rich visual features from a large amount of unlabeled data. The model then processes the input image, extracts key features, and encodes these features into high-dimensional feature vectors. These feature vectors contain a variety of information, including texture, edges, shape, and color, accurately describing the image content. The encoded feature information output serves as the basis for subsequent feature matching.

[0116] S612: Using an approximate nearest neighbor algorithm, perform feature matching on the encoded feature information and feature template information of each video material template in the video material database, and output a matching result;

[0117] Specifically, the coded feature information is then matched using an approximate nearest neighbor (ANN) algorithm. The approximate nearest neighbor algorithm can quickly find the feature template that is most similar to the target feature in a high-dimensional feature space. The coded feature information of each frame is compared with the feature template information of each video material template pre-stored in the video material database, and the similarity is calculated. The ANN algorithm finds the feature template that is closest to the coded feature information and outputs a matching result, which includes the matching degree and related information of each frame image with the most similar material template in the database.

[0118] S613: According to the matching result, the video material template corresponding to the feature template information matched with the encoding feature information is output as the target video material information.

[0119] Specifically, based on the feature matching results, the target video material information is determined. For each frame, the feature template information with the highest matching degree is selected and its corresponding video material template is output as the target video material information. This ensures that even in the absence of blind watermark information, the source and specific content of the video clip can be accurately identified through feature matching. In this way, each frame of the image can be efficiently and accurately associated with the original video material, achieving precise identification and management of video content.

[0120] S62: If the blind watermark information is identified, the video material information embedded in the blind watermark information is obtained as the target video material information;

[0121] Specifically, after successfully identifying the blind watermark information, the embedded video material information is extracted from the identified watermark data. The blind watermark usually contains a unique material identifier, copyright information or other relevant metadata. Through this information, the source and specific content of the video material can be accurately determined and used as the target video material information to ensure that each frame of the image can be accurately associated with the corresponding original material.

[0122] S63: If the blind watermark information is not recognized, extract and match the features of each frame of the compressed image, and determine the target video material information based on the matching results;

[0123] Specifically, if the blind watermark information cannot be identified, the compressed image is further analyzed using feature extraction and matching methods. Features such as texture, edges, shape, and color are extracted to form a unified image feature code. The extracted image features are then matched with the material features in the database to determine the most similar material. This ensures that the source and content of the video clip can be accurately identified even in the absence of a blind watermark. Based on the matching results, the target video material information is determined and output, thereby achieving accurate identification and management of the material.

[0124] S64: Inputting each frame of compressed image into a pre-trained feature extraction model to output key feature information;

[0125] Specifically, each frame of compressed image is input into a pre-trained feature extraction model to output key feature information. This process involves processing each frame of image through a trained feature extraction model in the field of computer vision, such as a convolutional neural network (CNN) or a Transformer model. The model has been trained on a large amount of image data and can automatically extract important visual features in the image, such as edges, textures, color distribution, and shapes. After the image is input, the model will extract high-level features from the image layer by layer through a series of convolutional layers, activation functions, and pooling operations. These features will be encoded into a high-dimensional feature vector that contains the core information of the image. These encoded key feature information can not only accurately describe the image content, but can also be used for subsequent feature matching, classification, or retrieval tasks to ensure the effective use and analysis of image data.

[0126] In one embodiment, see Figure 6 , the S64 includes:

[0127] S641: Decode the compressed video data to obtain audio data;

[0128] Specifically, the compressed video data is first decoded to extract the audio data. This process involves restoring the compressed video file to its original uncompressed state and extracting the audio portion of the video stream through a decoder. The decoding process typically involves converting the compressed encoded data stream (such as MP4 or H.264 format) into an uncompressed audio format (such as WAV or PCM) that can be further processed. The extracted audio data contains all the sound information in the video, including dialogue, background music, and ambient sounds, providing the basis for subsequent audio analysis and processing.

[0129] S642: Input the compressed image into a pre-trained face recognition classification model, classify and label the facial features in the recognized compressed image, and determine the person information;

[0130] S643: Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting primary color information, wherein the primary color information at least includes the advertising brand color tone detected in the compressed image or the main color tone extracted from the advertising landscape painting;

[0131] S644: Input the compressed image into the object detection model, locate and classify objects in the compressed image, and determine item information, wherein the item information includes at least item category and item location;

[0132] Specifically, the compressed image is fed into a pre-trained face recognition and classification model, a color analysis model, and an object detection model, respectively, to output person information, color information, and object information. Specifically, this step involves passing each frame of the compressed image into multiple specially trained deep learning models to extract and analyze different types of information. First, the image is fed into the face recognition and classification model. This model, trained using a large amount of pre-labeled facial data, accurately identifies facial features in the image and performs classification and labeling. For example, it can identify actors or other known people appearing in the image, such as the facial features of a reporter in a news report. Second, the image is fed into the color analysis model, which analyzes the color distribution in the image and extracts primary color information, such as the brand hue detected in an advertising image or the dominant color in a landscape painting. Finally, the image is fed into the object detection model, which locates and classifies objects in the image, such as identifying products in a shopping ad or traffic signs in a video, and determining their category and location. These models each process different aspects of the image, and by combining this information, they can obtain a comprehensive description of the image content, such as identifying actors, background colors, and major objects in a scene in a movie clip, providing key details for in-depth understanding and analysis of videos.

[0133] S645: Inputting the audio data into a pre-trained speech transcription model, and outputting audio feature information in the audio data;

[0134] Specifically, the extracted audio data is fed into a pre-trained speech transcription model to convert the spoken content in the audio into text. The speech transcription model uses natural language processing technology and speech recognition algorithms to identify the language, intonation, and syllables in the audio and transcribe it into text. This process can identify and record conversations, speeches, or other spoken information and generate corresponding text descriptions. The audio feature information includes the textual form of the speech content, allowing the information in the audio to be further analyzed and processed.

[0135] S646: Input the personnel information, color information, object information and audio feature information into the multimodal large language model respectively, and output the key feature information.

[0136] Specifically, the person information, color information, and object information extracted from the image, as well as the audio feature information extracted from the audio, are input into a multimodal large language model. The multimodal large language model can fuse data from different sources and understand and analyze the relationships between various types of information. The model will conduct a comprehensive analysis of these multimodal data to generate key feature information that describes the video content. This includes integrating the visual features in the image and the speech content in the audio to provide a comprehensive understanding and summary of the content. This fusion analysis method enables the various information extracted from the video to be effectively integrated and interpreted, generating deep insights into the video content.

[0137] S65: Input the key feature information into a multimodal large language model and output the text information.

[0138] Specifically, the key feature information is input into the multimodal large language model, and the text information is output. First, the key feature information extracted from the image (such as personnel information, color information, object information) and audio (such as speech transcription text) is integrated into a unified data format. These feature information will be input into the multimodal large language model, which uses deep learning technology to fuse data of different modalities (visual, audio, text, etc.) to understand the relationship between these information. By analyzing and contextually modeling the input multimodal data, the model can generate a coherent text description that accurately expresses the main points of the video content. The output text information usually includes a summary of the video scene, the content of the dialogue, a detailed description of the characters and objects, and other relevant information. This comprehensive analysis capability enables the multimodal large language model to provide a comprehensive and accurate summary of the video content, enhancing the understanding and utilization of video data.

[0139] Example 2

[0140] In the above embodiment 1, the analysis results related to the content of each video clip are obtained through content analysis, including personnel information, color information, object information and audio feature information. In the real advertising video editing scene, the relevance of the video clips is evaluated based on the content analysis results, and the video highlights are edited accordingly, which can greatly improve the quality of the video and the viewing experience. To this end, in one implementation, please refer to Figure 7 , after S6, further comprising:

[0141] Obtaining the scene type of the advertisement to be delivered, wherein the scene types include: introduction scene, product display scene, user experience scene, and problem-solving scene;

[0142] Specifically, the scene type of the advertisement to be delivered is obtained. Advertisements are generally divided into multiple scene types, such as introduction scenes, product display scenes, user experience scenes, and problem-solving scenes. Introduction scenes are usually used to attract the audience's attention and present the core information or brand of the advertisement; product display scenes focus on displaying the product's functions, appearance, and usage methods, helping the audience intuitively understand the product's value; user experience scenes show how users use the product and the experience and feedback they gain during use; problem-solving scenes are used to show how the product helps users solve practical problems, emphasizing the product's effectiveness and practicality. The step of obtaining scene types can be completed by analyzing the advertising planning content, brand requirements, and market demand. The advertising editing team usually presets these scene types to ensure that the advertisement can be presented according to the set plot or marketing strategy. At the same time, video analysis technology is used to extract key features in the video (such as products, characters, and scene elements) and match these features with predefined scene types to quickly determine the scene category to which the video clip belongs.

[0143] Based on the scene type, using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity to determine the target similarity corresponding to each video clip;

[0144] Specifically, after confirming the scene type, the video clip is further analyzed based on these scene types. Specifically, a text matching algorithm is used to determine the similarity between the clip and the target scene. First, the content of the video clip is usually converted into text form through a speech transcription model or subtitle extraction technology. This text information reflects the core content of the video clip, such as product name, function description, user experience, etc. Using text matching algorithms such as TF-IDF and BERT, the similarity between the text information in the video clip and the preset scene text template is calculated. Each scene type has a set of preset text templates, which contain keywords or sentence patterns related to the scene. The core of the similarity assessment is to calculate the semantic similarity between the text in the video clip and the template text, and then evaluate the matching degree between the video clip and the scene type.

[0145] In one embodiment, based on the scene type, using a preset text matching algorithm, performing similarity evaluation on text information related to the video clip content and a preset text template to determine the target similarity corresponding to each video clip includes:

[0146] Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined, respectively;

[0147] Specifically, a text matching algorithm is first used to process the textual information associated with the video clip. This textual information includes information about people, colors, objects, and audio features. This textual information is then evaluated for similarity with a pre-set text template. The text template typically represents a specific scene type, such as an introduction or product display, and contains keywords and related sentences. Using algorithms such as TF-IDF or BERT, the semantic similarity between the video textual information and the template is calculated, determining the first, second, third, and fourth similarities corresponding to the information about people, colors, objects, and audio features, respectively. These similarities reflect the degree of match between each video clip and the relevant features in the scene.

[0148] Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1;

[0149] Specifically, in order to adjust the evaluation of different similarities, a preset weight correction factor is obtained. This correction factor is usually greater than 1, indicating that certain similarities need to be amplified. The setting of the weight correction factor depends on the specific scenario requirements. For example, in some scenarios, color information may be more important than other features, so the weight of color similarity needs to be increased. The weight correction factor is used to dynamically adjust the importance of different similarities so that the final result can better meet the scenario requirements and advertising planning goals.

[0150] Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1;

[0151] Specifically, after similarity calculations are performed, an initial weight is assigned to each similarity. This initial weight determines the influence of each similarity in the overall similarity calculation, and the sum of all initial weights equals 1. Depending on the advertising planning objectives, the initial weight assignment can be preset to different values. For example, a product display scenario might give a higher weight to item information, while a user experience scenario might give a higher weight to person information.

[0152] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0153] Specifically, in the introduction scene, the color information and audio feature information of the video are usually key factors in attracting the audience's attention. Therefore, the weight correction factor is used to correct the similarities related to these features, namely the second similarity and the fourth similarity. The correction process includes multiplying the weight correction factor by the second similarity to obtain a new second similarity, and multiplying the weight correction factor by the fourth similarity to obtain a new fourth similarity. This means that the initial weights of color and audio will be amplified, so that these features occupy a larger proportion in the similarity evaluation, ensuring that the visual and auditory effects of the introduction scene are more attractive.

[0154] If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor;

[0155] Specifically, in product display scenarios, color and object information in videos are particularly important. Therefore, a weight correction factor is used to modify the second and third similarities. This adjustment ensures that the visual effects and object details of the product display are more prominent, attracting the audience's attention and prompting them to focus on the product's characteristics.

[0156] If the scenario type is a user experience scenario, modifying the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity using the weight modification factor;

[0157] Specifically, user experience scenarios focus on showcasing the interaction between people and products. Therefore, a weight correction factor is used to modify the initial weights corresponding to the first and third similarities. In these scenarios, the weights of these two types of information are amplified to ensure that the interaction between the user and the product and the product itself are more prominent, enhancing the user's sense of immersion.

[0158] If the scenario type is a problem-solving scenario, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0159] Specifically, in problem-solving scenarios, it's often necessary to highlight information about people, items, and audio features to demonstrate how the product solves the user's real-world problem. Therefore, using weight correction factors to modify the first, third, and fourth similarities can better demonstrate how the product addresses the user's concerns through interaction with the user, the problem-solving process, and related audio prompts, enhancing the persuasiveness and authenticity of the scenario.

[0160] Performing weighted averaging processing on the first similarity, the second similarity, the third similarity, and the fourth similarity based on the corrected initial weights to determine a target similarity;

[0161] Specifically, the first, second, third, and fourth similarities obtained previously are processed. Each similarity represents the degree of matching between different types of information. In the previous steps, each similarity was weighted based on the needs of the scene to ensure that the content that is important to the scene is highlighted. The weighted average processing refers to multiplying each modified initial weight by the corresponding similarity value, and then adding and averaging these results to calculate the overall target similarity. The target similarity indicates the degree of match between the video clip and the specific scene after integrating multimodal information. Through this process, the system can comprehensively evaluate the fit of each video clip with the preset scene under the multimodal information.

[0162] According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlights.

[0163] Specifically, after determining the target similarity, the value is compared with a preset similarity threshold. The similarity threshold is a pre-set standard used by the system to determine whether a video clip meets the minimum requirements of the scene. This threshold can be adjusted based on the needs of different scenes. For example, a product demonstration scene may require a higher similarity threshold, while an introduction scene may allow for a lower threshold. The similarity threshold is often set based on user expectations or advertising goals. If the target similarity is greater than or equal to the threshold, the video clip is highly relevant and well-matched to the scene requirements, making it suitable for further processing. Clips below the threshold are considered mismatched with the current scene and excluded from the video collection selection. After determining which video clips meet the similarity requirements, the system synthesizes these clips. The synthesis process involves combining multiple highly similar video clips according to a certain logic or strategy, such as chronological order, plot development, and visual consistency, to form a complete advertising collection. This step can be adjusted according to the design of the marketing strategy, such as combining product demonstration clips with user experience clips, or arranging problem-solving clips with product demonstration clips to ensure a coherent narrative structure for the collection.

[0164] Example 3

[0165] See Figure 8 Embodiment 3 of the present invention further provides a device for identifying and extracting target video material, the device comprising:

[0166] Video material acquisition module, used to obtain the original video material to be parsed;

[0167] A transcoding processing module, configured to perform transcoding processing on the original video material to determine standard video data in a preset encoding format;

[0168] A shot segmentation module, configured to decompose the standard video data into multiple video segments using a shot segmentation technique;

[0169] A blind watermark embedding module is used to embed blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm;

[0170] A compression processing module is used to compress each video segment after the blind watermark information is embedded to determine the compressed video data;

[0171] The feature extraction module is used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information.

[0172] Specifically, the target video material identification and extraction device provided by the embodiment of the present invention is adopted, and the device includes: a video material acquisition module, which is used to obtain the original video material to be parsed; a transcoding processing module, which is used to transcode the original video material and determine standard video data in a preset encoding format; a shot segmentation module, which is used to decompose the standard video data into multiple video segments using shot segmentation technology; a blind watermark embedding module, which is used to embed blind watermark information containing video material information into each frame image in each video segment using a preset feature processing algorithm; a compression processing module, which is used to compress each video segment after the blind watermark information is embedded to determine compressed video data; a feature extraction module, which is used to use a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information. This device achieves precise association between finished films and source material, and accurately identifies and analyzes video content through a series of orderly processing steps. First, by acquiring and transcoding the original video material, all data formats are ensured to be uniform. Then, using shot segmentation technology, the standard video data is broken down into multiple easily manageable segments, and blind watermark information is embedded, so that each source material segment can be accurately identified in the finished film, even after editing and compression. Finally, a feature extraction algorithm is used to identify and analyze the compressed video, outputting detailed text information related to the video content and target video material information. This method not only improves the matching accuracy of video material and finished films, ensuring that each segment of material can be accurately located and identified, but also efficiently extracts and summarizes video content, provides more detailed analysis and annotation, and solves the efficiency and accuracy problems of traditional video processing and analysis.

[0173] Example 4

[0174] In addition, combined Figure 1 The target video material identification and extraction method described in the first embodiment of the present invention can be implemented by an electronic device. Figure 9 A schematic diagram of the hardware structure of an electronic device provided in Example 4 of the present invention is shown.

[0175] An electronic device may include a processor and a memory storing computer program instructions.

[0176] Specifically, the processor may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits for implementing the embodiments of the present invention.

[0177] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In suitable cases, the memory may include a removable or non-removable (or fixed) medium. In suitable cases, the memory may be inside or outside the data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). In suitable cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0178] The processor implements any one of the target video material identification and extraction methods in the above embodiments by reading and executing computer program instructions stored in the memory.

[0179] In one example, the electronic device may further include a communication interface and a bus. Figure 9 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.

[0180] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.

[0181] Bus comprises hardware, software or both, couples the parts of described equipment together.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus can comprise one or more buses.Although the embodiment of the present invention describes and shows specific bus, the present invention considers any suitable bus or interconnection.

[0182] Example 5

[0183] In addition, in conjunction with the target video material identification and extraction method in the first embodiment, the fifth embodiment of the present invention may further provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when executed by a processor, the computer program instructions implement any of the target video material identification and extraction methods in the above embodiments.

[0184] In summary, the embodiments of the present invention provide a method, apparatus, device, and medium for identifying and extracting target video material.

[0185] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0186] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card or the like. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0187] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant location, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0188] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.

[0189] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.

Claims

1. A method for identifying and extracting target video material, characterized in that: The method comprises: Compressing each video segment after embedding the blind watermark information to determine compressed video data, wherein the blind watermark information includes video material information; Performing blind watermark information recognition on each compressed image frame in the compressed video data, and if the blind watermark information is recognized, obtaining the video material information embedded in the blind watermark information as target video material information related to the video content; If the blind watermark information is not recognized, inputting each frame of the compressed image into a pre-trained self-supervised visual transformation model to output encoding feature information; Using an approximate nearest neighbor algorithm, the encoding feature information is matched with feature template information of each video material template in the video material database, and a matching result is output; According to the matching result, the video material template corresponding to the feature template information matched with the encoding feature information is output as the target video material information.

2. The target video material identification and extraction method according to claim 1, characterized in that: Before compressing each video segment after embedding the blind watermark information, the method further includes: Transcoding the original video material to be parsed to determine standard video data in a preset encoding format; Performing grayscale conversion on each frame image in the standard video data to obtain a grayscale image of each frame; Using an edge detection algorithm, edge detection is performed on the grayscale image of each frame, and edge feature information is output; Determine the edge feature change difference based on the edge feature information corresponding to the grayscale images of adjacent frames; Determining video boundary information based on the edge feature change difference and a preset change threshold; Decomposing the standard video data according to the video boundary information to determine each of the video segments; Using a preset feature processing algorithm, blind watermark information containing video material information is embedded into each frame image in each video clip.

3. The target video material identification and extraction method according to claim 2, characterized in that: The method of embedding blind watermark information containing video material information into each frame image in each video clip by using a preset feature processing algorithm includes: Convert the video material information to be embedded into binary format and determine the encoded blind watermark information; Decompose each frame image in each video clip to obtain a regional image; Performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image; Adjusting the high-frequency components in the frequency domain image and embedding blind watermark information into the frequency domain image; The frequency domain images are subjected to inverse discrete cosine transformation and then reassembled to complete the embedding of blind watermark information in each video segment.

4. The target video material identification and extraction method according to claim 1, characterized in that: After outputting the video material template corresponding to the feature template information matched with the encoding feature information as the target video material information according to the matching result, the method further includes: Input each frame of compressed image into the pre-trained feature extraction model and output key feature information; The key feature information is input into a multimodal large language model, and text information related to the video content is output.

5. The target video material identification and extraction method according to claim 4, characterized in that: Inputting each frame of compressed image into a pre-trained feature extraction model, outputting key feature information includes: Decoding the compressed video data to obtain audio data; The compressed image is input into a pre-trained face recognition classification model, and the facial features in the recognized compressed image are classified and labeled to determine the person information; Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting primary color information, wherein the primary color information at least includes the advertising brand hue detected in the compressed image or the main hue extracted from the advertising landscape; Inputting the compressed image into the object detection model, locating and classifying objects in the compressed image, and determining object information, wherein the object information includes at least the object category and the object location; Inputting the audio data into a pre-trained speech transcription model and outputting audio feature information in the audio data; The personnel information, color information, object information and audio feature information are respectively input into a multimodal large language model, and the key feature information is output.

6. The target video material identification and extraction method according to claim 4, characterized in that: After inputting the key feature information into the multimodal large language model and outputting the text information, the method further includes: Obtaining the scene type of the advertisement to be delivered, wherein the scene types include: introduction scene, product display scene, user experience scene, and problem-solving scene; Based on the scene type, using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity to determine the target similarity corresponding to each video clip; According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlights.

7. The target video material identification and extraction method according to claim 6, characterized in that: The method of evaluating the similarity between text information related to the video clip content and a preset text template using a preset text matching algorithm based on the scene type to determine the target similarity corresponding to each video clip includes: Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined, respectively; Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1; Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1; If the scene type is an introduction scene, amplifying the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity by using the weight correction factor; If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are amplified using the weight correction factor; If the scenario type is a user experience scenario, amplifying the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity by using the weight correction factor; If the scenario type is a problem-solving scenario, amplifying the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity by using the weight correction factor; According to the initial weights after the amplification process, a weighted average process is performed on the first similarity, the second similarity, the third similarity and the fourth similarity to determine the target similarity.

8. A target video material identification and extraction device, characterized in that: The device comprises: A compression module, configured to compress each video segment after embedding the blind watermark information to determine compressed video data, wherein the blind watermark information includes video material information; a first target video material determination module, configured to perform blind watermark information recognition on each compressed image frame in the compressed video data, and if the blind watermark information is recognized, obtain the video material information embedded in the blind watermark information as target video material information related to the video content; A feature extraction module is configured to input each frame of the compressed image into a pre-trained self-supervised visual transformation model and output encoding feature information if the blind watermark information is not recognized; A feature matching module is used to perform feature matching between the encoded feature information and feature template information of each video material template in the video material database using an approximate nearest neighbor algorithm, and output a matching result; The second target video material determination module is configured to output the video material template corresponding to the feature template information matched with the encoding feature information as the target video material information based on the matching result.

9. An electronic device, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, which implement the method according to any one of claims 1 to 7 when the computer program instructions are executed by the processor.

10. A storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • News video information extraction method for global deep learning

    CN112004111A

  • Method and device for detecting advertisements

    CN103235956A

  • Video secret photographing and secret recording traceability method based on digital watermarking

    CN118540555A