Video content analysis method, device and equipment based on multimodal data processing

Through multimodal data processing technology, including lens slicing and blind watermark embedding, the problem of inability to accurately correlate films and materials in the existing technology is solved, and the accurate identification and analysis of video content is achieved, and the analysis efficiency and accuracy are improved.

CN119182920BActive Publication Date: 2025-05-23BEIJING LIANSHI LEGEND NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411279794.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-05-23
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The prior art cannot accurately correlate information from films and materials, and cannot accurately disassemble content, resulting in insufficient accuracy and robustness in video content analysis.

Method used

The video content analysis method based on multimodal data processing is adopted to accurately identify and analyze video content by obtaining original video materials, transcoding processing, lens slicing, blind watermark embedding, compression processing and feature extraction algorithms.

Benefits of technology

It realizes accurate correlation of films and materials, and accurately recognizes and analyzes video content, improves the matching accuracy between video materials and films, ensures that each piece of material can be accurately positioned and identified, efficiently extracts and summarizes video content, and provides more detailed analysis and labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119182920B_ABST
    Figure CN119182920B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, solves the problem in the prior art that it is impossible to accurately identify and analyze video content while accurately associating finished films with materials, and provides a video content analysis method, device and equipment based on multimodal data processing. The method includes: obtaining the original video material to be analyzed; transcoding the original video material to determine standard video data in a preset encoding format; using shot segmentation technology to decompose the standard video data into multiple video clips; using a preset feature processing algorithm to embed blind watermark information into each frame image in each video clip; compressing each video clip to determine compressed video data; identifying and analyzing the compressed video data, and outputting text information related to the video content and target video material information. The present invention realizes the precise association between finished films and materials, and improves the accuracy of video content recognition and analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a video content analysis method, device and equipment based on multimodal data processing. Background Art

[0002] With the explosive growth of video data, video content analysis plays an increasingly important role in modern information processing and management. Efficient and accurate video content analysis can not only improve user experience, but also provide support for various application scenarios, such as video recommendation systems, content review, copyright protection, and advertising. Especially in the digital media advertising industry, we can use automated video content analysis to quickly obtain key information in the video, save human resources, improve information processing efficiency, and help better understand and utilize video data. At present, common solutions for video content analysis are mainly concentrated in the following aspects: Manual labeling and classification: relying on manual labeling and classification of video content. Although it can achieve high accuracy, the labor cost is high and the efficiency is low; preliminary content recognition is performed through color histograms, motion detection and other means. This method can quickly identify and summarize all the information appearing in the video screen. It is more effective when dealing with simple scenes, but it lacks accuracy and robustness when faced with complex and changeable content such as advertising video materials and edited films. Using existing large model algorithms for content analysis has a high learning cost for users, and requires repeated manual training of the prompt model, and subsequent manual elimination of useless information to obtain the desired results, which often takes up a lot of time for the system and users. With the development of machine learning and technology, regardless of the time and cost paid, the above scheme can achieve the purpose of parsing video content. However, in the process of implementing the above scheme, the following problems cannot be accurately solved: It is impossible to accurately associate the information of the film and the material: the correlation between the film and the original material is not strong, resulting in the inability to accurately track the use effect of the material; it is impossible to accurately disassemble the content: the existing technology lacks accuracy and robustness when dealing with the changeable content of digital media advertising videos, and cannot accurately disassemble the meaning of the video.

[0003] The existing Chinese patent CN112004111A discloses a news video information extraction method based on global deep learning, including: at the video decoding layer, the shot labeling module labels each dynamic shot through the TSM spatiotemporal model to generate a label for each dynamic shot; the similarity calculation module calculates the similarity of all labels through the BM25 algorithm, and the shot stitching module stitches dynamic shots with similar labels into a theme video; the image processing module obtains the theme video, processes each frame of the theme video using the optical flow method, the grayscale histogram method, the Lucas-Kanade algorithm and the image entropy calculation method to obtain the key frame, and It is sent to the key frame cache module for caching; in the image parsing layer, the famous person detection module retrieves the key frame, uses the YOLOv3 model for target object detection and occupation detection, and uses the Facenet model to identify famous people; the key target detection module uses the Facenet model to identify the target object in the key frame; although the above patent also discloses color distribution recognition and character subject recognition, as well as logically connecting the visual content of multiple shots, it still cannot accurately associate the information of the film and the material: the correlation between the film and the original material is not strong, which makes it impossible to accurately track the use effect of the material; at the same time, it is impossible to accurately decompose the content.

[0004] Therefore, how to accurately identify and analyze video content while accurately associating finished films and materials is an urgent problem to be solved. Summary of the invention

[0005] In view of this, the present invention provides a video content analysis method, device and equipment based on multimodal data processing, so as to solve the problem in the prior art that it is impossible to accurately identify and analyze the video content while accurately associating the film with the material.

[0006] The technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides a video content analysis method based on multimodal data processing, the method comprising:

[0008] S1: Obtain the original video material to be parsed;

[0009] S2: performing transcoding processing on the original video material to determine standard video data in a preset encoding format;

[0010] S3: Decomposing the standard video data into multiple video segments by using a shot segmentation technology;

[0011] S4: using a preset feature processing algorithm, embedding blind watermark information containing video material information into each frame image in each video clip;

[0012] S5: compressing each video segment after the blind watermark information is embedded to determine compressed video data;

[0013] S6: Using a preset feature extraction algorithm, the compressed video data is identified and analyzed, and text information related to the video content and target video material information are output.

[0014] Preferably, S3 includes:

[0015] S31: Perform grayscale conversion on each frame image in the standard video data to obtain each frame grayscale image;

[0016] S32: using an edge detection algorithm to perform edge detection on the grayscale image of each frame, and output edge feature information;

[0017] S33: determining edge feature change difference according to edge feature information corresponding to grayscale images of adjacent frames;

[0018] S34: determining video boundary information according to the edge feature change difference and a preset change threshold;

[0019] S35: Decomposing the standard video data according to the video boundary information to determine each of the video segments.

[0020] Preferably, the S4 includes:

[0021] S41: converting the video material information to be embedded into a binary format, and determining the encoded blind watermark information;

[0022] S42: Decompose each frame image in each video segment to obtain a regional image;

[0023] S43: performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image;

[0024] S44: adjusting the high frequency components in the frequency domain image, and embedding the blind watermark information into the frequency domain image;

[0025] S45: performing inverse discrete cosine transformation on the frequency domain images and then recombining them to complete embedding of blind watermark information in each video segment.

[0026] Preferably, the S6 includes:

[0027] S61: performing blind watermark information recognition on the compressed image of each frame, and outputting the recognition result;

[0028] S62: if the blind watermark information is identified, obtaining the video material information embedded in the blind watermark information as the target video material information;

[0029] S63: If the blind watermark information is not identified, extract and match the features of each frame of the compressed image, and determine the target video material information according to the matching result;

[0030] S64: inputting each frame of compressed image into a pre-trained feature extraction model, and outputting key feature information;

[0031] S65: Input the key feature information into a multimodal large language model and output the text information. Preferably, S61 includes:

[0032] S611: Inputting the compressed image of each frame into a pre-trained self-supervised visual transformation model, and outputting encoding feature information;

[0033] S612: using an approximate nearest neighbor algorithm, performing feature matching between the encoding feature information and feature template information of each video material template in a video material database, and outputting a matching result;

[0034] S613: According to the matching result, the video material template corresponding to the feature template information matching the encoding feature information is output as the target video material information.

[0035] Preferably, the S64 includes:

[0036] S641: Decode the compressed video data to obtain audio data;

[0037] S642: Input the compressed image into a pre-trained face recognition classification model, classify and label the facial features in the recognized compressed image, and determine the person information;

[0038] S643: Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting main color information, wherein the main color information at least includes the advertising brand color tone detected in the compressed image or the main color tone extracted from the advertising landscape painting;

[0039] S644: Input the compressed image into the target detection model, locate and classify the objects in the compressed image, and determine the object information, wherein the object information at least includes the object category and the object location;

[0040] S645: Input the audio data into a pre-trained speech transcription model, and output audio feature information in the audio data;

[0041] S646: Input the personnel information, color information, object information and audio feature information into a multimodal large language model respectively, and output the key feature information.

[0042] Preferably, after S6, the method further includes:

[0043] Obtaining the scene type of the advertisement to be placed, wherein the scene type includes: introduction scene, product display scene, user experience scene and problem solving scene;

[0044] Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined respectively;

[0045] Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1;

[0046] Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1;

[0047] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0048] If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor;

[0049] If the scenario type is a user experience scenario, using the weight correction factor to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity;

[0050] If the scenario type is a problem-solving scenario, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0051] According to the corrected initial weights, the first similarity, the second similarity, the third similarity and the fourth similarity are weighted averaged to determine the target similarity;

[0052] According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlight.

[0053] In a second aspect, the present invention provides a video content analysis device based on multimodal data processing, the device comprising:

[0054] A video material acquisition module is used to acquire the original video material to be parsed;

[0055] A transcoding processing module, used for performing transcoding processing on the original video material to determine standard video data in a preset coding format;

[0056] A shot segmentation module, used to decompose the standard video data into multiple video segments by using a shot segmentation technology;

[0057] A blind watermark embedding module is used to embed blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm;

[0058] A compression processing module, used to compress each video segment after the blind watermark information is embedded, and determine the compressed video data;

[0059] The feature extraction module is used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information.

[0060] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect in the above-mentioned embodiment.

[0061] In a fourth aspect, an embodiment of the present invention further provides a storage medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method of the first aspect in the above-mentioned embodiment is implemented.

[0062] In summary, the beneficial effects of the present invention are as follows:

[0063] The present invention provides a video content analysis method, device and equipment based on multimodal data processing, the method comprising: obtaining original video material to be analyzed; performing transcoding processing on the original video material to determine standard video data of a preset encoding format; utilizing shot segmentation technology to decompose the standard video data into multiple video segments; utilizing a preset feature processing algorithm to embed blind watermark information containing video material information into each frame image in each video segment; performing compression processing on each video segment after the blind watermark information is embedded to determine compressed video data; utilizing a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information. The present invention realizes accurate association between finished films and materials, and accurately identifies and analyzes video content through a series of orderly processing steps. First, by acquiring and transcoding the original video material, all data formats are ensured to be unified. Then, the standard video data is decomposed into multiple easy-to-manage segments using the shot segmentation technology, and blind watermark information is embedded, so that each material segment can be accurately identified in the finished film later, even after editing and compression processing. Finally, the compressed video is identified and analyzed through the feature extraction algorithm, and detailed text information related to the video content and target video material information are output. This method not only improves the matching accuracy of video materials and finished films, ensures that each piece of material can be accurately located and identified, but also can efficiently extract and summarize video content, provide more detailed analysis and annotation, and solve the efficiency and accuracy problems in traditional video processing and analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the drawings required for use in the embodiment of the present invention will be briefly introduced below. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work, and these are all within the protection scope of the present invention.

[0065] Figure 1 Schematic diagram of the overall working process of the video content analysis method based on multimodal data processing in Example 1 of the present invention;

[0066] Figure 2 is a schematic diagram of a process of decomposing the standard video data into multiple video segments in Embodiment 1 of the present invention;

[0067] Figure 3 It is a schematic diagram of a process of embedding blind watermark information containing video material information into each frame image in each video clip in Embodiment 1 of the present invention;

[0068] Figure 4 This is a schematic diagram of a process for identifying and analyzing the compressed video data in Embodiment 1 of the present invention;

[0069] Figure 5 It is a schematic diagram of a process of blindly identifying watermark information of each frame of compressed image in compressed video data in Embodiment 1 of the present invention;

[0070] Figure 6 It is a schematic diagram of the process of extracting and matching features of the compressed image of each frame in Embodiment 1 of the present invention;

[0071] Figure 7 This is a flow chart of inputting each frame of compressed image into a pre-trained feature extraction model and outputting key feature information in Embodiment 1 of the present invention;

[0072] Figure 8 is a structural block diagram of a video content analysis device based on multimodal data processing in Embodiment 3 of the present invention;

[0073] Fig. 9 It is a schematic diagram of the structure of an electronic device in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described in conjunction with the drawings in the embodiment of the present invention. It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. In the description of the present invention, it should be understood that the orientation or position relationship indicated by the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. In the absence of further restrictions, the elements defined by the phrase "comprising..." do not exclude the existence of other identical elements in the process, method, article or device comprising the elements. If there is no conflict, the embodiments of the present invention and the various features in the embodiments can be combined with each other, all within the protection scope of the present invention.

[0075] Example 1

[0076] See also Figure 1Embodiment 1 of the present invention discloses a video content analysis method based on multimodal data processing, the method comprising:

[0077] S1: Obtain the original video material to be parsed;

[0078] Specifically, obtaining the raw video material to be parsed involves collecting and importing all the raw video files that need to be analyzed and processed, which come from different camera devices, recording environments or external resources. This process ensures that all the raw video materials to be processed are concentrated in a manageable storage location, which is convenient for subsequent transcoding, segmentation and analysis operations. By collecting and managing these raw video materials, high-quality input data can be provided for each subsequent processing step, laying the foundation for the entire video processing and analysis process.

[0079] S2: performing transcoding processing on the original video material to determine standard video data in a preset encoding format;

[0080] Specifically, the original video material is transcoded to ensure the consistency of the video file format and data security. After the user uploads the original video material, the original video material will be uniformly transcoded to convert videos of various formats into a preset standard encoding format, for example: first identify the original format and encoding information of the uploaded video, decode the video file into uncompressed raw data, re-encode the decoded data into a preset standard format, such as H.264 or H.265 encoding, set a unified resolution and frame rate, add encryption and error detection mechanisms in the encoding process, and improve the security of video data. This not only realizes the unified management of videos of different formats and simplifies the subsequent processing steps, but also reduces the storage space requirements by optimizing the encoding format and compression parameters, while ensuring that the video quality is not significantly affected. The transcoding process also adds encryption and error detection mechanisms to further improve the security of video data and prevent data from being damaged or unauthorized access during transmission and storage. Through this process, video materials can be managed more effectively to ensure the compatibility and security of all data in subsequent processing.

[0081] S3: Decomposing the standard video data into multiple video segments by using a shot segmentation technology;

[0082] Specifically, the standard video data is decomposed into multiple video clips using the shot segmentation technology. First, the shot switching points are identified. These switching points are usually determined by significant changes in the picture content, scene switching or significant changes in audio. In order to ensure the accuracy of the segmentation, advanced algorithms such as edge detection and motion analysis are used. When the change amplitude of the picture edge is detected to exceed a preset threshold, it is identified as a shot switching point; then, the standard video data is segmented at these switching points to generate multiple continuous and coherent video clips. Each clip retains important scene information in the original video, which is convenient for subsequent processing and analysis.

[0083] In one embodiment, see Figure 2 , said S3 comprises:

[0084] S31: Perform grayscale conversion on each frame image in the standard video data to obtain each frame grayscale image;

[0085] Specifically, first, each frame of the standard video data is gray-converted, that is, the color image is converted into a gray-scale image. Gray-scale conversion converts the RGB (red, green, blue) color information of the image into gray-scale values, and the gray-scale value is calculated by weighted average of the color value of each pixel. Specifically, the commonly used conversion formula is: gray-scale value = 0.299*R+0.587*G+0.114*B. This processing can simplify the image data, reduce the computational complexity, and highlight the structure and shape features in the image, which is convenient for subsequent edge detection.

[0086] S32: using an edge detection algorithm to perform edge detection on the grayscale image of each frame, and output edge feature information;

[0087] Specifically, next, edge detection algorithms such as Sobel, Canny, Laplacian, etc. are applied to process each frame of grayscale image to detect edge information in the image. The purpose of edge detection is to find areas in the image where pixel values ​​change significantly, which usually correspond to the boundaries and contours of objects. The edge detection algorithm calculates the image gradient or uses a differential operator to identify pixels in the image where grayscale values ​​change dramatically, and marks these pixels as edge points, thereby generating a binary image containing edge feature information.

[0088] S33: determining edge feature change difference according to edge feature information corresponding to grayscale images of adjacent frames;

[0089] Specifically, after edge detection is completed, the edge feature information of adjacent frames is compared to determine the change of edge features. The specific operation is to calculate the edge image difference between adjacent frames, and identify the area of ​​edge change in the image through pixel-level difference operation. The edge feature change difference is obtained by calculating the difference in edge pixel values ​​at corresponding positions in adjacent frames. These differences can be quantified as the amplitude of change, reflecting the dynamic change of scene content over time.

[0090] S34: determining video boundary information according to the edge feature change difference and a preset change threshold;

[0091] Specifically, the calculated edge feature change difference is compared with a preset change threshold. If the edge feature change difference between a frame and its adjacent frame exceeds the preset threshold, it is considered that there is a significant scene switch between the frame and the previous frame, that is, it is determined to be a video boundary. The preset change threshold is used to filter out subtle and insignificant changes to ensure that only obvious content changes (such as scene switches) are identified as video boundaries. In this way, the boundary information of each shot in the video can be accurately determined.

[0092] S35: Decomposing the standard video data according to the video boundary information to determine each of the video segments.

[0093] Specifically, finally, according to the determined video boundary information, the standard video data is segmented at these boundary points to generate multiple independent video clips, each of which represents a continuous scene or shot, preserving the coherence and integrity of the content. The decomposed video clips are convenient for subsequent processing and analysis, such as embedding blind watermarks, feature extraction, content annotation, etc. This method can effectively manage and process video data and improve the efficiency of video content analysis and processing.

[0094] Specifically, from grayscale conversion to edge detection, to the calculation of edge feature changes and the determination of video boundaries, the video is finally decomposed into multiple segments. This process realizes the refined processing and management of video content, laying a solid foundation for subsequent in-depth analysis and application.

[0095] S4: using a preset feature processing algorithm, embedding blind watermark information containing video material information into each frame image in each video clip;

[0096] Specifically, the blind watermark information containing the video material information is embedded into each frame image in each video clip using a preset feature processing algorithm, in order to ensure that the original video material can be accurately identified and associated in the subsequent film editing. The specific operations include: first, selecting a suitable feature processing algorithm, such as discrete cosine transform (DCT) or discrete wavelet transform (DWT), to convert each frame image into the frequency domain. Then, a specific frequency coefficient (usually a low-frequency or medium-frequency coefficient) is selected in the frequency domain for slight adjustment, and the blind watermark information is embedded therein. These embedded blind watermark information can contain key data such as a unique material identifier, ensuring that the image is slightly adjusted without affecting the visual quality of the video, so that the watermark information is invisible to the human eye, but can be extracted by a specific algorithm in the subsequent process. Finally, the image is restored to the spatial domain using an inverse transform to generate a frame image containing the blind watermark. In this way, each frame of each video clip carries unique identification information, achieving accurate material tracking and identification, and effectively solving the problem of positioning the video material in the film.

[0097] In one embodiment, see Figure 3 , said S4 comprises:

[0098] S41: converting the video material information to be embedded into a binary format, and determining the encoded blind watermark information;

[0099] Specifically, first, the video material information to be embedded (such as unique identifiers, copyright information, etc.) is converted into binary format, and the text or other forms of information are converted into a binary data stream through an encoding algorithm, such as ASCII encoding or UTF-8 encoding. After the information is converted into binary format, it is encrypted or hashed to generate the final blind watermark information. The purpose of this processing is to ensure that the information remains stable and secure during the embedding process, so as to facilitate subsequent embedding and extraction in the frequency domain image.

[0100] S42: Decompose each frame image in each video segment to obtain a regional image;

[0101] Specifically, each frame of each video clip is decomposed, usually by dividing the image into several small blocks (such as 8x8 or 16x16 macroblocks) to facilitate more fine-grained frequency domain processing. The decomposition process allows each small block to be processed independently, ensuring that the embedding process has the least impact on the overall image. These decomposed regional images can better analyze and process local features, which is conducive to improving the embedding efficiency and concealment of blind watermarks.

[0102] S43: performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image;

[0103] Specifically, after obtaining the regional image, discrete cosine transform (DCT) is applied to each regional image to convert it from the spatial domain to the frequency domain. DCT decomposes the image data into different frequency components so that the energy of the image is concentrated on a few low-frequency components. These low-frequency components often correspond to the main features of the image. Through this transformation, blind watermark information can be more effectively embedded in the frequency domain without significantly affecting the visual quality of the original image.

[0104] S44: adjusting the high frequency components in the frequency domain image, and embedding the blind watermark information into the frequency domain image;

[0105] Specifically, high-frequency components are selected in the frequency domain image for adjustment and the blind watermark information is embedded therein. The specific method is to make slight adjustments to the selected DCT coefficients according to the binary blind watermark information, such as increasing or decreasing the value of a specific coefficient to represent binary 0 or 1. The low-frequency components contain the main information of the image, and slight adjustments will not have a significant impact on the image quality, while the mid-frequency components balance the concealment and robustness of information embedding. In this way, the blind watermark information is securely and covertly embedded in the frequency domain characteristics of the image.

[0106] S45: performing inverse discrete cosine transformation on the frequency domain images and then recombining them to complete embedding of blind watermark information in each video segment.

[0107] Specifically, finally, the adjusted frequency domain image is subjected to inverse discrete cosine transform (IDCT) to convert it back to the spatial domain image and restore the visual effect of the original image. After IDCT processing, each regional image is reassembled back into the original frame image to ensure that each frame retains the blind watermark information completely. In this way, the blind watermark information is seamlessly embedded into each frame image in the video clip, and the stable embedding and subsequent extractability of the information are ensured without affecting the overall visual quality. Each video clip processed in this way carries unique identification information, realizing accurate tracking and protection of the material.

[0108] S5: compressing each video segment after the blind watermark information is embedded to determine compressed video data;

[0109] Specifically, each video clip after the blind watermark information is embedded is compressed to determine the compressed video data, the original film is stored in the company's intranet environment, and a low-resolution version of the video is compressed and stored in the cloud for labeling and analysis. Video compression technology significantly reduces the size of video files by optimizing encoding and reducing resolution, thereby reducing storage and transmission costs. The original film is stored in the intranet to ensure the high quality of the video and data security, and prevent unauthorized access. The low-resolution version of the video is sufficient to meet the needs of content recognition and annotation, and its smaller file size makes cloud storage and access more efficient. Storing low-resolution videos in the cloud enables remote access and sharing, facilitates collaboration and data processing among team members in different locations, improves work efficiency and flexibility, and ensures data security and consistency. This solution takes into account the quality, security and ease of use of video data through efficient compression and distributed storage, providing the team with a flexible working environment and strong technical support.

[0110] S6: Using a preset feature extraction algorithm, the compressed video data is identified and analyzed, and text information related to the video content and target video material information are output.

[0111] Specifically, the compressed video data is identified and analyzed using a preset feature extraction algorithm, and text information and target video material information related to the video content are output. This process first analyzes the compressed video frame by frame through advanced feature extraction algorithms, such as multimodal large language models, speech transcription models, face recognition models, target detection models, and color analysis models. These algorithms can recognize various contents in the video, such as text, speech, faces, objects, and colors. Then, the identified information is integrated and analyzed to generate detailed text descriptions and annotations, including scene descriptions, dialogue content, character information, and object locations. This information can not only be used for understanding and retrieving video content, but also help generate accurate identification and descriptions related to video materials, ensuring that these materials can be accurately associated and used in subsequent use. This process greatly improves the efficiency and accuracy of video content analysis, and provides strong technical support for content management and application.

[0112] In one embodiment, see Figure 4 , the S6 comprises:

[0113] S61: performing blind watermark information recognition on the compressed image of each frame, and outputting the recognition result;

[0114] Specifically, the extraction algorithm corresponding to the preset blind watermark embedding algorithm is used to process each frame of the compressed video data, extract the embedded blind watermark information, and determine the target video material information based on the extraction result. This step ensures that even in the case of compressed video, the source of the video clip and its specific content can still be accurately located and identified, thereby ensuring the traceability of the material and the effectiveness of content management.

[0115] In one embodiment, see Figure 5 , the S61 comprises:

[0116] S611: Inputting the compressed image of each frame into a pre-trained self-supervised visual transformation model, and outputting encoding feature information;

[0117] Specifically, first, each frame of compressed image is input into a pre-trained self-supervised visual transformation model (such as DINO-V2). The model learns rich visual features from a large amount of unlabeled data through self-supervised learning. The model processes the input image, extracts the key features of the image, and encodes these features into high-dimensional feature vectors. These feature vectors contain a variety of information such as the texture, edge, shape, and color of the image, and can accurately describe the image content. The output encoded feature information will serve as the basis for subsequent feature matching.

[0118] S612: using an approximate nearest neighbor algorithm, performing feature matching between the encoding feature information and feature template information of each video material template in a video material database, and outputting a matching result;

[0119] Specifically, the approximate nearest neighbor algorithm (ANN) is then used to perform feature matching on the coded feature information. The approximate nearest neighbor algorithm can quickly find the feature template that is most similar to the target feature in a high-dimensional feature space. The coded feature information of each frame of the image is compared with the feature template information of each video material template pre-stored in the video material database to calculate the similarity. Through the ANN algorithm, the feature template closest to the coded feature information is found, and the matching result is output. The matching result includes the matching degree and related information of each frame of the image with the most similar material template in the database.

[0120] S613: According to the matching result, the video material template corresponding to the feature template information matching the encoding feature information is output as the target video material information.

[0121] Specifically, finally, based on the feature matching results, the target video material information is determined. For each frame of the image, the feature template information with the highest matching degree is selected, and its corresponding video material template is output as the target video material information, ensuring that even in the absence of blind watermark information, the source and specific content of the video clip can be accurately identified through feature matching. In this way, each frame of the image can be efficiently and accurately associated with the original video material, realizing accurate identification and management of video content.

[0122] S62: if the blind watermark information is identified, obtaining the video material information embedded in the blind watermark information as the target video material information;

[0123] Specifically, after successfully identifying the blind watermark information, the embedded video material information is extracted from the identified watermark data. The blind watermark usually contains a unique material identifier, copyright information or other related metadata. Through this information, the source and specific content of the video material can be accurately determined and used as the target video material information to ensure that each frame of the image can be accurately associated with the corresponding original material.

[0124] S63: If the blind watermark information is not identified, extract and match the features of each frame of the compressed image, and determine the target video material information according to the matching result;

[0125] Specifically, if the blind watermark information cannot be identified, the compressed image is further analyzed using feature extraction and matching methods to extract features including texture, edge, shape, color and other information to form a unified image feature code. Then, the extracted image features are matched with the material features in the database to determine the most similar material, ensuring that the source and content of the video clip can be accurately identified even in the absence of a blind watermark. Based on the matching results, the target video material information is determined and output, thereby achieving accurate identification and management of the material.

[0126] S64: inputting each frame of compressed image into a pre-trained feature extraction model, and outputting key feature information;

[0127] Specifically, each frame of compressed image is input into a pre-trained feature extraction model to output key feature information. This process involves processing each frame of image through a trained feature extraction model in the field of computer vision, such as convolutional neural network CNN and Transformer model. The model has been trained on a large amount of image data and can automatically extract important visual features in the image, such as edges, textures, color distribution and shapes. After the image is input, the model will extract high-level features in the image layer by layer through a series of convolutional layers, activation functions and pooling operations. These features will be encoded into a high-dimensional feature vector that contains the core information of the image. These encoded key feature information can not only accurately describe the image content, but also can be used for subsequent feature matching, classification or retrieval tasks to ensure the effective use and analysis of image data.

[0128] In one embodiment, see Figure 6 , the S64 comprises:

[0129] S641: Decode the compressed video data to obtain audio data;

[0130] Specifically, first, the compressed video data is decoded to extract the audio data. This process involves restoring the compressed video file to its original uncompressed state and extracting the audio portion of the video stream through a decoder. The decoding process usually involves converting the compressed encoded data stream (such as MP4 or H.264 format) into an uncompressed audio format (such as WAV or PCM) that can be further processed. The extracted audio data contains all the sound information in the video, including dialogue, background music, and ambient sound, providing a basis for subsequent audio analysis and processing.

[0131] S642: Input the compressed image into a pre-trained face recognition classification model, classify and label the facial features in the recognized compressed image, and determine the person information;

[0132] S643: Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting main color information, wherein the main color information at least includes the advertising brand color tone detected in the compressed image or the main color tone extracted from the advertising landscape painting;

[0133] S644: Input the compressed image into the target detection model, locate and classify the objects in the compressed image, and determine the object information, wherein the object information at least includes the object category and the object location;

[0134] Specifically, the compressed image is respectively input into a pre-trained face recognition classification model, a color analysis model, and an object detection model, and outputs personnel information, color information, and object information. Specifically, this step involves passing each frame of the compressed image to multiple specially trained deep learning models to extract and analyze different types of information. First, the image is input into a face recognition classification model, which is trained using a large amount of pre-labeled facial data and can accurately identify facial features in the image, classify and label them, for example, identifying celebrities, actors or other known people appearing in the image, such as identifying the facial features of a reporter in a news report. Secondly, the image is input into a color analysis model, which analyzes the color distribution in the image and extracts the main color information, such as the brand tone detected in the advertising image or the main tone extracted in the landscape painting. Finally, the image is input into the object detection model, which determines the category and location of the object in the image by locating and classifying it, such as identifying the goods in the shopping advertisement or the traffic signs in the video. These models each process different aspects of an image, and by combining this information, a comprehensive description of the image content can be obtained, such as identifying actors, background colors, and major objects in a scene in a movie clip, providing key details for in-depth understanding and analysis of videos.

[0135] S645: Input the audio data into a pre-trained speech transcription model, and output audio feature information in the audio data;

[0136] Specifically, the extracted audio data is input into a pre-trained speech transcription model to convert the speech content in the audio into text information. The speech transcription model uses natural language processing technology and speech recognition algorithms to identify the language, intonation, and syllables in the audio and transcribe it into text. This process can recognize and record conversations, speeches, or other voice information and generate corresponding text descriptions. The audio feature information includes the text form of the speech content, so that the information in the audio can be further analyzed and processed.

[0137] S646: Input the personnel information, color information, object information and audio feature information into a multimodal large language model respectively, and output the key feature information.

[0138] Specifically, the person information, color information, object information extracted from the image, and the audio feature information extracted from the audio are input into the multimodal large language model. The multimodal large language model can fuse data from different sources and understand and analyze the relationship between various types of information. The model will conduct a comprehensive analysis of these multimodal data to generate key feature information that describes the video content. This includes integrating the visual features in the image and the speech content in the audio to provide a comprehensive understanding and summary of the content. This fusion analysis method enables the various information extracted from the video to be effectively integrated and interpreted, generating deep insights into the video content.

[0139] S65: Input the key feature information into a multimodal large language model, and output the text information.

[0140] Specifically, the key feature information is input into the multimodal large language model, and the text information is output. First, the key feature information extracted from the image (such as personnel information, color information, object information) and audio (such as speech transcription text) is integrated into a unified data format. These feature information will be input into the multimodal large language model, which uses deep learning technology to fuse data of different modalities (visual, audio, text, etc.) to understand the relationship between these information. By analyzing and contextualizing the input multimodal data, the model can generate a coherent text description that accurately expresses the main points of the video content. The output text information usually includes a summary of the video scene, the content of the dialogue, a detailed description of the characters and objects, and other relevant information. This comprehensive analysis capability enables the multimodal large language model to provide a comprehensive and accurate summary of the video content, enhancing the understanding and utilization of video data.

[0141] Example 2

[0142] In the above-mentioned embodiment 1, the analysis results related to the content of each video clip are obtained through content analysis, including personnel information, color information, object information and audio feature information. In the real advertising video editing scene, the relevance of the video clips is evaluated according to the content analysis results, and the video clips are edited to form a video collection, which can greatly improve the quality of the video and the viewing experience. For this reason, in one implementation, please refer to Figure 7 , after S6, further comprising:

[0143] Obtaining the scene type of the advertisement to be placed, wherein the scene type includes: introduction scene, product display scene, user experience scene and problem solving scene;

[0144] Specifically, the scene type of the advertisement to be placed is obtained. Advertisements are usually divided into multiple scene types, such as introduction scenes, product display scenes, user experience scenes, and problem-solving scenes. Introduction scenes are usually used to attract the audience's attention and present the core information or brand of the advertisement; product display scenes focus on displaying the functions, appearance, and usage methods of the product to help the audience intuitively understand the value of the product; user experience scenes show how users use the product, as well as the experience and feedback obtained during use; problem-solving scenes are used to show how the product helps users solve practical problems, emphasizing the effectiveness and practicality of the product. The step of obtaining scene types can be completed by analyzing the content of advertising planning, brand requirements, and market demand. The advertising editing team usually presets these scene types to ensure that the advertisement can be displayed according to the set plot or marketing strategy. At the same time, video analysis technology is used to extract key features in the video (such as products, characters, and scene elements), and these features are matched with predefined scene types to quickly determine the scene category to which the video clip belongs.

[0145] Based on the scene type, using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity to determine the target similarity corresponding to each video clip;

[0146] Specifically, after confirming the scene type, the video clips are further analyzed based on these scene types, and the similarity between the clips and the target scene is determined by a text matching algorithm. First, the content of the video clip is usually converted into text form through a speech transcription model or subtitle extraction technology. These text information reflects the core content of the video clip, such as product name, function description, user experience, etc. Using text matching algorithms, such as TF-IDF, BERT, etc., the text information in the video clip is calculated with the preset scene text template. Each scene type has a set of preset text templates, and the text template contains keywords or sentence patterns related to the scene. The core of the similarity evaluation is to calculate the semantic similarity between the text in the video clip and the template text, and then evaluate the matching degree between the video clip and the scene type.

[0147] In one embodiment, based on the scene type, using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and determining the target similarity corresponding to each video clip includes:

[0148] Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined respectively;

[0149] Specifically, the text matching algorithm is first used to process the text information related to the video clip, which includes personnel information, color information, object information and audio feature information. The text information is evaluated for similarity with the preset text template. The text template usually represents a specific scene type, such as an introduction scene, a product display scene, etc., and contains keywords and related sentences. The semantic similarity between the video text information and the template is calculated through an algorithm, such as TF-IDF or BERT, and the first similarity, second similarity, third similarity and fourth similarity corresponding to the personnel information, color information, object information and audio feature information are determined respectively. These similarities reflect the degree of matching between each video clip and the relevant features in the scene.

[0150] Obtaining a preset weight correction factor, wherein the weight correction factor is greater than 1;

[0151] Specifically, in order to adjust the evaluation of different similarities, a preset weight correction factor is obtained. The correction factor is usually greater than 1, indicating that some similarities need to be amplified. The weight correction factor is set according to specific scene requirements. For example, in some scenes, color information may be more important than other features, so the weight of color similarity needs to be increased. The weight correction factor is to dynamically adjust the importance of different similarities so that the final result can better meet the scene requirements and advertising planning goals.

[0152] Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1;

[0153] Specifically, after similarity calculation, an initial weight is assigned to each similarity, which determines the influence of each similarity in the total similarity calculation, and the sum of all initial weights is equal to 1. Depending on the advertising planning goals, the allocation of initial weights can be preset to different values. For example, a product display scenario may give a higher weight to item information, while a user experience scenario may give a higher weight to personnel information.

[0154] If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0155] Specifically, in the introduction scene, the color information and audio feature information of the video are usually key factors in attracting the audience's attention. Therefore, the weight correction factor is used to correct the similarities related to these features, namely the second similarity and the fourth similarity. The correction process includes multiplying the weight correction factor by the second similarity to obtain a new second similarity, and multiplying the weight correction factor by the fourth similarity to obtain a new fourth similarity. This means that the initial weights of color and audio will be amplified, so that these features occupy a larger proportion in the similarity evaluation, ensuring that the visual and auditory effects of the introduction scene are more attractive.

[0156] If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor;

[0157] Specifically, in the product display scene, the color information and item information in the video are particularly important, so the weight correction factor is used to correct the second and third similarities. This adjustment ensures that the visual effects and item details of the product are more prominent during display, which can attract the audience's attention and prompt them to focus on the characteristics of the product.

[0158] If the scenario type is a user experience scenario, using the weight correction factor to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity;

[0159] Specifically, the user experience scenario focuses on showing the interactive experience between people and products, so the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor. In this type of scenario, the weights of these two types of information are amplified to ensure that the interactive process between users and products and the display of the product itself are more prominent, thereby enhancing the user's sense of involvement.

[0160] If the scenario type is a problem-solving scenario, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor;

[0161] Specifically, in problem-solving scenarios, it is usually necessary to highlight personal information, item information, and audio feature information to show how the product solves the user's actual problem. Therefore, the weight correction factor is used to correct the first similarity, the third similarity, and the fourth similarity, so that it can better show how the product solves the user's troubles through interaction with the user, the problem-solving process, and related audio prompts, and enhance the persuasiveness and authenticity of the scene.

[0162] According to the corrected initial weights, the first similarity, the second similarity, the third similarity and the fourth similarity are weighted averaged to determine the target similarity;

[0163] Specifically, the first similarity, the second similarity, the third similarity, and the fourth similarity obtained previously are processed, and each similarity represents the degree of matching of different types of information. In the previous steps, each similarity is weighted according to the scene requirements to ensure that the content that the scene attaches importance to is highlighted. The weighted average processing refers to multiplying the corrected initial weights by the corresponding similarity values, and then adding and averaging these results to calculate the overall target similarity. The target similarity indicates the degree of matching between the video clip and the specific scene after integrating multimodal information. Through this process, the system can comprehensively evaluate the fit of each video clip with the preset scene under multimodal information.

[0164] According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlight.

[0165] Specifically, after obtaining the target similarity, the value is compared with a preset similarity threshold. The similarity threshold is a standard pre-set by the system to determine whether a video clip meets the minimum requirements of the scene. This threshold can be adjusted according to the needs of different scenes. For example, a product display scene may require a higher similarity threshold, while an introduction scene may allow a lower threshold. The similarity threshold is usually a standard set based on user expectations or advertising goals. If the target similarity is greater than or equal to this threshold, it means that the video clip has a high relevance and can well match the needs of the scene and is suitable for the next step of processing. Those clips below the threshold will be considered to be mismatched with the current scene and thus excluded from the selection of video highlights. After determining which video clips meet the similarity requirements, the system will synthesize these clips. The synthesis process includes combining multiple video clips with high similarity according to certain logic or strategies, such as time sequence, plot development, visual consistency, etc., to form a complete advertising collection. This step can be adjusted according to the design of the marketing strategy, such as combining product display clips with user experience clips, or arranging problem solving clips with product display clips in an orderly manner to ensure that the collection has a coherent narrative structure.

[0166] Example 3

[0167] See also Figure 8 Embodiment 3 of the present invention further provides a video content analysis device based on multimodal data processing, the device comprising:

[0168] A video material acquisition module is used to acquire the original video material to be parsed;

[0169] A transcoding processing module, used for performing transcoding processing on the original video material to determine standard video data in a preset coding format;

[0170] A shot segmentation module, used to decompose the standard video data into multiple video segments by using a shot segmentation technology;

[0171] A blind watermark embedding module is used to embed blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm;

[0172] A compression processing module, used to compress each video segment after the blind watermark information is embedded, and determine the compressed video data;

[0173] The feature extraction module is used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information.

[0174] Specifically, a video content analysis device based on multimodal data processing provided by an embodiment of the present invention is adopted, and the device includes: a video material acquisition module, used to acquire the original video material to be analyzed; a transcoding processing module, used to perform transcoding processing on the original video material, and determine standard video data of a preset encoding format; a shot segmentation module, used to decompose the standard video data into multiple video segments using a shot segmentation technology; a blind watermark embedding module, used to embed blind watermark information containing video material information into each frame image in each video segment using a preset feature processing algorithm; a compression processing module, used to perform compression processing on each video segment after the blind watermark information is embedded, and determine compressed video data; a feature extraction module, used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information. This device achieves accurate association between finished films and materials, and accurately identifies and analyzes video content through a series of orderly processing steps. First, by acquiring and transcoding the original video material, all data formats are ensured to be unified. Then, the standard video data is decomposed into multiple easy-to-manage segments using shot segmentation technology, and blind watermark information is embedded, so that each material segment can be accurately identified in the finished film, even after editing and compression. Finally, the compressed video is identified and analyzed through a feature extraction algorithm, and detailed text information related to the video content and target video material information are output. This method not only improves the matching accuracy of video materials and finished films, ensuring that each piece of material can be accurately located and identified, but also can efficiently extract and summarize video content, provide more detailed analysis and annotation, and solve the efficiency and accuracy problems in traditional video processing and analysis.

[0175] Example 4

[0176] In addition, combined Figure 1 The video content analysis method based on multimodal data processing of the embodiment 1 of the present invention described above can be implemented by an electronic device. Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in Embodiment 4 of the present invention is shown.

[0177] The electronic device may include a processor and a memory storing computer program instructions.

[0178] Specifically, the processor may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present invention.

[0179] The memory may include a large capacity memory for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory may be inside or outside a data processing device. In a specific embodiment, the memory is a non-volatile solid-state memory. In a specific embodiment, the memory includes a read-only memory (ROM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM) or a flash memory or a combination of two or more of these.

[0180] The processor implements any one of the video content analysis methods based on multimodal data processing in the above embodiments by reading and executing computer program instructions stored in the memory.

[0181] In one example, the electronic device may further include a communication interface and a bus. Fig. 9 As shown, the processor, memory, and communication interface are connected via a bus and communicate with each other.

[0182] The communication interface is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiments of the present invention.

[0183] The bus includes hardware, software or both, and the parts of the device are coupled to each other. For example, but not limitation, the bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In suitable cases, the bus may include one or more buses. Although the embodiment of the present invention describes and shows a specific bus, the present invention considers any suitable bus or interconnection.

[0184] Example 5

[0185] In addition, in combination with the video content analysis method based on multimodal data processing in the above embodiment 1, embodiment 5 of the present invention can also provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the video content analysis methods based on multimodal data processing in the above embodiments is implemented.

[0186] In summary, the embodiments of the present invention provide a method, apparatus and device for video content analysis based on multimodal data processing.

[0187] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.

[0188] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0189] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant location, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0190] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in a different order from the embodiments, or several steps can be performed simultaneously.

[0191] The above is only a specific implementation of the present invention. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system, module and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be covered within the protection scope of the present invention.

Claims

1. A video content analysis method based on multimodal data processing, characterized in that: The method comprises: S1: Obtain the original video material to be parsed; S2: performing transcoding processing on the original video material to determine standard video data in a preset encoding format; S3: Decomposing the standard video data into multiple video segments by using a shot segmentation technology; S4: using a preset feature processing algorithm, embedding blind watermark information containing video material information into each frame image in each video clip; S5: compressing each video segment after the blind watermark information is embedded to determine compressed video data; S6: using a preset feature extraction algorithm to identify and analyze the compressed video data, and output text information related to the video content and target video material information; The S6 includes: S61: performing blind watermark information recognition on each compressed image frame in the compressed video data, and outputting the recognition result; S62: if the blind watermark information is identified, obtaining the video material information embedded in the blind watermark information as the target video material information; S63: If the blind watermark information is not identified, extract and match the features of each frame of the compressed image, and determine the target video material information according to the matching result; S64: inputting each frame of compressed image into a pre-trained feature extraction model, and outputting key feature information; S65: Input the key feature information into a multimodal large language model, and output the text information.

2. The video content analysis method based on multimodal data processing according to claim 1, characterized in that: The S3 includes: S31: Perform grayscale conversion on each frame image in the standard video data to obtain each frame grayscale image; S32: using an edge detection algorithm to perform edge detection on the grayscale image of each frame, and output edge feature information; S33: determining edge feature change difference according to edge feature information corresponding to grayscale images of adjacent frames; S34: determining video boundary information according to the edge feature change difference and a preset change threshold; S35: Decomposing the standard video data according to the video boundary information to determine each of the video segments.

3. The video content analysis method based on multimodal data processing according to claim 1, characterized in that: The S4 includes: S41: converting the video material information to be embedded into a binary format, and determining the encoded blind watermark information; S42: Decompose each frame image in each video segment to obtain a regional image; S43: performing discrete cosine transform on the regional image to convert the regional image from a spatial domain image to a frequency domain image; S44: adjusting the high frequency components in the frequency domain image, and embedding the blind watermark information into the frequency domain image; S45: performing inverse discrete cosine transformation on the frequency domain images and then recombining them to complete embedding of blind watermark information in each video segment.

4. The video content analysis method based on multimodal data processing according to claim 1, characterized in that: The S63 includes: S631: Inputting the compressed image of each frame into a pre-trained self-supervised visual transformation model, and outputting encoding feature information; S632: using an approximate nearest neighbor algorithm, performing feature matching between the encoded feature information and feature template information of each video material template in a video material database, and outputting a matching result; S633: According to the matching result, the video material template corresponding to the feature template information matching the encoding feature information is output as the target video material information.

5. The video content analysis method based on multimodal data processing according to claim 1, characterized in that: The S64 includes: S641: Decode the compressed video data to obtain audio data; S642: Input the compressed image into a pre-trained face recognition classification model, classify and label the facial features in the recognized compressed image, and determine the person information; S643: Inputting the compressed image into a pre-trained color analysis model, analyzing the color distribution in the compressed image, and extracting main color information, wherein the main color information at least includes the advertising brand color tone detected in the compressed image or the main color tone extracted from the advertising landscape painting; S644: Input the compressed image into the target detection model, locate and classify the objects in the compressed image, and determine the object information, wherein the object information at least includes the object category and the object location; S645: Input the audio data into a pre-trained speech transcription model, and output audio feature information in the audio data; S646: Input the personnel information, color information, object information and audio feature information into a multimodal large language model respectively, and output the key feature information.

6. The video content analysis method based on multimodal data processing according to claim 1, characterized in that: After S6, the method further includes: Obtaining the scene type of the advertisement to be placed, wherein the scene type includes: introduction scene, product display scene, user experience scene and problem solving scene; Using a preset text matching algorithm, the text information related to the video clip content and the preset text template are evaluated for similarity, and a first similarity, a second similarity, a third similarity, and a fourth similarity corresponding to the person information, the color information, the object information, and the audio feature information are determined respectively; Obtaining a preset weight correction factor, wherein the weight correction factor is used to amplify the similarity, and the weight correction factor is greater than 1; Obtaining initial weights corresponding to the first similarity, the second similarity, the third similarity, and the fourth similarity, respectively, wherein the sum of the initial weights is equal to 1; If the scene type is an introduction scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor; If the scene type is a product display scene, the initial weight corresponding to the second similarity and the initial weight corresponding to the third similarity are corrected using the weight correction factor; If the scenario type is a user experience scenario, using the weight correction factor to correct the initial weight corresponding to the first similarity and the initial weight corresponding to the third similarity; If the scenario type is a problem-solving scenario, the initial weight corresponding to the first similarity, the initial weight corresponding to the third similarity, and the initial weight corresponding to the fourth similarity are corrected using the weight correction factor; According to the corrected initial weights, the first similarity, the second similarity, the third similarity and the fourth similarity are weighted averaged to determine the target similarity; According to the target similarity and a preset similarity threshold, video clips whose target similarity is greater than the similarity threshold are synthesized into a video highlight.

7. A video content analysis device based on multimodal data processing, characterized in that: The device comprises: A video material acquisition module is used to acquire the original video material to be parsed; A transcoding processing module, used for performing transcoding processing on the original video material to determine standard video data in a preset coding format; A shot segmentation module, used to decompose the standard video data into multiple video segments by using a shot segmentation technology; A blind watermark embedding module is used to embed blind watermark information containing video material information into each frame image in each video clip using a preset feature processing algorithm; A compression processing module, used to compress each video segment after the blind watermark information is embedded, and determine the compressed video data; A feature extraction module, used to identify and analyze the compressed video data using a preset feature extraction algorithm, and output text information related to the video content and target video material information; The method of using a preset feature extraction algorithm to identify and analyze the compressed video data and outputting text information related to the video content and target video material information includes: Perform blind watermark information recognition on each frame of compressed image in the compressed video data and output the recognition result; If the blind watermark information is identified, the video material information embedded in the blind watermark information is obtained as the target video material information; If the blind watermark information is not identified, feature extraction and matching are performed on each frame of the compressed image, and the target video material information is determined based on the matching result; Input each frame of compressed image into the pre-trained feature extraction model and output key feature information; The key feature information is input into a multimodal large language model, and the text information is output.

8. An electronic device, characterized in that: include: At least one processor, at least one memory and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, implement the method according to any one of claims 1 to 6.

9. A storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • News video information extraction method for global deep learning

    CN112004111A

  • Method for resisting video compression robustness blind watermarking based on DST

    CN115150627A

  • Intelligent image-text to video conversion method and system based on video structured data

    CN115272533A