Video content understanding method of multi-mode advertisement inventory intelligent matching system
Through lens segmentation and dynamic search window technology, the asynchronous problem of multimodal signal timing is solved, accurate alignment and efficient semantic analysis of multimodal signals are realized, and the accuracy and calculation efficiency of advertising matching are improved.
Patent Information
- Application Number
- CN202510880057.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Traditional video content understanding methods fail to effectively deal with the timing asynchronous problem of multimodal signals, resulting in cross-modal feature misalignment, affecting the accuracy and efficiency of semantic analytics of advertising matching.
Through lens segmentation and dynamic search window technology, the similarity between video frames, speech segments and text segments is analyzed, the dynamic balance factor and search window are determined, the precise alignment of multimodal signals is achieved, and multi-level semantic analysis is performed through deep learning models.
It improves the matching accuracy and efficiency of multimodal data, reduces computing resource consumption, and enhances the accuracy of content understanding and user experience.
Smart Images

Figure CN120388324A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video content understanding, and specifically to a video content understanding method for a multi-modal advertising inventory intelligent matching system. Background Art
[0002] Multi-modal advertising is an advertising form that integrates multiple media forms such as vision, audition, and text, and conveys brand value and product characteristics by fusing multi-dimensional information. The multi-modal advertising inventory intelligent matching system is an automated platform built based on artificial intelligence technology. Its core function is to accurately match a large amount of advertising inventory with media-side video traffic. By analyzing the multi-modal characteristics of video content, it realizes the efficient docking of advertisements, content, and users, and improves the relevance and conversion efficiency of advertising placement.
[0003] As the core carrier of multi-modal advertising, the unstructured information contained in video needs to be converted into computable structured features in order to perform semantic matching with the label system of the advertising inventory, avoiding the decline of user experience and waste of placement budget caused by the disconnection between advertisements and video scenes. Traditional video content understanding methods usually first extract visual, audio, and text features separately, then perform feature fusion through a model, and finally perform semantic parsing and matching by relying on keyword matching or a rule engine.
[0004] However, traditional methods do not consider the problem of temporal asynchrony that often exists in multi-modal signals such as visual images, audio voices, and text captions in videos due to reasons such as post-editing and device acquisition differences. It usually assumes that multi-modal signals are synchronous. If processed directly without alignment, it will cause cross-modal feature misalignment, resulting in deviation in semantic parsing, and further causing the matching logic of advertisement context orientation to fail. Summary of the Invention
[0005] In view of the above, it is necessary to provide a video content understanding method for a multi-modal advertising inventory intelligent matching system to solve the above problems.
[0006] An embodiment of this application provides a video content understanding method for a multi-modal advertising inventory intelligent matching system, and the method includes: Performing video frame extraction and shot segmentation on the original advertising video; extracting audio signals from the original advertising video to obtain speech segments; reading text segments from the original advertising video; Dividing the video frames, speech segments, and text segments corresponding to the corresponding timestamps based on the start and end times of each shot segmentation; Taking the video frames as a reference, analyzing the similarity between adjacent video frames and the dispersion degree of the spectral information corresponding to the speech segments in the shot where the timestamp corresponding to each video frame is located, determining the dynamic balance factor of the timestamp corresponding to each video frame, and adjusting the preset reference window size of the timestamp corresponding to each video frame to obtain the first search window; In the shot corresponding to the timestamp of each video frame, analyze the occurrence frequency and duration of the text segment corresponding to the timestamp per unit time, confirm the text dynamic factor of the timestamp corresponding to each video frame, and adjust the preset reference window size of the timestamp corresponding to each video frame to obtain a second search window; Based on the first search window and the second search window, analyze the feature similarity degree between the speech segment and the text segment and the video frame within the corresponding window to obtain a first sparse matrix and a second sparse matrix, and obtain a first optimal alignment path and a second optimal alignment path; Based on the obtained optimal alignment path, align the timestamps of the speech segment and the text segment with the video frame timestamp, integrate the information of the video frame, speech segment, and text segment with aligned timestamps through multimodal fusion technology, and implement multi-level semantic parsing through a deep learning model.
[0007] Preferably, the dynamic balance factor of the timestamp corresponding to each video frame is specifically: Obtain the average value of the cosine similarity between all adjacent video frames in the shot corresponding to the timestamp of each video frame as the visual motion intensity of the timestamp corresponding to each video frame; Extract the Mel frequency spectrum matrix of each speech segment, and use the sequence composed of the Mel frequency spectrum matrices of all speech segments as the audio feature sequence; Obtain the standard deviation of each element in the audio feature sequence corresponding to the shot where the timestamp of each video frame is located and calculate the mean value to obtain the audio energy change rate of the timestamp corresponding to each video frame; Calculate the absolute value of the difference and the sum value between the visual motion intensity and the audio energy change rate corresponding to each video frame respectively, and perform positive fusion on the negative correlation mapping result of the sum value and the absolute value of the difference to obtain the dynamic balance factor of the timestamp corresponding to each video frame.
[0008] Preferably, the adjustment of the preset reference window size of the timestamp corresponding to each video frame to obtain the first search window is specifically calculated by the formula: , where represents the width of the first search window of the timestamp corresponding to the t-th video frame; t represents the timestamp corresponding to the t-th video frame, represents the preset reference window size; represents the Sigmoid function; represents the floor value; represents the visual motion intensity of the timestamp corresponding to the t-th video frame; represents the audio energy change rate of the timestamp corresponding to the t-th video frame; represents the dynamic balance factor of the timestamp corresponding to the t-th video frame.
[0009] Preferably, the text dynamic factor for confirming the timestamp corresponding to each video frame is specifically as follows: Taking the sequence composed of all text segments as the text segment sequence; Obtaining the occurrence frequency of the text segment corresponding to the timestamp within the shot corresponding to the timestamp of each video frame per unit time, and the duration of the text segment corresponding to the timestamp, multiplying the two to obtain a first product; Obtaining the maximum occurrence frequency and the maximum duration of the text segment corresponding to the timestamp in the text segment sequence, multiplying them to obtain a second product; Taking the result of positively fusing the negative correlation mapping result of the second product with the first product as the text dynamic factor for the timestamp corresponding to each video frame.
[0010] Preferably, adjusting the preset reference window size for the timestamp corresponding to each video frame to obtain a second search window, specifically as follows: In the shot corresponding to the timestamp of each video frame, analyzing the similarity degree of the feature information between the video frame and the text segment to determine the semantic similarity of the timestamp corresponding to each video frame; Denoting the width of the second search window for the timestamp corresponding to the t-th video frame as , and its formula form is: , where: represents; represents the semantic similarity of the timestamp corresponding to the t-th video frame; represents a preset activation factor; represents an activation function; represents the floor value; represents the text dynamic factor of the timestamp corresponding to the t-th video frame; represents the preset reference window size.
[0011] Preferably, the process for specifically obtaining the semantic similarity is as follows: For each video frame, extracting a visual embedding vector through a pre-trained visual model, and taking the sequence composed of all visual embedding vectors as the visual feature vector; Generating a context semantic vector for each text segment, and taking the sequence composed of all context semantic vectors as the text segment feature sequence; Calculating the mean value of the cosine similarities of all elements at the same positions of the text segment feature sequence and the visual feature sequence within the shot corresponding to the timestamp of each video frame as the semantic similarity of the timestamp corresponding to each video frame.
[0012] Preferably, obtaining the first sparse matrix and the second sparse matrix is specifically as follows: For each video frame, traverse the speech segments within its first search window, and calculate the cosine similarity between the elements of the visual feature sequence within the window and the elements of the audio feature sequence, which serves as the elements of the first sparse matrix; For each video frame, traverse the text segments within its second search window, and use the normalized Euclidean distance between the elements of the visual feature sequence within the window and the elements of the text segment feature sequence as the elements of the second sparse matrix.
[0013] Preferably, the first optimal alignment path is obtained by applying the optimal path algorithm to the first sparse matrix; the second optimal alignment path is obtained by applying the optimal path algorithm to the second sparse matrix.
[0014] Preferably, the first optimal path is specifically the optimal alignment path composed of the timestamp indexes of the video frame and the speech segment; the second optimal path is specifically the optimal alignment path composed of the timestamp indexes of the video frame and the text segment.
[0015] Preferably, based on the obtained optimal alignment path, align the timestamps of the speech segment and the text segment with the timestamp of the video frame, specifically: Based on the optimal alignment path, map the timestamps of the speech segment and the text segment to the time axis of all video frames respectively; If the timestamps of the speech segment and the text segment in the optimal alignment path exactly match the video frame, directly take the corresponding speech segment and text segment; if the timestamps do not match, generate new speech segments and text segments through linear interpolation to obtain the speech segment and text segment aligned with the timestamp of the video frame.
[0016] This application has at least the following beneficial effects: This application is based on the timestamps of video frames. In each shot where the timestamp of a video frame is located, it analyzes the similarity between adjacent video frames and the dispersion degree of the corresponding feature matrix elements of the speech segments to determine the first search window, which helps improve the temporal synchronization accuracy of video frames and speech segments, optimize the matching effect of multimodal data, and thus enhance the accuracy and efficiency of video content understanding. At the same time, it analyzes the occurrence frequency and duration per unit time of text segments in each shot where the timestamp of a video frame is located to confirm the second search window, more accurately capture the temporal alignment relationship between the text and video frames, improve the accuracy of multimodal fusion, and make the matching between text information and visual content more natural and smooth. Based on the search window, it analyzes the feature similarity degree between the speech segments and text segments and the video frames within the corresponding window to obtain a sparse matrix and acquire the optimal alignment path, which helps accurately synchronize the temporal relationships among audio, text, and video frames, reduce data redundancy, and optimize the computational efficiency, thereby providing a more accurate basis for multimodal information fusion and subsequent semantic analysis. Based on the obtained optimal alignment path, it aligns the timestamps of the speech segments and text segments with the timestamps of the video frames, and integrates the aligned video frames, speech segments, and text segment information through multimodal fusion technology. It realizes multi-level semantic parsing through a deep learning model, which helps improve the fusion accuracy of multimodal data, ensure semantic consistency among different information sources, enhance the content understanding ability, and further improve the data-driven decision support and application effects.
[0017] This application transforms the global dense calculation of the traditional dynamic time warping algorithm into local sparse calculation through shot segmentation and dynamic search windows, significantly reducing the time complexity. At the same time, the design of shot boundary constraints and adaptive windows ensures the precise alignment of cross-modal signals within a semantically coherent local range, not only maintaining the temporal correlation of multimodal signals but also greatly reducing memory usage and computational resource consumption, providing efficient and accurate temporal alignment support for application scenarios such as real-time advertising placement and preprocessing of massive video inventories. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of the video content understanding method for the multimodal advertisement inventory intelligent matching system provided by this application; Figure 2 It is a flowchart for obtaining the dynamic balance factor provided by this application; Figure 3 It is a flowchart for obtaining the text dynamic factor provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example", etc. are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "or", "for example", etc. is intended to present relevant concepts in a specific manner.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0021] In addition, it should be noted that the terms "first", "second" in this application and its accompanying drawings are used to distinguish similar objects and are not used to describe a specific order or sequence. For the methods disclosed in the embodiments of this application or shown in the flowcharts, including one or more steps for implementing the methods, without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0023] This application proposes a video content understanding method for a multi-modal advertising inventory intelligent matching system, which is applied to the technical field of video content understanding. Refer to the attached Figure 1 , and the method includes the following steps: S1: Extract video frames from the original advertising video and perform shot segmentation; extract audio signals from the original advertising video to obtain speech segments; read text fragments from the original advertising video.
[0024] In the video content understanding task, the original video is unstructured data, and its visual images, audio signals, and text information exist in a mixed form, lacking a unified time reference and modal division, and cannot be directly used for cross-modal temporal analysis. Therefore, first, through multi-modal signal parsing, the unstructured video needs to be converted into structured temporal data that can be processed by a computer.
[0025] S101: Visual modality processing, the specific steps are as follows: (1)Video frame sequence parsing and timestamp extraction: Traverse the original video at a preset fixed frame rate and extract frames one by one from the beginning to the end of the video; in this embodiment, it is set to 30fps, that is, one frame is extracted every 1 / 30 second; generate an absolute timestamp for each extracted frame, with the unit of millisecond. For variable frame rate videos, synchronize timestamps through linear interpolation or nearest neighbor method to ensure time continuity; finally, obtain a video frame sequence, where each element is a video frame and its corresponding timestamp.
[0026] (2)Shot segmentation and scene boundary marking: Based on histogram difference or deep learning model for shot segmentation, record the start timestamp and end timestamp of each shot, generate shot-level time intervals (for example, shot 1 corresponds to timestamp [0ms, 2000ms], shot 2 corresponds to [2001ms, 5000ms]), and obtain a shot segmentation sequence, where each element is the start and end timestamps of a shot.
[0027] (3)Visual feature extraction: For the extracted video frames, extract feature vectors through a pre-trained visual model. In this embodiment, ResNet-18 pre-trained on ImageNet is used. Scale the video frames to 224×224 pixels and output the visual embedding vector corresponding to the video frame; obtain a visual feature sequence, where each element is the visual embedding vector corresponding to the video frame.
[0028] S102: Perform audio modality processing, and the specific steps are as follows: (1)Voice activity detection and speech segment segmentation: Extract the mono audio signal from the video, in WAV format, with a sampling rate of 16kHz. Apply a band-pass filter to remove low-frequency noise and high-frequency interference. In this embodiment, the frequency range of the band-pass filter used is 300Hz to 8kHz; use the energy detection method or a machine learning model to detect the time period when speech exists. In this embodiment, voice activity detection is performed by combining the PyAudio and WebRTC VAD modules; merge consecutive speech frames to generate speech segments with start and end timestamps, and filter out segments with too short duration. In this embodiment, segments less than 50ms are filtered; finally, obtain a speech segment sequence, where each element is a speech segment and its corresponding start and end timestamps.
[0029] (2)Speech segment feature extraction: Perform short-time Fourier transform on each speech segment to generate a spectrogram; convert linear frequency to Mel frequency through a Mel filter bank (usually 40 - 80 filters), and use the Mel spectrogram as the underlying acoustic representation of the audio to obtain an audio feature sequence, where each element is the Mel spectrogram matrix of a speech segment, with the dimension of the number of frames Mel channel number, where the number of frames refers to the number of speech frames in a speech segment, and the Mel channel number refers to the number of Mel filter banks.
[0030] S103: Perform text modality processing, and the specific steps are as follows: (1) Subtitle text segment acquisition: Read the subtitle file, split each subtitle segment by item, obtain the start and end timestamps of the subtitle segment, and remove special symbols and duplicate spaces.
[0031] (2) OCR text segment acquisition: Perform OCR on each video frame, identify the text content in the picture, sort the OCR text according to the timestamp of the video frame, and remove duplicate text.
[0032] (3) Fusion and feature generation of text segment sequences: Merge the subtitle and OCR texts, sort them according to the timestamp to generate a unified text segment sequence, where each element is a text segment and its corresponding start and end timestamps; Use a pre-trained language model to perform embedding encoding on each text segment to generate a context semantic vector, and obtain a text segment feature sequence, where each element is a semantic embedding vector of a text segment and its corresponding start and end timestamps. In this embodiment, the pre-trained language model used is BERT ((Bidirectional Encoder Representations from Transformers)), which is a well-known existing technology and will not be elaborated in this application. In other embodiments, the RoBERTa (A Robustly Optimized BERT Pretraining Approach) model can also be used.
[0033] S2: Based on the start and end times of each shot segmentation, divide the video frames, speech segments, and text segments with corresponding timestamps.
[0034] Traditional temporal alignment methods only perform temporal modeling on single modalities, do not explicitly handle cross-modal asynchronous problems, or sample modal features at fixed intervals for alignment, ignoring the actual time misalignment, resulting in incorrect semantic associations. However, this solution uses an improved DTW algorithm to perform temporal alignment on multi-modal videos. Through dynamic elastic matching, it allows local bending of the time axis, adaptively processes various asynchronous scenarios, and has a small alignment accuracy error, providing a more reliable time calibration basis for the intelligent matching of multi-modal advertising inventories.
[0035] A shot is the basic semantic unit of video content. The visual, audio, and text signals within the same shot have strong spatio-temporal correlations, while shot transitions are usually accompanied by breaks in scenes, actions, or semantics. Restricting the Dynamic Time Warping (DTW) alignment within a single shot can avoid invalid cross-shot matching, effectively reducing the computational complexity of traditional DTW. The semantic consistency of multi-modal signals within a shot is higher, and the dynamic time warping path is more likely to converge to the true temporal correspondence, avoiding distortion errors caused by cross-shot semantic discontinuities. Combining shot boundary forced alignment and dynamic search window restriction can further compress the invalid calculation area, optimizing the computational efficiency while ensuring accuracy.
[0036] According to the start and end timestamps of each shot in the shot segmentation sequence, the elements in the video frame sequence, speech segment sequence, and text segment sequence are respectively matched to the corresponding shot segments according to the timestamps, ensuring that each segment only contains multi-modal signals within the same shot. For example, the timestamp of an audio segment completely falls within the time interval of a certain shot, or an audio segment spanning multiple shots is split into independent sub-segments at the shot boundaries, finally obtaining multi-modal segments for each shot.
[0037] S3: Taking the video frames as the reference, within the shot corresponding to the timestamp of each video frame, analyze the similarity between adjacent video frames and the dispersion degree of the spectral information corresponding to the speech segments, determine the dynamic balance factor for the timestamp of each video frame, and adjust the preset reference window size for the timestamp of each video frame to obtain the first search window.
[0038] When calculating the similarity of time series, the dynamic search window defines an allowed horizontal offset range for each alignment point, ensuring that the alignment path only extends within a band-shaped area centered on the diagonal. This application prunes invalid alignment paths through the dynamic search window, reducing the time complexity of traditional DTW and significantly reducing the computational amount while ensuring the rationality of alignment.
[0039] Since time series usually have local temporal dependencies, that is, the elements at the current moment only have reasonable matching relationships with the elements at adjacent moments. The dynamic window effectively eliminates meaningless long-distance matches by restricting the horizontal offset range of the alignment path, reducing the occurrence of redundant calculations in global search. The window width is dynamically adjusted according to the characteristics of the sequence to ensure that the alignment paths that conform to the temporal logic are retained within the allowed offset range, avoiding alignment distortion caused by excessive pruning.
[0040] This application uses video frames as a benchmark to align the timing of audio and text. Visual information plays a central role in ad contextual targeting because it provides semantic representation of the scene, and video frames, with their fixed frame rate and continuous timestamps, provide a stable temporal reference. Therefore, ad contextual targeting primarily relies on visual semantic representations, with audio and text serving as supplementary information.
[0041] When the video frames are aligned with the speech segments, the average value of the cosine similarities between all adjacent video frames in the shot where the timestamp of each video frame is located is obtained as the visual motion intensity of the timestamp corresponding to each video frame, which is used to reflect the dynamic degree of the picture. The range is 0~2. The larger the value, the more drastic the picture change. The standard deviation of each element in the audio feature sequence corresponding to the shot where the timestamp of each video frame is located is obtained and the mean is calculated to obtain the audio energy change rate of the timestamp corresponding to each video frame, which is used to reflect the dynamic range of the audio signal. The larger the value, the more drastic the audio energy fluctuation. The absolute value and sum of the difference between the visual motion intensity and the audio energy change rate corresponding to each video frame are calculated respectively, and the negative correlation mapping result of the sum value is positively fused with the absolute value of the difference to obtain the dynamic balance factor of the timestamp corresponding to each video frame.
[0042] In this embodiment, the absolute value of the difference is recorded as A, and the sum is recorded as B. Then the formula of the dynamic balance factor is: Where, It represents the preset parameter, which is set to 0.01 to prevent the denominator from being zero. It should be understood that the window expansion amplitude is adjusted by the difference between the visual and audio dynamic characteristics. When the difference between the visual motion intensity and the audio energy change rate is large, the dynamic balance factor approaches 1, and the window width increases significantly to adapt to the large span distortion of the single mode. When the two are close, the dynamic balance factor approaches 0, and the window width is smoothly adjusted based on the average dynamic characteristics. Among them, the flow chart for obtaining the dynamic balance factor is as follows Figure 2 shown.
[0043] The width of the first search window is calculated based on the visual motion intensity and the audio energy change rate. The formula is: ,in, represents the width of the first search window corresponding to the timestamp of the t-th video frame; t represents the timestamp corresponding to the t-th video frame, Indicates the reference window size, which is used to define the minimum allowable search range for temporal alignment to ensure basic alignment accuracy in static scenarios. Its value range is 30 to 100 frames. In this embodiment, the value is 50 frames. Represents the Sigmoid function, The value range of is mapped to 0~1 to avoid sudden changes in window width and ensure smooth transition; Denotes the floor value; Denotes the visual motion intensity corresponding to the time stamp of the t-th video frame; Denotes the audio energy change rate corresponding to the time stamp of the t-th video frame; Denotes the dynamic balance factor corresponding to the time stamp of the t-th video frame.
[0044] It should be understood that by fusing the visual motion intensity and the audio energy change to dynamically adjust the search window width, when the picture switches rapidly or the audio changes violently, that is, the larger the value of the visual motion intensity or the audio energy change rate, the width of the first search window Expands, allowing a larger range of time warping to capture the true temporal correspondence in complex clips (such as slow motion, voiceovers); when the scene is static, the window width approaches the base width, reducing the invalid search area and lowering the calculation. The dynamic balance factor balances the dynamic differences between vision and audio, avoiding excessive window expansion caused by violent changes in a single modality, and finding the optimal solution between efficiency and accuracy; the Sigmoid function ensures that the window width is gradually adjusted with the change of the scene, avoiding frequent recalculation caused by small fluctuations within the shot, and improving the algorithm stability.
[0045] S4: In the shot where the time stamp corresponding to each video frame is located, analyze the appearance frequency and duration of the text segment corresponding to the time stamp per unit time, confirm the text dynamic factor of the time stamp corresponding to each video frame, and adjust the preset reference window size of the time stamp corresponding to each video frame to obtain a second search window.
[0046] When aligning the video frame with the text segment, obtain the appearance frequency of the text segment corresponding to the time stamp per unit time within the shot where the time stamp corresponding to each video frame is located, and the duration of the text segment corresponding to the time stamp, multiply the two to obtain a first product; obtain the maximum appearance frequency and the maximum duration of the text segment corresponding to the time stamp in the text segment sequence, multiply them to obtain a second product; use the result of positively fusing the negative correlation mapping result of the second product and the first product as the text dynamic factor of the time stamp corresponding to each video frame. In this embodiment, the first product is denoted as C, the second product is denoted as D, and the formula form of the text dynamic factor is: ; where Denotes a preset parameter, with a value of 0.01, to prevent the denominator from being zero. Among them, the flowchart for obtaining the text balance factor is as Figure 3 shown.
[0047] Next, calculate the width of the second search window according to the semantic relevance of the text, and its formula is: , where: Denotes the width of the second search window corresponding to the time stamp of the t-th video frame; represents the semantic similarity corresponding to the time stamp of the t-th video frame, that is, the average cosine similarity of the elements at the same positions of the text segment feature sequence and the visual feature sequence within the shot where the time stamp is located. The larger the value, the stronger the semantic association between the text and the picture content; represents a preset activation factor, with a range of 0 to 1. In this embodiment, the value is 0.5. Only when the window is extended to avoid invalid calculations caused by meaningless text; represents the text dynamic factor corresponding to the time stamp of the t-th video frame; represents a preset reference window size; represents an activation function. When is less than 0, it outputs 0. When the result is greater than 0, it outputs ; represents the floor value. Among them, for the calculation of semantic similarity, it is first necessary to perform L2 normalization on the embedded semantic vector of the text segment and the visual embedding vector of the video frame. To eliminate the difference in modulus length and make the cosine similarity only reflect the consistency of the vector directions, then the text and visual features are mapped to the same dimension through a projection layer (such as a fully connected layer), so as to ensure the effectiveness of the cosine similarity calculation of the corresponding vectors at the same positions.
[0048] It should be understood that when the text and the visual content are strongly semantically related, the width of the second search window is linearly extended with the part exceeding the threshold, allowing a larger range of temporal distortion to adapt to the non-strict synchronization of the text description and the action. In the weak association scenario, the basic window is maintained to limit the calculation range; for high-frequency long texts, the window expansion amplitude is amplified through to cope with the continuous picture changes corresponding to dense semantics. For low-frequency short texts, the window expansion is restricted to avoid excessive search.
[0049] It should be noted that since the dynamic search is performed within the shot segment, when the width of the dynamic search window exceeds the range of the shot segment, only the content within the shot segment is retained.
[0050] S5: Based on the first search window and the second search window, analyze the feature similarity degrees of the speech segment and the text segment with the video frame within the corresponding windows, obtain the first sparse matrix and the second sparse matrix, and acquire the first optimal alignment path and the second optimal alignment path.
[0051] For each video frame in each shot segment of the video frame sequence, calculate its corresponding first search window and second search window according to steps S2 and S3 respectively, map the speech segment or the text segment to the video frame time axis, and determine the corresponding speech segment index range and text segment index range within the window.
[0052] To save storage space, this application adopts a sparse matrix format and only stores similarity values that are not infinite, avoiding storing all-zero elements. First, initialize the visual-audio matrix and the visual-text matrix as sparse matrix structures, where the number of rows of the matrix is the number of video frames, and the number of columns is the number of speech segments or text segments. All elements are initially set to infinity, indicating that cross-window alignment is not allowed by default.
[0053] Then, for each video frame, traverse the speech segments within its first search window, calculate the cosine similarity between the elements of the visual feature sequence and the audio feature sequence within the window, and store the result in the corresponding position of the first sparse matrix. For each video frame, traverse the text segments within its second search window, calculate the normalized Euclidean distance between the elements of the visual feature sequence and the text segment feature sequence within the window, and store it in the corresponding position of the second sparse matrix.
[0054] Using dynamic programming to search for the optimal path in the first sparse matrix and the second sparse matrix is a well-known technique, and the specific steps will not be elaborated here. Finally, output the optimal alignment path composed of the timestamp indexes of video frames and speech segments and the timestamp indexes of video frames and text segments, which satisfies temporal continuity and has the maximum cumulative similarity.
[0055] S6: Based on the obtained optimal alignment path, align the timestamps of the speech segments and text segments with the video frame timestamps, integrate the information of the video frames, speech segments, and text segments with aligned timestamps through multimodal fusion technology, and implement multi-level semantic parsing through a deep learning model.
[0056] Map the timestamps of the speech segment sequence and the text segment sequence to the time axis of the video frame sequence based on the optimal alignment path, that is, establish a mapping for each point in the path. If the timestamps of the speech segment and the text segment exactly match the video frame, directly take the corresponding elements in the sequence; if the timestamps do not match, generate new elements through linear interpolation. For speech segments and text segments outside the optimal alignment path, use the nearest boundary value to fill or discard them, and finally output a synchronized multimodal sequence aligned with the video frame sequence timestamp.
[0057] According to the synchronized multimodal sequence, integrate cross-modal semantic information through multimodal fusion technologies (such as Transformer, gating mechanism). Visual features capture the main body, scene, and actions of the picture, audio features analyze speech content, emotional rhythm, and ambient sound, and text features supplement the semantic logic of subtitles and OCR text.
[0058] Furthermore, multi-level semantic parsing is implemented through a deep learning model, including video topic classification, dynamic scene recognition, emotional tone judgment, and target user profile mapping. Finally, these fine-grained semantic representations provide an accurate scene matching basis for the context targeting of programmatic advertising, improving the user experience and delivery efficiency while ensuring the relevance of the advertisement. Multimodal fusion technology and semantic parsing are well-known technologies, and their specific content and detailed steps will not be elaborated here.
[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0060] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for video content understanding of a multi-modal advertising inventory intelligent matching system, characterized in that The method includes the following steps: Performing video frame extraction and shot segmentation on the original advertising video; extracting an audio signal from the original advertising video to obtain speech segments; reading text segments from the original advertising video; Dividing the video frames, speech segments, and text segments with corresponding timestamps based on the start and end times of each shot segmentation; Taking the video frames as a reference, analyzing the similarity degree between adjacent video frames and the dispersion degree of the spectral information corresponding to the speech segments in the shot where the timestamp corresponding to each video frame is located, determining the dynamic balance factor for the timestamp corresponding to each video frame, and adjusting the preset reference window size for the timestamp corresponding to each video frame to obtain a first search window; In the shot where the timestamp corresponding to each video frame is located, analyzing the occurrence frequency and duration of the text segment corresponding to the timestamp per unit time, confirming the text dynamic factor for the timestamp corresponding to each video frame, and adjusting the preset reference window size for the timestamp corresponding to each video frame to obtain a second search window; Based on the first search window and the second search window, analyzing the feature similarity degree between the speech segments and text segments and the video frames within the corresponding windows to obtain a first sparse matrix and a second sparse matrix, and obtaining a first optimal alignment path and a second optimal alignment path; Based on the obtained optimal alignment paths, aligning the timestamps of the speech segments and text segments with the timestamps of the video frames, integrating the information of the video frames, speech segments, and text segments with aligned timestamps through multi-modal fusion technology, and implementing multi-level semantic parsing through a deep learning model.
2. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 1, wherein The dynamic balance factor for the timestamp corresponding to each video frame is specifically: Obtaining the average value of the cosine similarities between all adjacent video frames in the shot where the timestamp corresponding to each video frame is located as the visual motion intensity for the timestamp corresponding to each video frame; Extracting the Mel spectral matrix of each speech segment and using the sequence composed of the Mel spectral matrices of all speech segments as the audio feature sequence; Obtaining the standard deviation of each element in the audio feature sequence corresponding to the shot where the timestamp corresponding to each video frame is located and calculating the mean value to obtain the audio energy change rate for the timestamp corresponding to each video frame; Calculating the absolute value of the difference and the sum value between the visual motion intensity and the audio energy change rate corresponding to each video frame respectively, and performing positive fusion of the negative correlation mapping result of the sum value with the absolute value of the difference to obtain the dynamic balance factor for the timestamp corresponding to each video frame.
3. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 2, wherein, Adjusting the preset reference window size corresponding to the timestamp of each video frame to obtain a first search window, the specific formula is: , where represents the width of the first search window corresponding to the timestamp of the t-th video frame; t represents the timestamp corresponding to the t-th video frame, represents the preset reference window size; represents the Sigmoid function; represents the floor value; represents the visual motion intensity corresponding to the timestamp of the t-th video frame; represents the audio energy change rate corresponding to the timestamp of the t-th video frame; represents the dynamic balance factor corresponding to the timestamp of the t-th video frame.
4. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 1, wherein, The confirmation of the text dynamic factor for the timestamp corresponding to each video frame is specifically: Taking the sequence composed of all text segments as the text segment sequence; Obtaining the occurrence frequency of the text segment corresponding to the timestamp per unit time and the duration of the text segment corresponding to the timestamp within the shot where the timestamp corresponding to each video frame is located, and multiplying the two to obtain a first product; Obtaining the maximum occurrence frequency and the maximum duration of the text segment corresponding to the timestamp in the text segment sequence, and multiplying them to obtain a second product; Taking the result of positive fusion of the negative correlation mapping result of the second product with the first product as the text dynamic factor for the timestamp corresponding to each video frame.
5. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 2, characterized in that, Adjusting the preset reference window size corresponding to the timestamp of each video frame to obtain a second search window, specifically: In the shot where the timestamp corresponding to each video frame is located, analyze the similarity degree of the feature information between the video frame and the text segment, and determine the semantic similarity of the timestamp corresponding to each video frame; Denote the width of the second search window at the corresponding timestamp of the t-th video frame as , and its formula form is: , where: represents; represents the semantic similarity at the timestamp corresponding to the t-th video frame; represents a preset activation factor; represents an activation function; represents the floor value; represents the text dynamic factor at the timestamp corresponding to the t-th video frame; represents a preset reference window size.
6. The video content understanding method of the multimodal advertising inventory intelligent matching system according to claim 5, characterized in that, The specific acquisition process of the semantic similarity is as follows: For each video frame, extract the visual embedding vector through a pre-trained visual model, and use the sequence composed of all visual embedding vectors as the visual feature vector; Generate the context semantic vector for each text segment, and use the sequence composed of all context semantic vectors as the text segment feature sequence; Within the shot where the timestamp corresponding to each video frame is located, calculate the mean value of the cosine similarities of the elements at the same positions in the text segment feature sequence and the visual feature sequence as the semantic similarity of the timestamp corresponding to each video frame.
7. The video content understanding method of the multimodal advertisement inventory intelligent matching system according to claim 6, characterized in that, The specific method for obtaining the first sparse matrix and the second sparse matrix is as follows: For each video frame, traverse the speech segments within its first search window, and calculate the cosine similarity between the elements of the visual feature sequence within the window and the elements of the audio feature sequence as the elements of the first sparse matrix; For each video frame, traverse the text segments within its second search window, and use the normalized Euclidean distance between the elements of the visual feature sequence within the window and the elements of the text segment feature sequence as the elements of the second sparse matrix.
8. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 1, characterized in that The first optimal alignment path is obtained by using the optimal path algorithm for the first sparse matrix; the second optimal alignment path is obtained by using the optimal path algorithm for the second sparse matrix.
9. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 1, characterized in that The first optimal path is specifically the optimal alignment path composed of the timestamp index pairs of the video frame and the speech segment; the second optimal path is specifically the optimal alignment path composed of the timestamp index pairs of the video frame and the text segment.
10. The video content understanding method of the multi-modal advertisement inventory intelligent matching system according to claim 1, wherein Based on the obtained optimal alignment path, align the timestamps of the speech segment and the text segment with the video frame timestamp, specifically: Based on the optimal alignment path, map the timestamps of the speech segment and the text segment to the time axis of all video frames respectively; If the timestamps of the speech segment and the text segment in the optimal alignment path exactly match the video frame, directly take the corresponding speech segment and text segment; if the timestamps do not match, generate new speech segments and text segments through linear interpolation to obtain the speech segment and text segment aligned with the video frame timestamp.
Citation Information
Patent Citations
Audio and video content retrieval system and method based on artificial intelligence
CN118332140A
Multi-modal video understanding method based on large language model
CN118675092A
Video segmentation method and device based on multi-modal features, equipment and storage medium
CN119478763A
Video timestamp event identification and reasoning method based on multi-modal large model
CN119723431A
Video feature extraction and multi-dimensional matching-based movie advertisement real-time pushing system
CN120181922A
Cited By
Video content description method and device, electronic equipment and storage medium
CN120956982A
Automatic subtitle translation method and device based on large model and storage medium
CN121189338A
Method and device for automatic translation of subtitles based on large models, and storage medium
CN121189338B
Advertisement content automatic monitoring method and system
CN121660749A