A video stream multimodal content security dynamic audit system and method

By combining interactive behavior analysis of visual and audio modalities, the modal jump behavior segments in the video stream are identified, solving the problem of inaccurate modal disconnection judgment in existing technologies and achieving efficient compliance review of video content.

CN120455652BActive Publication Date: 2025-09-05SHAANXI TAODING IND GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510922318.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-05
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

The existing dynamic review system for multimodal content security of video streams has difficulty capturing time series correlation features when processing sensitive clips that switch quickly or appear asynchronously, resulting in inaccurate judgment of modal disconnection and easy omission of violation signals generated by multimodal interaction. It also lacks the ability to continuously evaluate dynamic features before and after emergencies, and the review robustness is insufficient.

Method used

Through the visual emergence candidate indicator extraction module, the violation picture judgment module, the modal mutation behavior recognition module and the modal interleaving interruption node recognition module, combined with the interactive behavior analysis of image and audio bimodal data, the modal jump behavior segments are identified and potential violation fragments are located, and a multi-dimensional measurement and scoring mechanism is constructed.

Benefits of technology

It achieves efficient review of video content compliance, improves the recognition coverage and response capabilities of multimodal mutation behaviors, and ensures the accurate restoration of abnormal interruption behaviors in complex video streams and the effective characterization of behavioral boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455652B_ABST
    Figure CN120455652B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of video analysis technology, specifically a system and method for dynamic security review of multimodal content of video streams. In the present invention, by jointly analyzing the changes in the hue ratio of continuous image frames and the heat map of the focus area, it is possible to accurately identify the areas in the video screen where the hue changes suddenly and the focus is focused, and effectively locate potential illegal fragments. On this basis, an interactive behavior analysis mechanism of image and audio bimodal data is introduced, so that the recognition of sudden behaviors is no longer limited to a single visual or auditory clue, but a modal jump behavior segment is constructed when the image action and voice changes are synchronously highlighted, and then the temporal overlapping points in the modal mutation process are further captured to enhance the positioning accuracy of abnormal interruption behaviors, and the trend information of the frame-level dominant weight changes is combined to realize multi-dimensional measurement and scoring of the sudden area, thereby improving the recognition coverage and response capability of dynamic content review to illegal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video analysis technology, and in particular to a system and method for dynamically reviewing the security of multimodal content in video streams. Background Art

[0002] The field of video analytics involves the identification, analysis, and understanding of targets, events, and behaviors in image sequences. It primarily encompasses core areas such as target detection and tracking, behavior recognition, scene analysis, structured video representation, content filtering, and real-time monitoring. This technology is widely used in security monitoring, intelligent transportation, media content management, and public safety scenarios.

[0003] Among them, the traditional video stream multimodal content security dynamic review system refers to a system that performs compliance and security testing on video stream content transmitted in real time on the Internet. It is used to dynamically identify and review sensitive content in modalities such as fused images and text in multi-source video information.

[0004] Since existing technologies rely solely on the fusion analysis of static modes of images and text, they find it difficult to capture the time series correlation characteristics of sudden changes in video streams. When processing sensitive clips that switch quickly or appear asynchronously, they often encounter problems with modal disconnection and inaccurate judgment. In particular, in scenarios where image color changes and voice changes are not synchronized, it is easy to miss violation signals generated by multimodal interactions. As a result, when the review system identifies videos with drastic picture jumps or audio masking, it is difficult to perceive and mark abnormal paragraphs in a timely manner, resulting in the risk of delays or misjudgments in content compliance determination results. In addition, due to the lack of segmented modeling of the behavioral evolution path of the modal mutation process, there is a lack of continuous evaluation capabilities for dynamic characteristics before and after emergencies. The review robustness is obviously insufficient in scenarios where high-density, short-term violation information accumulates. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a system and method for dynamic security review of multimodal content of video streams.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions: A video stream multimodal content security dynamic review system includes:

[0007] The visual emergence candidate indicator extraction module extracts consecutive frames of the audit video and constructs the image's hue ratio vector and focus heat map. It then selects image frames with sudden hue ratio mutations in local frame segments of the sequence to obtain a set of visual emergence candidate indicators.

[0008] The illegal picture determination module calculates the spatial distribution difference of the sudden change color area in the focus heat map based on the visual salience candidate indicator set, and determines the illegal picture salience frame set;

[0009] The modal mutation behavior recognition module extracts the image and audio multimodal sequence in the time segment covered by the set of frames with the offending images, identifies the continuous frame segments where the modality dominant undergoes a jump change, and constructs a set of modal mutation behavior segments;

[0010] The modal interleaving interruption node identification module identifies modal interleaving interruption frame locations where the image action and the voice activity time overlap in the modal mutation behavior segment set to form a modal interleaving interruption node set;

[0011] The video content review module reviews the interruption behavior in the video based on the modal interleaving interruption node set and outputs the video content review result.

[0012] As a further solution of the present invention, the visual emergence candidate indicator set specifically includes image frames with sudden changes in hue ratio, local frame segments, and hue ratio sudden changes; the illegal picture emergence frame set includes sudden hue areas, spatial distribution differences, and illegal pictures; the modal mutation behavior segment set specifically includes image movement changes, voice activity changes, and continuous frame segments; the modal interleaving interruption node set includes modal interleaving interruption frame sites, and frame sites where image movement and voice activity overlap in time; the video content review results include review conclusions and video interruption behaviors.

[0013] As a further solution of the present invention, the visual salience candidate indicator extraction module includes:

[0014] The tone proportion vector construction submodule obtains the complete frame image sequence of the audit video, divides each frame into a fixed-size grid, calculates the main tone type of each grid and the area ratio of the main tone type in the entire frame image, and generates a tone proportion vector set for each frame;

[0015] The focus heat map generation submodule converts each frame of image into a grayscale image, calculates the pixel clarity gradient using the Laplacian operator, and smoothes it through Gaussian filtering to generate a focus heat map for each frame of image with the pixel gradient amplitude as the weight;

[0016] The mutation frame screening submodule compares the hue proportion vector set of each frame and the focus heat map of each frame image for consecutive frames, screens frames with hue mutations, and obtains a set of visual salience candidate indicators.

[0017] As a further solution of the present invention, the illegal image determination module includes:

[0018] The focus region extraction submodule calls the focus heat map of each frame of the visual salience candidate index set, extracts the pixel coordinate set whose heat value in the previous frame is higher than the image heat mean, marks the target set as the focus reference region, and obtains the focus region coordinate set;

[0019] The mutation region positioning submodule locates the image region corresponding to the grid where the hue mutation is located in the current frame according to the focus region coordinate set, extracts the grid boundary coordinates of the image region and uniformly maps them to the image coordinate system to obtain the mutation region coordinate set;

[0020] The spatial overlap judgment submodule calls the focus area coordinate set and the mutation area coordinate set, counts the number of intersections of pixel coordinate points in the two sets, and calculates the proportion of the intersection number in the mutation area. The proportion is judged according to the set spatial overlap rate threshold to filter out the set of frames with illegal screen bursts.

[0021] As a further solution of the present invention, the modal mutation behavior recognition module includes:

[0022] The modal sequence extraction submodule obtains the time segment covered by all frame indexes in the set of frames where the illegal image appears, extracts the video image frame sequence and audio signal sequence corresponding to the time segment, extracts the displacement amplitude of the center point of the character's bounding box and the number of focused heat areas from the video image frame sequence, and extracts the speech start and end time periods and the amplitude of the sound source position change from the audio signal sequence to obtain image and audio modal features;

[0023] The dominant weight calculation submodule calculates the mutual information value between the image and audio modal features respectively, constructs a mutual information sequence at the corresponding frame level, assigns a dominant weight of the image modality or audio modality to each frame according to the mutual information value, and obtains a modality dominant weight data sequence;

[0024] The jump segment identification submodule detects the weight change trend of the modal dominant weight data sequence in continuous frame segments, identifies the time period in which continuous mutations occur in the weight dominant type, selects the frame segments with a change span greater than the jump judgment criterion as the jump area, and obtains the modal mutation behavior segment set.

[0025] As a further solution of the present invention, the modal interleaving interruption node identification module includes:

[0026] The image interruption extraction submodule obtains the image frame sequence in the modal mutation behavior segment set, extracts the center position coordinates of the character boundary box between adjacent frames, calculates the Euclidean distance change rate between frames, determines the image interruption frame location, and obtains the image interruption frame index set;

[0027] The voice interruption extraction submodule obtains the amplitude energy of each frame based on the audio signal sequence of the time period corresponding to the image interruption frame index set, determines the frame index with energy lower than the silence threshold as the voice interruption frame position, and obtains the voice interruption frame index set;

[0028] The time overlap screening submodule calls the image interruption frame index set and the voice interruption frame index set, compares the time tags of the two types of frame indexes, screens frame points with completely consistent indexes as interleaving interruption points, and generates a modal interleaving interruption node set.

[0029] As a further solution of the present invention, the video content review module includes:

[0030] The weight change identification submodule obtains the frame index in the modal interleaving interruption node set, extracts the corresponding modal dominant weight data before and after the corresponding frame, calculates the inter-frame change slope of the image modal weight, and identifies the frame segments with a continuous upward trend to obtain the image modal weight change interval;

[0031] The abnormal segment determination submodule determines whether the frame index of the image modal weight change interval is included in the set of illegal picture burst frames. If included, the target frame segment is recorded, and the modal dominant jump frequency and the frame index distribution density within the frame segment are calculated to obtain the modal abnormal jump score;

[0032] The audit result output submodule marks the corresponding start and end indexes of the frame segments on the time axis according to the modal abnormal jump score, constructs a multi-modal behavior abnormality identification structure, and summarizes and generates the video content audit results.

[0033] A method for dynamic security audit of multimodal content of video streams is provided. The method is based on the above-mentioned dynamic security audit system for multimodal content of video streams and includes the following steps:

[0034] S1: Extract consecutive frames of the audit video and construct the image's hue ratio vector and focus heat map. Filter the image frames with sudden hue ratio mutations in the local frame segments of the sequence to obtain a set of candidate visual salience indicators.

[0035] S2: Calculating the spatial distribution difference of the sudden change hue area in the focus heat map based on the visual salience candidate indicator set, and determining the set of illegal picture salience frames;

[0036] S3: extracting the image and audio multimodal sequences in the time segment covered by the set of frames with the offending images appearing suddenly, identifying the continuous frame segments where the modality dominant undergoes a jump change, and constructing a set of modality mutation behavior segments;

[0037] S4: identifying modal interleaving interruption frame locations where the image action and the voice activity time coincide in the modal mutation behavior segment set, and forming a modal interleaving interruption node set;

[0038] S5: Review the interruption behavior in the video based on the modal interleaving interruption node set, and output the video content review result.

[0039] Compared with the prior art, the advantages and positive effects of the present invention are:

[0040] In the present invention, by jointly analyzing the changes in the hue proportion of continuous image frames and the focus area heat map, it is possible to accurately identify the areas in the video screen where the hue changes suddenly and the focus is on, and effectively locate potential illegal fragments. On this basis, an interactive behavior analysis mechanism of image and audio bimodal data is introduced, so that the recognition of sudden behaviors is no longer limited to a single visual or auditory clue, but a modal jump behavior segment is constructed when the image action and voice changes are synchronously highlighted, and then the temporal overlapping points in the modal mutation process are further captured to enhance the positioning accuracy of abnormal interruption behaviors, and combined with the trend information of the frame-level dominant weight changes to realize multi-dimensional measurement and scoring of sudden areas, thereby improving the recognition coverage and response capability of dynamic content review for illegal behaviors, ensuring the accurate restoration of multimodal mutation behaviors and effective characterization of behavior boundaries in complex video streams, and ultimately achieving efficient review of video content compliance and structured result output. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a system flow chart of the present invention;

[0042] Figure 2 This is a flow chart of candidate visual salience indicators of the present invention;

[0043] Figure 3 This is a flow chart of the illegal image determination module of the present invention;

[0044] Figure 4 This is a flow chart of the modal mutation behavior recognition module of the present invention;

[0045] Figure 5 This is a flow chart of the modal interleaving interruption node identification module of the present invention;

[0046] Figure 6 This is a flowchart of the video content review module of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0048] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0049] See also Figure 1 The present invention provides a technical solution: a video stream multimodal content security dynamic review system comprising:

[0050] The visual emergence candidate indicator extraction module extracts consecutive frames of the audit video and constructs the image's hue ratio vector and focus heat map. It then selects image frames with sudden hue ratio mutations in local frame segments of the sequence to obtain a set of visual emergence candidate indicators.

[0051] The illegal image judgment module calculates the spatial distribution difference of the sudden tones in the focus heat map based on the visual salience candidate index set and determines the set of illegal image salience frames.

[0052] The modal mutation behavior recognition module extracts the image and audio multimodal sequences in the time segment covered by the set of frames where the offending images suddenly appear, identifies the continuous frame segments where the modality dominant undergoes a jump change, and constructs a set of modal mutation behavior segments;

[0053] The modal interleaving interruption node identification module identifies the modal interleaving interruption frame locations where the image action and the voice activity time coincide in the modal mutation behavior segment set, forming a modal interleaving interruption node set;

[0054] The video content review module reviews the interruption behavior in the video based on the modal interleaved interruption node set and outputs the video content review results;

[0055] The candidate indicator set for visual emergence includes image frames with sudden changes in hue ratio, local frame segments, and hue ratio sudden changes; the set of frames with sudden changes in illegal images includes sudden changes in hue area, spatial distribution differences, and illegal images; the set of modal mutation behavior segments includes image motion changes, voice activity changes, and continuous frame segments; the set of modal interleaving interruption nodes includes modal interleaving interruption frame sites and frame sites where image motion and voice activity overlap in time; the video content review results include review conclusions and video interruption behaviors.

[0056] See also Figure 2 ,The visual salience candidate index extraction module includes:

[0057] The tone proportion vector construction submodule obtains the complete frame image sequence of the audit video, divides each frame into a fixed-size grid, calculates the main tone type of each grid and the area ratio of the main tone type in the entire frame image, and generates a tone proportion vector set for each frame;

[0058] After obtaining a complete frame image sequence of the audit video, each frame is divided into a fixed-size grid. The division method is to divide each frame into 32×18 grid units in the horizontal and vertical directions respectively. At a resolution of 1920×1080, the size of each grid is 60×60 pixels. After the image is divided, the RGB values ​​of all pixels in each grid are obtained and the main color type is extracted according to the common color classification rules. The maximum channel value method is used to compare the average values ​​of the three RGB channels to determine the main color of each grid. Then, the number of grids covered by each main color in a frame is counted and its area ratio is calculated. The color ratio calculation formula is: ,in, Indicates the main color The proportion in the current frame, Indicates the main color The number of grids covered, Indicates the total number of grids in the current frame image (here is 576). Taking a frame image as an example, if the main color red appears on 72 grids, its area ratio is , repeat the operation for the remaining seven main colors to obtain a complete color ratio vector. This operation is performed on all frame images to obtain a complete set of color ratio vectors.

[0059] The focus heat map generation submodule converts each frame of image into a grayscale image, calculates the pixel clarity gradient using the Laplacian operator, and smoothes it through Gaussian filtering to generate a focus heat map for each frame of image with the pixel gradient amplitude as the weight;

[0060] Each frame of the acquired image is grayscale converted. The linear weighting method in the ITU-RBT.601 standard is used to assign weights of 0.299, 0.587 and 0.114 to the pixel values ​​of the three RGB channels respectively, and then perform weighted summation to calculate the grayscale value of each pixel and obtain a complete grayscale image. This operation is used to weaken the influence of color differences on the subsequent image clarity calculation. The focus area is extracted based on the brightness gradient in the image. Then, the Laplacian operator is applied to the grayscale image of each frame to obtain a clarity gradient map. The operator uses a 3×3 convolution kernel to perform a neighborhood difference convolution operation on each pixel in the image, thereby obtaining the second-order derivative value of the pixel grayscale change in the image. In actual operation, it is assumed that the grayscale distribution around a certain pixel point is a symmetrical area. , then the Laplacian value of the point will approach 0, while if it is at the edge, the absolute value of the value will be larger, which is used to reflect the clarity of the image. On this basis, a Gaussian filtering operation is performed on the obtained gradient map. The filtering uses a two-dimensional Gaussian kernel with a standard deviation of 1.2, and the convolution window size is set to 5×5. This processing reduces the high-frequency noise caused by tiny textures or background interference in image edge detection to improve the aggregation clarity of image edge information. The processed image is then normalized, and the gradient amplitude of all pixels is mapped to the [0, 1] interval according to the maximum and minimum values ​​of the entire image. The normalization uses a linear stretching method, that is, the ratio of all pixel values ​​to the maximum gradient value of the image is used as its weight. The image gradient map generated after normalization is the focus heat map, which reflects the relative focus degree of each area in the image.

[0061] The mutation frame screening submodule compares the hue ratio vector set of each frame and the focus heat map of each frame image for consecutive frames, filters out frames with hue mutations, and obtains a set of visual salience candidate indicators;

[0062] After constructing the hue proportion vector set and focus heatmap for each frame, a difference measure is performed on the hue proportion changes between consecutive frames to identify frames with potentially abnormal hue changes in the image sequence. The analysis window is set to three frames, meaning that the current frame is jointly compared with its two previous frames. First, the hue proportion vectors for each frame are arranged in order, with each vector dimension representing the proportion of area occupied by a dominant hue (e.g., red, green, or blue) in the frame. The current frame vector is then dimensionally subtracted from the average vector of the two previous frames. The absolute values ​​of the differences across all dimensions are summed to form a total difference measure. For example, if the red proportion of the current frame is 0.24, and the red proportions of the previous two frames are 0.10 and 0.11, respectively, then the average value of the red channel is 0.105, and the difference is 0.135. The same process is repeated for the other hue dimensions. Finally, the differences across all channels are summed to determine the hue change degree for the frame. If the total difference exceeds the set hue abrupt change threshold of 0.15, the frame is preliminarily judged to have a significant hue abrupt change. Taking frame 105 as an example, its calculated hue variation is 0.162, exceeding the set threshold and therefore marked as an abnormal candidate frame. After completing the hue mutation screening, spatial distribution matching is performed in conjunction with the focus heat maps of the previous and next frames to determine whether the hue mutation area is located outside the visual focus area. This determination is based on comparing the position coordinates of the image area corresponding to the grid where the hue mutation occurs in the current frame with the coordinates of the area with a focus weight greater than 0.6 in the focus heat map of the previous frame. The spatial overlap ratio of the two sets is calculated as the number of pixels in the intersection area divided by the total number of pixels in the hue mutation area. For example, if the hue mutation area contains 1800 pixels corresponding to the grid, and the coordinate intersection with the focus area of ​​the previous frame contains 370 pixels, the spatial overlap ratio is 370 ÷ 1800 ≈ 0.206, which is less than the set focus overlap threshold of 0.3. Therefore, the hue mutation area is determined not to be part of the visual focus area of ​​the previous frame, further confirming that the frame is an emergent risk frame. Finally, all image frames that meet the requirements of exceeding the hue mutation limit and having a spatial overlap rate lower than the threshold are sorted and output in sequence to form a set of candidate visual salience indicators.

[0063] See also Figure 3 , the illegal screen judgment module includes:

[0064] The focus region extraction submodule calls the focus heat map of each frame in the visual salience candidate index set, extracts the pixel coordinate set whose heat value in the previous frame is higher than the image heat mean, marks this set as the focus reference region, and obtains the focus region coordinate set;

[0065] To call the focus heat map for each frame in the visual salience candidate set, the corresponding focus intensity distribution information must first be extracted from each frame. Specifically, the grayscale values ​​of all pixels in the focus heat map are counted, and the mean heat value of all pixels in the frame is calculated as the focus reference value for the current frame. Based on this reference value, a pixel-by-pixel judgment is then performed, selecting a set of pixels with heat values ​​above this mean as the high-focus area. For example, a frame with a size of 1280×720, i.e., 921,600 pixels, is selected. Assuming the mean of all pixel values ​​in the image heat map is 0.305, pixels with heat values ​​greater than 0.305 are extracted from the image to form a focus pixel set. The pixel coordinates of this set are marked as the focus reference area for the frame. To further accurately locate the boundaries of the focused area, the density of focused pixels can be counted by horizontal and vertical coordinates. If the focused pixel density of more than five consecutive coordinate points in a certain direction is higher than twice the average focused pixel density of the entire image, the area is considered to have significant focused features and is retained. The remaining isolated focused points are removed. In actual operation, if the focus mean of a frame image is 0.305 and a total of 312,580 focused points are extracted, the focused area accounts for 33.91% of the entire frame image. The pixel coordinates of this area will form the focused area coordinate set of the frame, which will be used for the next spatial matching operation, and finally the focused area coordinate set is obtained.

[0066] The mutation region positioning submodule locates the image region corresponding to the grid where the hue mutation occurs in the current frame based on the focus region coordinate set, extracts the grid boundary coordinates of the image region and uniformly maps them to the image coordinate system to obtain the mutation region coordinate set;

[0067] The image region corresponding to the grid where the hue mutation occurs in the current frame is located according to the focused region coordinate set. First, the position of the hue mutation region in the image coordinate space needs to be determined. The operation process is to extract the grid index determined as the hue mutation in the previous step and map it to the actual pixel coordinate interval through the grid division structure of the image frame. Assume that the image has a resolution of 1280×720 and is divided into 32×18 grids, and each grid is 40×40 pixels. If the hue mutation occurs at the grid number (8, 6) in the current frame, then the corresponding upper left pixel in the image is (320, 240) and the lower right pixel is (359, 279). Therefore, the coordinate interval of the mutation region is determined to be [(320, 240), (359, 279)]. All pixels in this area are the coordinate set of the mutation region in the current frame. Then, based on the coordinate system consistency of the focused region coordinate set and the mutation region coordinate set, the coordinates of this area are unified to the overall image coordinate system, and the spatial boundary of the hue mutation region is established, providing basic data support for the spatial coincidence calculation, and finally the mutation region coordinate set is obtained.

[0068] The spatial overlap judgment submodule calls the focus area coordinate set and the mutation area coordinate set, counts the number of intersections of pixel coordinate points in the two sets, and calculates the proportion of the intersection in the mutation area. It then judges the proportion based on the set spatial overlap rate threshold to filter out the set of frames with sudden violations.

[0069] The focus region coordinate set and the mutation region coordinate set are called and their spatial relationship is compared. The number of intersections between the two coordinate sets is counted, and the degree of spatial overlap between the mutation region and the focus region is determined based on this. Specifically, each pixel coordinate in the mutation region coordinate set is iterated over to determine whether it exists in the focus region coordinate set. If so, it is counted as an intersection point. After counting all intersection points, the number is divided by the total number of pixels in the mutation region to calculate the spatial overlap ratio. For example, in a frame with 1600 pixels in the mutation region and 320 pixels in the intersection, the overlap ratio is 320 ÷ 1600 = 0.20, meaning that 20% of the mutation region in the current frame falls within the focus region of the previous frame. A threshold for the spatial overlap ratio is set at 0.3, and based on this threshold, if the calculated overlap ratio is less than 0.3, the mutation region is considered not to be part of the previous frame's focus region, and the current frame is marked as a non-focused frame. Through this judgment rule, frames with visual changes caused by lens movement or focus shift can be eliminated, and only frames that appear in non-focus areas are retained to form the illegal picture sudden frame set, and finally the illegal picture sudden frame set is output.

[0070] See also Figure 4 , the modal mutation behavior recognition module includes:

[0071] The modal sequence extraction submodule obtains the time segment covered by all frame indexes in the set of frames with the offending image appearing, extracts the video image frame sequence and audio signal sequence corresponding to the time segment, extracts the displacement amplitude of the center point of the character's bounding box and the number of focused heat areas from the video image frame sequence, and extracts the speech start and end time periods and the change amplitude of the sound source position from the audio signal sequence, thereby obtaining image and audio modal features.

[0072] After obtaining the time segment covered by all frame indices in the set of frames where the illegal screen appears, the corresponding start and end frame numbers in the video timeline should be determined based on the frame index. The corresponding video image frame sequence and audio signal sequence should be extracted using this time period as a window. The image frame sequence extraction must maintain frame rate consistency. If the video frame rate is 25 frames per second and the frame range is from the 100th frame to the 150th frame, the time period is 2 seconds and the number of image frames is 51. The audio signal corresponds to the continuous audio stream within this time period. When processing the image frame sequence, the target detection model is first used to obtain the bounding box of all human targets in the image, extract the position of the center point of each bounding box, and record the spatial displacement coordinates of the character in each frame. The inter-frame displacement amplitude is calculated by comparing the coordinates of the center position of the character in adjacent frames. For example, the center point of the character in the 101st frame is (520, 400) and that of the 102nd frame is (528, 408). The displacement between the two frames is Pixels are recorded frame by frame to form a time series of character displacement. At the same time, combined with the focus heat map of the frame image, the number of image areas with a heat value greater than 0.6 is counted as the number of focus heat areas, reflecting the density of the area where visual attention is concentrated. The extraction of audio modal features requires separating speech segments from the audio signal and detecting the start and end time of the speech. For example, in a 2-second audio segment, speech activity is detected from 0.3 seconds to 1.8 seconds, which is recorded as a speech interval. In addition, the spatial position of the sound source of each speech segment is estimated based on the microphone array or the sound arrival time difference method, and the amplitude of the position change of adjacent speech segments is compared to characterize the spatial dynamic behavior of the audio modality. Finally, the extraction of two types of behavioral features of image modality and audio modality is completed to obtain image and audio modal features.

[0073] The dominant weight calculation submodule calculates the mutual information value between the image and audio modal features respectively, constructs the mutual information sequence at the corresponding frame level, and assigns the dominant weight of the image modality or audio modality to each frame according to the size of the mutual information, thus obtaining the modality dominant weight data sequence;

[0074] To calculate the mutual information value between the image and audio modal features, it is necessary to construct a joint feature pair sequence based on the frame-level feature data, and calculate the inter-frame modal joint distribution and marginal distribution based on this to obtain the mutual information value. First, the image modal feature is represented as a two-dimensional vector ,in Indicates the inter-frame displacement amplitude of the center point of the character's bounding box in the frame, in pixels, obtained by calculating the Euclidean distance between the center coordinates of two adjacent frames. Indicates the number of regions with focus heat values ​​greater than 0.6 in this frame, which can be achieved by counting the pixel regions in the focus heat map. The audio modality feature is represented as ,in It is the speech activity flag. If the frame is within the speech start and end interval, it is assigned a value of 1, otherwise it is 0. Indicates the amplitude of the change in the sound source position, which can be estimated by the spatial coordinate distance between the sound source positioning points of the previous and next frames, in meters.

[0075] The formula for calculating mutual information is:

[0076] ;

[0077] in: is the complete set of image modality features, consisting of the image modality feature vectors of all frames constitute; is the complete set of audio modal features, consisting of the audio modal feature vectors of all frames constitute; Indicates from the complete works Select a certain image modality feature combination from , such as Indicates that the character is displaced by 12.3 pixels, corresponding to 5 focus heat areas; Indicates from the complete works Select a certain audio modal feature combination from Indicates that the frame is a speech interval and the sound source changes by 0.25 meters; The image modality feature is represented as And the audio modal characteristics are Joint probability of simultaneous occurrence; and They are image modality features and audio modal features The marginal probability distribution of .

[0078] These probability values ​​are obtained by the frequency of occurrence of statistical feature combinations in the full frame data. For example, in 50 frames of samples, a certain combination If it occurs 4 times, the joint probability is: , if the If it appears 10 times in all samples, , if the The frequency of occurrence is 8 times, then . Substitute the mutual information term: , for all and The same operation is performed on the combination of , and the results are accumulated to obtain the mutual information value of the frame . This value reflects the strength of the association between the image and audio modal features. In the mutual information sequence, if the mutual information value of a frame is less than 0.3, it means that the independence between the modalities is strong, and the frame can be preliminarily marked as a single-modal dominant frame; if the mutual information value is greater than 0.5, it means that the image and audio behavioral features have significant synchronization, and the frame is marked as a bimodal collaborative frame. The sequence generated by arranging the mutual information values ​​and dominant labels of all frames in sequence is the modal dominant weight data sequence.

[0079] The jump segment identification submodule detects the weight change trend of the modal dominant weight data sequence in continuous frame segments, identifies the time period when the weight dominant type has continuous mutations, selects the frame segments with a change span greater than the jump judgment criterion as the jump area, and obtains the set of modal mutation behavior segments;

[0080] After obtaining the modal dominant weight data sequence, the dominant weight values ​​of consecutive frames in the sequence need to be detected for changing trends in order to identify the temporal characteristics of modal dominant hopping behavior. First, the modal dominant weight data sequence is organized into a one-dimensional time series by frame number. The dominant mode of each frame is labeled as image mode, audio mode, or bimodal mode based on the dominant judgment result of the mutual information value, and encoded as +1, -1, and 0, respectively. This encoded sequence is then traversed and analyzed using a sliding window of fixed length (typically 5 frames). The difference between the maximum and minimum values ​​in each window is calculated and used as a measure of hopping strength. For example, if a segment of the encoded sequence is [+1, +1, -1, -1, -1], then the maximum value of this segment is +1, the minimum value is -1, and the hopping span is 2. The jump determination criterion is set to 1.5. This value is derived from experimental data statistics. During the experiment, it was found that in conventional background static images, the modal change range is generally within the range of [-1, 1], with the jump amplitude not exceeding 1.2. However, in scenes with rapid camera cuts or sudden voice intrusion, the amplitude exceeds 1.5. Therefore, 1.5 is used as the limit for determining sudden behavior. When a jump span greater than 1.5 is detected within a window, the start and end frame numbers of the frame segment covered by the window are recorded, and the window is continued to slide backward to detect the next segment. Taking actual data as an example, if the encoding values ​​of frames 75 to 79 are [-1, -1, +1, +1, +1], the jump span is 2, meeting the sudden behavior condition, and the frame segment [75, 79] is recorded. This detection process is repeated, and finally all frame segments with a jump amplitude greater than the threshold are extracted and formed into a structured frame segment set, named the modal sudden behavior segment set.

[0081] See also Figure 5 , the modal interleaving interruption node identification module includes:

[0082] The image interruption extraction submodule obtains the image frame sequence in the modal mutation behavior segment set, extracts the center position coordinates of the character bounding box between adjacent frames, calculates the Euclidean distance change rate between frames, determines the image interruption frame location, and obtains the image interruption frame index set;

[0083] After obtaining the image frame sequence in the modal mutation behavior segment set, first perform person detection on each frame to extract the bounding box information. The bounding box represents the spatial range of the target area in the form of a rectangle. The bounding box is defined by the coordinates of its upper left corner. and the coordinates of the lower right corner The only certainty is that is the starting pixel position in the horizontal and vertical directions, is the end position of the bounding box, and the center point of the box is used as the target center coordinate. The formula is:

[0084] ;

[0085] in, is the horizontal coordinate of the center point, is the vertical coordinate of the center point, which represents the visual center of gravity of the character in the image frame. and The character center coordinates are used to calculate the character movement distance between the two frames. The Euclidean distance metric is used. The calculation formula is as follows:

[0086] ;

[0087] in, Indicates the Frame to The displacement distance of the character's center position between frames, 、 It is The horizontal and vertical coordinates of the frame center point, 、 It is The horizontal and vertical coordinates of the center point of the frame. Furthermore, the rate of change of displacement between frames is defined. , used to determine the degree of mutation of the position change, and the calculation method is:

[0088] ;

[0089] in, For the The frame displacement change rate, in pixels / second, and Respectively represent arrive frames, and arrive The character displacement distance between frames, is the time interval between frames. If the video frame rate is 25 frames per second, then Second.

[0090] Assume that the upper left corner of the character's bounding box in frame 100 is (490, 280) and the lower right corner is (510, 320), then the center point is , the bounding box of the 101st frame is (505, 290)-(525, 330), and the center point is , the 102nd frame is (540, 320)-(560, 380), and the center point is , substituting into:

[0091] ;

[0092] ;

[0093] .

[0094] Set the image interruption judgment threshold is 850 pixels / second, if , it is considered that there is a sudden change in image motion in this frame. Therefore, the 100th frame is determined to be an image interruption frame, and finally all frame indexes that meet this condition are aggregated to form an image interruption frame index set.

[0095] The speech interruption extraction submodule obtains the amplitude energy of each frame based on the audio signal sequence of the time period corresponding to the image interruption frame index set, determines the frame index with energy lower than the silence threshold as the speech interruption frame location, and obtains the speech interruption frame index set;

[0096] Based on the time period corresponding to the image interrupt frame index set, the audio signal sequence is extracted and time-synchronized and framed according to the video frame rate. Assuming the video frame rate is 25 frames per second, each frame corresponds to 40 milliseconds of audio. When the sampling rate is 16kHz, each frame of audio contains 640 sample points, that is, , the energy of the sample data in each frame of audio segment is estimated, and the root mean square energy (RMS) is used as the detection basis. The RMS energy calculation formula is as follows:

[0097] ;

[0098] in, For the Frame audio energy, For the frame The amplitude value of the sample point, is the number of sampling points per frame. Taking a frame of audio sample as an example, let the sum of the square values ​​of the amplitude of the sample points in a certain frame be 1638.4, that is , then the RMS energy is:

[0099] ;

[0100] Calculate the energy for all frames sequentially , and then with the mute threshold The setting of silence threshold needs to be combined with the distribution of environmental noise and the statistics of voice signal strength. Generally, the RMS energy of background environmental noise is concentrated between 0.1 and 0.3, and the RMS energy of ordinary voice is mostly between 0.5 and 2.0. During the experiment, energy analysis of multiple voice segments and silence segments is performed to set the silence threshold. It can effectively divide silent and sound frames. , it is marked as a speech interruption frame. For example, if the audio energy of the 120th frame is 0.22, it is lower than the silence threshold and is identified as a speech interruption frame.

[0101] To ensure the stability of the results, an extended detection range of ±2 frames before and after each video interruption frame is used to form a 5-frame audio window. Within each window, the detection is performed to determine whether consecutive frames within the window have energy levels below the threshold. If the condition is met, the video interruption frame is used as the reference to confirm a voice interruption. This process is repeated to complete the silence determination for the time period corresponding to all video interruption frames. Finally, the frame numbers that meet the determination criteria are summarized to obtain the voice interruption frame index set.

[0102] The time overlap screening submodule calls the image interruption frame index set and the speech interruption frame index set, compares the time tags of the two types of frame indexes, and screens the frame points with completely consistent indexes as interleaving interruption points to generate a modal interleaving interruption node set;

[0103] After obtaining the image interruption frame index set and the voice interruption frame index set, first define the two as sets and ,in Indicates the The frame number of the image interruption frame, Indicates the The frame number of the voice interruption frame, The two sets are then sorted in ascending order so that matching operations can be performed by double pointer traversal or set operations. Indicates that Indicates the The frame number detected as the synchronization of image and voice mode interruption, the formation condition is , that is, the frame number exists in both the image interruption set and the voice interruption set. In order to consider the case where there is a slight delay in the modal response, the allowable error tolerance is introduced , the set value is usually 1 frame. This value is derived from the frame rate inversion. At 25fps, each frame time is 0.04 seconds. In practical applications, the time difference of ±1 frame is acceptable. The expanded judgment condition is written as:

[0104] ;

[0105] in, :The modal interleaving interruption node set Synchronous interrupt frame number; : Image interrupt frame index set Frame number; : The first index in the voice interruption frame set Frame number; : Existential symbol, indicating "something exists"; : The absolute value of the difference between the image frame and the audio frame number does not exceed the set error tolerance ; : Frame number tolerance range, set to 1 frame.

[0106] If the image is interrupted frame index set , voice interruption frame index set , under the condition of no error tolerance, the intersection is , but with frame error allowed Under the conditions, it is acceptable and ,because ,Will Also included in the collection For frame points whose difference is within 1 but not completely overlapped, a strategy can also be used to specify the representative frame number, such as taking the smaller number , or use weighted average. Finally, we get the set of modal interleaving interruption nodes , which is used for subsequent multimodal structure alignment and accurate identification of behavioral fracture areas.

[0107] See also Figure 6 , the video content review module includes:

[0108] The weight change identification submodule obtains the frame index in the modal interleaving interruption node set, extracts the corresponding modal dominant weight data before and after the corresponding frame, calculates the inter-frame change slope of the image modal weight, and identifies the frame segments with a continuous upward trend to obtain the image modal weight change interval;

[0109] After obtaining the frame index of each frame in the modal interleaving interruption node set, it is necessary to extract the modal dominant weight value sequence of several frames before and after the frame index, and calculate the slope of the image modal weight change with the frame sequence based on this, so as to identify the continuously rising frame segment interval. First, define the modal dominant weight sequence as ,in Indicates the The image modality dominance weight value of the frame is derived from the mutual information analysis process and is used to measure the degree of influence of the image modality on the current frame. The value range is [0, 1]. The higher the value, the more the frame is dominated by the image modality. (Right now ,in Set a fixed window length for the modal interleaving interrupt node set , from the frame To Frame Extraction Image modality weights of consecutive frames, weight change subsequence , and establish the corresponding time series ,in Indicates the frame number.

[0110] In order to describe the trend of the dominant weight of the image modality over time, the extracted sequence is fitted using linear regression. The core calculation in the fitting process is the slope , represents the rate of change of the weight value per frame. The specific formula is as follows:

[0111] ;

[0112] in: , is the number of frames involved in the calculation; For the The frame number of the sample; For the The image modality dominant weight of samples; , is the average value of the time series; , is the average value of the weight value sequence; Represents the slope. A value greater than 0 indicates that the weight increases with the number of frames.

[0113] set up , considering extracting the sequence centered on frame number 130, the corresponding frame numbers are {128, 129, 130, 131, 132}, and the corresponding image modality weight values ​​are {0.35, 0.42, 0.55, 0.61, 0.68}. Calculate the average frame number: ;

[0114] Calculate the average weight value: ;

[0115] Construct the numerator term: 0.522) + (130-130) (0.55-0.522) + (131-130) (0.61-0.522) + (132-130) (0.68-0.522) = 0.85;

[0116] Construct the denominator term: ;

[0117] Finally, we get the slope: .

[0118] It can be obtained that the slope of the image modality dominant weight of frame segment 128–132 is 0.085, which is a positive value and greater than the set change sensitivity threshold of 0.05, indicating that the image modality dominant weight of this frame segment shows a significant upward trend. If consecutive frame segments meet this condition, the frame segment number sequence is determined to be the image modality weight change interval, recorded as a set and used as the input of the subsequent anomaly detection module.

[0119] The abnormal segment determination submodule determines whether the frame index of the image modal weight change interval is included in the set of frames with illegal images. If included, the target frame segment is recorded and the modal dominant jump frequency and the frame index distribution density within the frame segment are calculated to obtain the modal abnormal jump score.

[0120] When obtaining the image modality weight change interval set After that, we need to process the start and end frame numbers of each frame segment in the set one by one. Let a segment be represented as ,in Indicates the starting frame number of the segment, Indicates the end frame number of the segment. First, all frame numbers in the segment are combined , and the set of frames with the offending screen pop-up Perform intersection comparison. Each element in Indicates the frame number that was identified as having visual abrupt changes, color jumps, or focus shifts in the previous analysis. There is at least one frame number that satisfies , that is, the frame segment contains a sudden illegal frame, then the segment is retained and enters the abnormality score calculation process.

[0121] First, calculate the modal dominant hopping frequency of the frame segment. It is necessary to extract the dominant modal label of each frame in the frame segment and construct the modal label sequence ,in Indicates the The dominant mode type of the frame. The number of inconsistent mode labels between consecutive frames in the sequence is counted as the number of jumps. . Hopping frequency It is defined as the frequency of the dominant change of the mode, that is: ,in Indicates the length of the frame segment, that is, the total number of frames. For example, if the frame segment is 140 to 149, the modal sequence is [image, image, audio, audio, bimodal, image, image, audio, image, image], where the modal change occurs in frames 141→142, 143→144, 144→145, 146→147, 147→148, totaling times, frame length , then the hopping frequency is: , Secondly, calculate the distribution density of the illegal frames in the frame segment. Let the frame segment and the illegal set The set of overlapping frame numbers is , which means: at the start frame and end frame Within the defined range, extract all the violations that also belong to the set A new set of frame numbers ,in Indicates the first frame number, as long as the frame number satisfies and , then the frame is included in the set , which is used for subsequent statistics. The number of its elements is recorded as .density It is defined as the proportion of illegal frames in the segment, that is: , assuming that there are 7 frames in segments 140 to 149 that appear simultaneously in the violation set In, then: Finally, the modal hopping frequency and the illegal frame density are combined to form an abnormality score , set the weight coefficient 、 , then the scoring formula is: , substituting the above data into 、 ,but: ,score , which is used to measure the comprehensive deviation of the frame segment in terms of modal dominant volatility and violation concentration. For each frame segment in the image modal weight change interval set After executing this process, all calculated score values ​​are combined into a modal abnormal transition score set.

[0122] The audit result output submodule marks the corresponding start and end indexes of the frame segments on the timeline based on the modal abnormal jump score, constructs a multi-modal behavior abnormality identification structure, and summarizes and generates the video content audit results;

[0123] After obtaining the score of each modal abnormal jump and its corresponding frame segment start and end numbers, these frame segments need to be annotated according to the score value to construct the final video anomaly identification structure. First, a scoring judgment baseline is set to distinguish whether a frame segment needs to be marked as an abnormal frame segment. The scoring baseline is fitted by actual statistical data, usually derived from the analysis results of a large number of video manual review samples. For example, by comparing the scores of frame segments that have been judged as abnormal in the past with the score distribution of normal frame segments, a reasonable critical score is determined. Assuming that the value is 0.5, all frame segments with a score higher than 0.5 will be marked as abnormal frame segments. Subsequently, all frame segments that meet the conditions are processed one by one in chronological order, and each frame segment is structurally annotated, including the frame segment start frame number, end frame number and the abnormal score value of the segment.

[0124] For example, a certain frame sequence from frame 200 to frame 212 was identified as having a continuously increasing image modality dominance weight, and 9 of the frames were frames where illegal images suddenly appeared. At the same time, the modality dominance frequently switched between image, audio, and dual modalities during this period. The final anomaly score of this segment was 0.68, which exceeded the score baseline, so this segment was recorded in the anomaly output structure. Another frame sequence from frame 280 to frame 285 had a score of 0.48, which was lower than the baseline and was not output marked. Frame segments marked as abnormal will have an abnormal identification added to the timeline. The identification structure will be indexed by the frame segment number and can be used for subsequent video content risk analysis, prompt interface highlighting, intelligent review interface linkage, and other operations. Multiple frame segment identifications will be sequentially organized into a complete video abnormal frame segment list, which will be output as part of the review structure. For example, during one review, the system output three abnormal segments: frames 200–212, frames 340–350, and frames 420–430, with corresponding scores of 0.68, 0.74, and 0.82. The system aggregates these segments and identifies them as key areas where the video exhibits abrupt changes in multimodal behavior and a concentration of illegal content. This structure can also be converted into interface data or JSON format.

[0125] A method for dynamic security review of multimodal content of video streams is provided. The method is based on the above-mentioned dynamic security review system for multimodal content of video streams and includes the following steps:

[0126] S1: Extract consecutive frames of the audit video and construct the image's hue ratio vector and focus heat map. Filter the image frames with sudden hue ratio mutations in the local frame segments of the sequence to obtain a set of candidate visual salience indicators.

[0127] S2: Calculate the spatial distribution difference of the sudden hue area in the focus heat map based on the visual salience candidate index set to determine the set of illegal sudden hue frames;

[0128] S3: Extract the image and audio multimodal sequences in the time segment covered by the set of frames where the offending images appear, identify the continuous frame segments where the modality dominance undergoes a jump change, and construct a set of modality mutation behavior segments;

[0129] S4: Identify the modal interleaving interruption frame locations where the image action and voice activity time coincide in the modal mutation behavior segment set, and form a modal interleaving interruption node set;

[0130] S5: Review the interruption behavior in the video based on the modal interleaving interruption node set, and output the video content review result.

[0131] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A video stream multimodal content security dynamic audit system, characterized by: The system comprises: The visual emergence candidate indicator extraction module extracts consecutive frames of the audit video and constructs the image's hue ratio vector and focus heat map. It then selects image frames with sudden hue ratio mutations in local frame segments of the sequence to obtain a set of visual emergence candidate indicators. The illegal picture determination module calculates the spatial distribution difference of the sudden change color area in the focus heat map based on the visual salience candidate indicator set, and determines the illegal picture salience frame set; The modal mutation behavior recognition module extracts the image and audio multimodal sequences in the time segment covered by the set of frames with sudden appearance of the offending images, identifies the continuous frame segments where the modality dominant undergoes a jump change, and constructs a set of modal mutation behavior segments; The modal interleaving interruption node identification module identifies modal interleaving interruption frame locations where the image action and the voice activity time overlap in the modal mutation behavior segment set to form a modal interleaving interruption node set; The video content review module reviews the interruption behavior in the video based on the modal interleaving interruption node set and outputs the video content review result.

2. The video stream multimodal content security dynamic audit system according to claim 1 is characterized in that: The candidate indicator set for visual emergence specifically includes image frames with sudden changes in hue ratio, local frame segments, and hue ratio sudden changes; the set of frames with sudden changes in illegal images includes sudden changes in hue area, spatial distribution differences, and illegal images; the set of modal mutation behavior segments specifically includes image motion changes, voice activity changes, and continuous frame segments; the set of modal interleaving interruption nodes includes modal interleaving interruption frame sites and frame sites where image motion and voice activity overlap in time; the video content review results include review conclusions and video interruption behaviors.

3. The video stream multimodal content security dynamic audit system according to claim 1 is characterized in that: The visual salience candidate index extraction module includes: The tone proportion vector construction submodule obtains the complete frame image sequence of the audit video, divides each frame into a fixed-size grid, calculates the main tone type of each grid and the area ratio of the main tone type in the entire frame image, and generates a tone proportion vector set for each frame; The focus heat map generation submodule converts each frame of image into a grayscale image, calculates the pixel clarity gradient using the Laplacian operator, and smoothes it through Gaussian filtering to generate a focus heat map for each frame of image with the pixel gradient amplitude as the weight; The mutation frame screening submodule compares the hue proportion vector set of each frame and the focus heat map of each frame image for consecutive frames, screens frames with hue mutations, and obtains a set of visual salience candidate indicators.

4. The video stream multimodal content security dynamic audit system according to claim 3 is characterized in that: The illegal screen determination module includes: The focus region extraction submodule calls the focus heat map of each frame of the visual salience candidate index set, extracts the pixel coordinate set whose heat value in the previous frame is higher than the image heat mean, marks the target set as the focus reference region, and obtains the focus region coordinate set; The mutation region positioning submodule locates the image region corresponding to the grid where the hue mutation is located in the current frame according to the focus region coordinate set, extracts the grid boundary coordinates of the image region and uniformly maps them to the image coordinate system to obtain the mutation region coordinate set; The spatial overlap judgment submodule calls the focus area coordinate set and the mutation area coordinate set, counts the number of intersections of pixel coordinate points in the two sets, and calculates the proportion of the intersection number in the mutation area. The proportion is judged according to the set spatial overlap rate threshold to filter out the set of frames with illegal screen bursts.

5. The video stream multimodal content security dynamic audit system according to claim 4 is characterized in that: The modal mutation behavior recognition module includes: The modal sequence extraction submodule obtains the time segment covered by all frame indexes in the set of frames where the illegal image appears, extracts the video image frame sequence and audio signal sequence corresponding to the time segment, extracts the displacement amplitude of the center point of the character's bounding box and the number of focused heat areas from the video image frame sequence, and extracts the speech start and end time periods and the amplitude of the sound source position change from the audio signal sequence to obtain image and audio modal features; The dominant weight calculation submodule calculates the mutual information value between the image and audio modal features respectively, constructs a mutual information sequence at the corresponding frame level, assigns a dominant weight of the image modality or audio modality to each frame according to the mutual information value, and obtains a modality dominant weight data sequence; The jump segment identification submodule detects the weight change trend of the modal dominant weight data sequence in continuous frame segments, identifies the time period in which continuous mutations occur in the weight dominant type, selects the frame segments with a change span greater than the jump judgment criterion as the jump area, and obtains the modal mutation behavior segment set.

6. The video stream multimodal content security dynamic audit system according to claim 5 is characterized in that: The modal interleaving interruption node identification module includes: The image interruption extraction submodule obtains the image frame sequence in the modal mutation behavior segment set, extracts the center position coordinates of the character boundary box between adjacent frames, calculates the Euclidean distance change rate between frames, determines the image interruption frame location, and obtains the image interruption frame index set; The voice interruption extraction submodule obtains the amplitude energy of each frame based on the audio signal sequence of the time period corresponding to the image interruption frame index set, determines the frame index with energy lower than the silence threshold as the voice interruption frame position, and obtains the voice interruption frame index set; The time overlap screening submodule calls the image interruption frame index set and the voice interruption frame index set, compares the time tags of the two types of frame indexes, screens frame points with completely consistent indexes as interleaving interruption points, and generates a modal interleaving interruption node set.

7. The video stream multimodal content security dynamic audit system according to claim 6 is characterized in that: The video content review module includes: The weight change identification submodule obtains the frame index in the modal interleaving interruption node set, extracts the corresponding modal dominant weight data before and after the corresponding frame, calculates the inter-frame change slope of the image modal weight, and identifies the frame segments with a continuous upward trend to obtain the image modal weight change interval; The abnormal segment determination submodule determines whether the frame index of the image modal weight change interval is included in the set of illegal picture burst frames. If included, the target frame segment is recorded, and the modal dominant jump frequency and the frame index distribution density within the frame segment are calculated to obtain the modal abnormal jump score; The audit result output submodule marks the corresponding start and end indexes of the frame segments on the time axis according to the modal abnormal jump score, constructs a multi-modal behavior abnormality identification structure, and summarizes and generates the video content audit results.

8. A method for dynamic security audit of multimodal content of video streams, characterized in that: The video stream multimodal content security dynamic audit system according to any one of claims 1 to 7 is implemented, comprising the following steps: S1: Extract consecutive frames of the audit video and construct the image's hue ratio vector and focus heat map. Filter the image frames with sudden hue ratio mutations in the local frame segments of the sequence to obtain a set of candidate visual salience indicators. S2: Calculating the spatial distribution difference of the sudden change hue area in the focus heat map based on the visual salience candidate indicator set, and determining the set of illegal picture salience frames; S3: extracting the image and audio multimodal sequences in the time segment covered by the set of frames with the offending images appearing suddenly, identifying the continuous frame segments where the modality dominant undergoes a jump change, and constructing a set of modality mutation behavior segments; S4: identifying modal interleaving interruption frame locations where the image action and the voice activity time coincide in the modal mutation behavior segment set, and forming a modal interleaving interruption node set; S5: Review the interruption behavior in the video based on the modal interleaving interruption node set, and output the video content review result.

Citation Information

Patent Citations

  • A multimodal video scene segmentation method based on sound and vision

    CN109344780A

  • Multi-modal data automatic cleaning and labeling method and system

    CN111767805A