Internet video live broadcast sensitive information monitoring method based on artificial intelligence
By combining dynamically manipulated differential convolutional basis and Riemannian manifold mapping techniques with sparse masking and frame-level violation classification, the problem of identifying rotation-sensitive information in Internet video live streaming is solved, achieving high-precision, real-time violation target detection.
Patent Information
- Application Number
- CN202511912634.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to effectively identify sensitive information such as rotation, tilt, or non-rigid transformations in live internet video streaming, and are susceptible to interference from high-saturation colors and dynamic noise, leading to frequent missed or false detections.
A dynamically manipulated differential convolutional base module is used to predict the dominant gradient direction of pixels, generate a manipulation coefficient matrix, and generate a direction-adaptive differential feature map through differential convolution. Combined with Riemannian manifold mapping and sparse masking, key regions are extracted, and a high-precision detection is achieved using an instantaneous frame-level violation classification subnetwork.
It achieves high-fidelity, zero-latency detection of rotation and tilt-sensitive information, suppresses background noise interference, improves the recognition rate and real-time performance of illegal targets in live streams, and meets the real-time monitoring needs of live content security.
Smart Images

Figure CN121708528A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Internet video live streaming technology, and in particular to an artificial intelligence-based method for monitoring sensitive information in Internet video live streaming. Background Technology
[0002] With the rapid development of online live streaming platforms and the continuous improvement of content review requirements, real-time compliance monitoring of live video streams has become an important research direction in the field of internet content security. In actual live streaming scenarios, gray and black market operators and creators of illegal content have used a variety of technical means to bypass automated review systems.
[0003] Common circumvention methods include rotating, tilting, or non-rigidly deforming sensitive text, QR codes, and illegal icons at arbitrary angles, making them difficult for existing models to accurately identify. Furthermore, illegal information is often inserted in a fleeting manner, appearing briefly for only 1-2 frames, significantly increasing the difficulty of detection. Meanwhile, the background of live streams is usually extremely complex, containing highly saturated colors, frequently flashing LED lights, and interfering factors such as bullet comments, further increasing the intensity of noise interference.
[0004] In existing technologies, most mainstream deep learning object detection methods employ convolutional neural network structures. However, traditional convolutional kernels lack inherent invariance to rotational and tilted geometric transformations, resulting in limited ability to identify illegal targets under arbitrary angle changes. To improve the model's adaptability to rotating targets, existing methods mainly rely on data augmentation techniques or spatial transformation networks. However, the former significantly increases the number of model parameters and training complexity, while the latter easily introduces interpolation errors and inference delays, affecting recognition accuracy and real-time performance. Meanwhile, conventional attention mechanisms are primarily based on pixel statistical characteristics in Euclidean space, making them susceptible to the effects of live stream background brightness and dynamic noise, leading to attention drift and difficulty in accurately focusing on fine-grained illegal targets that appear instantaneously in the live stream, resulting in frequent missed or false detections. Summary of the Invention
[0005] One objective of this invention is to propose an artificial intelligence-based method for monitoring sensitive information in live internet video streaming. This invention significantly improves the effective frame capture capability in high-noise live streaming environments.
[0006] According to an embodiment of the present invention, a method for monitoring sensitive information in live internet video streaming based on artificial intelligence includes: Acquire a continuous frame sequence of images from a live video stream and perform preprocessing to generate preprocessed frame images; The dominant gradient direction of each pixel in the preprocessed frame image is predicted using the orientation prediction subnetwork, and the manipulation coefficient matrix is output. The manipulation coefficient matrix and the preprocessed frame image are synchronously input into the dynamic manipulatory differential convolutional base module. The preprocessed frame image is convolved in parallel using the differential base function set and then linearly combined at the pixel level according to the manipulation coefficient matrix to generate an orientation adaptive differential feature map. Extract multi-scale fusion features from the directional adaptive difference feature map to generate a multi-scale fusion feature map; The multi-scale fused feature map is projected onto the Riemannian manifold space through nonlinear mapping, and the Riemannian manifold mapping feature map is output. A background reference manifold is constructed based on the Riemannian manifold mapping feature map, and the geometric deviation between each local region and the background reference manifold is calculated to generate a geometric deviation mapping map. A topological anomaly saliency map is generated based on the geometric deviation mapping map. A sparse mask generation module is then applied to the topological anomaly saliency map to extract the Top-K key regions and output a sparse mask. The sparse mask is then applied to a multi-scale fusion feature map to obtain a sparse candidate region feature map. The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork. Local classification reasoning is performed on each frame, and violation labeling results are output. When the violation labeling results indicate the existence of a violation target, the instantaneous frame locking mechanism is triggered, the corresponding frame and its preceding and following context frames are recorded, violation frame capture results are generated, and the violation frame capture results are encapsulated and output as an evidence chain to complete the instantaneous violation frame capture of the live video stream.
[0007] Optionally, the preprocessing includes performing normalization, brightness normalization, and region mask initialization on the continuous frame sequence images of the live video stream to generate preprocessed frame images.
[0008] Optionally, the step of using a direction prediction sub-network to predict the dominant gradient direction of each pixel in the preprocessed frame image includes: The preprocessed frame image tensor is input into the orientation prediction subnetwork, and the orientation feature map is output. Based on the orientation feature map, a dominant gradient vector is constructed at each spatial location, and the dominant gradient vector is converted into the dominant gradient direction angle on the two-dimensional plane through the arctangent function. Define a set of direction basis angles that correspond one-to-one with the set of difference basis functions; Based on the set of dominant gradient direction angles and direction basis angles, a manipulation coefficient matrix is constructed.
[0009] Optionally, the step of synchronously inputting the manipulation coefficient matrix and the preprocessed frame image into the dynamically manipulated differential convolutional basis module includes: The preprocessed frame image tensor and the manipulation coefficient matrix are synchronously input into the dynamically manipulated differential convolutional base module; In the dynamically manipulated differential convolutional basis module, local directional complexity weights based on the perturbation of the dominant gradient direction are constructed; In the dynamically manipulated differential convolution basis module, a set of differential basis functions is set, and multi-directional differential convolution operations are performed on the preprocessed frame image tensor to obtain the differential basis response tensor. By combining local directional complexity weights, instantaneous violation sensitivity modulation is applied to the differential basis response tensor to generate a complexity-modulated differential basis response tensor; Pixel-level linear combination of the complexity modulation differential basis response tensor is performed based on the manipulation coefficient matrix to generate an orientation-adaptive differential feature map.
[0010] Optionally, the multi-scale fusion features of the extracted orientation adaptive difference feature map include: The orientation-adaptive differential feature map is input into a multi-scale fusion network to construct a feature fusion structure with different scale branches, forming a set of multi-scale differential feature sub-maps. Independent convolutional encoding is performed on the multi-scale differential feature sub-map set to extract local spatial features within the scale, forming a set of coded feature maps within the scale. In a multi-scale fusion network, the molecular feature map of orientation difference with spatial resolution higher than the threshold is mapped to a low-scale level through downsampling operation, and the molecular feature map of orientation difference with spatial resolution lower than the threshold is mapped to a high-scale level through upsampling operation. Within each scale level, the original scale orientation difference molecular feature map is fused with all the mapped cross-scale features to obtain the inter-scale fusion feature map. Based on the intra-scale encoded feature map set and the inter-scale fused feature map, a cross-scale residual connection mechanism is used for joint fusion, and a fused activation map is constructed by feature weighting superposition. The fused activation map is input into the channel compression module, which compresses the number of feature channels along the channel dimension using a set of one-dimensional convolutional kernels with learnable parameters, and outputs a multi-scale fused feature map.
[0011] Optionally, the construction of the background reference manifold based on the Riemannian manifold mapping feature map includes: The multi-scale fused feature map is input into the nonlinear mapping module to construct a symmetric positive definite matrix feature tensor, which completes the mapping from Euclidean space to Riemannian manifold space and outputs the Riemannian manifold mapping feature map. Based on the statistical characteristics of the spatial distribution in the feature map of the Riemannian manifold, a background reference manifold is constructed; Based on the geodesic distance between the Riemannian manifold mapping feature map and the background reference manifold, the geometric deviation at the spatial location of each pixel is calculated, and a geometric deviation mapping map is constructed.
[0012] Optionally, the output sparse mask applied to the multi-scale fused feature map includes: The geometric deviation mapping is normalized to obtain a normalized geometric deviation mapping. Generate a topological anomaly saliency map based on the normalized geometric deviation map; Local consistency aggregation is performed on the topological anomaly saliency map to suppress isolated noise and generate candidate response maps; A sparse mask generation module is applied to the candidate response map to extract the Top-K key regions and output the sparse mask. The sparse mask is applied to the multi-scale fused feature map to obtain the sparse candidate region feature map.
[0013] Optionally, the local classification reasoning for each frame includes: The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork to extract frame-level semantic features and perform local classification inference, and output the violation labeling result. When the violation marking result indicates that the current frame is a violation frame, the instantaneous frame locking mechanism is triggered, the current frame and several frames before and after it are recorded as context, the violation frame capture result is generated, and the evidence chain of the violation frame capture result is encapsulated and output.
[0014] Optionally, the violation labeling result includes: A frame is marked as a violation if any of the following conditions are met: The sparse candidate region feature map contains a region with enhanced high-frequency directional response greater than the preset value, and the area of the region exceeds the first percentage threshold of the total area of the Top-K regions. At the same time, the magnitude of the deviation of the dominant gradient direction is greater than the angle threshold. The maximum activation value of at least one channel in the feature channel exceeds the set channel threshold, and the corresponding spatial region overlaps with the suspected violation region in the historical judgment with a degree greater than the second percentage threshold. The average saliency score of the Top-K region in the current frame is higher than the historical average by a preset multiple; All other cases are marked as non-violation frames.
[0015] The beneficial effects of this invention are: (1) This invention introduces a dynamic and manipulable differential convolutional basis. By predicting the dominant gradient direction and dynamically weighting at the pixel level, it can directly and adaptively combine differential basis functions in multiple directions in the feature space without increasing the number of convolutional kernels or without data augmentation or spatial transformation networks. This enables texture response sensitive to any angle, effectively solving the problem that traditional convolutional kernels can only handle limited directions and are difficult to capture rotation or tilt sensitive information. When detecting various rotating, tilted, and non-rigid deformation illegal QR codes, sensitive texts or patterns in live streams, no physical rotation interpolation is required. This enables high-fidelity, zero-latency high-precision capture, significantly improving the model's recognition rate of geometrically transformed targets in complex scenes.
[0016] (2) This invention maps the multi-scale fused feature map to the Riemannian manifold space through nonlinear transformation, and measures the geometric deviation of each region based on the background reference manifold. It effectively utilizes the topological distribution characteristics of the data in the manifold space. On this basis, it generates a topological anomaly saliency map by normalizing the geometric deviation, and focuses on the key regions with the most anomalous features in Top-K by combining sparse masking. It suppresses the interference of background noise unrelated to highlights, clutter, and bullet comments, and achieves highly sensitive detection of extremely small and highly disguised illegal targets that appear instantaneously in the live stream. The effective frame capture capability in high-noise live streaming environment is greatly improved.
[0017] (3) Based on sparse region focusing, this invention designs an instantaneous frame-level violation classification sub-network. Through multiple standards, it achieves fine discrimination of the Top-K key regions of each frame. Combined with the violation frame locking mechanism, the system can automatically record the violation frame and its context, output the complete violation frame capture result and evidence chain, greatly reducing the overall floating-point operation volume, and ensuring that the algorithm can complete automatic detection, discrimination and tracking with millisecond-level latency in high-concurrency live streaming scenarios, effectively meeting the engineering requirements of real-time monitoring and compliant evidence storage of live streaming content security. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an artificial intelligence-based method for monitoring sensitive information in live internet video streaming, as proposed in this invention. Figure 2 This is a schematic diagram of the structure of the dynamically manipulated differential convolutional base module in the artificial intelligence-based Internet video live streaming sensitive information monitoring method proposed in this invention. Detailed Implementation
[0019] Example 1: Reference Figures 1-2 A method for monitoring sensitive information in live internet video streaming based on artificial intelligence, comprising: Acquire a continuous frame sequence of images from a live video stream and perform preprocessing to generate preprocessed frame images; In this embodiment, preprocessing includes performing normalization, brightness normalization, and region mask initialization on the continuous frame sequence images of the live video stream to generate preprocessed frame images.
[0020] The dominant gradient direction of each pixel in the preprocessed frame image is predicted using the orientation prediction subnetwork, and the manipulation coefficient matrix is output. In this embodiment, the dominant gradient direction of each pixel in the preprocessed frame image is predicted using a direction prediction sub-network, including: The preprocessed frame image tensor is input into the orientation prediction subnetwork, and the orientation feature map is output. The preprocessed frame image tensor is a three-dimensional pixel intensity tensor formed by normalizing and luminance standardizing the continuous frame sequence images of the live video stream. It is used to represent the pixel intensity of a single frame image in the live video stream in terms of spatial location and channels. The spatial location includes the number of pixel rows in the vertical direction and the number of pixel columns in the horizontal direction. The number of channels is the number of color or feature channels.
[0021] The orientation prediction subnetwork consists of multiple convolutional layers, normalization layers, and nonlinear activation functions stacked sequentially. It performs multi-scale convolutional feature extraction on the preprocessed frame image tensor, capturing local texture change features of the preprocessed frame image tensor at various spatial locations through convolutional kernels of different scales and receptive fields. The feature representation capability is enhanced by nonlinear activation functions, and orientation feature maps are output through a set of orientation feature channels. The orientation feature maps are consistent with the preprocessed frame image tensor in spatial dimension, and a preset number of orientation feature channels are output in the feature channel dimension to encode the local dominant texture orientation information of each spatial location of the preprocessed frame image tensor.
[0022] Based on the orientation feature map, a dominant gradient vector is constructed at each spatial location, and the dominant gradient vector is converted into the dominant gradient direction angle on the two-dimensional plane through the arctangent function. The dominant gradient vector is obtained by linearly mapping the directional feature map to the corresponding spatial location. The dominant gradient vector includes horizontal gradient components and vertical gradient components. The dominant gradient vector describes the two-dimensional directional information of the local texture change of the tensor of the preprocessed frame image in the live video stream scenario. The dominant gradient direction angle represents the dominant gradient direction of each pixel of the tensor of the preprocessed frame image.
[0023] Define a set of direction basis angles that correspond one-to-one with the set of difference basis functions; In Example 1, the number of differential basis functions is set, and uniform sampling is performed in the interval between 0 and 280 degrees or between 0 and 360 degrees according to the same equal-interval discretization method as the number of differential basis functions, to obtain the directional basis angle set. The value of each directional basis angle in the directional basis angle set is equal to the starting angle plus its index minus one multiplied by the angle interval. The angle interval is determined by the number of directional basis angles. The directional basis angle set is used to discretize the differential response in different directions in the live video stream scenario, and corresponds one-to-one with each differential basis function in the differential basis function set.
[0024] Based on the set of dominant gradient direction angles and direction basis angles, construct the manipulation coefficient matrix; In Example 1, the manipulation coefficient matrix is constructed as follows: For each spatial location of the preprocessed frame image tensor, the angle difference between the dominant gradient direction angle of the spatial location and each direction base angle in the set of direction base angles is calculated. This difference is used to measure the degree of directional deviation between the local texture direction of the spatial location and each direction base angle. Each angle difference is used as input and weighted by a uniform negative weight parameter to obtain the manipulation coefficient value corresponding to each direction base angle. This value serves as the manipulation coefficient of the spatial location at that direction base angle. The above method is used to calculate the manipulation coefficient matrix for each spatial location of the preprocessed frame image tensor, thus obtaining the manipulation coefficient matrix defined in the entire preprocessed frame image tensor space.
[0025] ; in, Indicates the pixel position in the preprocessed frame image. Place, No. Individual directions The corresponding manipulation coefficient, Indicates the pixel position of the preprocessed frame image. The dominant gradient direction angle at that location. Represents the first in a dynamically manipulated differential convolutional basis. A base angle in each direction, with a total of [number] set [number] angles. Individual directions, The values are evenly distributed in interval, This represents the direction sensitivity adjustment parameter, used to adjust the degree of influence of different dominant gradient directions and the difference between the direction basis angles on the manipulation coefficients. This represents the total number of basis directions in a dynamically manipulated differential convolution basis. , These are the index variables for the base directions.
[0026] The manipulation coefficient matrix is defined at each pixel position of the preprocessed frame image tensor and at each basis direction index in the dynamically manipulated differential convolution basis. It is used to encode the combined weights of different directional basis functions at the pixel level in the scenario of capturing instantaneous illegal frames in a live video stream.
[0027] The manipulation coefficient matrix and the preprocessed frame image are synchronously input into the dynamic manipulatory differential convolutional base module. The preprocessed frame image is convolved in parallel using the differential base function set and then linearly combined at the pixel level according to the manipulation coefficient matrix to generate an orientation adaptive differential feature map. In this embodiment, the manipulation coefficient matrix and the preprocessed frame image are synchronously input into the dynamically manipulated differential convolutional base module, including: The preprocessed frame image tensor and the manipulation coefficient matrix are synchronously input into the dynamically manipulated differential convolutional base module; In the dynamically manipulated differential convolutional basis module, local directional complexity weights based on the perturbation of the dominant gradient direction are constructed; In Example 1, for each pixel position of the preprocessed frame image tensor, based on the dominant gradient direction angle, the absolute values of the differences between the dominant gradient direction angles of all pixels in the local neighborhood pixel set centered on the pixel position and with a neighborhood radius of a positive integer are summed and averaged to obtain the local direction complexity weight of the pixel position. The local direction complexity weight is used to measure the perturbation intensity of the dominant gradient direction angle in the neighborhood of the pixel position in the instantaneous illegal frame capture scenario of the live video stream. The larger the local direction complexity weight value, the more drastic the change of the dominant gradient direction in the neighborhood of the corresponding pixel. The smaller the local direction complexity weight value, the more consistent the dominant gradient direction in the neighborhood.
[0028] In the dynamically manipulated differential convolution basis module, a set of differential basis functions is set, and multi-directional differential convolution operations are performed on the preprocessed frame image tensor to obtain the differential basis response tensor. The set of differential basis functions consists of two-dimensional discrete differential kernel weight matrices corresponding to M predefined basis direction indices. Each differential basis function targets a specific basis direction. Specifically, a first- or second-order discrete differential kernel is rotated in the spatial domain according to the angle corresponding to the basis direction index, forming differential basis functions in M directions. The sum of the weights of each kernel is normalized to zero, ensuring that the convolution output reflects local intensity changes rather than absolute intensity. The spatial size of each differential basis function matches the local neighborhood size of the tensor of the preprocessed frame image. M represents the total number of basis directions, corresponding one-to-one with the set of direction basis angles in the manipulation coefficient matrix.
[0029] The multi-directional differential convolution operation is as follows: for each pixel spatial position (x,y) and each base direction index m, the differential basis function and the preprocessed frame image tensor are subjected to two-dimensional discrete convolution in the neighborhood of (x,y) channel by channel. The convolution results of all channels are weighted and accumulated to obtain the differential basis response tensor. The differential basis response tensor is used to characterize the local high-frequency texture change of the preprocessed frame image tensor at pixel position (x,y) and the direction corresponding to the base direction index m in the instantaneous illegal frame capture scenario of the live video stream.
[0030] By combining local directional complexity weights, instantaneous violation sensitivity modulation is applied to the differential basis response tensor to generate a complexity-modulated differential basis response tensor; The instantaneous violation sensitivity modulation method involves multiplying the differential basis response tensor by a directional complexity adjustment parameter plus one for each pixel position and each base direction index, and then multiplying it by a local directional complexity weight. The directional complexity adjustment parameter is a non-negative real number used to adjust the amplification or suppression intensity of the local directional complexity weight on the differential basis response tensor, resulting in a complexity-modulated differential basis response tensor. This is used to enhance the response of regions with large perturbations in the local dominant gradient direction to the differential basis function in the instantaneous violation frame capture scenario of live video streams, highlighting the fine-grained violation texture structure that appears instantaneously.
[0031] Pixel-level linear combination of the complexity modulation differential basis response tensor is performed based on the manipulation coefficient matrix to generate an orientation-adaptive differential feature map.
[0032] The pixel-level linear combination involves multiplying the corresponding manipulation coefficient matrix with the complexity modulation differential basis response tensor one by one for all base direction indices at each pixel location and summing the results to obtain the orientation adaptive differential response scalar at the corresponding pixel location. The orientation adaptive differential response scalars at all pixel locations are spatially arranged to construct an orientation adaptive differential feature map. The orientation adaptive differential feature map provides the multi-scale fusion network with orientation adaptive differential features that combine dominant gradient direction information and local orientation complexity information in the scenario of capturing instantaneous illegal frames in live video streams.
[0033] Extract multi-scale fusion features from the directional adaptive difference feature map to generate a multi-scale fusion feature map; In this embodiment, the multi-scale fusion features of the orientation adaptive difference feature map are extracted, including: The orientation-adaptive differential feature map is input into a multi-scale fusion network to construct a feature fusion structure with different scale branches, forming a set of multi-scale differential feature sub-maps. Different scale branches receive multi-scale downsampled versions of the directional adaptive differential feature map, forming a set of multi-scale differential feature sub-maps. The set of multi-scale differential feature sub-maps represents the directional texture change information at different resolution levels in the instantaneous violation frame capture scene of the live video stream.
[0034] Independent convolutional encoding is performed on the multi-scale differential feature sub-map set to extract local spatial features within the scale, forming a set of coded feature maps within the scale. Each scale of convolutional coding consists of multiple convolutional layers, normalization layers, and nonlinear activation functions in sequence. The set of encoded feature maps within each scale retains the dominant response pattern of the directional adaptive differential features and enhances the stability of the local structure.
[0035] In a multi-scale fusion network, the molecular feature map of orientation difference with spatial resolution higher than the threshold is mapped to a low-scale level through downsampling operation, and the molecular feature map of orientation difference with spatial resolution lower than the threshold is mapped to a high-scale level through upsampling operation. Within each scale level, the original scale orientation difference molecular feature map is fused with all the mapped cross-scale features to obtain the inter-scale fusion feature map. The orientation adaptive difference feature maps are subjected to max pooling operations with different scale factors to obtain a set of orientation difference molecular feature maps at multiple scale levels. The total number of scales is set to S. The orientation difference molecular feature maps at different scale levels are obtained through different spatial downsampling factors, and each scale index corresponds to an orientation difference molecular feature map.
[0036] Based on the intra-scale encoded feature map set and the inter-scale fused feature map, a cross-scale residual connection mechanism is used for joint fusion, and a fused activation map is constructed by feature weighting superposition. In Example 1, all spatially aligned intra-scale encoded feature maps and inter-scale fused feature maps of the same spatial size are concatenated along the channel dimension to obtain a preliminary fused feature tensor. The preliminary fused feature tensor is then weighted and superimposed along the channel dimension, with each channel weighting coefficient being a learnable parameter. The weighted superimposed output is then directly added to the original intra-scale encoded feature map using a short-circuit connection to obtain a fused activation map. The fused activation map not only integrates spatial structural information within each scale and global feature semantics between scales, but also enhances the stability of feature representation and adaptability to instantaneous violations in live video streams through a residual connection mechanism. This ensures that the fused activation map provides rich and effective direction-sensitive features and differential texture features to support the extraction of multi-scale fused feature maps.
[0037] The fused activation map is input into the channel compression module, which compresses the number of feature channels along the channel dimension using a set of one-dimensional convolutional kernels with learnable parameters, and outputs a multi-scale fused feature map.
[0038] Multi-scale fusion feature map preservation and orientation-adaptive differential features Figure 1 The spatial resolution and number of channels are obtained by fusion activation map compression, preserving the most representative orientation-sensitive features, scale-invariant features, and differential texture features in the instantaneous violation frame capture task of live video stream.
[0039] The multi-scale fused feature map is projected onto the Riemannian manifold space through nonlinear mapping, and the Riemannian manifold mapping feature map is output. A background reference manifold is constructed based on the Riemannian manifold mapping feature map, and the geometric deviation between each local region and the background reference manifold is calculated to generate a geometric deviation mapping map. In this embodiment, a background reference manifold is constructed based on the Riemannian manifold mapping feature map, including: The multi-scale fused feature map is input into the nonlinear mapping module to construct a symmetric positive definite matrix feature tensor, which completes the mapping from Euclidean space to Riemannian manifold space and outputs the Riemannian manifold mapping feature map. In Example 1, the nonlinear mapping module is based on the construction method of symmetric positive definite matrix. It takes the local feature vector at each spatial position in the multi-scale fusion feature map of the instantaneous violation frame capture task of the live video stream as input. The local feature vector is a column vector composed of all feature channels of the multi-scale fusion feature map at a certain pixel position. For each pixel position, the column vector is multiplied by its own transpose to obtain the feature autocorrelation matrix of the pixel position. Then, the product of the regularization constant and the identity matrix is added to ensure that the feature autocorrelation matrix is a symmetric positive definite matrix, which is used to represent the local topological structure information of the pixel position in the instantaneous violation frame capture task of the live video stream. The symmetric positive definite matrix features generated by all pixel positions are combined according to spatial position to form the Riemannian manifold mapping feature map.
[0040] Based on the statistical characteristics of the spatial distribution in the feature map of the Riemannian manifold, a background reference manifold is constructed; In Example 1, the set of symmetric positive definite matrix features for all pixel locations in the Riemannian manifold map feature map is used as input. Each element in the set of symmetric positive definite matrix features represents a local manifold feature at a spatial location. Based on the criterion of minimizing the average distance of the symmetric positive definite matrix features at all spatial locations under the log-Euclidean metric, a global symmetric positive definite matrix is searched such that the sum of the squared distances between the global symmetric positive definite matrix and the symmetric positive definite matrix features at all spatial locations under the log-Euclidean metric is minimized. The optimal global symmetric positive definite matrix is used as the background reference manifold to represent the global manifold structure statistical features of the multi-scale fused feature map in the normal background region in the live video stream scene.
[0041] Based on the geodesic distance between the Riemannian manifold mapping feature map and the background reference manifold, the geometric deviation at the spatial location of each pixel is calculated, and a geometric deviation mapping map is constructed.
[0042] In Example 1, for each pixel location, the symmetric positive definite matrix feature of the corresponding pixel in the Riemann manifold mapping feature map is extracted, and the background reference manifold corresponding to the same image frame is extracted as the global reference matrix.
[0043] For the symmetric positive definite matrix features of the pixel and the background reference manifold, respectively, a matrix logarithmic mapping operation is performed to obtain their representations in the matrix logarithmic space. The representation of the symmetric positive definite matrix features of the pixel position in the matrix logarithmic space is subtracted from the representation of the background reference manifold in the same space to obtain a difference matrix. All elements of the difference matrix are squared, and then summed element by element. The square root of the sum is taken to obtain the geometric deviation value of the pixel position. The geometric deviation values calculated for all pixel positions are arranged according to their spatial positions in the original frame image to form a geometric deviation map. The value of each pixel in the geometric deviation map represents the geodesic distance between its local features and the global background structure in the Riemannian manifold logarithmic space.
[0044] A topological anomaly saliency map is generated based on the geometric deviation mapping map. A sparse mask generation module is then applied to the topological anomaly saliency map to extract the Top-K key regions and output a sparse mask. The sparse mask is then applied to a multi-scale fusion feature map to obtain a sparse candidate region feature map. In this embodiment, the output sparse mask is applied to the multi-scale fused feature map, including: The geometric deviation mapping is normalized to obtain a normalized geometric deviation mapping. Generate a topological anomaly saliency map based on the normalized geometric deviation map; The significance value of topological anomalies is a dimensionless real number between zero and one. The spatial distribution of the topological anomaly significance map is consistent with that of the normalized geometric deviation map. It is used to transform the geometric deviation intensity in the logarithmic space of the Riemannian manifold into a dimensionless significance response that can be thresholded.
[0045] ; in, The significance value of topological anomalies. For the Sigmoid function, Adjusting the saliency steepness, Adjusting the saliency center, This is a normalized geometric deviation mapping diagram.
[0046] Local consistency aggregation is performed on the topological anomaly saliency map to suppress isolated noise and generate candidate response maps; For each spatial location in the topological anomaly saliency map, all spatial locations within a square neighborhood with a positive integer radius are selected, centered on the current spatial location. The saliency values of all topological anomalies within the neighborhood are summed and divided by the number of pixels in the neighborhood to obtain the candidate response value for the current spatial location. The candidate response values of all spatial locations form a candidate response map. The candidate response map is used to retain stable topological anomaly responses within the neighborhood in the scenario of capturing instantaneous violation frames in a live video stream.
[0047] A sparse mask generation module is applied to the candidate response map to extract the Top-K key regions and output the sparse mask. In Example 1, the sparse mask generation module divides the candidate response map into a set of non-overlapping rule blocks according to space. The size of each rule block is a fixed number of pixels. The saliency score of each rule block is obtained by summing all candidate response values in each rule block and dividing by the number of pixels in the rule block. The saliency scores of all rule blocks are sorted by size, and the top-ranked rule blocks are taken as Top-K key regions. A binary sparse mask with the same resolution as the multi-scale fusion feature map is constructed. All spatial positions within the Top-K key regions are assigned a value of one in the sparse mask, and the rest are assigned a value of zero. The sparse mask is used to retain the key regions with the highest saliency in a block-level sparse form in the scenario of instantaneous violation frame capture in live video stream.
[0048] The sparse mask is applied to the multi-scale fused feature map to obtain the sparse candidate region feature map.
[0049] For each spatial location and each channel, the value of the sparse mask at the corresponding spatial location is multiplied element-wise with the feature values of the corresponding spatial location and channel of the multi-scale fusion feature map to obtain a sparse candidate region feature map. The sparse candidate region feature map and the multi-scale fusion feature map have the same spatial resolution and number of channels, retaining the feature information of the Top-K key regions, which are concentrated in the key regions in the instantaneous violation frame capture scenario of live video stream.
[0050] The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork. Local classification reasoning is performed on each frame, and violation labeling results are output. When the violation labeling results indicate the existence of a violation target, the instantaneous frame locking mechanism is triggered, the corresponding frame and its preceding and following context frames are recorded, violation frame capture results are generated, and the violation frame capture results are encapsulated and output as an evidence chain to complete the instantaneous violation frame capture of the live video stream.
[0051] In this embodiment, local classification reasoning is performed on each frame, including: The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork to extract frame-level semantic features and perform local classification inference, and output the violation labeling result. In Example 1, the sparse candidate region feature map is fed as input into the instantaneous frame-level violation classification subnetwork: The high-frequency directional response enhancement region detection module analyzes the directional response value of each spatial location within the sparse candidate region feature map, calculates the high-frequency directional response value of all pixels within the Top-K key regions, and marks pixels with directional response values greater than a preset response threshold as enhancement regions. It then calculates the area of all enhancement regions and compares it to the total area of the Top-K key regions. Simultaneously, it analyzes the dominant gradient direction deviation within the enhancement regions, calculates the magnitude of its angular deviation from the mean of the dominant directions in the neighborhood, and compares it to an angle threshold. If the area of the enhancement region accounts for more than a first percentage threshold of the total Top-K area, and the magnitude of the dominant gradient direction deviation exceeds the angle threshold, then the criterion is established.
[0052] The channel activation and spatial overlap determination module calculates the channel response values of all spatial locations in the sparse candidate region feature map of the current frame for each feature channel, extracts the maximum activation value of each channel, and compares it with the preset channel threshold. If the maximum activation value of at least one channel exceeds the channel threshold, the module further detects the spatial overlap between the spatial region corresponding to that channel and the suspected violation region in the historical frame. It uses spatial intersection-union ratio or other similarity measurement methods. If the overlap is greater than the second percentage threshold, the criterion is established.
[0053] The regional saliency score and historical comparison module extracts the saliency score of each key region in the current frame's Top-K key regions and calculates the average saliency score of the Top-K regions; it retrieves the average saliency scores of the Top-K key regions in historical frames, takes their mean as the historical reference value, and compares the current frame's Top-K average saliency score with the historical average; if the current frame's average saliency score is higher than a preset multiple of the historical average, the criterion is established.
[0054] The violation labeling output module integrates the following: if any criterion is met, the instantaneous frame-level violation classification subnetwork outputs the current frame as a violation frame; if none of the criteria are met, the current frame is output as a non-violation frame.
[0055] In this embodiment, the violation marking results include: A frame is marked as a violation if any of the following conditions are met: The sparse candidate region feature map contains a region with enhanced high-frequency directional response greater than the preset value, and the area of the region exceeds the first percentage threshold of the total area of the Top-K regions. At the same time, the magnitude of the deviation of the dominant gradient direction is greater than the angle threshold. The maximum activation value of at least one channel in the feature channel exceeds the set channel threshold, and the corresponding spatial region overlaps with the suspected violation region in the historical judgment with a degree greater than the second percentage threshold. The average saliency score of the Top-K region in the current frame is higher than the historical average by a preset multiple; All other cases are marked as non-violation frames.
[0056] When the violation marking result indicates that the current frame is a violation frame, the instantaneous frame locking mechanism is triggered, the current frame and several frames before and after it are recorded as context, the violation frame capture result is generated, and the evidence chain of the violation frame capture result is encapsulated and output.
[0057] In Example 1, the instantaneous frame locking mechanism is based on the current violation frame index. Backtracking frame index Expand the frame index backward ,in Extract the live video stream within the preset context frame span. to A frame-level capture window is constructed from a sequence of consecutive frames within the frame index range. All corresponding frame images and their violation area masks are recorded, and a set of violation frame capture results is generated.
[0058] Each frame image in the set of captured violation frames, its corresponding violation area mask, classification probability result, Top-K region coordinates and significance score are uniformly encapsulated to construct a frame-level violation evidence structure. All frame-level evidence structures are organized in chronological order to form a complete violation frame capture evidence chain. The evidence chain is output as a standardized structured data format and stored in the audit record system.
[0059] Example 2: During the deployment of the live streaming content security system, the operations and maintenance team conducted real-time monitoring of a certain high-definition video live stream. The video stream resolution was 1920×1080, with 30 frames per second. Each frame's raw data was normalized to generate a pre-processed frame image. One day, during the live stream, the system automatically detected a suspected abnormal signal in frame 83210 of the video stream numbered "Stream-07": The system first performs normalization and brightness normalization on frame 83210 to obtain preprocessed frame image A. This frame is then input into the orientation prediction subnetwork, where the system automatically calculates the dominant gradient direction for each pixel. (Pixels in the frame center region in Example 2...) The predicted dominant direction is 34.5°, with a mean dominant direction of 33.8° and a standard deviation of 2.4° within a 16×16 pixel neighborhood. The system outputs a manipulation coefficient matrix. This matrix is distributed for 24 directional base angles (every 15° from 0° to 345°). The maximum control factor at the point is 0.21, corresponding to a direction of 30°.
[0060] The system synchronously inputs the manipulation coefficient matrix and the preprocessed frame image A into the dynamically manipulated differential convolutional basis module. The differential basis response tensor in the central region is calculated in 24 directions. The convolutional response in the 30° direction is 1.26, in the 45° direction it is 0.41, and in other directions it is below 0.18. The local directional complexity weights in the neighborhood are statistically calculated to be 0.87. The directional complexity adjustment parameter is set to 0.6. After complexity modulation, the convolutional response increases to 1.82. Based on all directional and complexity modulation results, the system performs weighted summation on the pixels to generate an adaptive directional differential feature map B.
[0061] Feature map B is input into the multi-scale fusion network, automatically obtaining sub-feature maps at three scales with spatial resolutions of 1920×1080, 960×540, and 480×270. After max pooling of the directional difference sub-feature maps at each scale, the activation values in the central region are 1.58, 1.22, and 0.97, respectively. After cross-scale residual fusion, the fused activation map outputs a weighted activation value of 1.41 in the central region. After channel compression, the multi-scale fused feature map C is output, with a main channel activation value of 1.12 in the central region.
[0062] The multi-scale fusion feature map C is input to the nonlinear mapping module, and the central region is... The local feature channel vector is Construct an autocorrelation matrix with a length of 64. To ensure positive definiteness, a total of 2,073,600 symmetric positive definite matrices are generated across the entire frame. The mean of all matrices is calculated in log-Euclidean space, and the system obtains the background reference manifold. .
[0063] The system further calculates the geometric deviation at each pixel, in the central region. The background mean was 0.09. After normalization, the mean value was 0.89. The significance value of topological anomalies was then transformed using a Sigmoid function. .
[0064] In the candidate response map stage, the average response value of the 16×16 blocks in the central region is 0.81. After Top-K extraction, the entire region is included in the sparse mask with a mask value of 1, while the mask value of most other regions is 0. The system uses the mask to filter the multi-scale fused feature map and obtains the sparse candidate region feature map D. The maximum response of feature map D in the main channel of this region is 2.02.
[0065] The system inputs the sparse candidate region feature map D into the instantaneous frame-level violation classification subnetwork, automatically aggregates spatial features, and outputs the classification results.
[0066] The area of the high-frequency directional response enhancement region is 312 pixels, and the total area of the Top-K region is 420 pixels, accounting for 74.3% of the total area (threshold is 50%). The dominant gradient direction deviation is 20.1° (threshold is 12°), the maximum activation value of the main channel is 2.02, the spatial overlap of historical suspected violation regions is 79.6% (threshold is 70%), the average saliency score of the Top-K region in the current frame is 0.85, the historical average is 0.36, the preset multiplier is 2, and the score multiplier is 2.36.
[0067] After comprehensive evaluation, the system marks frame 83210 as a violation frame and automatically locks the three frames before and after it (a total of seven frames) as context to generate a violation frame capture result. The evidence chain encapsulates structured data such as the original image of each frame, the violation area mask, the classification output probability (0.98 for this frame), the Top-K region coordinates (center region range), and response statistics, and sends it to the security audit backend.
[0068] During this live stream detection, the system processed 900,000 frames of data and automatically detected a total of 46 instances, including tilted QR codes, flashing sensitive logos, and color-distorted patterns. The system missed 3 cases, all of which were edge frames with extremely low resolution and no dominant direction; it also made 2 false positives, both of which were strong bright areas in the bullet screen, but the significance score did not reach twice the threshold; the average processing time per frame was 12.4ms; the detection accuracy was 94.2%, the missed detection rate was 0.33%, and the false detection rate was 0.22%.
[0069] Compared to traditional methods (convolutional object detection + channel attention) under the same conditions: 28 cases were hit, 18 cases were missed, mainly due to high-angle tilt and bright noise obscuring the target; 7 cases were false positives, mainly due to barrage interference and bright light areas; the average processing time was 27.9ms; the detection accuracy rate was 60.9%, the missed detection rate was 2.0%, and the false positive rate was 0.78%.
[0070] Regarding the training samples, the method of this invention uses 50,000 live broadcast violation materials covering multiple angles, deformations, and complex backgrounds (including 11,000 rotating QR code samples, 7,000 instantaneous flashing targets, and 20,000 LED background / bullet screen interference samples). Traditional methods have the same sample size, but the proportion of rotating and instantaneous flashing targets is only 18% of the total.
[0071] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for monitoring sensitive information in live internet video broadcasts based on artificial intelligence, characterized in that, include: Acquire a continuous frame sequence of images from a live video stream and perform preprocessing to generate preprocessed frame images; The dominant gradient direction of each pixel in the preprocessed frame image is predicted using the orientation prediction subnetwork, and the manipulation coefficient matrix is output. The manipulation coefficient matrix and the preprocessed frame image are synchronously input into the dynamic manipulatory differential convolutional base module. The preprocessed frame image is convolved in parallel using the differential base function set and then linearly combined at the pixel level according to the manipulation coefficient matrix to generate an orientation adaptive differential feature map. Extract multi-scale fusion features from the directional adaptive difference feature map to generate a multi-scale fusion feature map; The multi-scale fused feature map is projected onto the Riemannian manifold space through nonlinear mapping, and the Riemannian manifold mapping feature map is output. A background reference manifold is constructed based on the Riemannian manifold mapping feature map, and the geometric deviation between each local region and the background reference manifold is calculated to generate a geometric deviation mapping map. A topological anomaly saliency map is generated based on the geometric deviation mapping map. A sparse mask generation module is then applied to the topological anomaly saliency map to extract the Top-K key regions. The sparse mask is then output and applied to the multi-scale fusion feature map to obtain a sparse candidate region feature map. The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork. Local classification reasoning is performed on each frame, and violation labeling results are output. When the violation labeling results indicate the existence of a violation target, violation frame capture results are generated, and the violation frame capture results are encapsulated into an evidence chain and output.
2. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The preprocessing includes performing normalization, brightness normalization, and region mask initialization on the continuous frame sequence of the live video stream to generate preprocessed frame images.
3. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The method of predicting the dominant gradient direction of each pixel in the preprocessed frame image using a direction prediction sub-network includes: The preprocessed frame image tensor is input into the orientation prediction subnetwork, and the orientation feature map is output. Based on the orientation feature map, a dominant gradient vector is constructed at each spatial location, and the dominant gradient vector is converted into the dominant gradient direction angle on the two-dimensional plane through the arctangent function. Define a set of direction basis angles that correspond one-to-one with the set of difference basis functions; Based on the set of dominant gradient direction angles and direction basis angles, a manipulation coefficient matrix is constructed.
4. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The module that synchronously inputs the manipulation coefficient matrix and the preprocessed frame image into the dynamically manipulated differential convolution base module includes: The preprocessed frame image tensor and the manipulation coefficient matrix are synchronously input into the dynamically manipulated differential convolutional base module; In the dynamically manipulated differential convolutional basis module, local directional complexity weights based on the perturbation of the dominant gradient direction are constructed; In the dynamically manipulated differential convolution basis module, a set of differential basis functions is set, and multi-directional differential convolution operations are performed on the preprocessed frame image tensor to obtain the differential basis response tensor. By combining local directional complexity weights, instantaneous violation sensitivity modulation is applied to the differential basis response tensor to generate a complexity-modulated differential basis response tensor; Pixel-level linear combination of the complexity modulation differential basis response tensor is performed based on the manipulation coefficient matrix to generate an orientation-adaptive differential feature map.
5. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The multi-scale fusion features of the extracted directional adaptive difference feature map include: The orientation-adaptive differential feature map is input into a multi-scale fusion network to construct a feature fusion structure with different scale branches, forming a set of multi-scale differential feature sub-maps. Independent convolutional encoding is performed on the multi-scale differential feature sub-map set to extract local spatial features within the scale, forming a set of coded feature maps within the scale. In a multi-scale fusion network, the molecular feature map of orientation difference with spatial resolution higher than the threshold is mapped to a low-scale level through downsampling operation, and the molecular feature map of orientation difference with spatial resolution lower than the threshold is mapped to a high-scale level through upsampling operation. Within each scale level, the original scale orientation difference molecular feature map is fused with all the mapped cross-scale features to obtain the inter-scale fusion feature map. Based on the intra-scale encoded feature map set and the inter-scale fused feature map, a cross-scale residual connection mechanism is used for joint fusion, and a fused activation map is constructed by feature weighting superposition. The fused activation map is input into the channel compression module, which compresses the number of feature channels along the channel dimension using a set of one-dimensional convolutional kernels with learnable parameters, and outputs a multi-scale fused feature map.
6. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The construction of the background reference manifold based on the Riemannian manifold mapping feature map includes: The multi-scale fused feature map is input into the nonlinear mapping module to construct a symmetric positive definite matrix feature tensor, which completes the mapping from Euclidean space to Riemannian manifold space and outputs the Riemannian manifold mapping feature map. Based on the statistical characteristics of the spatial distribution in the feature map of the Riemannian manifold, a background reference manifold is constructed; Based on the geodesic distance between the Riemannian manifold mapping feature map and the background reference manifold, the geometric deviation at the spatial location of each pixel is calculated, and a geometric deviation mapping map is constructed.
7. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The output sparse mask is applied to the multi-scale fused feature map, including: The geometric deviation mapping is normalized to obtain a normalized geometric deviation mapping. Generate a topological anomaly saliency map based on the normalized geometric deviation map; Local consistency aggregation is performed on the topological anomaly saliency map to suppress isolated noise and generate candidate response maps; A sparse mask generation module is applied to the candidate response map to extract the Top-K key regions and output the sparse mask. The sparse mask is applied to the multi-scale fused feature map to obtain the sparse candidate region feature map.
8. The method for monitoring sensitive information in live internet video streaming based on artificial intelligence according to claim 1, characterized in that, The local classification reasoning for each frame includes: The sparse candidate region feature map is input into the instantaneous frame-level violation classification subnetwork to extract frame-level semantic features and perform local classification inference, and output the violation labeling result. When the violation marking result indicates that the current frame is a violation frame, the instantaneous frame locking mechanism is triggered, the current frame and several frames before and after it are recorded as context, the violation frame capture result is generated, and the evidence chain of the violation frame capture result is encapsulated and output.
9. A method for monitoring sensitive information in live internet video streaming based on artificial intelligence, as described in claim 8, characterized in that, The violation marking results include: A frame is marked as a violation if any of the following conditions are met: The sparse candidate region feature map contains a region with enhanced high-frequency directional response greater than the preset value, and the area of the region exceeds the first percentage threshold of the total area of the Top-K regions. At the same time, the magnitude of the deviation of the dominant gradient direction is greater than the angle threshold. The maximum activation value of at least one channel in the feature channel exceeds the set channel threshold, and the corresponding spatial region overlaps with the suspected violation region in the historical judgment with a degree greater than the second percentage threshold. The average saliency score of the Top-K region in the current frame is higher than the historical average by a preset multiple; All other cases are marked as non-violation frames.