AI-based real-time audit and automatic replacement method for illegal content in broadcast programs
By using multimodal deep learning and generative artificial intelligence technologies, we can identify and replace illegal elements in broadcast television programs, generating visually consistent and semantically coherent replacement content. This solves the problems of artistic integrity and technical indicators in program review in existing technologies, and achieves high-quality real-time synthesis effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN CABLE RADIO & TELEVISION NETWORK CO LTD
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-03
AI Technical Summary
Existing broadcast television program review technologies are insufficient to maintain the artistic integrity and viewing continuity of programs while ensuring security. Furthermore, existing solutions are highly destructive and cannot guarantee the artistic quality and technical specifications of the output video. They also struggle to achieve precise pixel-level positioning, context-aware content generation, and real-time synthesis of cinematic visual effects.
A pre-trained multimodal deep learning model is used to identify and deeply analyze illegal elements. Generative artificial intelligence models are used to generate visually consistent and semantically coherent replacement content. Combined with multi-objective optimization and advanced synthesis techniques, the replacement content is deeply integrated with the original scene.
It achieves the goal of eradicating illegal information while preserving the original visual artistic effect and narrative fluency of the program to the greatest extent, and outputting video streams that meet broadcast-grade technical quality and security standards.
Smart Images

Figure CN122120492B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of broadcast television program production and broadcast control technology, specifically involving an AI-based method for real-time review and automatic replacement of illegal content in broadcast programs. Background Technology
[0002] Content review of radio and television programs is a core element in ensuring the safety of cultural dissemination and compliance with industry regulatory requirements. With the massive growth of media content and the real-time and streaming development of broadcast formats, the task of reviewing inappropriate content presents stringent requirements of high concurrency, high real-time performance, and high accuracy. Existing review and handling technologies mainly rely on two modes: manual review and rule-driven automatic blocking. The specific process can be summarized as: manually or using simple models to identify inappropriate areas—applying destructive visual coverage or directly cutting off the signal. This makes it difficult to maintain the artistic integrity and viewing continuity of the program while ensuring safety.
[0003] Existing technologies can meet basic needs when dealing with the demand for high-quality and real-time broadcasting programs, but there are still some shortcomings: (1) Existing solutions focus on discovering violations, but the means of handling violations are extremely limited and highly destructive, lacking the constructive repair capabilities to start from a deep understanding of the scene, intelligently generate compliant content and seamlessly integrate it.
[0004] (2) The existing solution does not take into account the three-dimensional geometric environment, dynamic lighting conditions, camera movement and narrative semantic context of the violating element, resulting in the disposal result being disconnected from the original scene at both the physical and logical levels.
[0005] (3) The existing solution is a pixel-destruction-based processing method that cannot guarantee the artistic quality and technical specifications of the output video, and may introduce new visual flaws that do not meet professional broadcasting standards.
[0006] (4) Under broadcast-level latency constraints, existing technologies cannot simultaneously achieve accurate pixel-level positioning, context-aware content generation, and real-time synthesis of cinematic visual effects. Summary of the Invention
[0007] In view of this, in order to solve the problems mentioned in the background technology above, we now propose an AI-based method for real-time review and automatic replacement of illegal content in broadcast programs.
[0008] The objective of this invention can be achieved through the following technical solution: This invention provides an AI-based method for real-time review and automatic replacement of illegal content in broadcast programs, including: S1, identifying illegal elements in the video stream of broadcast programs based on a pre-trained multimodal deep learning model and determining their pixel-level spatial position in the video frame; and simultaneously performing deep analysis on the current video scene containing the illegal elements to obtain scene deep analysis results.
[0009] S2. Input the scene deep analysis results as constraints into the generative artificial intelligence model to generate at least one candidate replacement content for replacing the illegal element. The generation process must meet the constraints, including visual consistency constraints and semantic coherence constraints.
[0010] S3. Perform a comprehensive evaluation of the generated candidate replacement content for multi-objective optimization, select the optimal replacement scheme, and perform multi-scale real-time synthesis with the compliant background area of the original video stream to generate and output the compliant program video stream after replacement.
[0011] Compared with existing technologies, the beneficial effects of the present invention are as follows: 1. The present invention integrates multimodal recognition, scene deep analysis and generative artificial intelligence to generate visually consistent and semantically coherent replacement content in real time according to the specific context of the violation element, and achieves pixel-level accurate synthesis, thereby eliminating the violation information while preserving the original visual artistic effect and narrative fluency of the program to the greatest extent.
[0012] 2. This invention uses the multi-dimensional analysis results of the scene's geometric structure, lighting parameters, emotional state, and dialogue semantics as strong constraints for the generative AI model to guide the generation of replacement content. This ensures that the generated content not only matches the background in terms of texture and color, but also deeply integrates with the original scene in terms of three-dimensional spatial relationships, physical lighting interaction, and plot logic.
[0013] 3. This invention proposes a multi-index evaluation model covering visual fidelity, artistic style matching, coherence, and thoroughness of risk elimination. It automatically selects the optimal replacement scheme through weighted optimization and combines advanced compositing techniques such as Poisson editing, high-precision image matting, and motion blur compensation to ensure that the final output video stream meets broadcast-grade technical quality and safety standards, thus solving the problems of low quality and uncontrollable results in existing technologies. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating the implementation steps of the method of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Please see Figure 1 As shown, the present invention provides an AI-based method for real-time review and automatic replacement of illegal content in broadcast programs. The specific steps are as follows: S1, Identify illegal elements in the video stream of broadcast programs based on a pre-trained multimodal deep learning model and determine their pixel-level spatial positions in the video frame; at the same time, perform deep analysis on the current video scene containing the illegal elements to obtain the scene deep analysis results.
[0018] It should be noted that the multimodal deep learning model refers to a fusion encoder based on the Transformer architecture. Its visual branch uses Vision Transformer (ViT) and its audio branch uses Wav2Vec 2.0. It performs feature alignment and fusion through a cross-modal attention mechanism and is trained end-to-end on a large-scale labeled dataset of illegal content in broadcasting programs.
[0019] In one specific example, the training dataset contains 500,000 1080p / 25fps broadcast program clips, covering 5 major categories and 200 subcategories, including restricted signs and prohibited items. Each frame is labeled with a pixel-level mask and violation category label. The total loss function is a weighted sum of FocalLoss and Dice Loss. The initial learning rate is 1e-4, Cosine annealing is used, the batch size is 32, and the training is conducted for 100 rounds on an 8×A100 GPU.
[0020] Furthermore, to achieve real-time review processing, this application employs a GPU cluster-based stream processing architecture. Multiple stages, including video decoding, feature extraction, violation detection, content generation, and synthesis, constitute an asynchronous pipeline. The generative artificial intelligence model utilizes a lightweight latent diffusion model variant and leverages TensorRT for inference optimization. Experiments have verified that, on a hardware environment configured with NVIDIA A100 GPUs, the average processing latency for 1080p video streams can be controlled to within 400 milliseconds.
[0021] As an exemplary embodiment of the present invention, the specific implementation of the pixel-level spatial position of the violation element includes: decomposing the input broadcast program video stream into a continuous video frame sequence and simultaneously extracting the audio track signal.
[0022] Specifically, the extraction of audio track signals aims to decompose the input original video stream of broadcast programs into independent video frames and audio signals that can be analyzed in real time. The specific implementation process includes: using a high-performance video decoder based on hardware acceleration to decode the received RTMP or HLS protocol video stream into an uncompressed continuous video frame sequence, with the frame rate set at 25 to 30 fps to match the broadcast standard, and simultaneously extracting the original audio track signal in PCM format.
[0023] The key parameter in the decoding process is the decoding buffer size, which is set within a low range to ensure that the end-to-end processing latency is less than 500 milliseconds, meeting real-time requirements. This process is triggered by the detection of a valid video stream input and continues until the video stream is interrupted.
[0024] The spatial feature vectors of video frames are extracted using a visual Transformer, and the motion feature vectors between video frames are extracted using a Long Short-Term Memory network.
[0025] Specifically, the extraction of spatial and motion features from video frames aims to extract multi-dimensional feature vectors that can characterize static content and dynamic changes from a video frame sequence. The specific implementation process includes: employing a dual-path parallel feature extraction architecture, utilizing the Visual Transformer (ViT) model to process each video frame. ViT captures local texture and global spatial layout information by segmenting the image into 16x16 pixel patches and calculating the self-attention relationships between them, thereby generating spatial feature vectors. The spatial feature vector sequence corresponding to 16 consecutive frames is then fed into a Long Short-Term Memory (LSTM) network containing 512 hidden units. The LSTM, through its internal gating mechanism, learns and encodes motion patterns and temporal dependencies in the video segments, outputting motion feature vectors that integrate spatiotemporal information.
[0026] The violation content classifier identifies whether there are restricted symbols, prohibited items, inappropriate captions, controversial figures, and violations in video frames, and collectively refers to them as violation targets. It also outputs the initial bounding box coordinates of the detected violation targets.
[0027] Specifically, a well-trained classifier for illegal content is integrated at the end of the object detection framework. When the feature vector output from the previous step is input, the classifier will perform probability calculations for predefined illegal categories such as restricted signs, prohibited items, inappropriate captions, controversial figures, and illegal behaviors. If the confidence score of any category is higher than the preset activation threshold of 0.85, it is determined that there is a suspected violation.
[0028] For detected violations, an instance segmentation algorithm is used to generate a pixel-level semantic mask to define the geometric boundaries of the violation element in each frame.
[0029] Specifically, defining the geometric boundaries of the violation element in each frame aims to upgrade the localization of the violation target from a coarse bounding box to a pixel-level contour. The specific implementation process includes: after receiving the initial bounding box coordinates of the violation target, using these as guidance, initiating an instance segmentation algorithm such as Mask R-CNN or DeepLabV3+ within the defined region. This algorithm can not only distinguish different individuals within the same category, but also generate a pixel-level binary semantic mask for each detected violation element instance.
[0030] The semantic mask is a two-dimensional matrix with the same size as the original video frame, where pixels with a value of 1 precisely correspond to the area occupied by the illegal element, and pixels with a value of 0 represent the background, thereby achieving accurate definition of the geometric boundaries of the illegal element.
[0031] Lock the motion path of the violating element on the timeline of the video stream and generate a spatiotemporal bounding box sequence of its complete lifecycle.
[0032] Specifically, the spatiotemporal bounding box sequence aims to lock and track the motion trajectory of the same violation element in consecutive video frames, forming a complete spatiotemporal existence record. Its implementation includes: employing a multi-target tracking algorithm to assign a unique tracking ID to the mask of each violation element segmented for the first time; and using a Kalman filter model to predict the possible position and size of the violation element in the next frame. This prediction process can be described by a state prediction formula, which is as follows: ,in In the current frame Predicted values for the status of violating elements; In the previous frame The optimal estimate of the state of the element is calculated by the filter update step; This is the state transition matrix, which describes the evolution of the state over time based on a physical motion model such as uniform motion. This matrix is set during the initialization phase according to the expected target motion characteristics; the state vector... It typically includes the centroid coordinates, dimensions, and first-order time derivatives of the target mask.
[0033] Through continuous iterative prediction and updating, a spatiotemporal bounding box sequence is generated. This sequence records in detail the precise mask information of each violation element in every frame from its appearance to its disappearance, providing a stable and continuous spatiotemporal localization basis for the subsequent replacement module.
[0034] As an exemplary embodiment of the present invention, the scene depth analysis result includes the obtained scene geometric structure and camera motion parameters, light and shadow and color gamut parameters, character emotional state feature parameters, and dialogue context semantic logic feature parameters.
[0035] Optical flow estimation and 3D reconstruction algorithms are used to extract the geometric structure and camera motion parameters of the scene.
[0036] Specifically, the process of extracting the 3D geometric structure and camera motion parameters of the scene includes: employing a deep learning-based optical flow estimation algorithm to densely predict pixel-level displacements between consecutive video frames, generating an optical flow field. This optical flow field is then fed into a monocular vision simultaneous localization and mapping (SLAM) system. This system triangulates and continuously adjusts the feature points in the continuous optical flow field, calculating the camera's rotation and translation matrices in each frame in real time, thereby obtaining camera motion parameters and simultaneously constructing a sparse 3D point cloud of static objects in the scene, which constitutes the scene's geometric structure.
[0037] The lighting parameters and color gamut distribution of a scene are obtained using a colorimeter model and a global illumination analysis algorithm.
[0038] Specifically, the process of obtaining the scene's lighting and color gamut parameters includes: processing the current video frame using a global illumination analysis algorithm. This algorithm estimates the direction vector and intensity of the main light source by analyzing the shadow areas and highlight distribution in the image. A colorimeter model is used to analyze the overall color tendency of the image, quantifying the scene's color temperature by calculating the average chromaticity value in the CIELAB color space, typically ranging from 2800K to 7000K, and extracting a color histogram to characterize the color gamut distribution.
[0039] We use an affective computing model to extract feature vectors of a person's emotional state.
[0040] Specifically, the extraction process of the emotional state feature vector includes: locating the face of a person in the image using a face detection algorithm, and inputting the cropped facial image into a pre-trained emotion computing model. This model is a deep convolutional neural network, and its output is a two-dimensional emotional state feature vector. This vector maps the person's emotions to a valence-arousal (VA) emotional space, where the valence dimension represents the positive or negative nature of the emotion, and the arousal dimension represents the intensity of the emotion. In this way, the dynamic emotional changes of a person are quantified as a continuous trajectory in the emotional space.
[0041] We use natural language processing techniques to analyze the semantic and logical context of the dialogue in the current audio track.
[0042] Specifically, the analysis process of the semantic logical context of the dialogue includes: applying Automatic Speech Recognition (ASR) technology to the synchronously extracted audio track signal to convert it into a text subtitle stream. Then, using a large language model in Natural Language Processing (NLP), the text stream is analyzed in real time, performing named entity recognition and relation extraction tasks to identify key characters, events, and objects in the dialogue and construct the logical relationships between them. The output of these operations is a structured semantic logical context of the dialogue, which provides a semantic basis for understanding the plot development and character interactions.
[0043] This invention integrates multimodal recognition, deep scene analysis, and generative artificial intelligence to generate visually consistent and semantically coherent replacement content in real time based on the specific context of the violating element, and achieves pixel-level precise synthesis. This eliminates the violating information while preserving the original visual artistic effect and narrative fluency of the program to the greatest extent.
[0044] S2. Input the scene deep analysis results as constraints into the generative artificial intelligence model to generate at least one candidate replacement content for replacing the illegal element. The generation process must meet the constraints, including visual consistency constraints and semantic coherence constraints.
[0045] It should be noted that the generative artificial intelligence model specifically adopts the latent diffusion model, and its conditional input channel is used to receive the conditional vector encoded by the scene deep analysis result.
[0046] As an exemplary embodiment of the present invention, the specific method for generating the candidate replacement content includes: retrieving relevant materials from a preset compliant material library based on the attribute tags of the violating element, and calculating the matching degree score between the material and the violating element in terms of shape, texture, and size; if the matching degree score is higher than a first threshold, the material is used first; if the matching degree score is lower than the first threshold but higher than a second threshold, the aforementioned material is used as a base to start a real-time generation process for adaptation adjustment; if the matching degree score is lower than the second threshold, the real-time generation process is activated to synthesize a new image that conforms to the current scene description.
[0047] In one specific example, the compliance material library contains tens of thousands of authorized 2D / 3D compliance elements, classified according to the ISO21000 standard, with attribute tags using predefined 100-dimensional one-hot vectors such as beverage bottles, no text, and cylinders.
[0048] The formula for calculating the matching score is: ,in The intersection-union ratio of the mask for the violation area and the projected outline of the material. Calculations in CIELAB space This represents the relative difference in area. This represents the maximum area difference.
[0049] The selected material from the compliant material library or the new image is mapped into a high-dimensional latent space, and the latent variables are adjusted to make them distributed in the neighborhood of the original scene feature vector.
[0050] Specifically, selected materials from the compliant material library or the new image are input into the encoder and mapped to a high-dimensional latent space to obtain initial latent variables. Simultaneously, the extracted spatial feature vector representing the visual style of the original scene background is also mapped into this latent space as the target anchor point. Further, an iterative optimization algorithm fine-tunes the initial latent variables, aiming to minimize the Euclidean distance between the latent variables and the target anchor point in the latent space. The number of iterations in this optimization process is limited to between 50 and 100 to ensure that style alignment is achieved without sacrificing too much real-time performance. Ultimately, the image generated by the decoder based on the adjusted latent variables is highly consistent in texture and color with the neighborhood represented by the original scene feature vector.
[0051] An optical flow guidance mechanism is introduced to constrain the consistency of replacement content generated in consecutive frames, thereby eliminating flickering and artifacts commonly found in generated videos.
[0052] Specifically, to ensure a smooth transition between replacement content generated in consecutive frames and eliminate flickering and jitter artifacts common in video, an optical flow guidance mechanism is introduced. This mechanism utilizes the calculated scene optical flow field. In generating the first... When replacing the content of a frame, first replace the first frame. The replacement content area of the frame has been generated and composited, according to Frame to The optical flow vector of the frame undergoes a pixel-level forward warping transform to obtain a predicted value. Frame replacement content. This prediction result will be used as the conditional diffusion model in the generation process. The strong initial noise map of the frame content constrains the generation result, ensuring that it maintains a high degree of continuity with the previous frame in terms of motion and shape.
[0053] Calculate the depth occlusion relationship between other objects in the scene and the illegal area, and generate a depth map of the replacement area through a depth estimation model.
[0054] Specifically, a monocular depth estimation model is invoked to generate a pixel-by-pixel depth map for the current video frame containing the violation element. This depth map is a single-channel image, and the pixel value is inversely proportional to the distance of that point from the camera in the real world. During final compositing, the depth value of each pixel within the violation area is compared with the specified depth value of the replacement content. If other objects exist in the scene, such as the arm of a foreground person, and their depth value indicates that they are closer to the camera than the violation area, the original pixels of the arm will be preserved in the final image. This achieves correct occlusion of the replacement content, ensuring that the replacement is naturally embedded into the correct position in the 3D scene.
[0055] As an exemplary embodiment of the present invention, the visual consistency constraint is specifically defined as the generated replacement content maintaining continuous consistency with the original video background in terms of texture, lighting, shadow, color temperature, and motion trajectory.
[0056] The specific implementation of the visual consistency constraint includes: by extracting the style features and frequency distribution of the original background, adjusting the statistical moments of the generated replacement content so that the global histogram and local contrast of the replacement area completely match the background.
[0057] Specifically, several image patches are randomly sampled from the original background adjacent to the violation area mask, and these patches are input into a pre-trained convolutional neural network such as VGG-19 to extract activation response maps at different depths. These activation response maps collectively constitute the style features of the background. Further, statistical moment alignment is performed between the activation response maps of the generated replacement content and the background style features, specifically by applying adaptive instance normalization. The calculation of this process can be expressed by the following formula: ,in This represents the feature map of the generated replacement content at a certain layer of the network. and These are feature maps The channel-wise mean and standard deviation are obtained by calculating the spatial dimension of the feature map; These are the mean and standard deviation of the corresponding channels extracted from the background style features, respectively. It is a new feature map after style transfer.
[0058] By applying this operation to multiple layers of the network and ultimately decoding it back to the image space, the global color histogram and local texture contrast of the replaced content can be highly matched with the background area.
[0059] By using ghosting detection and edge smoothing algorithms to process the boundary between the replacement content and the background, a transition buffer is established to achieve a natural transition at the pixel level.
[0060] Specifically, a transition buffer is established by extending 5 to 15 pixels inwards and outwards from the boundary of the semantic mask of the generated violation element. Within this buffer, pixels are not directly copied. Instead, a Laplace equation is solved to construct an optimization objective: minimizing the difference between the gradient field within the synthetic region and the gradient field of the original replacement content, while forcing the boundary pixel values to be equal to the original background pixel values. By iteratively solving this partial differential equation, the optimal intensity value for all pixels within the buffer is calculated, allowing the edge gradient of the replacement content to decay smoothly and seamlessly connect with the background gradient.
[0061] Based on the calculated physical lighting parameters, virtual shadows and environmental occlusions are projected in real time onto the surface of the replacement content, and sensor noise and lens blur effects are simulated to match the imaging characteristics of the original camera.
[0062] Specifically, by utilizing the scene's geometry and the estimated direction of the main light source, virtual shadows and ambient occlusion (AO) are projected onto the surface of the replacement content. This process is achieved through screen-space ambient occlusion technology, simulating the soft shadow effect produced by light being blocked by nearby objects with low computational cost. Furthermore, to simulate the imaging characteristics of the original camera, the intensity and distribution parameters of sensor noise are analyzed and estimated from flat areas of the original video frames, and a layer of intensity-matched procedural Gaussian noise is superimposed on the replacement content. Based on the generated depth map, a lens blur effect (Bokeh effect) consistent with the original lens depth of field is applied to the replacement content to ensure its sharpness on the visual focal plane is consistent with other parts of the scene, thereby eliminating visual incongruity caused by excessive sharpness.
[0063] As an exemplary embodiment of the present invention, the semantic coherence constraint is specifically defined as the generated replacement content maintaining a strong correlation with the original video context in terms of plot logic, character interaction, and emotional expression.
[0064] The specific implementation methods of the semantic coherence constraint include: constructing a domain knowledge graph and extracting key entities and logical relationships from the program script or dialogue context.
[0065] Specifically, the construction of the domain knowledge graph aims to ensure that the replacement content conforms to the inherent logic of the plot. The domain knowledge graph is a graph-structured data structure where nodes represent key entities in the script or dialogue, such as characters, items, and locations, while edges represent relationships between these entities, such as ownership, use, or conversation. When generating candidate replacement content, a constraint is enforced: the semantic tags of any generated virtual objects or patched areas must maintain logical consistency with the context entities within the domain knowledge graph.
[0066] For example, if the domain knowledge graph indicates that a character is drinking water, then the candidate content to replace the non-compliant brand bottle in their hand must be an object with the semantic tag of container or beverage.
[0067] The extracted emotional state feature parameters of the characters are used as emotional benchmarks and mapped to a three-dimensional emotional space such as VA space.
[0068] When generating candidate replacement content, the semantic labels of the generated virtual objects or patched areas are logically consistent with the context entities in the knowledge graph, and the Euclidean distance change of the sentiment vector is within the preset smoothing threshold range.
[0069] Specifically, to maintain the continuity and authenticity of characters' emotional expressions, the extracted emotional state feature parameters are used as the emotional baseline. When generating replacement content, the potential impact of the replacement behavior on the characters' emotional states is assessed. This assessment is achieved through the following constraints: ,in The Euclidean distance between the replaced sentiment vector and the original sentiment vector in the VA space. The specific calculation method is as follows: ,in , These represent the valence and arousal of the predicted emotion evoked by the replaced content, respectively, which are predicted from the replaced image using a pre-trained emotion regression model. , These are the baseline valence and arousal extracted from the original image, respectively. It is a preset smoothing threshold, which is set in a small range, such as 0.1, to ensure that the changes in the emotion vector remain within a small and smooth range, avoiding abrupt emotional jumps.
[0070] This invention uses the multi-dimensional analysis results of the scene's geometric structure, lighting parameters, emotional state, and dialogue semantics as strong constraints for the generative AI model to guide the generation of replacement content. This ensures that the generated content not only matches the background in terms of texture and color, but also deeply integrates with the original scene in terms of three-dimensional spatial relationships, physical lighting interaction, and plot logic.
[0071] S3. Perform a comprehensive evaluation of the generated candidate replacement content for multi-objective optimization, select the optimal replacement scheme, and perform multi-scale real-time synthesis with the compliant background area of the original video stream to generate and output the compliant program video stream after replacement.
[0072] As an exemplary embodiment of the present invention, the multi-objective optimization includes a coherence index, an art style matching index, a visual fidelity index, and a risk elimination thoroughness index.
[0073] As an exemplary embodiment of the present invention, the specific calculation method of the coherence index, art style matching index, visual fidelity index and risk elimination thoroughness index includes: calculating the structural similarity and peak signal-to-noise ratio between each candidate replacement content and the edge of the original scene, normalizing them and then obtaining the visual fidelity index by weighted averaging.
[0074] Specifically, structural similarity (SSIM) is calculated at the boundary between the replacement region and the original background to measure the degree of matching in terms of brightness, contrast, and structure. Peak signal-to-noise ratio (PSNR) of the replacement region is also calculated; the reference image here is the ideal background obtained after repairing the violation area using the original background texture.
[0075] The Euclidean distance between the replacement region and the neighboring background feature maps on multiple convolutional layers is calculated, and the difference value obtained by averaging the distances is mapped to a score interval of 0 to 1 as an art style matching index.
[0076] Specifically, the image patch of the replacement region and the image patch of the neighboring background are each input into a pre-trained deep convolutional network, and feature maps of both are extracted across multiple convolutional layers. The style difference is quantified by calculating the Euclidean distance between these feature maps, and the average of the multiple Euclidean distances is used to obtain the difference value, which is mapped to a score range of 0 to 1. The smaller the difference, the higher the artistic style matching index.
[0077] A deep discriminant network is used to detect forgery traces in the replaced image, and a coherence index is obtained based on this.
[0078] Specifically, to detect whether the replacement operation introduces perceptible synthetic traces, a deep discriminative network is used for evaluation. This network is specifically trained to identify forgery traces in images. Its input is a complete video frame containing candidate replacement content, and its output is a probability value between 0 and 1, representing the likelihood that the frame is judged as forged. Subtracting this probability value from 1 yields a coherence index. The closer the coherence index is to 1, the more natural and seamless the fusion effect.
[0079] Establish a compliance secondary verification mechanism and calculate the risk elimination thoroughness indicator.
[0080] Specifically, the synthesized replacement area will be sent back to the violation content classifier. The violation content classifier will scan all violation categories and output the highest confidence score. The risk elimination thoroughness index is 1 minus this highest confidence score. When the risk elimination thoroughness index is close to 1, it indicates that the risk has been completely eliminated, thus ensuring that the replacement content itself does not contain any hidden violation elements.
[0081] As an exemplary embodiment of the present invention, the specific selection method of the optimal replacement scheme includes: weighting and summing the coherence index, art style matching index, visual fidelity index and risk elimination thoroughness index to obtain a comprehensive evaluation utility index for multi-objective optimization.
[0082] Specifically, the calculation formula for the comprehensive evaluation utility index of the multi-objective optimization is as follows: ,in Representative comprehensive evaluation utility indicators; These represent the normalized visual fidelity index, art style matching index, coherence index, and risk elimination thoroughness index, respectively. These are the weighting coefficients for the visual fidelity index, artistic style matching index, coherence index, and risk elimination thoroughness index, respectively. Furthermore, it can be dynamically adjusted according to the security level and artistic requirements of the broadcast content.
[0083] It should be noted that the weighting coefficients are preset values, and different weighting configuration templates are available for different types of programs, such as news, TV dramas, and variety shows. Furthermore, the system provides a management interface that allows reviewers to manually fine-tune the weights of specific items under special circumstances.
[0084] For example, in automatic mode, when the risk elimination thoroughness indicator is below the safety threshold for multiple consecutive frames, its weight will be automatically increased. To ensure content security is the top priority.
[0085] The candidate replacement with the highest value in the comprehensive evaluation utility index is selected as the final optimal replacement solution.
[0086] As an exemplary embodiment of the present invention, the specific synthesis method of the replaced compliant program video stream includes: establishing a transition buffer at the boundary between the optimal replacement scheme and the original video background, and smoothly connecting the gradient field of the replacement region and the gradient field of the background region at the boundary by solving the Poisson equation.
[0087] Specifically, to address the abrupt color and brightness changes at the boundary between the replacement region and the original background, a Poisson editing algorithm is employed for gradient domain fusion. First, a transition buffer is formed by extending 3 to 8 pixels inwards and outwards from the boundary of the identified violation element semantic mask. Instead of simply mixing pixel values directly within this region, the problem is transformed into solving a Poisson equation. The color values of all pixels are recalculated within this buffer, aiming to ensure that the color gradient within the newly generated region is as consistent as possible with the gradient of the optimal replacement scheme, while simultaneously requiring that the pixel values at the outermost boundary of the buffer be exactly equal to the pixel values of the original video background. By iteratively solving this equation, a pixel filling scheme with continuously changing gradients at the boundary is generated, thus achieving a smooth transition in color and brightness.
[0088] We apply deep learning-based automatic matting technology to generate high-precision alpha channels, achieving hair-level precision edge blending.
[0089] Specifically, to achieve pixel-level precision in handling complex edges, especially hair or transparent objects, a deep learning-based automatic matting technique is applied. This technique involves training a neural network model specifically to predict pixel transparency. The original video frame and a coarse three-region mask (Trimap) automatically generated based on the semantic mask of the violation elements are taken as input. This Trimap divides the image into three regions: absolute foreground, absolute background, and uncertain boundary. The model focuses on analyzing the uncertain boundary region and outputs an alpha value between 0 and 1 for each pixel within that region. This pixel-by-pixel alpha value ultimately constitutes a high-precision alpha channel, which is used in subsequent compositing steps to control the blending ratio between the optimal replacement scheme and the original background, thereby achieving fine edge blending at the hair-thin level.
[0090] A motion blur compensation algorithm is introduced to add motion blur effect to the generated replacement content based on the extracted lens motion parameters, so as to eliminate the visual incongruity caused by the generated content being too clear.
[0091] Specifically, to eliminate the visual disharmony caused by the inconsistency in motion blur between the statically generated replacement content and the dynamically captured original video, a motion blur compensation algorithm is introduced. This algorithm uses extracted lens motion parameters to calculate the motion trajectory and velocity of each pixel within the replacement area on the imaging plane during a single frame exposure time—that is, the screen-space motion vector. Based on this motion vector, a directional blur convolution kernel is applied to the synthesized replacement content area. The length and direction of this convolution kernel are precisely matched to the motion vector, and its intensity is adjusted according to simulated shutter speed parameters, such as 1 / 60 second. This operation adds a motion blur effect to the replacement content that is completely consistent with the original video, keeping its visual dynamics synchronized with the entire scene, thus avoiding a floating or pasted appearance.
[0092] This invention proposes a multi-index evaluation model covering visual fidelity, artistic style matching, coherence, and thoroughness of risk elimination. By automatically selecting the optimal replacement scheme through weighted optimization and combining advanced compositing techniques such as Poisson editing, high-precision image matting, and motion blur compensation, it ensures that the final output video stream meets broadcast-grade technical quality and safety standards, solving the problems of low quality and uncontrollable results in existing technologies.
[0093] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0094] Those skilled in the art will recognize that the algorithmic steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0095] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0097] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for real-time review and automatic replacement of illegal content in broadcast programs based on AI, characterized by: include: S1. Identify illegal elements in the video stream of broadcast programs based on a pre-trained multimodal deep learning model and determine their pixel-level spatial positions in the video frame; at the same time, perform deep analysis on the current video scene containing the illegal elements to obtain the scene deep analysis results; S2. Input the scene deep analysis results as constraints into the generative artificial intelligence model to generate at least one candidate replacement content for replacing the illegal element. The generation process must meet the constraints, including visual consistency constraints and semantic coherence constraints. S3. Perform a comprehensive evaluation of the generated multiple candidate replacement contents for multi-objective optimization, select the optimal replacement scheme, and perform multi-scale real-time synthesis with the compliant background area of the original video stream to generate and output the compliant program video stream after replacement. The scene depth analysis results include the obtained scene's geometric structure and camera motion parameters, lighting and color gamut parameters, character emotional state feature parameters, and dialogue context semantic logic feature parameters. Among them, optical flow estimation and 3D reconstruction algorithms are used to extract the geometric structure and camera motion parameters of the scene; The lighting parameters and color gamut distribution of a scene are obtained using a colorimeter model and a global illumination analysis algorithm. Extracting emotional state feature vectors of individuals using an affective computing model; We use natural language processing techniques to analyze the semantic and logical context of the dialogue in the current audio track.
2. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 1, characterized in that: The specific implementation methods for the pixel-level spatial location of the violating element include: The input broadcast program video stream is broken down into a continuous sequence of video frames, and the audio track signal is extracted simultaneously. The spatial feature vectors of video frames are extracted using a visual Transformer, and the motion feature vectors between video frames are extracted using a long short-term memory network. The violation content classifier identifies whether there are restricted symbols, prohibited items, inappropriate subtitles, controversial figures, and violations in video frames, and collectively refers to them as violation targets. At the same time, it outputs the initial bounding box coordinates of the detected violation targets. For detected illegal targets, an instance segmentation algorithm is used to generate a pixel-level semantic mask to define the geometric boundaries of the illegal elements in each frame. Lock the motion path of the violating element on the timeline of the video stream and generate a spatiotemporal bounding box sequence of its complete lifecycle.
3. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 1, characterized in that: The specific methods for generating the candidate replacement content include: Based on the attribute tags of the violating element, relevant materials are retrieved from the preset compliant material library, and their matching scores with the violating element in terms of shape, texture, and size are calculated. If the matching score is higher than the first threshold, the material is used first. If the matching score is lower than the first threshold but higher than the second threshold, the aforementioned material is used as the base to start the real-time generation process for adaptation adjustment. If the matching score is lower than the second threshold, the real-time generation process is activated to synthesize a new image that matches the current scene description. The selected material from the compliant material library or the new image is mapped into a high-dimensional latent space, and the latent variables are adjusted to make them distributed in the neighborhood of the original scene feature vector; An optical flow guidance mechanism is introduced to constrain the consistency of replacement content generated in consecutive frames, thereby eliminating flickering and artifacts common in generated videos. Calculate the depth occlusion relationship between other objects in the scene and the illegal area, and generate a depth map of the replacement area through a depth estimation model.
4. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 3, characterized in that: The visual consistency constraint is specifically defined as the generated replacement content maintaining continuous consistency with the original video background in terms of texture, lighting, shadow, color temperature, and motion trajectory. The specific implementation of the visual consistency constraint includes: by extracting the style features and frequency distribution of the original background, adjusting the statistical moments of the generated replacement content so that the global histogram and local contrast of the replacement area are completely matched with the background. Ghosting detection and edge smoothing algorithms are used to process the boundary between the replacement content and the background, and a transition buffer is established. Based on the calculated physical lighting parameters, virtual shadows and environmental occlusions are projected in real time onto the surface of the replacement content, and sensor noise and lens blur effects are simulated to match the imaging characteristics of the original camera.
5. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 4, characterized in that: The specific definition of the semantic coherence constraint is that the generated replacement content maintains a strong correlation with the original video context in terms of plot logic, character interaction, and emotional expression. The specific implementation methods of the semantic coherence constraint include: constructing a domain knowledge graph and extracting key entities and logical relationships from the program script or dialogue context; The extracted emotional state feature parameters of the characters are used as emotional benchmarks and mapped onto a three-dimensional emotional space. When generating candidate replacement content, the semantic labels of the generated virtual objects or patched areas are logically consistent with the context entities in the knowledge graph, and the Euclidean distance change of the sentiment vector is within the preset smoothing threshold range.
6. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 1, characterized in that: The multi-objective optimization includes coherence indicators, artistic style matching indicators, visual fidelity indicators, and risk elimination thoroughness indicators.
7. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 6, characterized in that: The specific calculation methods for the coherence index, artistic style matching index, visual fidelity index, and risk elimination thoroughness index include: Calculate the structural similarity and peak signal-to-noise ratio between each candidate replacement content and the edge of the original scene, normalize them, and then obtain the visual fidelity index by weighted averaging. The Euclidean distance between the replacement region and the neighboring background feature maps on multiple convolutional layers is calculated, and the difference value obtained by averaging them is mapped to a score interval of 0 to 1 as an art style matching index. A deep discriminant network is used to detect forgery traces in the replaced image, and a coherence index is obtained based on this. Establish a compliance secondary verification mechanism and calculate the risk elimination thoroughness indicator.
8. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 7, characterized in that: The specific selection methods for the optimal replacement scheme include: By weighting and summing the coherence index, art style matching index, visual fidelity index, and risk elimination thoroughness index, a comprehensive evaluation utility index for multi-objective optimization is obtained. The candidate replacement with the highest value in the comprehensive evaluation utility index is selected as the final optimal replacement solution.
9. The AI-based real-time review and automatic replacement method for illegal content in broadcast programs according to claim 1, characterized in that: The specific synthesis method of the replaced compliant program video stream includes: A transition buffer is established at the boundary between the optimal replacement scheme and the original video background. The gradient field of the replacement region and the gradient field of the background region are smoothly connected at the boundary by solving the Poisson equation. A deep learning-based automatic matting technique is applied to generate a high-precision alpha channel, achieving hair-level precision edge fusion. A motion blur compensation algorithm is introduced to add motion blur effect to the generated replacement content based on the extracted lens motion parameters, so as to eliminate the visual incongruity caused by the generated content being too clear.
Citation Information
Patent Citations
CN119815073A
CN121278126A