Live broadcast behavior tracking system based on deep learning

Through deep learning technology, combined with multimodal feature extraction and dynamic weighted fusion of video, audio and text data, the problems of detection lag and low recognition accuracy in live content supervision are solved, and real-time and intelligent control of violations is achieved.

CN120708001AInactive Publication Date: 2025-09-26GUANGZHOU QUNGE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510726701.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in live content supervision have problems such as delayed detection, untimely response, and low recognition accuracy. They are particularly difficult to deal with hidden violations and multi-dimensional violation factors in high-concurrency, multi-channel environments.

Method used

A live broadcast behavior tracking system based on deep learning is adopted. The multimodal feature extraction module extracts feature vectors from video, audio and text data, and uses the attention mechanism for dynamic weighted fusion, combined with the real-time response module to identify and handle violations.

Benefits of technology

It achieves high concurrency, low latency, and full-link intelligent control of live broadcast content, significantly improves the ability to identify complex scenarios and hidden violations, can output violation category labels and confidence scores in real time, and execute automatic alarms and masking measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708001A_ABST
    Figure CN120708001A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of live broadcast behavior monitoring, in particular to a live broadcast behavior tracking system based on deep learning, which obtains high-quality multi-source information and improves the accuracy of feature analysis by synchronously extracting image and audio data from a live broadcast video stream and combining frame extraction, image enhancement and voice recognition. According to the method, image features are extracted through a pre-trained convolutional neural network, audio features are extracted through a deep learning model, voice transliteration texts are fused, the weight of each modal feature is dynamically adjusted based on an attention mechanism, precise recognition of complex scenes and hidden violation behaviors is achieved, and image camouflage and latent language expression risks are effectively coped with. And the illegal type and confidence are output in real time, once suspected illegal behaviors are detected, alarm, interruption or shielding operation is triggered immediately, and related evidences are uploaded to an auditing database. And efficient, accurate and full-process management and control of the live broadcast violation behaviors are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of live broadcast behavior monitoring, and in particular to a live broadcast behavior tracking system based on deep learning. Background Art

[0002] Livestreaming has become a mainstream form of social interaction and content dissemination, widely used in entertainment, education, livestreaming, e-commerce, and other fields. The openness and real-time nature of livestreaming platforms greatly enriches the user experience, but also brings with it the risk of disseminating various illegal content, such as vulgar performances, violent acts, political rhetoric, and false propaganda. The emergence of such content not only harms the health of the online ecosystem but can also lead to platform penalties and negative social impact. Therefore, how to efficiently and accurately identify illegal content in livestreams has become a critical issue that the industry urgently needs to address.

[0003] Traditional methods for regulating live streaming content rely primarily on manual inspections, keyword filtering, or automated detection based on a single modality, such as text or images. These methods suffer from significant flaws, including delayed detection, untimely responses, and low recognition accuracy. First, manual review struggles with the high-concurrency, multi-channel live streaming environment, and can be difficult to detect subtle violations such as borderline content manipulation and scene disguises. Second, single-modality detection algorithms are limited by data noise, diverse expression methods, and adversarial disguises, often failing to fully identify complex scenarios and multi-dimensional violation factors. Summary of the Invention

[0004] To solve the above problems, the present invention provides a live broadcast behavior tracking system based on deep learning.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] A live streaming behavior tracking system based on deep learning, comprising:

[0007] The data preprocessing module is used to extract raw video data and audio data based on the live video stream, perform frame extraction and image enhancement processing on the raw video data to obtain an image frame sequence; and perform speech recognition on the audio data to obtain a transcribed text;

[0008] A multimodal feature extraction module, which extracts image feature vectors based on image frame sequences, audio feature vectors from denoised audio data, and text feature vectors from transcribed text.

[0009] The fusion discrimination module receives image feature vectors, audio feature vectors, and text feature vectors, performs dynamic weighted fusion based on the attention mechanism, generates a fusion discrimination result, and outputs a violation category label and confidence score;

[0010] The real-time response module is used to determine whether there is a suspected violation based on the violation category label and confidence score, trigger real-time alarms, automatic interruption and masking operations, and upload relevant image frames, audio clips and transcribed text as violation evidence to the audit database.

[0011] Furthermore, the extracting of original video data and audio data based on the live video stream and performing frame extraction and image enhancement processing on the original video data includes the following steps:

[0012] Extracting continuous video frames from the input live video stream according to a preset time interval;

[0013] De-noising, brightness equalization and sharpening are performed on the continuous video frames to obtain an image frame sequence.

[0014] Furthermore, extracting the image feature vector based on the image frame sequence includes:

[0015] The image frame sequence is used as input, a pre-trained convolutional neural network is output, features are extracted for each frame of image, and an image feature vector corresponding to each frame of image is output.

[0016] Furthermore, the convolutional neural network is trained by the following steps:

[0017] constructing training samples based on image frames containing combined labels, wherein the combined labels include character clothing type, body movement type, and background scene type;

[0018] Initialize the network structure and parameters of the convolutional neural network;

[0019] The image frames in the training set are used as input to the convolutional neural network, which performs forward propagation on each frame and outputs the image feature vector;

[0020] Based on the combined labels and image feature vectors in the training set, the network parameters are optimized according to the loss function.

[0021] Furthermore, extracting the audio feature vector based on the denoised audio data includes:

[0022] The denoised audio data is used as input, and the Mel-frequency cepstral coefficients are used to perform feature transformation on the audio signal. The obtained audio feature sequence is input into a pre-trained deep learning model, and an audio feature vector for characterizing the illegal audio content is output.

[0023] Furthermore, the dynamic weighted fusion based on the attention mechanism includes the following steps:

[0024] The image feature vector, audio feature vector, and text feature vector are input into the multimodal interactive attention layer respectively, and the correlation weights between the features of each modality are calculated to obtain the initial weighting coefficients.

[0025] Based on historical discrimination results and current discrimination confidence, dynamically adjusting the weighting coefficient to obtain an updated weighting coefficient;

[0026] The multimodal feature vectors weighted by the updated weighting coefficients are fused, and the fused feature vectors are input into the discriminant decision network to output the fused discrimination result.

[0027] Furthermore, the dynamic adjustment of the weighting coefficient based on the historical discrimination results and the current discrimination confidence level includes the following steps:

[0028] Based on the historical discrimination confidence sequence within a preset time window, the historical confidence mean and standard deviation corresponding to the image feature vector, audio feature vector and text feature vector are calculated;

[0029] Based on the discrimination confidence of the current image feature vector, audio feature vector and text feature vector and their corresponding historical confidence mean and standard deviation, the initial weighting coefficient of each modal feature is adjusted to obtain the updated weighting coefficient.

[0030] Furthermore, the formula for adjusting the initial weighting coefficient of each modal feature is as follows:

[0031] w′ i =w i +α·(c i -μ i )-β·σ i ;

[0032] Among them, w i is the initial weighting coefficient of the i-th modal feature; w′ i is the adjusted weighting coefficient of the i-th modal feature; c i is the confidence level of the i-th modal feature at the current moment; μ i is the historical confidence mean of the i-th modal feature within the preset time window; σ i is the historical confidence standard deviation of the i-th modal feature within the preset time window; α and β are weight adjustment parameters.

[0033] Furthermore, the discrimination decision network is constructed by the following steps:

[0034] Construct a training set based on the fused feature vectors and the corresponding violation category labels;

[0035] Constructing a feedforward neural network, inputting the fused feature vectors in the training set into the feedforward neural network to obtain an intermediate feature representation, and outputting a probability distribution of violation categories after passing a softmax function based on the intermediate feature representation;

[0036] Based on the probability distribution of violation categories and the violation category labels in the training set, a cross entropy loss function is used to perform supervised training on the feedforward neural network to obtain a discriminant decision network.

[0037] Furthermore, the formula of the cross entropy loss function is as follows:

[0038]

[0039] Where L is the cross entropy loss value; N is the number of training samples in the training set; C is the total number of categories; y ij is the jth class true label of the i-th sample; p ij The predicted probability that the i-th sample output by the feedforward neural network belongs to the j-th category.

[0040] The beneficial effects of the present invention are as follows: the present invention synchronously extracts multi-source image and audio data from live video streams, performs frame extraction and image enhancement respectively, and obtains high-quality text information through speech recognition, significantly improving the reliability of subsequent feature analysis. Feature analysis of image frame sequences is performed through a pre-trained convolutional neural network, sound features are extracted through a deep learning audio model, and text semantics are mined in combination with speech transcription results, fully capturing behavioral clues in different dimensions such as clothing, action, scene, and speech. A dynamic weighting strategy based on an attention mechanism is further introduced to adaptively fuse each modal feature based on actual discrimination confidence and historical discrimination performance, significantly improving the ability to identify complex scenarios and hidden violations, especially addressing new risks such as image disguise, cryptic expressions, and multimodal combination violations. Violation category labels and confidence scores are output in real time, and through a real-time response module, when suspected violations are detected, alarms, interruptions, or masking measures are automatically executed, and the corresponding multimodal evidence is uploaded to the audit database, achieving intelligent management and control capabilities with high concurrency, low latency, and full-link evidence retention. Through the above-mentioned technological innovations, this solution effectively overcomes the shortcomings of traditional methods in regulating live content, and achieves accurate detection and control of real-time violations. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a structural diagram of a live broadcast behavior tracking system based on deep learning in the present invention.

[0042] Figure 2 This is a flowchart of the steps of dynamic weighted fusion based on the attention mechanism in the present invention. DETAILED DESCRIPTION

[0043] See also Figure 1-Figure 2 As shown, the present invention relates to a live broadcast behavior tracking system based on deep learning, comprising:

[0044] The data preprocessing module is used to extract raw video data and audio data based on the live video stream, perform frame extraction and image enhancement processing on the raw video data to obtain an image frame sequence; and perform speech recognition on the audio data to obtain a transcribed text;

[0045] A multimodal feature extraction module, which extracts image feature vectors based on image frame sequences, audio feature vectors from denoised audio data, and text feature vectors from transcribed text.

[0046] The fusion discrimination module receives image feature vectors, audio feature vectors, and text feature vectors, performs dynamic weighted fusion based on the attention mechanism, generates a fusion discrimination result, and outputs a violation category label and confidence score;

[0047] The real-time response module is used to determine whether there is a suspected violation based on the violation category label and confidence score, trigger real-time alarms, automatic interruption and masking operations, and upload relevant image frames, audio clips and transcribed text as violation evidence to the audit database.

[0048] In some embodiments, the system extracts frames from the video stream at preset time intervals to ensure that video frames at key moments are effectively captured, and applies image enhancement algorithms such as denoising, brightness equalization, and sharpening to the extracted original video frames to improve the robustness and discrimination accuracy of subsequent feature extraction. For audio data, the system uses an end-to-end deep learning noise reduction algorithm (such as an audio enhancement model based on a U-Net or DCCRN structure) to perform real-time noise reduction on the original audio stream, effectively suppressing environmental noise and echo interference. Subsequently, the system calls an adaptive speech recognition (ASR) model to transcribe the denoised audio data into structured text, laying the foundation for text feature extraction. In the multimodal feature extraction stage, the image frame sequence is input into a pre-trained and multi-label fine-tuned convolutional neural network (such as ResNet-50 or EfficientNet). The network outputs a high-dimensional vector representation of the character's clothing type, body movement state, and background scene features corresponding to each frame of the image. After undergoing Mel-Frequency Cepstral Coefficient (MFCC) feature transformation, the audio data is fed into a BiGRU or Transformer-based audio feature extraction model, which generates multidimensional feature vectors describing speech emotions, keywords, and unusual acoustic events. For text, the transcribed text is encoded using contextual semantic models such as BERT to mine text feature vectors that indicate directive, inductive, or concealed violations. These three types of features are normalized to form a time-synchronized multimodal feature sequence. During the fusion and discrimination phase, the system employs a multimodal interactive attention mechanism to dynamically weight the features of each modality. Specifically, the feature vectors of image, audio, and text are fed into a multi-head self-attention layer. The weights of each modality are adaptively adjusted based on the current moment and the historical discriminant confidence sequence, generating a fused discriminant feature. This fused feature is then fed into a feedforward neural network (FNN) discriminator, which outputs the probability distribution and confidence level of each violation through a softmax layer. Based on the output class labels and confidence levels, the system determines whether the live broadcast segment contains suspected violations. Once the judgment result shows that the probability of violation exceeds the threshold, the real-time response module automatically triggers corresponding measures, including pushing alarm information to the platform management end and the anchor end, interrupting the live stream in real time, or blocking the illegal screen. At the same time, the system automatically archives relevant video frames, audio clips and text content, and uploads them to the audit database to achieve full-link traceability and subsequent review of violation evidence. This embodiment achieves efficient, intelligent and real-time control of live broadcast violations through the organic collaboration of multimodal deep feature extraction, attention fusion and adaptive judgment decision-making. Compared with the existing analysis based only on single frames or static images, this system combines temporal context to perform sequential cascade modeling of multi-dimensional behavioral information of images, audio and text, which significantly improves the ability to identify hidden, combined and progressive violations.In addition, the dynamic weight adjustment of the attention mechanism enables adaptive response to scene changes, disguise techniques, and multimodal attacks, effectively reducing the risk of misjudgment and missed judgments. It also achieves real-time identification of violation categories and simultaneously completes full-link evidence archiving, providing data support for subsequent tracing and review. Overall, this system provides a new technical paradigm for intelligent supervision of live content through algorithmic multimodal deep feature analysis, dynamic fusion, and real-time response. It greatly improves the intelligence and automation level of violation detection in complex scenarios, significantly different from existing technical systems.

[0049] Furthermore, the extracting of original video data and audio data based on the live video stream and performing frame extraction and image enhancement processing on the original video data includes the following steps:

[0050] Extracting continuous video frames from the input live video stream according to a preset time interval;

[0051] De-noising, brightness equalization and sharpening are performed on the continuous video frames to obtain an image frame sequence.

[0052] It should be noted that to ensure that the extracted frames are both representative and temporally continuous, the frame rate is set to 3 to 5 frames per second, which can be dynamically adjusted based on scene complexity and platform resources. Frame extraction is implemented using a sliding window mechanism to avoid missing key frames. Optical flow analysis is also combined to eliminate redundant frames, resulting in a structurally stable image frame sequence. During the image enhancement stage, the system sequentially performs denoising, brightness equalization, and sharpening on each extracted frame. Image denoising is based on a fusion strategy of non-local means filtering (NLM) and a residual autoencoder (Denoising Autoencoder). The former smooths local noise, while the latter preserves edge structure and texture details. Brightness equalization utilizes adaptive histogram equalization (CLAHE) to enhance local contrast in certain regions, preventing interference from global exposure deviations on feature extraction. To enhance key image regions such as human silhouettes, gestures, and boundary textures, the system uses a multi-scale Laplacian filter for image sharpening, integrating image gradient information to preserve detailed variations. The resulting processed image frame sequence not only exhibits spatial clarity, lighting stability, and edge sharpness, but also possesses high semantic recognizability, effectively supporting the subsequent convolutional neural network model for deep representation learning of clothing features, action recognition, and scene classification. Compared to the simple image cropping and mean filtering used in the prior art, this processing chain provides greater discriminative feature retention and noise robustness, significantly improving the accuracy of multimodal feature extraction and the overall recognition precision of the system.

[0053] Furthermore, extracting the image feature vector based on the image frame sequence includes:

[0054] The image frame sequence is used as input, a pre-trained convolutional neural network is output, features are extracted for each frame of image, and an image feature vector corresponding to each frame of image is output.

[0055] Furthermore, the convolutional neural network is trained by the following steps:

[0056] constructing training samples based on image frames containing combined labels, wherein the combined labels include character clothing type, body movement type, and background scene type;

[0057] Initialize the network structure and parameters of the convolutional neural network;

[0058] The image frames in the training set are used as input to the convolutional neural network, which performs forward propagation on each frame and outputs the image feature vector;

[0059] Based on the combined labels and image feature vectors in the training set, the network parameters are optimized according to the loss function.

[0060] In some embodiments, the image frame sequence after image enhancement processing is used as input and sequentially fed into a pre-trained convolutional neural network model. The convolutional neural network preferably uses the ResNet-50 structure as the backbone network, and on the basis of pre-training, it is fine-tuned using a multi-label data set containing character clothing type, body action category and background scene labels to enable it to have stronger multi-task discrimination capabilities. The network input layer receives image frames of fixed size (for example, 224×224 pixels) and extracts spatial semantic information in the image through multiple convolutional layers, batch normalization layers, nonlinear activation functions ReLU and residual connection structures. After the convolution feature map is generated, the global average pooling layer (GAP) is used to compress the spatial dimension to obtain a high-dimensional image feature vector with semantic concentration capability. The vector dimension is generally set to 2048 dimensions. In order to further improve the discriminability and multi-label expression ability of the features, the network output layer is designed as a multi-channel parallel output structure, corresponding to multiple sub-label dimensions such as clothing, action and scene, and the Sigmoid activation function is used to independently model each type of attribute. During the training phase, the loss function takes the form of a summation of weighted cross-entropy losses, calculating errors for different label branches separately and jointly optimizing network parameters through a unified back-propagation process. During the inference phase, the network generates a corresponding image feature vector for each frame, which serves as the key visual semantic representation input for the subsequent fusion module. This feature not only encodes the underlying texture and structure of the image but also implicitly integrates high-level semantic information such as clothing exposure, movement amplitude, and scene complexity through multi-task learning, helping to improve the system's ability to identify complex, blurred, or well-disguised violations.

[0061] Furthermore, extracting the audio feature vector based on the denoised audio data includes:

[0062] The denoised audio data is used as input, and the Mel-frequency cepstral coefficients are used to perform feature transformation on the audio signal. The obtained audio feature sequence is input into a pre-trained deep learning model, and an audio feature vector for characterizing the illegal audio content is output.

[0063] In some embodiments, first, the input noise reduction audio segment is time-windowed and short-time Fourier transform (STFT) processed, and on this basis, the Mel-frequency cepstral coefficients (MFCC) are extracted. This feature extraction process simulates the frequency response characteristics of the human auditory perception system. By mapping the spectrum to the Mel scale and calculating the discrete cosine transform (DCT) of the logarithmic power spectrum, the redundancy of the frequency domain information is compressed while retaining the acoustic structural characteristics of the speech. The extracted MFCC sequence is in the form of a two-dimensional matrix, with each frame containing several dimensional frequency coefficients, representing the continuous feature evolution of the audio signal in the time dimension. The sequence is then input into a pre-trained deep learning model, preferably an audio encoder based on a bidirectional gated recurrent unit (BiGRU) or Transformer structure. For the BiGRU structure, it can simultaneously model the forward and backward time dependencies of the audio signal and capture the context-related acoustic patterns in the speech segment; for the Transformer-based structure, the multi-head attention mechanism is used to enhance the feature representation of the key time period, which is particularly suitable for locating and modeling potential violation characteristics such as abnormal emotions, sudden changes in speech speed, and strong changes in tone. During the training phase, the audio model receives audio samples labeled with illegal semantic categories (such as violent language, sexual innuendo, abusive tone, drinking sound effects, etc.), and uses supervised learning with cross-entropy or focal loss to optimize parameters. After training, the model can map any noise-reduced audio segment into an audio feature vector with high-level behavioral semantic expression capabilities. The vector dimension is generally set to 128 to 512 dimensions, which can be adjusted according to the model architecture and task requirements. Compared with traditional audio analysis methods that only rely on keyword matching or emotion intensity scoring, this embodiment combines low-level acoustic signals with high-level semantic labels to construct a deep audio feature space with behavioral representation capabilities, enabling the system to have the ability to recognize illegal speech content that is semantically ambiguous, implicitly expressed, or masked by background sounds.

[0064] Furthermore, the dynamic weighted fusion based on the attention mechanism includes the following steps:

[0065] The image feature vector, audio feature vector, and text feature vector are input into the multimodal interactive attention layer respectively, and the correlation weights between the features of each modality are calculated to obtain the initial weighting coefficients.

[0066] Based on historical discrimination results and current discrimination confidence, dynamically adjusting the weighting coefficient to obtain an updated weighting coefficient;

[0067] The multimodal feature vectors weighted by the updated weighting coefficients are fused, and the fused feature vectors are input into the discriminant decision network to output the fused discrimination result.

[0068] In some embodiments, the image feature vector of each frame of image, the audio feature vector of each audio segment, and the text feature vector corresponding to each speech transcription are extracted respectively, and the three types of features are simultaneously input into the multimodal interactive attention layer. In this attention layer, the system measures the discriminant contribution of a certain modality in the current scene by calculating the similarity relationship between each modality, and generates the initial weighting coefficients of image, audio and text based on this, which are used to represent the basic importance of each modality in the overall fusion. The above initial weighting coefficients are further dynamically revised in combination with the current discrimination confidence and historical discrimination results. Specifically, the system records the confidence performance of each modality output in the previous time window, and calculates its mean and fluctuation; when the current confidence of a modality is higher than its historical average level and the volatility is low, the system will appropriately increase the weighted ratio of the modality; on the contrary, if the confidence of a modality fluctuates greatly or is continuously lower than the average level, its fusion weight will be automatically reduced to suppress its misleading influence on the final result. The corrected weighting coefficients will act on the image, audio and text feature vectors respectively, and a fused multimodal feature representation will be generated by linear weighting. The fused feature not only retains the key discriminant information of each in the multimodal data, but also avoids the interference caused by modal failure or quality fluctuation, thereby providing a stable and representative input for the subsequent discriminant network. Finally, the fused feature is input into the feedforward neural network, and the violation category and the corresponding confidence score are output. Compared with the existing methods, the attention mechanism of this embodiment can not only establish a collaborative relationship between modalities, but also dynamically adjust the fusion strategy according to the actual recognition situation, so as to realize flexible modeling and accurate discrimination of behavioral characteristics in different scenarios and different content structures, and improve the system's detection ability for complex, changeable and weak feature violations.

[0069] Furthermore, the dynamic adjustment of the weighting coefficient based on the historical discrimination results and the current discrimination confidence level includes the following steps:

[0070] Based on the historical discrimination confidence sequence within a preset time window, the historical confidence mean and standard deviation corresponding to the image feature vector, audio feature vector and text feature vector are calculated;

[0071] Based on the discrimination confidence of the current image feature vector, audio feature vector and text feature vector and their corresponding historical confidence mean and standard deviation, the initial weighting coefficient of each modal feature is adjusted to obtain the updated weighting coefficient.

[0072] It should be noted that the system first calculates the mean and standard deviation of the confidence scores for the three modalities—image, audio, and text—over the time window. The mean measures the modality's recent average contribution, while the standard deviation reflects the volatility and uncertainty of its discrimination results. The system then obtains the modality's actual discrimination confidence at the current moment and calculates the difference between it and the historical mean to determine whether the modality's performance in the current scenario is better or worse than its recent average. If the current confidence score of a modality is significantly higher than the historical mean and has low volatility (i.e., a low standard deviation), the system considers that modality an important source of information at the current moment and appropriately increases its weighting coefficient. Conversely, if the current confidence score is lower than the historical mean or the standard deviation is large, indicating that the modality's recent discrimination is unstable, the system decreases its weighting coefficient to reduce its interference with the final fusion result. This operation is performed separately for each modality, ultimately resulting in adaptively adjusted weighting coefficients for the three modalities: image, audio, and text. This mechanism ensures that, in different live broadcast scenarios, the system dynamically optimizes the fusion strategy based on actual fluctuations in modal discrimination performance. Unlike traditional fixed-weight or average fusion, this embodiment integrates discrimination confidence with historical trends. This not only improves the model's efficiency in utilizing heterogeneous modal information, but also enhances its adaptability to complex situations such as camouflaged behavior, signal interference, or modal loss, thereby improving the robustness and accuracy of overall behavior recognition.

[0073] Furthermore, the formula for adjusting the initial weighting coefficient of each modal feature is as follows:

[0074] w′ i =w i +α·(c i -μ i )-β·σ i ;

[0075] Among them, w i is the initial weighting coefficient of the i-th modal feature; w′ i is the adjusted weighting coefficient of the i-th modal feature; c i is the confidence level of the i-th modal feature at the current moment; μ i is the historical confidence mean of the i-th modal feature within the preset time window; σ i is the historical confidence standard deviation of the i-th modal feature within the preset time window; α and β are weight adjustment parameters.

[0076] It should be noted that the current discrimination confidence of each modality is first obtained. This confidence is then combined with the modality's historical confidence mean and standard deviation recorded within a preset time window to assess its performance trend and stability during the recent discrimination process. The adjustment strategy follows the following logic: When a modality's current confidence is significantly higher than its historical average and its historical discrimination results have little fluctuation (i.e., a low standard deviation), it indicates that the modality has high credibility and stability in the current context. The system will appropriately increase its weighting coefficient, giving it a higher weight in the fusion process. Conversely, if the current confidence is lower than the historical mean or has large historical fluctuations, it indicates that the modality has uncertainty or interference. The system will accordingly reduce its weighting coefficient to minimize its impact on the final fusion result. To achieve this adjustment process, the system introduces two adjustment parameters, which are used to control the degree of influence of confidence changes and historical fluctuations on the weighted results. The system dynamically adjusts the weighting coefficients for each modality by adding an increment proportional to the difference between the current confidence level and the historical mean, and reducing a penalty term proportional to the historical standard deviation, based on the initial weighting coefficients. This adjustment mechanism, based on a joint decision-making process based on current confidence and historical performance, enables real-time optimization of the multimodal fusion process based on the reliability of modal performance in the actual scenario. This is particularly applicable to live broadcast environments with uneven modal quality, frequent information conflicts, or camouflage interference, significantly improving the stability and accuracy of the final judgment results.

[0077] Furthermore, the discrimination decision network is constructed by the following steps:

[0078] Construct a training set based on the fused feature vectors and the corresponding violation category labels;

[0079] Constructing a feedforward neural network, inputting the fused feature vectors in the training set into the feedforward neural network to obtain an intermediate feature representation, and outputting a probability distribution of violation categories after passing a softmax function based on the intermediate feature representation;

[0080] Based on the probability distribution of violation categories and the violation category labels in the training set, a cross entropy loss function is used to perform supervised training on the feedforward neural network to obtain a discriminant decision network.

[0081] In some embodiments, a training dataset is constructed by combining feature vectors formed by fusing image, audio, and text features through an attention mechanism with corresponding violation labels. Each data item represents a live streaming behavior sample at a specific moment and its corresponding judgment result. These fused features contain high-level, abstract information from the visual, auditory, and semantic levels, representing the multiple dimensions of the current behavior. Based on the training data, the system constructs a feedforward neural network as the core structure for behavior classification. This neural network receives the fused feature vector as input and progressively applies nonlinear transformations to the input data through a multi-layer, fully connected network. At each layer, potential feature combinations and class boundaries are extracted. The final layer of the network uses a softmax activation function, resulting in a probability distribution representing the probability of the current sample belonging to each violation category. To train the model, the system uses a supervised learning approach, comparing the class probability distribution output by the model with the ground-truth label and quantifying the difference between the two using a cross-entropy loss function. This loss function effectively measures the uncertainty of the model's predictions, guiding the network to continuously adjust its internal weights to learn the optimal decision boundary for the training samples. During the training process, the network updates parameters through the back-propagation mechanism and optimization algorithm, iterates and converges, and finally forms a discriminant decision model with generalization ability. After the training is completed, the feedforward neural network can receive any new fused feature vector and quickly output its corresponding violation category and confidence score as the final judgment basis of the system. This embodiment realizes the unified modeling and discriminant learning of multimodal information in high-dimensional space through a deep neural structure, significantly improving the recognition ability of complex behavior combinations, cross-modal expressions and weak feature violations. It has strong adaptability and real-time processing performance, and is suitable for multi-scenario and multi-category live content compliance detection tasks.

[0082] Furthermore, the formula of the cross entropy loss function is as follows:

[0083]

[0084] Where L is the cross entropy loss value; N is the number of training samples in the training set; C is the total number of categories; y ij is the jth class true label of the i-th sample; p ij The predicted probability that the i-th sample output by the feedforward neural network belongs to the j-th category.

[0085] Specifically, for each training sample, its true category is represented in the form of one-hot encoding, and the predicted probability of the corresponding category output by the model is obtained. For each category, the product of the true label value of the category and the model predicted probability is calculated, and then the logarithm is taken. Finally, the sum is taken for all categories and averaged for all training samples to obtain the overall loss value of the current training batch. This loss value can effectively reflect the current degree of classification deviation of the model. During the entire training process, the system will continuously update the weight parameters of the neural network through the backpropagation mechanism based on the cross-entropy loss value, pushing the model prediction results to gradually approach the true label. Since cross-entropy can impose higher penalties for situations where the prediction deviation is large, it has better convergence and classification discrimination capabilities in multi-category violation identification tasks. Using the cross-entropy loss function as the optimization target ensures that the feedforward neural network can effectively learn the mapping relationship between the fusion feature vector and the behavior category, thereby improving the discrimination stability and classification accuracy of the model when facing complex, ambiguous or unclear boundary behavior features.

[0086] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A live streaming behavior tracking system based on deep learning, characterized by: include: A data preprocessing module is used to extract raw video data and audio data based on the live video stream, perform frame extraction and image enhancement processing on the raw video data, and obtain an image frame sequence; Perform speech recognition on the audio data to obtain transcribed text; A multimodal feature extraction module, which extracts image feature vectors based on image frame sequences, audio feature vectors from denoised audio data, and text feature vectors from transcribed text. The fusion discrimination module receives image feature vectors, audio feature vectors, and text feature vectors, performs dynamic weighted fusion based on the attention mechanism, generates a fusion discrimination result, and outputs a violation category label and confidence score; The real-time response module is used to determine whether there is a suspected violation based on the violation category label and confidence score, trigger real-time alarms, automatic interruption and masking operations, and upload relevant image frames, audio clips and transcribed text as violation evidence to the audit database.

2. A live streaming behavior tracking system based on deep learning according to claim 1, characterized in that: Extracting original video data and audio data based on the live video stream and performing frame extraction and image enhancement processing on the original video data includes the following steps: Extracting continuous video frames from the input live video stream according to a preset time interval; De-noising, brightness equalization and sharpening are performed on the continuous video frames to obtain an image frame sequence.

3. A live streaming behavior tracking system based on deep learning according to claim 1, characterized in that: Extracting the image feature vector based on the image frame sequence includes: The image frame sequence is used as input, a pre-trained convolutional neural network is output, features are extracted for each frame of image, and an image feature vector corresponding to each frame of image is output.

4. A live broadcast behavior tracking system based on deep learning according to claim 3, characterized in that: The convolutional neural network is trained by the following steps: constructing training samples based on image frames containing combined labels, wherein the combined labels include character clothing type, body movement type, and background scene type; Initialize the network structure and parameters of the convolutional neural network; The image frames in the training set are used as input to the convolutional neural network, which performs forward propagation on each frame and outputs the image feature vector; Based on the combined labels and image feature vectors in the training set, the network parameters are optimized according to the loss function.

5. The live broadcast behavior tracking system based on deep learning according to claim 1, characterized in that: Extracting the audio feature vector based on the denoised audio data includes: The denoised audio data is used as input, and the Mel-frequency cepstral coefficients are used to perform feature transformation on the audio signal. The obtained audio feature sequence is input into a pre-trained deep learning model, and an audio feature vector for characterizing the illegal audio content is output.

6. A live streaming behavior tracking system based on deep learning according to claim 1, characterized in that: The dynamic weighted fusion based on the attention mechanism includes the following steps: The image feature vector, audio feature vector, and text feature vector are input into the multimodal interactive attention layer respectively, and the correlation weights between the features of each modality are calculated to obtain the initial weighting coefficients. Based on historical discrimination results and current discrimination confidence, dynamically adjusting the weighting coefficient to obtain an updated weighting coefficient; The multimodal feature vectors weighted by the updated weighting coefficients are fused, and the fused feature vectors are input into the discriminant decision network to output the fused discrimination result.

7. A live streaming behavior tracking system based on deep learning according to claim 6, characterized in that: The dynamically adjusting the weighting coefficient based on the historical discrimination results and the current discrimination confidence comprises the following steps: Based on the historical discrimination confidence sequence within a preset time window, the historical confidence mean and standard deviation corresponding to the image feature vector, audio feature vector and text feature vector are calculated; Based on the discrimination confidence of the current image feature vector, audio feature vector and text feature vector and their corresponding historical confidence mean and standard deviation, the initial weighting coefficient of each modal feature is adjusted to obtain the updated weighting coefficient.

8. A live broadcast behavior tracking system based on deep learning according to claim 7, characterized in that: The formula for adjusting the initial weighting coefficient of each modal feature is as follows: w′ i =w i +a·(c i -m i )-b·s i ; Among them, w i is the initial weighting coefficient of the i-th modal feature; w′ i is the adjusted weighting coefficient of the i-th modal feature; c i is the confidence level of the i-th modal feature at the current moment; μ i is the historical confidence mean of the i-th modal feature within the preset time window; σ i is the historical confidence standard deviation of the i-th modal feature within the preset time window; α and β are weight adjustment parameters.

9. The live broadcast behavior tracking system based on deep learning according to claim 6, characterized in that: The discriminant decision network is constructed by the following steps: Construct a training set based on the fused feature vectors and the corresponding violation category labels; Constructing a feedforward neural network, inputting the fused feature vectors in the training set into the feedforward neural network to obtain an intermediate feature representation, and outputting a probability distribution of violation categories after passing a softmax function based on the intermediate feature representation; Based on the probability distribution of violation categories and the violation category labels in the training set, a cross entropy loss function is used to perform supervised training on the feedforward neural network to obtain a discriminant decision network.

10. A live broadcast behavior tracking system based on deep learning according to claim 9, characterized in that: The formula of the cross entropy loss function is as follows: Where L is the cross entropy loss value; N is the number of training samples in the training set; C is the total number of categories; y ij is the jth class true label of the i-th sample; p ij The predicted probability that the i-th sample output by the feedforward neural network belongs to the j-th category.

Citation Information

Cited By

  • Deep learning-fused exploration scene monitoring illegal behavior automatic identification method and system

    CN121659076A

  • Anchor intention recognition method and device

    CN121744017A

  • Incremental learning live broadcast multi-type violation early warning method and system

    CN122340285A