Network audio-visual content auditing method and system based on large model
By dynamically adjusting the model sampling interval and risk prediction, the problems of resource waste and lag in the online audiovisual content review system have been solved, realizing the rational allocation of resources and real-time risk monitoring, thereby improving review efficiency and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-13
AI Technical Summary
Existing online audiovisual content review systems suffer from resource waste due to fixed-frequency sampling and lag in threshold triggering strategies, making it difficult to balance review costs with real-time security.
By acquiring multi-dimensional datasets of live video data and combining them with changes in risk, the model sampling interval is dynamically adjusted. The sampling frequency is reduced in low-risk scenarios and the sampling frequency is increased in high-risk scenarios to capture key information in a timely manner. Feature data is extracted using a lightweight BERT model and a pre-trained sentiment analysis model to calculate the comprehensive risk entropy and risk change acceleration, thereby enabling risk prediction and timely intervention.
It effectively solves the problems of resource waste and lag, saves computing resources, promptly detects violations, reduces the spread of harmful information, and improves review efficiency and real-time performance.
Smart Images

Figure CN121665029A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology. More specifically, this invention relates to a method and system for reviewing online audiovisual content based on a large model. Background Technology
[0002] With the rapid development of mobile internet technology, online audiovisual content has become the core carrier of information dissemination. In order to maintain the health and security of cyberspace, platforms need to conduct real-time review of massive live broadcasts and video streams in order to identify and block various harmful information. At present, the mainstream solution in the industry is to introduce large language models or multimodal large models and use their powerful semantic understanding capabilities to perform frame extraction analysis on video streams.
[0003] However, there is a significant contradiction between the high inference cost of large models and the real-time requirements of massive concurrency. Existing review systems mainly adopt fixed-frequency sampling or simple threshold triggering strategies. The fixed-frequency sampling strategy refers to mechanically capturing images and submitting them for review at set time intervals. This will cause a huge waste of computing power in the monotonous content of the live broadcast room. For sudden instantaneous violations, the sparsity of the sampling interval makes it easy to miss detection.
[0004] Simple threshold triggering strategies typically rely on the current value of a single indicator such as volume or bullet screen speed for judgment. Only when the indicator exceeds a preset threshold is a large-scale model review triggered. This linear judgment method has obvious phase lag. That is, when the system detects an abnormal indicator and mobilizes computing resources, the violation has often already occurred or even ended. The system is in a state of passively chasing risks, making it difficult to achieve pre-event prediction or timely blocking during the event. It cannot meet the dual requirements of real-time performance and cost control for industrial applications. Summary of the Invention
[0005] To address the technical problems of existing technologies, such as the large waste of computing power in fixed-frequency sampling and the lag in threshold triggering strategies, which makes it difficult for the system to balance audit costs and real-time security, this invention provides solutions in the following aspects.
[0006] In a first aspect, the present invention provides a method for reviewing online audiovisual content based on a large model, comprising: The system acquires a set of video information from live streams to be tested within a preset time period. Based on the phonological features and interaction popularity of the information in the video information set, it performs multidimensional extraction and analysis to obtain a multidimensional data set of the live streams to be tested. This multidimensional data set is then normalized to form a current multidimensional data profile representing the current operating status of the live stream. The system analyzes the disorder and fluctuation of the live stream's operating status using this current multidimensional data profile and calculates the comprehensive risk entropy. Based on the fluctuation intensity of the comprehensive risk entropy, it obtains the risk change acceleration. Based on the continuity between the current state and historical predictions of the live stream's operating status, and combined with the risk change acceleration, it calculates the predicted risk entropy value. Based on the magnitude of the predicted risk entropy value, the live stream's operating status is classified into low-risk and high-risk states. Based on the difference between the predicted risk entropy value and the mean comprehensive risk entropy, the minimum sampling interval is adjusted to obtain the model sampling interval for a high-risk live stream. A countdown is performed based on the model sampling interval. At the end of the countdown, audiovisual content data is extracted, and a large model is invoked to perform security inference on the live stream.
[0007] This invention effectively solves the resource waste caused by fixed-frequency sampling in existing review methods and the lag problem of threshold triggering strategies. By extracting live broadcast-related information from multiple dimensions and making predictions based on changes in risk, the model sampling interval is dynamically adjusted. In low-risk scenarios, the sampling frequency is reduced to avoid unnecessary resource consumption; in high-risk scenarios, sampling is encrypted to capture key information in a timely manner. This approach does not require high-frequency sampling of all live broadcasts, saving computing resources, and can predict risks in advance, which helps to intervene in a timely manner in the early stages of violations and avoid missed detections or the spread of violations due to lag.
[0008] Preferably, the multidimensional dataset of the live video to be tested is obtained, including: A lightweight BERT model was used to statistically analyze the sensitive word hit density sequence in the text data of the live video under test. Simultaneously, the text data sequence was input into a pre-trained sentiment analysis model to obtain a sequence of the degree of radicalism in speech. The energy value of each frame of the audio buffer stream was calculated using a short-time energy analysis algorithm, denoted as the audio short-time energy sequence. The number of times the temporal waveform of a single frame of the audio buffer stream crossed the zero horizontal axis was counted, denoted as the speech zero-crossing rate sequence. The number of newly added bullet comments and the number of new user IDs entering the live room within the analysis window were statistically analyzed and denoted as the bullet comment addition rate sequence and the user entry rate sequence, respectively. The obtained multi-dimensional data sequences were time-aligned and denoted as the multi-dimensional data set of the live video under test.
[0009] Preferably, forming a current multidimensional data profile representing the current operating status of the live broadcast room includes: Starting from the current moment, the preset length is... Within a time window, the maximum-minimum normalization algorithm is applied to each sequence in the multidimensional data set within the time window to map the data corresponding to the current moment to the standard interval [0,1]. The normalized results are then spliced together in a preset order to form a current multidimensional data profile representing the current operating status of the live broadcast room. The current multidimensional data profile includes feature data from multiple dimensions, such as the density of sensitive words, the degree of radicalism of speech, the short-term energy of audio, the zero-crossing rate of voice, the rate of new bullet comments, and the rate of users entering the room.
[0010] Preferably, the comprehensive risk entropy satisfies the following expression: ; In the formula, This represents the overall risk entropy at the current moment; This indicates the number of feature data categories participating in the calculation in the current multidimensional data profile; , Feature data representing the current time and the previous time. ; Representing feature data In length of The average value within the time window; This represents the pre-acquired feature data. Weighting coefficients; This represents the pre-obtained mutation sensitivity coefficient; Represents a logarithmic function; Represents a very small positive number, ensuring that the denominator is not zero; Represents the maximum value function; This represents an exponential function with the natural constant as its base.
[0011] This invention effectively reflects the changes and anomalies of various situations in live broadcasts by calculating comprehensive risk entropy. It can keenly capture sudden situations in live broadcasts. The comprehensive risk entropy expression takes into account the influence of different factors and the changing trends of situations, making risk assessment more comprehensive. It avoids focusing only on the situation at a single moment and ignoring the process of change, and can discover potential risks more promptly, thus buying time for subsequent risk warnings and responses.
[0012] Preferably, obtaining the risk change acceleration includes: Obtain the comprehensive risk entropy at the current time, the previous time, and the two time points before the current time; calculate the difference between the comprehensive risk entropy at the current time and the comprehensive risk entropy at the previous time, and denote it as the risk change rate; calculate the difference between the comprehensive risk entropy at the previous time and the comprehensive risk entropy at the two time points before the current time, and denote it as the risk change rate at the previous time; calculate the difference between the risk change rate and the risk change rate at the previous time to obtain the risk change acceleration, which characterizes the speed of change in the intensity of the comprehensive risk entropy fluctuation.
[0013] Preferably, the predicted risk entropy value satisfies the following expression: ; In the formula, and This represents the predicted risk entropy values for the next and current time points; and The rate of risk change and the acceleration of risk change represent the overall risk entropy at the current moment; This represents the pre-obtained smoothing coefficient. ; This represents the pre-obtained trend enhancement coefficient; This represents the absolute value function.
[0014] This invention integrates current and historical risk situations, while also considering the trend of risk changes, making the prediction results more stable and realistic. It avoids prediction deviations caused by data anomalies at a single moment, and can keenly capture the sudden change trend of risk, making the prediction results more valuable.
[0015] Preferably, the operational status of the live streaming room is classified into low-risk and high-risk statuses, including: Get the latest data up to the current time. The comprehensive risk entropy of each sampling point is denoted as the average comprehensive risk entropy. The maximum sampling interval and the minimum sampling interval are obtained from the database. If the predicted risk entropy value at the next moment is less than the average comprehensive risk entropy, the current live video under test is considered to be a low-risk video, and the live room operation status is low-risk. The maximum sampling interval is used for sampling until the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy. If the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy, the current live video under test is considered to be a high-risk video, and the live room operation status is high-risk.
[0016] This invention classifies the risk level of the live broadcast room's operation based on risk prediction results, and adopts different sampling strategies for different risk levels. When the risk is low, the sampling frequency is reduced to avoid unnecessary resource consumption; when the risk is high, it is given special attention to ensure timely monitoring. This classification and processing method allows resources to be allocated rationally.
[0017] Preferably, the model sampling interval satisfies the following expression: ; In the formula, Indicates the model sampling interval; Indicates the minimum sampling interval; Indicates the maximum sampling interval; This represents the predicted risk entropy value for the next moment; This represents the average of the overall risk entropy; This indicates a pre-acquired risk sensitivity threshold; This represents the pre-obtained step response coefficient.
[0018] This invention flexibly adjusts the model's sampling interval based on the degree to which the risk of the live streaming room deviates from the normal level. The higher the risk, the more intensive the sampling, enabling more detailed monitoring of risk changes; when the risk decreases, the sampling interval is appropriately widened to avoid wasting resources. This dynamic adjustment method makes sampling more targeted.
[0019] Preferably, at the end of the countdown, the audiovisual content data is captured and a large model is invoked to perform security inference on the live stream, including: The system counts down based on the model sampling interval. When the countdown ends, it immediately captures the current video keyframe or audio slice and calls the cloud-based big model API to perform content security inference. If the big model determines that the content is safe, the system continues to run with the current logic. If the big model determines that the content is in violation, the system immediately triggers the blocking mechanism. If the big model determines that the content is suspected or gives a high level of risk confidence, the system uses this feedback signal as an instruction to maintain a high-frequency review status for subsequent monitoring of the live stream until the big model determines that the content is safe multiple times in a row.
[0020] Secondly, the present invention provides a network audiovisual content review system based on a large model, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned network audiovisual content review method based on a large model is implemented.
[0021] By adopting the above technical solution, a computer program is generated from the above-mentioned network audiovisual content review method based on a large model and stored in a memory so that it can be loaded and executed by a processor. In this way, a terminal device can be made based on the memory and the processor for convenient use.
[0022] The beneficial effects of this invention are as follows: This invention provides a more reasonable and efficient solution for online audiovisual content review. By comprehensively collecting various information from live streams, scientifically analyzing the risk status of live streams and predicting their changing trends, it dynamically adjusts the investment of review resources, making the review work more targeted and flexible. Under the premise of ensuring the health and security of cyberspace, it rationally allocates computing resources, avoiding resource waste and improving overall review efficiency. This review method can not only promptly detect and handle illegal content and reduce the spread of harmful information, but also reduce interference with normal live streams, providing users with a good network environment. At the same time, its scientific logic and flexible adjustment mechanism enable it to adapt to the review needs of different types and scenarios of live streams, possessing strong practicality and promotional value, and providing strong support for the standardized development of the online audiovisual industry. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a network audiovisual content review method based on a large model according to the present invention; Figure 2 This diagram schematically illustrates the multimodal real-time data stream input of a network audiovisual content review method based on a large model according to the present invention. Figure 3 This is an illustrative comparison of risk entropy manifold calculation and trend prediction in a network audiovisual content review method based on a large model according to the present invention. Figure 4 This diagram illustrates the dynamic scheduling results of computing power based on risk prediction in a network audiovisual content review method based on a large model, as described in this invention. Detailed Implementation
[0024] This invention discloses a method for reviewing online audiovisual content based on a large model, referring to... Figure 1 This includes steps S1-S4: S1: Obtain the video information set of the live video to be tested within a preset time period. Based on the phonetic features and interaction popularity of the information in the video information set, obtain the multidimensional data set of the live video to be tested through multidimensional extraction and analysis. Normalize the multidimensional data set to form the current multidimensional data profile representing the current operating status of the live room.
[0025] It should be noted that at the data access layer, the audio and video buffer streams and metadata streams of the live video to be tested on the network are acquired in real time. In order to avoid the heavy image processing overhead and ensure millisecond-level response, this step focuses on extracting lightweight statistical and acoustic features to transform unstructured audio-visual streams into structured time-series data.
[0026] Obtain a set of video information from the live stream to be tested within a preset time period. Based on the phonetic features and interaction popularity of the information in the video information set, obtain a multidimensional data set of the live stream to be tested through multidimensional extraction and analysis, including: Within a preset time period, the real-time text stream and bullet screen text of each frame of the live video under test, automatically transcribed from speech recognition, are recorded as text data. A lightweight BERT model is used to statistically analyze the sensitive word hit density sequence in the text data. Simultaneously, the text data sequence is input into a pre-trained sentiment analysis model to obtain a sequence representing the degree of radicalism of the speech, which represents the sentiment polarity score. The audio buffer stream of the live video under test within the same preset time period is obtained, and the energy value of each frame of the audio buffer stream is calculated using a short-time energy analysis algorithm, which is recorded as the audio short-time energy sequence representing the change in sound loudness. The number of times the time-domain waveform of a single frame of the audio buffer stream crosses the zero horizontal axis is counted, which is recorded as the speech zero-crossing rate sequence representing the abruptness and complexity of the speech. The metadata stream within the same preset time period is obtained, and an analysis window with a fixed time length of L is set. The number of newly added bullet screens and the number of user IDs newly entering the live room within the analysis window are counted, which are recorded as the bullet screen addition rate sequence and the user entry rate sequence, respectively. The obtained multi-dimensional data sequences are time-aligned and recorded as the multi-dimensional data set of the live video under test.
[0027] It should be noted that, Figure 2 This diagram illustrates the input of multimodal real-time data streams, showing the sensitive word hit density sequence and short-time audio energy sequence extracted from the live video under test, along with the depicted bullet screen density flow curves and audio energy flow curves. (Timeline...) Within the interval, Figure 2 The video showed two consecutive instances of violations within the live stream, with the barrage density flow curve first appearing in... The first surge occurred, peaking at approximately 0.6, followed by a slight decline, and then... Another even larger surge occurred, peaking at approximately 1.0. This bimodal pattern accurately reflects the characteristic of online public opinion surging one wave after another. and In the non-risk range, the data only presents low-amplitude background noise that conforms to physical laws.
[0028] At this point, a multi-dimensional dataset of the live video to be tested has been obtained.
[0029] It should be noted that in the calculation process of multidimensional data fusion, different feature data have different physical meanings and numerical dimensions. If these original data that span several orders of magnitude are directly used for calculation, the feature data with increased values will dominate the calculation weight. Therefore, this invention compresses all features to the same order of magnitude by establishing a unified data standard.
[0030] Specifically, the multidimensional dataset is normalized to form a current multidimensional data profile representing the current operational status of the live streaming room, including: Starting from the current moment, the preset length is... Within a time window, the sensitive word hit density sequence, speech radicalism sequence, audio short-time energy sequence, voice zero-crossing rate sequence, bullet screen addition rate sequence, and user entry rate sequence of the live video to be tested are processed using a maximum-minimum normalization algorithm. The data corresponding to the current moment is mapped to the standard interval [0,1]. The normalized results are then spliced together in a preset order to form a current multidimensional data profile representing the current operating status of the live room. The current multidimensional data profile includes feature data from multiple dimensions, including sensitive word hit density, speech radicalism, audio short-time energy, voice zero-crossing rate, bullet screen addition rate, and user entry rate.
[0031] Thus, a multidimensional data profile representing the current operational status of the live streaming room has been obtained.
[0032] S2: Combine the current multidimensional data profile analysis to determine the chaotic state and volatility of the live broadcast room's operation, and calculate the comprehensive risk entropy; based on the volatility intensity of the comprehensive risk entropy, obtain the risk change acceleration; according to the continuity between the current state of the live broadcast room's operation and the historical predicted state, and combined with the risk change acceleration, calculate the predicted risk entropy value.
[0033] It should be noted that the risks of online audiovisual content are often accompanied by drastic changes in information or the aggregation of specific signals, such as a sudden change in volume due to heated arguments or a surge in comments due to violations. Simple linear weighting cannot reflect the risk indicators brought about by sudden changes. Therefore, this invention introduces an exponential term to nonlinearly amplify the first-order difference of the features, i.e., the rate of change, making the system extremely sensitive to sudden situations. In order to calculate the degree of chaos or change in the current operating state of the live broadcast room, this invention constructs a nonlinear comprehensive risk entropy.
[0034] Preferably, by combining the current multi-dimensional data profile analysis of the chaotic state and volatility of the live broadcast room's operation, a comprehensive risk entropy is calculated, including: Based on business experience, a mutation sensitivity coefficient is set to amplify sudden changes in the live broadcast room, and preset weight coefficients corresponding to all multidimensional feature data are obtained from the database in advance.
[0035] The comprehensive risk entropy satisfies the following expression: ; In the formula, This represents the overall risk entropy at the current moment; This indicates the number of feature data categories participating in the calculation in the current multidimensional data profile; , Feature data representing the current time and the previous time. ; Representing feature data In length of The average value within the time window; This represents the pre-acquired feature data. Weighting coefficients; This represents the pre-obtained mutation sensitivity coefficient; Represents a logarithmic function; Represents a very small positive number, ensuring that the denominator is not zero; Represents the maximum value function; This represents an exponential function with the natural constant as its base.
[0036] In the formula, This represents the feature data in the current multidimensional data profile. The first-order difference represents the feature data. The rate of change at the current moment; Indicates only for feature data The sudden increase is magnified to ensure that small sudden increases can be identified; Representing feature data Deviation from its length The degree of normality within the time window; Representing feature data Risk contribution value; This represents the sum of the risk contribution values of all categories of feature data in the current multidimensional data profile; This represents the comprehensive risk entropy at the current moment, calculated as a logarithmic sum of the current and historical states of all categories of feature data. It measures the degree of abnormality in the operation of the live streaming room. The larger the value, the more significant the overall deviation and mutation of all categories of feature data in the current multidimensional data profile, and the higher the comprehensive risk entropy.
[0037] For example, in the expression for comprehensive risk entropy and , It is recommended that the value be greater than or equal to 5. If only the short-time audio energy in the current multidimensional data profile is considered, its , , ,like , Suddenly increased to ,but This shows that when the operational risks of a live streaming room undergo drastic changes, via After amplification and exponential operations, a significant gain is generated, which increases rapidly, thus sensitively reflecting abnormal fluctuations in the content. Round to two decimal places.
[0038] It should be noted that if scheduling is based solely on the comprehensive risk entropy at the current moment, the system will still be reactive and exhibit lag. To achieve preventative scheduling, this invention needs to predict the risk entropy value at the next moment. Traditional moving average prediction algorithms assume that the data has a certain degree of stability, but when faced with the occasional pulse-like risk fluctuations in live streams, they often exhibit significant phase lag. That is, when the risk begins to increase, the prediction result is often half a beat slower, leading to untimely review and intervention. This invention introduces the concepts of momentum and acceleration from physics to correct the linear prediction. It not only considers the current rate of change, i.e., the first-order difference, but also focuses on the acceleration of change, i.e., the second-order difference. When the risk entropy prediction value shows an accelerating upward trend, an additional positive gain is applied to the risk entropy prediction value.
[0039] Preferably, the acceleration of risk change is obtained based on the volatility intensity of the comprehensive risk entropy, including: Obtain the comprehensive risk entropy at the current time, the previous time, and the two time points before the current time; calculate the difference between the comprehensive risk entropy at the current time and the comprehensive risk entropy at the previous time, and denote it as the risk change rate; calculate the difference between the comprehensive risk entropy at the previous time and the comprehensive risk entropy at the two time points before the current time, and denote it as the risk change rate at the previous time; calculate the difference between the risk change rate and the risk change rate at the previous time to obtain the risk change acceleration, which characterizes the speed of change in the intensity of the comprehensive risk entropy fluctuation.
[0040] Specifically, based on the continuity between the current status of the live streaming room and its historical predictions, and combined with the acceleration of risk changes, the predicted risk entropy value is calculated, including: The smoothing coefficient used to balance the current risk situation with the historical forecast situation, and the trend enhancement coefficient used to characterize the weight of the impact of risk change trends on the forecast results are obtained in advance from the database.
[0041] The predicted risk entropy value satisfies the following expression: ; In the formula, and This represents the predicted risk entropy values for the next and current time points; and The rate of risk change and the acceleration of risk change represent the overall risk entropy at the current moment; This represents the pre-obtained smoothing coefficient. ; This represents the pre-obtained trend enhancement coefficient; This represents the absolute value function.
[0042] In the formula, To balance the continuity between the current state of the live broadcast room and the historical predictions, and to avoid the distortion of prediction results caused by abnormal fluctuations in the actual situation at a single moment; Dynamic correction of risk trends is achieved by multiplying the rate of risk change, the acceleration of risk change, and the trend enhancement coefficient. The more drastic the change in risk conditions, the more significant the correction. The stronger the impact on the predicted value of risk entropy, This represents the risk change acceleration correction factor. Its core function is to dynamically adjust the risk level based on the acceleration of risk changes. The intensity is moderately amplified through a square root function, avoiding interference from extreme values of risk change acceleration. When the risk changes smoothly or uniformly, the acceleration... With the acceleration correction factor close to 1, the model degenerates into a conventional trend prediction when the risk accelerates. The magnitude is relatively large, resulting in an acceleration correction factor significantly greater than 1, which directly amplifies the acceleration. , making This resulted in a relatively sharp increase, ensuring that intervention measures from the large-scale model were in place before the risks of operating the live streaming room reached their peak. By combining the smoothed results of the current status of the live broadcast room operation with the historical prediction status, and superimposing the trend term of the product of the risk change rate and acceleration, the risk entropy of the next moment can be predicted. This prediction method not only satisfies the stability of the risk entropy prediction value, but also captures the sudden change trend of the risk status, which meets the needs of dynamic monitoring of the risk of live broadcast room operation.
[0043] For example, , ,like and Close, all values are In scenarios where the risks of operating a live streaming room are rapidly increasing, , , ,calculate , ;calculate , ;but However, if we do not consider the corrective effect of the acceleration of risk changes, then If it is 1, then It is evident that after introducing the risk change acceleration correction function, The value improved from 3.0 to 3.095, generating a positive gain, indicating that the system effectively predicted the upward trend of operational risks in the live streaming room, ensuring intervention measures were taken before the risks peaked. The above results... , All values are rounded to three decimal places.
[0044] It should be noted that, Figure 3 This is a comparison chart of risk entropy manifold calculation and trend prediction. The chart includes the real-time calculated risk curve representing the comprehensive risk entropy and the predicted risk curve representing the predicted risk entropy value. The real-time calculated risk curve and the predicted risk curve clearly show two independent peaks. At the moment when the first wave of risk situation breaks out at t=60, the predicted risk curve rises rapidly and provides an accurate warning. When the second wave of stronger risk situation arrives at t=70, the predicted risk curve rises again with a steeper slope, and the peak value is significantly higher than the first peak. This proves that the system has extremely high sensitivity to continuous and escalating risks. At the same time, in the non-risk period, the curve quickly approaches the 0 axis without chaotic fluctuations.
[0045] S3: Based on the predicted risk entropy value, the operation status of the live broadcast room is classified into low-risk and high-risk conditions; according to the difference between the predicted risk entropy value and the average comprehensive risk entropy, the minimum sampling interval is adjusted to obtain the model sampling interval when the live broadcast room is in a high-risk condition.
[0046] It should be noted that this invention uses the average comprehensive risk entropy as the core threshold and combines the magnitude of the predicted risk entropy value to classify the risk level of the live broadcast room's operation. In low-risk situations, the maximum sampling interval is used to save computing resources, while in high-risk situations, a non-linear adjustment mechanism for the sampling interval is activated. By dynamically adjusting the sampling frequency, accurate monitoring of high-risk live broadcast rooms is achieved, balancing the system's computing power consumption with the real-time requirements of live broadcast supervision.
[0047] Specifically, based on the predicted risk entropy value, the operational status of the live streaming room is classified into low-risk and high-risk statuses, including: Get the latest data up to the current time. The comprehensive risk entropy of each sampling point is denoted as the average comprehensive risk entropy. The maximum sampling interval and the minimum sampling interval are obtained from the database. If the predicted risk entropy value at the next moment is less than the average comprehensive risk entropy, the current live video under test is considered to be a low-risk video, and the live room operation status is low-risk. The maximum sampling interval is used for sampling until the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy. If the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy, the current live video under test is considered to be a high-risk video, and the live room operation status is high-risk.
[0048] It should be noted that, The value ranges from 100 to 200, and in this embodiment, the value is 150.
[0049] It should be noted that this invention designs a nonlinear sampling interval adjustment logic for live streaming rooms in high-risk situations. The core basis is the degree of deviation between the predicted risk entropy value and the average comprehensive risk entropy value. The deviation degree is made dimensionless by using a risk sensitivity threshold. The step response coefficient is used to control the adjustment intensity of the sampling interval as the risk changes, so that the sampling interval approaches the minimum sampling interval as the risk increases and approaches the maximum sampling interval as the risk decreases. This achieves a fine allocation of computing resources and a dynamic balance between real-time monitoring in live streaming room risk monitoring scenarios.
[0050] Specifically, based on the difference between the predicted risk entropy value and the mean comprehensive risk entropy, the minimum sampling interval is adjusted to obtain the model sampling interval for a live broadcast room in a high-risk state, including: The risk sensitivity threshold used to define the benchmark for the degree of risk deviation is obtained in advance from the database, as well as the step response coefficient used to characterize the response strength as the sampling interval changes with the degree of risk deviation.
[0051] The model sampling interval satisfies the following expression: ; In the formula, Indicates the model sampling interval; Indicates the minimum sampling interval; Indicates the maximum sampling interval; This represents the predicted risk entropy value for the next moment; This represents the average of the overall risk entropy; This indicates a pre-acquired risk sensitivity threshold; This represents the pre-obtained step response coefficient.
[0052] In the formula, This means dividing the difference between the predicted risk entropy value and the mean comprehensive risk entropy by the risk sensitivity threshold to achieve dimensionless processing of the model sampling interval, thus avoiding interference from differences in the magnitude of risk entropy values in different scenarios on the adjustment logic. The larger the value, the smaller the normalized deviation corresponding to the same difference, and the lower the sensitivity of the model sampling interval to changes in risk. The smaller the value, the higher the sensitivity. This indicates that the nonlinearity of the nonlinear transformation is determined by the step response coefficient, when... Initially, the denominator grows slowly, and the model sampling interval decreases gradually; after the risk deviation exceeds the threshold, the denominator grows rapidly, and the model sampling interval decreases sharply, approaching... To achieve rapid and intensive monitoring during sudden risk changes; when At that time, the denominator grows rapidly, the sampling interval decreases rapidly, and after the risk deviation increases, the growth of the denominator slows down, and the model sampling interval tends to stabilize. This represents the dynamic variation of the model sampling interval; its magnitude determines the relationship between the model sampling interval and... The gap; When the predicted risk entropy value for the next moment is higher than or equal to the average comprehensive risk entropy, it indicates that the live broadcast room is in a medium-to-high risk state, and the sampling interval needs to be adjusted nonlinearly according to the degree of risk deviation.
[0053] For example, ; , Second, Second, , , Seconds, when , Seconds, it can be seen that as the predicted risk entropy exceeds the average comprehensive risk entropy, the model sampling interval will rapidly decrease from... Seconds down to This ensures the system's ability to accurately detect high risks. The above results... , All values are rounded to two decimal places.
[0054] It should be noted that, Figure 4 The graph shows the results of dynamic computing power scheduling based on risk prediction. It includes a fixed sampling curve and a dynamic sampling interval curve. During the non-risk periods t=0~60 and t=90~200, the dynamic sampling interval curve is a horizontal straight line, stable at the bottom 10-second mark. This intuitively demonstrates the computing power savings of this invention during risk-free periods and effectively eliminates high-amplitude fluctuations caused by misjudgments. Corresponding to the double peaks of risk periods, the dynamic sampling interval curve exhibits… In the sinking action, the system first rapidly compresses the sampling interval to 0.5 seconds at t=60; during the brief risk decline period in the middle, the sampling interval attempts to adjust back, but as the second wave of stronger risk arrives, it drops back to the ultra-fast mode of 0.5 seconds. Compared with the fixed sampling curve and the fixed 5-second sampling of the existing technology, this solution achieves intelligent scheduling that is more economical in normal times and faster in times of danger.
[0055] S4: Count down the time based on the model sampling interval, capture the audiovisual content data when the countdown ends, and call the large model to perform security inference on the live broadcast.
[0056] It should be noted that the dynamic adjustment logic of the model sampling interval is deeply integrated with the live streaming security inference process. In low-risk conditions, the maximum sampling interval is used to reduce computing power consumption, while in medium- and high-risk conditions, the sampling interval is shortened through nonlinear adjustment to increase the monitoring frequency. This achieves fine-grained scheduling of computing resources, ensuring basic coverage of live streaming rooms in low-risk conditions while focusing on real-time control of live streaming rooms in high-risk conditions, thus providing relatively efficient data support for the large model to accurately infer the security of live streaming content.
[0057] Specifically, a countdown is performed based on the model's sampling interval. At the end of the countdown, audiovisual content data is extracted, and a large model is invoked to perform security inference on the live stream, including: The system counts down based on the model sampling interval. When the countdown ends, it immediately captures the current video keyframe or audio slice and calls the cloud-based big model API to perform content security inference. If the big model determines that the content is safe, the system continues to run with the current logic. If the big model determines that the content is in violation, the system immediately triggers the blocking mechanism. If the big model determines that the content is suspected or gives a high level of risk confidence, the system uses this feedback signal as an instruction to maintain a high-frequency review status for subsequent monitoring of the live stream until the big model determines that the content is safe multiple times in a row.
[0058] This completes the real-time review of online audiovisual content.
[0059] This invention also discloses a large-scale model-based online audiovisual content review system, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a large-scale model-based online audiovisual content review method according to this invention.
[0060] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0061] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.
Claims
1. A method for reviewing online audiovisual content based on a large model, characterized in that, include: The system acquires a set of video information of the live video to be tested within a preset time period. Based on the phonetic features and interaction popularity of the information in the video information set, it extracts and analyzes the data to obtain a multidimensional data set of the live video to be tested. The multidimensional data set is then normalized to form a current multidimensional data profile that represents the current operating status of the live room. By combining the current multidimensional data profile analysis of the chaotic state and volatility of the live broadcast room's operation, a comprehensive risk entropy is calculated; Based on the volatility intensity of the comprehensive risk entropy, the acceleration of risk change is obtained; Based on the continuity between the current status of the live broadcast room and the historical prediction status, and combined with the acceleration of risk change, the predicted value of risk entropy is calculated. Based on the predicted risk entropy value, the operation status of the live broadcast room is classified into low-risk and high-risk status. Based on the difference between the predicted risk entropy and the mean comprehensive risk entropy, the minimum sampling interval is adjusted to obtain the model sampling interval when the live broadcast room is in a high-risk state. A countdown is performed based on the model's sampling interval. At the end of the countdown, audiovisual content data is captured and the large model is invoked to perform security inference on the live stream.
2. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The multidimensional data set of the live video to be tested includes: A lightweight BERT model was used to statistically analyze the sensitive word hit density sequence in the text data of the live video under test. Simultaneously, the text data sequence was input into a pre-trained sentiment analysis model to obtain a sequence of the degree of radicalism in speech. The energy value of each frame of the audio buffer stream was calculated using a short-time energy analysis algorithm, denoted as the audio short-time energy sequence. The number of times the temporal waveform of a single frame of the audio buffer stream crossed the zero horizontal axis was counted, denoted as the speech zero-crossing rate sequence. The number of newly added bullet comments and the number of new user IDs entering the live room within the analysis window were statistically analyzed and denoted as the bullet comment addition rate sequence and the user entry rate sequence, respectively. The obtained multi-dimensional data sequences were time-aligned and denoted as the multi-dimensional data set of the live video under test.
3. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The formation of the current multidimensional data profile representing the current operating status of the live broadcast room includes: Starting from the current moment, the preset length is... Within a time window, the maximum-minimum normalization algorithm is applied to each sequence in the multidimensional data set within the time window to map the data corresponding to the current moment to the standard interval [0,1]. The normalized results are then spliced together in a preset order to form a current multidimensional data profile representing the current operating status of the live broadcast room. The current multidimensional data profile includes feature data from multiple dimensions, such as the density of sensitive words, the degree of radicalism of speech, the short-term energy of audio, the zero-crossing rate of voice, the rate of new bullet comments, and the rate of users entering the room.
4. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The comprehensive risk entropy satisfies the following expression: ; In the formula, This represents the overall risk entropy at the current moment; This indicates the number of feature data categories participating in the calculation in the current multidimensional data profile; , Feature data representing the current time and the previous time. ; Representing feature data In length of The average value within the time window; This represents the pre-acquired feature data. Weighting coefficients; This represents the pre-obtained mutation sensitivity coefficient; Represents a logarithmic function; Represents a very small positive number, ensuring that the denominator is not zero; Represents the maximum value function; This represents an exponential function with the natural constant as its base.
5. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The obtained risk change acceleration includes: Obtain the comprehensive risk entropy at the current time, the previous time, and the two time points before the current time; calculate the difference between the comprehensive risk entropy at the current time and the comprehensive risk entropy at the previous time, and denote it as the risk change rate; calculate the difference between the comprehensive risk entropy at the previous time and the comprehensive risk entropy at the two time points before the current time, and denote it as the risk change rate at the previous time; calculate the difference between the risk change rate and the risk change rate at the previous time to obtain the risk change acceleration, which characterizes the speed of change in the intensity of the comprehensive risk entropy fluctuation.
6. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The predicted risk entropy value satisfies the following expression: ; In the formula, and This represents the predicted risk entropy values for the next and current time points; and The rate of risk change and the acceleration of risk change represent the overall risk entropy at the current moment; This represents the pre-obtained smoothing coefficient. ; This represents the pre-obtained trend enhancement coefficient; This represents the absolute value function.
7. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The classification of live streaming room operation status into low-risk and high-risk status includes: Get the latest data up to the current time. The comprehensive risk entropy of each sampling point is denoted as the average comprehensive risk entropy. The maximum sampling interval and the minimum sampling interval are obtained from the database. If the predicted risk entropy value at the next moment is less than the average comprehensive risk entropy, the current live video under test is considered to be a low-risk video, and the live room operation status is low-risk. The maximum sampling interval is used for sampling until the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy. If the predicted risk entropy value at the next moment is greater than or equal to the average comprehensive risk entropy, the current live video under test is considered to be a high-risk video, and the live room operation status is high-risk.
8. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The model sampling interval satisfies the following expression: ; In the formula, Indicates the model sampling interval; Indicates the minimum sampling interval; Indicates the maximum sampling interval; This represents the predicted risk entropy value for the next moment; This represents the average of the overall risk entropy; This indicates a pre-acquired risk sensitivity threshold; This represents the pre-obtained step response coefficient.
9. The method for reviewing online audiovisual content based on a large model according to claim 1, characterized in that, The step of capturing audiovisual content data at the end of the countdown and calling a large model to perform security inference on the live stream includes: The system counts down based on the model sampling interval. When the countdown ends, it immediately captures the current video keyframe or audio slice and calls the cloud-based big model API to perform content security inference. If the big model determines that the content is safe, the system continues to run with the current logic. If the big model determines that the content is in violation, the system immediately triggers the blocking mechanism. If the big model determines that the content is suspected or gives a high level of risk confidence, the system uses this feedback signal as an instruction to maintain a high-frequency review status for subsequent monitoring of the live stream until the big model determines that the content is safe multiple times in a row.
10. A network audiovisual content review system based on a large model, characterized in that, include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement a network audiovisual content review method based on a large model according to any one of claims 1-9.