Method and system for monitoring internet live streaming for image and voice recognition violations
By combining image and voice recognition technologies, this technology acquires and processes internet live streaming data, identifies violations, and optimizes monitoring parameters based on historical data. This solves the problem of accurately identifying violations in existing technologies, achieving efficient violation monitoring and platform order maintenance.
Patent Information
- Application Number
- CN202510449135.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing technologies cannot simultaneously utilize image and voice recognition technologies to accurately acquire and process internet live streaming data to identify violations and optimize monitoring parameters based on historical violation data.
By combining image and speech recognition, live stream data is acquired and its integrity is verified. Specific data processing modes and models are used to identify violations. Data collection parameters are corrected based on historical violation data to dynamically optimize monitoring.
It enables accurate identification and timely alerts for violations in internet live streaming, safeguarding the order of live streaming platforms and user experience, and optimizing the targeting and adaptability of the monitoring system.
Smart Images

Figure CN120302088B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of live broadcast monitoring, in particular to a monitoring method and system for internet live broadcast irregular behavior based on image and voice recognition. BACKGROUND
[0002] With the development of internet technology, live broadcast in various scenarios has developed rapidly. However, there are various irregularities in the live broadcast process of current live broadcast platforms, such as smoking behavior in live broadcast scenarios, i.e. anchors smoking in live broadcast scenarios. Smoking is not only harmful to health, but also brings bad influence, which is not conducive to the healthy growth of young people. The Chinese utility model patent with publication number CN113705370A discloses a detection method and device for live broadcast room irregular behavior, electronic equipment and storage medium, which relates to the technical field of internet mobile terminal application. After frame extraction is performed on the to-be-detected video to obtain a plurality of video frames, the plurality of video frames are sequentially input into an irregular behavior detection model to determine whether there is irregular behavior in the video frames and to mark the irregular behavior as abnormal. If it is determined according to the abnormal marking of the first video frame that the first video frame has a first irregular behavior, it is determined according to the abnormal marking of the second video frame whether the second video frame has a second irregular behavior, wherein the second video frame is the previous video frame of the first video frame among the plurality of video frames. Finally, after it is determined that the second video frame has the second irregular behavior, it is determined according to the first irregular behavior and the second irregular behavior that the to-be-detected video has live broadcast room irregular behavior. The detection method ensures the accuracy of live broadcast room irregular behavior detection through a multi-frame fusion mechanism.
[0003] However, the prior art has the problem that it is difficult to simultaneously utilize image and voice recognition technology, accurately acquire and process live broadcast data to identify internet live broadcast irregular behavior, and optimize monitoring parameters according to historical irregular data. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a monitoring method and system for internet live broadcast irregular behavior based on image and voice recognition, which solves the problem that the prior art is difficult to simultaneously utilize image and voice recognition technology, accurately acquire and process live broadcast data to identify internet live broadcast irregular behavior, and optimize monitoring parameters according to historical irregular data.
[0005] To achieve the above object, the present application is implemented by the following technical solutions: a kind of image and speech recognition internet live broadcast illegal behavior monitoring method, comprising the following steps: based on the data acquisition parameter determined from live broadcast platform the live stream data of each host is obtained and the data integrity is verified, and live stream data includes original video stream data and original audio stream data;By determining the data processing mode, live stream data is processed to obtain preprocessed live stream data, including video frame set and speech text data;Image illegal processing model is used to identify the features of video frame set, and image illegal mark information is output, including the illegal type of each image frame, illegal probability and image illegal timestamp;Text processing model is used to process text of speech text data, and text illegal mark information is output, including illegal text paragraph and text illegal timestamp;Based on image illegal mark information and text illegal mark information, it is determined whether the host is illegal and alarmed when illegal;Based on the historical illegal data of host, data acquisition parameter is corrected.
[0006] Further, the data integrity of the live stream data of each host is verified, comprising the following steps: obtaining the data integrity coefficient I comp And data format correctness index I form : Wherein, D act Actual returned data amount, D exp Data amount expected to be obtained, N corr The number of fields in the data with correct format, N total Total field number in data;If data integrity coefficient is greater than the data integrity threshold value determined for the corresponding host and data format correctness index is greater than the data format correctness threshold value determined for the corresponding host, then the data integrity of the host is normal, otherwise, the data integrity is abnormal, and the live stream data of the host is reacquired;Wherein, data acquisition parameters include API application frequency and data acquisition request interval time of web crawler program;Data acquisition parameters, data integrity threshold value and data format correctness threshold value, data processing mode, image illegal processing model and text processing model are determined by obtaining the corresponding influence factor of each host.
[0007] Further, the determination mode of data acquisition parameter, data integrity threshold value and data format correctness threshold value is: obtaining the set influence factor-data processing parameter mapping set stored in database, comparing each set influence factor with influence factor one by one, calculating similarity based on Euclidean distance, determining the set influence factor corresponding to the minimum similarity, and obtaining the corresponding data processing parameters from the database, including data acquisition parameters, data integrity threshold value and data format correctness threshold value;
[0008] The data processing mode includes a video stream processing mode and an audio stream processing mode, wherein the video stream processing mode includes an adaptive histogram equalization processing mode and a Gaussian filter denoising processing mode, and the audio stream processing mode includes a first-level text processing mode and a second-level text processing mode; the image violation processing model includes a Mask R-CNN model and a Faster R-CNN model, and the text processing model includes a Transformer architecture text processing model and a keyword matching processing model; and the determination manners of the data processing mode, the image violation processing model and the text processing model are as follows:
[0009] The influence factor corresponding to each anchor is obtained, and the influence factor is compared with an influence threshold value; if the influence factor is greater than the influence threshold value, the video stream processing mode is the adaptive histogram equalization processing mode, the audio stream processing mode is the first-level text processing mode, the image violation processing model is the Mask R-CNN model, and the text processing model is the Transformer architecture text processing model; if the influence factor is not greater than the influence threshold value, the video stream processing mode is the Gaussian filter denoising processing mode, the audio stream processing mode is the second-level text processing mode, the image violation processing model is the Faster R-CNN model, and the text processing model is the keyword matching processing model.
[0010] Further, the influence factor is obtained as follows: live broadcast platform influence data and anchor influence data corresponding to each anchor are obtained, a live broadcast platform influence coefficient λ1 and anchor influence coefficients λi of each anchor are determined i is the number of the anchor; the live broadcast platform influence coefficient λ1 and the anchor influence coefficients λi of each anchor are used to determine the initial influence factor The initial influence factor is obtained A correction factor is determined based on real-time live broadcast data of each anchor The initial influence factor is corrected in real time to obtain the influence factor e is a natural constant.
[0011] Further, the live broadcast platform influence coefficient λ1 is obtained as follows:
[0012] λ1 = λ Sc + λ PF + λ B + λ TH + λ QL ;
[0013] wherein λ Sc is a market share coefficient, which is obtained based on a logistic regression model, λ PF is a user reputation coefficient, which is obtained based on a calculation model of a score mean and a score standard deviation, and λ B is a media exposure coefficient, determined based on a chromatographic analysis method, λ TH is a technology input-output ratio coefficient, determined based on technology innovation funds and achievement transformation effects, λ QL is a user growth potential coefficient, determined based on historical user growth rates;
[0014] anchor influence coefficient The acquisition method is as follows:
[0015]
[0016] wherein, is an anchor fan size factor, is an anchor live broadcast activity factor, is an anchor content quality factor, X is the total number of anchors in the live broadcast platform.
[0017] Further, the correction factor is obtained as follows: real-time data in the anchor live broadcast process is obtained, including the growth rate of the number of viewers per unit time xZ in the live broadcast process, the keyword density of the audience per unit time ρ gc and the total number of online people P zx in the current unit time in the live broadcast process, and the correction factor is obtained based on a correction factor calculation model wherein, is the average value of the number of online people in the historical live broadcast process of the anchor.
[0018] Further, based on image violation marking information and text violation marking information, it is determined whether the anchor violates the rules and alarms when violating the rules, including the following steps: if the image violation timestamp and the text violation timestamp are consistent, the violation probability is greater than the set violation threshold, and the violation text paragraph contains a violation keyword, it is determined that the anchor violates the rules, and alarm information is sent to the live broadcast platform operator; if the image violation timestamp and the text violation timestamp are consistent, the violation probability is greater than the set violation threshold, or the violation text paragraph contains a violation keyword, it is determined that the anchor violates the rules, and alarm information is sent to the live broadcast platform operator.
[0019] Further, based on the historical violation data of the anchor, the data collection parameters are corrected, including the following steps: based on the historical violation data of each anchor, a platform-wide risk baseline matrix is constructed, and the matrix elements are calculated, wherein the row dimension of the risk baseline matrix represents the violation type violation risk baseline, and the column dimension represents the time window; the real-time violation feature vector of the live broadcast platform in the monitoring time window is obtained, and the dynamic violation risk coefficient is calculated by combining the platform risk baseline matrix and the anchor influence coefficient of each anchor; the platform violation risk level is determined based on the risk level mapping relationship, and the corrected data collection parameters stored in the database are matched based on the platform violation risk level.
[0020] Further, the calculation method of the matrix elements is as follows:
[0021]
[0022] wherein w i′j′ is a matrix element in the full-platform risk baseline matrix, δ(C i′,j′ ) is the number of violations of the i'th type of violation at the time window j', is the maximum number of violations of the i'th type of violation, i' is the number of the type of violation, j' is the number of the time window, ξ is the time decay factor, t now is the current time, t wg is the time of the most recent violation, and J is the total number of types of violation.
[0023] The method for obtaining the dynamic violation risk coefficient is as follows:
[0024]
[0025] wherein λ risk is the dynamic violation risk coefficient, F eral is the real-time violation feature vector, is the transpose matrix of the full-platform risk baseline matrix M base , and is the average of the influence factors.
[0026] The dynamic violation risk coefficient is compared with each risk threshold interval stored in the database to determine the risk threshold interval to which the dynamic violation risk coefficient belongs, and then the platform violation risk level is determined according to the mapping relationship between the risk threshold interval and the risk level. Subsequently, the corresponding stored modified data acquisition parameters are obtained from the database according to the platform violation risk level.
[0027] The application discloses an image and voice identification internet live broadcast violation behavior monitoring system and method.
[0028] The application has the following advantages:
[0029] The image and voice identification internet live broadcast violation behavior monitoring method and system can obtain live stream data by determining data collection parameters and verify the completeness of the data, ensure the availability of the data, obtain video frame sets and voice text data by a specific data processing mode, provide a basis for subsequent analysis, identify violation information by using image and text processing models, comprehensively capture violation clues, determine whether a host is in violation and alarm according to the information, timely stop the violation behavior, and correct the data collection parameters according to historical violation data of the host, realize dynamic optimization monitoring, effectively maintain the order of a live broadcast platform and guarantee a good experience of users.
[0030] Of course, any product implementing the present application does not necessarily achieve all the advantages mentioned above. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 The application discloses an image and voice identification internet live broadcast violation behavior monitoring method flow chart.
[0032] Figure 2 The application discloses an image and voice identification internet live broadcast violation behavior monitoring system flow chart. DETAILED DESCRIPTION
[0033] Please refer to Figure 1The embodiment of the application provides a technical scheme: a monitoring method for live broadcast rule violation of image and voice recognition, comprising the following steps: obtaining live broadcast stream data of each host from a live broadcast platform based on determined data collection parameters and verifying data integrity, wherein the live broadcast stream data comprises original video stream data and original audio stream data;
[0034] obtaining a data integrity coefficient I of each host comp and a data format correctness index I form :
[0035]
[0036] wherein D act is an actual returned data amount, D exp is a data amount expected to be obtained, N corr is a number of fields with correct format in the data, and N total is a total number of fields in the data; the two indexes evaluate data quality in a quantitative manner from two key dimensions of data amount and data format, thereby providing accurate basis for subsequent judgment on whether the data meet monitoring requirements, and making data quality evaluation more scientific and objective.
[0037] If the data integrity coefficient is greater than a determined data integrity threshold value of the corresponding host and the data format correctness index is greater than a determined data format correctness threshold value of the corresponding host, the data integrity of the host is normal, otherwise, the data integrity is abnormal, and the live broadcast stream data of the host is reobtained; the live broadcast stream data used for subsequent rule violation analysis can be effectively ensured to be complete and correct in format, and the live broadcast stream data used for subsequent rule violation analysis can be effectively ensured to be complete and correct in format, so that false negative or false positive of rule violation is avoided, and the accuracy of the monitoring result is ensured.
[0038] wherein the data collection parameters comprise API application frequency and data acquisition request interval time of a web crawler program; the data collection parameters, the data integrity threshold value and the data format correctness threshold value, the data processing mode, the image rule violation processing model and the text processing model are determined by obtaining corresponding influence factors of each host. The data collection parameters (such as API application frequency and data acquisition request interval time of a web crawler program), the data integrity threshold value, the data format correctness threshold value, the data processing mode, the image rule violation processing model and the text processing model are determined by obtaining corresponding influence factors of each host. This enables the monitoring method to dynamically adjust related parameters and modes of data collection and processing according to specific conditions of different live broadcast platforms and hosts, improves the pertinence and adaptability of the monitoring system, and further effectively identifies and processes various live broadcast rule violations.
[0039] The determination mode of the data collection parameters, the data integrity threshold value and the data format correctness threshold value is as follows:
[0040] Obtain the set of setting influence factor-data processing parameter mappings stored in the database, compare each influence factor with each setting influence factor one by one, calculate the similarity based on the Euclidean distance, determine the setting influence factor corresponding to the minimum similarity, and obtain the corresponding data processing parameter from the database, including the data acquisition parameter, the data integrity threshold and the data format correctness threshold; This way makes the data acquisition link can be accurately set according to the comprehensive situation of the live broadcast platform and the host. For example, for hosts with different influence, the API application frequency and the data acquisition request interval time of the web crawler program are reasonably adjusted, which not only ensures the efficiency of data acquisition, but also avoids causing too much burden to the live broadcast platform, while ensuring that the quality of the acquired data meets the monitoring requirements.
[0041] The data processing mode includes a video stream processing mode and an audio stream processing mode, wherein the video stream processing mode includes an adaptive histogram equalization processing mode and a Gaussian filter denoising processing mode, and the audio stream processing mode includes a first-level text processing mode and a second-level text processing mode; the image violation processing model includes a Mask R-CNN model and a Faster R-CNN model, and the text processing model includes a Transformer architecture text processing model and a keyword matching processing model; the determination mode of the data processing mode, the image violation processing model and the text processing model is:
[0042] Compare the influence factor corresponding to each host with the influence threshold;
[0043] If the influence factor is greater than the influence threshold, the video stream processing mode is the adaptive histogram equalization processing mode, the audio stream processing mode is the first-level text processing mode, the image violation processing model is the Mask R-CNN model, and the text processing model is the Transformer architecture text processing model;
[0044] The first level text processing mode: first format conversion, converted into a standard format supported by the speech recognition engine, such as the common PCM (Pulse-Code Modulation) format. In the preprocessing stage, noise reduction algorithm is used to remove background noise in the audio, and more powerful and more accurate speech recognition engine is selected, such as the enhanced version of Baidu speech recognition engine based on deep learning. The engine has a more complex neural network architecture and can better handle high complexity audio data. To further improve the recognition accuracy, the amount of language model training data is increased. New training data is selected from a large amount of text data related to live content, including live scripts, relevant professional literature, and discussions on live topics on social media, to enrich the language model's learning ability for language expression in live scenarios. The preprocessed audio signal is input into the speech recognition engine with adjusted parameters for recognition. The engine outputs the recognition result text, which is the speech-to-text content corresponding to the high-impact audio stream, used for subsequent text keyword matching and semantic analysis.
[0045] The Mask R-CNN model is developed based on FasterR-CNN and can realize target detection and instance segmentation simultaneously. When initializing the model, pre-trained weight parameters are loaded, and a set of pre-trained weight parameters is set. At the same time, according to the characteristics of live images, such as size, color channel number, etc., the input layer parameters of the model are set. Assuming that the live image size is HxWxC, where H is the image height (pixels), W is the image width (pixels), and C is the color channel number (such as RGB image C=3), the input layer of the model is adjusted to adapt to the structure of this size. The image frames of the high-impact platform and the anchor after image frame extraction and preprocessing (such as adaptive histogram equalization processing) are input into the Mask R-CNN model. Before input, the image frames are normalized to map the image pixel value from the range of 0-255 to the range of 0-1, which helps to speed up the model training and improve the stability of the model.
[0046] The Mask R-CNN model performs forward propagation operation on the input normalized image frame. The model first extracts deep features of the image through a backbone network (such as ResNet or Feature Pyramid Network, FPN). During feature extraction, different layers of convolution operations are performed by convolution kernels and feature maps. Next, the Region Proposal Network (RPN) generates candidate regions that may contain targets. Then, the Region of Interest (RoI) Align operation is performed on the feature map for each candidate region to obtain a fixed-size feature vector, which is used for subsequent classification and segmentation prediction. The Mask R-CNN model performs classification and segmentation prediction on the feature vector of each candidate region to determine whether it is a violation image feature. At the same time, for the detected violation region, instance segmentation is performed to obtain a segmentation mask, marking the accurate outline of the violation target. For the detected violation image frame, the corresponding timestamp is recorded to store the violation image frame marking information.
[0047] The speech-to-text content of the high-influence category is preprocessed. The text is converted to lowercase, and irrelevant information such as punctuation marks, special characters, etc. is removed. For example, to remove punctuation marks, a set of punctuation marks can be defined, and a string replacement operation can be used. For each character in the text, if it belongs to the set of punctuation marks, it is replaced with a null character to obtain the preprocessed text. A semantic analysis model based on the Transformer architecture is selected, such as the BERT (Bidirectional Encoder Representations from Transformers) model. The pre-trained BERT model parameters are loaded, which are trained on a large-scale text corpus and contain rich language semantic knowledge. According to the characteristics of the live content, the input parameters of the model are configured, such as truncating or padding the text to a fixed length (e.g., 128 words). Based on the semantic representation output by the model, a full connection layer and a classifier are used for violation semantic mining. The output probability of the output classifier is output. If the output result is greater than the set text violation threshold, it is determined that the text contains violation semantics, and the corresponding text paragraph and timestamp are marked.
[0048] If the influence factor is not greater than the influence threshold, the video stream processing mode is the Gaussian filter denoising processing mode, the audio stream processing mode is the second-level text processing mode, the image violation processing model is the Faster R-CNN model, and the text processing model is the keyword matching processing model.
[0049] The second level text processing mode: also converts the format to PCM format, and uses a simple noise reduction algorithm, such as a mean filter-based noise reduction method. A basic configuration of a speech recognition engine is used, such as an open source Kaldi speech recognition engine basic version. The engine has lower computing resource requirements and is suitable for processing audio streams with relatively low requirements for recognition accuracy. The preprocessed audio signal is input into the basic configuration of the speech recognition engine for recognition. The engine outputs the recognition result text, which is the voice-to-text content corresponding to other live streams, for subsequent text keyword matching and semantic analysis.
[0050] If the Faster R-CNN model is selected, the initialization process is similar to Mask R-CNN, and the pre-trained weight parameters are loaded, and the input layer parameters are set according to the live image size. The low-impact live stream image frame after image frame extraction and preprocessing (such as Gaussian filtering) is input into the selected model. The image frame is also normalized to obtain a normalized image frame, and the normalization formula is consistent with the high-impact image frame processing. The Faster R-CNN model operates on the input normalized image frame. The deep features of the image are extracted through the backbone network, and the feature map region proposal network (RPN) generates a candidate region set on the feature map. Then, the candidate regions are subjected to RoIPooling operation (similar to RoIAlign, but slightly different in calculation method), to obtain fixed-size feature vectors for subsequent classification and bounding box regression. The model classifies and predicts the feature vectors of the candidate regions, records the timestamps of the image frames detected with violation features, and stores the violation image frame label information for subsequent comprehensive judgment with the voice-to-text content.
[0051] When the text processing model is a keyword matching processing model, the preprocessed other type voice-to-text content is subjected to keyword matching. A string matching algorithm (such as a naive string matching algorithm or a more efficient KMP algorithm) is used to find keywords in the keyword library in the text. The keyword library is a pre-defined keyword related to violation content, such as "violence" and "pornography". The keyword library can be continuously expanded and optimized according to the type of live content and common violation scenarios.
[0052] The live streaming data is processed through a determined data processing mode to obtain preprocessed live streaming data including a video frame set and voice-to-text data; and according to a comparison result of the influence factor and the influence threshold, a video stream processing mode, an audio stream processing mode, an image violation processing model and a text processing model are selected for different anchors. For example, a high-influence anchor uses more complex and accurate processing modes and models, such as an adaptive histogram equalization processing mode to improve image quality to assist image violation identification, and a Transformer architecture text processing model to mine complex semantic violations; a low-influence anchor uses relatively simple and efficient modes and models. Such configuration can make full use of resources, improve the efficiency and accuracy of violation identification, and reduce system operation cost while ensuring monitoring effect.
[0053] The influence factor is obtained as follows: live platform influence data and anchor influence data corresponding to each anchor are obtained, live platform influence coefficient λ1 and anchor influence coefficient λi of each anchor are determined i is the number of the anchor; based on the live platform influence coefficient λ1 and the anchor influence coefficient λi of each anchor An initial influence factor is obtained A correction factor is determined based on real-time live data of each anchor The initial influence factor is corrected in real time to obtain the influence factor
[0054] e is a natural constant.
[0055] The live platform influence coefficient λ1 is obtained as follows: λ1 = λ Sc + λ PF + λ B + λ TH + λ QL ; wherein λ Sc is a market share coefficient, which is obtained based on a logistic regression model, λ PF is a user reputation coefficient, which is obtained based on a calculation model of score mean and score standard deviation, λ B is a media exposure coefficient, which is determined based on a chromatography analysis method, λ TH is a technology input-output ratio coefficient, which is determined based on technology innovation funds and achievement transformation effect, and λ QL is a user growth potential coefficient, which is determined based on historical user growth rate.
[0056] wherein U p is the total number of registered users of the live platform, U t is the total number of users in the live industry, M is the user proportion, and a and b are parameters obtained by training of the logistic regression model.
[0057] λ PF = cR - dσ, c is a weight factor of the score mean R, and d is a weight factor of the score standard deviation σ.
[0058] E comp is a comprehensive media exposure index, E max is the maximum value in the comprehensive media exposure index of all live streaming platforms, Nm is the number of reports by mainstream media (such as television stations, well-known newspapers, and large news websites), Nn is the number of reports by industry professional media, Ns is the number of hot topics on social media (such as Weibo and Douyin), w m is a weight factor of Nm, w n is a weight factor of Nn, w s is a weight factor of Ns.
[0059] λ TH = k*T + m*A, k is an adjustment factor of the technology R&D expense ratio, T is the technology R&D expense ratio, I tech is the technology R&D expense, R rev is the total operating income in the same period, m is a weight factor of the technology input result conversion coefficient A, which is the weighted sum of the number of user growth after the application of new technology and the improvement of user activity.
[0060] U p,new is the predicted number of new users in the prediction time period T new , and τ is the average user growth rate in the historical period.
[0061] The calculation of the live streaming platform influence coefficient comprehensively considers multiple dimensions such as the market share coefficient, user reputation coefficient, media exposure coefficient, technology input-output ratio coefficient, and user growth potential coefficient. The market share coefficient is obtained based on the logistic regression model, the user reputation coefficient is calculated based on the score mean and standard deviation, the media exposure coefficient is determined using the chromatography analysis method, the technology input-output ratio coefficient is obtained by combining the technology innovation fund and the result conversion effect, and the user growth potential coefficient is determined according to the historical user growth rate. This multi-dimensional calculation method can comprehensively and accurately evaluate the position of live streaming platforms in the market, user evaluation, industry influence, and development potential, providing reliable platform-level data support for subsequent determination of influence factors.
[0062] Anchor influence coefficient The acquisition method is as follows: wherein, is the anchor fan size factor, is the anchor live streaming activity factor, is the anchor content quality factor, and X is the total number of anchors in the live streaming platform.
[0063] U f is the number of fans of the host, U f,avg is the average value of the number of fans of all hosts on the platform.
[0064] D i,g is the average number of online viewers of the gth live broadcast of the host, H i,g is the duration of the gth live broadcast of the host, G is the number of live broadcasts of the host in the statistical period, D iz is the average number of online viewers of the zth live broadcast of the ith host, H iz is the duration of the zth live broadcast of the ith host.
[0065] V i,g is the content quality score of the gth live broadcast video of the host, scored by a professional review team from multiple dimensions such as content innovation, professionalism, and entertainment, R i,g is the audience feedback score of the gth live broadcast video of the host, obtained by weighted sum based on the number of likes, comments, shares, etc., R i,g,min and R i,g,max are the minimum and maximum values of the audience feedback scores of all live broadcast videos of all hosts in the statistical period.
[0066] The host influence coefficient is calculated by the host fan scale factor, the host live broadcast activity factor, and the host content quality factor, and is standardized in combination with the total number of hosts on the live broadcast platform. Among them, the fan scale factor reflects the audience base of the host, the live broadcast activity factor reflects the live broadcast frequency and the degree of interaction with the audience, and the content quality factor evaluates the quality of the host's live broadcast content from professional review and audience feedback. This calculation method can accurately measure the influence of each host in the platform, and help the monitoring system to allocate monitoring resources more reasonably and choose appropriate monitoring strategies according to the actual influence of the host, improve the accuracy and effectiveness of monitoring different host violation behaviors.
[0067] The correction factor is obtained as follows: real-time data during the host's live broadcast is obtained, including the growth rate of the number of viewers per unit time xZ during the live broadcast, the keyword density of the audience per unit time ρ gc and the total number of online people P zx in the current unit time during the live broadcast, and the correction factor is obtained based on the correction factor calculation model wherein, is the average value of the number of online people in the host's historical live broadcast.
[0068] The real-time data in the anchor live broadcast process is obtained, such as the viewing population growth rate per unit time, the audience keyword density, and the total number of online people per unit time, and a correction factor is calculated in combination with the average value of the online population in the historical live broadcast process of the anchor. This method can reflect the heat change, audience feedback, and content orientation in real time when the anchor is live broadcasting. When the viewing population growth rate is high, the audience keyword density is abnormal, or the total number of online people fluctuates greatly, the correction factor will be adjusted accordingly, and the influence factor will be corrected in real time. This enables the monitoring system to adjust the data acquisition parameters, data processing mode, and violation judgment standard in real time according to the real-time dynamics of the anchor live broadcast, enhances the timeliness and accuracy of monitoring, and more effectively identifies and responds to violations during live broadcast.
[0069] The determined image violation processing model is used for feature recognition on the video frame set, and image violation mark information is output, including the violation type, violation probability, and image violation timestamp of each image frame;
[0070] The determined text processing model is used for text processing on the voice-to-text data, and text violation mark information is output, including the violation text paragraph and the text violation timestamp;
[0071] Based on the image violation mark information and the text violation mark information, it is determined whether the anchor violates the rules and alarms when violating the rules. If the image violation timestamp and the text violation timestamp are consistent, the violation probability is greater than the set violation threshold, and the violation text paragraph contains a violation keyword, it is determined that a violation has occurred, and an alarm information is sent to the live broadcast platform operator. If the image violation timestamp and the text violation timestamp are inconsistent, the violation probability is greater than the set violation threshold, or the violation text paragraph contains a violation keyword, it is determined that a violation has occurred, and an alarm information is sent to the live broadcast platform operator. Two violation determination situations are defined, and by comprehensively considering the image violation timestamp, the text violation timestamp, the violation probability, and the keywords in the violation text paragraph, the violation behavior can be identified comprehensively and accurately. This multi-dimensional determination method reduces the possibility of misjudgment and omission, and ensures that the anchor can be discovered in time as long as there is a violation sign in the image or voice. Once a violation is determined, an alarm information is immediately sent to the live broadcast platform operator, so that the operator can take timely measures, such as interrupting the live broadcast, warning or punishing the anchor, etc., effectively curbing the continuous spread of violation behavior, maintaining the good order and healthy environment of the live broadcast platform, and protecting the viewing experience of the audience.
[0072] The data acquisition parameters are corrected based on the historical violation data of the anchor.
[0073] A full-platform risk baseline matrix is constructed based on the historical violation data of each anchor, and the matrix elements are calculated, wherein the row dimension of the risk baseline matrix represents the violation type violation risk baseline, and the column dimension represents the time window; the real-time violation feature vector F of the live broadcast platform in the monitoring time window is obtained eralFeral = [f1,...,fi ′ ,...], where where, is the violation probability of the k'th violation event corresponding to the i'th violation type, γ time is the first time decay factor, ΔT is the monitoring time window, i' is the number of violation type.
[0074] The platform risk baseline matrix and the anchor influence coefficient of each anchor are combined to calculate the dynamic violation risk coefficient; the platform violation risk level is determined based on the risk level mapping relationship, and the corrected data collection parameters stored in the database are matched based on the platform violation risk level. The full-platform risk baseline matrix is constructed, and the violation situation is recorded from two dimensions of violation type and time window, providing a systematic data framework for analyzing violation trends. By calculating the matrix elements, the risk degree of different violation types in each time window can be intuitively presented. The dynamic violation risk coefficient is calculated by combining the real-time violation feature vector and the anchor influence coefficient, which can accurately evaluate the current violation risk level of the platform. According to the risk level mapping relationship, the platform violation risk level is determined, and the corrected data collection parameters are matched accordingly, realizing the adaptive adjustment of the monitoring system. For example, when the violation risk is high, the data collection frequency is increased to capture violation behavior more timely; when the risk is low, the collection frequency is reasonably reduced to save system resources, thereby improving monitoring efficiency, optimizing resource allocation, and enhancing the effectiveness and sustainability of the monitoring system.
[0075] The full-platform risk baseline matrix is denoted as
[0076] The calculation method of the matrix elements is as follows:
[0077]
[0078] where, w i′j′ is the matrix element in the full-platform risk baseline matrix, δ(C i′,j′ ) is the number of violations of the i'th violation type at time window j', C is the maximum number of violations of the i'th violation type, j' is the number of time window, ξ is the second time decay factor, t now is the current time, t wg is the time of the last violation, and J is the total number of violation types; the sliding window update mechanism: delete historical data more than 3 months at 0 o'clock every day. For the calculation of matrix elements, factors such as the number of violations of the violation type within the time window, the maximum number of violations, and time decay are considered. Through this calculation method, the risk status of different violation types at different times can be accurately reflected, the historical trend and timeliness of violation behavior are embodied, and an accurate data basis is provided for evaluating the overall risk of the platform.
[0079] The method for obtaining the dynamic violation risk coefficient is:
[0080]
[0081] Wherein, λ risk is the dynamic violation risk coefficient, is the transpose matrix of the full-platform risk baseline matrix M base , and is the average of the influence factors. The calculation of the dynamic violation risk coefficient combines the real-time violation feature vector, the transpose of the full-platform risk baseline matrix, and the average of the influence factors, and integrates the real-time violation situation with the historical risk data and the anchor influence factors, so as to dynamically and comprehensively evaluate the current violation risk degree of the platform.
[0082] An image and voice recognition-based monitoring system for internet live broadcast violation behaviors, used for the image and voice recognition-based monitoring method for internet live broadcast violation behaviors as described above, as shown in Figure 2 The system comprises a live stream data acquisition module configured to acquire live stream data of anchors from a live broadcast platform based on determined data collection parameters and verify data integrity, wherein the live stream data comprises original video stream data and original audio stream data; a data preprocessing module configured to process the live stream data through a determined data processing mode to obtain preprocessed live stream data, including a video frame set and voice-to-text data; an image processing module configured to utilize a determined image violation processing model to perform feature recognition on the video frame set and output image violation marking information, including violation types of each image frame, violation probabilities, and image violation time stamps; a text processing module configured to utilize a determined text processing model to perform text processing on the voice-to-text data and output text violation marking information, including violation text passages and text violation time stamps; a violation judgment module configured to determine whether an anchor violates rules based on the image violation marking information and the text violation marking information and alarm when a violation occurs; and a collection parameter correction module configured to correct the data collection parameters based on historical violation data of the anchors.
[0083] An electronic device, comprising a processor and a memory having computer program instructions stored therein, wherein the computer program instructions, when executed by the processor, cause the processor to perform the image and voice recognition-based monitoring method for internet live broadcast violation behaviors as described above.
[0084] A computer-readable storage medium for storing a program, wherein the program, when executed by a processor, implements the image and voice recognition-based monitoring method for internet live broadcast violation behaviors as described above.
[0085] Those skilled in the art will appreciate that embodiments of the present application can be devised for a variety of applications. It is therefore intended that the present application cover all such modifications and variations of the application disclosed herein provided they come within the scope of the appended claims and their equivalents. It is intended to
[0086] The present application is described in reference to the drawings using a flowchart and / or a block diagram of an embodiment of the system, apparatus (system), and computer program product according to the present application. It will be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0087] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0089] While the preferred embodiments of the application have been described, additional variations and modifications can be employed by those skilled in the art. Therefore, the appended claims are intended to cover all such modifications and variations as fall within the scope of the present application.
[0090] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for monitoring violations of internet live streaming regulations using image and voice recognition, characterized in that, Includes the following steps: Based on the determined data collection parameters, the live stream data of each streamer is obtained from the live streaming platform and the integrity of the data is verified. The live stream data includes raw video stream data and raw audio stream data. Among them, the data collection parameters include the API application frequency and the data acquisition request interval of the web crawler program; Data acquisition parameters, data integrity threshold, data format correctness threshold, data processing mode, image violation processing model, and text processing model are all determined by obtaining the corresponding influence factors for each anchor. Obtain the set of mappings between the set impact factors and data processing parameters stored in the database, compare the impact factors with each set impact factor one by one, calculate the similarity based on Euclidean distance, determine the set impact factor corresponding to the minimum similarity, and obtain the corresponding data processing parameters from the database, including data collection parameters, data integrity threshold and data format correctness threshold. The data processing modes include video stream processing and audio stream processing. The video stream processing mode includes adaptive histogram equalization and Gaussian filtering for noise reduction, while the audio stream processing mode includes first-level and second-level text processing. The image violation handling models include Mask R-CNN and Faster R-CNN models, and the text processing models include Transformer architecture text processing and keyword matching models. The data processing modes, image violation handling models, and text processing models are determined as follows: The influence factors for each streamer are compared with the influence thresholds. If the impact factor is greater than the impact threshold, the video stream processing mode is the adaptive histogram equalization processing mode, the audio stream processing mode is the first-level text processing mode, the image violation processing model is the Mask R-CNN model, and the text processing model is the Transformer architecture text processing model. If the impact factor is not greater than the impact threshold, the video stream processing mode is Gaussian filtering noise reduction processing mode, the audio stream processing mode is second-level text processing mode, the image violation processing model is Faster R-CNN model, and the text processing model is keyword matching processing model. The impact factor is obtained as follows: Obtain the influence data of the live streaming platform and the influence data of each streamer to determine the influence coefficient of the live streaming platform. The influence coefficient of each streamer 'i' is the broadcaster's ID; Based on the influence coefficient of live streaming platforms The influence coefficient of each streamer Obtain the initial impact factor ; The correction factor is determined based on the real-time live streaming data of each streamer. Regarding the initial impact factor Real-time corrections are performed to obtain the impact factor. : e is the natural constant; Influence coefficient of live streaming platform The method to obtain it is as follows: ; Among them, is the market share coefficient, obtained based on the logistic regression model, is the user reputation coefficient, obtained based on the calculation model of the average score and the standard deviation of the score, is the media exposure coefficient, determined based on the analytic hierarchy process, is the technical input-output ratio coefficient, determined based on the technical innovation funds and the effect of achievement transformation, Influence coefficient of the streamer The method to obtain it is as follows: ; in, Factors related to the number of streamers' fans As a factor affecting the activity level of live streamers, The quality factor of the content broadcaster is X, where X is the total number of broadcasters on the live streaming platform. The correction factor is obtained as follows: Acquire real-time data during the streamer's live broadcast, including the growth rate (xZ) of viewers per unit time and the keyword density of viewers per unit time. The total number of online users per unit of time during the live stream. The correction factor is obtained based on the correction factor calculation model. : ; in, This is the average number of viewers during the streamer's historical live broadcasts; By using a defined data processing model, the live stream data is processed to obtain preprocessed live stream data, including a set of video frames and speech-to-text data. The defined image violation processing model is used to identify features of a set of video frames and output image violation marking information, including the violation type, violation probability and image violation timestamp for each image frame. Using a defined text processing model, speech-to-text data is processed to output text violation marker information, including violation text paragraphs and text violation timestamps; Based on image and text violation marker information, determine whether the broadcaster has violated regulations and issue an alert when a violation occurs; The data collection parameters were adjusted based on the streamer's historical violation data.
2. The method for monitoring violations of internet live streaming based on image and voice recognition according to claim 1, characterized in that, Obtain the data integrity coefficient of each streamer and data format correctness index : ; in, This represents the actual amount of data returned. To request the amount of data you expect to obtain, This represents the number of correctly formatted fields in the data. This represents the total number of fields in the data. If the data integrity coefficient is greater than the determined data integrity threshold for the corresponding streamer and the data format correctness index is greater than the determined data format correctness threshold for the corresponding streamer, then the streamer's data integrity is normal; otherwise, the data integrity is abnormal, and the streamer's live stream data is re-acquired.
3. The method for monitoring violations of internet live streaming based on image and voice recognition according to claim 2, characterized in that, The data collection parameters were adjusted based on the streamer's historical violation data, including the following steps: A platform-wide risk baseline matrix is constructed based on the historical violation data of each streamer. The matrix elements are calculated, where the row dimension of the risk baseline matrix represents the violation type and the violation risk base, and the column dimension represents the time window. Obtain the real-time violation feature vector of the live streaming platform within the monitoring time window, and calculate the dynamic violation risk coefficient by combining the platform risk baseline matrix and the influence coefficient of each streamer. The platform's violation risk level is determined based on the risk level mapping relationship, and the corrected data collection parameters stored in the database are matched based on the platform's violation risk level.
4. The method for monitoring violations of Internet live streaming based on image and voice recognition according to claim 3, characterized in that, The matrix elements are calculated as follows: ; in, These are matrix elements in the platform-wide risk baseline matrix. In the time window Time The number of violations of each type For the first The maximum number of violations for each type of violation The number represents the type of violation. The time window number, The time decay factor, For the current time, J represents the time of the most recent violation, and J represents the total number of violation types. The method for obtaining the dynamic violation risk coefficient is as follows: ; in, For dynamic violation risk coefficient, This is a real-time violation feature vector. For the entire platform risk baseline matrix The transpose of the matrix, This represents the mean of the impact factors.
5. A monitoring system for violations of internet live streaming regulations using image and speech recognition, used in conjunction with the monitoring method for violations of internet live streaming regulations using image and speech recognition as described in any one of claims 1-4, characterized in that, include: The live stream data acquisition module is used to acquire live stream data from each streamer on the live streaming platform based on determined data acquisition parameters and to verify the integrity of the data. The live stream data includes raw video stream data and raw audio stream data. The data preprocessing module is used to process the live stream data according to a defined data processing mode to obtain preprocessed live stream data, including a set of video frames and speech-to-text data. The image processing module is used to perform feature recognition on a set of video frames using a defined image violation processing model, and output image violation marking information, including the violation type, violation probability and image violation timestamp for each image frame; The text processing module is used to process speech-to-text data using a defined text processing model and output text violation marker information, including the violation text paragraphs and text violation timestamps. The violation detection module is used to determine whether the anchor has violated regulations based on image violation marker information and text violation marker information, and to issue an alarm when a violation occurs. The data collection parameter correction module is used to correct data collection parameters based on the streamer's historical violation data.
Citation Information
Patent Citations
Method and device for detecting violation behavior of live streaming room, electronic equipment and storage medium
CN113705370A
Block chain-based live broadcast data monitoring method and system
CN110740356A
Live video monitoring method and related device
CN112492343A