Live Content Risk Control Processing System and Supervision Method Based on Real-Time Data

Through real-time data analysis and natural language processing technology, the text content of live broadcast users is reviewed and divided and the growth continuous analysis is carried out, which solves the problems of inaccurate and illegal expansion of audits in live broadcast risk control, and achieves efficient and timely risk control strategies and security controls.

CN120151568BActive Publication Date: 2025-07-2251VV COM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510623220.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-22
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing live broadcast risk control audit technology has inaccurate audits, inability to promptly detect the surge in illegal users, and failure to deeply explore the user impact relationship, making it difficult to effectively control live broadcast risks.

Method used

Through real-time data analysis, keyword comparison and natural language processing technology are used to review and divide the text content of audience users, combine technology and manual review to identify intuitive and potential violation users, and conduct continuous growth analysis to judge the surge in violations and chain reactions, and determine the focus of users.

Benefits of technology

It improves the efficiency and reliability of live broadcast risk control review, realizes timely safety control and accurate risk control strategies, prevents the expansion of violations, and identifies and focuses on violating users who have a serious impact on other users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151568B_ABST
    Figure CN120151568B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of Internet live broadcasting. The present invention provides a live broadcast content risk control processing system and a supervision method based on real-time data, including: conducting review division and division coordination of technical review and manual review for the text content sent by audience users during the live broadcast, so as to improve the efficiency and reliability of the live broadcast risk control review. For the audience users who violate the rules, a growth continuity analysis is conducted to determine whether there is a surge in the number of audience users who violate the rules. If so, the safety control of the live broadcast is conducted to improve the timeliness of the live broadcast risk control. Through the identity comparison and analysis of the audience users who violate the rules, it is determined whether the audience users who violate the rules have a serious impact on the conversion of the violating users. If so, the key users who violate the rules and the attention priority of the key users who violate the rules are determined, and the influence relationship between the live broadcast violating users is deeply excavated to identify the key users who not only violate the rules themselves, but also drive other users to violate the rules and trigger a chain reaction of violations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet live broadcast, and specifically relates to a live broadcast content risk control processing system and supervision method based on real-time data. Background Art

[0002] With the rapid development of the live broadcast industry, the review and risk control of live broadcast content have become key links to ensure the healthy operation of the platform. At present, there are many defects in the existing live broadcast risk control review technologies. First, in terms of the review of live broadcast text content, due to the influence of semantic complexity, it is difficult to identify the accurate information in the audience's text simply relying on technical reviews, and it is extremely easy to have inaccurate reviews, affecting the review efficiency and reliability. Second, the monitoring of live broadcast violations lacks continuous analysis of the growth trend of violating users, and it is impossible to detect the sudden increase in violating users in a timely manner, resulting in difficulty in quickly taking effective control measures when the violation phenomenon seriously expands and missing the best risk control opportunity. Third, it fails to deeply explore the influence relationship among live broadcast violating users, and it is impossible to identify key users who not only violate regulations themselves but also drive other users to violate regulations and trigger a chain reaction of violations, making it difficult to formulate targeted risk control strategies and unable to meet the timeliness and accuracy requirements of live broadcast risk prevention and control. These problems have severely restricted the effective management and healthy development of live broadcast platforms for violating content.

[0003] Therefore, the present invention provides a live broadcast content risk control processing system and supervision method based on real-time data. Summary of the Invention

[0004] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0006] A live broadcast content risk control supervision method based on real-time data, including:

[0007] Through keyword comparison, the text sending content of audience users within a preset first monitoring period is reviewed and classified.

[0008] Based on the review results after review and classification, the audience users within the preset first monitoring period are classified to determine normal audience users and violating audience users. Among them, different violation handling measures are taken for the intuitive violating users and potential violating users included in the violating audience users.

[0009] Perform a growth persistence analysis on the violating audience users within a preset second monitoring period. If there is a continuous growth, by comparing and analyzing the violating audience users within the preset second monitoring period and the preset first monitoring period, determine whether there is a sudden increase in violations. If so, perform safety control of the live broadcast.

[0010] Perform a non - overlapping comparison of the user identities of the violating audience users during a preset second monitoring period and the potential violating users during a preset first monitoring period to determine newly added violating users. Then, perform an identity overlap comparison between the newly added violating users and normal audience users to determine converted violating users. Conduct a similarity analysis of the text sending content between the converted violating users and the violating audience users during the preset first monitoring period to determine whether the violating audience users during the preset first monitoring period trigger a violation chain reaction. If so, determine the key - focus violating users and their corresponding focus priorities.

[0011] Further, the audit classification method is as follows:

[0012] Use natural language processing technology to extract keywords from the text sending content. If there are keywords corresponding to the text sending content in the keyword library, classify the text sending content into technical audit;

[0013] If there are no keywords corresponding to the text sending content in the keyword library, use semantic recognition technology in natural language processing to determine whether there are potential risks in the text sending content. If there are, classify the text sending content into manual audit.

[0014] Further, the audit results after classification include manual audit results and technical audit results;

[0015] The process of classifying the audience users is as follows:

[0016] If the technical audit result shows a violation, mark the audience user corresponding to the text sending content as an obvious violating user;

[0017] If the manual audit result shows a violation, mark the audience user corresponding to the text sending content as a potential violating user.

[0018] Further, the adoption of different violation handling measures includes:

[0019] For obvious violating users, adopt corpus security control measures such as prohibiting comments and sending bullet screens. For potential violating users, adopt corpus security control measures such as platform warnings.

[0020] Further, the process of growth sustainability analysis is as follows:

[0021] Obtain the violating audience users at different times during the preset second monitoring period and count their numbers, integrate them into an audience violation sequence, perform linear regression on the audience violation sequence, and determine whether there is a continuous increase;

[0022] The method of determining whether there is a violation surge is as follows:

[0023] Obtain the illegal audience users within the preset second monitoring period and perform a difference calculation with the illegal audience users within the preset first monitoring period to obtain the illegal increase value;

[0024] If the illegal increase value is greater than or equal to the illegal increase threshold, there is an illegal increase phenomenon.

[0025] Furthermore, the method for determining new illegal users is as follows:

[0026] If the identity of the illegal audience user is different from that of all potential illegal users within the preset first monitoring period, mark the illegal audience user as a new illegal user;

[0027] The process for determining converted illegal users is as follows:

[0028] If the identity of a new illegal user is the same as that of any normal audience user within the preset first monitoring period, mark the new illegal user as a converted illegal user.

[0029] Furthermore, the method for determining whether the illegal audience user triggers an illegal chain reaction is as follows:

[0030] Obtain the text sending content showing violations in the audit results of converted illegal users and integrate them into a converted illegal text library;

[0031] Obtain the text sending content showing violations in the audit results of illegal audience users within the preset first monitoring period and integrate them into an initial illegal text library;

[0032] If at least one initial illegal text content in the initial illegal text library has a similarity with the converted illegal text content reaching the threshold, mark the converted illegal text content as illegal similar text content and mark the initial illegal text content as initial impact text content;

[0033] Identify the users affected by the violation based on the illegal similar text content;

[0034] Count the proportion of users affected by the violation among the converted illegal users,

[0035] If the proportion of users affected by the violation among the converted illegal users is greater than or equal to the preset proportion threshold, it indicates that the illegal audience user triggers an illegal chain reaction.

[0036] Furthermore, the converted illegal text content is the text sending content included in the converted illegal text library;

[0037] The initial illegal text content is the text sending content included in the initial illegal text library;

[0038] The method for identifying the users affected by the violation is as follows:

[0039] Statistically calculate the proportion of the number of similar text contents of violation in the converted violation text library. If it is greater than or equal to the preset threshold of the number proportion, mark the converted violation user as a user affected by the violation.

[0040] Further, the method for determining the key users of concern about violations and the corresponding concern priorities is as follows:

[0041] Obtain the violation audience users corresponding to the initial impact text content, and mark them as the key users of concern about violations;

[0042] Based on any key user of concern about violations, determine the concern priority of the key user of concern about violations according to the number of the initial impact text contents corresponding to the key user of concern about violations.

[0043] A live content risk control processing system based on real-time data includes:

[0044] Live broadcast audit division module: Through keyword comparison, audit and divide the text sending contents of audience users within a preset first monitoring period;

[0045] Audience user classification and processing module: Classify the audience users within a preset first monitoring period through the audit results after audit division, determine normal audience users and violation audience users. Among them, for the direct violation users and potential violation users included in the violation audience users, different violation handling measures are taken;

[0046] Live broadcast security control module: Conduct a growth persistence analysis on the violation audience users within a preset second monitoring period. If there is a continuous growth, judge whether there is a phenomenon of a sharp increase in violations by comparing and analyzing the violation audience users within the preset second monitoring period and the preset first monitoring period. If so, conduct security control of the live broadcast;

[0047] Live broadcast violation impact analysis module: Conduct a non-overlap comparison of user identities between the violation audience users within a preset second monitoring period and the potential violation users within a preset first monitoring period to determine new violation users, and conduct an identity overlap comparison between the new violation users and normal audience users to determine converted violation users. Conduct a similarity analysis of the text sending contents between the converted violation users and the violation audience users within a preset first monitoring period to judge whether the violation audience users within the preset first monitoring period trigger a violation chain reaction. If so, determine the key users of concern about violations and the corresponding concern priorities.

[0048] The beneficial effects of the present invention are as follows:

[0049] 1. Obtain the text content sent by audience users during the preset first monitoring period during the live broadcast, and conduct review and division of the text content through keyword extraction, comparison and analysis, including technical review and manual review, to prevent the technical review from being inaccurate due to the inability to identify ambiguity or uncertainty in the text content sent by audience users during the live broadcast. Therefore, the technical review and manual review are divided and coordinated to improve the efficiency and reliability of the live broadcast risk control review.

[0050] 2. Conduct growth continuity analysis on the number of violating viewers at different times during the preset second monitoring period. In the case of sustained growth, determine whether there is a surge in the number of violating viewers. If so, conduct safety control of the live broadcast, thereby achieving effective safety control of the live broadcast when the violation phenomenon seriously expands, and improving the timeliness of the risk control of the live broadcast.

[0051] 3. Perform a non-overlapping comparison of user identities between the violating audience users within the preset second monitoring period and the potential violating users within the preset first monitoring period to determine newly added violating users, and perform an identity overlap comparison between the newly added violating users and normal audience users to determine converted violating users, perform a similarity analysis on the text sent by the converted violating users and the violating audience users within the preset first monitoring period to determine whether the violating audience users within the preset first monitoring period trigger a chain reaction of violations, and if so, mark the violating audience users within the preset first monitoring period to determine the key violating users and the attention priority of the key violating users according to the number of initial affected text content items. The present invention identifies the violating users who not only violate the live broadcast but also have a serious violation impact on other audience users and drive the violation to affect the rhythm during the live broadcast, and determines the key attention priority of these live broadcast violating audiences in subsequent live broadcasts, while providing a risk control strategy for subsequent live broadcasts, and further improves the timeliness of live broadcast risk control. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The present invention will be further described below in conjunction with the accompanying drawings.

[0053] Figure 1 is a flowchart of the steps of the live content risk control and supervision method based on real-time data according to an embodiment of the present invention;

[0054] Figure 2 It is a flowchart of a live content risk control processing system based on real-time data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0056] Example 1: Please refer toFigure 1 As shown in Figure 1 , the live content risk control and supervision method based on real-time data according to the embodiments of the present invention includes the following steps:

[0057] Step 1: Obtain the text sending content of the audience users during a preset first monitoring period in the live broadcast process, and conduct review and classification on the text sending content through keyword extraction and comparison analysis;

[0058] In Step 1, the text sending content includes the comments of the audience users and the bullet screen text sent during the live broadcast process;

[0059] In Step 1, the methods for reviewing the text sending content include technical review and manual review. Among them, the technical review is implemented through the intelligent review system during the live broadcast;

[0060] In Step 1, the process of reviewing and classifying the text sending content is as follows:

[0061] Obtain the text sending content of the audience users during a preset first monitoring period in the live broadcast process, and use natural language processing technology to extract the keywords in the text sending content, and compare the keywords with the keyword library in the intelligent review system. The specific comparison process is as follows:

[0062] If there is a keyword corresponding to the text sending content in the keyword library, the text sending content is classified into technical review;

[0063] If there is no keyword corresponding to the text sending content in the keyword library, use the semantic recognition technology in natural language processing to determine whether there is a potential risk in the text sending content;

[0064] If there is, the text sending content is classified into manual review. If not, no review process is performed on the text sending content;

[0065] Among them, the methods for using the semantic recognition technology in natural language processing to determine whether there is a potential risk in the text sending content include but are not limited to:

[0066] Method 1: Analyze the semantic structure of the text, including:

[0067] Syntax parsing: Use natural language processing tools to analyze the syntax of the text, determine the subject-predicate-object, attributive-adverbial-complementary, etc. structures of the sentence, understand the syntactic relationships between words. For example, in analyzing the sentence "He told a somewhat negative story", it is determined that "negative" is the attributive modifying "story", and the core semantics of the sentence is clarified;

[0068] Semantic role labeling: Identify the semantic roles played by each word in the text, such as agent, patient, time, place, etc. For example, in "In the live broadcast yesterday, he mentioned some information", it is clear that "he" is the person who mentioned the information and "live broadcast yesterday" is a time adverbial, which helps to accurately understand the semantics;

[0069] Dependency syntactic analysis: Analyze the dependency relationship between words and judge the modification, dominance and other relationships between words. For example, in "The description of that scene has a negative tendency", dependency analysis shows that "negative tendency" is a supplement to "description", which further understands the semantic level of the text.

[0070] Method 2: Evaluate the sentiment of the text, including:

[0071] Sentiment analysis: Use sentiment analysis models to determine the sentiment polarity of the text to see whether there are negative emotions such as negativity, anger, and dissatisfaction, which are related to potential risks;

[0072] Topic model analysis: Analyze the topic distribution of the text through topic models such as LDA to determine whether it is related to potential risks;

[0073] Step 2: Classify the audience users in the preset first monitoring period through the audit results after audit division, determine normal audience users and illegal audience users, where illegal audience users include direct illegal users and potential illegal users, and take different illegal handling measures for different illegal audience users after classification;

[0074] In step 2, the audit results after audit division include manual audit results and technical audit results, wherein both manual audit results and technical audit results include two types of audit results: violation and non-violation;

[0075] In step 2, the process of classifying the audience users in the preset first monitoring period according to the audit results after audit division is as follows:

[0076] If the technical review results show a violation, the viewer user corresponding to the text content will be marked as an intuitively violating user;

[0077] If the manual review result shows a violation, the viewer user corresponding to the text content will be marked as a potential violator;

[0078] The viewer user corresponding to the text-sent content indicates that the text-sent content is sent by the viewer user during the live broadcast;

[0079] It should be noted that since the technical review result shows a violation, it means that the keywords in the keyword library are directly present in the text sent by the audience user; then the audience user corresponding to the text sent is marked as an intuitive violation user. Since the manual review result shows a violation, although the keywords in the keyword library are not present in the text sent, but there is a potential risk, then the audience user corresponding to the text sent is marked as a potential violation user;

[0080] In step two, the way to take different violation handling measures for different violated audience users after classification is as follows:

[0081] For intuitive violation users, corpus security control measures such as prohibiting comments and sending bullet screens are adopted. For potential violation users, corpus security control measures such as platform warnings are adopted;

[0082] The technical solution of the embodiment of the present invention is: obtain the text sent by the audience user during the preset first monitoring period in the live broadcast, and through the extraction and comparison analysis of keywords, review and classify the text sent, including technical review and manual review, to prevent the problem that the technical review is inaccurate due to the inability to identify the fuzzy or uncertain situation in the text sent by the audience user during the live broadcast. Therefore, the division and cooperation of technical review and manual review are carried out to improve the efficiency and reliability of live broadcast risk control review.

[0083] Embodiment 2: Please refer to Figure 1 As shown, on the basis of Embodiment 1 of the present invention, considering the subsequent risk control development of the live broadcast, therefore, a growth analysis is carried out on the violated audience users during the subsequent detection period to prevent the expansion of live broadcast violation phenomena, so as to carry out risk control processing of the live broadcast and realize effective control of live broadcast content violations. Therefore, the method for live broadcast content risk control supervision based on real-time data described in the embodiment of the present invention further includes the following steps:

[0084] Step three: Obtain the violated audience users at different times during the preset second monitoring period in the live broadcast after taking different violation handling measures, and conduct a growth persistence analysis. If there is a continuous growth, by comparing and analyzing the violated audience users during the preset second monitoring period and the preset first monitoring period, judge whether there is a phenomenon of rapid increase in violations. If so, carry out safety control of the live broadcast;

[0085] In step three, the process of conducting a growth persistence analysis is as follows:

[0086] Obtain the violated audience users at different times during the preset second monitoring period and count the number, and integrate them into an audience violation sequence;

[0087] Perform a linear regression on the sequence of viewer violations to determine the linear regression equation y = kx + b for viewer user violations, where y is the number of violating viewer users, k is the slope, x is the time, and b is the intercept;

[0088] Calculate the significance P-value for the linear regression equation regarding viewer user violations;

[0089] If the slope of the linear regression equation regarding viewer user violations is positive and the significance P-value is greater than the significance P-threshold, it indicates that the number of violating viewer users at different times during the preset second monitoring period shows a continuous increase;

[0090] If the slope of the linear regression equation regarding viewer user violations is positive and the significance P-value is greater than the significance P-threshold, it indicates that the number of violating viewer users at different times during the preset second monitoring period does not show a continuous increase. Then, different violation handling measures are taken for the obvious violating users and potential violating users during the preset second monitoring period;

[0091] In step three, the method for determining whether there is a phenomenon of a sharp increase in violations is as follows:

[0092] Obtain the violating viewer users during the preset second monitoring period and perform a difference ratio calculation with the violating viewer users during the preset first monitoring period to obtain the violation increase value;

[0093] Among them, the violation increase value reflects the growth situation of the violating viewer users during the preset second monitoring period compared to those during the preset first monitoring period. The violation increase value = (the number of violating viewer users during the preset second monitoring period - the number of violating viewer users during the preset first monitoring period) ÷ the number of violating viewer users during the preset first monitoring period;

[0094] Compare the violation increase value with the violation increase threshold;

[0095] If the violation increase value is greater than or equal to the violation increase threshold, it indicates that there is a phenomenon of a sharp increase in the number of violating viewer users during the preset second monitoring period compared to those during the preset first monitoring period. Then, perform safety control for the live broadcast, including but not limited to muting the live broadcast and closing the live broadcast, to prevent the further deterioration and expansion of the live broadcast violation situation;

[0096] If the violation increase value is less than the violation increase threshold, it indicates that there is no phenomenon of a sharp increase in the number of violating viewer users during the preset second monitoring period compared to those during the preset first monitoring period;

[0097] The technical solution of the embodiment of the present invention is as follows: For the illegal audience users at different times within a preset second monitoring period, growth sustainability analysis is carried out. In the case of continuous growth, it is judged whether there is a sudden increase in the illegal audience users. If so, security control of the live broadcast is carried out, realizing effective live broadcast security control in the case of serious expansion of live broadcast violations and improving the timeliness of live broadcast risk control.

[0098] Embodiment 3: Please refer to Figure 1 As shown, on the basis of Embodiment 1 and Embodiment 2 of the present invention, identity comparison and analysis are carried out on newly added illegal users subsequently, aiming to identify some illegal users with serious illegal impacts among the illegal users, and conduct key attention to violations. In subsequent live broadcasts, key attention to violations of illegal users with serious illegal impacts is conducive to improving the timeliness of live broadcast risk control. Therefore, the method for live broadcast content risk control and supervision based on real-time data described in the embodiment of the present invention further includes the following steps:

[0099] Step 4: Perform non-overlapping comparison of user identities between the illegal audience users within the preset second monitoring period and the potential illegal users within the preset first monitoring period to determine newly added illegal users, and perform overlapping comparison of identities between the newly added illegal users and normal audience users to determine converted illegal users. Perform text sending content similarity analysis on the converted illegal users and the illegal audience users within the preset first monitoring period to judge whether the illegal audience users within the preset first monitoring period trigger an illegal chain reaction. If so, mark the illegal audience users within the preset first monitoring period to determine the key attention users for violations and the attention priority levels of the key attention users for violations;

[0100] In Step 4, the process of performing non-overlapping comparison of user identities to determine newly added illegal users is as follows:

[0101] Based on any one illegal audience user within the preset second monitoring period;

[0102] If the identity of the illegal audience user is the same as that of any one potential illegal user within the preset first monitoring period, no processing is performed;

[0103] If the identity of the illegal audience user is different from that of all potential illegal users within the preset first monitoring period, mark the illegal audience user as a newly added illegal user;

[0104] It should be noted that the reason for performing non-overlapping comparison of user identities between the illegal audience users within the preset second monitoring period and the potential illegal users within the preset first monitoring period is that since intuitive illegal users within the preset first monitoring period have already adopted language corpus security control measures such as prohibiting comments and sending bullet screens, the illegal audience users within the preset second monitoring period will not have the same identity as the intuitive illegal users within the preset first monitoring period;

[0105] In step four, for the comparison of overlapping user identities to determine the process of converting a non - compliant user:

[0106] Based on any newly added non - compliant user;

[0107] If the identity of the newly added non - compliant user is the same as that of any normal viewer user within the preset first monitoring period, mark the newly added non - compliant user as a converted non - compliant user;

[0108] If the identity of the newly added non - compliant user is different from that of all normal viewer users within the preset first monitoring period, no processing is performed;

[0109] It should be noted that a converted non - compliant user means that within the preset first monitoring period, the user did not violate the regulations, but violated the regulations within the preset second monitoring period;

[0110] In step four, for the analysis of the similarity of the text sending content between the converted non - compliant users and the non - compliant viewer users within the preset first monitoring period to determine whether the non - compliant viewer users within the preset first monitoring period trigger a chain of non - compliance:

[0111] Obtain the text sending content that shows non - compliance in the audit result of the converted non - compliant user and integrate it into a converted non - compliant text library;

[0112] Among them, each converted non - compliant user corresponds to a converted non - compliant text library, and the converted non - compliant text library contains all the non - compliant text sending content of the corresponding converted non - compliant user;

[0113] Obtain the text sending content that shows non - compliance in the audit result of the non - compliant viewer users within the preset first monitoring period and integrate it into an initial non - compliant text library. Among them, the initial non - compliant text library contains all the text sending content that shows non - compliance in the audit results of all non - compliant viewer users within the preset first monitoring period;

[0114] Mark the text sending content included in the converted non - compliant text library as converted non - compliant text content;

[0115] Mark the text sending content included in the initial non - compliant text library as initial non - compliant text content;

[0116] Based on any one piece of converted non - compliant text content in any converted non - compliant text library;

[0117] If in the initial non - compliant text library, there is at least one piece of initial non - compliant text content whose similarity to the converted non - compliant text content reaches the threshold, mark the converted non - compliant text content as non - compliant similar text content and mark the initial non - compliant text content as initial impact text content;

[0118] If none of the initial violation text contents in the initial violation text library has a similarity with the converted violation text content reaching the threshold, no operation will be performed on the converted violation text content;

[0119] Among them, the way to perform similarity analysis is to convert the text into vectors by using a pre-trained semantic model (such as BERT, Word2Vec) and calculate the similarity between the vectors;

[0120] Count the proportion of the number of violation similar text contents in the converted violation text library and compare it with the preset proportion threshold of the number of texts;

[0121] If the proportion of the number of violation similar text contents in the converted violation text library is greater than or equal to the preset proportion threshold of the number of texts, it means that there are multiple text sending contents in the converted violation text library that are similar to the text sending contents in the initial violation text library, which means that the possibility that the converted violation user is affected by the violations of the violation audience users within the preset first monitoring period is relatively high, then mark the converted violation user as a user affected by violations;

[0122] If the proportion of the number of violation similar text contents in the converted violation text library is less than the preset proportion threshold of the number of texts, no processing will be performed;

[0123] Count the proportion of the number of users affected by violations among the converted violation users and compare it with the preset proportion threshold of the number;

[0124] If the proportion of the number of users affected by violations among the converted violation users is greater than or equal to the preset proportion threshold of the number, it means that among the converted violation users, there are a majority of converted violation users who have live broadcast violations due to being affected by the violations of the violation audience users within the preset first monitoring period, indicating that the violation audience users within the preset first monitoring period trigger a violation chain reaction;

[0125] If the proportion of the number of users affected by violations among the converted violation users is less than the preset proportion threshold of the number, no operation will be performed;

[0126] In step four, the process of marking the violation audience users within the preset first monitoring period and determining the violation key attention users and the attention priority levels of the violation key attention users is as follows:

[0127] Obtain the violation audience users corresponding to the initial impact text content and mark them as violation key attention users;

[0128] Based on any violation key attention user, determine the attention priority level of the violation key attention user according to the number of initial impact text contents corresponding to the violation key attention user;

[0129] The technical solution of the embodiment of the present invention is as follows: non-overlapping comparison of user identities is performed between the illegal audience users within a preset second monitoring period and the potential illegal users within a preset first monitoring period to determine newly added illegal users, and then identity overlapping comparison is performed between the newly added illegal users and normal audience users to determine converted illegal users. Similarity analysis of the text sending content is carried out between the converted illegal users and the illegal audience users within the preset first monitoring period to determine whether the illegal audience users within the preset first monitoring period trigger an illegal chain reaction. If so, mark the illegal audience users within the preset first monitoring period, determine the key illegal users to be concerned about, and determine the attention priority of the key illegal users to be concerned about according to the number of initial impact text content items. The present invention identifies the illegal users who not only violate the regulations during the live broadcast but also have a serious illegal impact on other audience users and drive the rhythm of the illegal impact, and determines the key attention priority of these live broadcast illegal audience users in subsequent live broadcasts, providing a risk control strategy for subsequent live broadcasts and further improving the timeliness of live broadcast risk control.

[0130] Embodiment 4: Please refer to Figure 2 As shown in the figure, a live broadcast content risk control processing system based on real-time data according to an embodiment of the present invention includes:

[0131] Live broadcast review classification module: Obtain the text sending content of audience users within a preset first monitoring period during the live broadcast, and perform review classification on the text sending content through keyword extraction and comparison analysis;

[0132] The process of performing review classification on the text sending content is as follows:

[0133] Obtain the text sending content of audience users within a preset first monitoring period during the live broadcast, use natural language processing technology to extract keywords in the text sending content, and compare the keywords with the keyword library in the intelligent review system. Among them, the specific comparison process is as follows:

[0134] If there are keywords corresponding to the text sending content in the keyword library, classify the text sending content into technical review;

[0135] If there are no keywords corresponding to the text sending content in the keyword library, use semantic recognition technology in natural language processing to determine whether there are potential risks in the text sending content. If so, classify the text sending content into manual review;

[0136] Audience user classification and processing module: Classify the audience users within a preset first monitoring period according to the review results after review classification to determine normal audience users and illegal audience users. Among them, the illegal audience users include direct illegal users and potential illegal users, and different illegal treatment measures are taken for different classified illegal audience users;

[0137] The review results after review classification include manual review results and technical review results. Among them, both the manual review results and the technical review results include two review results: violation and non-violation;

[0138] The process of classifying the audience users within a preset first monitoring period based on the review results after review classification is as follows:

[0139] If the technical review result shows a violation, mark the audience user corresponding to the text sending content as an intuitive violation user;

[0140] If the manual review result shows a violation, mark the audience user corresponding to the text sending content as a potential violation user;

[0141] Among them, the audience user corresponding to the text sending content means that the text sending content was sent by this audience user during the live broadcast;

[0142] The method of taking different violation handling measures for different classified violation audience users is as follows:

[0143] For intuitive violation users, adopt language material security control measures such as prohibiting comments and sending bullet screens. For potential violation users, adopt language material security control measures such as platform warnings;

[0144] Live broadcast security control module: Obtain the violation audience users at different times within a preset second monitoring period during the live broadcast after taking different violation handling measures, and conduct growth persistence analysis. If there is a continuous growth, through comparative analysis of the violation audience users within the preset second monitoring period and the preset first monitoring period, judge whether there is a phenomenon of rapid increase in violations. If so, perform live broadcast security control;

[0145] The process of conducting growth persistence analysis is as follows:

[0146] Obtain the violation audience users at different times within the preset second monitoring period and count the number, and integrate them into an audience violation sequence;

[0147] Perform linear regression on the audience violation sequence to determine the linear regression equation y = kx + b for audience user violations. Among them, y is the number of violation audience users, k is the slope, x is the time, and b is the intercept;

[0148] Calculate the significance P value for the linear regression equation for audience user violations;

[0149] If the slope of the linear regression equation for audience user violations is positive and the significance P value is greater than the significance P threshold, it means that the violation audience users at different times within the preset second monitoring period show continuous growth;

[0150] If the slope of the linear regression equation regarding the violations of the audience users is positive and the significance P-value is greater than the significance P-threshold, it indicates that there is no continuous growth in the number of violating audience users at different times during the preset second monitoring period. Different violation handling measures will be taken for the obvious violating users and potential violating users during the preset second monitoring period;

[0151] The method for determining whether there is a phenomenon of a sharp increase in violations is as follows:

[0152] Obtain the violating audience users during the preset second monitoring period and perform a difference ratio calculation with the violating audience users during the preset first monitoring period to obtain the violation sharp increase value;

[0153] Among them, the violation sharp increase value reflects the growth situation of the violating audience users during the preset second monitoring period compared to those during the preset first monitoring period. The violation sharp increase value = (the violating audience users during the preset second monitoring period - the violating audience users during the preset first monitoring period) ÷ the violating audience users during the preset first monitoring period;

[0154] Compare the violation sharp increase value with the violation sharp increase threshold;

[0155] If the violation sharp increase value is greater than or equal to the violation sharp increase threshold, it indicates that there is a phenomenon of a sharp increase in the violating audience users during the preset second monitoring period compared to those during the preset first monitoring period. Then, perform safety control for the live broadcast, including but not limited to muting the live broadcast and closing the live broadcast, to prevent the further deterioration and expansion of the live broadcast violation situation;

[0156] If the violation sharp increase value is less than the violation sharp increase threshold, it indicates that there is no phenomenon of a sharp increase in the violating audience users during the preset second monitoring period compared to those during the preset first monitoring period;

[0157] Live broadcast violation impact analysis module: Compare the violating audience users during the preset second monitoring period with the potential violating users during the preset first monitoring period to determine non-overlapping user identities, identify new violating users, and then compare the new violating users with the normal audience users to determine converted violating users. Analyze the similarity of the text sent by the converted violating users and the violating audience users during the preset first monitoring period to determine whether the violating audience users during the preset first monitoring period trigger a violation chain reaction. If so, mark the violating audience users during the preset first monitoring period to determine the key violation attention users and the attention priority levels of the key violation attention users;

[0158] The process of performing non-overlapping user identity comparison to determine new violating users is as follows:

[0159] Based on any one of the violating audience users during the preset second monitoring period;

[0160] If the identity of the violating audience user is different from that of all potential violating users within the preset first monitoring period, the violating audience user is marked as a newly added violating user;

[0161] The process of determining the converted violating user by performing user identity coincidence comparison is as follows:

[0162] Based on any one of the newly added violating users;

[0163] If the identity of the newly added violating user is the same as that of any one of the normal audience users within the preset first monitoring period, the newly added violating user is marked as a converted violating user;

[0164] If the identity of the newly added violating user is different from that of all normal audience users within the preset first monitoring period, no processing is performed;

[0165] The process of analyzing the similarity of the text sending content between the converted violating user and the violating audience users within the preset first monitoring period to determine whether the violating audience users within the preset first monitoring period trigger a violation chain reaction is as follows:

[0166] Obtain the text sending content that shows violations in the audit results of the converted violating users and integrate it into a converted violation text library;

[0167] Among them, each converted violating user corresponds to a converted violation text library, and the converted violation text library contains all the violating text sending content of the corresponding converted violating user;

[0168] Obtain the text sending content that shows violations in the audit results of the violating audience users within the preset first monitoring period and integrate it into an initial violation text library. Among them, the initial violation text library contains all the text sending content that shows violations in the audit results of all violating audience users within the preset first monitoring period;

[0169] Mark the text sending content included in the converted violation text library as converted violation text content;

[0170] Mark the text sending content included in the initial violation text library as initial violation text content;

[0171] Based on any one converted violation text content in any one converted violation text library;

[0172] If in the initial violation text library, there is at least one initial violation text content whose similarity to the converted violation text content reaches the threshold, mark the converted violation text content as violation similar text content and mark the initial violation text content as initial impact text content;

[0173] Among them, the way to perform similarity analysis is to convert the text into vectors by using pre-trained semantic models (such as BERT, Word2Vec), and calculate the similarity between the vectors;

[0174] Count the proportion of the number of illegal similar text contents in the converted illegal text library, and compare it with the preset proportion threshold of the number of texts;

[0175] If the proportion of the number of illegal similar text contents in the converted illegal text library is greater than or equal to the preset proportion threshold of the number of texts, it means that there are multiple text sending contents in the converted illegal text library that are similar to the text sending contents in the initial illegal text library, indicating that the converted illegal users are more likely to be affected by the violations of the illegal audience users within the preset first monitoring period. Then, mark the converted illegal users as users affected by violations;

[0176] Count the proportion of the number of users affected by violations among the converted illegal users, and compare it with the preset proportion threshold of the number;

[0177] If the proportion of the number of users affected by violations among the converted illegal users is greater than or equal to the preset proportion threshold of the number, it means that among the converted illegal users, most of the converted illegal users are affected by the violations of the illegal audience users within the preset first monitoring period and have live broadcast violations, indicating that the illegal audience users within the preset first monitoring period have a serious impact on the converted illegal users' live broadcast violations;

[0178] The process of marking the illegal audience users within the preset first monitoring period and determining the key users to be concerned about violations and the attention priority of the key users to be concerned about violations is as follows:

[0179] Obtain the illegal audience users corresponding to the initial impact text content, and mark them as key users to be concerned about violations;

[0180] Based on any key user to be concerned about violations, determine the attention priority of the key user to be concerned about violations according to the number of initial impact text contents corresponding to the key user to be concerned about violations.

[0181] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A live content risk control and supervision method based on real-time data, characterized in that: include: By comparing keywords, the text content sent by the audience users in the preset first monitoring period is reviewed and divided; The audience users in the preset first monitoring period are classified according to the audit results after audit division, and normal audience users and illegal audience users are determined, wherein different violation handling measures are taken for the illegal audience users, including direct illegal users and potential illegal users; The audit results after the audit division include manual audit results and technical audit results; The process of classifying the audience users is as follows: If the technical review results show a violation, the viewer user corresponding to the text content will be marked as an intuitively violating user; If the manual review result shows a violation, the viewer user corresponding to the text content will be marked as a potential violator; Conduct growth continuity analysis on the violating viewers in the preset second monitoring period. If there is a continuous growth, compare and analyze the violating viewers in the preset second monitoring period and the preset first monitoring period to determine whether there is a surge in violations. If so, conduct safety control on the live broadcast. A non-overlapping comparison of user identities is performed between the violating audience users within the preset second monitoring period and the potential violating users within the preset first monitoring period to determine newly added violating users, and an identity overlap comparison is performed between the newly added violating users and normal audience users to determine converted violating users, and a similarity analysis of the text sent by the converted violating users and the violating audience users within the preset first monitoring period is performed to determine whether the violating audience users within the preset first monitoring period trigger a chain reaction of violations, and if so, the key violating users and the corresponding attention priorities are determined.

2. The live broadcast content risk control and supervision method based on real-time data according to claim 1 is characterized by: The audit division method is as follows: Use natural language processing technology to extract keywords from text content. If the keywords corresponding to the text content exist in the keyword library, the text content will be classified for technical review. If the keyword corresponding to the text content does not exist in the keyword library, the semantic recognition technology in natural language processing is used to determine whether the text content has potential risks. If so, the text content will be classified for manual review.

3. The live broadcast content risk control and supervision method based on real-time data according to claim 1 is characterized in that: The different violation handling measures include: For users who obviously violate the rules, the corpus security control measure of prohibiting comments and sending barrages is adopted. For users who may violate the rules, the corpus security control measure of platform warning is adopted.

4. The live broadcast content risk control and supervision method based on real-time data according to claim 1 is characterized in that: The process of growth sustainability analysis is: Obtain and count the number of audience users who violate the rules at different times during the preset second monitoring period, integrate them into an audience violation sequence, perform linear regression on the audience violation sequence, and determine whether there is a sustained increase; The method for determining whether there is a surge in violations is as follows: Obtain the illegal audience users in the preset second monitoring period and perform difference calculation with the illegal audience users in the preset first monitoring period to obtain the illegal surge value; If the illegal surge value is greater than or equal to the illegal surge threshold, there is an illegal surge phenomenon.

5. The live content risk control and supervision method based on real-time data according to claim 1, characterized in that: The method for determining newly added illegal users is as follows: If the illegal audience users are not identical to all potential illegal users within the preset first monitoring period, mark the illegal audience users as newly added illegal users; The process for determining converted illegal users is as follows: If a newly added illegal user is identical to any normal audience user within the preset first monitoring period, mark the newly added illegal user as a converted illegal user.

6. The live content risk control and supervision method based on real-time data according to claim 1, characterized in that: The method for judging whether the illegal audience users trigger an illegal chain reaction is as follows: Obtain the text sending content with illegal audit results shown for the converted illegal users, and integrate it into the converted illegal text library; Obtain the text sending content with illegal audit results shown for the illegal audience users within the preset first monitoring period, and integrate it into the initial illegal text library; If there is at least one initial illegal text content in the initial illegal text library whose similarity to the converted illegal text content reaches the threshold, mark the converted illegal text content as the illegal similar text content, and mark the initial illegal text content as the initial influencing text content; Identify the users affected by the illegal behavior according to the illegal similar text content; Count the proportion of the number of users affected by the illegal behavior among the converted illegal users, If the proportion of the number of users affected by the illegal behavior among the converted illegal users is greater than or equal to the preset quantity proportion threshold, it means that the illegal audience users trigger an illegal chain reaction.

7. The live content risk control and supervision method based on real-time data according to claim 6, characterized in that: The converted illegal text content is the text sending content included in the converted illegal text library; The initial illegal text content is the text sending content included in the initial illegal text library; The method for identifying the users affected by the illegal behavior is as follows: Count the proportion of the number of the illegal similar text content in the converted illegal text library. If it is greater than or equal to the preset number proportion threshold, mark the converted illegal users as the users affected by the illegal behavior.

8. The live content risk control and supervision method based on real-time data according to claim 1, characterized in that: The method for determining the illegal key attention users and the corresponding attention priorities is as follows: Obtain the illegal audience users corresponding to the initial influencing text content, and mark them as the illegal key attention users; Based on any one of the illegal key attention users, determine the attention priority of the illegal key attention users according to the number of the initial influencing text content corresponding to the illegal key attention users.

9. A live content risk control processing system based on real-time data, characterized in that: Including: Live audit division module: Through keyword comparison, audit and divide the text sending content of the audience users within the preset first monitoring period; Audience user classification and processing module: Classify the audience users within the preset first monitoring period according to the audit results after audit division, determine the normal audience users and the illegal audience users. Among them, for the direct illegal users and potential illegal users included in the illegal audience users, different illegal handling measures are taken; The review results after review division include manual review results and technical review results; The process of classifying the audience users is as follows: If the technical review result shows a violation, the audience user corresponding to the text sending content is marked as an intuitive violation user; If the manual review result shows a violation, the audience user corresponding to the text sending content is marked as a potential violation user; Live broadcast security control module: Analyze the growth persistence of the violating audience users within the preset second monitoring period. If there is a continuous increase, determine whether there is a surge in violations by comparing and analyzing the violating audience users within the preset second monitoring period and the preset first monitoring period. If so, perform security control on the live broadcast; Live broadcast violation impact analysis module: Compare the identities of the violating audience users within the preset second monitoring period with the potential violation users within the preset first monitoring period to determine new violating users, and then compare the identities of the new violating users with the normal audience users to determine converted violating users. Analyze the similarity of the text sending content between the converted violating users and the violating audience users within the preset first monitoring period to determine whether the violating audience users within the preset first monitoring period trigger a violation chain reaction. If so, determine the key violation attention users and the corresponding attention priorities.

Citation Information

Patent Citations

  • Live broadcast control method and device, equipment and storage medium

    CN113613026A

  • Illegal self-media user identification method and system based on transverse graph federal learning

    CN117688535A