Content auditing method, computer equipment and computer storage medium
By reviewing the content of reading operation requests of Internet platform users, extracting data characteristics and determining the probability of random inspections, identifying and manually reviewing illegal content, the problem of insufficient algorithm identification in the existing technology is solved, and more efficient content review is achieved.
Patent Information
- Application Number
- CN202510643389.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technology cannot identify new illegal content due to insufficient algorithm capabilities in the Internet platform, resulting in the missed review of risk content, and the risk of massive existing data cannot be effectively solved. The focus of the review is to ensure the security of daily new content. UGC content that has not been actively reported is risky.
By reviewing the user's read operation request, extracting the data characteristics of the content to be reviewed, determining the probability of random inspection based on the characteristics, selecting the random inspection content for identification and manual review, reducing business invasiveness and covering massive read operation content.
It reduces the risk of content missed reporting and missed review, improves content review efficiency, reduces review business volume, and improves the accuracy of review of user requested content.
Smart Images

Figure CN120492758A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of content review, and specifically to a content review method, computer equipment, and computer storage medium. Background Art
[0002] On internet platforms, the amount of user-generated content (UGC) continues to grow, and the scenarios are becoming increasingly complex. This content must be reviewed and approved by the review system before it can be published to the internet platform and viewed by platform users. When a user uploads or creates UGC content, the review system writes the user's UGC content to the platform's storage system through a write operation. The uploaded UGC content is then reported to the content security system for both machine and human review.
[0003] However, the existing method reviews the content uploaded by write operations, and its focus is on ensuring the security of daily new UGC content. However, due to insufficient algorithm capabilities or the inability to identify new types of illegal content, risky content is missed during upload, and the risk of massive stock data cannot be resolved. Summary of the Invention
[0004] The embodiments of the present application provide a content review method, computer equipment, and computer storage medium. By reviewing the content of read operation requests, the content review business does not need to be actively reported, so it is less invasive to the business, reduces the risk of content underreporting and underreview, and can cover massive amounts of read operation content, thereby improving content review efficiency.
[0005] A first aspect of an embodiment of the present application provides a content review method, the method comprising:
[0006] Obtaining content requested by a user's read operation request, where the content requested by the read operation request includes content to be reviewed;
[0007] Extracting data features of the content to be reviewed, and determining a probability of random inspection of the content to be reviewed based on the data features of the content to be reviewed;
[0008] Determining, among the plurality of pieces of content to be reviewed, the content to be reviewed based on the probability of random inspection of the content to be reviewed;
[0009] The sample inspection content is identified according to the content features of the sample inspection content to obtain a sample inspection content identification result, and the sample inspection content identification result is used for manual review.
[0010] A second aspect of an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0011] A third aspect of an embodiment of the present application provides a computer storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer executes the method of the first aspect.
[0012] A fourth aspect of the embodiments of the present application provides a computer program product, which, when running on a computer device, enables the computer device to execute the method of the first aspect.
[0013] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0014] The computer device obtains the content requested by the user's read operation request, which includes content to be reviewed. The data features of the content to be reviewed are extracted, and the probability of random inspection of the content to be reviewed is determined based on the data features of the content to be reviewed. Among the multiple pieces of content to be reviewed, the random inspection content among the multiple pieces of content to be reviewed is determined based on the random inspection probability of the content to be reviewed, and the random inspection content is identified based on the content features of the random inspection content to obtain the random inspection content identification result. Therefore, by reviewing the content of the read operation request, the content review business does not need to be actively reported, so it is less invasive to the business, reduces the risk of content underreporting and under-review, and can cover a large amount of read operation content, thereby improving content review efficiency.
[0015] Secondly, through random sampling, massive amounts of content are screened and reviewed. For example, content with a high probability of violation (the sampling probability represents the probability of violation) is randomly reviewed, so that review resources can be concentrated on content with a higher probability of violation, avoiding the problem of large review workload caused by reviewing every piece of content. It can reduce the data scale of the review content, reduce the review workload, and thus improve the review efficiency of user-requested content. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a schematic diagram of the network framework in the embodiment of this application;
[0017] Figure 2 This is a flow chart of the content review method in the embodiment of the present application;
[0018] Figure 3 This is another flowchart of the content review method in the embodiment of the present application;
[0019] Figure 4A schematic diagram of an exemplary application scenario of the content review method in an embodiment of the present application;
[0020] Figure 5 This is a structural diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The embodiments of the present application provide a content review method, computer equipment, and computer storage medium. By reviewing the content of read operation requests, the content review business does not need to be actively reported, so it is less invasive to the business, reduces the risk of content underreporting and underreview, and can cover massive amounts of read operation content, thereby improving content review efficiency.
[0022] See also Figure 1 In the embodiment of the present application, the network framework includes:
[0023] The service server 100 and the terminal cluster; the terminal cluster may include: terminal device 200a, terminal device 200b, terminal device 200c, ..., terminal device 200n and other terminal devices.
[0024] The business server 100 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal devices 200a, 200b, 200c, ..., 200n) may be smart phones, tablet computers, laptop computers, desktop computers, PDAs, mobile internet devices (MIDs), wearable devices (such as smart watches, smart bracelets, etc.), smart computers, smart cars, and other smart terminals.
[0025] Among them, the business server 100 can establish a communication connection with each terminal device in the terminal cluster, and a communication connection can also be established between each terminal device in the terminal cluster. In other words, the business server 100 can establish a communication connection with each terminal device in the terminal device 200a, terminal device 200b, terminal device 200c, ..., terminal device 200n. For example, a communication connection can be established between the terminal device 200a and the business server 100. A communication connection can be established between the terminal device 200a and the terminal device 200b, and a communication connection can also be established between the terminal device 200a and the terminal device 200c. Among them, the above-mentioned communication connection does not limit the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, etc. The specific method can be determined according to the actual application scenario, and this application does not impose any restrictions on this.
[0026] It should be understood that Figure 1 Each terminal device in the terminal cluster shown can be installed with an application client. When the application client runs in each terminal device, it can interact with the business server 100 respectively, so that the business server 100 can receive business data from each terminal device (such as user identity data uploaded by the user through the terminal device). Among them, the application client can be a music player application, a karaoke software application, a browser application, a social application, an instant messaging application, a live broadcast application, a game application, a short video application, a video application, a shopping application, a novel application, a payment application, etc., which has the function of displaying text, images, audio, and video and other data information. The specific application can be determined according to the actual application scenario requirements and is not limited here. Among them, the application client can be an independent client or an embedded sub-client integrated in a client (such as a music application, a karaoke software application, etc.). The specific application can be determined according to the actual application scenario and is not limited here.
[0027] On internet platforms, the amount of user-generated content (UGC) continues to grow, and the scenarios are becoming increasingly complex. In related solutions, the review system processes all types of UGC produced / uploaded by users on the platform through algorithms such as image security check algorithms, text security check algorithms, audio security check algorithms, and content blacklist matching. Violations are immediately deleted, while content deemed risky by the algorithm is sent to a human review platform for re-examination. Content deemed safe by the algorithm is directly published on the platform. Suspected risky content is also published directly on the platform after being confirmed safe by human review.
[0028] However, this review approach still carries potential risks. For one thing, insufficient algorithmic capabilities or an inability to identify new types of illegal content can lead to risky content being missed. Furthermore, only a small portion of this massive amount of data is actively accessed by users, while the vast majority of data has minimal access. Furthermore, the focus of review is on ensuring the security of newly added UGC content daily. This makes backtracking the entire volume of UGC content extremely costly and inefficient, and fails to address the risks posed by the massive amount of existing data.
[0029] On the other hand, the UGC content uploaded by users is written into the platform's storage system through write operations. With the increasing complexity of product functions and the increase in UGC scenarios, there is a risk that UGC content is not actively reported to the review system.
[0030] In order to solve the above technical problems, a content review method of the embodiment of the present application is proposed. Figure 1 The content review method in the embodiment of the present application is described with the network framework and the application scenario of the embodiment of the present application:
[0031] See also Figure 2 In one embodiment of the content review method of the present application, the following steps are included:
[0032] 201. Obtain content requested by a user's read operation request, where the content requested by the read operation request includes content to be reviewed;
[0033] The method of this embodiment can be applied to a computer device, which can be Figure 1 The business server 100 or each terminal device in the network framework shown. In some embodiments, the computer device can implement the content review method provided in this embodiment by running a computer program. The computer device can serve as an audit system to review the content requested by the user through a read operation request. Among them, a read operation request refers to an access request for UGC content initiated by a user to the platform through a terminal when accessing the Internet platform, and the content requested to be accessed by the read operation request is the UGC content that the user requests to access the Internet platform. Among them, UGC content refers to the content form independently created, published and disseminated by users, for example, it may include text data, picture data, video data, audio data, live broadcast data and other data content, and its data type is not limited.
[0034] The content requested by the user read operation request obtained by the computer device includes content that is pending review. This content is content that needs to be reviewed for violations, such as whether it contains pornography, vulgarity, insults, violence, terrorism, or other negative content, or whether it violates relevant platform regulations.
[0035] 202. Extracting data features of the content to be reviewed, and determining a probability of random inspection of the content to be reviewed based on the data features of the content to be reviewed;
[0036] 203. Determine, among the plurality of pieces of content to be reviewed, the content to be sampled based on the probability of sampling the content to be reviewed;
[0037] Computer equipment can extract data features from the content to be reviewed and determine the probability of random inspection of the content to be reviewed based on the data features of the content to be reviewed. The data features of the content to be reviewed are used to characterize the characteristics or attributes of the content to be reviewed, and other information related to the content to be reviewed. For example, data features can be information that characterizes the uploader of the content to be reviewed, the device information used by the uploader to publish the content, or the type of content to be reviewed, etc. There is a certain correlation between these data features and the likelihood of violation of the content to be reviewed. Therefore, the probability of random inspection of the content to be reviewed can be determined based on the data features of the content to be reviewed. That is, the greater the likelihood of violation of the content to be reviewed corresponding to the data features, the greater the probability of random inspection of the content to be reviewed; conversely, the smaller the probability of random inspection of the content to be reviewed.
[0038] For example, the uploader of the pending content may have a high violation rate among the multiple contents uploaded in the past, indicating that the content uploaded by the uploader is more likely to violate the rules. Therefore, the content currently uploaded by the uploader may be given a higher probability of random inspection so that the content uploaded by the uploader can be randomly inspected and reviewed as much as possible.
[0039] The computer device, acting as an audit system, can receive a large number of read operation requests from users and, based on these read operation requests, identify multiple pieces of content to be audited from the requested content. Among the multiple pieces of content to be audited, the computer device can determine which content to audit based on the audit probability. The greater the audit probability of a piece of content, the more likely it is to be audited.
[0040] For example, a computer device determines 100 pieces of content to be reviewed from content requested for access by multiple read operation requests. If a preset sampling probability of greater than 60% is required for sampling, the content to be reviewed with a sampling probability greater than 60% can be determined from the 100 pieces of content to be reviewed as the sampling content for inspection and review. If a preset sampling probability of the top 70% of a batch of content to be reviewed with the highest sampling probability is required for inspection, the top 70 pieces of content with the highest sampling probability can be determined from the currently determined 100 pieces of content to be reviewed for inspection and review. This embodiment does not limit the specific method for determining the sampling content from the multiple pieces of content to be reviewed based on the sampling probability, as long as the higher the sampling probability, the more likely the content to be reviewed is to be inspected and reviewed.
[0041] The sampling probability of each content to be reviewed is determined by the data characteristics of each content to be reviewed, and the sampling content to be reviewed is determined from multiple content to be reviewed based on the sampling probability of the content to be reviewed. This allows UGC content with various characteristics to have a certain probability of being sampled. For example, frequently accessed high-hot data can be given a higher sampling probability, making it more likely to be sampled, thereby allocating review resources according to data popularity.
[0042] 204. Identify the sample inspection content according to the content characteristics of the sample inspection content to obtain a sample inspection content identification result;
[0043] After determining the sample inspection content, the sample inspection content can be identified based on its content features to obtain a sample inspection content identification result. This content feature is used to represent characteristic information of the substantive content of the sample inspection content, which can be received and acquired by users visually or aurally. For example, it can be text in text content, images / captions / sound in video content, text in image content, or image information represented by images, etc. Therefore, by identifying the content features of the sample inspection content, it can be determined whether the sample inspection content contains illegal information that cannot be received and acquired by users.
[0044] The spot check content identification result may indicate that the spot check content is compliant, or that the spot check content is in violation of regulations. The spot check content identification result may be further sent for manual review, and personnel may review the content identified as in violation by the computer device. Content identified as compliant by the computer device does not need to be sent for manual review. For example, when the computer device, as an image review platform, identifies that the image content is in violation of regulations, it may output the review result of the image violation for further confirmation by the image reviewer; when the computer device, as a text review platform, identifies that the text content is in violation of regulations, it may output the review result of the text content violation for review and confirmation by the text content reviewer; when the computer device, as an audio review platform, identifies that the audio content is in violation of regulations, it may output the review result of the audio content violation for review and confirmation by the audio reviewer; when the computer device, as a video review platform, identifies that the video content is in violation of regulations, it may output the review result of the video content violation for review and confirmation by the video reviewer; when the computer device, as a multimedia data review platform, identifies that the multimedia data content is in violation of regulations, it may output the review result of the multimedia data content violation for review and confirmation by the multimedia data reviewer; when the computer device, as a live data review platform, identifies that the live data content is in violation of regulations, it may output the review result of the live data content violation for review and confirmation by the live data reviewer.
[0045] In this embodiment, the computer device obtains the content requested to be accessed by the user's read operation request, and the content requested to be accessed by the read operation request includes the content to be reviewed, extracts the data features of the content to be reviewed, and determines the probability of random inspection of the content to be reviewed based on the data features of the content to be reviewed. Among the multiple contents to be reviewed, the random inspection content among the multiple contents to be reviewed is determined based on the random inspection probability of the content to be reviewed, and the random inspection content is identified based on the content features of the random inspection content to obtain the random inspection content identification result, which is used for manual review. Therefore, by reviewing the content of the read operation request, the content review business does not need to be actively reported, so it is less invasive to the business, reduces the risk of content underreporting and underreview, and can cover a large amount of read operation content, thereby improving content review efficiency.
[0046] The following will be Figure 1 The network framework shown and Figure 2 Based on the embodiment shown, the embodiment of the present application is further described in detail. Figure 3 Another embodiment of the content review method in the embodiment of the present application includes:
[0047] 301. Obtain content requested by a user's read operation request, where the content requested by the read operation request includes content to be reviewed.
[0048] In this embodiment, the computer device serves as the audit system. This audit system can be an audit system deployed within the Internet platform where users upload UGC content, or it can be an audit system independent of the Internet platform. The service gateway of the Internet platform can forward the network requests (req) and network responses (rsp) of the user's large number of read operation requests to the computer device serving as the audit system through bypass forwarding. The computer device can then obtain the content requested by the large number of read operation requests.
[0049] The computer device may further pre-process the content requested for access by a large number of read operation requests. An optional pre-processing method is that the computer device may obtain the user's read operation request and remove the read operation request for non-creative content from the read operation request, thereby obtaining the creative content requested for access by the read operation request for creative content. The access time of multiple creative contents falls within multiple first time windows. Furthermore, duplicate creative contents within each first time window may be deduplicated, and the non-duplicate creative contents retained within each first time window may be determined as content to be reviewed.
[0050] Read requests for non-generated content, for example, can be requests to read (query) user permissions, account levels, song listening time, video viewing time, and other read operations unrelated to UGC. Conversely, generated content is human-generated substantive content, such as text, audio, and video content.
[0051] For example, assuming a computer device obtains content requested by a user read operation within a 5-minute time window between 17:00 and 17:05, the computer device can deduplicate the content requested by the user within that time window, removing duplicate content. Duplicate content could be the same content accessed by multiple different users within the time window, which can be deduplicated using the above mechanism.
[0052] Therefore, by eliminating read operation requests for non-original content and the above-mentioned deduplication mechanism, the amount of content can be greatly reduced, the burden of content review can be reduced, and review resources can be saved.
[0053] 302. Extracting data features of the content to be reviewed, and determining a probability of random inspection of the content to be reviewed based on the data features of the content to be reviewed;
[0054] In this embodiment, the probability of random inspection of the content to be reviewed is determined based on the data characteristics of the content to be reviewed. An optional implementation method is to adopt a rule-based random inspection strategy. If the first data characteristic of the content to be reviewed matches the preset data characteristic, then the probability of random inspection of the content to be reviewed is determined to be greater than the preset random inspection probability; if the first data characteristic of the content to be reviewed does not match the preset data characteristic, then the probability of random inspection of the content to be reviewed is determined to be less than the preset random inspection probability.
[0055] Among them, when adopting a rule-based sampling strategy, the first data feature of the content to be reviewed refers to the feature related to the user or device requesting access to the content to be reviewed, or the feature information related to the access matters of the content to be reviewed. There is a certain correlation between these feature information and the possibility of violation of the content to be reviewed. For example, it may include but is not limited to: the user ID of the user who initiates the read operation request for the content to be reviewed, the device ID of the device that initiates the read operation request for the content to be reviewed, the source IP address of the user who initiates the read operation request for the content to be reviewed, the access frequency of the content to be reviewed, or the access time of the content to be reviewed. Among them, IP refers to the Internet Protocol address, which is used to identify devices in the network.
[0056] For example, in practice, it has been found that there are certain patterns in the browsing of illegal content on the platform, such as some users are easily deceived and browse more fraudulent content; and in certain time periods, users browse more illegal content. Therefore, based on this prior knowledge, random inspection rules can be formulated, for example, including:
[0057] Build and maintain a database of high-risk user IDs. These users are prone to accessing illegal content, which can increase the probability of random inspections for content accessed by these users, thereby increasing the probability of their accessed content being inspected. Therefore, the user ID of the user who initiated the read operation request for the content to be reviewed can be extracted. If this user ID matches a user ID in the high-risk user ID database, the probability of random inspection for the content requested by this user is determined to be greater than the preset random inspection probability, thereby increasing the probability of random inspection of the content.
[0058] A high-risk device ID library can also be built and maintained, which can extract the device ID of the device that initiated the read operation request for the content to be reviewed. If the device ID matches a device ID in the high-risk device ID library, the probability of random inspection of the content accessed by the device is determined to be greater than the preset random inspection probability, thereby increasing the probability of random inspection of the content accessed by the device. This strategy can be used for devices with no user account logged in or devices with multiple user accounts logged in simultaneously. In these two cases, the random inspection probability of the content to be reviewed cannot be determined based on the user ID. In this case, the random inspection probability of the content to be reviewed can be determined based on the device ID.
[0059] A high-risk IP library can also be built and maintained. For example, the high-risk IP library can record overseas IPs, and the source IP that initiates the read operation request for the content to be reviewed can be extracted. If the source IP matches an IP in the high-risk IP library, it can be determined that the probability of random inspection of the content requested by the source IP is greater than the preset random inspection probability, thereby increasing the probability of the content being randomly inspected.
[0060] In other scenarios, in the field of content security, unpopular content with low access frequency is often more prone to risks (popular content is often safer because it is often inspected or reported), but it is difficult to pick out unpopular content if random inspections are conducted. Therefore, the access frequency of the content to be reviewed can also be extracted. If the access frequency of the content to be reviewed is less than the preset access frequency, the content to be reviewed is determined to be unpopular content, and the probability of random inspection of the content to be reviewed can be determined to be greater than the preset random inspection probability, thereby increasing the chance of the content to be reviewed being inspected. Therefore, through this strategy, the probability of random inspection of unpopular content with low access volume can also be increased, making it more likely to be inspected.
[0061] In other scenarios, if the access time of the content to be reviewed matches a preset time period, the probability of random inspection for the content to be reviewed can be determined to be greater than the preset random inspection probability. For example, during late night hours, pornographic and vulgar content is more likely to be viewed. Therefore, the random inspection probability of content accessed by read requests during this time period can be increased, making it more likely to be inspected to eliminate risky content.
[0062] In practice, the above-mentioned high-risk user ID library, high-risk device ID library and high-risk IP library can be dynamically configured according to the audit results of the content audit business, and the above-mentioned preset access frequency and preset time period can be dynamically changed. The preset sampling probability under each strategy can also be adjusted according to the needs of the audit business.
[0063] Another optional implementation involves using a model-based sampling strategy, inputting the second data feature of the content to be reviewed into a pre-trained model to obtain the sampling probability of the content to be reviewed predicted by the pre-trained model based on the second data feature of the content to be reviewed. The pre-trained model is trained using a machine learning algorithm based on the data features and labels of the content in each set of training samples. The content label is used to indicate the review result of whether the content violates the rules.
[0064] Among them, when adopting a model-based sampling strategy, the second data feature of the content to be reviewed refers to feature information related to the creator of the content to be reviewed, the visitors of the content to be reviewed, or the content type and theme of the content to be reviewed. There is a certain correlation between these feature information and the possibility of violation of the content to be reviewed. For example, it may include but is not limited to: the historical violation rate of the creator of the content to be reviewed, the preferences of the visitors of the content to be reviewed and the proportion of the visitors who access the illegal content, the content label of the content to be reviewed, the IP risk level of the user who initiates the read operation request for the content to be reviewed, and the device risk level of the user who initiates the read operation request for the content to be reviewed.
[0065] In a model-based spot check strategy, the model predicts the probability of spot checks for content under review based on a wealth of data features. For example, the model can utilize machine learning models or deep learning models, such as the XGBoost algorithm, deep neural networks, decision trees, SVMs, and logistic regression models. These data features are extracted from the content and used as training samples along with the content's labels. The model then learns the correlation between the data features of the content under review and its likelihood of violation. After the model is trained, it can be used to identify the data features of the content under review to predict the probability of spot checks for the content under review.
[0066] Therefore, through the model-based sampling strategy, the complex patterns hidden in the big data of compliant content and non-compliant content can be mined to improve the accuracy of identifying content to be reviewed; and there is no need to configure various data feature libraries (such as the above-mentioned high-risk user ID library, high-risk device ID library, etc.), reducing configuration costs and maintenance costs.
[0067] 303. Determine, among the plurality of pieces of content to be reviewed, the content to be sampled based on the probability of sampling the content to be reviewed;
[0068] In this embodiment, when determining the sampled content from multiple pieces of content to be reviewed based on the sampled probability of the content to be reviewed, candidate sampled content from the multiple pieces of content to be reviewed can be determined based on the sampled probability of the content to be reviewed. The access times of the multiple candidate sampled content fall within multiple second time windows. Furthermore, duplicate candidate sampled content within each second time window can be deduplicated, and the non-duplicate candidate sampled content retained within each second time window can be determined as the sampled content.
[0069] This second time window can be longer than the first time window described above, so as to further remove more duplicate content from the sampled content. For example, using one day as a time window, candidate sampled content with access times within a certain date can be identified from multiple candidate sampled content. The computer device can then deduplicate the candidate sampled content within that date, i.e., remove duplicate candidate sampled content, thereby reducing the data size of the content that needs to be sampled, further reducing the burden of content review, and saving review and processing resources of the review system.
[0070] 304. Identify the sample inspection content according to the content characteristics of the sample inspection content to obtain a sample inspection content identification result;
[0071] In this embodiment, the sample inspection content determined in the above steps can be audited and filtered through a variety of audit and filtering mechanisms to obtain compliant content and illegal content. One optional way to identify and audit the sample inspection content is to determine whether the words in the sample inspection content match the preset keywords. If so, it is determined that the sample inspection content is illegal content; if not, it is determined that the sample inspection content is not illegal content. The preset keywords can be sensitive words that touch on pornography, vulgarity, insults, violence and terrorism and other bad behaviors. By matching keywords, it is possible to quickly determine whether the sample inspection content is illegal, and the audit results can be more accurate.
[0072] When determining whether the words in the sampled content match the preset keywords, if the sampled content is text data, it is possible to directly determine whether the words in the sampled content match the preset keywords; if the sampled content is audio data or video data, the audio data or video data can be converted into text data first, and then the words in the text data converted from the audio data or video data can be determined whether they match the preset keywords.
[0073] The method of reviewing the sampled content based on preset keywords is low-cost and fast, and can be used on a large scale. Keywords can also be dynamically updated according to actual review business and changes in sensitive words to comprehensively and effectively identify the sampled content in various scenarios, thereby improving the accuracy and efficiency of content review.
[0074] Another alternative method for identifying and reviewing sampled content is to use computer recognition to obtain a content recognition model. This content recognition model can be trained using a machine learning algorithm based on the content features and content recognition results of each set of training samples. The content recognition results can be the results of manual review to determine whether the content is illegal. Furthermore, the content features of the sampled content can be input into the content recognition model to obtain the sampled content recognition results predicted by the content recognition model based on the content features of the sampled content.
[0075] For example, text models can be deployed to review and identify text-based spot checks. A text model (TextModel) refers to a machine learning model specifically designed to process and understand text data. Text models can include traditional machine learning models such as naive Bayes models, support vector machine models, and logistic regression models; they can also include deep learning models such as recurrent neural networks, long short-term memory networks, gated recurrent units, and Transformer models. Using text models to identify and filter spot checks is low-cost, quicker to review, and has stronger recall capabilities.
[0076] For example, large language models can be deployed to review and identify text-based random inspection content. Large language models (LLMs) are deep learning-based natural language processing models capable of understanding and generating human language. These models, based on the Transformer architecture and trained on large amounts of text data, possess powerful language understanding and generation capabilities. Using large language models to identify and filter random inspection content offers strong recall capabilities, enabling the identification of additional illegal content. They also offer the ability to identify new types of illegal or rare content, and possess enhanced generalization capabilities.
[0077] Therefore, by deploying a content recognition model to identify and filter the sampled content, we can accurately identify the illegal content in the sampled content, and can accurately identify new and rare illegal content, thereby covering a wider and more comprehensive range of illegal content and improving content review efficiency.
[0078] 305. Among the sampled content for which the sampled content identification result indicates compliance, the portion of the sampled content determined to be compliant is used for manual review;
[0079] To further ensure the accuracy and reliability of content review, this embodiment incorporates a supplementary sampling mechanism. Specifically, within the sampled content identified as compliant, a portion of the sampled content can be identified as compliant for manual review. This supplementary sampling mechanism can compensate for computer equipment's misidentification of illegal content as compliant, thereby improving the accuracy of content review overall and preventing users from accessing illegal content through missed content.
[0080] This embodiment, through a read-based content sampling model, can identify vulnerabilities in traditional write-based content review mechanisms, driving the optimization of content security algorithms, systems, and policies. In practice, after the review system went live, it discovered numerous issues, including violations in large amounts of historical data, content review errors caused by write operations without reporting new data, and security service vulnerabilities.
[0081] Secondly, as a risk observation mechanism, it can detect changes in risk across the entire platform and promote the optimization of content security policies. For example, in practice, the review mechanism based on write-based content deactivated a long audio review policy, resulting in an increase in platform risk. However, the spot check interface for read-based content can quickly detect an increase in the number of illegal content being sampled. Therefore, the method of this embodiment can promote the optimization of content review policies.
[0082] For example, Figure 4 An application scenario of this embodiment is shown. In the scenario shown in the figure, the service gateway of the Internet platform can forward the network requests (request, req) and network responses (response, rsp) of a large number of read operation requests of users to a computer device serving as an audit system through bypass forwarding, and the computer device can obtain the content requested to be accessed by the large number of read operation requests.
[0083] The computer device may further pre-process the content accessed by the large number of read operation requests. This pre-processing may include data cleaning and deduplication. For example, read operation requests for non-creative content may be removed from the large number of read operation requests, while read operation requests for creative content may be retained, or duplicate creative content repeated within the same time window may be deduplicated. After pre-processing is completed, the retained content may be prepared for review.
[0084] Afterwards, rule-based spot checks and model-based spot checks can be conducted on the content to be reviewed, respectively, to determine the spot-checked content from multiple pieces of content to be reviewed. Among them, key monitoring sources can be configured, and key monitoring sources are configured with monitoring indicators such as high-risk UID (User Identification), high-risk IP, and high-risk equipment, so that a rule-based spot check strategy can be executed on the content to be reviewed based on the above-mentioned multiple monitoring indicators to determine whether the data characteristics of the content to be reviewed match the monitoring indicators in the key monitoring source. Among them, the specific implementation of rule-based spot checks and model-based spot checks has been described in detail in the previous article and will not be repeated here.
[0085] After the spot-checked content is determined for the content to be reviewed, the spot-checked content can be deduplicated a second time. For example, duplicate content in the spot-checked content can be deduplicated by setting a time window that is longer than the time window based on the first deduplication. Then, the spot-checked content obtained after the second deduplication is sent to keyword filtering, text model filtering, and large language model filtering in turn, and the proportion and concentration of illegal content in the spot-checked content are increased through layer-by-layer screening, exclusion, and release of compliant content. A supplementary spot-checking mechanism is also included here, which determines part of the spot-checked content from the spot-checked content identified as compliant content for manual review. Among them, the specific implementation of the above-mentioned multiple filtering algorithms and the supplementary spot-checking mechanism has been described in detail in the previous article and will not be repeated here.
[0086] After multiple filtering algorithms identify and filter out illegal content, it can be sent for further manual review and re-examination to confirm the content. Furthermore, manual review can provide feedback on new monitoring indicator data, such as the emergence of new high-risk UIDs and new high-risk IP addresses, which can be updated to key monitoring sources to continuously improve the accuracy of rule-based spot checks and ensure more comprehensive coverage of illegal content.
[0087] Therefore, based on the entire review system, on the one hand, intelligent sampling is used to reduce the amount of data in the review content and increase the proportion and concentration of illegal content; on the other hand, various technical layers of filtering are used to further increase the concentration and proportion of illegal content in the sampled content. These two aspects ensure that the review system can efficiently discover and detect more illegal content.
[0088] The computer device in the embodiment of the present application is described below. Figure 5 In one embodiment of the present application, a computer device includes:
[0089] The computer device 500 may include one or more central processing units (CPUs) 501 and a memory 505 , wherein the memory 505 stores one or more application programs or data.
[0090] Memory 505 may be volatile or persistent storage. The program stored in memory 505 may include one or more modules, each of which may include a series of instruction operations on the computer device. Furthermore, CPU 501 may be configured to communicate with memory 505 and execute the series of instruction operations in memory 505 on computer device 500.
[0091] The computer device 500 may also include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input and output interfaces 504, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0092] The CPU 501 can execute the aforementioned Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment will not be described in detail here.
[0093] The present application also provides a computer storage medium, wherein one embodiment includes: the computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the aforementioned Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment.
[0094] The present application also provides a computer program product, wherein one embodiment includes: when the computer program product is run on a computer device, the computer device executes the aforementioned Figures 2 to 3 The operations performed by the computer device in the illustrated embodiment.
[0095] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0097] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0098] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0099] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk, and other media that can store program code.
Claims
1. A content review method, characterized in that: The method comprises: Obtaining content requested by a user's read operation request, where the content requested by the read operation request includes content to be reviewed; Extracting data features of the content to be reviewed, and determining a probability of spot checking the content to be reviewed based on the data features of the content to be reviewed; Determining, among the plurality of pieces of content to be reviewed, the content to be reviewed based on the probability of random inspection of the content to be reviewed; The sample inspection content is identified according to the content characteristics of the sample inspection content to obtain a sample inspection content identification result.
2. The method according to claim 1, characterized in that The data features of the content to be reviewed include the first data feature and / or the second data feature; Determining the probability of spot checking the content to be reviewed based on the data characteristics of the content to be reviewed includes: If the first data feature of the content to be reviewed matches the preset data feature, then the probability of random inspection of the content to be reviewed is determined to be greater than the preset random inspection probability; if the first data feature of the content to be reviewed does not match the preset data feature, then the probability of random inspection of the content to be reviewed is determined to be less than the preset random inspection probability; and / or, The second data feature of the content to be reviewed is input into a pre-trained model to obtain the sampling probability of the content to be reviewed predicted by the pre-trained model based on the second data feature of the content to be reviewed; wherein, the pre-trained model is obtained by training the data features of the content in each group of training samples and the label of the content by a machine learning algorithm, and the label of the content is used to indicate the review result of whether the content is in violation of regulations.
3. The method according to claim 2, characterized in that The first data feature of the content to be reviewed includes: a user identifier of a user initiating a read operation request for the content to be reviewed, a device identifier of a device initiating a read operation request for the content to be reviewed, a source IP address of a device initiating a read operation request for the content to be reviewed, and a frequency of access to the content to be reviewed or a time of access to the content to be reviewed; and / or, The second data features of the content to be reviewed include: the historical violation rate of the creator of the content to be reviewed, the visitor preferences of the content to be reviewed and the proportion of visitors who access illegal content, the content tags of the content to be reviewed, the IP risk level of the user who initiates the read operation request for the content to be reviewed, and the device risk level of the user who initiates the read operation request for the content to be reviewed.
4. The method according to claim 1, wherein The step of identifying the sampled content according to the content features of the sampled content to obtain a sampled content identification result includes: Determining whether the words in the sampled content match preset keywords; If so, it is determined that the sample inspection content is illegal; If not, it is determined that the sample inspection content does not constitute illegal content.
5. The method according to claim 1, wherein The step of identifying the sampled content according to the content features of the sampled content to obtain a sampled content identification result includes: Obtaining a content recognition model, wherein the content recognition model is obtained by training the content features and content recognition results of the content in each set of training samples using a machine learning algorithm; The content features of the sampled content are input into the content recognition model to obtain the sampled content recognition result predicted and output by the content recognition model based on the content features of the sampled content.
6. The method according to claim 1, characterized in that The obtaining of the content requested by the user's read operation request includes: Obtaining a read operation request from a user, and removing read operation requests for non-creative content from the read operation request, thereby obtaining the creative content requested for access by the read operation request for creative content; wherein access times of the plurality of creative contents fall within a plurality of first time windows; The repeated creative contents within each of the first time windows are deduplicated, and the non-repeated creative contents retained within each of the first time windows are determined as the content to be reviewed.
7. The method according to claim 1, characterized in that The step of determining the sampling contents of the plurality of contents to be reviewed based on the sampling probability of the contents to be reviewed includes: Determining candidate random inspection contents from the plurality of contents to be reviewed based on the random inspection probability of the contents to be reviewed; wherein the access time of the plurality of candidate random inspection contents falls within a plurality of second time windows; The candidate sampling contents repeated in each second time window are deduplicated, and the candidate sampling contents that are not repeated and retained in each second time window are determined as the sampling contents.
8. The method according to any one of claims 1 to 7, characterized in that After obtaining the spot check content identification result, the method further includes: Among the sampled contents whose identification results indicate that the contents are compliant, the portion of the sampled contents determined to be compliant is used for manual review.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed on a computer, enable the computer to perform the method according to any one of claims 1 to 8.