Method for improving manual auditing efficiency of text content
By using group hashing and unique identification technology, large-scale text content is classified, stored, and deduplicated, and historical review results are automatically inherited. This solves the problem of low efficiency in traditional manual review and achieves efficient and accurate text content review.
Patent Information
- Application Number
- CN202511438133.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Traditional manual review methods are inefficient and AI-powered intelligent review has the problem of misjudgment, which makes it impossible to complete the review of large-scale text content within a reasonable time. In particular, the processing of duplicate content is inefficient and there is a risk of misjudgment.
By employing group hashing and unique identifier technologies, text content is classified, stored, and deduplicated. It automatically inherits historical review results, merges similar content through group hash values, and displays only representative data for batch review. This combination of automatic and manual review improves efficiency and accuracy.
It significantly improved the efficiency of manual review, shortened the review time by 97%, ensured the consistency and accuracy of review conclusions, reduced repetitive work, and lowered operational risks.
Smart Images

Figure CN120910036A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and particularly relates to a method for improving the efficiency of manual review of text content. BACKGROUND
[0002] The core efficiency bottleneck of current large-scale text content review is that when the amount of data to be reviewed reaches hundreds of thousands or even millions, the traditional manual review method cannot complete the work within a reasonable time, especially for scenarios such as reprinted content, repeated sensitive information, attachment content or the same external link, the reviewer needs to repeatedly process the same content, resulting in a lot of invalid labor.
[0003] Although the existing review technology uses keyword filtering or simple hash deduplication to some extent, there are still some deficiencies. First, the traditional manual review method is inefficient, and when the amount of data to be reviewed reaches hundreds of thousands or even millions, it is basically impossible to perform normal review, and the reviewer needs to process the same or similar content one by one, resulting in a lot of repetitive labor. Second, the existing AI intelligent review method has the problem of false review, and cannot meet the requirement of 100% accuracy. Once a problem occurs, it may lead to the notification of the client being reviewed by the superior agency.
[0004] Therefore, in view of the above problems, the present application provides a method for improving the efficiency of manual review of text content. SUMMARY
[0005] In order to overcome the problem of low efficiency of the traditional manual review method and the problem of false review of the existing AI intelligent review method, the present application provides a method for improving the efficiency of manual review of text content. By combining classification storage, grouping hash and automatic inheritance, the reviewer can focus on the differentiated content that really needs manual judgment, and the efficiency of manual review is significantly improved.
[0006] The technical scheme of the present application is as follows: a method for improving the efficiency of manual review of text content, comprising the following steps: S1, detecting the text content to be reviewed to generate a detection result; S2, for each detection result, calculating a grouping hash value according to its detection type and key information, which is used to group similar content into the same group; S3, for each detection result, calculating a unique identifier according to its detection type, media identifier, web link and key information, which is used for deduplication processing; S4, before storing the detection result, querying the database to determine whether there is an identical record according to the unique identifier, if there is, discarding the current detection result, if not, continuing to determine whether there is a historical manual review result that can be inherited; S5, if there is an inheritable historical audit result, automatically inherit the latest manual audit result, including the audit state, the risk type after the audit and the risk level after the audit, and the auditor is marked as system, and the audit time is set as the current system time; S6, store the detection result, the grouping hash, the unique identifier and the inherited audit result in the database; S7, in the manual audit link, the de-duplication display is performed on the to-be-audited data according to the grouping hash, and the batch audit is performed on the de-duplicated data by the auditor; S8, after the audit is completed, the audit state of all detection results under the grouping hash is updated.
[0007] As preferred, the calculation manner of the grouping hash is different according to different detection types, including: For sensitive information, the "detection type + sentence + sensitive word + recommended word" is calculated by hash; For privacy information, the "detection type + sentence + privacy content" is calculated by hash; For sensitive person, the "detection type + sentence + name + position" is calculated by hash; For risk external link, the "detection type + external link" is calculated by hash; For custom word, the "detection type + sentence + custom word + recommended word" is calculated by hash.
[0008] As preferred, the calculation manner of the unique identifier is different according to different detection types, including: For sensitive information, the "detection type + media Id + web link + sentence + sensitive word + recommended word" is calculated by hash; For privacy information, the "detection type + media Id + web link + sentence + privacy content" is calculated by hash; For sensitive person, the "detection type + media Id + web link + sentence + name + position" is calculated by hash; For risk external link, the "detection type + media Id + web link + external link" is calculated by hash; For custom word, the "detection type + media Id + web link + sentence + custom word + recommended word" is calculated by hash.
[0009] As preferred, the step S5 specifically includes: S51, according to the grouping hash of the current detection result, query the database for the existing manual audit record; S52, arrange in reverse order according to the audit time, and select the latest manual audit result for inheritance; S53, set the audit state, the risk type after audit, and the risk level after audit of the current detection result to the corresponding values of the inherited record; S54, mark the auditor as system, and set the audit time as the current system time, for distinguishing whether the partition is automatically audited by the system or manually audited by the human.
[0010] Preferably, the audit state is set to the corresponding field value of the last manual audit record, the risk type after audit is set to the corresponding field value of the last manual audit record, and the risk level after audit is set to the corresponding field value of the last manual audit record.
[0011] Preferably, the batch audit step in step S7 comprises: S71, the auditor queries the data to be audited according to the detection type; S72, the system displays the data to be audited after deduplication according to the grouping hash; S73, the auditor selects one or more representative records for audit operation; S74, the system marks all detection results in the group as audited according to the grouping hash of the selected record.
[0012] Preferably, for the detection results of the risk external link type, three batch audit modes are supported during audit, which are audit by main domain name, audit by Host, and audit by link, wherein the audit by main domain name is to audit all external link detection results under the same main domain name together, the audit by Host is to audit all external link detection results under the same Host together, and the audit by link is to audit according to the grouping hash, which is consistent with other detection types.
[0013] Preferably, the detection result comprises one or more of sensitive information, private information, sensitive person, risk external link, and custom word, but is not limited thereto.
[0014] Preferably, the audit state comprises two states of "confirming that it is a problem that needs to be modified" and "confirming that it is not a problem that needs to be modified".
[0015] Preferably, the risk type and risk level after audit are dynamically set according to the specific detection type and audit result.
[0016] The beneficial effects of the present application are: 1. The present application intelligently merges massive text content to be audited through the grouping hash mechanism, so that the auditor does not need to process repeated or highly similar content one by one, but only needs to audit the representative data after deduplication by grouping, and when processing data of the order of 100,000, the audit time can be shortened from tens of hours in the traditional way to several hours, and the efficiency can be improved by up to 97%, thereby solving the industry problem that large-scale text audit tasks are difficult to complete within a reasonable time.
[0017] 2. The application creates an automatic inheritance mechanism of audit results, ensuring the consistency of audit conclusions. The system can automatically query and apply historical artificial audit results to newly detected similar problems based on grouping hash, which not only avoids accidental deviations that may occur when the same auditor makes repeated judgments on the same content, but also eliminates different rulings that different auditors may make on essentially the same problem, thereby greatly improving the accuracy of the overall audit work and reducing the operational risks caused by misjudgment or inconsistent standards.
[0018] 3. The application realizes accurate deduplication of detection results through unique identification (unique_id), avoiding repeated detection results into the database due to multiple scans of the same media and the same page. This mechanism reduces the total amount of the list to be audited from the data source, so that each record displayed to the auditor represents an independent and unique specific problem, further reducing unnecessary repetition.
[0019] 4. The auditor can select the most suitable batch processing combination according to the risk characteristics and distribution of external links, realizing efficient disposal of a large number of associated external links under the same domain. This not only greatly simplifies the operation process, but also ensures the consistency of risk control for the entire site or batch of associated external links. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 The workflow diagram of the application is shown.
[0021] Figure 2 The audit link process diagram of the application is shown. DETAILED DESCRIPTION
[0022] To make the purpose, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0023] The present application provides an embodiment: a method for improving the efficiency of manual review of text content, which classifies and stores the detection results of text content, calculates the group hash and unique id for each detection result and saves them together with the detection results, removes duplicates according to the unique id during the storage of the detection results, and automatically matches and inherits the historical manual review results according to the group hash, avoiding the need for manual repeated review of the same problem. In addition, when processing the data to be reviewed, the reviewer no longer needs to display all the original data to be reviewed, but only needs to display the data after deduplication according to the group hash, that is, only one piece of data is displayed for the same group of data, and one piece of data is reviewed to review a batch of data, realizing fast manual review.
[0024] The present application relies on the following text content review system implementation, specifically: A content detection engine is used to preliminarily and automatically detect the original text content (such as web articles, attachment texts, user comments, etc.) grabbed or submitted, and identify detection results that may contain sensitive information, private information, sensitive persons, risk external links or custom words.
[0025] A data processing and calculation module is responsible for receiving the preliminary detection results from the detection engine and performing the calculation logic of group hash (group_hash) and unique identification (unique_id).
[0026] A database is used to store all detection results and their related metadata, including at least the following fields: id (auto-increment primary key), media_id (media identification), page_url (web link), detection_type (detection type), content_snippet (sentence or content snippet), keyword (sensitive word / privacy content / name / external link / custom word), suggested_word (recommended word, if applicable), position (position, if applicable), group_hash (group hash value), unique_id (unique identification), review_status (review status), reviewed_risk_type (reviewed risk type), reviewed_risk_level (reviewed risk level), reviewer (reviewer), review_time (review time), create_time (creation time).
[0027] An audit operation module provides a user interface for the reviewer to display the list of data to be reviewed after deduplication by group, and provides a batch review operation interface.
[0028] The auditing result inheritance and updating module executes automatic inheritance logic in a data storage link and executes batch updating logic after manual auditing.
[0029] The present application provides an embodiment for detailing the workflow of the present application: The workflow of the present method mainly includes two links, including a detection result storage link and a manual auditing link.
[0030] Please refer to Figure 1 , first, the detection result storage link is described, this link is to clean, deduplicate and pretreat the massive original results generated by the detection engine, to lay a foundation for efficient manual auditing: In the first step, after the data processing module receives a detection result, it first calculates the group_hash and unique_id according to the detection_type (detection type) of the detection result, and the calculation adopts a standard hash function to convert the spliced string into a fixed-length hash value.
[0031] Among them, the calculation rule of the group_hash is described, which is used for content grouping and batch auditing: Sensitive information: calculate hash("Sensitive information"+content_snippet+keyword+suggested_word). This rule ensures that the same sentence fragments containing the same sensitive words and recommended words appearing in different pages will be grouped into the same group.
[0032] Private information: calculate hash("Private information"+content_snippet+keyword). This rule ensures that the same sentence fragments containing the same private content (such as ID number, phone number) appearing in different pages will be grouped into the same group.
[0033] Sensitive person: calculate hash("Sensitive person"+content_snippet+keyword+position). This rule ensures that the same sentence fragments reporting the same person (same name and position) appearing in different pages will be grouped into the same group.
[0034] Risk external link: calculate hash("Risk external link"+keyword). This rule ensures that all results detected with the same URL external link will be grouped into the same group.
[0035] Custom word: calculate hash("Custom word"+content_snippet+keyword+suggested_word). This rule is similar to sensitive information, which ensures that the same custom word appearing in the same context is grouped into a group.
[0036] Sensitive information: compute hash("Sensitive information" + media_id + page_url + content_snippet + keyword + suggested_word).
[0037] Private information: compute hash("Private information" + media_id + page_url + content_snippet + keyword).
[0038] Sensitive person: compute hash("Sensitive person" + media_id + page_url + content_snippet + keyword + position).
[0039] Risk external link: compute hash("Risk external link" + media_id + page_url + keyword).
[0040] Custom word: compute hash("Custom word" + media_id + page_url + content_snippet + keyword + suggested_word).
[0041] The unique_id incorporates the media_id and page_url, thus determining a specific question on a specific page under a certain media.
[0042] Step 2: The system queries the database using the computed unique_id as a condition. If there already exists a record with the same unique_id in the database, it indicates that the same specific question on the same page under the same media has already been detected, and the current result is a duplicate detection (e.g., multiple scans of the same page). Therefore, the result is discarded, and the process ends.
[0043] Step 3: If there does not exist, the system queries all records in the database with the same detection_type and existing manual review results using the group_hash of the current detection result as a condition. The query results are sorted in descending order of review_time.
[0044] Step 4: If a record is found in Step 3, the most recent one (i.e., the latest manual review result) is selected.
[0045] Step 5: The review-related fields of the current detection result are set to the values of the inherited record: review_status = the review_status of the record being inherited (e.g. "Confirm Issue" or "Confirm No Issue"); reviewed_risk_type = the reviewed_risk_type of the record being inherited (e.g. an external chain being reviewed as "Gambling"); reviewed_risk_level = the reviewed_risk_level of the record being inherited (e.g. private information being reviewed as "High Risk"); reviewer ='system' (explicitly marking this as a system automatic review); review_time = the current system time; Step 6: Then store the current detection results (already containing inherited review results) into the database.
[0046] Step 7: If no historical manual review records are found in Step 3, directly execute Step 6 to store the current detection results (audit-related fields are empty) into the database.
[0047] Please refer to Figure 2 Further, the manual review link is described, which is a process directly participated by the reviewer, and its interface and operation logic are optimized based on the group_hash generated in the storage link: Step 1: The reviewer logs in to the review system and enters the manual review workbench.
[0048] Step 2: The reviewer first selects the filtering conditions, such as querying by detection_type (sensitive information, private information, etc.). The system operation is not simply listing all records with review_status as "unreviewed", but grouping the unreviewed data by group_hash and detection_type, then only showing one representative record in each group, for example, the latest or earliest one, and the interface will clearly display the problem content corresponding to the group_hash and the total number of records to be reviewed in the group.
[0049] Step 3: The reviewer browses the de-duplicated list and can select one or more representative records, then clicks the "review" button.
[0050] Step 4: The system determines the review mode according to the detection_type of the selected records.
[0051] Step 5: If it is a risk external chain, the system provides three batch review granularity options for the reviewer: According to the main domain name auditing: the system parses the main domain name (such as the main domain name of https: / / job.xidian.edu.cn / is xidian.edu.cn) of the external link, and after the auditing personnel confirm the operation, the system will update all the detection results of the external link main domain name in the database which are the same as the main domain name and in the un-audited state to the results of this auditing.
[0052] According to the Host auditing: the system parses the Host of the external link (the external link is https: / / job.xidian.edu.cn / campus / view / id / 751457, and the corresponding Host is job.xidian.edu.cn, and when batch auditing, the external links with the same Host such as https: / / job.xidian.edu.cn / campus / view / id / 751457 and https: / / job.xidian.edu.cn / campus / view / id / 751456 will be audited together), After the auditing personnel confirm the operation, the system will update all the detection results of the external link Host in the database which are the same as the Host and in the un-audited state to the results of this auditing.
[0053] According to the link auditing: the mode is consistent with the group_hash auditing logic, after the auditing personnel confirm the operation, the system will update all the detection results of the group_hash in the database which are the same as the selected record and in the un-audited state to the results of this auditing.
[0054] Step 6, if it is not a risk external link type, the system directly adopts the batch auditing of the group_hash, after the auditing personnel make the auditing decision and confirm, the system will update all the detection results of the group_hash in the database which are the same as the selected record, the detection_type is the same and in the un-audited state to the results of this auditing.
[0055] After any batch auditing operation is successfully executed, the system page will automatically refresh, and all the records of the one or more groups just audited will disappear from the un-audited list.
[0056] Step 7, the auditing personnel judge whether to continue auditing, if yes, return to step 2, query and audit new groups, if no, the process is terminated.
[0057] Further, the application provides example 1: This embodiment is applied in the Xi'an Boda software content security scanning and monitoring platform, and takes the supervision of 1154 websites under Xiamen University as an example, and the following table compares the auditing workload and time consumption of the full scanning results before and after the application of the application. Table 1 Comparison between traditional method and the present application
[0058] According to Table 1, taking the "risk outside chain" as an example, more than 840,000 original records are left after grouping and hashing merging, only more than 28,000 unique groups, and the auditor only needs to process these 28,000 groups of data, the workload is directly reduced by 97%, and the total audit time is reduced from about 992 hours to about 125 hours, the efficiency is obviously improved.
[0059] Further, the present application provides Example 2: This embodiment is based on a large news portal website, and the compliance of the articles on the platform is audited, and a large number of articles are reprinted from other media. The same article reprinted using the traditional method may be published by multiple channel editors or recommended multiple times, so the auditor needs to perform repeated auditing operations on each reprinted article, which makes the efficiency very low.
[0060] After using the present application, this embodiment gives an example. The system detects that a reprinted article about "a major event" contains sensitive word A, calculates its group_hash (based on "sensitive information" + fragment + sensitive word A + recommended word), and assumes that the article is first appeared and has no historical audit record. The detection result is marked as "to be audited", the auditor audits the article and determines that it is "non-problem" and confirms the release. After that, another channel reprints the same article, the system detects and calculates its group_hash, and finds that the value is exactly the same as the value when the system starts detecting. At this time, the system automatically marks the audit status of the new detection result as "non-problem" through the automatic inheritance mechanism, and the auditor marks it as "system" and stores it in the database. This article does not need to be audited again. At this time, the auditor will never see this reprinted article in the background "to be audited" list, which has been audited.
[0061] The above are only preferred embodiments of the present application, not other forms of limitations on the present application. Any skilled person in the art can use the above disclosed technical content to make changes or modifications to equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made to the above embodiments without departing from the technical solution content of the present application, according to the technical essence of the present application, still belongs to the protection scope of the technical solution of the present application.
Claims
1. A method for improving the efficiency of manual review of text content, characterized in that, The method comprises the following steps: S1, detecting the content of the text to be audited to generate a detection result; S2, for each detection result, calculating a grouping hash value according to the detection type and key information thereof, which is used to group similar contents into the same group; S3, for each detection result, calculating a unique identifier according to the detection type, media identifier, webpage link and key information, which is used for deduplication processing; S4, before storing the detection result in the database, querying whether the same record exists in the database according to the unique identifier, if the same record exists, discarding the current detection result, if the same record does not exist, continuing to determine whether there is a historical artificial auditing result that can be inherited; S5, if there is a historical artificial auditing result that can be inherited, automatically inheriting the most recent artificial auditing result, including the auditing state, risk type after auditing and risk level after auditing, and marking the auditor as system and setting the auditing time as the current system time; S6, storing the detection result, the grouping hash, the unique identifier and the inherited auditing result in the database; S7, in the artificial auditing link, deduplicating and displaying the data to be audited according to the grouping hash, and batch auditing the deduplicated data by the auditor; S8, after the auditing is completed, updating the auditing state of all detection results under the grouping hash.
2. The method for improving the efficiency of human review of text content of claim 1, wherein, The calculation method of the grouping hash is different according to the detection type, comprising: for sensitive information, calculating the hash of "detection type + sentence + sensitive word + recommended word"; for privacy information, calculating the hash of "detection type + sentence + privacy content"; for sensitive person, calculating the hash of "detection type + sentence + name + position"; for risk external link, calculating the hash of "detection type + external link"; for custom word, calculating the hash of "detection type + sentence + custom word + recommended word".
3. The method for improving the efficiency of human review of text content of claim 1, wherein, The calculation method of the unique identifier is different according to the detection type, comprising: for sensitive information, calculating the hash of "detection type + media Id + webpage link + sentence + sensitive word + recommended word"; for privacy information, calculating the hash of "detection type + media Id + webpage link + sentence + privacy content"; for sensitive person, calculating the hash of "detection type + media Id + webpage link + sentence + name + position"; for risk external link, calculating the hash of "detection type + media Id + webpage link + external link"; for custom word, calculating the hash of "detection type + media Id + webpage link + sentence + custom word + recommended word".
4. The method for improving the efficiency of human review of text content of claim 1, wherein, The step S5 specifically comprises: S51, querying the database for the artificial auditing record according to the grouping hash of the current detection result; S52, arranging the artificial auditing records in descending order of auditing time, and selecting the most recent artificial auditing result for inheritance; S53, setting the auditing state, risk type after auditing and risk level after auditing of the current detection result as the corresponding values of the inherited record; S54, marking the auditor as system, which is used to distinguish whether the auditing is automatically performed by the system or manually performed by the auditor, and setting the auditing time as the current system time.
5. The method for improving the efficiency of human review of text content of claim 4, wherein: The audit status is set to the corresponding field value of the last manual audit record; the post-audit risk type is set to the corresponding field value of the last manual audit record; and the post-audit risk level is set to the corresponding field value of the last manual audit record.
6. The method for improving the efficiency of human review of text content of claim 1, wherein, The batch audit step in the step S7 comprises: S71, the auditor queries the data to be audited according to the detection type; S72, the system displays the data to be audited after deduplication according to the grouping hash; S73, the auditor selects one or more representative records for auditing operation; S74, the system marks all detection results in the group as audited according to the grouping hash of the selected records.
7. The method for improving the efficiency of human review of text content of claim 6, wherein: For the detection results of the risk external link type, three batch audit modes are supported during auditing, namely, auditing according to the main domain name, auditing according to the Host, and auditing according to the link, wherein the auditing according to the main domain name is to audit all external link detection results under the same main domain name together; the auditing according to the Host is to audit all external link detection results under the same Host together; and the auditing according to the link is to audit according to the grouping hash, which is consistent with other detection types.
8. The method for improving the efficiency of human review of text content of claim 1, wherein: The detection results comprise but are not limited to one or more of the sensitive information, the private information, the sensitive person, the risk external link, and the custom word.
9. The method for improving the efficiency of human review of text content of claim 1, wherein: The audit status comprises two states of "confirming that it is a problem requiring modification" and "confirming that it is not a problem requiring modification".
10. The method for improving the efficiency of human review of text content of claim 1, wherein: The post-audit risk type and risk level are dynamically set according to the specific detection type and the audit result.
Citation Information
Patent Citations
Content checking method and system
CN103885964A
Auditing data duplication eliminating method of distributed deploying and auditing platform
CN105721256A
Assurance of security rules in network
CN112219382A
Search result display method and device, equipment and medium
CN116069995A
Information display method and device, equipment, storage medium and program product
CN118484611A