Text review method and related devices, equipment and storage media

By performing character filtering and matching link assembly on the input text, the existing text review methods solve the inaccuracy and incomplete problems in the face of complex evasion methods, and achieve more efficient violation identification and processing.

CN119416790BActive Publication Date: 2025-05-30IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510026504.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When faced with complex evasion methods, the existing text review methods are inaccurate and incomplete, making it difficult to effectively identify and deal with illegal content such as substitution of similar characters, splitting Chinese characters, and homophone replacement.

Method used

By filtering the input text, including character culling and variant conversion, the text to be reviewed expressed in natural language is obtained, and multiple matching methods are assembled according to the application scenario to form a matching link. Combined with series and parallel connection methods, the text to be reviewed is reviewed.

Benefits of technology

It improves the comprehensiveness and accuracy of text review, can more effectively identify and deal with complex violations, and adapt to violations in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416790B_ABST
    Figure CN119416790B_ABST
Patent Text Reader

Abstract

The present application discloses a text review method and related devices, equipment, and storage media. Among them, the text review method includes: performing character filtering on the input text to obtain the text to be reviewed expressed in natural language, and assembling several matching methods based on the application scenario of the input text to obtain a matching link; wherein, the character filtering includes at least one of character elimination and variant conversion, the matching link includes at least one matching method, and the matching methods in the matching link are assembled in at least one connection method of series and parallel; reviewing the text to be reviewed based on the matching link to obtain a review result; wherein, the review result includes whether the input text involves illegal content. The above solution can improve the comprehensiveness and accuracy of text review.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A text review method, characterized in that: include: Character filtering is performed based on the input text to obtain a text to be reviewed expressed in a natural language, and a plurality of matching methods are assembled based on the application scenario of the input text to obtain a matching chain; wherein the character filtering includes at least one of character elimination and variant conversion, the matching chain includes at least one of the matching methods, and the matching methods in the matching chain are assembled in at least one connection mode of series connection and parallel connection, and the complexity of the application scenario is positively correlated to the number of the matching methods in the matching chain and the complexity of the connection mode; The text to be reviewed is reviewed based on the matching link to obtain a review result; wherein the review result includes whether the input text involves illegal content, and when the matching link includes a regular match, the method includes: Obtain a regular expression library; wherein the regular expression library contains a plurality of regular expressions, and the regular expressions are used to hit illegal content; Based on the regular expression library, an index tree for filtering the regular expression is constructed; wherein different paths in the index tree represent different regular expressions respectively; Traversing the index tree based on the text to be reviewed to select the regular expression as a candidate expression; Reconstructing based on the candidate expression to obtain an uncertain finite automaton; Performing conversion based on the uncertain finite automaton to obtain a deterministic finite automaton; The text to be reviewed is matched based on the determined finite automaton to obtain the output result of the regular match; wherein the output result includes whether illegal content is hit.

2. The method according to claim 1, characterized in that In the case where the matching link includes a sensitive word match, the method includes: Based on the sensitive word library, a sensitive word matching tree is constructed; wherein the sensitive word library contains a number of sensitive words involving illegal content, and the pronunciation units represented by each node on any path in the sensitive word matching tree are sequentially combined into a pronunciation sequence of the sensitive word; In response to the pronunciation unit represented by the node being an initial consonant unit of a flat tongue consonant and a retroflex consonant, completing the pronunciation unit represented by the node to be an initial consonant unit containing both a flat tongue consonant and a retroflex consonant; The text to be reviewed is matched based on the sensitive word matching tree to obtain an output result of the sensitive word matching; wherein, when the sensitive word matching turns on fuzzy matching of front and back nasals, in the process of matching the text to be reviewed based on the sensitive word matching tree, the pronunciation unit and the back nasals in the text to be reviewed are automatically ignored, and the output result includes whether illegal content is hit.

3. The method according to claim 1, characterized in that The matching of the to-be-audited text based on the determined finite automaton to obtain the output result of the regular matching includes: Initialize the current state to the initial state, and read the text to be reviewed character by character; Inquiring in the finite automaton the target state to which the initial state is transferred after encountering the current character, and updating the initial state to the target state; Detect whether the latest initial state is an accepting state; In response to the latest initial state being the acceptance state, determining that the output result includes hitting illegal content; In response to the latest initial state not being the accepting state, returning to the step of reading the text to be reviewed character by character.

4. The method according to claim 1, characterized in that: The step of constructing an index tree for filtering the regular expression based on the regular expression library comprises: Extracting a deterministic substring of the regular expression as an index key of the regular expression; wherein the deterministic substring is a string that is screened to determine whether it is suspected to match the regular expression; The index tree is constructed based on the index keys of the regular expressions.

5. The method according to claim 1, characterized in that In the case where the matching link includes a semantic match, the method includes: The text to be reviewed is identified based on a semantic recognition model to obtain an output result of the semantic matching; wherein the output result includes whether illegal content is found.

6. The method according to claim 1, characterized in that In the process of reviewing the text to be reviewed based on the matching link, the method further includes: In response to the output result of any of the matching methods on the matching link including hitting illegal content, the matching link is jumped out, and the audit result is directly determined to be that the input text involves illegal content.

7. The method according to any one of claims 1 to 6, characterized in that: The matching link includes sensitive word matching, regular matching and semantic matching connected in series; And / or, the input text includes at least one of input content and output content of a large language model.

8. A text review device, characterized in that: include: A filtering and assembling module, used for performing character filtering based on an input text to obtain a text to be reviewed expressed in a natural language, and assembling a plurality of matching methods based on an application scenario of the input text to obtain a matching chain; wherein the character filtering includes at least one of character elimination and variant conversion, the matching chain includes at least one of the matching methods, and the matching methods in the matching chain are assembled in at least one connection mode of series connection and parallel connection, and the complexity of the application scenario is positively correlated to the number of the matching methods in the matching chain and the complexity of the connection mode; A result acquisition module is used to audit the text to be audited based on the matching link to obtain an audit result; wherein the audit result includes whether the input text involves illegal content, and when the matching link contains a regular match, it includes: Obtain a regular expression library; wherein the regular expression library contains a plurality of regular expressions, and the regular expressions are used to hit illegal content; Based on the regular expression library, an index tree for filtering the regular expression is constructed; wherein different paths in the index tree represent different regular expressions respectively; Traversing the index tree based on the text to be reviewed to select the regular expression as a candidate expression; Reconstructing based on the candidate expression to obtain an uncertain finite automaton; Performing conversion based on the uncertain finite automaton to obtain a deterministic finite automaton; The text to be reviewed is matched based on the determined finite automaton to obtain the output result of the regular match; wherein the output result includes whether illegal content is hit.

9. An electronic device, characterized in that: The invention at least comprises a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the text review method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the text review method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text auditing method and device, equipment and medium

    CN111506708A

  • Medical insurance medication auditing method and device, electronic equipment and storage medium

    CN116525136A

  • Text sensitive word detection method and device, equipment and storage medium

    CN117435720A