Prohibited word auditing system based on B2B platform

By introducing technical means of multi-level prohibited word management, context analysis, deformation word processing and positive and anti-word selection support, the problem of insufficient flexibility in detection of prohibited words in the existing technology is solved, and higher detection accuracy and flexibility are achieved to meet the needs of different platforms and cultures.

CN120046609APending Publication Date: 2025-05-27河南嵩网信息科技有限公司

Patent Information

Application Number
CN202411955889.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-28
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology lacks flexibility in the detection of prohibited words, and cannot effectively identify illegal content related to complex language structures, deformed words and context, resulting in high false alarm and missed alarm rates, and it is difficult to adapt to the diverse needs of different platforms, cultures and application scenarios.

Method used

Multi-level banned word management, context analysis, deformation word processing, and positive word selection and anti-word selection support are introduced. Through multi-level management, the weight and priority of banned words are distinguished, context analysis is used to identify the differences between normal and illegal use, identify and deal with the deformation forms of banned words, and allow administrators to configure positive word selection and anti-word selection to adapt to complex contexts.

Benefits of technology

It improves the accuracy and flexibility of banned word detection, reduces false alarms and missed reports, enhances the ability to identify complex violation scenarios, adapts to the needs of different cultures, languages ​​and platforms, and improves the intelligence and humanization of content review.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A B2B platform-based prohibited word auditing system specifically comprises the following steps: step 1, preliminary prohibited word detection: firstly, the system performs preliminary prohibited word detection on a text submitted by a user, and updates a prohibited word library at the same time; 2, anagrams are detected, wherein after prohibited words pass detection, the text is sent into a special anagrams detection module to be further processed; and in the deformed word detection module, the text is subjected to alphabetic processing, and more complex font and symbol deformation matching is carried out, so that all potential prohibited behaviors can be identified, and meanwhile, a prohibited word bank is updated. According to the method, multi-level prohibited word management is adopted, the multi-level management allows the system to carry out distinguishing and processing according to the severity of the content, different weight and priority settings are provided, the accuracy and flexibility of detection are improved, different degrees of illegal content can be effectively coped with, and misinformation and missing report are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a prohibited word review system based on a B2B platform. Background Art

[0002] In modern digital communication, especially on content generation and interaction platforms such as B2B platforms, prohibited word detection is an important part of ensuring the legality and compliance of platform content. The main task of a prohibited word detection system is to identify and filter content containing sensitive, illegal, or inappropriate words to maintain the legality of the platform and the user experience.

[0003] Implementation methods and deficiencies of the prior art: 1. Text analysis and pattern matching: Many existing systems rely on simple text analysis and pattern matching techniques, such as keyword filtering. Although this method can detect explicitly appearing prohibited words, it is not flexible enough in dealing with complex language structures. Moreover, the detection process usually lacks an understanding of the text context, resulting in an inability to distinguish between normal and illegal uses of prohibited words, with high false positive and false negative rates. For example, some words are harmless in a specific context but may involve sensitive content in other situations.

[0004] 2. Machine learning: Some systems attempt to use machine learning algorithms to improve the intelligence of detection, but this usually requires a large amount of labeled data for training and has limited recognition effects for specific contexts.

[0005] 3. Single prohibited word list: Most current systems use a single prohibited word list, treating each word as an independent unit. The system can only detect single-occurring prohibited words and performs poorly in dealing with cases of multiple prohibited word combinations or semantically related expressions, lacking an analysis of the relevance between words. Therefore, it cannot identify illegal content that needs to be determined by a combination of multiple prohibited words, resulting in some complex illegal scenarios being ignored by the system.

[0006] 4. Insufficient handling of variant words: When faced with variant forms of words (such as spelling mistakes, stemming, synonyms, etc.), the system often lacks effective recognition capabilities, and detection can be easily bypassed. Users can avoid detection by using spelling mistakes, stemming, or synonyms. In this case, the system cannot effectively identify these variant words, resulting in insufficient comprehensiveness of detection.

[0007] 5. Lack of support for positive and negative selected words: The prior art lacks support for positive and negative selected words and cannot make more detailed judgments based on context conditions. For example, a prohibited word should be regarded as a normal use in a certain context, but the existing system cannot recognize this situation and cannot accurately judge whether the content is prohibited in a specific context, resulting in misjudgments.

[0008] Therefore, the prior art lacks flexibility in detecting prohibited words. In particular, the configuration of prohibited words in existing systems is usually difficult to adjust and cannot meet the diverse needs of different platforms, cultures, and application scenarios.

[0009] CN118536506A discloses a multi-modal prohibited word detection method and system. The method includes: extracting image, voice, and text data from an advertising video, obtaining each prohibited word in a prohibited word library, obtaining a set of word segmentation results of the text data, constructing a multi-dimensional character-word frequency vector for each character of each prohibited word in the prohibited word library, obtaining a positive judgment function and a negative judgment function for new words of each character within each prohibited word in the prohibited word library, determining a positive index and a negative index for each character of each prohibited word in the prohibited word library, obtaining a character reconfirmation feature index for each character of each prohibited word in the prohibited word library, obtaining a non-prohibited determination factor for each prohibited word in the prohibited word library in the text data, and combining a neural network model to complete multi-modal prohibited word detection in the advertising video. In this method, context analysis is adopted and combined with multi-dimensional characters for judgment. However, it still cannot meet the diverse needs of different platforms, cultures, and application scenarios.

[0010] CN107977423A discloses an automatic filtering processing system for Internet articles containing illegal words, including an illegal word library collection module, a word library manual verification module, a word segmentation processing module, an illegal word content conversion module, a foreground trigger-based access filtering module, and a background editing and publishing detection module. Such a technical solution can effectively and automatically filter illegal words from Internet products and article content and achieve long-term and effective automatic detection and processing of product and article content data. This method uses illegal words for judgment and also cannot meet the diverse needs of different platforms, cultures, and application scenarios. Summary of the Invention

[0011] The technical problem to be solved by the invention is: The present invention aims to solve the defects of the prior art by introducing features such as multi-level prohibited word management, context analysis, variant word processing, positive and negative selection word support, etc., and proposes a prohibited word review method and system based on a B2B platform to improve the accuracy and flexibility of prohibited word detection.

[0012] The technical solution of the present invention is specifically as follows: A prohibited word review system based on a B2B platform specifically includes the following steps: Step 1, preliminary prohibited word detection: First, the system will perform a preliminary prohibited word check on the text submitted by the user and update the prohibited word library at the same time; Step 2, Variant Word Detection: After the prohibited word detection passes, the text will be sent to a dedicated variant word detection module for further processing; in the variant word detection module, the text is phoneticized and more complex glyph and symbol deformation matching is performed to ensure that all potential prohibited behaviors can be identified, and the prohibited word library is updated.

[0013] The preliminary prohibited word detection includes a multi-level prohibited word management module, a positive and negative selection word module, and a context analysis module.

[0014] The multi-level prohibited word management module classifies prohibited words into multiple levels, and each level corresponds to different weights and priorities.

[0015] In the positive and negative selection word module, the administrator is allowed to configure positive selection words and negative selection words for each prohibited word; the positive selection word means that when the content contains a certain prohibited word and its related positive selection word at the same time, the system can determine that the content is prohibited to ensure the satisfaction of the context conditions; the negative selection word means that when the content contains a certain prohibited word but also contains its related negative selection word at the same time, the system will ignore the prohibited determination, and this configuration helps the system to identify non-prohibited situations in a specific context.

[0016] The context analysis module uses context analysis technology to consider not only individual words but also the overall relationships of phrases, sentences, and paragraphs to help the system identify prohibited words used in normal situations and distinguish them from real violations.

[0017] Step 2 specifically includes the following steps: Step 2.1, After the preliminary prohibited word detection, obtain the specific text content from the database; Step 2.2, Text preprocessing, including text cleaning, removing invalid characters to reduce noise; Step 2.3, Phoneticization: Convert all Chinese characters and symbols in the text into pinyin form.

[0018] In Step 2.3, the text is phoneticized in the order of each character from front to back.

[0019] It also includes Step 2.4, After the text is phoneticized, the phoneticized entries in the prohibited word library are permuted and combined, and the phoneticized entries in the prohibited word library are compared by the way of permutation and combination to identify potential violations in the text.

[0020] In the prohibited word review system, a unique identifier is set for each piece of text, and each piece of text is sent to the message queue according to this unique identifier, and multiple background processes simultaneously read data from the queue and perform the processing of Step 1 and Step 2.

[0021] The beneficial effects of the present invention are: 1. Multi-level banned word management: Multi-level management allows the system to differentiate and process content based on its severity, providing different weights and priority settings, improving the accuracy and flexibility of detection, and being able to effectively respond to content of varying degrees of violations, reducing false positives and missed negatives.

[0022] 2. Use inflected word processing. By introducing inflected word processing capabilities, the system can identify variants such as spelling errors, stem changes and synonyms, which significantly improves the comprehensiveness of detection, reduces missed reports caused by word form changes, and prevents users from bypassing detection through word form variants.

[0023] 3. Support for positive and negative word selection: Support for positive and negative word selection configuration. The system can accurately determine whether the content is illegal in complex contexts, reduce false positives, and judge the compliance of the content through more complex rules and conditions, thereby improving the accuracy and flexibility of detection.

[0024] 4. Contextual analysis: The system uses contextual analysis to evaluate the context in the text, identify the differences between normal and illegal use of banned words, significantly reduce false positives and false negatives, help the system distinguish between legal and illegal usage, and make content review more intelligent and humane.

[0025] 5. Flexibility and configurability: The system allows administrators to customize and adjust banned word lists, detection levels, and related parameters, providing a high degree of flexibility to adapt to the needs of different cultures, languages, and platforms, reducing operational challenges caused by lack of adaptability. DETAILED DESCRIPTION

[0026] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0027] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0028] A banned word review system based on a B2B platform. Since the matching logic of inflected words is relatively complex and the computational overhead is relatively large, in order to improve processing efficiency and reduce user waiting time, the present invention adopts a phased detection strategy, which specifically includes the following steps: Step 1, Preliminary Prohibited Word Detection: First, the system will conduct a preliminary check for prohibited words in the text submitted by the user. In this stage, only the basic matching of prohibited words is concerned, without involving the recognition of complex variant words, and the prohibited word library is updated simultaneously.

[0029] Step 2, Variant Word Detection: After the prohibited word detection passes, the text will be sent to a dedicated variant word detection module for further processing. In the variant word detection module, the text is phoneticized and more complex glyph and symbol variant matching is performed to ensure that all potential prohibited behaviors can be identified, and the prohibited word library is updated simultaneously.

[0030] In Step 1, prohibited words refer to words or phrases that are prohibited from being used in a specific platform or application scenario, and these words may involve inappropriate, sensitive or illegal content.

[0031] In the present invention, the preliminary prohibited word detection includes a multi-level prohibited word management module, a positive and negative selection word module, and a context analysis module.

[0032] The present invention adopts a multi-level prohibited word management module. The multi-level prohibited word management module classifies prohibited words into multiple levels, such as first level, second level and third level, and each level corresponds to different weights and priorities. This hierarchical management enables the system to flexibly process according to the prohibited level of the content. For example, the first-level prohibited words correspond to the most serious content, while the third-level prohibited words may involve minor violations. It should be noted that according to different industries or fields, the content of prohibited word classification is different, and those skilled in the art can classify the prohibited words according to the actual situation of the industry. In this way, the system of the present invention can more carefully control and manage the strategies of prohibited word detection, can adopt different processing strategies according to the severity of the content, and improves the flexibility and accuracy of detection.

[0033] In the positive and negative selection word module, the administrator is allowed to configure positive selection words and negative selection words for each prohibited word. The positive selection word means that when the content contains a certain prohibited word and its related positive selection word at the same time, the system can determine that the content is prohibited to ensure the satisfaction of the context conditions. The negative selection word means that when the content contains a certain prohibited word but also contains its related negative selection word at the same time, the system will ignore the prohibited determination, and this configuration helps the system to identify non-prohibited situations in a specific context.

[0034] The present invention configures positive and negative selection words for prohibited words, supports the configuration of positive selection words and negative selection words, enabling the system to determine whether the content is prohibited according to conditions in complex scenarios. The positive selection words ensure that the prohibition is determined only when the specific context is met, while the negative selection words exclude the prohibition determination when the conditions are met. This mechanism increases the accuracy of system detection, especially in scenarios where context conditions need to be identified, reducing misjudgments and false alarms.

[0035] In the context analysis module, more accurate detection is achieved by analyzing the text context to distinguish between normal and illegal usage. The context analysis module uses context analysis technology to not only consider individual words, but also analyze the overall relationship between phrases, sentences and paragraphs to help the system identify banned words used in normal situations and distinguish them from real violations. The present invention evaluates the context in the text to determine whether the content is banned, solving the problem of false positives and false negatives caused by the lack of context analysis in the prior art.

[0036] In step 2, the inflected words are detected.

[0037] Deformed words include various variant forms of banned words, such as misspellings, synonyms, and stem changes, and the system can identify and detect these variants. The present invention can detect various variant forms of illegal content by identifying and processing variants of banned words (such as misspellings, stem changes, and synonyms), thereby enhancing the comprehensiveness of detection and reducing the possibility of bypassing detection.

[0038] The specific steps include: Step 2.1, after preliminary banned word detection, obtain specific text content from the database; Step 2.2: Text preprocessing, including text cleaning, removing invalid characters (such as spaces, special symbols, etc.) to reduce noise and ensure the accuracy of subsequent analysis; Step 2.3, Pinyin processing: Convert all Chinese characters and symbols in the text (such as plus sign "+", asterisk "*", etc.) into Pinyin form. For example, "+" will be converted to "jia", "v" will be converted to "wei", and the capital letter "V" will also be Pinyinized to "wei" to achieve matching of Pinyin deformation.

[0039] For example, the text submitted by the user "You can +V to contact me" will be processed as: ke yi jia wei lailianxi wo ou. The system will compare this pinyin sequence and find that it contains the banned word pinyin "jia wei" and mark it as a violation. This allows the system to detect banned words in different spellings, whether the user enters "jiawei" or expresses it in pinyin or in a deformed way, and can accurately identify it.

[0040] Preferably, in step 2.3, the text is pinyinized in a word-by-word order from front to back.

[0041] Furthermore, the method further includes step 2.4, after the text is pinyinized, the pinyinized entries in the banned word library are arranged and combined, and the pinyinized entries in the banned word library are compared by the arrangement and combination method, so as to identify potential violations in the text.

[0042] The present invention performs pinyin processing on all prohibited words in the prohibited word library, and can identify and process various forms of prohibited words, ensuring that even if the prohibited words are rewritten through spelling transformation, symbol replacement or other means, the system can still accurately detect violations. This function significantly improves the comprehensiveness and accuracy of detection, effectively preventing content creators from bypassing the system's review mechanism through simple word form changes, and improving the matching accuracy and the universality of the system.

[0043] Preferably, in the prohibited word review system of the present invention, a unique identifier is set for each piece of text, and each piece of text is sent to the message queue according to the unique identifier. Multiple background processes simultaneously read data from the queue and perform the processing of steps 1 and 2. In this way, we can effectively allocate computing resources, shorten the response time, and prevent the system from being overloaded when processing a large amount of text.

[0044] The present invention has high configurability. Administrators can adjust the prohibited word list, detection level, and parameters of positive and negative selected words according to requirements. This flexibility allows the system to be easily adapted to different application scenarios, such as content review requirements for different languages, cultures or platforms. In addition, the system supports real-time configuration updates, further enhancing the convenience and flexibility of operation. Additionally, the system has high configurability, allowing administrators to customize the prohibited word list, level, and other detection parameters. This flexibility enables the system to adapt to different platforms, cultures and application scenarios, meeting different review requirements.

[0045] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the overall concept of the present invention, several changes and improvements can still be made, and these should also be regarded as the protection scope of the present invention.

Claims

1. A banned word review system based on a B2B platform, characterized by: The specific steps include: Step 1: Preliminary banned word detection: First, the system will perform a preliminary banned word check on the text submitted by the user and update the banned word library at the same time; Step 2: Deformed word detection: After the banned word detection is passed, the text will be sent to a special deformed word detection module for further processing; in the deformed word detection module, the text is pinyinized and more complex glyph and symbol deformation matching is performed to ensure that all potential banned behaviors can be identified and the banned word library is updated at the same time.

2. According to claim 1, a banned word review system based on a B2B platform is characterized in that: The preliminary banned word detection includes a multi-level banned word management module, a positive and negative word selection module, and a context analysis module.

3. According to claim 2, a banned word review system based on a B2B platform is characterized in that: The multi-level banned word management module divides banned words into multiple levels, each level corresponding to different weights and priorities.

4. According to claim 2, a banned word review system based on a B2B platform is characterized in that: In the positive and negative word selection module, administrators are allowed to configure positive and negative words for each banned word; positive words mean that when the content contains both a banned word and its related positive words, the system will determine the content as banned to ensure that the context conditions are met; negative words mean that when the content contains a banned word but also contains its related negative words, the system will ignore the banned judgment. This configuration helps the system identify non-banned situations in specific contexts.

5. The banned word review system based on a B2B platform according to claim 2 is characterized in that: The context analysis module adopts context analysis technology, which not only considers individual words but also analyzes the overall relationship of phrases, sentences and paragraphs to help the system identify banned words used in normal situations and distinguish them from real violations.

6. The banned word review system based on a B2B platform according to claim 1, characterized in that: Step 2 specifically includes the following steps: Step 2.1, after preliminary banned word detection, obtain specific text content from the database; Step 2.2, text preprocessing, including text cleaning, removing invalid characters to reduce noise; Step 2.3: Pinyin processing: convert all Chinese characters and symbols in the text into pinyin form.

7. The banned word review system based on a B2B platform according to claim 6, characterized in that: In step 2.3, the text is pinyinized in a word-by-word order from front to back.

8. The banned word review system based on a B2B platform according to claim 6, characterized in that: The method also includes step 2.4, after the text is pinyinized, the pinyinized entries in the banned word library are arranged and combined, and the pinyinized entries in the banned word library are compared by the arrangement and combination method, so as to identify potential violations in the text.

9. The banned word review system based on a B2B platform according to claim 1, characterized in that: In the banned word review system, a unique identifier is set in each text, and each text is sent to the message queue according to the unique identifier. Multiple background processes read data from the queue at the same time and perform steps 1 and 2.

Citation Information

Patent Citations

  • Method and system for automatically filtering and processing illegal word-containing internet article

    CN107977423A

  • Multi-mode forbidden word detection method and system

    CN118536506A

Cited By

  • Internet-oriented illegal advertisement identification method, device and system

    CN120338884A

  • Video speech recognition method and system for illegal short video

    CN120727040A