Filtering method and system based on text feature analysis

Through a filtering method based on text feature analysis, using standardized point mutual information and word frequency statistics to calculate phrase cohesion, combined with a custom dictionary, the problem of expensive finished data sets and difficulty in adapting to professional needs in large language model training is solved, achieving low-cost and efficient corpus screening and information security assurance.

CN120706423APending Publication Date: 2025-09-26BEIJING ELECTRONICS SCI & TECH INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510914127.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-26

Smart Images

  • Figure CN120706423A_ABST
    Figure CN120706423A_ABST
Patent Text Reader

Abstract

The invention provides a filtering method and system based on text feature analysis, and the method comprises the steps: carrying out the filtering and deletion of sensitive words of a received input text, and obtaining a primary filtering text; performing primary word segmentation on the primary filtered text, and performing primary word frequency statistics on a primary word segmentation result; calculating a cohesion degree score of adjacent words in the primary word segmentation result, determining the adjacent words for word combination based on a comparison result of a cohesion degree threshold value and the cohesion degree score, and executing combination operation to obtain a secondary filtering text; and carrying out secondary word frequency statistics on the secondary filtered text, screening words with word frequencies higher than a preset threshold value in a secondary word frequency statistics result, and carrying out matching search in a self-defined field dictionary to determine a field scene to which the input text belongs. The method serves as a personalized scheme for optimizing the large language model corpus, the cost is reduced, meanwhile, the safety of the large language model corpus is guaranteed, and an effective and convenient filtering way is provided for developers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large language model information security technology, and relates to a filtering method and system based on text feature analysis. Background Art

[0002] Faced with the craze for large language models, many companies and platforms have jumped on the bandwagon. However, as this field is still emerging, supporting software and resources are still relatively scarce. The size of the corpus dataset used in large language model training directly determines its performance. Larger datasets generally result in better model performance. Currently, many platforms primarily provide ready-made datasets. While this seemingly simple approach of providing only ready-made datasets can actually harbor potential problems.

[0003] First, they are expensive and offer limited options. Datasets are generally large, requiring extensive manpower to screen and sift through, creating a high barrier to entry. For smaller development teams, limited manpower and funding often make it difficult to fully screen and customize datasets for their specific domains, making it difficult to create their own. Purchasing datasets is also expensive. Since there are only a few vendors and platforms offering ready-made training sets, and because scarcity makes things valuable, some vendors charge exorbitant prices. If trainees want custom training sets, the cost can become even more prohibitive. High training set prices increase training costs. While some open-source, ready-made training sets exist, these are mostly designed for general-purpose applications and are unable to meet diverse training needs. These factors make it difficult for trainees to select a training set that precisely meets their needs.

[0004] Second, they are difficult to adapt to professional or specific needs. These datasets are mainly designed to respond to general user instructions or target a few specific fields. They lack understanding of other special industries and professional knowledge, making it difficult to meet the needs of some professional fields.

[0005] Third, the source of the dataset may be insecure. Since developers receive off-the-shelf datasets, they lack visibility into the dataset selection process. This could lead to subtle manipulation of the dataset, such as using biased corpus, causing it to output poorly designed and biased responses. This issue can have even more serious consequences when it comes to coding. If someone deliberately modifies the dataset, hiding malicious code within a commonly used function, the AI ​​could mistakenly believe the code is part of a common function and output code containing a malicious backdoor. Purchasing off-the-shelf datasets can reduce the resources allocated to the dataset, leading to a lack of scrutiny of its internal content, potentially allowing malicious code to enter the model. Furthermore, if the large language model is not specifically designed for programming but rather a general-purpose model or is intended for other non-computer science fields, its users may lack programming expertise and be unable to identify the security of the code, potentially exposing their computers to attacks.

[0006] Therefore, how to provide a low-cost, effective and convenient filtering method and system based on text feature analysis is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0007] In light of this, the present invention proposes a filtering method and system based on text feature analysis. This method calculates phrase cohesion through standardized point mutual information and word frequency statistics. Combined with custom dictionary loading, it identifies and filters domain phrases, performs text analysis and filtering, and thus rapidly filters out harmful information and ensures information security. Furthermore, a user-friendly front-end is provided to facilitate user experience.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] The present invention discloses a filtering method based on text feature analysis, comprising the following steps:

[0010] S1: Recognize sensitive words in the received input text, and delete the recognized sensitive words from the input text to obtain a filtered text;

[0011] S2: performing initial word segmentation on the filtered text and performing initial word frequency statistics on the initial word segmentation results;

[0012] S3: Based on the initial word frequency statistics, adaptively segmenting the first filtered text, including: calculating cohesion scores of adjacent words in the initial word segmentation results, and determining adjacent words to be merged based on a comparison result of a cohesion threshold and the cohesion scores, and performing a merging operation to obtain a second filtered text;

[0013] S4: Perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, perform matching searches in a custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

[0014] Preferably, the S1 comprises the following steps:

[0015] S11: Recognize sensitive words in the received input text, and replace the recognized sensitive words with unified identifiers to obtain replaced text;

[0016] S12: uniformly deleting the identifiers in the replaced text to obtain a filtered text.

[0017] Preferably, the formula for calculating the cohesion score of adjacent words in the initial word segmentation result in S3 is:

[0018]

[0019] In the formula, score represents the cohesion score; time represents the number of times two adjacent words appear consecutively in a filtered text; len(ls) represents the total number of words in a filtered text; key_m represents the number of times the first word in the adjacent words appears in a filtered text, and key_n represents the number of times the second word in the adjacent words appears in a filtered text.

[0020] Preferably, the method further comprises the following steps:

[0021] Set the threshold time_y for the number of times two adjacent words appear consecutively in a filtered text;

[0022] Determine whether the time value of the current two adjacent words is greater than the value of time_y. If so, calculate the cohesion score of the current two adjacent words.

[0023] Preferably, the method further includes the step of performing sentiment analysis on the secondary filtered text and outputting sentiment tendency conclusion information.

[0024] Preferably, the steps of operating the customized domain dictionary include:

[0025] Receive the domain dictionary of the required domain uploaded by the user in real time and perform storage operations; receive the matching instructions sent by the user in real time, and perform matching searches in the stored domain dictionary for the words whose frequency is higher than the preset threshold in the secondary word frequency statistics results; or

[0026] Receive the user's deletion instructions for the uploaded domain dictionary in real time and perform the deletion operation so that the currently deleted domain dictionary is excluded when executing the matching instructions sent by the user.

[0027] Preferably, the method further comprises the following steps:

[0028] Receive domain keywords uploaded by users in real time and perform storage operations; receive matching instructions sent by users in real time, and select words with a frequency higher than a preset threshold in the secondary word frequency statistics results for matching with domain keywords; or

[0029] Receive the user's deletion instructions for the uploaded domain keywords in real time and perform the deletion operation so that the currently deleted domain keywords are excluded when executing the matching instructions sent by the user.

[0030] Preferably, the method further includes a step of displaying the domain scene to which the secondary filtered text and the input text belong.

[0031] The present invention also provides a filtering system based on text feature analysis according to the filtering method based on text feature analysis, comprising: a sensitive word filtering unit, a primary word segmentation unit, an adaptive word segmentation unit and a domain matching unit; wherein,

[0032] The sensitive word filtering unit is used to identify sensitive words in the received input text and delete the identified sensitive words from the input text to obtain a filtered text;

[0033] The primary word segmentation unit is used to perform primary word segmentation on the primary filtered text and perform primary word frequency statistics on the primary word segmentation result;

[0034] The adaptive word segmentation unit is used to perform adaptive word segmentation on the first filtered text according to the initial word frequency statistics result, including: calculating the cohesion score of adjacent words in the initial word segmentation result, and determining the adjacent words to be merged based on the comparison result of the cohesion threshold and the cohesion score, and performing the merging operation to obtain the second filtered text;

[0035] The domain matching unit is used to perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, perform matching searches in a custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

[0036] Preferably, it also includes a user interaction interface; the user interaction interface includes a text upload button, a word segmentation dictionary word add button, a word segmentation dictionary word delete button and a filter result display area; wherein,

[0037] The text upload button receives and responds to the input text upload instruction sent by the user, receives the input text and stores it;

[0038] The word segmentation dictionary word addition button receives and responds to a word segmentation dictionary word addition instruction sent by a user, and stores the added word segmentation dictionary word for being called by the initial word segmentation unit to perform an initial word segmentation operation, and / or for being called by the domain matching unit to perform a domain matching operation;

[0039] The word segmentation dictionary word deletion button receives and responds to the word segmentation dictionary word deletion instruction sent by the user, and deletes the stored word segmentation dictionary words;

[0040] The filtering result display area is used to display the domain scenario to which the secondary filtered text and the input text belong.

[0041] It can be seen from the above technical solution that, compared with the prior art, the beneficial effects of the present invention include:

[0042] In order to solve the problem of text analysis and filtering of corpora in large language models, the present invention aims to provide screening for the corpus, reduce the workload of developers in reviewing and selecting corpora, and thus reduce the cost of maintaining the security of large language models.

[0043] The present invention can provide a preliminary corpus screening and scoring function for the training of a large language model in a specific field according to the different fields that the corpus is aimed at. Developers can choose different modules according to different academic fields. According to the different field modules selected by the developer, articles with less relevance to the knowledge in the field can be screened out for a specific field, and the relevance of the screened articles can be scored, thereby realizing a preliminary filtering function of the text. As a large language model corpus screening software, the present invention provides a new solution for filtering bad corpus involved in large language model training.

[0044] While filtering, this invention innovatively incorporates a word segmentation system, allowing users to customize filtering rules and conditions based on actual needs without the need for additional programming. Simply inputting a selection of specialized articles during use enables adaptive word segmentation. This invention combines standardized point mutual information and word frequency statistics to calculate phrase cohesion, achieving domain-adaptive word segmentation and significantly improving the effectiveness of word segmentation. This innovative approach allows for accurate word segmentation of text without the machine understanding its meaning.

[0045] The present invention realizes the cleaning of sensitive words in articles, and understands that the unit of text is words, not sentences or contexts. Therefore, it is necessary to quickly identify illegal words and block them from the perspective of sensitive words, whether in advertising, law or other fields.

[0046] The present invention uses a text sentiment scoring algorithm to implement simple sentiment analysis of text. Through the UI interface of the work, users can independently choose to use various functions.

[0047] Overall, this invention, with its intelligent corpus screening, low-cost investment, and adaptable flexibility, has injected new vitality into the development of big data models and NLP. In today's society, the field of AI is experiencing rapid development and widespread application, and there is an urgent need for a highly efficient and adaptive corpus screening solution to achieve the goal of filtering out completely healthy and safe corpora for AI training. Therefore, this invention clearly has irreplaceable contemporary value and provides new vitality for the development of the artificial intelligence industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. Those skilled in the art can also derive other drawings based on the provided drawings without inventive effort.

[0049] Figure 1 A flowchart of a filtering method based on text feature analysis provided by an embodiment of the present invention;

[0050] Figure 2 A flow chart for generating a filtered text provided by an embodiment of the present invention;

[0051] Figure 3 A flow chart for generating secondary filtered text provided by an embodiment of the present invention;

[0052] Figure 4 A schematic diagram of a user interaction interface provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0054] The first aspect of the embodiment of the present invention provides a filtering method based on text feature analysis. The filtering process is carried out in two steps. The first step is to detect and clean the sensitive words in the text; the second step is to analyze the topic of the article through word frequency statistics according to the training requirements of the large language model, so as to select the required field corpus. From coarse to fine, efficient filtering of the corpus is achieved. Figure 1 As shown, the following steps are included:

[0055] S1: Identify sensitive words in the received input text and delete the identified sensitive words from the input text to obtain a filtered text;

[0056] S2: Perform initial word segmentation on the filtered text and perform initial word frequency statistics on the initial word segmentation results;

[0057] S3: Based on the initial word frequency statistics, adaptively segmenting the first filtered text, including: calculating the cohesion scores of adjacent words in the initial segmentation results, and determining adjacent words to be merged based on the comparison result of the cohesion threshold and the cohesion scores, and performing the merging operation to obtain the second filtered text;

[0058] S4: Perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, and perform matching searches in the custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

[0059] In one embodiment, S1 includes the following steps:

[0060] S11: Recognize sensitive words in the received input text, and replace the recognized sensitive words with unified identifiers to obtain replaced text;

[0061] S12: uniformly delete the identifiers in the replaced text to obtain a filtered text.

[0062] In this embodiment, Figure 2 As shown, a certain amount of sensitive word libraries have been collected in advance, and the import of sensitive words can be achieved by calling the os library and using its file operation function.

[0063] The sensitive word detection algorithm can be implemented through Python's built-in replace and in. If the sensitive word library is too large, you can also choose an efficient sensitive word filter based on DFA and implemented through the tire tree.

[0064] The effect achieved by this embodiment is to clean the text by detecting sensitive words and illegal words, replacing them with "*", and deleting them in the next step without participating in the word frequency statistical process.

[0065] In one embodiment, Figure 3 As shown in S2, before the word frequency statistics are performed, a pre-processing operation of the filtered text is also included:

[0066] Unify the format of the filtered text. Specifically, collect existing punctuation statistics and stop word lists, and then remove punctuation from the filtered text through screening and comparison, without removing stop words. Stop words are not removed because when S3 calculates cohesion, stop words may be part of a vocabulary in the domain. Retaining these words prevents misjudgments when merging adjacent words and matching domains.

[0067] In one embodiment, the specific execution steps of the initial word segmentation include: implementing it by calling jieba.lcut(content, cut_all=False, HMM=True). Jieba is a commonly used Chinese word segmentation library, and the lcut method will directly return a list of segmented words. The parameter cut_all=False means using the precise mode for word segmentation, which attempts to cut the sentence most accurately and is suitable for text analysis; HMM=True means using the Hidden Markov Model (Hidden Markov Model), which can improve the accuracy of word segmentation, especially for the recognition of some unregistered words (that is, words that do not appear in the word segmentation dictionary).

[0068] In one embodiment, after the initial word segmentation, S3-S4 will adaptively segment the text domain and count the corresponding word frequencies to obtain the topic of the article. The formula for calculating the cohesion score of adjacent words in the initial word segmentation result is:

[0069]

[0070] In the formula, score represents the cohesion score; time represents the number of times two adjacent words appear consecutively in a filtered text; len(ls) represents the total number of words in a filtered text; key_m represents the number of times the first word in the adjacent words appears in a filtered text, and key_n represents the number of times the second word in the adjacent words appears in a filtered text.

[0071] For example, if we only use Jieba word segmentation without importing a cryptography-specific vocabulary, academic terms like "differential attack" will be divided into two words: "check score" and "attack." Therefore, the purpose of the formula calculation is to achieve adaptive word segmentation within a domain. Even without importing a specific domain vocabulary, we can still achieve the effect of separating domain terms. The specific principle is that these domain terms often appear in pairs in articles. Based on this, we use this formula to quantify this characteristic.

[0072] In one embodiment, the following steps are also included:

[0073] Set the threshold time_y for the number of times two adjacent words appear consecutively in a filtered text;

[0074] Determine whether the time value of the current two adjacent words is greater than the value of time_y. If so, calculate the cohesion score of the current two adjacent words.

[0075] In this embodiment, the co-occurrence threshold time_y is set to 2 to reduce the possibility of occasional phrases. This part screens potential segmentation objects with an occurrence frequency greater than 2, and then applies the cohesion algorithm based on them to reduce computational complexity.

[0076] In one embodiment, in S4, after a series of adjustments, a preset threshold of 0.5 for the secondary word frequency is selected. When the cohesion score of a combined word exceeds the threshold, the whole word is retained. Otherwise, the word combination is not performed. This achieves a balance between word segmentation accuracy and efficiency.

[0077] In one embodiment, S4 obtains the secondary word frequency statistics results, and the word segmentation is adjusted according to the results. During the specific execution, only the word segmentation results need to be modified, and there is no need to modify the original dictionary: for the words newly recognized by the secondary word segmentation, the corresponding combined words are added to the output results, and the number of times the two original adjacent words occur together (i.e., time) is subtracted.

[0078] In one embodiment, determining whether the secondary filtered text belongs to a certain field in S4 is achieved by checking whether the high-frequency words in the secondary filtered text appear in a predefined field dictionary. The specific steps include:

[0079] Load the domain dictionary into a collection, count the most frequent words in the text (e.g., the top 20), and then calculate the number of matches between these frequent words and the domain dictionary. If at least one domain word matches, the text is considered to belong to that domain; otherwise, it is considered not. This method, based on domain dictionary matching, is suitable for text classification scenarios with clear domain characteristics.

[0080] In one embodiment, the method further includes performing sentiment analysis on the secondary filtered text and outputting sentiment tendency conclusion information.

[0081] In one embodiment, the steps for creating a custom domain dictionary include:

[0082] Receive the domain dictionary of the required domain uploaded by the user in real time and perform storage operations; receive the matching instructions sent by the user in real time, and perform matching searches in the stored domain dictionary for the words whose frequency is higher than the preset threshold in the secondary word frequency statistics results; or

[0083] Receive the user's deletion instructions for the uploaded domain dictionary in real time and perform the deletion operation so that the currently deleted domain dictionary is excluded when executing the matching instructions sent by the user.

[0084] In one embodiment, the following steps are also included:

[0085] Receive domain keywords uploaded by users in real time and perform storage operations; receive matching instructions sent by users in real time, and select words with a frequency higher than a preset threshold in the secondary word frequency statistics results for matching with domain keywords; or

[0086] Receive the user's deletion instructions for the uploaded domain keywords in real time and perform the deletion operation so that the currently deleted domain keywords are excluded when executing the matching instructions sent by the user.

[0087] In one embodiment, the method further includes a step of displaying the domain scene to which the secondary filtered text and the input text belong.

[0088] The second aspect of the embodiment of the present invention further discloses a code change impact intelligent analysis system based on a filtering method based on text feature analysis according to the first aspect of the embodiment, comprising: a sensitive word filtering unit, a primary word segmentation unit, an adaptive word segmentation unit and a domain matching unit; wherein,

[0089] The sensitive word filtering unit is used to identify sensitive words in the received input text and delete the identified sensitive words from the input text to obtain a filtered text;

[0090] The initial word segmentation unit is used to perform initial word segmentation on a filtered text and perform initial word frequency statistics on the initial word segmentation results;

[0091] The adaptive word segmentation unit is used to perform adaptive word segmentation on the first filtered text according to the initial word frequency statistics, including: calculating the cohesion score of adjacent words in the initial word segmentation result, and determining the adjacent words to be merged based on the comparison result of the cohesion threshold and the cohesion score, and performing the merging operation to obtain the second filtered text;

[0092] The domain matching unit is used to perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, and perform matching searches in the custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

[0093] In one embodiment, it further includes a user interaction interface; the user interaction interface includes a text upload button, a word segmentation dictionary word addition button, a word segmentation dictionary word deletion button and a filter result display area; wherein,

[0094] The text upload button receives and responds to the input text upload instruction sent by the user, receives the input text and stores it;

[0095] The word segmentation dictionary word addition button receives and responds to the word segmentation dictionary word addition instruction sent by the user, and stores the added word segmentation dictionary word for being called by the initial word segmentation unit to perform the initial word segmentation operation, and / or for being called by the domain matching unit to perform the domain matching operation;

[0096] The word segmentation dictionary word deletion button receives and responds to the word segmentation dictionary word deletion instruction sent by the user, and deletes the stored word segmentation dictionary words;

[0097] The filter result display area is used to display the secondary filter text and the domain scenario to which the input text belongs.

[0098] In this embodiment, Figure 4 As shown, at the text upload button "Open File", the user can click it to select an existing text file in txt format to import.

[0099] After the program reads the file content, the system will automatically perform preliminary analysis and screening, and promptly display the results in the "Text Analysis Results" text box in the filter result display area.

[0100] The "Add Segmentation Dictionary Term" button and the "Delete Segmentation Dictionary Term" button allow users to modify the initial segmentation dictionary, giving users room for customization. This can also be developed to be called by the domain matching unit to perform domain matching operations.

[0101] At the "Exit Program" button, users can exit the program by clicking it.

[0102] The user interface is designed to allow users who are not familiar with the code compiler to use it directly, thereby lowering the usage threshold and improving user operation efficiency and satisfaction.

[0103] The second aspect of the embodiment of the present invention is used to execute all the steps of the first aspect of the embodiment.

[0104] The above is a detailed introduction to the filtering method and system based on text feature analysis provided by the present invention. In this embodiment, specific examples are used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

[0105] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in this embodiment may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown in this embodiment, but is intended to conform to the widest scope consistent with the principles and novel features disclosed in this embodiment.

Claims

1. A filtering method based on text feature analysis, characterized in that: The steps include: S1: Recognize sensitive words in the received input text, and delete the recognized sensitive words from the input text to obtain a filtered text; S2: performing initial word segmentation on the filtered text and performing initial word frequency statistics on the initial word segmentation results; S3: Based on the initial word frequency statistics, adaptively segmenting the first filtered text, including: calculating cohesion scores of adjacent words in the initial word segmentation results, and determining adjacent words to be merged based on a comparison result of a cohesion threshold and the cohesion scores, and performing a merging operation to obtain a second filtered text; S4: Perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, perform matching searches in a custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

2. A filtering method based on text feature analysis according to claim 1, characterized in that: The S1 comprises the following steps: S11: Recognize sensitive words in the received input text, and replace the recognized sensitive words with unified identifiers to obtain replaced text; S12: uniformly deleting the identifiers in the replaced text to obtain a filtered text.

3. The filtering method based on text feature analysis according to claim 1, characterized in that: The formula for calculating the cohesion score of adjacent words in the initial word segmentation result in S3 is: In the formula, score represents the cohesion score; time represents the number of times two adjacent words appear consecutively in a filtered text; len(ls) represents the total number of words in a filtered text; key_m represents the number of times the first word in the adjacent words appears in a filtered text, and key_n represents the number of times the second word in the adjacent words appears in a filtered text.

4. A filtering method based on text feature analysis according to claim 3, characterized in that: The following steps are also included: Set the threshold time_y for the number of times two adjacent words appear consecutively in a filtered text; Determine whether the time value of the current two adjacent words is greater than the value of time_y. If so, calculate the cohesion score of the current two adjacent words.

5. The filtering method based on text feature analysis according to claim 1, characterized in that: The method also includes the step of performing sentiment analysis on the secondary filtered text and outputting sentiment tendency conclusion information.

6. The filtering method based on text feature analysis according to claim 1, characterized in that: The steps of operating the customized domain dictionary include: Receive the domain dictionary of the required domain uploaded by the user in real time and perform storage operations; receive the matching instructions sent by the user in real time, and perform matching searches in the stored domain dictionary for the words whose frequency is higher than the preset threshold in the secondary word frequency statistics results; or Receive the user's deletion instructions for the uploaded domain dictionary in real time and perform the deletion operation so that the currently deleted domain dictionary is excluded when executing the matching instructions sent by the user.

7. The filtering method based on text feature analysis according to claim 1, characterized in that: The following steps are also included: Receive domain keywords uploaded by users in real time and perform storage operations; receive matching instructions sent by users in real time, and select words with a frequency higher than a preset threshold in the secondary word frequency statistics results for matching with domain keywords; or Receive the user's deletion instructions for the uploaded domain keywords in real time and perform the deletion operation so that the currently deleted domain keywords are excluded when executing the matching instructions sent by the user.

8. The filtering method based on text feature analysis according to claim 1, characterized in that: The method also includes the step of displaying the domain scene to which the secondary filtered text and the input text belong.

9. A filtering system based on text feature analysis of a filtering method based on text feature analysis according to any one of claims 1 to 8, characterized in that: include: Sensitive word filtering unit, initial word segmentation unit, adaptive word segmentation unit and domain matching unit; among them, The sensitive word filtering unit is used to identify sensitive words in the received input text and delete the identified sensitive words from the input text to obtain a filtered text; The primary word segmentation unit is used to perform primary word segmentation on the primary filtered text and perform primary word frequency statistics on the primary word segmentation result; The adaptive word segmentation unit is used to perform adaptive word segmentation on the first filtered text according to the initial word frequency statistics result, including: calculating the cohesion score of adjacent words in the initial word segmentation result, and determining the adjacent words to be merged based on the comparison result of the cohesion threshold and the cohesion score, and performing the merging operation to obtain the second filtered text; The domain matching unit is used to perform secondary word frequency statistics on the secondary filtered text, filter out words with a frequency higher than a preset threshold in the secondary word frequency statistics results, perform matching searches in a custom domain dictionary, and determine the domain scenario to which the input text belongs based on the matching results.

10. The filtering system based on text feature analysis according to claim 9, characterized in that: It also includes a user interaction interface; the user interaction interface includes a text upload button, a word segmentation dictionary word add button, a word segmentation dictionary word delete button and a filter result display area; wherein, The text upload button receives and responds to the input text upload instruction sent by the user, receives the input text and stores it; The word segmentation dictionary word addition button receives and responds to a word segmentation dictionary word addition instruction sent by a user, and stores the added word segmentation dictionary word for being called by the initial word segmentation unit to perform an initial word segmentation operation, and / or for being called by the domain matching unit to perform a domain matching operation; The word segmentation dictionary word deletion button receives and responds to the word segmentation dictionary word deletion instruction sent by the user, and deletes the stored word segmentation dictionary words; The filtering result display area is used to display the domain scenario to which the secondary filtered text and the input text belong.

Citation Information

Patent Citations

  • Text filtering method and device and computer storage medium

    CN112818110A

  • Text review method and device, equipment and storage medium

    CN118378631A

  • Text processing method and device and related equipment

    CN119940358A