A method and system for constructing a sensitive word library based on natural language processing

Through natural language processing technology, identify and analyze sensitive words in text and determine their topic types and sensitivity values, it solves the problem that traditional methods are difficult to fully cover sensitive words and comprehension context, and achieves more efficient and accurate construction of sensitive thesaurus.

CN118885556BActive Publication Date: 2025-05-13BEIJING ZHONGKE RUITU TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410923783.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-05-13
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

Traditional sensitive thesaurus construction methods rely on manual experience and rules, making it difficult to fully cover various possible sensitive words and expressions, and lack of understanding of the context and context, resulting in timeless access to sensitive content.

Method used

Using a natural language processing method, by acquiring and preprocessing text data, identifying sensitive words and their characteristic information, analyzing characteristic information to determine topic types, classifying sensitive words, evaluating sensitivity values, and determining storage areas based on sensitivity values ​​to construct sensitive thesaurus.

Benefits of technology

It improves the efficiency and accuracy of sensitive word recognition, can cover various sensitive words and their characteristics more comprehensively, adapt to different contexts and expression methods, and achieve faster and more accurate construction of sensitive thesaurus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118885556B_ABST
    Figure CN118885556B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for constructing a sensitive word library based on natural language processing, which includes: obtaining and preprocessing text data containing sensitive words, and obtaining sensitive words from the text data through natural language processing technology, and determining characteristic information of the sensitive words; determining the subject type corresponding to the characteristic information of the sensitive words, and classifying the sensitive words according to the subject type to obtain a sensitive word set; determining the sensitivity evaluation index of the subject type corresponding to the sensitive word set, and evaluating the sensitivity value of the subject type corresponding to the sensitive word set according to the sensitivity evaluation index; determining the storage area of ​​the sensitive word set according to the sensitivity value, storing the sensitive word set belonging to the same sensitivity value in the corresponding storage area, and completing the construction of the sensitive word library. The present invention determines the sensitive words in the text through natural language processing technology, improves the accuracy and efficiency of sensitive word recognition, and determines the construction logic of its word library through comprehensive analysis of the sensitive word situation, so as to more accurately construct a sensitive word library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of word library construction, and in particular to a sensitive word library construction method and system based on natural language processing. Background Art

[0002] The construction of sensitive word libraries is to effectively manage and filter inappropriate content that may appear on digital platforms. The motivations behind this construction include protecting users from insults, abuse, discrimination, pornography and other bad content, improving users' experience and sense of security on the Internet, meeting compliance requirements of laws and regulations, maintaining corporate brand image and protecting the network security of minors. By building sensitive word libraries, platforms and applications can better manage user-generated content, ensure that communication and interaction on the platform are more positive and healthy, and thus create a safer and more friendly digital environment.

[0003] However, the traditional method of constructing a sensitive word library is usually based on manual experience and rules. This method is easily limited by personal experience and subjective judgment, and it is difficult to fully cover all possible sensitive words and expressions. At the same time, the traditional sensitive word library is difficult to adapt to sensitive content in different contexts and expressions, and lacks the ability to understand and analyze context and context, resulting in the inability to obtain sensitive content in a timely manner. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a method and system for constructing a sensitive word library based on natural language processing, including:

[0005] Obtain text data containing sensitive words and preprocess the text data;

[0006] Obtain sensitive words from pre-processed text data through natural language processing technology and determine the characteristic information of sensitive words;

[0007] Analyze the characteristic information of sensitive words and determine the subject type of sensitive words;

[0008] Classify sensitive words according to topic type, and classify sensitive words of the same topic type into the same sensitive word set;

[0009] Determine the sensitivity evaluation index of the topic type corresponding to the sensitive word set, and evaluate the sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index;

[0010] The storage area of ​​the sensitive word set is determined according to the sensitivity value, and the sensitive word set belonging to the same sensitivity value is stored in the corresponding storage area to complete the construction of the sensitive word library.

[0011] Furthermore, the preprocessing of the text data includes:

[0012] Clean the text data and remove non-text content in the text, including special characters, punctuation marks, and HTML tags;

[0013] Split the text in the cleaned text data into words or phrases.

[0014] Furthermore, the method of obtaining sensitive words from the pre-processed text data by natural language processing technology and determining feature information of the sensitive words includes:

[0015] Identify sensitive words and their original feature information in pre-processed text data through natural language processing technology;

[0016] The correlation between the sensitive word and its original feature information is calculated, and the original feature information with a correlation greater than a preset threshold is screened out as the feature information of the sensitive word.

[0017] Furthermore, the analyzing of the characteristic information of the sensitive words to determine the subject type of the sensitive words includes:

[0018] Obtain feature information corresponding to sensitive words, and determine feature meaning and information meaning corresponding to each feature information;

[0019] Obtain known topic types of sensitive words, and respectively calculate the correlation between the feature meaning and information meaning corresponding to each feature information and each known topic type to obtain a first correlation and a second correlation;

[0020] Setting fixed weight coefficients for the first correlation and the second correlation respectively, and performing weighted addition calculation on the first correlation and the second correlation and their weight coefficients to obtain a meaning association value of the feature information;

[0021] The meaning association values ​​of all feature information corresponding to each sensitive word are added together to obtain the association degree value of the sensitive word, and the known topic type with the highest association degree value is used as the topic type of the sensitive word.

[0022] Furthermore, the sensitive words are classified according to the subject type, and the sensitive words belonging to the same subject type are classified into the same sensitive word set, including:

[0023] Each sensitive word is labeled with the corresponding topic type, and the sensitive words are classified according to the labels using a preset label classification model, and sensitive words belonging to the same topic type are classified into the same sensitive word set.

[0024] Furthermore, determining a sensitivity evaluation index of a topic type corresponding to the sensitive word set, and evaluating a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index, includes:

[0025] Determine the keywords of the topic type corresponding to the sensitive word set, and analyze the semantic sensitivity, social sensitivity, and historical sensitivity of the keywords;

[0026] Determine semantic sensitivity, social sensitivity, and historical sensitivity as sensitive evaluation indicators of the topic type corresponding to the sensitive word set, and calculate the relevance of each sensitive evaluation indicator to the topic type corresponding to the sensitive word set;

[0027] Calculate the sum of the relevance of all sensitive evaluation indicators, and use the ratio of the relevance of each sensitive evaluation indicator to the sum as the weight coefficient of each sensitive evaluation indicator;

[0028] According to each sensitive evaluation index, the topic type corresponding to the sensitive word set is evaluated and valued, and the evaluation value corresponding to each sensitive evaluation index is obtained;

[0029] The evaluation value and weight coefficient corresponding to each sensitive evaluation indicator are weightedly added to obtain the sensitivity value of the topic type corresponding to the sensitive word set.

[0030] Furthermore, determining the storage area of ​​the sensitive word set according to the sensitivity value includes:

[0031] A storage area-sensitivity value interval correspondence relationship is preset, and the storage area-sensitivity value interval correspondence relationship is associated with a corresponding storage area for each sensitivity value interval of the sensitive word set;

[0032] The sensitivity value of the sensitive word set is obtained, and based on the mapping relationship between the sensitivity value interval to which the sensitivity value belongs and the storage area corresponding to the sensitive word set in the storage area-sensitivity value interval correspondence relationship, the storage area corresponding to the sensitive word set is determined.

[0033] The present invention also provides a sensitive word library construction system based on natural language processing, comprising:

[0034] The acquisition module is used to obtain text data containing sensitive words and pre-process the text data;

[0035] A determination module is used to obtain sensitive words from pre-processed text data through natural language processing technology and determine feature information of sensitive words;

[0036] An analysis module is used to analyze the characteristic information of sensitive words and determine the subject type of sensitive words;

[0037] A classification module is used to classify sensitive words according to the subject type and classify sensitive words belonging to the same subject type into the same sensitive word set;

[0038] An evaluation module, used to determine a sensitivity evaluation index of a topic type corresponding to a sensitive word set, and evaluate a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index;

[0039] The storage module is used to determine the storage area of ​​the sensitive word set according to the sensitivity value, store the sensitive word set belonging to the same sensitivity value in the corresponding storage area, and complete the construction of the sensitive word library.

[0040] Compared with the prior art, the method and system for constructing a sensitive word library based on natural language processing in the embodiment of the present invention have the following beneficial effects:

[0041] The present invention uses natural language processing technology to determine sensitive words in a text, without spending a lot of manpower and time, and can accurately and quickly identify sensitive words, thereby improving the efficiency of sensitive word identification. It also comprehensively analyzes the sensitive word situation, covers various sensitive words and their characteristics as much as possible, to adapt to sensitive content in different situations, and comprehensively analyzes the logical relationship between sensitive words, determines the construction logic of its word library, and constructs a sensitive word library faster and more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic diagram of the process structure of a method for building a sensitive word library based on natural language processing in an embodiment of the present invention;

[0043] Figure 2 It is a schematic diagram of the composition of a sensitive word library construction system based on natural language processing in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0045] In the description of the present application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the platform or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0046] The terms "second" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, a feature defined with "second" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more.

[0047] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0048] like Figure 1 As shown, in an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, including: S100: obtaining text data containing sensitive words and preprocessing the text data; S200: obtaining sensitive words from the preprocessed text data through natural language processing technology, and determining the characteristic information of the sensitive words; S300: analyzing the characteristic information of the sensitive words and determining the subject type of the sensitive words; S400: classifying the sensitive words according to the subject type, and classifying the sensitive words belonging to the same subject type into the same sensitive word set; S500: determining the sensitivity evaluation index of the subject type corresponding to the sensitive word set, and evaluating the sensitivity value of the subject type corresponding to the sensitive word set according to the sensitivity evaluation index; S600: determining the storage area of ​​the sensitive word set according to the sensitivity value, storing the sensitive word set belonging to the same sensitivity value in the corresponding storage area, and completing the construction of the sensitive word library.

[0049] Furthermore, the present invention uses natural language processing technology to determine sensitive words in the text, without spending a lot of manpower and time, and can accurately and quickly identify sensitive words, thereby improving the efficiency of sensitive word identification. It also comprehensively analyzes the sensitive word situation and covers various sensitive words and their characteristics as much as possible to adapt to sensitive content in different situations. It also comprehensively analyzes the logical relationship between sensitive words and determines the construction logic of its word library, so as to construct a sensitive word library faster and more accurately.

[0050] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein the preprocessing of text data includes: cleaning the text data to remove non-text content in the text, the non-text content including special characters, punctuation marks, and HTML tags; and segmenting the text in the cleaned text data into words or phrases.

[0051] Specifically, regular expressions are used to match and remove special characters, punctuation marks, and HTML tags in the text. Removing special characters and punctuation marks can make the text cleaner and easier to process, which helps with subsequent text analysis. Removing HTML tags can convert the text from web page format to plain text format, facilitating subsequent text processing and analysis. The cleaned text data is segmented into words or phrases through a word segmenter, providing a basis for subsequent text analysis and processing.

[0052] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein sensitive words are obtained from preprocessed text data through natural language processing technology, and feature information of the sensitive words is determined, including: identifying sensitive words and their original feature information in the preprocessed text data through natural language processing technology; calculating the correlation between the sensitive words and their original feature information, and screening out the original feature information with a correlation greater than a preset threshold as the feature information of the sensitive words.

[0053] Specifically, natural language processing techniques, such as text classification, named entity recognition (NER) or sentiment analysis, are used to analyze the preprocessed text data and identify sensitive words therein. Once sensitive words are identified, contextual information of these words in the original text can be extracted, such as the words or phrases surrounding the sensitive words, and the sentences or paragraphs in which the sensitive words appear, which will be used as the original feature information of the sensitive words. Text similarity calculation methods in natural language processing techniques, such as cosine similarity, are used to calculate the correlation between sensitive words and their original feature information. For each sensitive word, its correlation score with the original feature information is calculated. A correlation threshold is set to filter out the original feature information with a correlation greater than the threshold as the feature information of the sensitive words.

[0054] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein the feature information of the sensitive words is analyzed to determine the subject type of the sensitive words, including: obtaining feature information corresponding to the sensitive words, and determining the feature meaning and information meaning corresponding to each feature information; obtaining known subject types of the sensitive words, and respectively calculating the correlation between the feature meaning and information meaning corresponding to each feature information and each known subject type to obtain a first correlation and a second correlation; respectively setting fixed weight coefficients for the first correlation and the second correlation, and performing weighted addition calculation on the first correlation and the second correlation and their weight coefficients to obtain a meaning association value of the feature information; adding the meaning association values ​​of all feature information corresponding to each sensitive word to obtain a correlation degree value of the sensitive word, and taking the known subject type with the highest correlation degree value as the subject type of the sensitive word.

[0055] Specifically, each sensitive word is traversed to obtain its corresponding feature information; for each feature information, its meaning in the text needs to be determined, including feature meaning and information meaning; for each sensitive word, its known topic type is obtained, and these topic types are pre-defined sensitive word classifications; using a text similarity calculation method, the correlation between the feature meaning and information meaning corresponding to each feature information and each known topic type is calculated to obtain a first correlation and a second correlation; fixed weight coefficients are set for the first correlation and the second correlation respectively, and then the first correlation and the second correlation and their weight coefficients are weightedly added to obtain a meaning association value of the feature information; the meaning association values ​​of all feature information corresponding to each sensitive word are added to obtain a correlation degree value of the sensitive word, and the known topic type with the highest correlation degree value is used as the topic type of the sensitive word.

[0056] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein sensitive words are classified according to topic types, and sensitive words belonging to the same topic type are classified into the same sensitive word set, including: labeling each sensitive word with a label corresponding to the topic type, classifying the sensitive words according to the labels through a preset label classification model, and classifying sensitive words belonging to the same topic type into the same sensitive word set.

[0057] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein the method determines a sensitive evaluation index of a topic type corresponding to a sensitive word set, and evaluates a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitive evaluation index, including: determining keywords of the topic type corresponding to the sensitive word set, and analyzing the semantic sensitivity, social sensitivity, and historical sensitivity of the keywords; determining the semantic sensitivity, social sensitivity, and historical sensitivity as sensitive evaluation indexes of the topic type corresponding to the sensitive word set, and calculating the relevance of each sensitive evaluation index to the topic type corresponding to the sensitive word set; calculating the sum of the relevances of all sensitive evaluation indicators, and taking the ratio of the relevance of each sensitive evaluation indicator to the sum as the weight coefficient of each sensitive evaluation indicator; evaluating and valuing the topic type corresponding to the sensitive word set according to each sensitive evaluation indicator, and obtaining an evaluation value corresponding to each sensitive evaluation indicator; and performing weighted addition calculation on the evaluation value and the weight coefficient corresponding to each sensitive evaluation indicator, and obtaining a sensitivity value of the topic type corresponding to the sensitive word set.

[0058] Specifically, for each sensitive word set, determine the keywords of the corresponding topic type, and these keywords are obtained through the keyword extraction algorithm in natural language processing technology; for each keyword, perform semantic sensitivity, social sensitivity and historical sensitivity analysis, semantic sensitivity refers to whether the intrinsic meaning of the word is sensitive, social sensitivity refers to the sensitivity of the word in the social environment, and historical sensitivity refers to the sensitivity of the word in historical culture; determine semantic sensitivity, social sensitivity and historical sensitivity as sensitive evaluation indicators of the topic type corresponding to the sensitive word set, and use the text similarity calculation method to calculate the relevance of each sensitive evaluation indicator to the topic type corresponding to the sensitive word set; calculate the sum of the relevance of all sensitive evaluation indicators, and use the ratio of the relevance of each sensitive evaluation indicator to the total as the weight coefficient of each sensitive evaluation indicator; evaluate and take values ​​for the topic type corresponding to the sensitive word set according to each sensitive evaluation indicator, and obtain the evaluation value corresponding to each sensitive evaluation indicator; perform weighted addition calculation on the evaluation value and weight coefficient corresponding to each sensitive evaluation indicator, and obtain the sensitivity value of the topic type corresponding to the sensitive word set.

[0059] In an embodiment of the present application, a method for constructing a sensitive word library based on natural language processing is provided, wherein the storage area of ​​a sensitive word set is determined according to a sensitivity value, including: presetting a storage area-sensitivity value interval correspondence relationship, wherein the storage area-sensitivity value interval correspondence relationship is associated with a corresponding storage area for each sensitivity value interval of the sensitive word set; obtaining the sensitivity value of the sensitive word set, and determining the storage area corresponding to the sensitive word set based on a mapping relationship between the sensitivity value interval to which the sensitivity value belongs within the storage area-sensitivity value interval correspondence relationship.

[0060] Specifically, a sensitivity value interval is defined and associated with a corresponding storage area. For example, the sensitivity value interval is defined as [0, 0.2] for low sensitivity, (0.2, 0.5] for medium sensitivity, and (0.5, 1] ​​for high sensitivity, and a corresponding storage area is specified for each sensitivity value interval. Based on the mapping relationship between the sensitivity value interval to which the sensitivity value belongs in the storage area-sensitivity value interval correspondence relationship, the storage area corresponding to the sensitive word set is determined. For example, if the sensitivity value of a sensitive word set is 0.3, it will be mapped to the medium sensitivity interval according to the interval correspondence, thereby determining the corresponding storage area.

[0061] like Figure 2As shown, in an embodiment of the present application, a system for constructing a sensitive word library based on natural language processing is provided, including: an acquisition module, used to acquire text data containing sensitive words and preprocess the text data; a determination module, used to acquire sensitive words from the preprocessed text data through natural language processing technology, and determine the characteristic information of the sensitive words; an analysis module, used to analyze the characteristic information of the sensitive words and determine the subject type of the sensitive words; a classification module, used to classify the sensitive words according to the subject type, and classify the sensitive words belonging to the same subject type into the same sensitive word set; an evaluation module, used to determine the sensitivity evaluation index of the subject type corresponding to the sensitive word set, and evaluate the sensitivity value of the subject type corresponding to the sensitive word set according to the sensitivity evaluation index; a storage module, used to determine the storage area of ​​the sensitive word set according to the sensitivity value, store the sensitive word set belonging to the same sensitivity value in the corresponding storage area, and complete the construction of the sensitive word library.

[0062] In summary, the embodiment of the present invention provides a method and system for constructing a sensitive word library based on natural language processing, which includes: obtaining and preprocessing text data containing sensitive words, and obtaining sensitive words from it through natural language processing technology, and determining the characteristic information of sensitive words; determining the topic type corresponding to the sensitive word characteristic information, and classifying the sensitive words according to the topic type to obtain a sensitive word set; determining the sensitivity evaluation index of the sensitive word set corresponding to the topic type, and evaluating the sensitivity value of the sensitive word set corresponding to the topic type according to the sensitivity evaluation index; determining the storage area of ​​the sensitive word set according to the sensitivity value, and storing the sensitive word set belonging to the same sensitivity value in the corresponding storage area to complete the construction of the sensitive word library. The present invention determines the sensitive words in the text through natural language processing technology, improves the accuracy and efficiency of sensitive word recognition, and determines the construction logic of its word library through comprehensive analysis of the sensitive word situation, so as to more accurately construct a sensitive word library.

[0063] Finally, it should be noted that it is obvious that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

[0064] The above is only an example of implementation of the present invention, but it cannot be used to limit the scope of the present invention. Any structural changes made according to the present invention, as long as they do not lose the essence of the present invention, should be regarded as falling within the scope of protection of the present invention and being restricted. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and related instructions of the platform described above can refer to the corresponding process in the aforementioned platform embodiment, and will not be repeated here.

[0065] The term "comprises" or any other similar term is intended to cover a non-exclusive inclusion such that a process, platform, article, or apparatus / platform that includes a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, platform, article, or apparatus / platform.

[0066] So far, the technical solutions of the present invention have been described in conjunction with the further embodiments shown in the accompanying drawings, but it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0067] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A method for constructing a sensitive word library based on natural language processing, characterized in that: include: Obtain text data containing sensitive words and preprocess the text data; Obtain sensitive words from pre-processed text data through natural language processing technology and determine the characteristic information of sensitive words; Analyze the characteristic information of sensitive words and determine the subject type of sensitive words; Classify sensitive words according to topic type, and classify sensitive words of the same topic type into the same sensitive word set; Determine the sensitivity evaluation index of the topic type corresponding to the sensitive word set, and evaluate the sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index; Determine the storage area of ​​the sensitive word set according to the sensitivity value, store the sensitive word sets with the same sensitivity value in the corresponding storage area, and complete the construction of the sensitive word library; The preprocessing of text data includes: Clean the text data and remove non-text content in the text, including special characters, punctuation marks, and HTML tags; Segment the text in the cleaned text data into words or phrases; The method of obtaining sensitive words from the pre-processed text data by natural language processing technology and determining feature information of the sensitive words includes: Identify sensitive words and their original feature information in pre-processed text data through natural language processing technology; Calculate the correlation between the sensitive word and its original feature information, and select the original feature information with a correlation greater than a preset threshold as the feature information of the sensitive word; The analyzing the characteristic information of the sensitive words to determine the subject type of the sensitive words includes: Obtain feature information corresponding to sensitive words, and determine feature meaning and information meaning corresponding to each feature information; Obtain known topic types of sensitive words, and respectively calculate the correlation between the feature meaning and information meaning corresponding to each feature information and each known topic type to obtain a first correlation and a second correlation; Setting fixed weight coefficients for the first correlation and the second correlation respectively, and performing weighted addition calculation on the first correlation and the second correlation and their weight coefficients to obtain a meaning association value of the feature information; The meaning association values ​​of all feature information corresponding to each sensitive word are added together to obtain the association degree value of the sensitive word, and the known topic type with the highest association degree value is used as the topic type of the sensitive word; The sensitive words are classified according to the subject type, and the sensitive words belonging to the same subject type are classified into the same sensitive word set, including: Each sensitive word is labeled with a corresponding topic type, and the sensitive words are classified according to the labels using a preset label classification model, and sensitive words belonging to the same topic type are classified into the same sensitive word set; The step of determining a sensitivity evaluation index of a topic type corresponding to the sensitive word set, and evaluating a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index, includes: Determine the keywords of the topic type corresponding to the sensitive word set, and analyze the semantic sensitivity, social sensitivity, and historical sensitivity of the keywords; Determine semantic sensitivity, social sensitivity, and historical sensitivity as sensitive evaluation indicators of the topic type corresponding to the sensitive word set, and calculate the relevance of each sensitive evaluation indicator to the topic type corresponding to the sensitive word set; Calculate the sum of the relevance of all sensitive evaluation indicators, and use the ratio of the relevance of each sensitive evaluation indicator to the sum as the weight coefficient of each sensitive evaluation indicator; According to each sensitive evaluation index, the topic type corresponding to the sensitive word set is evaluated and valued, and the evaluation value corresponding to each sensitive evaluation index is obtained; The evaluation value and weight coefficient corresponding to each sensitive evaluation indicator are weighted and added together to obtain the sensitivity value of the topic type corresponding to the sensitive word set; Determining the storage area of ​​the sensitive word set according to the sensitivity value includes: A storage area-sensitivity value interval correspondence relationship is preset, and the storage area-sensitivity value interval correspondence relationship is associated with a corresponding storage area for each sensitivity value interval of the sensitive word set; The sensitivity value of the sensitive word set is obtained, and based on the mapping relationship between the sensitivity value interval to which the sensitivity value belongs and the storage area corresponding to the sensitive word set in the storage area-sensitivity value interval correspondence relationship, the storage area corresponding to the sensitive word set is determined.

2. A sensitive word library construction system based on natural language processing, characterized in that: include: The acquisition module is used to obtain text data containing sensitive words and pre-process the text data; A determination module is used to obtain sensitive words from pre-processed text data through natural language processing technology and determine feature information of sensitive words; An analysis module is used to analyze the characteristic information of sensitive words and determine the subject type of sensitive words; A classification module is used to classify sensitive words according to the subject type and classify sensitive words belonging to the same subject type into the same sensitive word set; An evaluation module, used to determine a sensitivity evaluation index of a topic type corresponding to a sensitive word set, and evaluate a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index; A storage module is used to determine the storage area of ​​the sensitive word set according to the sensitivity value, store the sensitive word set belonging to the same sensitivity value in the corresponding storage area, and complete the construction of the sensitive word library; The preprocessing of text data includes: Clean the text data and remove non-text content in the text, including special characters, punctuation marks, and HTML tags; Segment the text in the cleaned text data into words or phrases; The method of obtaining sensitive words from the pre-processed text data by natural language processing technology and determining feature information of the sensitive words includes: Identify sensitive words and their original feature information in pre-processed text data through natural language processing technology; Calculate the correlation between the sensitive word and its original feature information, and select the original feature information with a correlation greater than a preset threshold as the feature information of the sensitive word; The analyzing the characteristic information of the sensitive words to determine the subject type of the sensitive words includes: Obtain feature information corresponding to sensitive words, and determine feature meaning and information meaning corresponding to each feature information; Obtain known topic types of sensitive words, and respectively calculate the correlation between the feature meaning and information meaning corresponding to each feature information and each known topic type to obtain a first correlation and a second correlation; Setting fixed weight coefficients for the first correlation and the second correlation respectively, and performing weighted addition calculation on the first correlation and the second correlation and their weight coefficients to obtain a meaning association value of the feature information; The meaning association values ​​of all feature information corresponding to each sensitive word are added together to obtain the association degree value of the sensitive word, and the known topic type with the highest association degree value is used as the topic type of the sensitive word; The sensitive words are classified according to the subject type, and the sensitive words belonging to the same subject type are classified into the same sensitive word set, including: Each sensitive word is labeled with a corresponding topic type, and the sensitive words are classified according to the labels using a preset label classification model, and sensitive words belonging to the same topic type are classified into the same sensitive word set; The step of determining a sensitivity evaluation index of a topic type corresponding to the sensitive word set, and evaluating a sensitivity value of the topic type corresponding to the sensitive word set according to the sensitivity evaluation index, includes: Determine the keywords of the topic type corresponding to the sensitive word set, and analyze the semantic sensitivity, social sensitivity, and historical sensitivity of the keywords; Determine semantic sensitivity, social sensitivity, and historical sensitivity as sensitive evaluation indicators of the topic type corresponding to the sensitive word set, and calculate the relevance of each sensitive evaluation indicator to the topic type corresponding to the sensitive word set; Calculate the sum of the relevance of all sensitive evaluation indicators, and use the ratio of the relevance of each sensitive evaluation indicator to the sum as the weight coefficient of each sensitive evaluation indicator; According to each sensitive evaluation index, the topic type corresponding to the sensitive word set is evaluated and valued, and the evaluation value corresponding to each sensitive evaluation index is obtained; The evaluation value and weight coefficient corresponding to each sensitive evaluation indicator are weighted and added together to obtain the sensitivity value of the topic type corresponding to the sensitive word set; Determining the storage area of ​​the sensitive word set according to the sensitivity value includes: A storage area-sensitivity value interval correspondence relationship is preset, and the storage area-sensitivity value interval correspondence relationship is associated with a corresponding storage area for each sensitivity value interval of the sensitive word set; The sensitivity value of the sensitive word set is obtained, and based on the mapping relationship between the sensitivity value interval to which the sensitivity value belongs and the storage area corresponding to the sensitive word set in the storage area-sensitivity value interval correspondence relationship, the storage area corresponding to the sensitive word set is determined.

Citation Information

Patent Citations

  • Sensitive word auditing method based on large language model, storage medium and electronic equipment

    CN116720515A