Sensitive word determination method and device and electronic equipment

By using emotion scores and feature vector input recognition models in the sensitive word recognition system, the problems of emerging sensitive words being ignored and difficult to understand context are solved, and sensitive word recognition with high accuracy is achieved.

CN120218064APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330412.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Emerging sensitive words in the prior art may be ignored, and the different meanings of words in different contexts can be understood, resulting in misjudgment or misjudgment.

Method used

By determining the emotional score of the text to be detected and inputting the emotion score and feature vectors into the pre-trained recognition model, the sensitive words contained in the text to be detected are determined based on the output of the recognition model.

Benefits of technology

The accurate and effective identification of sensitive words is achieved, which reduces the risk of emerging sensitive words being ignored, and understands the meaning in the text to be detected through emotional scores, which improves the accuracy of sensitive words recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218064A_ABST
    Figure CN120218064A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a sensitive word determination method and apparatus, and an electronic device, and are used for solving the problem of erroneous judgment or missed judgment caused by incapability of understanding different meanings of words in different contexts since emerging sensitive words may be neglected in related technologies. In the embodiment of the invention, the electronic equipment determines the emotion score of the to-be-detected text; wherein the emotion score is used for representing the emotion tendency of the to-be-detected text; the sentiment score and the feature vector, determined based on the encoder, of the to-be-detected text are input into the recognition model, the sensitive words contained in the to-be-detected text are determined based on output of the recognition model, the sensitive words can be accurately and effectively recognized through the recognition model which is trained in advance, the risk that emerging sensitive words are neglected is reduced, and the recognition efficiency is improved. By determining the emotion score of the to-be-detected text, the recognition model can understand the meaning in the to-be-detected text, and then sensitive words are accurately and effectively recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to technical fields such as natural language processing, machine learning, and data mining, and particularly relates to a method, apparatus, and electronic device for determining sensitive words. Background Art

[0002] Numerous industries and organizations are strictly bound by legal and regulatory frameworks, and it is crucial to ensure the compliance of communication content and document information. The application of sensitive word recognition technology can effectively assist organizations in following these regulations, thereby avoiding potential legal risks. At the same time, on platforms for user-generated content such as social media, forums, and blogs, sensitive words indicating inappropriate remarks, harassment, bullying, and other bad behaviors occur from time to time. By accurately identifying and filtering these sensitive words, it not only helps to maintain good order on the platform but also effectively protects the dignity and rights of users. Therefore, the implementation of sensitive word recognition is particularly crucial and necessary.

[0003] In related technologies, in order to identify sensitive words, first, a static sensitive word library is established, which contains various words and phrases considered to need filtering or monitoring. These words can be formulated based on laws and regulations, social norms, or company policies, etc. Through a simple string matching algorithm, it is checked whether the words in the sensitive word library are included in the text to be detected. Or, the sensitive words in the text to be detected are searched for manually. However, when searching for sensitive words based on the sensitive word library, since it is unable to update and adapt to newly emerging sensitive words and phrases in a timely manner, this may lead to the neglect of emerging sensitive words. Moreover, the meaning of sensitive words is often affected by the context. Traditional detection methods usually cannot understand the different meanings of words in different contexts, resulting in misjudgment or missed judgment. For example, certain words are neutral in specific situations, while they may have sensitive meanings in other situations. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, and electronic device for determining sensitive words, which are used to solve the problems in related technologies that emerging sensitive words may be ignored and the different meanings of words in different contexts cannot be understood, resulting in misjudgment or missed judgment.

[0005] In a first aspect, embodiments of this application provide a method for determining sensitive words, and the method includes:

[0006] Determine the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected;

[0007] Input the sentiment score and the feature vector of the determined text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0008] Second aspect, an embodiment of the present application further provides a sensitive word determination device, and the device includes:

[0009] A determination module, configured to determine an emotion score of the text to be detected; wherein, the emotion score is used to represent the emotion tendency of the text to be detected;

[0010] A processing module, configured to input the emotion score and the feature vector of the text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0011] Third aspect, an embodiment of the present application further provides an electronic device, including:

[0012] A memory, configured to store program instructions;

[0013] A processor, configured to call the program instructions stored in the memory and execute the steps included in the method for determining sensitive words according to any one of the above-mentioned program instructions.

[0014] In the embodiment of the present application, the electronic device determines the emotion score of the text to be detected; wherein, the emotion score is used to represent the emotion tendency of the text to be detected; and inputs the emotion score and the feature vector of the text to be detected determined based on the encoder into the recognition model, and determines the sensitive words included in the text to be detected based on the output of the recognition model. The pre-trained recognition model can accurately and effectively identify sensitive words, reducing the risk of emerging sensitive words being ignored. Moreover, by determining the emotion score of the text to be detected, the recognition model can understand the meaning in the text to be detected, and then accurately and effectively identify sensitive words. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic diagram of a process for determining sensitive words provided by an embodiment of the present application;

[0017] Figure 2 It is a detailed schematic diagram of a process for determining sensitive words provided by an embodiment of the present application;

[0018] Figure 3 It is a structural diagram of a sensitive word determination device provided by an embodiment of the present application;

[0019] Figure 4Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application shall fall within the protection scope of the present application.

[0021] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in this development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.

[0022] In the present application, referring to "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0023] The "connection", "connection", "coupling" and other similar terms involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in the present application refers to two or more. The "and / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0024] In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of the technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0025] To accurately and effectively identify sensitive words, embodiments of the present application provide a method, apparatus, and electronic device for determining sensitive words.

[0026] The method for determining sensitive words includes: the electronic device determines the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected; the sentiment score and the feature vector of the text to be detected are input into a pre-trained recognition model, and based on the output of the recognition model, the sensitive words included in the text to be detected are determined.

[0027] Embodiment 1:

[0028] Figure 1 It is a process schematic diagram of a method for determining sensitive words provided by an embodiment of the present application. The process includes the following steps:

[0029] S101: Determine the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected.

[0030] The method for determining sensitive words provided by the embodiments of the present application is applied to an electronic device, which can be an intelligent device such as a PC or a server, and specifically can be a World Wide Web (Web) server.

[0031] To accurately and effectively determine sensitive words, the electronic device can first receive the text to be detected that needs to be detected. Among them, the text to be detected may come from platforms such as social media, comment areas, and forums where user-generated content is located. Specifically, the data input / collection device will, according to the upstream service configuration, crawl data from the specified platforms and websites and extract the text to be detected from them. In this process, the following two methods can be used: Using a crawler tool: Using the Scrapy library to regularly crawl the specified platforms and websites. The crawled content is extensive, including but not limited to user comments, posts, news titles, etc. Application Programming Interface (API) access: Supports real-time data crawling through API interface calls to ensure the timeliness of the data. The data input / collection device usually runs on a cloud server or a local server. After successfully obtaining the text to be detected, the device will send the text to the specified electronic device for subsequent sensitive word detection work.

[0032] In addition, it can also be that when the user has a need for sensitive word detection, the user selects the text to be detected on the preset page of the device used by the user or the preset device and clicks the preset button, such as the "detection" button, then the electronic device can also obtain the text to be detected.

[0033] After obtaining the text to be detected, the electronic device can perform sentiment analysis on the text to be detected, add sentiment features, and help the recognition model understand the context of sensitive words. Specifically, the electronic device can use a sentiment analysis process, such as the TextBlob text processing toolkit, to process the text to be detected and obtain the sentiment score of the text to be detected. Among them, the sentiment score is used to represent the sentiment tendency of the text to be detected.

[0034] The electronic device can also choose VADER to determine the sentiment score. VADER can output four key sentiment scores, including Negative (negative), Neutral (neutral), Positive (positive), and a comprehensive sentiment score (Compound). Negative: represents the intensity or quantity of negative sentiment in the text. Neutral: represents the intensity or quantity of neutral sentiment in the text. Positive: represents the intensity or quantity of positive sentiment in the text.

[0035] The sentiment score output by the VADER sentiment analysis tool is a composite index containing four dimensions. Among them, Compound is a comprehensive index used to measure the overall sentiment tendency of the text to be detected. Its calculation formula is as follows:

[0036]

[0037] Among them, ∑sentiment score represents the sum of all sentiment scores, including Negative, Neutral, and Positive, and the denominator is the normalization process of these sentiment scores. Specifically, how to determine the corresponding sentiment scores based on Negative, Neutral, and Positive is the prior art and will not be elaborated here.

[0038] Through this formula, a Compound between a certain range, such as -1 to 1, can be obtained. This sentiment score can intuitively reflect the sentiment tendency of the text to be detected.

[0039] In the embodiments of this application, introducing sentiment analysis can help the recognition model better capture the meaning of sensitive words in the context and improve the accuracy of sensitive word recognition.

[0040] S102: Input the sentiment score and the feature vector of the text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0041] To accurately and effectively identify sensitive words, an electronic device can determine the feature vector of the text to be detected. In one possible implementation, the electronic device can input the text to be detected into an encoder and obtain the output of the encoder, which is the feature vector of the text to be detected.

[0042] To accurately and effectively identify sensitive words, the electronic device locally stores a pre-trained recognition model. Here, sensitive words refer to specific words or phrases that may involve personal privacy or social taboos, etc. The electronic device can input the sentiment score of the text to be detected and the feature vector of the text to be detected into the pre-trained recognition model, obtain the output of the recognition model, and determine the sensitive words contained in the text to be detected according to the output of the recognition model. Among them, the recognition model can design an intelligent feedback processing mechanism to automatically identify and give priority to processing the text to be detected with a negative sentiment tendency.

[0043] If sensitive words are detected, the electronic device can mark or filter out this content and may generate warnings or logs for subsequent review. Further processing of sensitive words may include deletion, blocking, or forwarding to manual review. In some cases, the detected content may be handed over to manual review to decide whether to take further actions. This can help reduce false positives and false negatives.

[0044] Under the dynamic monitoring mechanism, the electronic device can promptly identify sensitive words in the text to be detected. By keeping up with new trends and actively adopting user feedback, the electronic device can continuously update its sensitive word library to ensure continuous and effective identification of newly emerging sensitive content. For this purpose, the electronic device can build a real-time user feedback system that seamlessly receives and processes user feedback through an API. At the same time, Docker technology can also be used to deploy the API, and the recognition model can be carefully encapsulated into a Docker container. This measure not only significantly enhances the scalability and maintenance convenience of the system but also ensures that the recognition model can be deployed quickly and seamlessly in diverse environments. In this way, a rapid response to user feedback can be achieved, thereby greatly improving the user experience and satisfaction. More importantly, it enables the sensitive word library to closely match the actual needs of users and maintain its efficiency and accuracy.

[0045] An electronic device can utilize Prometheus and Grafana to build an efficient monitoring system that focuses on real-time monitoring of the request volume, response time, and prediction accuracy rate of APIs in the sensitive word recognition task. Here, the API request volume, response time, and prediction accuracy rate specifically refer to the key performance indicators of the electronic device when performing the sensitive word recognition function. The throughput of the API can also be determined, where throughput refers to the number of requests that can be processed within a given time period, reflecting the processing capacity of the system. This monitoring system has the following advantages: Excellent scalability: Based on the microservices architecture design, the system can quickly respond to expansion requirements, achieve flexible deployment, and greatly enhance the overall flexibility and adaptability of the system. Real-time monitoring and early warning function: Through the integrated monitoring system, it can capture and analyze changes in system performance in real time, issue early warnings in a timely manner, ensure the stable operation of the sensitive word recognition service, and effectively reduce potential risks.

[0046] Among them, the calculation formula for API performance indicators is:

[0047]

[0048] Among them, Latency is the average value of the response time, N is the total number of texts to be detected received, ResponseTime is the response time for identifying sensitive words in each text to be detected. Throughput is the throughput, and Total Time is the total duration for processing each text to be detected.

[0049] This application combines multiple modern technologies and methods for sensitive word determination and has the following advantages compared to traditional solutions: Dynamic data processing: By real-time scraping and enhancement technologies, the timeliness of data is maintained. Deep feature understanding: Using deep learning models to improve the text understanding ability. Integration and automation: Through ensemble learning and automatic parameter tuning, the accuracy and efficiency of the model are improved. User feedback mechanism: Continuously update the model to enhance its adaptability and accuracy. Modern deployment and monitoring: Using modern deployment and monitoring technologies to ensure system stability. This solution enhances intelligence and adaptability while maintaining the efficiency of sensitive word determination and is more suitable for rapidly changing network environments.

[0050] It should be noted that there are the following main problems in the related technical solutions: Due to the limitations of the static thesaurus and context understanding, the traditional method often has false positives, identifying non-sensitive content as sensitive words, and false negatives, affecting the accuracy of content review. This solution uses more advanced natural language processing technology to understand the context meaning of words, which can help the system more accurately determine whether a word is sensitive in a specific context, reducing false positives and false negatives. Many traditional detection systems lack support for multiple languages and cultural backgrounds and cannot effectively handle sensitive words in a multilingual environment, resulting in limited applications globally. This solution can be customized for different languages and cultural backgrounds to support the determination of sensitive words in multiple languages to meet the application requirements globally.

[0051] In the embodiment of the present application, the electronic device determines the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected; and inputs the sentiment score and the feature vector of the text to be detected determined based on the encoder into the recognition model, and determines the sensitive words included in the text to be detected based on the output of the recognition model. The recognition model that has been pre-trained can accurately and effectively identify sensitive words, reducing the risk of emerging sensitive words being ignored. Moreover, by determining the sentiment score of the text to be detected, the recognition model can understand the meaning in the text to be detected, and then accurately and effectively identify sensitive words.

[0052] Embodiment 2:

[0053] In order to accurately and effectively identify sensitive words, on the basis of the above embodiment, in the embodiment of the present application, before inputting the sentiment score and the determined feature vector of the text to be detected into the recognition model that has been pre-trained, the method further includes:

[0054] Obtain the target source information of the text to be detected, and determine the target removal characters corresponding to the target source information according to the corresponding relationship between the pre-saved source information and the removal characters; wherein, the removal characters include HyperText Markup Language (HTML) tags, special characters, links, and numbers;

[0055] Remove the target removal characters from the text to be detected.

[0056] Since the text to be detected usually contains characters without actual meaning, in order to improve the efficiency of sensitive word recognition, the electronic device can remove the characters without actual meaning from the text to be detected.

[0057] It should be noted that in the text to be detected from different sources, the characters without actual meaning are different. For example, in social media, emojis and abbreviations are used for relevant expressions. Therefore, in order to accurately and effectively identify sensitive words, the electronic device pre-saves the corresponding relationship between the source information and the characters to be removed. Among them, the characters to be removed include HTML tags, special characters, links, and numbers; that is to say, according to different data sources, different preprocessing logics are designed. For example, for social media text, emojis and abbreviations are retained.

[0058] The electronic device can obtain the target source information of the text to be detected. The target source information can be social media, etc. Specifically, when the data input / collection device sends the text to be detected to the electronic device, it can also send the target source information of the text to be detected to the electronic device. It can also be that the user selects the text to be detected and the target source information of the text to be detected on the preset page of the device or preset device used by the user, and clicks the preset button, such as the "detect" button, then the electronic device can also obtain the text to be detected and the target source information of the text to be detected.

[0059] The electronic device can determine the target characters to be removed corresponding to the target source information according to the target source information of the text to be detected, and remove the target characters to be removed in the text to be detected.

[0060] In order to accurately and effectively remove the target characters to be removed, based on the above embodiments, in the embodiments of the present application, removing the target characters to be removed in the text to be detected includes:

[0061] Screening out the target characters to be removed in the text to be detected based on a regular expression;

[0062] Removing the screened target characters to be removed.

[0063] In order to accurately and effectively remove the target characters to be removed, the electronic device can screen out the target characters to be removed in the text to be detected based on a regular expression, and remove the screened target characters to be removed.

[0064] This process is equivalent to denoising the text to be detected. The text denoising formula is:

[0065] Cleaned Text=Remove(HTML Tags)→Keep(Letters and Chinese Characters)

[0066] Among them, the step of Remove(HTML Tags) refers to removing all HTML tags from the text to be detected, and the step of Keep(Letters and Chinese Characters) ensures that only letters and Chinese characters in the text are retained. Among them, the letter usually refers to the Latin letter, but it can also include letters of other languages.

[0067] By using the method provided in the embodiments of the present application, customized cleaning methods can be adopted according to the different sources of the text to be detected, improving the understanding of text characteristics.

[0068] Embodiment 3:

[0069] In order to accurately and effectively identify sensitive words, based on the above embodiments, in the embodiments of the present application, determining the feature vector of the text to be detected includes:

[0070] Determining the TF-IDF feature vector of the term frequency-inverse document frequency (TF-IDF) feature of the text to be detected;

[0071] And based on the Bidirectional Encoder Representations from Transformers (BERT) model, obtaining the BERT feature vector of the text to be detected;

[0072] According to the preset splicing order, splicing the TF-IDF feature vector and the BERT feature vector to obtain the feature vector of the text to be detected.

[0073] In order to accurately and effectively identify sensitive words, the electronic device can determine the TF-IDF feature vector of the TF-IDF feature of the text to be detected. Specifically, how to determine the TF-IDF feature vector of the text is prior art and will not be elaborated here.

[0074] The electronic device can also calculate the BERT feature vector of the text to be detected. Specifically, the electronic device stores a BERT model. The electronic device can use the transformers library of Hugging Face to load the pre-trained BERT model, input the text to be detected into the BERT model. The tokenization layer of the BERT model divides the text to be detected into multiple words, the vector layer of the BERT model identifies the feature vector corresponding to each word, and the average pooling layer of the BERT model performs pooling processing on each feature vector and outputs the BERT feature vector. It should be noted that when using the BERT model to determine the BERT feature vector, usually the [CLS] token can be added to the beginning of the text to be detected and the corresponding output representation can be obtained. This output representation is the BERT feature vector encoded by the BERT model for the entire text to be detected. The electronic device can also fine-tune the BERT model.

[0075] After obtaining the TF-IDF feature vector and the BERT feature vector, the electronic device can concatenate the TF-IDF feature vector and the BERT feature vector according to a preset concatenation order to form a new feature vector as the input of the subsequent recognition model.

[0076] The method provided by the embodiments of this application has complementarity. The TF-IDF feature vector provides a simple and effective text representation based on word frequency and inverse document frequency, which can reflect the importance of words in the text to be detected, but does not consider the context and polysemy of word meanings. The BERT feature vector can provide richer semantic information through context modeling, taking into account the context dependence and polysemy of words. After combination, the recognition model can simultaneously utilize the statistical information of word frequency and the semantic information of context, improving the representation ability of the text to be detected. In some cases, the TF-IDF feature vector may perform well on small sample data sets, while the BERT feature vector performs excellently on large corpora. Combining the two can better adapt to different data situations. For tasks such as text classification, sentiment analysis, and information retrieval, combining the two features may improve the accuracy and robustness of the model. Combining the TF-IDF feature and the BERT feature can create a powerful text representation, making full use of the advantages of both. This combination can not only improve the richness and expressive ability of text features, but also may promote the performance improvement of downstream tasks. Combining the traditional TF-IDF feature and the BERT feature forms a hybrid feature vector. Using models such as BERT to capture context information greatly improves the understanding of implicit meanings compared with traditional methods. Combining traditional and deep learning features enables the model to have stronger expressive ability and can handle complex language phenomena.

[0077] The method provided by the embodiments of the present application implements a hybrid feature extraction method, which combines TF-IDF features with BERT embeddings to determine sensitive words. It improves the accuracy and timeliness of sensitive word determination, reduces the possibility of missed detection and false detection, and has stronger robustness in the face of diverse and complex language environments.

[0078] In the embodiments of the present application, pandas in Python can be used for data cleaning, Jieba for Chinese word segmentation, and combined with the Term Frequency-Inverse Document Frequency Vectorizer (TfidfVectorizer) for feature extraction.

[0079] Embodiment 4:

[0080] In order to accurately and effectively identify sensitive words, on the basis of the above embodiments, in the embodiments of the present application, the recognition model includes an eXtreme Gradient Boosting (XGBoost) sub-model and at least one other sub-model; determining the sensitive words included in the text to be detected based on the output of the recognition model includes:

[0081] Obtain the words determined to be sensitive words in the text to be detected by each sub-model of the recognition model, and the probabilities of being sensitive words.

[0082] For the words output by each sub-model, according to the probability that the word is a sensitive word output by each sub-model, determine the target probability that the word is a sensitive word. If the target probability is greater than the threshold, then determine that the word is a sensitive word included in the text to be detected.

[0083] In order to accurately and effectively identify sensitive words, the recognition model provided by the embodiments of the present application includes an XGBoost sub-model and at least one other sub-model. The XGBoost sub-model and at least one other sub-model respectively identify the words determined to be sensitive words in the text to be detected, and the probabilities of being sensitive words. The electronic device can, for the words output by each sub-model, according to the probability that the word is a sensitive word output by each sub-model, determine the target probability that the word is a sensitive word. In a possible implementation manner, the electronic device can save corresponding weight values for each sub-model. For each sub-model, determine the probability that the word output by the sub-model is a sensitive word, and determine the product of the probability and the weight value saved for the sub-model. Determine the sum value of the products of each sub-model as the target probability that the word is a sensitive word. After determining the target probability that the word is a sensitive word, the electronic device can determine whether the target probability is greater than the threshold. If it is greater than the threshold, then the word can be determined to be a sensitive word included in the text to be detected.

[0084] The embodiment of this application provides a system for dynamically determining sensitive words and optimizing feedback by combining multiple models. By combining the XGBoost sub-model with other sub-models, a detection mechanism that combines multiple models is formed. The XGBoost sub-model is used for preliminary classification, and other sub-models are used for secondary confirmation. The strategy of model fusion is adopted to improve the accuracy. This method improves the accuracy and reliability of sensitive word determination, and at the same time reduces the risk of overfitting of the model. It has strong adaptability and can be quickly adjusted and optimized according to different scenarios. By combining XGBoost with other models, the VotingClassifier method of ensemble learning can be used for training. The method provided by the embodiment of this application significantly improves the accuracy and stability of decision-making by integrating learning and combining the advantages of different sub-models. Moreover, automated tools can be used for hyperparameter tuning, which improves the efficiency and effect of model training.

[0085] In the embodiment of this application, by combining TF-IDF features and BERT features and using the XGBoost sub-model for training and ensemble learning, the performance of tasks such as text classification can be effectively improved. Using K-fold cross-validation and automatic parameter tuning can help improve the generalization ability and performance of the recognition model. The strategy of model fusion in the embodiment of this application is adopted to improve the accuracy. Traditional methods often rely only on a single model and cannot fully utilize the advantages of different models. The embodiment of this application improves the processing ability of complex text data, especially the diversity of user-generated content. The method of model fusion is adopted to improve the detection effect by using the advantages of multiple sub-models, while the prior art usually only uses a single model. By combining traditional and modern technologies, a more comprehensive recognition method is formed.

[0086] Example 5:

[0087] In order to accurately and effectively train the XGBoost sub-model, based on the above embodiments, in the embodiment of this application, the XGBoost sub-model is trained in the following manner:

[0088] Obtain any sample text in the sample set, and the sample sensitive word and sample probability that are sensitive words in the sample text saved for the sample text;

[0089] Input the sample text and the determined sentiment score of the sample text into the original XGBoost sub-model, and obtain the output sensitive word and output probability that are sensitive words in the sample text output by the original XGBoost sub-model;

[0090] Determine a deviation loss value according to the deviation between the sample probability and the output probability, and determine a regularization loss value according to the sum of squares of the weights of each current tree in the original XGBoost sub-model, the number of trees, the weight corresponding to the sum of squares of the weights, and the weight corresponding to the number of trees; determine a target loss value based on the sum of the regularization loss value and the deviation loss value;

[0091] Train the original XGBoost sub-model based on the target loss value.

[0092] To accurately and effectively train the XGBoost sub-model, the electronic device stores a sample set, which contains sample texts from different sources. And to implement the training of the XGBoost sub-model, for each sample text, the sample sensitive word in the sample text and the sample probability that the sample sensitive word is a sensitive word are also stored.

[0093] The electronic device can obtain any sample text in the sample set, obtain the sentiment score of the sample text, input the sample text and the determined sentiment score of the sample text into the original XGBoost sub-model, and obtain the output sensitive word in the sample text output by the original XGBoost sub-model and the output probability that the output sensitive word is a sensitive word.

[0094] The electronic device can determine a deviation loss value according to the deviation between the sample probability and the output probability, and determine a regularization loss value according to the sum of squares of the weights of each current tree in the XGBoost sub-model, the number of trees, the circle corresponding to the sum of squares of the weights, and the weight corresponding to the number of trees, where the regularization loss value is used to control the complexity of the XGBoost sub-model. After determining the regularization loss value, the electronic device can determine the target loss value according to the sum of the regularization loss value and the deviation loss value. After determining the target loss value, the electronic device can train the original XGBoost sub-model based on the target loss value.

[0095] Specifically, the electronic device can determine the target loss value through the following formula:

[0096] Objective = Loss + R

[0097]

[0098] where, Objective is the target loss value, Loss is the deviation loss value, R is the regularization loss value, n is the number of output sensitive words, i is the output probability of the i-th output sensitive word, is the sample probability of the sample sensitive word corresponding to the i-th output sensitive word, is the deviation between the output probability of the i-th output sensitive word and the corresponding sample probability, T is the number of trees, w j is the weight of the j-th tree, γ is the weight corresponding to the number of trees, is the weight corresponding to the sum of the squares of the weights.

[0099] In the embodiments of the present application, GridSearchCV can be used to perform hyperparameter search to find the parameters of the optimal recognition model. GridSearchCV is used to perform hyperparameter tuning on the XGBoost model. Such a tuning step can effectively improve the performance of the model, help find the optimal combination of hyperparameters, and thus enhance the generalization ability of the model on unknown data. According to the specific task and dataset, different hyperparameters and ranges can be selected for further adjustment and optimization.

[0100] cross_val_score can be used for K-fold cross-validation to make full use of the sample set and evaluate the performance of the recognition model. K-fold cross-validation is a method for evaluating the performance of machine learning models. Its main purpose is to improve the generalization ability of the model through more comprehensive training and testing: the entire sample set is divided into K equal or nearly equal subsets, and usually the value of K is 5 or 10. In K rounds of iteration, in each round, one of the subsets is selected as the test set, and the other K-1 subsets are used as the sample set. In this way, the recognition model will be trained on different sample sets and evaluated on the corresponding test sets. Result summary: The evaluation results (such as accuracy, loss, etc.) of each round will be recorded, and finally, by calculating the average of these results, the overall performance evaluation of the model is obtained.

[0101] Cross-validation formula:

[0102]

[0103] where K is the number of subsets, D i is the validation set of the i-th fold, and Score is the evaluation metric of the recognition model on this fold. The evaluation metrics include accuracy, F1 score, etc.

[0104] K-fold cross-validation enables each data point to be used once in both training and testing. This method avoids the information waste caused by splitting the data into a simple sample set and a test set and can make full use of the data. Through multiple trainings and tests, the parameters of the model can be more robust, thereby reducing the dependence on a specific sample set and reducing the risk of overfitting. Through the evaluation of multiple folds, K-fold cross-validation provides a more comprehensive and reliable estimate of the model performance and avoids the bias caused by the contingency of data division.

[0105] In the embodiments of the present application, MLflow or TensorBoard is used to monitor the loss and performance metrics during the training of the recognition model. The model is evaluated through K-fold cross-validation, which improves the generalization ability of the model. Continuously monitoring the performance of the recognition model helps to timely detect problems and adjust strategies to ensure the long-term effectiveness of the recognition model.

[0106] Embodiment 6:

[0107] In order to accurately and effectively train the recognition model, based on the above embodiments, in the embodiments of the present application, the sample set is obtained in the following manner:

[0108] Obtain any pre-generated text; if there is any vocabulary in the text that is included in the synonym list, within a preset numerical range, randomly determine a target value. If the target value is less than the preset value, then use any synonym of the vocabulary in the synonym list to replace the vocabulary; add the newly generated text, the sensitive words saved for the text, and the probability of being a sensitive word to the sample set.

[0109] In order to improve the generalization ability of the recognition model, the electronic device can perform data augmentation on the sample texts in the sample set. Specifically, synonym replacement can be performed on the sample texts in the sample set to increase semantic diversity.

[0110] Specifically, the electronic device can obtain any pre-generated text, perform word segmentation on the text to determine each vocabulary included in the text. In order to accurately and effectively perform synonym replacement, the electronic device locally stores a synonym list. If any of the vocabulary is included in the synonym list, then the synonym of the vocabulary can be used to replace the vocabulary. The electronic device can randomly determine a target value within a preset numerical range, and the preset numerical range can be [0, 1]. After determining the target value, it can be judged whether the target value is less than the preset value. If the target value is less than the preset value, then use any synonym of the vocabulary in the synonym list to replace the vocabulary, and add the newly generated text, the sensitive words saved for the text, and the probability of being a sensitive word to the sample set.

[0111] The electronic device can also randomly insert or delete some non-keywords in the generated text to improve the generalization ability of the recognition model.

[0112] Traditional methods often fail to recognize variants of sensitive words, such as spelling mistakes, different forms of expressions, and synonyms, which limits their recognition scope and flexibility. This application introduces multiple algorithms, such as synonym matching, root word analysis, and variant recognition, to improve the detection ability of sensitive words. It can handle spelling mistakes, different expressions, and synonyms, thus enhancing the accuracy. Compared with traditional solutions that rely on static data sets, this application improves the timeliness and diversity of data through dynamic data scraping. Through data augmentation techniques, the amount of data in the sample set for model training has increased significantly, which helps to improve the robustness and accuracy of the recognition model.

[0113] Embodiment 7:

[0114] To improve the accuracy of sensitive word recognition, based on the above embodiments, in the embodiments of this application, the method further includes:

[0115] If a correction instruction for a sensitive word in the text to be detected is received, obtain the target sensitive word carried in the correction instruction, and fine-tune the recognition model based on the target sensitive word.

[0116] In actual scenarios, there is usually no effective user feedback mechanism, and it is impossible to continuously optimize the recognition model according to the actual usage and feedback of users. To improve the accuracy of the recognition model, the embodiments of this application design a real-time detection method to quickly respond to user-generated content, and at the same time establish a user feedback mechanism to allow users to report the accuracy of the detection results, thereby optimizing the recognition model.

[0117] Specifically, if a correction instruction for a sensitive word in the text to be detected is received, the electronic device can obtain the target sensitive word identified in the correction instruction and fine-tune the recognition model based on the target sensitive word. To accurately and effectively fine-tune each sub-model in the recognition model, the electronic device can determine the preset probability value as the probability value that the target sensitive word is a sensitive word.

[0118] Figure 2 This is a detailed process schematic diagram for determining sensitive words provided by the embodiments of this application. The process includes the following steps:

[0119] S201: Receive the text to be detected.

[0120] S202: Obtain the target source information of the text to be detected.

[0121] S203: Determine the target removal characters corresponding to the target source information according to the pre-saved correspondence between the source information and the removal characters.

[0122] S204: Remove the target removal characters from the text to be detected.

[0123] S205: Determine the sentiment score of the text to be detected.

[0124] S206: Determine a feature vector obtained by concatenating the TF-IDF feature vector of the text to be detected and the BERT feature vector.

[0125] S207: Input the sentiment score and the feature vector into the pre-trained recognition model.

[0126] S208: Obtain the words that are sensitive words output by each sub-model of the recognition model, and the probability of being sensitive words.

[0127] S209: For each word output by each sub-model, determine the target probability of the word being a sensitive word according to the probability of the word being a sensitive word output by each sub-model. If the target probability is greater than a threshold, determine the word as a sensitive word contained in the text to be detected.

[0128] Embodiment 8:

[0129] The processing of user feedback in related technologies is often not timely enough, and it is difficult to quickly adjust the sensitive word library. The version application implementation example can achieve real-time updates to ensure that user feedback is responded to immediately. The dynamic and adaptive capabilities of the sensitive word library are significantly improved. By closely combining user feedback with sensitive word determination, the embodiment of the present application can update the sensitive word library in real time to cope with the ever-changing network environment. By responding to negative feedback in real time and giving priority to it, the user experience and satisfaction are significantly improved. The sensitive word library can be continuously updated according to the actual needs of users to ensure that it always maintains efficiency and accuracy.

[0130] In order to maintain the advancement and accuracy of the model, the embodiment of the present application implements a continuous learning strategy. The core of this strategy is to regularly integrate newly acquired data and user feedback to form an expanded sample set, and then use this new data set to retrain or fine-tune the model. This approach not only improves the dynamic adaptability of the model, enabling it to quickly adapt to and identify new sensitive content, but also greatly enhances the user's participation and the practicality of the system. Thanks to the introduction of the user feedback mechanism, the embodiment of the present application can continuously absorb new knowledge and flexibly respond to the increasingly complex and changeable sensitive content in the network environment, thereby maintaining its leading edge and accuracy. By providing users with convenient feedback channels and ensuring that their feedback can directly affect the improvement of the system, the user's sense of participation and belonging is greatly enhanced, and it also provides a continuous source of power for continuous optimization.

[0131] In the embodiments of this application, the data storage device uses the efficient relational database MySQL, which is specifically used for the text to be detected, the sensitive word library, and user feedback data. This data storage device not only provides powerful data retrieval and management capabilities, but also supports flexible deployment in local or cloud environments to ensure the security and convenient access of data. The data processing server, as a high-performance computing node, can be responsible for performing data preprocessing tasks on the text to be detected, including data cleaning, word segmentation, and feature extraction. In addition, it also undertakes the training and verification tasks of the recognition model to ensure the continuous optimization of the model performance. The server form is flexible and can be a cloud server or a local computer cluster, both of which are configured with powerful CPU and memory resources. API Service: Based on Web server technology, the API service becomes the core of sensitive word determination. It provides RESTful API interfaces to receive user feedback and immediately return the sensitive word determination results, and at the same time serves as a bridge to connect the front-end application and the back-end service. Whether deployed on a cloud or local server, the API service can ensure stable and efficient communication. The visualization dashboard is built with a front-end application (such as Prometheus). The visualization dashboard displays the sensitive word determination results, user feedback analysis, and monitoring data in real time. Its interactive interface provides an intuitive and convenient management tool for administrators, who can easily access it through a Web browser.

[0132] Data Crawling and Data Storage: The data crawler efficiently stores the newly crawled content into the MySQL database through the API or direct database connection to ensure the timeliness and integrity of the data. Data Processing and Storage: The data processing server regularly extracts new data from the database for preprocessing and redeposits the processing results into the database to form a closed loop of data processing. Data Processing and Machine Learning: The machine learning module obtains feature data from the data processing server, performs model training and prediction, and immediately feeds back the results to the data storage device to achieve seamless docking of the model and the data. API Service and Data Storage: The API service not only receives user feedback and stores it in the database, but also retrieves the sensitive word determination results from the database to provide instant feedback to users. Visualization Dashboard and API Service: Through the API, the visualization dashboard obtains monitoring data in real time, dynamically updates the front-end interface, and provides administrators with a real-time and comprehensive system monitoring and management view.

[0133] The embodiments of this application adopt a three - layer architecture: the data layer, the service layer, and the application layer. The data layer is responsible for data collection and storage. The service layer performs data processing and analysis. The application layer provides user interfaces and visualization. Clear functions are assigned to each layer, and RESTful APIs are used for communication between modules to ensure a loose - coupling design of the system. This can ensure the scalability and maintainability of the system, enabling different modules to be developed, updated, and replaced independently. Through a clear division of module responsibilities, the overall performance and response speed are improved. Traditional sensitive word determination systems often implement a single module, lacking a systematic design and being difficult to handle complex data processing requirements. The three - layer architecture of this system provides higher flexibility and adaptability. It enhances the maintainability of the system, facilitating the addition of new functions in the later stage. Each module can be optimized for performance independently to adapt to different data processing requirements.

[0134] In the embodiments of this application, a self - adaptive update mechanism for the sensitive word library is also designed. Through scheduled tasks and user feedback analysis, new sensitive words are automatically identified and added. An identification model is used to analyze user - generated content to identify potential sensitive words and update them. This ensures that the sensitive word library can always reflect the current usage of network language, improving the timeliness of content review. It has strong adaptability and can quickly respond to the trends and language changes of new monitoring targets. A real - time monitoring dashboard is built based on Dash to display key information such as sensitive word determination results and user feedback analysis. A dynamic update mechanism is designed to enable the dashboard to reflect the system status and data changes in real time. This achieves more efficient decision - making support, enabling managers to quickly identify potential problems. It improves transparency and operability, making content review more scientific and data - driven.

[0135] In the embodiments of the present application, Flask or FastAPI is used to build an API service, Scrapy is used for data scraping, and TensorFlow or PyTorch is used to build a deep learning model. Multiple open-source technologies are integrated to form a complete sensitive word determination system. It provides higher flexibility and adaptability, facilitating subsequent function expansion and technology iteration. It reduces the development and maintenance costs and improves the overall performance of the system. A friendly user feedback interface is designed to quickly submit feedback information through the API and display real-time analysis results. A dynamically updated visualization tool is adopted to enhance the user interaction experience. By optimizing the interaction design, the user's convenience of use is improved. The interaction between the user and the system is enhanced, and the user's sense of participation is improved. It helps the system obtain more effective feedback and promotes the continuous optimization of the sensitive word library. A self-adaptive update mechanism for the sensitive word library is designed to automatically identify and add new sensitive words through scheduled tasks and user feedback analysis. A machine learning model is used to analyze user-generated content to identify potential sensitive words and update them. The update of the traditional sensitive word library is often manual and periodic, making it difficult to cope with the rapidly changing online language environment. The existing determination of sensitive words relies on a static sensitive word library and a simple text matching algorithm. Although the implementation process is relatively straightforward, there are many limitations in aspects such as context understanding, real-time performance, and flexibility. The embodiments of the present application improve the real-time performance and accuracy of the sensitive word library to ensure effective response to newly emerging sensitive content. The dynamic update of the sensitive word library is realized, which is in sharp contrast to the traditional static update mechanism. The machine learning analysis method is adopted to provide more accurate sensitive word recognition.

[0136] Embodiment 9:

[0137] Based on the same inventive concept, the embodiments of the present application provide a sensitive word determination device. Please refer to Figure 3 , the device includes:

[0138] A determination module 301, configured to determine the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected;

[0139] A processing module 302, configured to input the sentiment score and the feature vector of the text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0140] In a possible implementation manner, the processing module 302 is further configured to obtain the target source information of the text to be detected, determine the target removal characters corresponding to the target source information according to the corresponding relationship between the pre-saved source information and the removal characters; wherein, the removal characters include HTML tags, special characters, links, numbers; and remove the target removal characters in the text to be detected.

[0141] In a possible implementation manner, the processing module 302 is specifically configured to filter out the target removal characters in the text to be detected based on a regular expression; and remove the filtered target removal characters.

[0142] In a possible implementation manner, the processing module 302 is specifically configured to determine a TF-IDF feature vector of the TF-IDF features of the text to be detected; and based on a BERT model, obtain a BERT feature vector of the text to be detected; and splice the TF-IDF feature vector and the BERT feature vector according to a preset splicing order to obtain a feature vector of the text to be detected.

[0143] In a possible implementation manner, the processing module 302 is specifically configured to obtain the words in the text to be detected that are sensitive words output by each sub-model of the recognition model, and the probability of being a sensitive word; for each word output by each sub-model, determine a target probability of the word being a sensitive word according to the probability of the word being a sensitive word output by each sub-model. If the target probability is greater than a threshold, determine that the word is a sensitive word included in the text to be detected.

[0144] In a possible implementation manner, the processing module 302 is further configured to train the XGBoost sub-model in the following manner: obtain any sample text in the sample set, and the sample sensitive words and the sample probability of being a sensitive word saved for the sample text; input the sample text and the determined sentiment score of the sample text into the original XGBoost sub-model to obtain the output sensitive words and the output probability of being a sensitive word in the sample text output by the original XGBoost sub-model; determine a deviation loss value according to the deviation between the sample probability and the output probability, and determine a regularization loss value according to the sum of the squares of the weights of each current tree, the number of trees, the weight corresponding to the sum of the squares of the weights, and the weight corresponding to the number of trees in the original XGBoost sub-model; determine a target loss value based on the sum value of the regularization loss value and the deviation loss value; and train the original XGBoost sub-model based on the target loss value.

[0145] In a possible implementation manner, the processing module 302 is further configured to obtain the sample set in the following manner: obtain any pre-generated text; if there is any word in the text included in the synonym list, randomly determine a target value within a preset numerical range. If the target value is less than the preset value, use any synonym of the word in the synonym list to replace the word; add the newly generated text and the sensitive words and the probability of being a sensitive word saved for the text to the sample set.

[0146] In a possible implementation, the processing module 302 is further configured to, if a correction instruction for a sensitive word in the text to be detected is received, obtain the target sensitive word carried in the correction instruction, and fine-tune the recognition model based on the target sensitive word.

[0147] Embodiment 10:

[0148] Based on the same inventive concept, an embodiment of the present application provides an electronic device that can implement the function of determining sensitive words as described above. Please refer to Figure 4 , the device includes a processor 401, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404.

[0149] A computer program is stored in the memory 403. When the program is executed by the processor 401, the processor 401 is caused to execute the following steps:

[0150] Determine the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected;

[0151] Input the sentiment score and the feature vector of the determined text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0152] In a possible implementation, before inputting the sentiment score and the feature vector of the determined text to be detected into a pre-trained recognition model, the method further includes:

[0153] Obtain the target source information of the text to be detected, and determine the target removal characters corresponding to the target source information according to the corresponding relationship between the pre-saved source information and the removal characters; wherein, the removal characters include HTML tags, special characters, links, and numbers;

[0154] Remove the target removal characters from the text to be detected.

[0155] In a possible implementation, removing the target removal characters from the text to be detected includes:

[0156] Filter out the target removal characters in the text to be detected based on a regular expression;

[0157] Remove the filtered target removal characters.

[0158] In a possible implementation, determining the feature vector of the text to be detected includes:

[0159] Determine the TF-IDF feature vector of the TF-IDF features of the text to be detected;

[0160] And based on the BERT model, obtain the BERT feature vector of the text to be detected;

[0161] According to the preset splicing order, splice the TF-IDF feature vector and the BERT feature vector to obtain the feature vector of the text to be detected.

[0162] In a possible implementation manner, the recognition model includes an XGBoost sub-model and at least one other sub-model; determining the sensitive words included in the text to be detected based on the output of the recognition model includes:

[0163] Obtain the words in the text to be detected that are sensitive words output by each sub-model of the recognition model, and the probabilities of being sensitive words;

[0164] For the words output by each sub-model, according to the probability that the word is a sensitive word output by each sub-model, determine the target probability that the word is a sensitive word. If the target probability is greater than the threshold, determine the word as the sensitive word included in the text to be detected.

[0165] In a possible implementation manner, the XGBoost sub-model is trained in the following way:

[0166] Obtain any sample text in the sample set, and the sample sensitive words and sample probabilities of being sensitive words in the sample text saved for the sample text;

[0167] Input the sample text and the determined sentiment score of the sample text into the original XGBoost sub-model, and obtain the output sensitive words and output probabilities of being sensitive words in the sample text output by the original XGBoost sub-model;

[0168] Determine the deviation loss value according to the deviation between the sample probability and the output probability, and determine the regularization loss value according to the sum of the squares of the weights of each current tree in the original XGBoost sub-model, the number of trees, the weights corresponding to the sum of the squares of the weights, and the weights corresponding to the number of trees; based on the sum value of the regularization loss value and the deviation loss value, determine the target loss value;

[0169] Train the original XGBoost sub-model based on the target loss value.

[0170] In a possible implementation manner, the sample set is obtained in the following way:

[0171] Obtain any pre-generated text; if there is any vocabulary in the text that is included in the synonym list, randomly determine a target value within a preset numerical range. If the target value is less than the preset value, use any synonym of the vocabulary in the synonym list to replace the vocabulary; add the newly generated text, the sensitive words saved for the text, and the probabilities of the sensitive words to the sample set.

[0172] In a possible implementation manner, the method further includes:

[0173] If a correction instruction for a sensitive word in the text to be detected is received, obtain the target sensitive word carried in the correction instruction, and fine-tune the recognition model based on the target sensitive word.

[0174] The communication bus mentioned in the above server may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0175] The communication interface 402 is used for communication between the above electronic device and other devices.

[0176] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0177] The above processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0178] Embodiment 11:

[0179] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which a computer program executable by an electronic device is stored. When the program runs on the electronic device, the electronic device is caused to execute the following steps when executed:

[0180] Determine the sentiment score of the text to be detected; wherein, the sentiment score is used to represent the sentiment tendency of the text to be detected;

[0181] Input the sentiment score and the feature vector of the text to be detected into a pre-trained recognition model, and determine the sensitive words included in the text to be detected based on the output of the recognition model.

[0182] In a possible implementation manner, before inputting the sentiment score and the feature vector of the text to be detected into a pre-trained recognition model, the method further includes:

[0183] Obtain the target source information of the text to be detected, and determine the target removal characters corresponding to the target source information according to the pre-saved correspondence between the source information and the removal characters; wherein, the removal characters include HTML tags, special characters, links, and numbers;

[0184] Remove the target removal characters in the text to be detected.

[0185] In a possible implementation manner, removing the target removal characters in the text to be detected includes:

[0186] Filter out the target removal characters in the text to be detected based on a regular expression;

[0187] Remove the filtered target removal characters.

[0188] In a possible implementation manner, determining the feature vector of the text to be detected includes:

[0189] Determine the TF-IDF feature vector of the TF-IDF features of the text to be detected;

[0190] And based on the BERT model, obtain the BERT feature vector of the text to be detected;

[0191] Concatenate the TF-IDF feature vector and the BERT feature vector according to a preset concatenation order to obtain the feature vector of the text to be detected.

[0192] In a possible implementation manner, the recognition model includes an XGBoost sub-model and at least one other sub-model; determining the sensitive words included in the text to be detected based on the output of the recognition model includes:

[0193] Obtain the words determined by each sub-model of the recognition model as sensitive words in the text to be detected, and the probabilities of being sensitive words;

[0194] For each word output by each sub - model, determine the target probability that the word is a sensitive word according to the probability that the word is a sensitive word output by each sub - model. If the target probability is greater than the threshold, determine that the word is a sensitive word included in the text to be detected.

[0195] In a possible implementation manner, the XGBoost sub - model is trained in the following way:

[0196] Obtain any sample text in the sample set, and the sample sensitive words and sample probabilities that are sensitive words in the sample text saved for the sample text;

[0197] Input the sample text and the determined sentiment score of the sample text into the original XGBoost sub - model, and obtain the output sensitive words and output probabilities that are sensitive words in the sample text output by the original XGBoost sub - model;

[0198] Determine the deviation loss value according to the deviation between the sample probability and the output probability, and determine the regularization loss value according to the sum of the squares of the weights of each current tree in the original XGBoost sub - model, the number of trees, the weights corresponding to the sum of the squares of the weights, and the weights corresponding to the number of trees; Based on the sum value of the regularization loss value and the deviation loss value, determine the target loss value;

[0199] Train the original XGBoost sub - model based on the target loss value.

[0200] In a possible implementation manner, the sample set is obtained in the following way:

[0201] Obtain any pre - generated text; If there is any word in the text that is included in the synonym list, randomly determine a target value within a preset numerical range. If the target value is less than the preset value, use any synonym of the word in the synonym list to replace the word; Add the newly generated text and the sensitive words and probabilities that are sensitive words saved for the text to the sample set.

[0202] In a possible implementation manner, the method further includes:

[0203] If a correction instruction for sensitive words in the text to be detected is received, obtain the target sensitive word carried in the correction instruction, and fine - tune the recognition model based on the target sensitive word.

[0204] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0205] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0206] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0207] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0208] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for determining sensitive words, characterized in that: The method comprises: Determine the sentiment score of the text to be detected; wherein the sentiment score is used to represent the sentiment tendency of the text to be detected; The sentiment score and the determined feature vector of the text to be detected are input into a pre-trained recognition model, and the sensitive words contained in the text to be detected are determined based on the output of the recognition model.

2. The method according to claim 1, characterized in that Before inputting the sentiment score and the determined feature vector of the text to be detected into the pre-trained recognition model, the method further includes: Obtaining target source information of the text to be detected, and determining target removal characters corresponding to the target source information according to the pre-saved correspondence between the source information and the removal characters; wherein the removal characters include hypertext markup language HTML tags, special characters, links, and numbers; The target removal characters are removed from the text to be detected.

3. The method according to claim 2, characterized in that The step of removing the target character from the text to be detected includes: Filter out the target removal characters in the text to be detected based on a regular expression; The filtered target removal characters are removed.

4. The method according to claim 1, characterized in that Determining the feature vector of the text to be detected includes: Determine a TF-IDF feature vector of a term frequency-inverse document frequency TF-IDF feature of the text to be detected; And based on the bidirectional transformer BERT model, obtaining the BERT feature vector of the text to be detected; The TF-IDF feature vector and the BERT feature vector are concatenated in a preset concatenation order to obtain a feature vector of the text to be detected.

5. The method according to claim 1, characterized in that: The recognition model includes an extreme gradient boosting XGBoost sub-model and at least one other sub-model; and determining the sensitive words contained in the text to be detected based on the output of the recognition model includes: Obtain the words that are sensitive words in the text to be detected and the probability of being sensitive words output by each sub-model of the recognition model; For each word output by each sub-model, the target probability of the word being a sensitive word is determined according to the probability of the word being a sensitive word output by each sub-model. If the target probability is greater than a threshold, the word is determined to be a sensitive word contained in the text to be detected.

6. The method according to claim 5, characterized in that The XGBoost sub-model is trained in the following way: Obtain any sample text in the sample set, and the sample sensitive words in the sample text that are sensitive words and the sample probability of being sensitive words that are stored for the sample text; The sample text and the determined sentiment score of the sample text are input into the original XGBoost sub-model, and the output sensitive words and the output probability of the sensitive words in the sample text output by the original XGBoost sub-model are obtained; Determine a deviation loss value according to the sample probability and the deviation of the output probability, and determine a regularization loss value according to the sum of the squares of the weights of each tree currently in the original XGBoost sub-model, the number of trees, the weight corresponding to the sum of the squares of the weights, and the weight corresponding to the number of trees; determine a target loss value based on the sum of the regularization loss value and the deviation loss value; The original XGBoost sub-model is trained based on the target loss value.

7. The method according to claim 6, characterized in that The sample set is obtained in the following way: Obtain any pre-generated text; if there is any word included in the synonym list in the text, randomly determine a target value within a preset value range, and if the target value is less than the preset value, use any synonym of the word in the synonym list to replace the word; add the newly generated text and the sensitive words and the probability of being sensitive words saved for the text to the sample set.

8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: If a correction instruction for a sensitive word in the text to be detected is received, a target sensitive word carried in the correction instruction is obtained, and the recognition model is fine-tuned based on the target sensitive word.

9. A sensitive word determination device, characterized in that: The device comprises: A determination module, used to determine the sentiment score of the text to be detected; wherein the sentiment score is used to represent the sentiment tendency of the text to be detected; The processing module is used to input the sentiment score and the determined feature vector of the text to be detected into a pre-trained recognition model, and determine the sensitive words contained in the text to be detected based on the output of the recognition model.

10. An electronic device, characterized in that: include: A memory for storing program instructions; The processor is used to call the program instructions stored in the memory, and execute the steps included in the method for determining sensitive words according to any one of claims 1 to 8 according to the obtained program instructions.