Model training method and sensitive information identification method
By clustering and complexity index analysis of sensitive information data sets, combining pre-trained language models and LSTM networks, the limitations of traditional sensitive word detection methods are solved, and high-precision recognition of sensitive words of different topics is achieved.
Patent Information
- Application Number
- CN202510686774.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional sensitive word detection methods rely on static vocabulary, difficult to cover new vocabulary and variants, and lack refined training for different topics, resulting in low recognition accuracy and recall.
By obtaining the data set carrying sensitive information labels, clustering them to obtain multiple data subsets, selecting sensitive words with high frequency, calculating vocabulary diversity and context dependence indicators, adjusting model training strategies, and using pre-trained language models and long-term memory networks for training.
The recognition accuracy and recall of sensitive words of different topics are improved, and the model's ability to understand multimorphic vocabulary and complex contexts is enhanced.
Smart Images

Figure CN120561590A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of text detection, and more specifically, to a model training method and a method for identifying sensitive information. Background Art
[0002] In the field of content detection technology, especially in the Chinese environment, traditional sensitive word detection methods have obvious limitations, mainly reflected in the following aspects:
[0003] 1. The static nature and limitations of sensitive word libraries: Traditional detection methods rely heavily on pre-defined sensitive word libraries, which are often constructed within a limited timeframe and scope. These libraries struggle to fully capture the ever-increasing number of new words, slang, and variations, especially those targeting sensitive topics like specific ones. With the explosive growth of internet text information, new sensitive words and expressions are constantly emerging. The pace of updating related word libraries has far outstripped this changing trend, resulting in a significant number of newly identified sensitive words being missed, leading to a significant under-detection rate.
[0004] 2. Sensitive words related to different topics have unique characteristics and expressions. For example, sensitive words related to gender discrimination and those related to regional discrimination show significant differences in word usage and emotional intensity. However, related detection models are generally general-purpose and lack specific training for different topics. Therefore, when detecting sensitive words related to specific topics, their recognition accuracy and recall rates are low, and they cannot effectively meet the needs of different scenarios.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] The present application provides a model training method and a method for identifying sensitive information to at least solve the technical problem of poor recognition accuracy of sensitive words for different topics due to the fact that the relevant technology has not trained a model for identifying sensitive words for different topics.
[0007] According to one aspect of the present application, a model training method is provided, including: obtaining a data set carrying sensitive information labels, and clustering data of the same topic in the data set to obtain multiple first data subsets; in each first data subset, selecting the first sensitive words with the largest occurrence frequency n, and determining the lexical diversity index and context dependency index of the first sensitive words, wherein n is a positive integer, the lexical diversity index is used to characterize the number of variants in the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on context; according to the lexical diversity index and context dependency index of the first sensitive words in each first data subset, determining the complexity index corresponding to each first data subset; according to the complexity index corresponding to each first data subset, determining the target ratio corresponding to each first data subset, and in each first.
[0008] Optionally, determining the lexical diversity index of the first sensitive word includes: obtaining a preconfigured variant text set corresponding to the first sensitive word, wherein the variant text set includes at least the first sensitive word and the variant word corresponding to the first sensitive word; determining the lexical diversity index based on the ratio of the number of types of variant words in the variant text set to a target number of times, wherein the target number of times the first sensitive word and the variant word appear in the variant text set, and the value range of the lexical diversity index is [0,1].
[0009] Optionally, determining the context dependency index of the first sensitive word includes: determining the original text in which the first sensitive word is located, and obtaining a context window centered on the first sensitive word; extracting the first text from the original text according to the context window; processing the first text using a pre-trained language model to obtain a first semantic vector corresponding to the first sensitive word and multiple second semantic vectors corresponding to multiple word segmentation results other than the first sensitive word in the first text; calculating the semantic similarity between the first semantic vector and the second semantic vector to obtain multiple semantic similarities, and determining the average value of the multiple semantic similarities as the context dependency index, wherein the value range of the context dependency index is [0,1].
[0010] Optionally, the complexity index corresponding to each first data subset is determined based on the lexical diversity index and context dependency index of the first sensitive words, including: when n is 1, multiplying the lexical diversity index and the context dependency index of the first sensitive words in the first data subset to obtain the complexity index corresponding to the first data subset; when n is not 1, determining the weight coefficient of each first sensitive word based on the frequency of occurrence of each first sensitive word in the first data subset; multiplying the lexical diversity index and the context dependency index of each first sensitive word in the first data subset to obtain the initial complexity index of each first sensitive word; and performing weighted summation of the initial complexity index of each first sensitive word based on the weight coefficient to obtain the complexity index corresponding to the first data subset.
[0011] Optionally, based on the complexity index corresponding to each first data subset, the target ratio corresponding to each first data subset is determined, including: determining the target interval in which the complexity index corresponding to each first data subset is located; determining the target ratio corresponding to the target interval based on a preset mapping relationship, to obtain the target ratio corresponding to each first data subset.
[0012] Optionally, the sensitive information identification model includes: a pre-trained language model, a pre-trained word segmenter, a long short-term memory network, and a fully connected layer.
[0013] Optionally, the sensitive information identification model is trained using the data set and different second data subsets respectively, including: mapping the data set and different second data subsets into a training data set and a verification data set respectively; dividing the training data set and the verification data set into batches, and randomly shuffling the division results to obtain multiple batches of training data sets and verification data sets, and inputting the multiple batches of training data sets and verification data sets into the sensitive information identification model in sequence for model training. During the training process, a binary cross entropy loss function is selected as the optimization target, and a gradient descent mechanism with a batch size of 32 is adopted. The loss gradient is calculated in parallel on each batch of training data sets and the momentum adaptive parameter update of the optimizer is triggered, and the verification data set is used to determine the evaluation index.
[0014] According to another aspect of the present application, a method for identifying sensitive information is also provided, including: obtaining text to be identified on different topics; identifying the text to be identified on different topics through different sensitive information identification models to obtain multiple identification results, wherein the different sensitive information identification models are obtained by training through the above model training method.
[0015] According to another aspect of the present application, a model training device is also provided, including: an acquisition module for acquiring a data set carrying sensitive information labels, and clustering data of the same topic in the data set to obtain multiple first data subsets; a first determination module for selecting the first sensitive words with the largest occurrence frequency in each first data subset, and determining the lexical diversity index and context dependency index of the first sensitive words, wherein n is a positive integer, the lexical diversity index is used to characterize the number of variants in the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context; a second determination module for determining the complexity index corresponding to each first data subset based on the lexical diversity index and context dependency index of the first sensitive words in each first data subset; a third determination module for determining the target ratio corresponding to each first data subset based on the complexity index corresponding to each first data subset, and selecting training data of the target ratio in each first data subset to obtain multiple second data subsets; a training module for training a sensitive information recognition model using the data set and different second data subsets respectively, wherein the sensitive information recognition model is used to identify sensitive information in the text to be identified.
[0016] According to another aspect of the present application, a non-volatile storage medium is also provided, which includes a stored program, wherein when the program runs, the device where the storage medium is located is controlled to execute the above model training method.
[0017] According to another aspect of the present application, an electronic device is provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the above model training method is executed when the program is run.
[0018] According to another aspect of the present application, a computer program is also provided, wherein the above model training method is implemented when the computer program is executed by a processor.
[0019] According to another aspect of the present application, a computer program product is provided, which includes a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above model training method is implemented.
[0020] In the present application, a data set carrying sensitive information labels is obtained, and data of the same topic are clustered in the data set to obtain multiple first data subsets; in each first data subset, the first sensitive words with the largest occurrence frequency of n are selected, and the lexical diversity index and context dependency index of the first sensitive words are determined, where n is a positive integer, the lexical diversity index is used to characterize the number of variants in the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context; based on the lexical diversity index and context dependency index of the first sensitive words in each first data subset, the complexity index corresponding to each first data subset is determined; based on the complexity index corresponding to each first data subset, the target ratio corresponding to each first data subset is determined, and in each first method, the purpose of determining different data amounts according to the complexity indicators of sensitive words of different topics, and training models for identifying sensitive words of different topics according to different data amounts is achieved, thereby achieving the technical effect of improving the recognition accuracy of sensitive words of different topics, and thus solving the technical problem of poor recognition accuracy of sensitive words of different topics caused by the fact that the relevant technology has not trained models for identifying sensitive words of different topics. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0022] Figure 1 is a flow chart of a model training method according to an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of a model training method according to an embodiment of the present application;
[0024] Figure 3 is a structural diagram of a model training device according to an embodiment of the present application;
[0025] Figure 4 This is a hardware structure block diagram of a computer terminal according to a model training method of an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] According to an embodiment of the present application, a method embodiment of a model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] Figure 1 is a flow chart of a model training method according to an embodiment of the present application, such as Figure 1 As shown, the method includes the following steps:
[0030] Step S101: obtain a data set carrying sensitive information labels, and cluster data with the same subject in the data set to obtain multiple first data subsets.
[0031] In step S101, text data containing sensitive information is first collected from multiple data sources, such as social media platforms, online forums, and news comments, to ensure data diversity and representativeness. The collected data is then preprocessed, including but not limited to text cleaning, removal of irrelevant characters and stop words, and text formatting. The preprocessed text data is then labeled with a topic, such as gender and region. A clustering algorithm is then used to partition the dataset into multiple first data subsets based on the topic tags.
[0032] In step S102, in each first data subset, the first sensitive word with the largest occurrence frequency of n is selected, and the lexical diversity index and context dependency index of the first sensitive word are determined, where n is a positive integer, the lexical diversity index is used to characterize the number of variants in the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context.
[0033] In step S102, a word frequency count is performed on the text in each first data subset, and the top n words with the highest frequency of occurrence are identified as first sensitive words, where n can be set according to actual needs. The lexical diversity index and context dependency index of the first sensitive words are determined.
[0034] The lexical diversity index primarily measures the range and richness of variation in a word or group of related words when expressing the same concept. This index identifies how many different variations or synonyms a particular sensitive word has that convey the same or similar negative connotations. A higher lexical diversity index indicates a greater diversity of sensitive word expressions, placing higher demands on the model's recognition capabilities.
[0035] The context-dependency metric measures the stability of the meaning of a word or a group of related words across different contexts. In online text, a word may be sensitive in some contexts and neutral or positive in others. A higher metric indicates that the meaning and sentiment of the sensitive word are more susceptible to the influence of surrounding context, requiring the model to possess strong semantic understanding and contextual analysis capabilities.
[0036] Step S103: determining a complexity index corresponding to each first data subset according to the lexical diversity index and the context dependency index of the first sensitive word in each first data subset.
[0037] In step S103 above, the lexical diversity index focuses on the number of different variants and synonyms of sensitive words, which is directly related to the model's learning difficulty and recognition accuracy. The context dependency index examines how sensitive words change with changes in context, which is a key factor in determining whether the model can accurately understand the semantics of the text. Based on these two indicators, the complexity index is determined. Through the complexity index, model training can specifically strengthen the understanding of polymorphic vocabulary and complex contexts, significantly improving the accuracy and comprehensiveness of sensitive word detection.
[0038] It's understandable that data subsets with high complexity indicators indicate that sensitive words in these subsets are more diverse in expression and more context-dependent. By identifying these subsets, we can more specifically adjust model training strategies, such as using more training data. This allows the model to more effectively learn the various variations of sensitive words and their behavior patterns in specific contexts, thereby improving recognition accuracy and recall.
[0039] Step S104 , determining a target ratio corresponding to each first data subset according to the complexity index corresponding to each first data subset, and selecting training data of the target ratio from each first data subset to obtain multiple second data subsets.
[0040] It is understood that based on the complexity index, the text data ratio used for training in each data subset is set. Subsets with high complexity require a higher target ratio to enhance the model's ability to recognize complex sensitive words. Within each first data subset, training data is randomly sampled according to the set target ratio to form multiple second data subsets. When selecting data, it is necessary to consider maintaining the natural distribution of sensitive words and context to avoid bias.
[0041] In step S105 , the sensitive information recognition model is trained using the data set and different second data subsets, respectively, wherein the sensitive information recognition model is used to recognize sensitive information in the text to be recognized.
[0042] During the model training phase, the dataset and different second data subsets are used to train the model. Specifically, the dataset and different second data subsets are first loaded using the pandas library to ensure the consistency of the data format. Then, a tf.data.Dataset is constructed. The model is trained using the Adam optimizer and the binary cross entropy loss function binary_crossentropy within the TensorFlow framework using gradient descent. Each subset of topics corresponds to a model instance, and each model is particularly good at identifying sensitive words within a certain topic. After model training is complete, each model is optimized by evaluating its performance on the test data, including accuracy, precision, recall, and F1 score.
[0043] According to the above steps, a data set carrying sensitive information labels is obtained, and data of the same topic are clustered in the data set to obtain multiple first data subsets; in each first data subset, the first sensitive words with the largest occurrence frequency of n are selected, and the lexical diversity index and context dependency index of the first sensitive words are determined, wherein n is a positive integer, the lexical diversity index is used to characterize the number of variants of the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context; according to the lexical diversity index and context dependency index of the first sensitive words in each first data subset, the complexity index corresponding to each first data subset is determined; according to the complexity index corresponding to each first data subset, the target ratio corresponding to each first data subset is determined, and in each first method, the purpose of determining different data amounts according to the complexity indicators of sensitive words of different topics and training models for identifying sensitive words of different topics according to different data amounts is achieved, thereby achieving the technical effect of improving the recognition accuracy of sensitive words of different topics.
[0044] The following Figure 1 The steps shown are exemplified and explained.
[0045] According to some optional embodiments of the present application, determining the lexical diversity index of the first sensitive word can be achieved by the following method: obtaining a pre-configured variant text set corresponding to the first sensitive word, wherein the variant text set includes at least the first sensitive word and the variant word corresponding to the first sensitive word; determining the lexical diversity index based on the ratio of the number of types of variant words in the variant text set to the target number of times, wherein the target number of times the first sensitive word and the variant word appear in the variant text set, and the value range of the lexical diversity index is [0,1].
[0046] Among them, variant words in the variant text set include but are not limited to misspelled versions, homophones, abbreviations, special expressions in dialects or Internet languages, etc.
[0047] Furthermore, determining the context dependency index of the first sensitive word can be achieved by the following method: determining the original text where the first sensitive word is located, and obtaining a context window centered on the first sensitive word; extracting the first text from the original text according to the context window; processing the first text using a pre-trained language model to obtain a first semantic vector corresponding to the first sensitive word and multiple second semantic vectors corresponding to multiple word segmentation results other than the first sensitive word in the first text; calculating the semantic similarity between the first semantic vector and the second semantic vector to obtain multiple semantic similarities, and determining the average value of the multiple semantic similarities as the context dependency index, wherein the value range of the context dependency index is [0,1].
[0048] Specifically, the original text containing the first sensitive word is located from the dataset through a database query or text search algorithm. A context window is established around the sensitive word. For example, the window size (i.e., the number of words before and after the sensitive word) includes 5 to 10 words before and after the sensitive word. The specific value of the window size depends on the total length of the original text. A text segment containing the sensitive word and its context window is cut from the original text, referred to as the first text. This portion of text will be the primary target of subsequent semantic analysis. Deep semantic encoding is performed on the first text using, for example, BERT or RoBERTa. A pre-trained language model can automatically adjust the representation vector of each word based on the context, which is crucial for understanding the meaning of sensitive words in a specific context. The pre-trained language model generates a semantic vector for the first sensitive word. This vector embodies the comprehensive semantic features of the sensitive word in the current context. Semantic vectors are also generated for all word segments in the first text except the first sensitive word, forming a second set of semantic vectors. These vectors are used to analyze the semantic relationship between the sensitive word and other words. For each vector in the second set of semantic vectors, the cosine similarity or other semantic similarity metric is calculated between it and the first semantic vector to obtain a series of semantic similarity values. All calculated semantic similarity values are averaged to create an indicator reflecting the contextual dependency of sensitive words. A high average similarity indicates that the meaning of a sensitive word is highly dependent on surrounding words, making it more difficult to identify. Conversely, a low average similarity indicates that the meaning of a sensitive word is relatively fixed across different contexts, making identification relatively easy.
[0049] Furthermore, the complexity index corresponding to each first data subset is determined based on the lexical diversity index and context dependency index of the first sensitive word, which can be achieved by the following method: when n is 1, the lexical diversity index and the context dependency index of the first sensitive word in the first data subset are multiplied to obtain the complexity index corresponding to the first data subset; when n is not 1, the weight coefficient of each first sensitive word is determined based on the frequency of occurrence of each first sensitive word in the first data subset; the lexical diversity index and the context dependency index of each first sensitive word in the first data subset are multiplied to obtain the initial complexity index of each first sensitive word; and the initial complexity index of each first sensitive word is weightedly summed according to the weight coefficient to obtain the complexity index corresponding to the first data subset.
[0050] It is worth noting that because both the lexical diversity and contextual dependency indices are normalized to the same range of values, the resulting complexity metric remains dimensionally consistent with the two original indices. This means that regardless of the data subset or computational context from which the metric originates, the values of the complexity metric can be directly compared and interpreted without the need for additional unit conversion or scaling, simplifying data analysis and model tuning. Furthermore, multiplication effectively emphasizes the importance and complexity of data subsets that exhibit both high lexical diversity and high contextual dependency. When both values are close to 1, the result of the multiplication will also be close to 1, indicating that the data subset has higher complexity and requires more computational resources and more refined model tuning to process. Conversely, if one metric is low, the resulting complexity metric will be lower even if the other is high.
[0051] According to other optional embodiments of the present application, based on the complexity index corresponding to each first data subset, the target ratio corresponding to each first data subset is determined, including: determining the target interval in which the complexity index corresponding to each first data subset is located; determining the target ratio corresponding to the target interval based on a preset mapping relationship, to obtain the target ratio corresponding to each first data subset.
[0052] For example, the complexity index is 0.678, and its target range is 0.65 to 0.7, and the target ratio corresponding to this target range is 80%.
[0053] According to other optional embodiments of the present application, the sensitive information identification model includes: a pre-trained language model, a pre-trained word segmenter, a long short-term memory network and a fully connected layer.
[0054] The sensitive information identification model integrates key components such as pre-trained language models (such as BERT), pre-trained word segmenters, long short-term memory networks (LSTMs), and fully connected layers to achieve accurate identification of sensitive information in text.
[0055] The self-attention mechanism formula in the Transformer architecture of the BERT model is:
[0056]
[0057] Q, K, V represent query, key and value matrices respectively, d k is the dimension of the key.
[0058] BERT, through its multi-layer Transformer architecture, is able to understand the meaning of words in different contexts. Semantic Representation: BERT generates rich semantic vector representations for each word. These vectors contain not only the semantics of the word itself, but also its position in the sentence and contextual information, providing high-quality features for subsequent classification tasks.
[0059] The update formula of LSTM unit:
[0060] i t =σ(W xi x t +W hi h t-1 +b i )
[0061] f t =σ(W xf x t +W hf h t-1 +b f )
[0062] g t =tanh(W xg x t +W hg h t-1 +b g )
[0063] o t =σ(W xo x t +W ho h t-1 +b o )
[0064] c t =f t ⊙c t-1 +i t ⊙g t
[0065] h t =o t ⊙tanh(c t )
[0066] i t 、f t 、o t Represent the activation values of the input gate, forget gate and output gate respectively, gt is the candidate cell state, ct is the cell state, ht is the hidden state, σ is the sigmoid function, and ⊙ represents element-by-element multiplication.
[0067] LSTM can analyze the emotional and semantic coherence of sentences, determine the overall tendency of sentences, and effectively process long-term dependencies in sequential data. LSTM can capture the sequential relationships and contextual dependencies between words in a text, helping to understand the overall structure and semantics of a sentence.
[0068] Furthermore, the sensitive information identification model is trained using the dataset and different second data subsets respectively, which can be achieved by the following method: mapping the dataset and different second data subsets into training datasets and verification datasets respectively; dividing the training dataset and verification dataset into batches, and randomly shuffling the division results to obtain multiple batches of training datasets and verification datasets, and inputting the multiple batches of training datasets and verification datasets into the sensitive information identification model in sequence for model training. During the training process, the binary cross entropy loss function is selected as the optimization target, and a gradient descent mechanism with a batch size of 32 is adopted. The loss gradient is calculated in parallel on each batch of training datasets and the momentum adaptive parameter update of the optimizer is triggered, and the verification dataset is used to determine the evaluation index.
[0069] Specifically, first, the dataset and different second data subsets are divided into two parts: one part is used as a training dataset for model learning, and the other part is used as a validation dataset to evaluate the performance of the model on unseen data. Next, in order to improve training efficiency and the generalization ability of the model, the training dataset is further divided into multiple small batches. Each batch contains 32 records. The choice of this batch size takes into account the efficient use of computing resources and the stability of model training. By randomly shuffling the dataset, it is ensured that the data batches received for each model training are randomly selected, which reduces the sequential effect of the data and avoids the model from learning unnecessary pattern deviations. Similarly, the validation dataset is also divided into multiple batches, but it does not need to be shuffled because the validation set is mainly used for evaluation rather than training.
[0070] During model training, a binary cross-entropy loss function is chosen as the optimization objective. This is because sensitive information identification is a binary classification problem; the model needs to determine whether a piece of text contains sensitive information. As the model performs forward propagation on each training batch, it generates a predicted probability regarding whether the input text is sensitive. The model then compares the predicted probability with the actual label (i.e., whether the text actually contains sensitive information) and calculates the loss. This loss signal is then transmitted to the model parameters via backpropagation, triggering the adaptive parameter update mechanism of the optimizer (such as Adam). After each batch of training, the optimizer calculates the loss gradient in parallel to efficiently update the model weights and adjust the model to reduce prediction error.
[0071] The validation dataset is used to monitor performance changes during model training and ensure that the model does not overfit the training data. After the model completes a round of learning on all training data, it is evaluated using data from the validation dataset, calculating key evaluation metrics such as accuracy, precision, recall, and F1 score. These metrics determine the model's performance when processing unseen data, particularly in terms of its accuracy in identifying sensitive information. By continuously adjusting training strategies, such as learning rate decay and regularization parameter optimization, the model's recognition and generalization capabilities can be gradually improved, enabling it to accurately judge sensitive information of various types and variations.
[0072] Figure 2 is a schematic diagram of a model training method according to an embodiment of the present application, such as Figure 2 As shown, the method can be implemented through the following steps.
[0073] Data collection and preprocessing. A large amount of Chinese text data is collected from multiple sources, such as social media, forums, and news comments. The data is collected on a wide range of scales to ensure broad coverage so that the model can learn instances of sensitive words in various contexts. The preprocessing steps include text cleaning, word segmentation, and converting the text into a format suitable for model input. Text cleaning removes irrelevant characters, labels, or noise, while word segmentation breaks the text into words or phrases for subsequent model analysis. Finally, a pre-trained word segmenter is used to convert the text into a vector representation that the model can understand.
[0074] Model construction. Use the Transformers library to load the pre-trained BERT model and its tokenizer. The BERT model, with its powerful language understanding capabilities, provides a contextually relevant vector representation for each input word, providing rich semantic information for subsequent LSTM model processing. Create a model that integrates BERT and LSTM. First, the BERT model acts as an encoder to encode the input text. Then, the LSTM model analyzes the encoded sequence to capture the temporal dependencies between words. Finally, a fully connected layer performs classification decisions to determine whether the text contains sensitive words and the topic to which they belong.
[0075] Model training. Use the pandas library to load the dataset and perform necessary preprocessing, such as unifying column names and removing irrelevant columns. Convert the processed dataset to the TensorFlow Dataset format to facilitate data batch processing and iteration during model training. Also, shuffle the dataset to ensure that the model receives random batches of data during each training session. Compile the model using the Adam optimizer and set the learning rate to 1e-5. During training, the model runs on TPU acceleration, training on the entire dataset and on subsets of data divided by topic to enhance the model's generalization capabilities across different topics.
[0076] Model evaluation and optimization. After training, evaluate model performance by calculating metrics such as accuracy, precision, recall, and F1 score. Additionally, plot evaluation results to visually demonstrate the model's performance on different topics. Based on the evaluation results, adjust model parameters, optimizer settings, or training strategies to improve the model's detection performance.
[0077] Sensitive word detection: Based on the model's prediction results, determine whether the text contains sensitive words and the topic category to which they belong, thereby achieving sensitive word detection.
[0078] The above steps integrate the self-attention mechanism to conduct in-depth semantic analysis of the text, which can better understand the semantic associations and contextual information in the text, thereby more accurately identifying violent speech and emotion classification.
[0079] The present invention also provides a method for identifying sensitive information, including the following steps:
[0080] Obtain texts to be identified of different topics; identify the texts to be identified of different topics through different sensitive information identification models to obtain multiple identification results, wherein different sensitive information identification models are Figure 1 The model training method is used to train the obtained
[0081] The text to be recognized is obtained by the following method: obtaining basic information, including the intent keywords and slots of the current conversation turn, historical conversation information stored in the conversation state, user profile, and user information such as current location and device. Based on this basic information, the specific scenario of the current conversation is identified, such as operations, recommendations, and traffic diversion. A prompt word template for the specific scenario is obtained, and the basic information and prompt word template are input into the intent recognition model to obtain the text to be recognized as output by the intent recognition model.
[0082] Figure 3 is a structural diagram of a model training device according to an embodiment of the present application, such as Figure 3 As shown, the device includes:
[0083] The acquisition module 31 is used to acquire a data set carrying sensitive information labels, and cluster data with the same subject in the data set to obtain multiple first data subsets.
[0084] The first determination module 32 is used to select the first sensitive words with the largest occurrence frequency n in each first data subset, and determine the lexical diversity index and context dependency index of the first sensitive words, where n is a positive integer, the lexical diversity index is used to characterize the number of variants of the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context.
[0085] The second determining module 33 is configured to determine a complexity index corresponding to each first data subset according to the lexical diversity index and the context dependency index of the first sensitive words in each first data subset.
[0086] The third determination module 34 is used to determine the target ratio corresponding to each first data subset according to the complexity index corresponding to each first data subset, and select training data of the target ratio from each first data subset to obtain multiple second data subsets.
[0087] The training module 35 is used to train the sensitive information recognition model using the data set and different second data subsets respectively, wherein the sensitive information recognition model is used to recognize sensitive information in the text to be recognized.
[0088] Optionally, determining the lexical diversity index of the first sensitive word includes the following steps: obtaining a pre-configured variant text set corresponding to the first sensitive word, wherein the variant text set includes at least the first sensitive word and the variant word corresponding to the first sensitive word; determining the lexical diversity index based on the ratio of the number of types of variant words in the variant text set to a target number of times, wherein the target number of times the first sensitive word and the variant word appear in the variant text set, and the value range of the lexical diversity index is [0,1].
[0089] Optionally, determining the context dependency index of the first sensitive word includes the following steps: determining the original text in which the first sensitive word is located, and obtaining a context window centered on the first sensitive word; extracting the first text from the original text according to the context window; processing the first text using a pre-trained language model to obtain a first semantic vector corresponding to the first sensitive word and multiple second semantic vectors corresponding to multiple word segmentation results other than the first sensitive word in the first text; calculating the semantic similarity between the first semantic vector and the second semantic vector to obtain multiple semantic similarities, and determining the average value of the multiple semantic similarities as the context dependency index, wherein the value range of the context dependency index is [0,1].
[0090] Optionally, the complexity index corresponding to each first data subset is determined based on the lexical diversity index and context dependency index of the first sensitive words, including the following steps: when n is 1, multiplying the lexical diversity index and the context dependency index of the first sensitive words in the first data subset to obtain the complexity index corresponding to the first data subset; when n is not 1, determining the weight coefficient of each first sensitive word based on the frequency of occurrence of each first sensitive word in the first data subset; multiplying the lexical diversity index and the context dependency index of each first sensitive word in the first data subset to obtain the initial complexity index of each first sensitive word; and performing weighted summation on the initial complexity index of each first sensitive word based on the weight coefficient to obtain the complexity index corresponding to the first data subset.
[0091] Optionally, based on the complexity index corresponding to each first data subset, determining the target ratio corresponding to each first data subset includes the following steps: determining the target interval in which the complexity index corresponding to each first data subset is located; determining the target ratio corresponding to the target interval based on a preset mapping relationship, and obtaining the target ratio corresponding to each first data subset.
[0092] Optionally, the sensitive information identification model includes: a pre-trained language model, a pre-trained word segmenter, a long short-term memory network, and a fully connected layer.
[0093] Optionally, the sensitive information identification model is trained using the data set and different second data subsets respectively, including the following steps: mapping the data set and different second data subsets into a training data set and a verification data set respectively; dividing the training data set and the verification data set into batches, and randomly shuffling the division results to obtain multiple batches of training data sets and verification data sets, and inputting the multiple batches of training data sets and verification data sets into the sensitive information identification model in sequence for model training. During the training process, the binary cross entropy loss function is selected as the optimization target, and a gradient descent mechanism with a batch size of 32 is adopted. The loss gradient is calculated in parallel on each batch of training data sets and the momentum adaptive parameter update of the optimizer is triggered, and the verification data set is used to determine the evaluation index.
[0094] It should be noted that the above Figure 3 The modules in the embodiment can be program modules (for example, a set of program instructions that implement a specific function) or hardware modules. For the latter, they can be expressed in the following forms, but are not limited to these: the expression form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0095] It should be noted that Figure 3 The preferred implementation of the embodiment shown can be found in Figure 1The relevant description of the illustrated embodiment will not be repeated here.
[0096] Figure 4 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a model training method. Figure 4 As shown, the computer terminal 40 may include one or more (402a, 402b, ..., 402n are shown in the figure) processors 402 (the processor 402 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission module 406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components than shown, or with Figure 4 Different configurations shown.
[0097] It should be noted that the one or more processors 402 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 40. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0098] The memory 404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiment of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, that is, realizing the above-mentioned model training method. The memory 404 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 404 may further include a memory remotely arranged relative to the processor 402, and these remote memories can be connected to the computer terminal 40 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0099] The transmission module 406 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 40. In one embodiment, the transmission module 406 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 406 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0100] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 40 .
[0101] It should be noted that, in some optional embodiments, the above Figure 4 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 4 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.
[0102] It should be noted that Figure 4 The computer terminal shown is used to execute Figure 1 The model training method shown, therefore the relevant explanations in the execution method of the above commands are also applicable to the electronic device and will not be repeated here.
[0103] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored program, wherein when the program runs, the device where the storage medium is located is controlled to execute the above model training method.
[0104] A program for a non-volatile storage medium to perform the following functions: obtaining a data set carrying a sensitive information label, and clustering data of the same topic in the data set to obtain multiple first data subsets; in each first data subset, selecting the first sensitive words with the largest occurrence frequency n, and determining the lexical diversity index and context dependency index of the first sensitive words, wherein n is a positive integer, the lexical diversity index is used to characterize the number of variants of the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context; determining the complexity index corresponding to each first data subset based on the lexical diversity index and context dependency index of the first sensitive words in each first data subset; determining the target ratio corresponding to each first data subset based on the complexity index corresponding to each first data subset, and in each first.
[0105] An embodiment of the present application also provides an electronic device, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the above model training method is executed when the program is running.
[0106] The processor is used to run a program that performs the following functions: obtaining a data set carrying sensitive information labels, and clustering data of the same topic in the data set to obtain multiple first data subsets; in each first data subset, selecting the first sensitive words with the largest occurrence frequency of n, and determining the lexical diversity index and context dependency index of the first sensitive words, wherein n is a positive integer, the lexical diversity index is used to characterize the number of variants in the expression of sensitive information, and the context dependency index is used to characterize the degree of dependence of sensitive information on the context; according to the lexical diversity index and context dependency index of the first sensitive words in each first data subset, determining the complexity index corresponding to each first data subset; according to the complexity index corresponding to each first data subset, determining the target ratio corresponding to each first data subset, and in each first.
[0107] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0108] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0109] In the above-mentioned embodiments of the present application, the collected information is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary protection measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0110] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0111] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0112] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0114] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A model training method, characterized in that: include: Obtaining a data set carrying a sensitive information label, and clustering data with the same subject in the data set to obtain a plurality of first data subsets; In each of the first data subsets, select the first sensitive words with the highest occurrence frequency n times, and determine the lexical diversity index and context dependency index of the first sensitive words, where n is a positive integer, the lexical diversity index is used to represent the number of variations in the expression of sensitive information, and the context dependency index is used to represent the degree of dependence of sensitive information on context; Determining a complexity index corresponding to each of the first data subsets according to the lexical diversity index and the context dependency index of the first sensitive words in each of the first data subsets; Determining a target ratio corresponding to each of the first data subsets according to the complexity index corresponding to each of the first data subsets, and selecting training data of the target ratio from each of the first data subsets to obtain a plurality of second data subsets; The data set and different second data subsets are respectively used to train a sensitive information recognition model, wherein the sensitive information recognition model is used to recognize sensitive information in the text to be recognized.
2. The method according to claim 1, characterized in that Determining the lexical diversity index of the first sensitive word includes: Obtaining a preconfigured variant text set corresponding to the first sensitive word, wherein the variant text set includes at least the first sensitive word and variant words corresponding to the first sensitive word; The lexical diversity index is determined based on the ratio of the number of types of variant words in the variant text set to a target number of times, wherein the target number of times is the total number of times the first sensitive word and the variant word appear in the variant text set, and the value range of the lexical diversity index is [0, 1].
3. The method according to claim 2, characterized in that Determining the context dependency index of the first sensitive word includes: Determining the original text containing the first sensitive word and obtaining a context window centered on the first sensitive word; extracting the first text from the original text based on the context window; Processing the first text using a pre-trained language model to obtain a first semantic vector corresponding to the first sensitive word and a plurality of second semantic vectors corresponding to a plurality of word segmentation results in the first text other than the first sensitive word; The semantic similarity between the first semantic vector and the second semantic vector is calculated to obtain multiple semantic similarities, and an average value of the multiple semantic similarities is determined as the context dependency index, wherein the value range of the context dependency index is [0,1].
4. The method according to claim 3, characterized in that Determining a complexity index corresponding to each of the first data subsets according to the lexical diversity index and the context dependency index of the first sensitive words includes: When n is 1, multiplying the lexical diversity index and the context dependency index of the first sensitive word in the first data subset to obtain the complexity index corresponding to the first data subset; When n is not 1, determining a weight coefficient for each of the first sensitive words according to the frequency of occurrence of each of the first sensitive words in the first data subset; performing multiplication calculation on the lexical diversity index and the context dependency index of each of the first sensitive words in the first data subset to obtain an initial complexity index of each of the first sensitive words; According to the weight coefficient, a weighted sum is performed on the initial complexity index of each of the first sensitive words to obtain the complexity index corresponding to the first data subset.
5. The method according to claim 1, characterized in that Determining a target ratio corresponding to each of the first data subsets according to a complexity index corresponding to each of the first data subsets includes: Determining a target interval of the complexity index corresponding to each of the first data subsets; According to a preset mapping relationship, a target ratio corresponding to the target interval is determined, and a target ratio corresponding to each of the first data subsets is obtained.
6. The method according to claim 1, characterized in that The sensitive information identification model includes: a pre-trained language model, a pre-trained word segmenter, a long short-term memory network and a fully connected layer.
7. The method according to claim 6, characterized in that The sensitive information recognition model is trained using the data set and different second data subsets respectively, including: Mapping the data set and different second data subsets into a training data set and a validation data set respectively; The training dataset and the validation dataset are divided into batches, and the division results are randomly shuffled to obtain multiple batches of training datasets and validation datasets. Multiple batches of training data sets and validation data sets are sequentially input into the sensitive information recognition model for model training. During the training process, the binary cross entropy loss function is selected as the optimization target, and a gradient descent mechanism with a batch size of 32 is adopted. The loss gradient is calculated in parallel on each batch of training data sets and the momentum adaptive parameter update of the optimizer is triggered. The validation data set is used to determine the evaluation indicators.
8. A method for identifying sensitive information, characterized in that: include: Obtain text to be recognized on different topics; The text to be identified on different topics is identified by different sensitive information identification models to obtain multiple identification results, wherein the different sensitive information identification models are obtained by training through the model training method described in any one of claims 1 to 7.
9. A model training device, characterized in that: include: an acquisition module, configured to acquire a data set carrying sensitive information labels, and cluster data of the same subject in the data set to obtain a plurality of first data subsets; A first determination module is configured to select, from each of the first data subsets, the first sensitive words with the highest occurrence frequencies (n), and determine a lexical diversity index and a context dependency index for the first sensitive words, where n is a positive integer, the lexical diversity index is used to represent the number of variations in the expression of sensitive information, and the context dependency index is used to represent the degree of dependence of sensitive information on context; a second determining module, configured to determine a complexity index corresponding to each of the first data subsets based on the lexical diversity index and the context dependency index of the first sensitive words in each of the first data subsets; a third determining module, configured to determine a target ratio corresponding to each of the first data subsets according to the complexity index corresponding to each of the first data subsets, and select training data of the target ratio from each of the first data subsets to obtain a plurality of second data subsets; A training module is used to train a sensitive information recognition model using the data set and different second data subsets respectively, wherein the sensitive information recognition model is used to identify sensitive information in the text to be identified.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein, when the program is running, the device where the non-volatile storage medium is located is controlled to execute the model training method described in any one of claims 1 to 7 and the sensitive information identification method described in claim 8.
11. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein when the program is run, the model training method described in any one of claims 1 to 7 and the sensitive information identification method described in claim 8 are executed.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the model training method described in any one of claims 1 to 7 and the sensitive information identification method described in claim 8.
Citation Information
Patent Citations
Risk classification model training method and system
CN111539612A
Natural language ambiguity resolution method and system based on deep learning and knowledge base
CN117648933A
Identification and prevention of sensitive information exposure in telephonic conversations
US20250030795A1
Native expansion of a sparse training dataset into a dense training dataset for supervised training of a synonymous variant sequence generator
WO2024086143A1