Text compliance detection method based on unbalanced data set

By preprocessing and encoding the large model data set, the classification model is constructed for text compliance detection, which solves the problem of insufficient detection speed and accuracy under the unbalanced data set, and achieves more efficient and safe information dissemination and improvement of large model adaptability.

CN120144767APending Publication Date: 2025-06-13INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510141591.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing large models are difficult to effectively deal with unbalanced data sets in text compliance detection, resulting in insufficient detection speed and accuracy, and cannot meet the compliance needs of specific fields.

Method used

By collecting and preprocessing the large model data set, determining the vocabulary, encoding and dimensionality reduction of the data, calculating the sub-vocabulary vector similarity values ​​and comprehensive similarity values, and building a classification model to achieve text compliance detection.

Benefits of technology

It improves the speed and accuracy of text processing, strengthens the compliance supervision of large models in the data training and content generation process, ensures the security and compliance of information dissemination, and improves the adaptability and generalization capabilities of large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144767A_ABST
    Figure CN120144767A_ABST
Patent Text Reader

Abstract

The invention provides a text compliance detection method based on an unbalanced data set, and belongs to the technical field of text processing, and the method comprises the steps: collecting and preprocessing a large model data set, determining first data, and determining a vocabulary; performing coding and dimension reduction processing on the first data based on the vocabulary, and determining a second coding vector of each first statement in the first data; determining a sub-vocabulary vector similarity value and a comprehensive similarity value of every two first statements in the first data, and classifying the first data to determine a first category, and first category data and second category data of each category in the first category; and constructing a classification model based on the second category data, and processing the large model data based on the classification model. According to the method, the text processing speed and accuracy can be improved, the compliance supervision of a large model service provider in the data training and content generation process is enhanced, the security and compliance of information propagation are guaranteed, and the adaptability and generalization ability of a large model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and particularly to text compliance detection based on an imbalanced data set. Background Art

[0002] Text compliance detection plays a crucial role in ensuring the compliance of large models and is widely used in fields such as fintech, medicine, military, and finance. As large model technology is gradually applied to various fields, although the large model itself will do some alignment work to prevent the appearance of some illegal words or sentences during the conversation, it cannot cover the requirements of certain specific fields.

[0003] Therefore, the present invention provides a text compliance detection method based on an imbalanced data set. Summary of the Invention

[0004] The present invention provides a text compliance detection method based on an imbalanced data set. By collecting and preprocessing the large model data set, the first data is determined, and the vocabulary is determined. According to the vocabulary, the first data is encoded and dimensionally reduced to determine the second encoding vector of each first statement in the first data. The sub-vocabulary vector similarity value and the comprehensive similarity value of every two first statements in the first data are calculated, the first category, the first category data of each category in the first category, and the second category data are determined. A classification model is constructed according to the second category data, and the classification processing of the large model data is realized according to the classification model, which can improve the speed and accuracy of text processing, strengthen the compliance supervision of large model service providers during the data training and content generation processes, serve society more safely and responsibly, ensure the security and compliance of information dissemination, and improve the adaptability and generalization ability of the large model.

[0005] The present invention provides a text compliance detection method based on an imbalanced data set, including:

[0006] 101: Collect and preprocess the large model data set, determine the first data, and determine the vocabulary based on the first data;

[0007] 102: Encode the first data based on the vocabulary, perform dimensionality reduction processing on the encoded first data, and determine the second encoding vector of each first statement in the first data;

[0008] 103: Determine the sub-vocabulary vector similarity value and the comprehensive similarity value of every two first statements in the first data, and classify the first data to determine the first category and the first category data of each category in the first category;

[0009] 104: Augment the data set of the first category data of each category in the first category to determine the second category data of each category;

[0010] 105: Construct a classification model based on the second category of data, and process the large model data based on the classification model.

[0011] According to a text compliance detection method based on an imbalanced data set provided by the present invention, the preprocessing includes format conversion, data cleaning, and stop word removal.

[0012] According to a text compliance detection method based on an imbalanced data set provided by the present invention, collect and preprocess the large model data set, and determine the first data, including:

[0013] Extract all data types in the collected large model data set, determine the data format with the most occurrences as the first format, and convert the data in the large model data set that is inconsistent with the first format;

[0014] Clean the large model data set after format conversion, wherein the data cleaning includes blank character standardization, special character removal, and punctuation removal;

[0015] Extract stop words from the large model data set after data cleaning, and remove all the extracted stop words;

[0016] Determine the large model data set after stop word removal as the first data, wherein the first data contains multiple first statements.

[0017] According to a text compliance detection method based on an imbalanced data set provided by the present invention, determine a vocabulary based on the first data, including:

[0018] Extract all the words that appear in the first data, determine the first vocabulary, and count the number of times each word in the first vocabulary appears in the first data, wherein the first vocabulary includes multiple words;

[0019] Sort all the words in the first vocabulary based on the number of times each word appears in the first data, determine the word sequence, and draw a word distribution graph based on the word sequence and the number of times each word appears in the first data;

[0020] Determine the cumulative coverage rate of each word based on the word distribution graph;

[0021] Compare the cumulative coverage rate of each word with the set coverage threshold one by one based on the word sequence, and determine the second vocabulary based on the comparison result;

[0022] Determine the dependency window based on the length of the first statement in the first data, split the first data based on the dependency window, and determine the word statements of each word in the first vocabulary for the split first data, wherein the word statements include multiple first statements containing the corresponding words;

[0023] Determine the lexical dependency value of each word in the first vocabulary based on the lexical sentences of each word in the first vocabulary;

[0024] Compare the lexical dependency value of each word in the first vocabulary with the set dependency value threshold, extract all words whose lexical dependency value is greater than the set dependency value threshold, and determine the third vocabulary;

[0025] Determine the vocabulary based on the second vocabulary and the third vocabulary.

[0026] According to a text compliance detection method based on an imbalanced dataset provided by the present invention, compare the cumulative coverage rate of each word with the set coverage rate threshold one by one based on the word sequence, and determine the second vocabulary based on the comparison result, including:

[0027] If the cumulative coverage rate of the word is less than the set coverage rate threshold, compare the cumulative coverage rate of the next word in the word sequence corresponding to the word whose cumulative coverage rate is less than the set coverage rate threshold with the set coverage rate threshold;

[0028] If the cumulative coverage rate of the word is greater than or equal to the set coverage rate threshold, the comparison stops, and determine the word corresponding to the cumulative coverage rate greater than or equal to the set coverage rate threshold and all words before the word in the word sequence as the second vocabulary, where the order of words in the second vocabulary is the same as the order of words in the word sequence.

[0029] According to a text compliance detection method based on an imbalanced dataset provided by the present invention, determine the vocabulary based on the second vocabulary and the third vocabulary, including:

[0030] Extract the words that appear in both the second vocabulary and the third vocabulary, determine them as common words, and adjust the order of the common words in the second vocabulary based on the lexical dependency value of each common word;

[0031] For all words in the third vocabulary except the common words, insert them into the second vocabulary with adjusted order based on the corresponding lexical dependency values in turn;

[0032] Determine the vocabulary based on all words and the word order in the inserted second vocabulary.

[0033] According to a text compliance detection method based on an imbalanced dataset provided by the present invention, encode the first data based on the vocabulary to determine the encoded data, including:

[0034] Determine the word vectors based on the vocabulary, determine the index value of each word in the vocabulary in the word vectors, and determine the corresponding sub-word vectors based on the index value of each word;

[0035] Find all the words contained in each first statement in the first data in the lexical vectors, and extract the sub-lexical vectors of all the words;

[0036] Add up the sub-lexical vectors of all the words contained in each first statement to determine the first encoding vectors of all the first statements;

[0037] Determine the encoding matrix based on the first encoding vectors of all the first statements in the first data, and perform centering processing on the encoding matrix;

[0038] Determine the covariance matrix of the centered encoding matrix, perform eigenvalue decomposition on the covariance matrix, and determine the principal components;

[0039] Project the first encoding vector of each first statement in the first data onto the principal components to determine the second encoding vector after dimensionality reduction of each first statement in the first data.

[0040] According to a text compliance detection method based on an imbalanced data set provided by the present invention, determine the sub-lexical vector similarity values and comprehensive similarity values of every two first statements in the first data, and classify the first data, including:

[0041] Based on all the sub-lexical vectors, lexical dependency values, and second encoding vectors of every two first statements in the first data, calculate the sub-lexical vector similarity value and comprehensive similarity value of the two first statements;

[0042]

[0043] Among them, Sw ij represents the sub-lexical vector similarity value of the i-th first statement and the j-th first statement in the first data, Vw ia and Vw jb respectively represent the a-th sub-lexical vector of the i-th first statement and the b-th sub-lexical vector of the j-th first statement in the first data, iN1 and jN1 respectively represent the number of sub-lexical vectors of the i-th first statement and the j-th first statement in the first data, Rw ia represents the lexical dependency value of the word corresponding to the a-th sub-lexical vector of the i-th first statement in the first data, Rw jb represents the lexical dependency value of the word corresponding to the b-th sub-lexical vector of the j-th first statement in the first data, S ij represents the comprehensive similarity value of the i-th first statement and the j-th first statement in the first data, Vs i represents the second encoding vector of the i-th first statement in the first data, Vs j represents the second encoding vector of the j-th first statement in the first data, Lw i and Lw jrespectively represent the number of index values of the i-th first statement and the j-th first statement in the first data, f(Lw i ,Lw j ) represents the regularization value based on the number of index values of the i-th first statement and the number of index values of the j-th first statement in the first data, W1 represents the weight of the statement similarity value, W2 represents the weight of the lexical similarity value, and δ represents the regularization parameter;

[0044] Based on the sub-lexical vector similarity values and comprehensive similarity values of each first statement in the first data and all the other first statements, perform clustering analysis on the first data to determine the first category and the first category data of each category.

[0045] According to a text compliance detection method based on an imbalanced data set provided by the present invention, perform data set augmentation on the first category data of each category in the first category to determine the second category data of each category, including:

[0046] Based on the comprehensive similarity values of all the first statements included in the first category data of each category in the first category, determine the inter-class similarity of all the first category data;

[0047] Sort the inter-class similarities of all the categories in the first category from smallest to largest to determine the inter-class similarity sequence, and determine the minimum value, lower quartile, upper quartile, and maximum value of the inter-class similarity sequence;

[0048] Determine that all the categories corresponding to the inter-class similarities within the range from the minimum value to the lower quartile are low-similarity categories, determine that all the categories corresponding to the inter-class similarities within the range from the lower quartile to the upper quartile are medium-similarity categories, and determine that all the categories corresponding to the inter-class similarities within the range from the upper quartile to the maximum value are high-similarity categories;

[0049] Swap the word orders of all the first statements included in each low-similarity category to determine all the first statements and the statements after swapping the word orders of each first statement as the second category data of the corresponding low-similarity category;

[0050] Perform synonym replacement on all the first statements included in each medium-similarity category to determine all the first statements and the statements after synonym replacement of each first statement as the second category data of the corresponding medium-similarity category;

[0051] Swap the word orders of all the first statements included in each high-similarity category to determine all the first statements and the statements after swapping the word orders of each first statement as the second category data of the corresponding high-similarity category.

[0052] According to a text compliance detection method based on an imbalanced data set provided by the present invention, process the large model data based on a classification model, including:

[0053] The classification model makes a compliance judgment on the classification of the input data of the large model data. At the same time, the classification model makes a compliance judgment on the classification of the output data of the large model data.

[0054] Compared with the prior art, the beneficial effects of the present application are as follows:

[0055] By collecting and preprocessing the large model data set, determining the first data, and determining the vocabulary, encoding and dimensionality reduction processing are performed on the first data according to the vocabulary to determine the second encoding vector of each first statement in the first data. Calculate the sub-vocabulary vector similarity value and the comprehensive similarity value of every two first statements in the first data, determine the first category, the first category data of each category in the first category and the second category data, construct a classification model according to the second category data, and realize the classification processing of the large model data according to the classification model, which can improve the speed and accuracy of text processing, strengthen the compliance supervision of large model service providers during the data training and content generation processes, serve the society more safely and responsibly, ensure the security and compliance of information dissemination, and improve the adaptability and generalization ability of the large model. Description of the Drawings

[0056] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of a text compliance detection method based on an imbalanced data set provided by an embodiment of the present invention. Detailed Embodiments

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0059] Embodiment 1:

[0060] The embodiment of the present invention provides a text compliance detection method based on an imbalanced data set, as Figure 1 shown, including:

[0061] 101: Collect and preprocess the large model dataset, determine the first data, and determine the vocabulary based on the first data;

[0062] 102: Encode the first data based on the vocabulary, perform dimensionality reduction on the encoded first data, and determine the second encoding vectors of each first statement in the first data;

[0063] 103: Determine the sub-vocabulary vector similarity values and comprehensive similarity values of every two first statements in the first data, and classify the first data to determine the first category and the first category data of each category in the first category;

[0064] 104: Augment the dataset of the first category data of each category in the first category to determine the second category data of each category;

[0065] 105: Build a classification model based on the second category data and process the large model data based on the classification model.

[0066] In this embodiment, the large model refers to a large-scale language model, which is a complex model trained based on large-scale data and deep learning methods, and is particularly outstanding in the fields of natural language processing (NLP), computer vision (CV), etc.

[0067] In this embodiment, the large model dataset includes the input data and output data of the large model, and can be news websites, social media, open text datasets (such as Wikipedia, Common Crawl), books, and academic papers, etc.

[0068] In this embodiment, the classification model is trained according to the second data using TextRNN (a text classification model based on recurrent neural network).

[0069] In this embodiment, during the training of the classification model, cross-entropy is used as the loss function, and the Adam optimizer is used for parameter update. Through multiple rounds of iterative training, the classification model learns how to classify the text content into different categories.

[0070] Beneficial effects of the above technical solution: By collecting and preprocessing the large model dataset, determining the first data, and determining the vocabulary, encoding and dimensionality reduction are performed on the first data according to the vocabulary to determine the second encoding vector of each first statement in the first data. Calculate the sub-vocabulary vector similarity value and the comprehensive similarity value of every two first statements in the first data, determine the first category, the first category data of each category in the first category, and the second category data. Construct a classification model according to the second category data, and realize the classification processing of the large model data according to the classification model, which can improve the speed and accuracy of text processing, strengthen the compliance supervision of large model service providers during the data training and content generation process, serve the society more safely and responsibly, ensure the security and compliance of information dissemination, and improve the adaptability and generalization ability of the large model.

[0071] Embodiment 2:

[0072] The embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. The preprocessing includes format conversion, data cleaning, and stop word removal.

[0073] In this embodiment, format conversion means unifying the data formats of all data in the large model dataset into the first format.

[0074] In this embodiment, data cleaning means screening and correcting the large model dataset, aiming to remove or correct errors, incompleteness, inconsistencies, or noise information in the data, thereby improving the quality of the data.

[0075] In this embodiment, stop word removal means removing those common words that contribute less to the text semantics during the text processing of the large model dataset to initially reduce the dimension of the large model dataset.

[0076] Beneficial effects of the above technical solution: By determining the preprocessing of the large model dataset, the quality of the first data can be improved.

[0077] Embodiment 3:

[0078] The embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. Collecting and preprocessing the large model dataset to determine the first data includes:

[0079] Extract all data types in the collected large model dataset, determine the data format with the most occurrences as the first format, and perform format conversion on the data in the large model dataset that is inconsistent with the first format;

[0080] Perform data cleaning on the large model dataset after format conversion, where data cleaning includes whitespace normalization, special character removal, and punctuation removal;

[0081] Extract stop words from the large model dataset after data cleaning, and remove all the extracted stop words;

[0082] Determine that the large model dataset after stop word removal is the first data, where the first data contains multiple first statements.

[0083] In this embodiment, whitespace normalization means uniformly processing redundant spaces and line breaks into a single space to maintain text consistency.

[0084] In this embodiment, special character removal means clearing unnecessary special characters, control characters, HTML tags, or other non-text symbols. For example, for data scraped from web pages or social media, a large amount of noisy information such as advertisements and script codes in the data is removed.

[0085] In this embodiment, punctuation removal means retaining useful punctuation marks such as full stops and commas, and removing irrelevant or noisy punctuation marks such as consecutive symbols.

[0086] In this embodiment, extracting and removing stop words from the large model dataset after data cleaning can reduce the noise in the large model dataset and highlight key information. For example, stop words can include "ah", "ne", "de", "zai", etc.

[0087] In this embodiment, the large model dataset includes multiple statements. Each statement in the large-scale dataset corresponds to a first statement in the first data, and the first statement represents the statement after preprocessing the corresponding statement in the large-scale dataset.

[0088] Beneficial effects of the above technical solution: By collecting and preprocessing the large model dataset and determining the first data, the quality of the first data can be ensured, providing a data basis for determining the vocabulary table and improving the accuracy of text processing.

[0089] Embodiment 4:

[0090] An embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. Determining a vocabulary table based on the first data includes:

[0091] Extract all the words that appear in the first data to determine the first vocabulary, and count the number of times each word in the first vocabulary appears in the first data, where the first vocabulary includes multiple words;

[0092] Sort all the words in the first vocabulary based on the number of times each word appears in the first data to determine a word sequence, and draw a word distribution diagram based on the word sequence and the number of times each word appears in the first data;

[0093] Determine the cumulative coverage rate of each word based on the word distribution diagram;

[0094] Compare the cumulative coverage rate of each word and the set coverage rate threshold one by one based on the word sequence, and determine the second word based on the comparison result.

[0095] Determine the dependency window based on the length of the first sentence in the first data, segment the first data based on the dependency window, and determine the word sentences of each word in the first word for the segmented first data, where the word sentence includes multiple first sentences containing the corresponding word.

[0096] Determine the word dependency value of each word in the first word based on the word sentences of each word in the first word.

[0097] Compare the word dependency value of each word in the first word with the set dependency value threshold, extract all the words whose word dependency value is greater than the set dependency value threshold, and determine the third word.

[0098] Determine the vocabulary based on the second word and the third word.

[0099] In this embodiment, the first word includes all the words included in all the first sentences in the first data.

[0100] In this embodiment, the earlier the position of the word in the word sequence, the greater the number of times the word appears in the first data.

[0101] In this embodiment, the cumulative coverage rate represents the ratio of the sum of the number of occurrences of the corresponding word and all the words to the left of the word in the word distribution diagram to the sum of the number of occurrences of all the words in the word distribution diagram.

[0102] In this embodiment, each word has a corresponding cumulative coverage rate. For example: there are 100 words in the word distribution diagram, the total number of occurrences of all the words is 500, the first word in the word distribution diagram is x1, and the number of times it appears in the first data is 50. Then the cumulative coverage rate of the first word is 50 / 500 = 10%. The second word in the word distribution diagram is x2, and the number of times it appears in the first data is 40. Then the cumulative coverage rate of the second word is 10% + 40 / 500 = 10% + 8% = 18%, or 90 / 500 = 18%, and so on.

[0103] In this embodiment, the abscissa of the word distribution diagram represents the words that appear in the first word, and the ordinate represents the number of times the corresponding word appears in the large model dataset. The word distribution diagram is a zigzag downward broken line.

[0104] In this embodiment, according to the word sequence, start comparing the cumulative coverage rate of the word and the set coverage rate threshold from the first word.

[0105] In this embodiment, the dependency window can be determined based on the sentence lengths in the large model dataset. The average sentence length of the large-scale dataset can be calculated to determine the average sentence length as the length of the dependency window.

[0106] In this embodiment, the lexical dependency value can be determined by calculating the word frequency and inverse document frequency of the corresponding word, constructing a Skip-Gram model, constructing a CBOW model, or constructing a Transformer model.

[0107] In this embodiment, each word in the first vocabulary corresponds to a word sentence, and the word sentence includes all the sentences containing the word in the first data.

[0108] In this embodiment, the fewer the number of the first sentences included in the word sentence of the word, the higher the semantic dissimilarity represented in each of the first sentences in the word sentence, and the greater the lexical dependency value of the word.

[0109] In this embodiment, the third vocabulary includes all the words whose lexical dependency values are greater than the set dependency value threshold.

[0110] Beneficial effects of the above technical solution: By determining the first vocabulary, the second vocabulary, and the third vocabulary from the first data and determining the vocabulary list, a lexical basis can be provided for determining the encoded data, improving the classification accuracy of the first data, and further enhancing the text processing accuracy.

[0111] Embodiment 5:

[0112] The embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. By comparing the cumulative coverage rate of each word with the set coverage rate threshold one by one based on the word sequence, the second vocabulary is determined, including:

[0113] If the cumulative coverage rate of the word is less than the set coverage rate threshold, compare the cumulative coverage rate of the next word in the word sequence corresponding to the word whose cumulative coverage rate is less than the set coverage rate threshold with the set coverage rate threshold;

[0114] If the cumulative coverage rate of the word is greater than or equal to the set coverage rate threshold, the comparison stops, and the word corresponding to the cumulative coverage rate greater than or equal to the set coverage rate threshold and all the words before the word in the word sequence are determined as the second vocabulary, where the order of the words in the second vocabulary is the same as the order of the words in the word sequence.

[0115] In this embodiment, the set coverage rate threshold represents the target cumulative rate set for the first data in advance, usually a percentage, such as 70%, 80%, or 90%, to ensure that the selected second vocabulary includes enough data content in all the words to cover the first data.

[0116] In this embodiment, if the cumulative coverage rate of the words compared this time is less than the set coverage rate threshold, then the cumulative coverage rate of the next word after the word corresponding to the cumulative coverage rate of the words compared this time in the word sequence is compared with the set coverage rate threshold.

[0117] In this embodiment, the comparison stops until the cumulative coverage rate of the words is greater than or equal to the set coverage rate threshold, and the word corresponding to the cumulative coverage rate of the words when the comparison stops, and all the words before this word in the word sequence are determined as the second words.

[0118] Beneficial effects of the above technical solution: By comparing the cumulative coverage rate of each word with the set coverage rate threshold according to the word sequence, and determining the second words based on the comparison results, it can provide a data basis for determining the vocabulary list and improve the representativeness of all the words included in the vocabulary list for the large model dataset.

[0119] Embodiment 6:

[0120] An embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. Determining a vocabulary list based on the second words and the third words includes:

[0121] Extract the words that appear in both the second words and the third words, and determine them as common words. Based on the word dependency values of each word in the common words, adjust the order of the common words in the second words;

[0122] For all the words in the third words except the common words, based on the corresponding word dependency values, insert them into the second words with the adjusted order in sequence;

[0123] Determine the vocabulary list based on all the words and the word order in the second words after insertion.

[0124] In this embodiment, based on the word dependency values of each word in the common words, adjusting the order of the common words in the second words can first standardize the word dependency values of all the words in the common words, and according to the magnitudes of the standardized word dependency values, adjust the order of the word in the second words forward. For example, the greater the word dependency value of a certain word, the greater the forward adjustment span, and vice versa, the smaller the forward adjustment span.

[0125] In this embodiment, for all the words in the third words except the common words, based on the corresponding word dependency values, inserting them into the second words with the adjusted order in sequence can standardize the word dependency values of all the words in the third words except the common words, and according to the standardized word dependency value of each word, insert them into the second words with the adjusted order. For example: the greater the standardized word dependency value, the more forward the position inserted into the second words with the adjusted order, and vice versa, the more backward the position.

[0126] Beneficial effects of the above technical solution: Determining a vocabulary based on the second vocabulary and the third vocabulary can provide a data basis for encoding the first data, and further improve the representativeness of all the vocabulary included in the vocabulary for the large model dataset.

[0127] Example 7:

[0128] An embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset. Encoding the first data based on a vocabulary to determine encoded data, including:

[0129] Determine a vocabulary vector based on the vocabulary, and determine the index value of each vocabulary in the vocabulary vector. Determine the corresponding sub-vocabulary vector based on the index value of each vocabulary;

[0130] Find all the vocabulary included in each first statement in the first data in the vocabulary vector, and extract the sub-vocabulary vectors of all the vocabulary;

[0131] Add up the sub-vocabulary vectors of all the vocabulary included in each first statement to determine the first encoding vector of all the first statements;

[0132] Determine an encoding matrix based on the first encoding vectors of all the first statements in the first data, and perform centering processing on the encoding matrix;

[0133] Determine the covariance matrix of the centered encoding matrix, perform eigenvalue decomposition on the covariance matrix, and determine the principal components;

[0134] Project the first encoding vector of each first statement in the first data onto the principal components to determine the second encoding vector after dimensionality reduction of each first statement in the first data.

[0135] In this embodiment, the number of vocabulary in the vocabulary is the vector length of the vocabulary vector.

[0136] In this embodiment, the index value represents the position value of the corresponding vocabulary in the vocabulary vector.

[0137] In this embodiment, the vector lengths of the encoding vectors of all the first statements in the first data are the same as the length of the vocabulary vector.

[0138] In this embodiment, the vector length of the sub-vocabulary vector is the same as the length of the vocabulary vector. The value of the sub-vocabulary vector at the index value position of the corresponding vocabulary is 1, and the rest of the positions are 0. For example, if the vocabulary vector is a vector with a length of 7 (including 7 vocabulary), the index value of a certain vocabulary in the vocabulary vector is a, and the value of the sub-vocabulary vector of this vocabulary at the a-th position is 1, and the rest of the positions are 0.

[0139] In this embodiment, an encoding vector of each first statement is determined based on all sub-lexical vectors of the first statement. For the encoding vector of the first statement, the value at the position of the index value of each existing lexical item is 1, and the value at the position of the index value of the non-recorded lexical item is 0. For example, a certain first statement contains 4 lexical items, and the index values of the 4 lexical items in the lexical vector are 1, 3, 4, and 6 respectively, and the corresponding sub-lexical vectors are [1,0,0,0,0,0,0], [0,0,1,0,0,0,0], [0,0,0,1,0,0,0], [0,0,0,0,0,1,0], and the encoding vector of this first statement is [1,0,1,1,0,1,0].

[0140] In this embodiment, the encoding matrix is a matrix of the length of the lexical vector × the number of first statements in the first data.

[0141] In this embodiment, the centering process means subtracting the element in all first encoding vectors from the mean value of its corresponding dimension, so that the mean value of each dimension feature is 0, ensuring that vectors of different dimensions are analyzed on the same scale and eliminating the offset.

[0142] In this embodiment, the covariance matrix represents the matrix calculated for the encoding matrix after centering processing, which describes the correlation between different dimensions, that is, which dimensions' changes are linearly correlated.

[0143] In this embodiment, the covariance matrix is decomposed to extract the eigenvalues and eigenvectors therein.

[0144] In this embodiment, the second encoding vector represents the low-dimensional vector obtained by projecting the first encoding vector of each statement onto the principal components.

[0145] Beneficial effects of the above technical solution: By encoding the first data according to the vocabulary table to determine the encoded data, the dimension of the data can be reduced, and at the same time, noise and redundant information can be removed, not only retaining the key features of the data, but also significantly improving the efficiency and speed of subsequent processing.

[0146] Embodiment 8:

[0147] An embodiment of the present invention provides a text compliance detection method based on an imbalanced data set, which determines the sub-lexical vector similarity value and the comprehensive similarity value of every two first statements in the first data, and classifies the first data, including:

[0148] Based on all sub-lexical vectors, lexical dependency values, and second encoding vectors of every two first statements in the first data, calculate the sub-lexical vector similarity value and the comprehensive similarity value of the two first statements;

[0149]

[0150] Among them, Swij Denotes the sub - vocabulary vector similarity value between the \(i\) - th first statement and the \(j\) - th first statement in the first data, \(V_w\) ia 、\(V_w\) jb Denote the \(a\) - th sub - vocabulary vector of the \(i\) - th first statement and the \(b\) - th sub - vocabulary vector of the \(j\) - th first statement in the first data respectively. \(iN_1\) and \(jN_1\) denote the number of sub - vocabulary vectors of the \(i\) - th first statement and the \(j\) - th first statement in the first data respectively, \(R_w\) ia Denotes the lexical dependency value of the word corresponding to the \(a\) - th sub - vocabulary vector of the \(i\) - th first statement in the first data, \(R_w\) jb Denotes the lexical dependency value of the word corresponding to the \(b\) - th sub - vocabulary vector of the \(j\) - th first statement in the first data, \(S\) ij Denotes the comprehensive similarity value between the \(i\) - th first statement and the \(j\) - th first statement in the first data, \(V_s\) i Denotes the second encoding vector of the \(i\) - th first statement in the first data, \(V_s\) j Denotes the second encoding vector of the \(j\) - th first statement in the first data, \(L_w\) i 、\(L_w\) j Denote the number of index values of the \(i\) - th first statement and the \(j\) - th first statement in the first data respectively, \(f(L_w\) i , \(L_w\) j ) denotes the regularization value based on the number of index values of the \(i\) - th first statement and the number of index values of the \(j\) - th first statement in the first data. \(W_1\) denotes the statement similarity value weight, \(W_2\) denotes the lexical similarity value weight, and \(\delta\) denotes the regularization parameter;

[0151] Based on the sub - vocabulary vector similarity values and comprehensive similarity values between each first statement in the first data and all the other first statements, perform clustering analysis on the first data to determine the first category and the first - category data of each category.

[0152] In this embodiment, Denotes the lexical similarity value between the \(a\) - th sub - vocabulary vector of the \(i\) - th first statement and the \(b\) - th sub - vocabulary vector of the \(j\) - th first statement in the first data.

[0153] In this embodiment, Denotes the statement similarity value between the second encoding vector of the \(i\) - th first statement and the second encoding vector of the \(j\) - th first statement in the first data.

[0154] In this embodiment, the sub - vocabulary vector similarity value represents the similarity value between all sub - vocabulary vectors included in the corresponding two statements.

[0155] In this embodiment, the comprehensive similarity value represents the similarity value determined by the sub - vocabulary vector similarity value, statement similarity value, and regularization value of the corresponding two statements.

[0156] In this embodiment, the calculation formula of the regularization value can be: where α ij represents the quantity adjustment parameter of the number of index values of the i-th first statement and the number of index values of the j-th first statement in the first data.

[0157] In this embodiment, the clustering analysis can be hierarchical clustering, K-means clustering, DBSCAN, etc.

[0158] In this embodiment, the first category includes multiple categories, and the first category data includes multiple first statements.

[0159] Advantages of the above technical solution: By determining the sub-lexical vector similarity value and the comprehensive similarity value of every two first statements in the first data and classifying the first data, while maintaining the semantic information of the original data set, it can provide a data basis for data enhancement and the extraction of constructing a classification model.

[0160] Embodiment 9:

[0161] The embodiment of the present invention provides a text compliance detection method based on an imbalanced data set. Augmenting the data set of the first category data of each category in the first category to determine the second category data of each category, including:

[0162] Based on the comprehensive similarity value of all the first statements included in the first category data of each category in the first category, determine the inter-class similarity of all the first category data;

[0163] Sort the inter-class similarities of all categories in the first category from small to large, determine the inter-class similarity sequence, and determine the minimum value, lower quartile, upper quartile, and maximum value of the inter-class similarity sequence;

[0164] Determine that all categories corresponding to the inter-class similarity within the range from the minimum value to the lower quartile are low similarity categories, determine that all categories corresponding to the inter-class similarity within the range from the lower quartile to the upper quartile are medium similarity categories, and determine that all categories corresponding to the inter-class similarity within the range from the upper quartile to the maximum value are high similarity categories;

[0165] Swap the word orders of all the first statements included in each low similarity category, and determine that all the first statements and the statements after the word order swapping of each first statement are the second category data of the corresponding low similarity category;

[0166] Perform synonym replacement on all the first statements included in each medium similarity category, and determine that all the first statements and the statements after the synonym replacement of each first statement are the second category data of the corresponding medium similarity category;

[0167] For all the first sentences included in each highly similar category, swap the word order, and determine that all the first sentences and the sentences after swapping the word order of each first sentence are the second category data corresponding to the highly similar category.

[0168] In this embodiment, the minimum value represents the first inter-class similarity of the inter-class similarity sequence, the lower quartile represents the th inter-class similarity of the inter-class similarity sequence, the lower quartile represents the th inter-class similarity of the inter-class similarity sequence, the maximum value represents the last inter-class similarity of the inter-class similarity sequence, where represents rounding up, and Nx represents the number of all categories in the first category.

[0169] In this embodiment, for highly similar categories, the back-translation method is adopted, that is, machine translation back to the original language to generate more diverse text samples.

[0170] In this embodiment, for moderately similar categories, data diversity is increased by synonym replacement while keeping the basic semantics of the sentence unchanged.

[0171] In this embodiment, for low-similarity categories, word order swapping is used to generate new text samples. This method can create different expressions without changing the meaning of the sentence.

[0172] Beneficial effects of the above technical solution: By augmenting the dataset of the first category data for each category in the first category and determining the second category data for each category, it is possible to increase the diversity and richness of the dataset while maintaining the semantic information of the original dataset, providing more comprehensive data support for the training of the classification model.

[0173] Example 10:

[0174] An embodiment of the present invention provides a text compliance detection method based on an imbalanced dataset, which processes large model data based on a classification model, including:

[0175] The classification model makes a compliance judgment on the input data classification of the large model data. At the same time, the classification model makes a compliance judgment on the output data classification of the large model data.

[0176] In this embodiment, before the user inputs the input data into the large model, the classification model will first make a compliance judgment on the input data to ensure that it does not contain any inappropriate content; similarly, after the large model outputs the data, the classification model will make a judgment on the output data to ensure that it meets the preset standards and specifications.

[0177] In this embodiment, if non-compliant content is detected, these contents will be automatically blocked or modified to ensure the security and compliance of information dissemination.

[0178] Beneficial effects of the above technical solution: Processing large model data according to the classification model can improve the speed and accuracy of text processing, strengthen the compliance supervision of large model service providers during data training and content generation, serve society more safely and responsibly, ensure the security and compliance of information dissemination, and improve the adaptability and generalization ability of large models.

[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0180] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A text compliance detection method based on an unbalanced data set, characterized in that: include: 101: Collect and preprocess a large model data set, determine first data, and determine a vocabulary based on the first data; 102: Encode the first data based on the vocabulary, perform dimensionality reduction processing on the encoded first data, and determine a second encoding vector for each first sentence in the first data; 103: Determine a sub-lexical vector similarity value and a comprehensive similarity value between every two first sentences in the first data, and classify the first data to determine a first category and first category data of each category in the first category; 104: performing data set augmentation on the first category data of each category in the first category, and determining the second category data of each category; 105: Building a classification model based on the second category data, and processing the large model data based on the classification model.

2. According to the text compliance detection method based on unbalanced data set according to claim 1, it is characterized in that: Preprocessing includes format conversion, data cleaning, and stop word removal.

3. A text compliance detection method based on an unbalanced data set according to claim 2, characterized in that: Collecting and preprocessing the large model data set to determine the first data includes: Extract all data types from the collected large model data set, determine that the data format with the largest number of occurrences is the first format, and perform format conversion on data in the large model data set that is inconsistent with the first format; Perform data cleaning on the large model data set after format conversion, where data cleaning includes whitespace standardization, special character removal, and punctuation removal; Perform stop word extraction on the large model data set after data cleaning and remove all extracted stop words; It is determined that the large model data set after the stop words are removed is the first data, wherein the first data includes a plurality of first sentences.

4. A text compliance detection method based on an unbalanced data set according to claim 3, characterized in that: Determining a vocabulary based on the first data includes: Extracting all words appearing in the first data, determining a first word, and counting the number of times each word in the first word appears in the first data, wherein the first word includes a plurality of words; Sort all words in the first vocabulary based on the number of times each word in the first vocabulary appears in the first data, determine a word sequence, and draw a word distribution map based on the word sequence and the number of times each word appears in the first data; Determine the cumulative coverage of each word based on the word distribution map; Comparing the cumulative coverage of each word one by one based on the word sequence and setting a coverage threshold, and determining the second word based on the comparison result; Determining a dependency window based on a length of a first sentence in the first data, segmenting the first data based on the dependency window, and determining a vocabulary sentence for each vocabulary in the first vocabulary for the segmented first data, wherein the vocabulary sentence includes a plurality of first sentences containing corresponding vocabulary; determining a lexical dependency value for each word in the first vocabulary based on the lexical sentence for each word in the first vocabulary; Comparing the vocabulary dependency value of each vocabulary in the first vocabulary with the set dependency value threshold, extracting all vocabulary whose vocabulary dependency value is greater than the set dependency value threshold, and determining the third vocabulary; A vocabulary is determined based on the second vocabulary and the third vocabulary.

5. A text compliance detection method based on an unbalanced data set according to claim 4, characterized in that: Comparing the cumulative coverage of each word one by one based on the word sequence and setting a coverage threshold, and determining the second word based on the comparison result, including: If the vocabulary cumulative coverage is less than the set coverage threshold, compare the vocabulary cumulative coverage of the vocabulary corresponding to the vocabulary cumulative coverage less than the set coverage threshold, the vocabulary cumulative coverage of the next vocabulary in the vocabulary sequence and the set coverage threshold; If the cumulative coverage of the vocabulary is greater than or equal to the set coverage threshold, the comparison stops, and the vocabulary corresponding to the vocabulary cumulative coverage greater than or equal to the set coverage threshold and all the vocabulary before the vocabulary in the vocabulary sequence are determined to be the second vocabulary, wherein the vocabulary order in the second vocabulary is the same as the vocabulary order in the vocabulary sequence.

6. A text compliance detection method based on an unbalanced data set according to claim 5, characterized in that: Determining a vocabulary based on the second vocabulary and the third vocabulary includes: Extracting words that appear simultaneously in the second vocabulary and the third vocabulary and determining them as common words, and adjusting the order of the common words in the second vocabulary based on the word dependency value of each word in the common vocabulary; All words in the third vocabulary except the common words are sequentially inserted into the adjusted second vocabulary based on the corresponding word dependency values; The vocabulary is determined based on all the words in the inserted second vocabulary and the order of the words.

7. A text compliance detection method based on an unbalanced data set according to claim 3, characterized in that: Encoding the first data based on the vocabulary to determine the encoded data includes: Determine a vocabulary vector based on the vocabulary table, determine the index value of each word in the vocabulary table in the vocabulary vector, and determine the corresponding sub-vocabulary vector based on the index value of each word; Searching the vocabulary vector for all the words contained in each first sentence in the first data, and extracting sub-vocabulary vectors of all the words; Adding sub-vocabulary vectors of all vocabulary contained in each first sentence to determine first encoding vectors of all first sentences; Determine a coding matrix based on the first coding vectors of all first sentences in the first data, and perform centralization processing on the coding matrix; Determine the covariance matrix of the coding matrix after centralization, perform eigenvalue decomposition on the covariance matrix, and determine the principal components; The first encoding vector of each first sentence in the first data is projected onto the principal component to determine a second encoding vector of each first sentence in the first data after dimension reduction.

8. A text compliance detection method based on an unbalanced data set according to claim 4, characterized in that: Determining a sub-lexical vector similarity value and a comprehensive similarity value of every two first sentences in the first data, and classifying the first data, including: Based on all sub-lexical vectors, lexical dependency values, and second encoding vectors of every two first sentences in the first data, calculating a sub-lexical vector similarity value and a comprehensive similarity value of the two first sentences; Among them, Sw ij Vw represents the similarity value of the sub-vocabulary vectors of the i-th first sentence and the j-th first sentence in the first data. ia 、Vw jb Respectively represent the ath sub-lexicon vector of the i-th first sentence and the bth sub-lexicon vector of the j-th first sentence in the first data, iN1 and jN1 respectively represent the number of sub-lexicon vectors of the i-th first sentence and the j-th first sentence in the first data, Rw ia Represents the vocabulary dependency value of the vocabulary corresponding to the a-th sub-lexicon vector of the i-th first sentence in the first data, Rw jb represents the vocabulary dependency value of the vocabulary corresponding to the b-th sub-lexicon vector of the j-th first sentence in the first data, S ij represents the comprehensive similarity value between the i-th first sentence and the j-th first sentence in the first data, Vs i Represents the second encoding vector of the i-th first sentence in the first data, Vs j represents the second encoding vector of the jth first sentence in the first data, Lw i , Lw j Respectively represent the number of index values ​​of the i-th first statement and the j-th first statement in the first data, f(Lw i , Lw j ) represents a regularization value based on the number of index values ​​of the i-th first sentence and the number of index values ​​of the j-th first sentence in the first data, W1 represents a sentence similarity value weight, W2 represents a vocabulary similarity value weight, and δ represents a regularization parameter; Based on the sub-lexical vector similarity values ​​and the comprehensive similarity values ​​between each first sentence and all other first sentences in the first data, a cluster analysis is performed on the first data to determine the first category and the first category data of each category.

9. The text compliance detection method based on an unbalanced data set according to claim 1, characterized in that: The first category data of each category in the first category is augmented with a data set to determine the second category data of each category, including: Determine the inter-class similarity of all first category data based on the comprehensive similarity values ​​of all first sentences contained in the first category data of each category in the first category; Sort the inter-class similarities of all categories in the first category from small to large, determine the inter-class similarity sequence, and determine the minimum value, lower quartile, upper quartile and maximum value of the inter-class similarity sequence; All categories corresponding to the inter-class similarity within the range from the minimum value to the lower quartile are determined as low similarity categories, all categories corresponding to the inter-class similarity within the range from the lower quartile to the upper quartile are determined as medium similarity categories, and all categories corresponding to the inter-class similarity within the range from the upper quartile to the maximum value are determined as high similarity categories; Transposing the word order of all first sentences included in each low-similarity category, and determining that all first sentences and each sentence after the word order of the first sentence is transformed are second category data of the corresponding low-similarity category; Perform synonym replacement on all first sentences included in each similar category, and determine that all first sentences and each sentence after the synonym replacement of the first sentence are the second category data of the corresponding similar category; The word order of all first sentences included in each high similarity category is swapped, and all first sentences and each sentence after the word order of the first sentence is swapped are determined as second category data of the corresponding high similarity category.

10. The text compliance detection method based on an unbalanced data set according to claim 1, characterized in that: Processing of large model data based on classification models, including: The classification model performs compliance judgment on the classification of input data of the large model data, and at the same time, the classification model performs compliance judgment on the classification of output data of the large model data.