Harmful corpus mining method and apparatus, and electronic device and storage medium

Through text similarity search, semantic feature classification and corpus sampling processing, the existing harmful corpus mining problems are solved, efficient and low-cost harmful corpus acquisition are achieved, and the quality of the safety classification model is improved.

WO2025124025A1PCT designated stage expired Publication Date: 2025-06-19SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/130510
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-11-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing harmful corpus mining methods are inefficient and have high manpower and material costs, resulting in low quality of safety classification models.

Method used

Through text similarity search, semantic feature classification processing and corpus sampling processing, we comprehensively consider the three dimensions of text similarity, semantics and corpus to obtain high-quality harmful corpus sample sets.

Benefits of technology

It improves the efficiency of harmful corpus mining, reduces manpower and material costs, helps to train high-quality safety classification models, and improves the efficiency of training corpus cleaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130510_19062025_PF_FP_ABST
    Figure CN2024130510_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present invention is a harmful corpus mining method. The method comprises: S1, acquiring a first harmful corpus seed sample and a corpus data set; S2, performing text similarity retrieval in the corpus data set on the basis of the first harmful corpus seed sample, so as to obtain a potential harmful corpus sample set; S3, performing semantic feature classification processing on potential harmful corpus samples in the potential harmful corpus sample set by means of a preset semantic feature classification model, so as to obtain a harmful corpus hard sample set; and S4, performing corpus sampling processing on the harmful corpus hard sample set by means of a preset corpus sampling large model, so as to obtain a first target harmful corpus sample set. The correlation between three dimensions, that is, text similarity, semantics and corpus, and harmful corpus determination is taken into comprehensive consideration, thereby improving the efficiency of harmful corpus mining and reducing manpower and material resource costs, further facilitating the training of a high-quality security classification model, and improving the efficiency of training corpus cleaning.
Need to check novelty before this filing date? Find Prior Art

Description

Harmful corpus mining method, device, electronic device and storage medium Technical Field

[0001] The present invention relates to the field of corpus mining, and in particular to a harmful corpus mining method, device, electronic device and storage medium.

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311705754.2 and invention name “Method, device, electronic device and storage medium for mining harmful corpus”, the entire contents of which are incorporated by reference into this application. Background Art

[0003] With the rapid development of software and hardware technology, language model technology has also become the focus of people's attention. The training of language models depends on the quality of the training corpus. If the training corpus contains text content that violates laws and morals, the language model trained can easily mislead and have adverse effects on users. Therefore, it is necessary to perform security cleaning on the training corpus. The existing technology mainly uses security classification models to clean the corpus. However, when the current security classification model is trained, the acquisition of harmful corpus samples mainly relies on manual retrieval and labeling on the Internet, which is extremely inefficient and has high manpower and material costs, resulting in low quality of the security classification model. Therefore, how to provide a harmful corpus mining method to solve the existing harmful corpus mining problems of extremely low efficiency and high manpower and material costs has become an urgent problem to be solved. Technical issues

[0004] The present invention provides a method for mining harmful corpus, aiming to address the current problems of extremely low efficiency and high labor and material costs associated with harmful corpus mining. A first target harmful corpus sample set is obtained through text similarity retrieval, semantic feature classification, and corpus sampling. The method comprehensively considers the relevance of three dimensions—text similarity, semantics, and corpus—to harmful corpus judgment. This method improves the efficiency of harmful corpus mining, reduces labor and material costs, and ultimately helps train a high-quality security classification model, improving the efficiency of training corpus cleaning.

[0005] In a first aspect, an embodiment of the present invention provides a method for mining harmful corpus, characterized in that the method comprises the following steps:

[0006] S1. Obtain a first harmful corpus seed sample and a corpus dataset, where the corpus dataset is a public pre-training corpus dataset for a large language model;

[0007] S2. Performing a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples;

[0008] S3. Performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set;

[0009] S4. Perform corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set.

[0010] Optionally, the step S2 of performing text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set includes:

[0011] Building a corpus data retrieval library based on the corpus data set;

[0012] Taking the first harmful corpus seed sample as a retrieval target, performing a text similarity search in the corpus data retrieval database to obtain a retrieval result for each first harmful corpus seed sample, wherein the retrieval result includes a text similarity matching score value for each retrieval result;

[0013] Sort the search results by similarity based on the text similarity matching score values ​​to obtain a matching result ranking table;

[0014] Based on the matching result ranking table, the potentially harmful corpus sample set is determined.

[0015] Optionally, before performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model in S3 to obtain a difficult harmful corpus sample set, the method further includes:

[0016] Obtain the semantic feature classification model to be trained;

[0017] Collecting a first safe corpus seed sample according to the data scale of the first harmful corpus seed sample;

[0018] The first harmful corpus seed sample is used as a positive sample, and the first safe corpus seed sample is used as a negative sample, and the samples are input into the semantic feature classification model to be trained for classification training to obtain the preset semantic feature classification model.

[0019] Optionally, the step S3 of performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set includes:

[0020] Inputting each potentially harmful corpus sample in the potentially harmful corpus sample set into the preset semantic feature classification model for classification processing to obtain a classification score value for each potentially harmful corpus sample;

[0021] The difficult sample set of harmful corpus is determined based on the classification score value of each of the potentially harmful corpus samples.

[0022] Optionally, the step S4 of performing corpus sampling processing on the difficult sample set of harmful corpus using a preset corpus sampling large model to obtain a first target harmful corpus sample set includes:

[0023] Determining the harmfulness type of each harmful corpus difficult sample in the harmful corpus difficult sample set based on the harmful corpus difficult sample set;

[0024] Determining, based on the harmful type, input prompt words corresponding to each difficult sample of the harmful corpus;

[0025] Each of the harmful corpus difficult samples and the input prompt words corresponding to each of the harmful corpus difficult samples are input into the preset corpus sampling large model for corpus sampling processing to obtain the first target harmful corpus sample set.

[0026] Optionally, in step S4, after performing corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set, the method further includes:

[0027] Comparing the data size of the first target harmful corpus sample set with a data size threshold;

[0028] If the data size of the first target harmful corpus sample set is smaller than the data size threshold, fusing the first target harmful corpus sample set with the first harmful corpus seed sample to obtain a second harmful corpus seed sample;

[0029] The steps S1 to S4 are iteratively updated based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set, where the data size of the second target harmful corpus sample set is greater than the data size threshold.

[0030] Optionally, iteratively updating steps S1 to S4 based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set includes:

[0031] Fusing all first target harmful corpus sample sets and the first harmful corpus seed samples in historical iteration rounds to obtain updated harmful corpus seed samples;

[0032] Collect a second safe corpus seed sample according to the data scale of the updated harmful corpus seed sample;

[0033] The updated harmful corpus seed sample is used as a positive sample, and the second safe corpus seed sample is used as a negative sample to input into the preset semantic feature classification model for update training to obtain an updated semantic feature classification model.

[0034] In a second aspect, an embodiment of the present invention further provides a harmful corpus mining device, the harmful corpus mining device comprising:

[0035] A first acquisition module is configured to acquire a first harmful corpus seed sample and a corpus dataset, wherein the corpus dataset is a public dataset of pre-trained corpus for a large language model;

[0036] a first retrieval module configured to perform a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes each potentially harmful corpus sample;

[0037] A first classification module is configured to perform semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set;

[0038] The first sampling module is used to perform corpus sampling processing on the harmful corpus difficult sample set through a preset corpus sampling large model to obtain a first target harmful corpus sample set.

[0039] In a third aspect, an embodiment of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the harmful corpus mining method provided in an embodiment of the present invention are implemented.

[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the harmful corpus mining method provided in the embodiment of the invention are implemented.

[0041] In an embodiment of the present invention, S1 obtains a first harmful corpus seed sample and a corpus dataset, wherein the corpus dataset is a public dataset of pre-trained corpus for a large language model; S2 performs a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples; S3 performs semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set; S4 performs corpus sampling processing on the difficult harmful corpus sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set. The first target harmful corpus sample set is obtained through text similarity retrieval processing, semantic feature classification processing, and corpus sampling processing, comprehensively considering the relevance of the three dimensions of text similarity, semantics, and corpus with harmful corpus judgment, thereby improving the efficiency of harmful corpus mining and reducing manpower and material costs, thereby helping to train a high-quality security classification model and improving the efficiency of training corpus cleaning. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] FIG1 is a flow chart of a harmful corpus mining method provided by an embodiment of the present invention;

[0044] FIG2 is a schematic diagram of the structure of a harmful corpus mining device provided in an embodiment of the present invention;

[0045] FIG3 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Modes for Carrying Out the Invention

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] As shown in FIG1 , FIG1 is a flowchart of a harmful corpus mining method provided by an embodiment of the present invention, including:

[0048] S1. Obtain a first harmful corpus seed sample and a corpus dataset.

[0049] In embodiments of the present invention, the aforementioned harmful corpus mining method can be applied within a corpus management platform, which can be constructed from a server or server cluster. These servers or server clusters can be electronic devices capable of text processing, text recognition, data transmission, and data storage. The corpus management platform can respond to user input of harmful corpus seed samples and data size requirements, and, based on the text in the corpus dataset, output a target harmful corpus dataset that meets the required data size.

[0050] The first harmful corpus seed sample may be text content that is contrary to mainstream opinion and violates legal ethics. The corpus dataset may be a public dataset of pre-trained corpus for a large language model, for example, the "WuDao dataset", the "Shusheng Wanjuan dataset", the "Common Crawl dataset", etc. The corpus dataset may include multiple corpus texts, and the corpus texts may be paragraphs, sentences, phrases, words, characters, and other texts.

[0051] Specifically, the first harmful corpus seed sample can be obtained by any easily accessible method. For example, the first harmful corpus seed sample can be obtained by searching for harmful corpus seed samples on the Internet using harmful keywords. The harmful keywords can be keywords such as violence and privacy, or can be artificially created or purchased through commercial channels.

[0052] In a possible embodiment, after the corpus platform obtains the harmful corpus seed sample and data scale requirement input by the user, it outputs a target harmful corpus dataset that meets the data scale requirement based on the text in the corpus dataset.

[0053] In another possible embodiment, when the above-mentioned corpus platform only obtains the data scale requirement input by the above-mentioned user, it can search for harmful corpus seed samples on the Internet using the above-mentioned harmful keywords to obtain the above-mentioned first harmful corpus seed samples, and perform harmful corpus mining based on the above-mentioned corpus data set according to the above-mentioned harmful corpus seed samples to obtain the above-mentioned first harmful corpus data set.

[0054] S2. Perform text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set.

[0055] In an embodiment of the present invention, the above-mentioned potentially harmful corpus sample set includes various potentially harmful corpus samples, and the above-mentioned text similarity retrieval can be a processing flow or processing method based on text similarity. The text similarity between the above-mentioned first harmful corpus seed sample and multiple corpus texts in the above-mentioned corpus data set can be calculated to obtain a text similarity matching score value between each corpus text and each first harmful corpus seed sample, and the above-mentioned potentially harmful corpus sample set is screened out from the above-mentioned corpus text according to the above-mentioned text similarity matching score value.

[0056] In a possible embodiment, the above-mentioned first harmful corpus seed sample and the corpus text in the above-mentioned corpus dataset can be subjected to text preprocessing respectively to obtain the preprocessed first harmful corpus seed sample and the preprocessed corpus text. The above-mentioned preprocessing can be word segmentation processing, stop word removal processing and stem extraction processing, etc. Through the above-mentioned preprocessing, the above-mentioned corpus text and the above-mentioned first harmful corpus seed sample can be converted from text form to word or phrase form.

[0057] The text similarity between each preprocessed first harmful corpus seed sample and each preprocessed corpus text is calculated using a similarity algorithm to obtain a text similarity matching score between each first harmful corpus seed sample and each corpus text. The potentially harmful corpus sample set is screened from the corpus text based on the text similarity matching score. The similarity algorithm can be a cosine similarity algorithm, a Jaccard similarity algorithm, or a Euclidean distance algorithm.

[0058] Specifically, the word frequency or TF-IDF value (i.e., Term Frequency-Inverse Document Frequency) of each word or phrase in the above-mentioned preprocessed first harmful corpus seed sample and the preprocessed corpus text can be calculated respectively, and a first word frequency vector is constructed according to the word frequency or TF-IDF value of each word or phrase in each preprocessed first harmful corpus seed sample, and a second word frequency vector is constructed according to the word frequency or TF-IDF value of each word or phrase in each preprocessed corpus text.

[0059] The cosine value of the angle between the first term frequency vector and the second term frequency vector is calculated using the cosine similarity algorithm, and the text similarity between each of the first harmful corpus seed samples and each of the corpus texts is determined to obtain the text similarity matching score. It should be noted that the closer the direction of the first term frequency vector and the second term frequency vector match, and the closer the angle is to zero, the closer the cosine value is to 1. This indicates that the first term frequency vector and the second term frequency vector are more similar, and the text similarity matching score between the first harmful corpus seed sample corresponding to the first term frequency vector and the corpus text corresponding to the second term frequency vector is higher, and vice versa.

[0060] In a possible embodiment, after the above-mentioned corpus management platform obtains the first harmful corpus seed sample and the corpus data set, the text similarity between each first harmful corpus seed sample and each corpus text in the corpus data set is calculated by the above-mentioned similarity algorithm to obtain the text similarity matching score value between each first harmful corpus seed sample and each corpus text in the corpus data set, and the above-mentioned potential harmful corpus sample set is screened out from the above-mentioned corpus data set based on the above-mentioned text similarity matching score value.

[0061] S3. Perform semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus.

[0062] In an embodiment of the present invention, the above-mentioned semantic features may include natural semantic features and accessory semantic features. The above-mentioned natural semantic features may be understood as objective semantic features decomposed from the text based on basic concepts and logical meanings, and the above-mentioned accessory semantic features may be understood as non-natural subjective semantic features in semantics, such as emotional color, stylistic color, and image color. The above-mentioned preset semantic feature classification model may be a semantic feature classification model obtained by re-classification training of a pre-trained BERT (Bidirectional Encoder Representation from Transformers) model, or a semantic feature classification model obtained by re-classification training of a pre-trained HRNet, Unet++, or SegFormer model.

[0063] The above-mentioned harmful corpus difficult sample set may include various harmful corpus difficult samples. The above-mentioned harmful corpus difficult samples may be potential harmful corpus samples that are difficult to classify by the above-mentioned preset semantic feature classification model. The above-mentioned preset semantic feature classification model is used to perform semantic feature classification processing on each potential harmful corpus sample in the above-mentioned potential harmful corpus sample set to obtain a classification score value for each potential harmful corpus sample. According to the above-mentioned classification score value, the potential harmful corpus samples that are difficult to classify by the above-mentioned preset semantic feature classification model are determined from the above-mentioned potential harmful corpus samples as the above-mentioned harmful corpus difficult samples, thereby obtaining the above-mentioned harmful corpus difficult sample set.

[0064] Specifically, the above-mentioned semantic feature classification can be a processing flow or processing method for binary classification based on semantic features. Taking the above-mentioned pre-trained BERT model as an example, the first safe corpus seed sample can be obtained according to the data scale of the above-mentioned first harmful corpus seed sample. The data scale of the above-mentioned first harmful corpus seed sample is consistent with the data scale of the above-mentioned first safe corpus seed sample. The above-mentioned first safe corpus seed sample can be the text content disclosed in public publications, the text content disclosed in official media, and the text content disclosed on Wikipedia. The above-mentioned first harmful corpus seed sample is used as a positive sample, and the above-mentioned first safe corpus seed sample is used as a negative sample. They are input into the above-mentioned pre-trained BERT model for classification training to obtain the above-mentioned preset semantic feature classification model.

[0065] In a possible embodiment, when the above-mentioned corpus management platform performs text similarity retrieval in the corpus data set based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, each potentially harmful corpus sample in the above-mentioned potentially harmful corpus sample set is input one by one into the above-mentioned preset semantic feature classification model for classification processing to obtain a classification score value for each potentially harmful corpus sample. According to the classification score value of each potentially harmful corpus sample, difficult harmful corpus samples are screened out from the above-mentioned potentially harmful corpus sample set to obtain the above-mentioned difficult harmful corpus sample set.

[0066] S4. Perform corpus sampling processing on the difficult sample set of harmful corpus using a preset corpus sampling large model to obtain a first target harmful corpus sample set.

[0067] In an embodiment of the present invention, the above-mentioned preset corpus sampling large model can be any large deep learning model that can perform text recognition and text processing, for example, it can be a large deep learning model such as the ChatGPT large model, the Wenxin Yiyan large model, and the Chang'e large model.

[0068] Specifically, input prompt words can be constructed according to the harmful type of the harmful corpus difficulty samples in the above-mentioned harmful corpus difficulty sample set, and each harmful corpus difficulty sample and the corresponding input prompt word can be input into the above-mentioned preset corpus sampling large model for corpus sampling processing. The above-mentioned preset corpus sampling large model outputs the output results corresponding to each harmful corpus difficulty sample, and the above-mentioned first target harmful corpus sample set is determined based on the above-mentioned output results.

[0069] The harmful types may include advertising harmful types, violence harmful types, personal privacy harmful types, hate speech harmful types, malicious attack harmful types, false information harmful types, malicious code software harmful types, etc. The output result may be "safe" or "harmful".

[0070] For example, when the harmful type of the difficult harmful corpus sample is hate speech, the corresponding input prompt word may be "I will input a text at a time. If the text contains hate speech, then you should output 'harmful'; otherwise, you should output 'safe'." The difficult harmful corpus sample and the input prompt word are simultaneously input into the preset corpus sampling large model for corpus sampling processing to obtain the output result. If the output result is harmful, the difficult harmful corpus sample is added to the first target harmful corpus sample set; otherwise, it is not added.

[0071] When the harmful type of the difficult harmful corpus sample is a malicious attack type, the corresponding input prompt word can be "I will input a text each time. If the text contains malicious attack content, then you should output 'harmful', otherwise you should output 'safe'." The difficult harmful corpus sample and the input prompt word are simultaneously input into the preset corpus sampling large model for corpus sampling processing to obtain the above output result. If the above output result is harmful, the difficult harmful corpus sample is added to the first target harmful corpus sample set; otherwise, it is not added.

[0072] In a possible embodiment, when the above-mentioned corpus management platform performs semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set through a preset semantic feature classification model to obtain a harmful corpus difficult sample set, according to the harmful type of each harmful corpus difficult sample in the above-mentioned harmful corpus difficult sample set, an input prompt word corresponding to each harmful corpus difficult sample is constructed, and each harmful corpus difficult sample and the corresponding input prompt word are input into the corpus sampling large model for corpus sampling processing to obtain the above-mentioned output result, and the harmful corpus difficult sample with the output result of harmful is determined as the first target harmful corpus sample to obtain the above-mentioned first target harmful corpus sample set.

[0073] In an embodiment of the present invention, S1 obtains a first harmful corpus seed sample and a corpus dataset, wherein the corpus dataset is a public dataset of pre-trained corpus for a large language model; S2 performs a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples; S3 performs semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set; S4 performs corpus sampling processing on the difficult harmful corpus sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set. The first target harmful corpus sample set is obtained through text similarity retrieval processing, semantic feature classification processing, and corpus sampling processing, comprehensively considering the relevance of the three dimensions of text similarity, semantics, and corpus with harmful corpus judgment, thereby improving the efficiency of harmful corpus mining and reducing manpower and material costs, thereby helping to train a high-quality security classification model and improving the efficiency of training corpus cleaning.

[0074] Optionally, in S2, the step of performing a text similarity search in the corpus data set based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, a corpus data retrieval library can also be constructed based on the corpus data set; taking the first harmful corpus seed sample as the retrieval target, performing a text similarity search in the corpus data retrieval library to obtain a retrieval result for each first harmful corpus seed sample, the retrieval result including a text similarity matching score value for each retrieval result; sorting the retrieval results by similarity based on the text similarity matching score value to obtain a matching result sorting table; and determining the potentially harmful corpus sample set based on the matching result sorting table.

[0075] In an embodiment of the present invention, the corpus data retrieval library can be constructed based on the corpus dataset. Specifically, a distributed search tool can be used as a search engine to construct the corpus data retrieval library. The search tool can be an Elasticsearch search tool or a Sphinx search tool. More specifically, the corpus text in the corpus dataset can be segmented to obtain segmented corpus text. An inverted index can be established based on the segmented corpus text and distributedly deployed to obtain the corpus data retrieval library.

[0076] The first harmful corpus seed sample is used as a search target, and a text similarity search is performed in the corpus data search database to obtain a text similarity matching score value between each corpus text and the first harmful corpus seed sample. Corpus samples with text similarity matching scores greater than 0 are used as search results corresponding to the first harmful corpus seed sample. The search results are sorted by similarity based on the text similarity matching scores to obtain a matching result ranking table. In the matching result ranking table, the higher the ranking, the more similar the corpus text corresponding to the corresponding first harmful corpus seed sample is, and vice versa. It should be noted that the first harmful corpus seed sample can be one seed sample or multiple seed samples, each seed sample can correspond to one search result or multiple search results, and each search result can include one corpus text and a text similarity matching score value corresponding to one corpus text.

[0077] After obtaining the aforementioned ranking table of matching results, a ranking threshold for matching results can be set based on the ranking table. Search results with text similarity matching scores greater than the ranking threshold for matching results can be identified as potentially harmful corpus samples, thereby obtaining the aforementioned potentially harmful corpus sample set. The ranking threshold for matching results can be set based on the size of the ranking table. The larger the size of the ranking table, the smaller the ranking threshold, and vice versa. The size of the ranking table can be understood as the data size of the ranking table.

[0078] Specifically, when the size of the ranked matching results table is large, the ranking threshold can be set to 30% (i.e., only the top 30% of search results in the ranked matching results table are selected). When the size of the ranked matching results table is small, the ranking threshold can be set to 60% (i.e., only the top 60% of search results in the ranked matching results table are selected). It should be noted that when the data size of the ranked matching results table is large enough, only a small number of search results with high rankings can be selected to ensure search accuracy. When the data size of the ranked matching results table is small, it is necessary to select search results with relatively low rankings to ensure the data size of the potentially harmful corpus dataset.

[0079] Optionally, before the step of performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set by using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus in S3, a semantic feature classification model to be trained can also be obtained; a first safe corpus seed sample is collected according to the data scale of the first harmful corpus seed sample; the first harmful corpus seed sample is used as a positive sample, and the first safe corpus seed sample is used as a negative sample, and the samples are input into the semantic feature classification model to be trained for classification training to obtain a preset semantic feature classification model.

[0080] In an embodiment of the present invention, the semantic feature classification model to be trained may be the pre-trained BERT (Bidirectional Encoder Representation from Transformers) model or a pre-trained HRNet, Unet++, SegFormer or other model. The first safe corpus seed sample may be a text content that is mainstream and does not violate legal ethics. The first safe corpus seed sample may be obtained from public publications, official media, Wikipedia and other channels. The first harmful corpus seed sample and the first safe corpus seed sample of the same data scale are simultaneously input into the semantic feature classification model to be trained. The first harmful corpus seed sample is used as a positive sample, and the first safe corpus seed sample is used as a negative sample. The semantic feature classification model to be trained is classified and trained to obtain the preset semantic feature classification model.

[0081] In a possible embodiment, the above-mentioned semantic feature classification can be a semantic feature binary classification. The first harmful corpus seed sample and the first safe corpus seed sample of the same data scale can be simultaneously input into the above-mentioned semantic feature classification model to be trained. The above-mentioned first harmful corpus seed sample is used as a positive sample, and the above-mentioned first safe corpus seed sample is used as a negative sample. The above-mentioned semantic feature classification model to be trained is trained in binary classification to obtain the above-mentioned preset semantic feature classification model.

[0082] Optionally, in S3, in the step of performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set through a preset semantic feature classification model to obtain a difficult sample set of harmful corpus, each potentially harmful corpus sample in the potentially harmful corpus sample set can also be input into a preset semantic feature classification model for classification processing to obtain a classification score value of each potentially harmful corpus sample; based on the classification score value of each potentially harmful corpus sample, the difficult sample set of harmful corpus is determined.

[0083] In an embodiment of the present invention, the classification score values ​​correspond to the potentially harmful corpus samples, with each potentially harmful corpus sample corresponding to a classification score value. The classification score values ​​may range from 0 to 1. It should be noted that the closer the classification score value is to 1 for a potentially harmful corpus sample, the more likely it is harmful corpus, while the closer the classification score value is to 0, the more likely it is safe corpus. In this case, a classification score threshold can be determined based on the score interval, and the classification score values ​​of each potentially harmful corpus sample can be compared with the classification score threshold to obtain a comparison result. Based on the comparison result, the difficult sample set of harmful corpus can be determined.

[0084] Specifically, when the score range is between 0 and 1, the median of the range, 0.5 + (±0.1), can be used as the classification score threshold (i.e., 0.4 to 0.6 can be used as the classification score threshold). The classification score of each potentially harmful corpus sample is compared with the classification score threshold. Potentially harmful corpus samples with scores below the classification score threshold are identified as safe corpus, potentially harmful corpus samples with scores above the classification score threshold are identified as harmful corpus, and potentially harmful corpus samples within the threshold range of the classification score threshold are identified as difficult harmful corpus samples, thereby obtaining the difficult harmful corpus sample set. It is understood that potentially harmful corpus samples within the threshold range of the classification score threshold have certain similarities with both the first safe corpus seed sample and the first harmful corpus seed sample in terms of semantic features. From the perspective of harmful corpus mining, potentially harmful corpus samples within the threshold range of the classification score threshold are more valuable.

[0085] Optionally, in S4, in the step of performing corpus sampling processing on the harmful corpus difficult sample set through a preset corpus sampling large model to obtain a first target harmful corpus sample set, the harmful type of each harmful corpus difficult sample in the harmful corpus difficult sample set can be determined based on the harmful corpus difficult sample set; based on the harmful type, the input prompt word corresponding to each harmful corpus difficult sample can be determined; each harmful corpus difficult sample and the input prompt word corresponding to each harmful corpus difficult sample can be input into the preset corpus sampling large model for corpus sampling processing to obtain the first target harmful corpus sample set.

[0086] In an embodiment of the present invention, an input prompt word corresponding to each harmful corpus difficult sample can be constructed according to the harmful type of each harmful corpus difficult sample. The above-mentioned input prompt word is used to prompt the above-mentioned preset corpus sampling large model to perform corpus sampling processing. The above-mentioned preset corpus sampling large model can perform corpus sampling processing on each of the above-mentioned harmful corpus difficult samples according to the above-mentioned input prompt word, and output the corpus sampling results of each harmful corpus difficult sample. The above-mentioned corpus sampling results can be "safe" or "harmful". According to the corpus sampling results of each of the above-mentioned harmful corpus difficult samples, each of the above-mentioned harmful corpus difficult samples is labeled to obtain a labeled harmful corpus difficult sample set. In the above-mentioned labeled harmful corpus difficult sample set, the harmful corpus difficult samples labeled as "harmful" are used as the first target harmful corpus sample to obtain the above-mentioned first target harmful corpus sample set.

[0087] In a possible embodiment, after taking the difficult harmful corpus samples with the label "harmful" from the difficult harmful corpus samples with the label as the first target harmful corpus samples, the first target harmful corpus sample set is obtained. A part of the difficult harmful corpus samples with the label "safe" can also be randomly sampled and added to the first target harmful corpus sample set to avoid sampling bias in the corpus sampling processing.

[0088] Optionally, after the step of performing corpus sampling processing on the harmful corpus difficult sample set through a preset corpus sampling large model to obtain a first target harmful corpus sample set in S4, the data scale of the first target harmful corpus sample set can also be compared with a data scale threshold; if the data scale of the first target harmful corpus sample set is smaller than the data scale threshold, the first target harmful corpus sample set is fused with the first harmful corpus seed sample to obtain a second harmful corpus seed sample; and iteratively updating steps S1 to S4 based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set, wherein the data scale of the second target harmful corpus sample set is larger than the data scale threshold.

[0089] In an embodiment of the present invention, the data size threshold can be determined based on user needs (which can be understood as the data size required by the user), and the fusion process can be fuzzy deduplication or superposition. The data size of the first target harmful corpus sample set is compared with the data size threshold. If the data size of the first target harmful corpus sample set is less than the data size threshold, the first target harmful corpus sample set is fused with the first harmful corpus seed sample to obtain a second harmful corpus seed sample. Based on the second harmful corpus seed sample, steps S1 to S4 are iteratively updated to obtain a second target harmful corpus sample set.

[0090] Specifically, if the data size of the first target harmful corpus sample set obtained by the t-th harmful corpus mining is smaller than the data size threshold, the first target harmful corpus sample set obtained by the t-th harmful corpus mining and the first harmful corpus seed sample obtained by the t-th harmful corpus mining can be fused to obtain the second harmful corpus seed sample required for the t+1-th harmful corpus mining. Based on the second harmful corpus seed sample required for the t+1-th harmful corpus mining, a text similarity search is performed in the above corpus data set to obtain the potential harmful corpus sample set obtained by the t+1-th harmful corpus mining. The above t+1-th harmful corpus mining is classified by a preset semantic feature classification model. The potential harmful corpus sample set obtained from the first harmful corpus mining is subjected to semantic feature classification processing to obtain the harmful corpus difficult sample set obtained from the t+1th harmful corpus mining. The harmful corpus difficult sample set obtained from the above t+1th harmful corpus mining is subjected to corpus sampling processing through a preset corpus sampling large model to obtain the first target harmful corpus sample set obtained from the t+1th harmful corpus mining. When the data scale of the above first target corpus sample set is greater than the above data scale threshold or reaches a preset number of iterations, the above iterative update is stopped, and the first target corpus sample set greater than the above data scale threshold is determined as the above second target corpus sample set.

[0091] It should be noted that the iterative update of the above-mentioned harmful corpus mining can be abstractly understood as expanding outward step by step from a small number of seed nodes in a graph structure. The edges between the nodes in the graph structure are composed of the above-mentioned text similarity matching score values ​​and the above-mentioned classification score values ​​weighted, and the above-mentioned corpus sampling model is used to perform node sampling to reduce the complexity of the graph structure.

[0092] Optionally, in the step of iteratively updating steps S1 to S4 based on the second harmful corpus seed sample to obtain the second target harmful corpus sample set, all first target harmful corpus sample sets and first harmful corpus seed samples in historical iteration rounds can also be fused to obtain updated harmful corpus seed samples; second safe corpus seed samples are collected according to the data scale of the updated harmful corpus seed samples; the updated harmful corpus seed samples are used as positive samples, and the second safe corpus seed samples are used as negative samples to input into the preset semantic feature classification model for update training to obtain an updated semantic feature classification model.

[0093] In an embodiment of the present invention, if the current iteration round is the t+1th iteration round, all iteration rounds before the t+1th harmful corpus mining are historical iteration rounds. All first harmful corpus sample sets and first harmful corpus seed samples obtained before the t+1th harmful corpus mining are fused to obtain updated harmful corpus seed samples. Second safe corpus seed samples are collected according to the data scale of the updated harmful corpus seed samples. The updated harmful corpus seed samples are used as positive samples, and the second safe corpus seed samples are used as negative samples. They are input into a preset semantic feature classification model for update training to obtain an updated semantic feature classification model. The updated semantic feature classification model can be used for the t+1th harmful corpus mining.

[0094] As shown in FIG2 , an embodiment of the present invention further provides a harmful corpus mining device, comprising:

[0095] A first acquisition module 201 is configured to acquire a first harmful corpus seed sample and a corpus dataset, wherein the corpus dataset is a public dataset of pre-trained corpus for a large language model;

[0096] A first retrieval module 202 is configured to perform a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples;

[0097] The first classification module 203 is configured to perform semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set;

[0098] The first sampling module 204 is configured to perform corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set.

[0099] Optionally, the first retrieval module 202 includes:

[0100] A first construction submodule is used to construct a corpus data retrieval library based on the corpus data set;

[0101] A first retrieval submodule is configured to use the first harmful corpus seed sample as a retrieval target, perform a text similarity search in the corpus data retrieval database, and obtain a retrieval result for each first harmful corpus seed sample, wherein the retrieval result includes a text similarity matching score value for each retrieval result;

[0102] A first sorting submodule is configured to sort the search results by similarity based on the text similarity matching score values ​​to obtain a matching result sorting table;

[0103] The first determination submodule is configured to determine the potentially harmful corpus sample set based on the matching result ranking table.

[0104] Optionally, the harmful corpus mining device further includes:

[0105] The second acquisition module is used to obtain the semantic feature classification model to be trained;

[0106] A first collection module is configured to collect a first safe corpus seed sample according to the data size of the first harmful corpus seed sample;

[0107] The first training module is used to input the first harmful corpus seed sample as a positive sample and the first safe corpus seed sample as a negative sample into the semantic feature classification model to be trained for classification training to obtain the preset semantic feature classification model.

[0108] Optionally, the first classification module 203 includes:

[0109] A first classification submodule is configured to input each potentially harmful corpus sample in the potentially harmful corpus sample set into the preset semantic feature classification model for classification processing, thereby obtaining a classification score value for each potentially harmful corpus sample;

[0110] The second determining submodule is configured to determine the difficult sample set of harmful corpus based on the classification score value of each of the potentially harmful corpus samples.

[0111] Optionally, the first sampling module 204 includes:

[0112] A third determining submodule is configured to determine the harmfulness type of each harmful corpus difficult sample in the harmful corpus difficult sample set based on the harmful corpus difficult sample set;

[0113] A fourth determination submodule is configured to determine, based on the harmful type, an input prompt word corresponding to each difficult sample of the harmful corpus;

[0114] The first sampling submodule is used to input each of the harmful corpus difficult samples and the input prompt words corresponding to each of the harmful corpus difficult samples into the preset corpus sampling large model for corpus sampling processing to obtain the first target harmful corpus sample set.

[0115] Optionally, the harmful corpus mining device further includes:

[0116] A first comparison module, configured to compare the data size of the first target harmful corpus sample set with a data size threshold;

[0117] a first fusion module, configured to fuse the first target harmful corpus sample set with the first harmful corpus seed sample to obtain a second harmful corpus seed sample if the data size of the first target harmful corpus sample set is smaller than the data size threshold;

[0118] The first iterative module is configured to iteratively update the steps S1 to S4 based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set, wherein the data size of the second target harmful corpus sample set is greater than the data size threshold.

[0119] Optionally, the first iteration module includes:

[0120] A first fusion submodule is configured to fuse all first target harmful corpus sample sets and the first harmful corpus seed samples in historical iteration rounds to obtain updated harmful corpus seed samples;

[0121] A first collection submodule is configured to collect a second safe corpus seed sample according to the data size of the updated harmful corpus seed sample;

[0122] The first training submodule is used to input the updated harmful corpus seed sample as a positive sample and the second safe corpus seed sample as a negative sample into the preset semantic feature classification model for update training to obtain an updated semantic feature classification model.

[0123] As shown in FIG3 , an embodiment of the present invention further provides an electronic device, characterized in that it includes a processor, and the processor can execute any one of the above-mentioned harmful corpus mining methods.

[0124] Specifically, the system includes a processor 301, a memory 302, and a computer program for executing a harmful corpus mining method stored in the memory 302 and capable of running on the processor 301, wherein:

[0125] The processor 301 runs the computer program of the harmful corpus mining method stored in the memory 302 and performs the following steps:

[0126] S1. Obtain a first harmful corpus seed sample and a corpus dataset, where the corpus dataset is a public pre-training corpus dataset for a large language model;

[0127] S2. Performing a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples;

[0128] S3. Performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set;

[0129] S4. Perform corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set.

[0130] Optionally, the step S2 performed by the processor 301 of performing a text similarity search in the corpus dataset based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set includes:

[0131] Building a corpus data retrieval library based on the corpus data set;

[0132] Taking the first harmful corpus seed sample as a retrieval target, performing a text similarity search in the corpus data retrieval database to obtain a retrieval result for each first harmful corpus seed sample, wherein the retrieval result includes a text similarity matching score value for each retrieval result;

[0133] Sorting the search results by similarity based on the text similarity matching score values ​​to obtain a matching result ranking table;

[0134] Based on the matching result ranking table, the potentially harmful corpus sample set is determined.

[0135] Optionally, before performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set in S3, the processor 301 further executes:

[0136] Obtain the semantic feature classification model to be trained;

[0137] Collecting a first safe corpus seed sample according to the data scale of the first harmful corpus seed sample;

[0138] The first harmful corpus seed sample is used as a positive sample, and the first safe corpus seed sample is used as a negative sample, and the samples are input into the semantic feature classification model to be trained for classification training to obtain the preset semantic feature classification model.

[0139] Optionally, the step S3 executed by the processor 301 of performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set using a preset semantic feature classification model to obtain a difficult harmful corpus sample set includes:

[0140] Inputting each potentially harmful corpus sample in the potentially harmful corpus sample set into the preset semantic feature classification model for classification processing to obtain a classification score value for each potentially harmful corpus sample;

[0141] The difficult sample set of harmful corpus is determined based on the classification score value of each of the potentially harmful corpus samples.

[0142] Optionally, the step S4 executed by the processor 301 of performing corpus sampling processing on the difficult harmful corpus sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set includes:

[0143] Determining the harmfulness type of each harmful corpus difficult sample in the harmful corpus difficult sample set based on the harmful corpus difficult sample set;

[0144] Determining, based on the harmful type, input prompt words corresponding to each difficult sample of the harmful corpus;

[0145] Each of the harmful corpus difficult samples and the input prompt words corresponding to each of the harmful corpus difficult samples are input into the preset corpus sampling large model for corpus sampling processing to obtain the first target harmful corpus sample set.

[0146] Optionally, in step S4, after performing corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set, the processor 301 may further execute:

[0147] Comparing the data size of the first target harmful corpus sample set with a data size threshold;

[0148] If the data size of the first target harmful corpus sample set is smaller than the data size threshold, fusing the first target harmful corpus sample set with the first harmful corpus seed sample to obtain a second harmful corpus seed sample;

[0149] The steps S1 to S4 are iteratively updated based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set, where the data size of the second target harmful corpus sample set is greater than the data size threshold.

[0150] Optionally, the iterative updating of steps S1 to S4 based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set performed by the processor 301 includes:

[0151] Fusing all first target harmful corpus sample sets and the first harmful corpus seed samples in historical iteration rounds to obtain updated harmful corpus seed samples;

[0152] Collect a second safe corpus seed sample according to the data scale of the updated harmful corpus seed sample;

[0153] The updated harmful corpus seed sample is used as a positive sample, and the second safe corpus seed sample is used as a negative sample to input into the preset semantic feature classification model for update training to obtain an updated semantic feature classification model.

[0154] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the various processes of the harmful corpus mining method or the application-side harmful corpus mining method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0155] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware using a computer program. The computer program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0156] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A harmful corpus mining method, characterized in that: The method comprises the following steps: S1. Obtain a first harmful corpus seed sample and a corpus dataset, where the corpus dataset is a public dataset of pre-trained corpus of a large language model; S2. Performing a text similarity search in the corpus data set based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes various potentially harmful corpus samples; S3, performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set by using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus; S4. Perform corpus sampling processing on the harmful corpus difficult sample set using a preset corpus sampling large model to obtain a first target harmful corpus sample set.

2. The harmful corpus mining method according to claim 1, characterized in that: The step S2, performing text similarity search in the corpus data set based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, comprises: Building a corpus data retrieval library based on the corpus data set; Taking the first harmful corpus seed sample as a retrieval target, performing a text similarity search in the corpus data retrieval database, and obtaining a retrieval result for each of the first harmful corpus seed samples, wherein the retrieval result includes a text similarity matching score value of each of the retrieval results; Sort the search results by similarity based on the text similarity matching score values ​​to obtain a matching result sorting table; Based on the matching result ranking table, the potentially harmful corpus sample set is determined.

3. The harmful corpus mining method according to claim 1, characterized in that: Before S3, performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set by using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus, the method further includes: Obtain the semantic feature classification model to be trained; Collecting a first safe corpus seed sample according to the data scale of the first harmful corpus seed sample; The first harmful corpus seed sample is used as a positive sample, and the first safe corpus seed sample is used as a negative sample, and is input into the semantic feature classification model to be trained for classification training to obtain the preset semantic feature classification model.

4. The harmful corpus mining method according to claim 1, characterized in that: S3, performing semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set by using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus, including: Inputting each potentially harmful corpus sample in the potentially harmful corpus sample set into the preset semantic feature classification model for classification processing to obtain a classification score value for each potentially harmful corpus sample; The difficult sample set of harmful corpus is determined based on the classification score value of each of the potentially harmful corpus samples.

5. The harmful corpus mining method according to claim 1, characterized in that: The step S4, performing corpus sampling processing on the harmful corpus difficult sample set by using a preset corpus sampling large model to obtain a first target harmful corpus sample set, includes: Based on the harmful corpus difficult sample set, determining the harmfulness type of each harmful corpus difficult sample in the harmful corpus difficult sample set; Based on the harmful type, determining the input prompt words corresponding to each of the harmful corpus difficult samples; Each of the harmful corpus difficult samples and the input prompt words corresponding to each of the harmful corpus difficult samples are input into the preset corpus sampling large model for corpus sampling processing to obtain the first target harmful corpus sample set.

6. The harmful corpus mining method according to any one of claims 1 to 5, characterized in that: After performing corpus sampling processing on the harmful corpus difficult sample set by using a preset corpus sampling large model in S4 to obtain a first target harmful corpus sample set, the method further includes: Comparing the data size of the first target harmful corpus sample set with a data size threshold; If the data size of the first target harmful corpus sample set is smaller than the data size threshold, fusing the first target harmful corpus sample set with the first harmful corpus seed sample to obtain a second harmful corpus seed sample; The steps S1 to S4 are iteratively updated based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set, wherein the data scale of the second target harmful corpus sample set is greater than the data scale threshold.

7. The harmful corpus mining method according to claim 6, characterized in that: The iterative updating of steps S1 to S4 based on the second harmful corpus seed sample to obtain a second target harmful corpus sample set includes: Fusing all first target harmful corpus sample sets and the first harmful corpus seed samples in historical iteration rounds to obtain updated harmful corpus seed samples; Collecting a second safe corpus seed sample according to the data scale of the updated harmful corpus seed sample; The updated harmful corpus seed sample is used as a positive sample, and the second safe corpus seed sample is used as a negative sample to be input into the preset semantic feature classification model for update training to obtain an updated semantic feature classification model.

8. A harmful corpus mining device, characterized in that: The harmful corpus mining device comprises: A first acquisition module is used to acquire a first harmful corpus seed sample and a corpus data set, where the corpus data set is a public data set of pre-trained corpus of a large language model; A first retrieval module is used to perform text similarity retrieval in the corpus data set based on the first harmful corpus seed sample to obtain a potentially harmful corpus sample set, wherein the potentially harmful corpus sample set includes each potentially harmful corpus sample; A first classification module is used to perform semantic feature classification processing on each potentially harmful corpus sample in the potentially harmful corpus sample set by using a preset semantic feature classification model to obtain a difficult sample set of harmful corpus; The first sampling module is used to perform corpus sampling processing on the harmful corpus difficult sample set through a preset corpus sampling large model to obtain a first target harmful corpus sample set.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the harmful corpus mining method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the harmful corpus mining method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Extended corpus generation method and device in target field and electronic equipment

    CN112541076A

  • Corpus processing method, related device and equipment

    CN113821593A

  • Dialogue generation method and device based on knowledge base, electronic equipment and storage medium

    CN117093698A

  • Method and apparatus for entering information, electronic device, computer readable storage medium

    US20220156611A1