Illegal information identification method, device and equipment, storage medium and computer program product
By recalling relevant cases from the case library and generating prompt words, and using big models to identify violation information, the problems of low manual monitoring accuracy and inconsistent penalty standards in the existing technology are solved, and higher recognition accuracy and punishment uniformity are achieved.
Patent Information
- Application Number
- CN202510168873.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-23
AI Technical Summary
The existing methods of identifying violation information through manual monitoring, with low overall accuracy and inconsistent penalty standards.
By recalling relevant cases from the case library based on the information text to be identified, prompt words for the big model are generated, and the violation identification results of the information to be identified are generated through the big model.
It improves the accuracy of identification of violation information and improves the unity of the judgment criteria of large models in similar cases.
Smart Images

Figure CN120030148A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, storage medium and computer program product for identifying illegal information. Background Art
[0002] At present, with the rapid development of information technology, the information available on the Internet is growing exponentially, and the identification of illegal information has become a difficult problem. The existing illegal information identification method usually relies on manual monitoring of illegal information, which has the defects of low overall accuracy and inconsistent penalty standards. Summary of the invention
[0003] The main purpose of this application is to provide a method, device, equipment, storage medium and computer program product for identifying illegal information, aiming to solve the technical problems that the related illegal information identification methods rely on manual monitoring of illegal information, the overall accuracy is not high, and the penalty standards are not unified.
[0004] To achieve the above purpose, the present application provides a method for identifying illegal information, which includes:
[0005] Recall relevant cases from the case library based on the information text of the information to be identified;
[0006] Generate prompt words of a large model according to the information text and the related cases;
[0007] The prompt word is input into the large model, and the violation identification result of the information to be identified is generated by the large model.
[0008] Optionally, the recalling of relevant cases from the case library based on the information text of the information to be identified includes:
[0009] Based on the information text of the information to be identified, case recall is performed in the case library to obtain candidate cases;
[0010] The candidate cases are sorted, and relevant cases are selected from the candidate cases according to the sorting result.
[0011] Optionally, the recalling of cases in a case library based on the information text of the information to be identified to obtain candidate cases includes:
[0012] Generate a text vector corresponding to the information text of the information to be identified through the embedding model;
[0013] Calculate the similarity between the text vector and the vectors of each case in the case library;
[0014] Vectors are recalled in the case library according to the vector similarity to obtain candidate cases.
[0015] Optionally, the recalling of cases in a case library based on the information text of the information to be identified to obtain candidate cases includes:
[0016] Extracting text keywords from the information text of the information to be identified;
[0017] Based on the text keywords, case recall is performed in the case library to obtain candidate cases.
[0018] Optionally, the sorting of the candidate cases and selecting relevant cases from the candidate cases according to the sorting result includes:
[0019] Calculating the similarity between the candidate case and the information text;
[0020] The candidate cases are sorted according to the similarities, and relevant cases are selected from the candidate cases according to the sorting result.
[0021] Optionally, before recalling relevant cases from the case library based on the information text of the information to be identified, the method further includes:
[0022] Obtaining the original text of the information to be identified;
[0023] The original text is optimized by a text optimization model to obtain the information text of the information to be identified.
[0024] Optionally, before recalling relevant cases from the case library based on the information text of the information to be identified, the method further includes:
[0025] Obtain a historical penalty dataset, and construct a supervised fine-tuning dataset based on the historical penalty dataset;
[0026] The preset base model is fine-tuned based on the supervised fine-tuning dataset to obtain a large model.
[0027] Optionally, the acquiring of a historical penalty dataset and constructing a supervised fine-tuning dataset based on the historical penalty dataset includes:
[0028] Acquire a historical penalty data set, and clean the historical penalty data set to obtain a cleaned data set;
[0029] The cleaned data set is disambiguated to obtain a disambiguated data set, and a supervised fine-tuning data set is constructed based on the disambiguated data set.
[0030] Optionally, the acquiring of the historical penalty data set and cleaning the historical penalty data set to obtain the cleaned data set includes:
[0031] Acquire a historical penalty data set, and parse the historical penalty data set to obtain at least one of the text meaning, text length, and penalty law provision of each historical penalty data in the historical penalty data set;
[0032] The historical penalty data set is cleaned according to at least one of the text meaning, the text length, and the penalty law provision to obtain a cleaned data set.
[0033] Optionally, disambiguating the cleaned data set to obtain a disambiguated data set, and constructing a supervised fine-tuning data set based on the disambiguated data set includes:
[0034] Calculating similarity indices between data in the cleaned data set, and dividing the cleaned data set into a plurality of similar text clusters according to the similarity indices;
[0035] Selecting target samples from each similar text cluster, wherein the target samples are samples with inconsistent regulations in the similar text cluster;
[0036] The target sample is disambiguated and verified to obtain a disambiguated dataset, and a supervised fine-tuning dataset is constructed based on the disambiguated dataset.
[0037] Optionally, the step of inputting the prompt word into the large model and generating a violation identification result of the information to be identified by using the large model includes:
[0038] Input the prompt word into the large model to obtain the penalty result and reasoning reason output by the large model;
[0039] The penalty result and the reasoning reason are analyzed to obtain a violation identification result of the information to be identified.
[0040] In addition, to achieve the above purpose, the present application also proposes a violation information identification device, the violation information identification device comprising:
[0041] A case recall module, used to recall relevant cases from the case library based on the information text of the information to be identified;
[0042] A prompt word generation module, used for generating prompt words of a large model according to the information text and the related cases;
[0043] The information identification module is used to input the prompt word into the large model and generate a violation identification result of the information to be identified through the large model.
[0044] Optionally, the case recall module is further used to recall cases in the case library based on the information text of the information to be identified to obtain candidate cases; sort the candidate cases, and select relevant cases from the candidate cases according to the sorting results.
[0045] Optionally, the case recall module is also used to generate a text vector corresponding to the information text of the information to be identified through an embedding model; calculate the vector similarity between the text vector and each case in the case library; and perform vector recall in the case library based on the vector similarity to obtain candidate cases.
[0046] Optionally, the case recall module is further used to extract text keywords from the information text of the information to be identified; and to perform case recall in the case library based on the text keywords to obtain candidate cases.
[0047] Optionally, the case recall module is further used to calculate the similarity between the candidate cases and the information text; sort the candidate cases according to the similarity, and select relevant cases from the candidate cases according to the sorting result.
[0048] Optionally, the violation information identification device further includes:
[0049] The text optimization module is used to obtain the original text of the information to be identified; optimize the original text through the text optimization model to obtain the information text of the information to be identified.
[0050] In addition, to achieve the above-mentioned purpose, the present application also proposes a violation information identification device, which includes a memory, a processor, and a violation information identification program stored in the memory and executable on the processor, and the violation information identification program is configured to implement the violation information identification method described above.
[0051] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, on which a violation information identification program is stored, and when the violation information identification program is executed by a processor, the violation information identification method described above is implemented.
[0052] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a violation information identification program, and when the violation information identification program is executed by a processor, it implements the violation information identification method described above.
[0053] One or more technical solutions proposed in this application have at least the following technical effects:
[0054] In the present application, it is disclosed that relevant cases are recalled from a case library based on the information text of the information to be identified, prompt words of a large model are generated according to the information text and the relevant cases, the prompt words are input into the large model, and violation identification results of the information to be identified are generated through the large model; since the present application combines the information text of the information to be identified and the relevant cases to generate the prompt words of the large model, and identifies the violation information through the large model based on the prompt words, it is possible to improve the accuracy of identifying the violation information and enhance the uniformity of the penalty standards of the large model on similar cases. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0057] Figure 1 This is a flow chart of the first embodiment of the method for identifying illegal information of this application;
[0058] Figure 2 This is a flow chart of the second embodiment of the method for identifying illegal information of this application;
[0059] Figure 3 This is a flow chart of the third embodiment of the method for identifying illegal information of this application;
[0060] Figure 4 This is a specific flow chart of an embodiment of the method for identifying illegal information of this application;
[0061] Figure 5 This is a schematic diagram of the module structure of the illegal information identification device according to an embodiment of the present application;
[0062] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the method for identifying illegal information in an embodiment of the present application.
[0063] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0064] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0065] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0066] At present, with the rapid development of information technology, the information available on the Internet is growing exponentially, and the identification of illegal information has become a difficult problem. The existing illegal information identification method usually relies on manual monitoring of illegal information, which has the defects of low overall accuracy and inconsistent penalty standards.
[0067] Therefore, in order to overcome the above-mentioned defects, the present application provides a solution, which includes: recalling relevant cases from the case library based on the information text of the information to be identified, generating prompt words of the large model according to the information text and the relevant cases, inputting the prompt words into the large model, and generating violation identification results of the information to be identified through the large model; since the present application combines the information text of the information to be identified and the relevant cases to generate the prompt words of the large model, and identifies the violation information through the large model based on the prompt words, it can improve the accuracy of the identification of the violation information and enhance the uniformity of the penalty standards of the large model on similar cases.
[0068] It should be noted that the executor of this embodiment can be a violation information identification device with data processing, network communication and program running functions, such as a server, etc., or other electronic devices that can achieve the same or similar functions, and this embodiment does not impose any restrictions on this.
[0069] Based on this, the present application embodiment provides a method for identifying illegal information. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the method for identifying illegal information of this application.
[0070] In a first embodiment, the violation information identification method includes:
[0071] Step S10: Recall relevant cases from the case library based on the information text of the information to be identified.
[0072] It should be understood that in order to improve the uniformity of the penalty standards of the big model in similar cases, in this embodiment, the relevant cases are first recalled from the case library based on the information text of the information to be identified, and then the prompt words of the big model are generated in combination with the information text of the information to be identified and the relevant cases, and the illegal information is identified through the big model based on the prompt words. Among them, the information to be identified may refer to text information that needs to be identified or classified as illegal, such as advertising copy, user comments, etc. The information text may refer to the specific text form of the information to be identified, which is the basis for subsequent processing and analysis. The case library may refer to a database that stores a large number of historical violation cases or related texts, which is used to provide comparison and reference for the information to be identified. Related cases may refer to historical cases related to the information to be identified that are screened out from the case library according to certain rules (such as similarity, keyword matching, etc.).
[0073] In a specific implementation, first, the information text of the information to be identified is obtained through user input, file reading, etc., and then the index or search algorithm in the case library is used to recall the historical cases related to the information text content (such as keywords, topics, etc.) of the information to be identified. The recall method includes but is not limited to keyword matching, vector similarity calculation, etc., which is not limited in this embodiment.
[0074] Furthermore, in order to solve the problems of improper segmentation, improper use of punctuation and semantic ambiguity in the original text and improve the accuracy of case recall and violation identification, before step S10, it also includes: obtaining the original text of the information to be identified; optimizing the original text through a text optimization model to obtain the information text of the information to be identified.
[0075] It is understandable that the original text of the information to be identified may refer to the text information to be directly obtained without any processing or optimization, such as the advertisement copy and comment content submitted by the user. The text optimization model may be a natural language processing model, which is used to optimize the original text, such as improper segmentation, improper punctuation, and semantic ambiguity, so that the subsequent steps can better understand and process the text information. The information text may refer to the more standardized and easy-to-understand text information obtained after being processed by the text optimization model, as the input of the subsequent steps.
[0076] In the specific implementation, the original text is input into the text optimization model. After the text optimization model performs grammar checking and semantic understanding on the text, it outputs more standardized and easy-to-understand text information.
[0077] Step S20: Generate prompt words for a large model according to the information text and the related cases.
[0078] It is understandable that a large model can refer to a natural language processing model that has been trained on a large scale and has strong text understanding and generation capabilities, and is used for tasks such as violation identification. Prompts can refer to text snippets used to guide a large model to perform a specific task or generate a specific output, and can include task descriptions, input examples, etc.
[0079] In the specific implementation, the information text and related cases are filled into the preset prompt word template to generate a complete prompt word. The preset prompt word template can be a task description, input example, etc., which is used to guide the large model to perform violation identification tasks.
[0080] Step S30: input the prompt word into the large model, and generate a violation identification result of the information to be identified through the large model.
[0081] It should be understood that the violation identification result may refer to the conclusion generated by the large model based on the prompt word and the information text as to whether the information to be identified violates the rules and the type of violation. In a specific implementation, the generated prompt word is passed as input to the large model for reasoning. The large model generates the violation identification result about the information to be identified based on the content of the prompt word, combined with its own knowledge base and reasoning ability.
[0082] Furthermore, in order to improve the reliability of the violation identification results, in this embodiment, when the large model outputs the penalty result, it will also output the reasoning reasons at the same time. The step S30 includes: inputting the prompt word into the large model to obtain the penalty result and reasoning reasons output by the large model; parsing the penalty result and the reasoning reasons to obtain the violation identification result of the information to be identified. Among them, the penalty result may refer to the violation identification conclusion made by the large model based on the input information (including information text and related cases), which may include information such as whether there is a violation and the type of violation. The reasoning reason may refer to the logical or factual basis provided by the large model to support its conclusion when giving the penalty result. The violation identification result may refer to the final conclusion on whether the information to be identified is in violation and the specific circumstances of the violation after parsing the penalty result and the reasoning reasons.
[0083] In the specific implementation, the generated prompt words are passed as input to the large model for reasoning. The large model generates the penalty results and reasoning reasons for the information to be identified based on the content of the prompt words, combined with its own knowledge base and reasoning ability. The reasoning reasons and penalty results output by the large model are parsed, and according to the preset strategies or rules, the reasoning reasons and penalty results are converted into the final violation identification results. The parsing process can include judging the rationality of the reasoning reasons, verifying the accuracy of the penalty results, etc.
[0084] This embodiment combines the information text of the information to be identified and related cases to generate prompt words for the big model, and identifies illegal information through the big model based on the prompt words, thereby improving the accuracy of illegal information identification and enhancing the uniformity of the penalty standards of the big model in similar cases.
[0085] Reference Figure 2 , Figure 2 This is a flow chart of the second embodiment of the method for identifying illegal information in this application. Figure 1 The first embodiment shown proposes a second embodiment of the method for identifying illegal information of the present application.
[0086] In the second embodiment, the step S20 includes:
[0087] Step S201: recall cases in the case library based on the information text of the information to be identified to obtain candidate cases.
[0088] It should be understood that in order to ensure that the cases most relevant or similar to the information to be identified are used in the subsequent steps, in this embodiment, the case recall is first performed in the case library based on the information text of the information to be identified to obtain candidate cases, and then the candidate cases are sorted, and relevant cases are selected from the candidate cases according to the sorting results. Among them, the candidate cases can refer to historical cases that are relevant or similar to the information text of the information to be identified and are screened from the case library according to certain rules, and the candidate cases can be used as candidates for subsequent sorting and selection.
[0089] In a specific implementation, the index or search algorithm in the case library is used to recall historical cases related to or similar to the information text content (such as keywords, themes, etc.) of the information to be identified. The recall method may include vector recall (calculating similarity based on vectors generated by the embedding model) and keyword recall (extracting text keywords based on strategies for matching), etc., which are not limited in this embodiment.
[0090] Further, the step S201 includes: generating a text vector corresponding to the information text of the information to be identified through an embedding model; calculating the vector similarity between the text vector and each case in the case library; and performing vector recall in the case library according to the vector similarity to obtain candidate cases. Among them, the embedding model may refer to a machine learning model that can convert text data into a vector representation in a high-dimensional space. These vectors can capture the semantic information in the text, so that similar texts are closer in the vector space. The text vector may refer to a vector generated by the embedding model that represents the position or feature of the text data in the vector space. Vector similarity may be a measure of the distance or similarity between two vectors in the vector space. The vector similarity calculation method may include cosine similarity, Euclidean distance, etc., which is not limited in this embodiment. Vector recall may refer to the process of retrieving the case that is most similar or related to the information text of the information to be identified from the case library based on vector similarity.
[0091] In the specific implementation, the information text of the information to be identified is converted into a vector representation using a pre-trained embedding model. This vector can capture the semantic information in the text and provide a basis for subsequent vector recall. The vector similarity calculation method (such as cosine similarity) is used to calculate the similarity between the vector of the information text of the information to be identified and the vector of each case in the case library. According to the vector similarity list, the top N cases with higher similarity are selected as candidate cases. Among them, N can be set in advance.
[0092] Furthermore, the step S201 includes: extracting text keywords from the information text of the information to be identified; and recalling cases in the case library based on the text keywords to obtain candidate cases. The text keywords may refer to words or phrases extracted from the information text of the information to be identified that can summarize or represent the text theme, content or characteristics.
[0093] In the specific implementation, natural language processing (NLP) technology, such as word frequency statistics and word embedding, is used to extract representative or important words or phrases from the information text to be identified as text keywords. The extracted text keywords are used as query conditions to search in the case library to find cases matching these keywords as candidate cases.
[0094] Of course, in order to increase the number of candidate cases, in this embodiment, vector recall and keyword recall can also be performed simultaneously. In a specific implementation, the candidate cases include a first candidate case and a second candidate case. Generate a text vector corresponding to the information text of the information to be identified through an embedding model; calculate the vector similarity between the text vector and each case in the case library; perform vector recall in the case library based on the vector similarity to obtain the first candidate case. Extract text keywords from the information text of the information to be identified; perform case recall in the case library based on the text keywords to obtain the second candidate case.
[0095] Step S202: sorting the candidate cases, and selecting relevant cases from the candidate cases according to the sorting result.
[0096] It is understood that sorting may refer to the process of sorting candidate cases by relevance, for example, based on the similarity or relevance between the candidate cases and the information text of the information to be identified. A related case may refer to a case selected from the sorted candidate cases that is most relevant or similar to the information text of the information to be identified, and is used to generate the prompt words of the large model, and a related case is at least one case.
[0097] In a specific implementation, sorting the candidate cases and selecting related cases from the candidate cases according to the sorting results may be calculating the similarity between the candidate cases and the information text; sorting the candidate cases according to the similarity, and selecting related cases from the candidate cases according to the sorting results.
[0098] This embodiment first recalls cases in the case library based on the information text of the information to be identified, obtains candidate cases, then sorts the candidate cases, and selects relevant cases from the candidate cases according to the sorting results, thereby ensuring that the cases most relevant or similar to the information to be identified are used in subsequent steps, thereby further improving the accuracy and efficiency of identifying illegal information.
[0099] Reference Figure 3 , Figure 3 This is a flow chart of the third embodiment of the method for identifying illegal information of the present application. Based on the above embodiments, the third embodiment of the method for identifying illegal information of the present application is proposed.
[0100] In the third embodiment, before step S10, the method further includes:
[0101] Step S01: Obtain a historical penalty dataset, and construct a supervised fine-tuning dataset based on the historical penalty dataset.
[0102] It should be understood that in order to enable the large model to better identify violation information, in this embodiment, a supervised fine-tuning dataset is constructed based on the historical penalty dataset, and the preset base model is fine-tuned according to the supervised fine-tuning dataset to obtain the large model. Among them, the historical penalty dataset may refer to a dataset formed by the review experts' judgments on advertisements, texts and other content in the past period of time, including information such as information text, penalty results and possible reasons for penalties. The supervised fine-tuning dataset (Supervised Fine-Tuning, SFT) may refer to a dataset for supervising fine-tuning the preset base model formed based on the historical penalty dataset after pre-processing steps such as cleaning and disambiguation.
[0103] In the specific implementation, the historical penalty data set is collected from the historical penalty data source, including information text, penalty results, etc. Then the historical penalty data set is preprocessed to remove invalid, repeated or contradictory data to ensure the quality and consistency of the data. Finally, a supervised fine-tuning data set (SFT data set) is constructed based on the processed data.
[0104] Furthermore, in order to improve the quality of the supervised fine-tuning dataset, the step S01 includes: obtaining a historical penalty dataset, and cleaning the historical penalty dataset to obtain a cleaned dataset; disambiguating the cleaned dataset to obtain a disambiguated dataset, and constructing a supervised fine-tuning dataset based on the disambiguated dataset. Among them, the cleaned dataset may refer to a dataset obtained by preprocessing the historical penalty dataset and removing invalid, redundant or erroneous data. The cleaning process may include removing meaningless text, too short text, text that is missing key information (such as penalty laws), etc. The disambiguated dataset may refer to processing ambiguous data in the cleaned dataset to ensure that texts with the same meaning have consistent penalty labels or law labels, thereby eliminating ambiguity. The disambiguation process may include text semantic analysis, manual review and other means.
[0105] In the specific implementation, historical penalty datasets are collected from historical data sources, and the historical penalty datasets are preprocessed to remove invalid, redundant or erroneous data. The cleaning process may involve steps such as data screening, deduplication, and missing value processing. The ambiguous data in the cleaned dataset is processed to ensure that texts with the same meaning have consistent penalty labels. Then, a supervised fine-tuning dataset is constructed based on the disambiguated dataset to fine-tune the preset base model.
[0106] Furthermore, in order to improve the data set cleaning effect, the historical penalty data set is obtained, and the historical penalty data set is cleaned to obtain the cleaned data set, including: obtaining the historical penalty data set, and parsing the historical penalty data set to obtain the text meaning, text length and at least one of the penalty laws of each historical penalty data in the historical penalty data set; cleaning the historical penalty data set according to the text meaning, the text length and at least one of the penalty laws to obtain the cleaned data set.
[0107] In the specific implementation, we collect penalty cases from historical data sources, parse each case, and extract key information such as text meaning, text length, and penalty law. Based on the clarity of text meaning, the rationality of text length, and the completeness of penalty law, we screen and clean the historical penalty data set to remove data that does not meet the requirements.
[0108] Further, in order to improve the data set disambiguation effect, the cleaned data set is disambiguated to obtain the disambiguated data set, and a supervised fine-tuning data set is constructed based on the disambiguated data set, including: calculating the similarity index between each data in the cleaned data set, and dividing the cleaned data set into multiple similar text clusters according to the similarity index; selecting a target sample from each similar text cluster, wherein the target sample is a sample with inconsistent regulations in the similar text cluster; performing disambiguation verification on the target sample to obtain the disambiguated data set, and constructing a supervised fine-tuning data set based on the disambiguated data set. Wherein, similarity index: a quantitative index used to measure the similarity between two data (such as advertising penalty cases), such as cosine similarity, Jaccard similarity, Rouge-L index, etc. Similar text clusters can refer to multiple subsets into which the data in the cleaned data set is divided according to similarity indexes, and the data in each subset has high similarity. Target samples can refer to samples in similar text clusters that are selected for further verification due to inconsistent application of regulations.
[0109] In the specific implementation, a text similarity algorithm is used to calculate the similarity index between each case in the cleaned data set, and then the cases are divided into multiple similar text clusters based on the similarity index. In the similar text clusters, samples with inconsistent application of regulations are selected as target samples. These samples may have ambiguity or differences in understanding of regulations. Manual or automated disambiguation is performed on the target samples to ensure consistency in the application of regulations for the same or similar texts. Then, a supervised fine-tuning dataset is constructed based on the disambiguated dataset.
[0110] Step S02: fine-tune the preset base model based on the supervised fine-tuning dataset to obtain a large model.
[0111] It is understood that the preset base model can refer to a deep learning model that has certain language understanding and generation capabilities before fine-tuning. In the specific implementation, the preset base model is trained using a supervised fine-tuning dataset to adjust the model's parameters to make it more consistent with the standards of the audit experts. During the training process, the performance of the model is optimized by minimizing the loss function.
[0112] This embodiment constructs a supervised fine-tuning dataset based on the historical penalty dataset, and fine-tunes the preset base model according to the supervised fine-tuning dataset to obtain a large model, so that the large model can better identify violation information. The large model will maintain a high degree of consistency in the process of illegal identification, and effectively avoid misjudgment due to inconsistent standards.
[0113] For ease of understanding, refer to Figure 4 This invention is provided for illustration, but is not intended to limit the present application. Figure 4 This is a specific flow chart of an embodiment of the method for identifying illegal information of this application. Figure 4 In the present invention, the illegal information identification method mainly includes the following steps: text optimization, case recall, case sorting and large model recognition. The text optimization process is used to solve the problems of improper punctuation, improper punctuation and semantic ambiguity in the information text. The information text optimized by the text optimization model has a good improvement in the model recall case and large model penalty recognition. The case recall process is used to recall relevant cases in the case library. In this solution, case recall includes vector recall and keyword recall. Vector recall is the vector generated by the embedding model, and the candidate cases are recalled by calculating the similarity. Keyword recall is to extract keywords from the information text by strategy and recall the candidate cases by keywords. The case sorting process is the process of sorting the candidate cases by relevance. After the candidate cases are sorted, the top 3 candidate cases are taken out according to the sorting results, and the cases are spliced into the prompt word (prompt). The large model (fine-tuned large model) is improved in the way of few-shot learning (Few-Shot Learning) to improve the accuracy of illegal identification and the unification of illegal standards.
[0114] For ease of understanding, the following examples are given, but are not intended to limit the present application. As an example, the method for identifying illegal information mainly includes the following steps:
[0115] 1. Advertisement text A will be optimized through the text optimization model, and the optimized text B will be output.
[0116] 2. Text B is recalled in two ways to recall candidate cases (number N).
[0117] 3. The candidate cases (number N) and text B are input into the sorting model to obtain sorted candidate cases (the higher the ranking, the higher the relevance).
[0118] 4. Take the TOP3 related cases and text B, assemble the prompts in a fixed format, and input them into the fine-tuned large model. The large model outputs the reasoning reasons and the penalty results.
[0119] 5. Analyze the reasoning and penalty results of the large model according to the strategy to obtain the final violation identification results in a regular format.
[0120] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the method for identifying illegal information in the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0121] This application also provides a violation information identification device, please refer to Figure 5 , the illegal information identification device comprises:
[0122] A case recall module 10, for recalling relevant cases from a case library based on the information text of the information to be identified;
[0123] A prompt word generation module 20, used to generate prompt words of a large model according to the information text and the related cases;
[0124] The information identification module 30 is used to input the prompt word into the large model and generate a violation identification result of the information to be identified through the large model.
[0125] The violation information identification device provided by the present application adopts the violation information identification method in the above embodiment, which can solve the technical problems that the overall accuracy of the related violation information identification method is low through manual monitoring of violation information and the penalty standards are not uniform. Compared with the prior art, the beneficial effects of the violation information identification device provided by the present application are the same as the beneficial effects of the violation information identification method provided by the above embodiment, and the other technical features of the violation information identification device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0126] The present application provides a violation information identification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the violation information identification method in the above-mentioned embodiment one.
[0127] Reference below Figure 6, which shows a schematic diagram of the structure of a violation information identification device suitable for implementing the embodiment of the present application. The violation information identification device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The violation information identification device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0128] like Figure 6 As shown, the illegal information identification device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to RAM (Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the illegal information identification device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the violation information identification device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a violation information identification device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0129] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0130] The violation information identification device provided by the present application adopts the violation information identification method in the above embodiment, which can solve the technical problems that the overall accuracy rate of the related violation information identification method is low through manual monitoring of violation information and the penalty standards are not uniform. Compared with the prior art, the beneficial effects of the violation information identification device provided by the present application are the same as the beneficial effects of the violation information identification method provided by the above embodiment, and the other technical features in the violation information identification device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0131] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0132] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0133] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the violation information identification method in the above-mentioned embodiment.
[0134] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory) or flash memory, optical fiber, CD-ROM (CD-Read Only Memory, portable compact disk read-only memory), optical storage device, magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0135] The above-mentioned computer-readable storage medium may be included in the violation information identification device; or it may exist independently without being assembled into the violation information identification device.
[0136] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the violation information identification device, the violation information identification device executes the violation information identification method.
[0137] The computer program code for performing the operation of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0138] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0139] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0140] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned violation information identification method, which can solve the technical problems that the overall accuracy of the related violation information identification method is low through manual monitoring of violation information and the penalty standards are not uniform. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the violation information identification method provided by the above-mentioned embodiment, and will not be repeated here.
[0141] The present application also provides a computer program product, including a computer program, which implements the above-mentioned violation information identification method when executed by a processor.
[0142] The computer program product provided by this application can solve the technical problems that the overall accuracy of the related violation information identification method is low through manual monitoring of violation information, and the penalty standards are not uniform. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the violation information identification method provided by the above embodiment, and will not be repeated here.
[0143] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
[0144] The present application discloses A1, a method for identifying illegal information, the illegal information identification method comprising:
[0145] Recall relevant cases from the case library based on the information text of the information to be identified;
[0146] Generate prompt words of a large model according to the information text and the related cases;
[0147] The prompt word is input into the large model, and the violation identification result of the information to be identified is generated by the large model.
[0148] A2. The method for identifying illegal information as described in A1, wherein the relevant cases are recalled from the case library based on the information text of the information to be identified, comprising:
[0149] Based on the information text of the information to be identified, case recall is performed in the case library to obtain candidate cases;
[0150] The candidate cases are sorted, and relevant cases are selected from the candidate cases according to the sorting result.
[0151] A3. The method for identifying illegal information as described in A2, wherein the case recall in the case library based on the information text of the information to be identified to obtain candidate cases includes:
[0152] Generate a text vector corresponding to the information text of the information to be identified through the embedding model;
[0153] Calculate the similarity between the text vector and the vectors of each case in the case library;
[0154] Vectors are recalled in the case library according to the vector similarity to obtain candidate cases.
[0155] A4. The method for identifying illegal information as described in A2, wherein the case is recalled from the case library based on the information text of the information to be identified to obtain candidate cases, including:
[0156] Extracting text keywords from the information text of the information to be identified;
[0157] Based on the text keywords, case recall is performed in the case library to obtain candidate cases.
[0158] A5. The method for identifying illegal information as described in A2, wherein the candidate cases are sorted and relevant cases are selected from the candidate cases according to the sorting result, including:
[0159] Calculating the similarity between the candidate case and the information text;
[0160] The candidate cases are sorted according to the similarities, and relevant cases are selected from the candidate cases according to the sorting result.
[0161] A6. The method for identifying illegal information as described in any one of A1 to A5, before recalling relevant cases from the case library based on the information text of the information to be identified, further comprising:
[0162] Obtaining the original text of the information to be identified;
[0163] The original text is optimized by a text optimization model to obtain the information text of the information to be identified.
[0164] A7. The method for identifying illegal information as described in any one of A1 to A5, before recalling relevant cases from the case library based on the information text of the information to be identified, further comprising:
[0165] Obtain a historical penalty dataset, and construct a supervised fine-tuning dataset based on the historical penalty dataset;
[0166] The preset base model is fine-tuned based on the supervised fine-tuning dataset to obtain a large model.
[0167] A8. The method for identifying violation information as described in A7, wherein the step of obtaining a historical penalty dataset and constructing a supervised fine-tuning dataset based on the historical penalty dataset includes:
[0168] Acquire a historical penalty data set, and clean the historical penalty data set to obtain a cleaned data set;
[0169] The cleaned data set is disambiguated to obtain a disambiguated data set, and a supervised fine-tuning data set is constructed based on the disambiguated data set.
[0170] A9. The method for identifying violation information as described in A8, wherein the step of obtaining a historical penalty data set and cleaning the historical penalty data set to obtain a cleaned data set includes:
[0171] Acquire a historical penalty data set, and parse the historical penalty data set to obtain at least one of the text meaning, text length, and penalty law provision of each historical penalty data in the historical penalty data set;
[0172] The historical penalty data set is cleaned according to at least one of the text meaning, the text length, and the penalty law provision to obtain a cleaned data set.
[0173] A10. The method for identifying illegal information as described in A8, wherein disambiguating the cleaned data set to obtain a disambiguated data set, and constructing a supervised fine-tuning data set based on the disambiguated data set, comprises:
[0174] Calculating similarity indices between data in the cleaned data set, and dividing the cleaned data set into a plurality of similar text clusters according to the similarity indices;
[0175] Selecting target samples from each similar text cluster, wherein the target samples are samples with inconsistent regulations in the similar text cluster;
[0176] The target sample is disambiguated and verified to obtain a disambiguated dataset, and a supervised fine-tuning dataset is constructed based on the disambiguated dataset.
[0177] A11. The method for identifying illegal information as described in any one of A1 to A5, wherein the inputting of the prompt word into the large model and the generation of the illegal identification result of the information to be identified by the large model comprises:
[0178] Input the prompt word into the large model to obtain the penalty result and reasoning reason output by the large model;
[0179] The penalty result and the reasoning reason are analyzed to obtain a violation identification result of the information to be identified.
[0180] The present application also discloses B12, a device for identifying illegal information, the device comprising:
[0181] A case recall module is used to recall relevant cases from the case library based on the information text of the information to be identified;
[0182] A prompt word generation module, used for generating prompt words of a large model according to the information text and the related cases;
[0183] The information identification module is used to input the prompt word into the large model and generate a violation identification result of the information to be identified through the large model.
[0184] B13. In the violation information identification device as described in B12, the case recall module is further used to recall cases in the case library based on the information text of the information to be identified to obtain candidate cases; sort the candidate cases, and select relevant cases from the candidate cases according to the sorting results.
[0185] B14. In the violation information identification device as described in B13, the case recall module is also used to generate a text vector corresponding to the information text of the information to be identified through an embedding model; calculate the vector similarity between the text vector and each case in the case library; and perform vector recall in the case library based on the vector similarity to obtain candidate cases.
[0186] B15. In the device for identifying illegal information as described in B13, the case recall module is further used to extract text keywords from the information text of the information to be identified; and to recall cases in the case library based on the text keywords to obtain candidate cases.
[0187] B16. In the violation information identification device as described in B13, the case recall module is further used to calculate the similarity between the candidate cases and the information text; sort the candidate cases according to the similarity, and select relevant cases from the candidate cases according to the sorting result.
[0188] B17. The device for identifying illegal information as described in any one of B12 to B16, wherein the device for identifying illegal information further comprises:
[0189] The text optimization module is used to obtain the original text of the information to be identified; optimize the original text through the text optimization model to obtain the information text of the information to be identified.
[0190] The present application also discloses C18, a violation information identification device, which includes: a memory, a processor, and a violation information identification program stored in the memory and executable on the processor, and when the violation information identification program is executed by the processor, the violation information identification method described above is implemented.
[0191] The present application also discloses D19, a storage medium, on which a violation information identification program is stored, and when the violation information identification program is executed by a processor, the violation information identification method as described above is implemented.
[0192] The present application also discloses E20, a computer program product, which includes a violation information identification program, and when the violation information identification program is executed by a processor, the violation information identification method as described above is implemented.
Claims
1. A method for identifying illegal information, characterized in that: The method for identifying illegal information includes: Recall relevant cases from the case library based on the information text of the information to be identified; Generate prompt words of a large model according to the information text and the related cases; The prompt word is input into the large model, and the violation identification result of the information to be identified is generated by the large model.
2. The method for identifying illegal information according to claim 1, characterized in that: The process of recalling relevant cases from the case library based on the information text of the information to be identified includes: Based on the information text of the information to be identified, case recall is performed in the case library to obtain candidate cases; The candidate cases are sorted, and relevant cases are selected from the candidate cases according to the sorting result.
3. The method for identifying illegal information according to claim 2, characterized in that: The process of recalling cases in the case library based on the information text of the information to be identified to obtain candidate cases includes: Generate a text vector corresponding to the information text of the information to be identified through the embedding model; Calculate the similarity between the text vector and the vectors of each case in the case library; Vectors are recalled in the case library according to the vector similarity to obtain candidate cases.
4. The method for identifying illegal information according to claim 2, characterized in that: The process of recalling cases in the case library based on the information text of the information to be identified to obtain candidate cases includes: Extracting text keywords from the information text of the information to be identified; Based on the text keywords, case recall is performed in the case library to obtain candidate cases.
5. The method for identifying illegal information according to claim 2, characterized in that: The step of sorting the candidate cases and selecting relevant cases from the candidate cases according to the sorting result includes: Calculating the similarity between the candidate case and the information text; The candidate cases are sorted according to the similarities, and relevant cases are selected from the candidate cases according to the sorting result.
6. The method for identifying illegal information according to any one of claims 1 to 5, characterized in that: Before recalling relevant cases from the case library based on the information text of the information to be identified, the method further includes: Obtaining the original text of the information to be identified; The original text is optimized by a text optimization model to obtain the information text of the information to be identified.
7. A device for identifying illegal information, characterized in that: The violation information identification device comprises: A case recall module is used to recall relevant cases from the case library based on the information text of the information to be identified; A prompt word generation module, used for generating prompt words of a large model according to the information text and the related cases; The information identification module is used to input the prompt word into the large model and generate a violation identification result of the information to be identified through the large model.
8. A device for identifying illegal information, characterized in that: The violation information identification device includes: a memory, a processor, and a violation information identification program stored in the memory and executable on the processor. When the violation information identification program is executed by the processor, the violation information identification method according to any one of claims 1 to 6 is implemented.
9. A storage medium, characterized in that: The storage medium stores a violation information identification program, and when the violation information identification program is executed by the processor, the violation information identification method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a violation information identification program, and when the violation information identification program is executed by a processor, the violation information identification method according to any one of claims 1 to 6 is implemented.