Public affair-oriented negative information discrimination method and system
Through the large-scale negative information scoring and screening mechanism and the fine-tuned Qwen2-1.5B model, combined with the specific task data set and keyword scoring mechanism, the problem of low accuracy of negative information discrimination for complex public affairs in the existing technology is solved, and efficient and accurate negative information discrimination is achieved.
Patent Information
- Application Number
- CN202411978982.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
When facing negative information on complex public affairs, the prior art has low accuracy and cannot effectively deal with complex contexts, resulting in misjudgment and misreport.
A large-scale negative information scoring and screening mechanism is adopted, combined with the fine-tuned Qwen2-1.5B model, by building a specific task data set and a keyword scoring mechanism, we will improve the ability to automatically distinguish negative information for complex public affairs.
It significantly improves the ability to automatically discriminate negative information for complex public affairs, can effectively deal with massive information on social platforms, and improves the accuracy and efficiency of discrimination.
Smart Images

Figure CN119988622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for identifying negative information in public affairs. Background Art
[0002] In existing social media information processing systems, automatic identification of negative information usually relies on traditional text analysis models or rule-based methods. Although these methods can identify negative emotions in text to a certain extent, they show obvious shortcomings when facing negative information about public affairs. Existing negative information identification technical solutions can be roughly divided into the following categories:
[0003] 1. Keyword matching-based rule systems: This type of system relies on predefined keywords or phrases to detect negative information in text. Although this method is simple to implement and effective in certain clear contexts, it is highly susceptible to language variants and has difficulty accurately identifying implicit negative emotions, especially when dealing with complex language structures such as metaphors, sarcasm, or irony, with a high rate of misjudgment.
[0004] 2. Traditional machine learning models: Classification algorithms such as SVM (support vector machine) and naive Bayes rely on manually designed features to identify negative information in text. Although these models are effective in simplified sentiment analysis tasks, their performance is limited by the quality of feature extraction and data annotation when facing negative information about public affairs. In particular, when dealing with implicit negative emotions, accuracy is difficult to guarantee.
[0005] 3. Sentiment discrimination model based on deep learning: In recent years, with the application of deep learning models such as LSTM and GRU, the accuracy of text sentiment discrimination has improved. However, these models still require a lot of computing resources when facing large-scale complex texts, and their ability to distinguish subtle negative emotions (such as criticism and badmouthing) is insufficient, especially when dealing with public affairs topics. It is easy to make misjudgments.
[0006] There is a clear semantic difference between negative information about public affairs and negative emotions. Traditional techniques for identifying negative emotions cannot accurately distinguish negative information about public affairs. Existing techniques for identifying negative information about public affairs have low accuracy and are unable to handle complex contexts, as follows:
[0007] 1. Existing methods are not very accurate in judging negative information about public affairs.
[0008] Existing methods for identifying negative information about public affairs mainly include keyword matching, traditional machine learning models, and deep learning models. Among them, keyword matching relies on simple matching of pre-set keywords or phrases to negative information. Although this method is simple to implement, it lacks understanding of the language context and has difficulty dealing with complex language structures, especially for implicit negative emotions, sarcastic remarks, etc. Traditional machine learning models rely on manually designed features to identify negative information, but there is a shortage of negative information samples about public affairs, and the semantics are complex and changeable, making it difficult to accurately learn features. When deep learning models deal with complex topics such as politics and ideology, existing models cannot effectively capture potential negative emotions or biases, resulting in underreporting or misjudgment.
[0009] 2. Existing methods have weak ability to identify complex negative information.
[0010] Models such as SVM and Naive Bayes rely on artificially designed features to identify negative information. These models perform well in simple negative sentiment identification, but their expressiveness is limited when dealing with negative information about public affairs, especially when faced with complex language expressions and multi-dimensional negative information. Such models have insufficient ability to distinguish, and their training and reasoning efficiency is low when processing large-scale data, making it difficult to meet the needs of real-time identification. Limitations of deep learning models: In recent years, deep learning methods based on models such as BERT and LSTM have been introduced into the field of sentiment analysis, but when these models deal with complex topics such as politics and ideology, existing models cannot effectively capture potential negative emotions or biases, resulting in underreporting or misjudgment.
[0011] Therefore, the existing technology lacks a solution that can automatically and efficiently identify negative information about complex public affairs on massive social platforms, especially in the identification of negative information in complex language contexts. Summary of the invention
[0012] The present invention discloses a method and system for identifying negative information about public affairs, aiming to improve the ability to identify negative information about complex public affairs through an advanced language model.
[0013] To achieve the above object, the technical solution of the present invention includes the following contents.
[0014] A method for identifying negative information in public affairs, the method comprising:
[0015] Collect and screen information samples to obtain negative information samples;
[0016] Embed the filtered negative information samples into the prompt template to fine-tune the large model;
[0017] The target information is embedded into the prompt template, and based on the fine-tuned large model, the negative discrimination result of the target information is obtained.
[0018] Furthermore, the screening of negative information samples includes:
[0019] Use the iFlytek Spark Cognitive Model to score the negative degree of the collected information samples and obtain the scoring results of the information samples;
[0020] The scoring result of the information sample is compared with a set threshold to obtain a negative information sample.
[0021] Furthermore, before using the iFlytek Spark cognitive model to score the negative degree of the collected information samples, the method further includes:
[0022] Remove irrelevant content from the information sample; wherein the irrelevant content includes: @username, links, forwarding logos and emoticons.
[0023] Furthermore, the content of the prompt template includes: instruction description, input information and output information format; wherein, the content of the instruction description includes: please determine whether the given text is extremely bad negative information on public affairs. Note: you can only output yes / no and cannot perform any additional output; only when the text is obviously public affairs content, contains extremely bad negative information, and has a potential risk of serious endangerment to safety can it be judged as yes. Ordinary negative text will still be judged as no.
[0024] Furthermore, the large model includes: Qwen2-1.5B model.
[0025] Furthermore, after obtaining the negative discrimination result of the target information based on the fine-tuned large model, the method further includes:
[0026] Extracting keywords from the information sample, the keywords include: bonus keywords, deduction keywords and filtering keywords;
[0027] Extract keywords from target information;
[0028] Based on the matching of the keywords in the information sample and the keywords in the target information, the negative discrimination result of the target information is scored, and the discrimination correctness of the negative discrimination result is obtained according to the scoring result.
[0029] A negative information identification system for public affairs, the system comprising:
[0030] A sample screening module is used to collect and screen information samples to obtain negative information samples;
[0031] The large model fine-tuning module is used to embed the filtered negative information samples into the prompt template to fine-tune the large model;
[0032] The information discrimination module is used to embed the target information into the prompt template and obtain the negative discrimination result of the target information based on the fine-tuned large model.
[0033] An electronic device, characterized in that the electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements any of the above-mentioned methods for distinguishing negative information for public affairs.
[0034] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, any of the above-mentioned methods for distinguishing negative information for public affairs is implemented.
[0035] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, any of the above-mentioned methods for distinguishing negative information for public affairs is implemented.
[0036] Compared with the existing technology, the present invention significantly improves the ability to automatically distinguish negative information on complex public affairs through a large-scale negative information scoring and screening mechanism combined with a fine-tuned Qwen2-1.5B model, and can efficiently deal with potential negative information in the massive amount of information on social platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The present invention is a flow chart of the method for identifying negative information about public affairs. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below through specific implementations and in conjunction with the accompanying drawings.
[0039] The present application provides a system for distinguishing negative information about public affairs. The system is based on large model fine-tuning technology and solves the shortcomings of existing technologies in processing complex negative information by automatically distinguishing negative information, thereby improving the accuracy and efficiency of processing.
[0040] The main idea of the system is to use large-scale language models to intelligently identify negative information about public affairs, thereby effectively identifying high-risk negative information that poses a potential threat to national security. Figure 1 As shown, the specific process is as follows:
[0041] 1. Data pre-processing: First, the initial data from Weibo is processed, including removing irrelevant content (such as @username, links, forwarding logos, emoticons, etc.) to obtain clean training data. This step ensures the standardization and purity of the training data and provides high-quality data input for model training.
[0042] 2. Scoring and screening of negative information: Use the iFlytek Spark cognitive model to score the degree of negativity in the training data, with a score range of 1-10. Based on the scoring results, texts with scores below 5 are discarded, and only high-risk negative information (5 points and above) is retained. This step ensures that the system can focus on negative information with greater potential threats, improving the pertinence of the judgment results.
[0043] 3. Fine-tuning the Qwen2-1.5B model: Construct a prompt dataset for the filtered negative information. The dataset format is {"instruction":"Please judge whether the given text is extremely bad public affairs negative information. Note: You can only output 'yes / no' and cannot make any additional output; only when the text is obvious public affairs content, contains extremely bad negative information, and has a potential risk of seriously endangering national security can it be judged as 'yes', and ordinary negative text will still be judged as 'no'.","input":"……","output":"yes / no"}, and use this dataset to fine-tune the Qwen2-1.5B model. The fine-tuned model can efficiently judge negative public affairs information and filter out high-risk content.
[0044] 4. Negative information identification: During the model identification process, the Qwen2-1.5B model will judge negative information with a score of 5 or above and output a "yes / no" result. The model can automatically identify potential negative information about public affairs and efficiently filter out text that does not meet high-risk standards.
[0045] 5. Keyword scoring mechanism: Use the keywords obtained in advance through entity extraction to score the results of the previous step. This scoring mechanism can further eliminate noise and standardize the output of the model, thereby ensuring more accurate results.
[0046] Through this technical solution, the system effectively solves the efficiency and accuracy problems in the existing technology in the process of identifying negative information in public affairs, improves the ability to quickly identify complex information, and is suitable for data monitoring tasks on large-scale social platforms.
[0047] In summary, the present invention cleans and processes the initial data from Weibo, removes irrelevant information (such as @usernames, links, emoticons, etc.), and ensures the standardization and consistency of the data. This process eliminates noise, ensures that the input data of the negative information discrimination task is clean and reliable, and provides high-quality data support for subsequent negative information discrimination.
[0048] The present invention uses the iFlytek Spark cognitive model to negatively score the pre-processed data, with a score range of 1-10. Data with a score of 5 or above is retained, and data with a score below 5 is discarded. This scoring mechanism helps the system prioritize information with a higher degree of negativity, ensuring that the model focuses on high-risk negative information on public affairs, and improving the accuracy and efficiency of discrimination.
[0049] The present invention converts the screened data with scores of 5 or above into a prompt data set in a specific format, and the format is {"instruction":"Please judge whether the given text is extremely bad public affairs negative information. Note: You can only output 'yes / no' and cannot perform any additional output; only when the text is obvious public affairs content, contains extremely bad negative information, and has a potential risk of seriously endangering national security can it be judged as 'yes'. Ordinary negative text will still be judged as 'no'.","input":"……","output":"yes / no"}. This step provides support for the fine-tuning of the Qwen2-1.5B model by constructing a specific task data set to ensure that the model is suitable for the task of discriminating negative public affairs information.
[0050] The present invention uses the constructed prompt data set to fine-tune the Qwen2-1.5B model, so that it has the ability to identify negative public affairs information. The model determines whether it is negative public affairs information based on the input text and outputs a "yes / no" result. The performance of this model in screening negative public affairs information is crucial, marking the text that meets the conditions for subsequent classification or processing.
[0051] Finally, the present invention regulates the output of the model through a keyword scoring mechanism. Keywords are divided into scoring keywords (add 1 point for each hit), subtracting keywords (subtract 1 point for each hit), and filtering keywords (-100 points after a hit). By limiting the rules in this way, irrelevant information can be further filtered to obtain expected results.
[0052] Next, a specific experiment is used to illustrate the method for identifying negative information on public affairs provided by the present invention.
[0053] The hardware configuration of the experiment of the present invention is shown in Table 1.
[0054] operating system CentOS Linux release 7.5.1804 Memory 125G CPU Intel(R)Xeon(R)CPU E5-2667v4@3.20GHz
[0055] Table 1
[0056] Experimental design: This paper uses a first-level annotation model - qwen2-1.5b, without additional fine-tuning parameters, based on three V100 graphics cards, to simulate the performance in the negative information discrimination task. The test data set includes 1,000 manually annotated negative samples and 1,500 positive samples from social platforms. The main content of the test is a two-classification task, that is, judging whether the input text is negative information.
[0057] The experimental results are shown in Table 2:
[0058] Accuracy 93.6% Negative Precision 96.4% Recall 87.3% Reasoning time About 10 minutes for 2500 data, using V100*3 inference resources
[0059] Table 2
[0060] As shown in Table 2, the method proposed in the present invention has achieved excellent performance in the task of discriminating negative information, especially in terms of accuracy, reaching 96.4%, which significantly improves the accuracy of the system in complex negative information scenarios. Although the reasoning speed of this method is relatively slow, and it takes about 0.24 seconds to process each piece of data, its high efficiency in screening and classifying large-scale negative data makes it an ideal choice for big data applications. In addition, the model achieves accurate identification of negative information through the support of large-scale parameter quantities, and achieves a high recall rate while ensuring accuracy, which is suitable for practical scenarios that require high accuracy and precision.
[0061] The above embodiments are provided only for the purpose of describing the present invention, and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention should all be included within the scope of the present invention.
Claims
1. A method for identifying negative information in public affairs, characterized in that: The method comprises: Collect and screen information samples to obtain negative information samples; Embed the filtered negative information samples into the prompt template to fine-tune the large model; The target information is embedded into the prompt template, and based on the fine-tuned large model, the negative discrimination result of the target information is obtained.
2. The method according to claim 1, characterized in that The screening of negative information samples includes: Use the iFlytek Spark Cognitive Model to score the negative degree of the collected information samples and obtain the scoring results of the information samples; The scoring result of the information sample is compared with a set threshold to obtain a negative information sample.
3. The method according to claim 2, characterized in that Before using the iFlytek Spark Cognitive Model to score the negative degree of the collected information samples, the following steps are also included: Remove irrelevant content from the information sample; wherein the irrelevant content includes: @username, links, forwarding logos and emoticons.
4. The method according to claim 1, characterized in that The content of the prompt template includes: instruction description, input information and output information format; wherein, the content of the instruction description includes: please judge whether the given text is extremely bad public affairs negative information, note: you can only output yes / no, and cannot perform any additional output; only when the text is obvious public affairs content, contains extremely bad negative information, and has the potential risk of serious endangerment to safety can it be judged as yes, and ordinary negative text will still be judged as no.
5. The method according to claim 1, characterized in that The large model includes: Qwen2-1.5B model.
6. The method according to any one of claims 1 to 5, characterized in that: After obtaining the negative discrimination result of the target information based on the fine-tuned large model, it also includes: Extracting keywords from the information sample, the keywords include: bonus keywords, deduction keywords and filtering keywords; Extract keywords from target information; Based on the matching of the keywords in the information sample and the keywords in the target information, the negative discrimination result of the target information is scored, and the discrimination correctness of the negative discrimination result is obtained according to the scoring result.
7. A negative information identification system for public affairs, characterized in that: The system comprises: A sample screening module is used to collect and screen information samples to obtain negative information samples; The large model fine-tuning module is used to embed the filtered negative information samples into the prompt template to fine-tune the large model; The information discrimination module is used to embed the target information into the prompt template and obtain the negative discrimination result of the target information based on the fine-tuned large model.
8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for identifying negative information for public affairs as described in any one of claims 1-6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method for identifying negative information for public affairs as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method for identifying negative information for public affairs as described in any one of claims 1 to 6.