Big model-based ai false information network propagation scenario analysis method and system

By generating multiple sets of AI information and randomly extracting subsets to train a large user model, the state transitions of different individual users are simulated. This solves the problems of low efficiency and poor generalization in user feature learning in existing technologies, and achieves more accurate analysis of the spread of false information and prediction of debunking strategies.

CN120567706BActive Publication Date: 2026-02-17SCHOOL OF INFORMATION & COMM TECH NAT UNIV OF DEFENSE TECH OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510669972.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2026-02-17
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency in learning user characteristics, poor generalization and flexibility when analyzing AI-driven dissemination of misinformation on the internet. They are unable to accurately simulate individual user background knowledge and reactions, resulting in inaccurate analysis of misinformation dissemination.

Method used

By generating multiple sets of AI information, randomly extracting subsets for large-scale model training, generating multiple large user models, and randomly placing these models in the network structure for simulation, the state transitions of different individual users are simulated. Combined with threat level labeling and debunking user models, the scenarios of false information dissemination are analyzed.

Benefits of technology

It improves the simulation accuracy and generalization of false information dissemination scenarios, and can simulate debunking strategies under the premise of controllable security risks, predict the spread trend of false information, provide decision support for social media and network supervision, and enhance the ability to identify false information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567706B_ABST
    Figure CN120567706B_ABST
Patent Text Reader

Abstract

The application provides an AI false information network propagation scene analysis method and system based on a large model, and relates to the technical field of deep learning; the method comprises the following steps: generating an AI information set; randomly extracting at least one group of AI information in the AI information set multiple times to obtain multiple subsets, and separately training the large model by taking the multiple subsets as training sets to generate multiple different user large models; randomly placing the different user large models on user nodes in a network structure for simulation, selecting an arbitrary user node in the network structure to put AI false information, simulating different user individuals by using the corresponding user large models on each user node, and realizing the analysis of AI false information in the network propagation scene by analyzing the state conversion of each user node in the network structure. The application can solve the problems of low efficiency of network user node feature learning and poor generalization and flexibility of false information propagation analysis methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a method and system for analyzing AI-based disinformation network propagation scenarios based on large models. Background Technology

[0002] With the development of artificial intelligence (AI) technology, the amount of AI-generated information is growing rapidly. The generation and dissemination of AI-generated information on the internet have undergone entirely new changes, and the use of AI technology to create false information has become a new type of cyber threat. The spread of AI-generated false information online is typically characterized by rapid dissemination, strong concealment, difficulty in identification, a focus on gaining clicks and traffic, and an attempt to manipulate the emotions and stances of the audience through emotional presentation. Social networks are rife with fake text, fake images, and fake audio and video generated using AI technology; the era of "a picture is worth a thousand words" is over. In the age of artificial intelligence, people use posting, commenting, and forwarding to rapidly spread false information in a short period of time. Various types of online false information are flooding netizens with an explosive surge, making it difficult to distinguish the authenticity of the content and greatly disrupting people's normal lives and social security and stability. Therefore, analyzing the dissemination scenarios of AI-generated false information online is of great significance for curbing its spread, maintaining a healthy online ecosystem, and ensuring social stability.

[0003] Traditional methods for analyzing the spread of misinformation primarily utilize the Susceptible-Infected-Recovered (SIR) model, based on propagation dynamics, to describe the state changes of user nodes during the spread of misinformation across networks. The main principle is to categorize user nodes in the network into three states: susceptible (S), infected (I), and immune (R). In the SIR model, there are transition probabilities between different user states; for example, a user node in susceptible state S will transition to infected state I after believing the misinformation. With the development of deep learning, many scholars have improved the SIR model to analyze the spread of AI information in networks through various modifications, such as adding memory mechanisms, trust thresholds, group effects, and alert mechanisms. The main principle is to collect misinformation from the network, including false topics, false content, user comments, and user propagation paths, and then use natural language processing, graph neural networks, and other techniques to learn and train the feature vectors of user nodes and misinformation to make the propagation scenario more concrete.

[0004] However, these methods still have many drawbacks: First, extracting and learning user features based on user comments is relatively inefficient. In real-world scenarios, most user comments are short, and some users who receive false information may not even comment, making it difficult for these methods to truly learn the individual user's perception of false information. Second, they have poor generalization ability. In large social networks, the dataset used to train the model is unlikely to cover the entire social network structure. In other words, many individual users have not received false information, so these users cannot be trained on the false information, resulting in poor generalization ability of these methods. Third, they lack flexibility. These methods only train and generate individual user features based on user comment information, which often only contains the user's opinion on false information, ignoring the uniqueness of the user's individual background knowledge, resulting in poor flexibility. Summary of the Invention

[0005] Therefore, it is necessary to provide an AI-based method and system for analyzing the spread of misinformation on the internet, based on a large model, to address the aforementioned technical problems and solve the issues of inefficient user feature learning and poor generalization and flexibility of the analysis method.

[0006] On the one hand, this invention provides a method and system for analyzing AI-driven online dissemination scenarios of misinformation based on large-scale models, the method comprising:

[0007] Generate an AI information set; the AI ​​information set includes multiple sets of AI information, each set of AI information includes one piece of false AI information generated by AI and one piece of real information corresponding to the false AI information;

[0008] At least one set of AI information is randomly extracted from the AI ​​information set multiple times to obtain multiple subsets. Each subset is used as a training set to train the large model separately. One subset corresponds to one user large model, thereby generating multiple different user large models.

[0009] Different user models are randomly placed on user nodes in the network structure for simulation. AI false information is then deployed to any user node in the network structure. The AI ​​false information propagates along the edges of the network structure to surrounding user nodes. Different user individuals are simulated using the corresponding user models on each user node. The analysis of the state transitions of each user node in the network structure is used to analyze the AI ​​false information in the network propagation scenario.

[0010] On the other hand, the present invention provides an AI-based system for analyzing the spread of misinformation on a large model network, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0011] Overall, this invention provides a method and system for analyzing the spread of AI-based misinformation on the internet, based on a large model, which achieves the following advantages compared to existing technologies:

[0012] (1) This invention obtains multiple subsets by randomly extracting at least one set of AI information from the AI ​​information set multiple times, and uses these subsets as training sets to train the large model separately. Each subset corresponds to a user large model, thereby generating multiple different user large models. The different user large models are randomly placed on user nodes in the network structure for simulation, and different user individuals are simulated using the user large models corresponding to each user node. The settings can be configured according to the actual needs of the simulation, so that each user node randomly has different background knowledge, rather than relying solely on user comments to extract and learn user features. This reduces the influence of the analysis method on the size and quality of the dataset, and avoids the problem of low efficiency in learning user features in previous information dissemination models. In addition, randomly extracting multiple subsets of AI information for separate training can yield more types of user individuals, thereby covering the entire network structure, which is closer to the real dissemination scenario and improves the generalization and flexibility of the simulation scenario.

[0013] (2) This invention marks AI false information with threat level and obtains multiple subsets by randomly extracting AI information sets from the threat level dataset. This allows for greater diversity of individual users. Even if many individual users have not received AI false information, they can still be trained, thereby improving the generalization and flexibility of the entire propagation scenario and improving the accuracy of AI false information network propagation scenario analysis results.

[0014] (3) This invention randomly places multiple different user big models into user nodes in the network structure for simulation. Different user big models are used to simulate different user individuals. The user big models are divided into ordinary user big models, rumor-refuting user big models, rumor-spreading user big models, and deceived user big models. This allows different user individuals in the network structure to change their node state according to the received false information and rumor-refuting information, thereby more realistically simulating and predicting the state changes of each user node and the impact on the network space when AI false information spreads in the network.

[0015] (4) Under the premise of controllable security risks, this invention can simulate the network propagation scenario of AI-related false information and simulate various debunking strategies in the simulation environment, such as official authoritative releases, expert interpretations, and user debunking. By setting different debunking times, methods, and channels, it can predict the changes in individual users' perceptions of false information, and thus effectively predict the debunking effect of each debunking strategy at different stages of propagation. This provides effective support for curbing the spread of AI-related false information on the network and formulating efficient debunking solutions, thereby preventing the spread and impact of AI-related false information.

[0016] (5) This invention can be used to predict the spread trend of AI misinformation in the future, make preparations in advance, and provide decision support for social media platforms and network regulatory departments. For example, it can predict which regions and groups of people may become hotspots for the spread of misinformation, thereby strengthening information monitoring and guidance in a targeted manner to prevent the large-scale spread of AI misinformation in these regions or groups of people.

[0017] (6) The debunking user big model proposed in this invention is trained by randomly extracting subsets of AI false information from the first threat dataset, the second threat dataset and the second threat dataset, which can improve the ability to identify AI false information. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. The drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating the steps of an AI-based method for analyzing the spread of misinformation on the internet, based on a large model, provided by this invention.

[0020] Figure 2 This is a diagram illustrating AI-generated false information and real information obtained from the Science Debunking Network in this invention.

[0021] Figure 3 This is a schematic diagram of AI-generated fake information using a large model, as described in this invention.

[0022] Figure 4 This is a schematic diagram illustrating the generation of AI-generated false information using a large model interactive interface according to the present invention;

[0023] Figure 5 This is a schematic diagram illustrating how the present invention trains multiple large models for multiple different users using multiple training sets;

[0024] Figure 6This is a schematic diagram of the LoRA adapter module of the present invention;

[0025] Figure 7 This is a schematic diagram showing the distribution of the LoRA adapter modules of the present invention;

[0026] Figure 8 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 1 ;

[0027] Figure 9 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 2 ;

[0028] Figure 10 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 3 ;

[0029] Figure 11 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 4 ;

[0030] Figure 12 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 5 ;

[0031] Figure 13 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 6 ;

[0032] Figure 14 This is a scenario analysis illustration of the present invention without a large user model for debunking rumors. Figure 7 ;

[0033] Figure 15 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 1 ;

[0034] Figure 16 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 2 ;

[0035] Figure 17 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 3 ;

[0036] Figure 18 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 4 ;

[0037] Figure 19 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 5 ;

[0038] Figure 20This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 6 ;

[0039] Figure 21 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 7 ;

[0040] Figure 22 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 8 ;

[0041] Figure 23 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 9 ;

[0042] Figure 24 This is a scenario analysis illustration of the present invention with a large user model for debunking rumors. Figure 10 . Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0044] It should be noted that, in the description of the embodiments of the present invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method, step, or system that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to the method, step, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the method, step, or system that includes said element.

[0045] To address the new challenges posed by the rapid development of AI technology and the spread of AI-generated misinformation within network structures, this invention proposes a method and system for analyzing AI-generated misinformation network propagation scenarios based on a large-scale model. Specifically, as follows... Figure 1 As shown, the method includes:

[0046] Step 101: Generate the AI ​​information set. The AI ​​information set includes multiple sets of AI information. Each set of AI information includes one piece of AI-generated false information and one piece of real information corresponding to the AI ​​false information.

[0047] As an example, the method for generating AI-generated misinformation includes: obtaining online misinformation and corresponding real information from a science debunking website; using the Qwen big data model, following a prompt word framework, inputting a false topic, and generating AI-generated misinformation corresponding to that topic. This AI-generated misinformation, along with the corresponding real information, forms a set of AI information; multiple sets of AI information constitute an AI information set.

[0048] For example, you can collect different kinds of online misinformation and their corresponding factual information on the Science Rumor Debunking Network (https: / / piyao.kepuchina.cn / ); such as Figure 2 As shown, in the Science Rumor Debunking Network, each set of information includes data such as false topic, false information (rumor), real information, and information source. Based on the prompt word framework, the false topic is input into the large model to generate corresponding false information, and the "false information" and the corresponding "real information" are treated as a set of AI information. Figure 3 As shown, this invention randomly selected two different sets of false topics from the Science Rumor Debunking Network, and generated corresponding false information based on the false topics by using the Qwen large model and following the prompt word framework.

[0049] Cue words refer to the designed and optimized input prompts used when using large models for generative tasks. These prompts guide the model to produce expected outputs, helping to control and adjust the model's output. In traditional large language models, users typically only need to provide a simple prompt or question, and the model generates a corresponding answer or sentence based on training data and pre-trained knowledge. However, this approach can lead to inaccurate, ambiguous, or even unsuitable content. Therefore, this invention utilizes a cue word framework to introduce more interaction and guidance, allowing for more precise control over the large model's output. This helps the model better understand user needs and generate more appropriate content.

[0050] It should be noted that the prompt word framework includes: ICIO framework or CRISPE framework.

[0051] The ICIO framework consists of an instruction section, a context section, an input data section, and an output guidance section. Different AI-generated false information can be output by setting the content of different sections.

[0052] The instruction section consists of explicit guidelines or questions in the form of text instructions that guide the model in generating responses. These can be brief statements or specific questions.

[0053] Context: Context refers to the additional information provided to the model during the dialogue, such as dialogue history, dialogue summary, keywords or tags.

[0054] The input data part refers to the raw or preprocessed text that is used as input to the model, which may include instruction text, context data, formatting tags, control terms, and other relevant information.

[0055] The output guidance section refers to the output indicators that inform the model how and in what format the text is generated. This helps the model understand the expected output type and generate the corresponding response.

[0056] For example, AI-generated misinformation can be generated using the ICIO framework, as shown in Table 1. The ICIO framework clearly defines the requirements for the generated AI-generated misinformation. The instructions describe the main content of the misinformation, the context provides supplementary information, the input data specifies the opening content, and the output guidance explicitly specifies the style of the output. This helps large models accurately understand the task and generate AI-generated misinformation that meets the requirements of a specific scenario.

[0057] Table 1 ICIO Framework

[0058]

[0059] The CRISPE framework includes information on capabilities and roles, insights, statements, personality traits, and experiments, making it suitable for writing more complex AI-generated misinformation.

[0060] The Capabilities and Roles section includes role descriptions and scope of responsibilities. The role descriptions detail the roles the large model needs to play, such as product manager, lawyer, doctor, educator, etc., to ensure that the large model understands its role responsibilities in a specific field. The scope of responsibilities clarifies the scope of the large model's responsibilities in the selected role, including what aspects need to be considered and involved.

[0061] The insights section includes industry background, target audience, and specific context. The industry background provides detailed information about the industry, including its history, trends, competitive landscape, and key players. The target audience is specifically defined, including its characteristics, needs, preferences, and behaviors. The specific context describes the context in which the content is generated, such as geographical location, time, and event, so that the large model can generate relevant content based on the context.

[0062] The statement section includes task breakdown, expected output, and key information. Task breakdown divides the task into smaller subtasks, which helps the large model understand and process the task and ensures that the task is clear and unambiguous. Expected output specifies the required output, such as a report, suggestion, solution, or answer to a specific question. Key information emphasizes key information to ensure that the output of the large model meets the needs of the task.

[0063] The personality section includes tone and emotion, and format requirements. The tone and emotion section describes the required tone, such as friendly, formal, or encouraging, and specifies the emotion requirements, such as optimistic, neutral, or serious. The format requirements define the format of the generated content, such as paragraphs, lists, tables, and charts, to ensure that the document conforms to specific format standards.

[0064] The experiment section includes generating example types, quantity and diversity, feedback and improvement, etc. The generating example types specify the required types of generated examples, such as paragraphs, question answers, suggestions, data analysis, etc.; quantity and diversity specify the required number of generated examples and the required diversity descriptions to ensure that the task requirements are met; feedback and improvement provide a feedback mechanism that allows the large model to be modified and improved to meet the requirements.

[0065] For example, AI-generated misinformation can be generated using the CRISPE framework, as shown in Table 2.

[0066] Table 2 CRISPE Framework

[0067]

[0068] As an example, to make the process of generating AI-generated fake information more convenient and intuitive, a web page for interacting with a large model is written using the gradio package in Python. This allows users to remotely access the web page to input commands and view the output results. Figure 4 The image shows the results of generating AI-generated misinformation using an interactive interface based on the Qwen large model.

[0069] Of course, you can also directly extract false information and corresponding real information from the science debunking website to obtain multiple sets of AI information.

[0070] Step 102: Randomly extract at least one set of AI information from the AI ​​information set multiple times to obtain multiple subsets, and use each subset as a training set to train the large model separately. Each subset corresponds to a user large model, thereby generating multiple different user large models.

[0071] It should be noted that the training of the large model is independent for each subset. That is, each subset will train a different user large model. In this way, each user large model can simulate different users, and each user has different background knowledge, thus avoiding the problem of inefficient learning of user features in previous propagation models.

[0072] For example, such as Figure 5As shown, the AI ​​information set is randomly divided into n subsets, and different subsets are allowed to contain the same false AI information. For each of the n subsets, n large user models are trained to simulate n individual users in the network. It should be noted that the number of large user models is usually equal to the number of nodes in the network structure.

[0073] As an example, the LoRA fine-tuning method is used to train a large model separately using multiple subsets as training sets. That is, the LoRA fine-tuning method uses a subset as a training set to train a large model separately, resulting in a single large user model.

[0074] It's important to note that LoRA fine-tuning is a crucial method. Deep neural networks involve numerous matrix multiplication operations, and the parameters of a neural network can be represented as a series of matrices of varying sizes. If a portion of the parameters in a neural network are matrices... After fine-tuning, it became Then matrix multiplication in neural networks ,in, Indicates input data, This represents the result of the original neural network operation. This represents the result after fine-tuning. In large language models, matrices... The number of parameters is very large, in order to reduce The number of parameters can be approximated using a low-rank matrix. ,Right now Assuming yes The matrix, and , They are and The matrix, when and Very large and When the size is small, then The structural principle of the LoRA adapter module is as follows: Figure 6 As shown, it can be seen that in LoRA, This can be viewed as a dimensionality reduction operation. This can be viewed as an operation to increase the dimensionality of something.

[0075] The LoRA adapter module can be placed in different locations within the Transformer layer. Specifically, it can be placed within the linear mapping module. , or It can also be placed next to the multi-head self-attention module. Next to, or placed in the feedforward network module , Next to. For example, Figure 7 As shown in the illustration, as an example, the present invention distributes the LoRA adapter modules evenly across... , or and The best position is where it works; preferably, place it in... and The location offers the best value for money.

[0076] As one embodiment, the method further includes: labeling the AI ​​misinformation in each group of AI information with a threat level, dividing the AI ​​information into different threat level datasets; and randomly extracting at least one group of AI information from the same or different threat level datasets as a training set. Further, the threat level datasets include a first threat level dataset, a second threat level dataset, and a third threat level dataset, arranged from low to high threat levels. For example, they may include a low threat level dataset, a medium threat level dataset, and a high threat level dataset.

[0077] The methods for marking threat levels include one or more of TER, BLEURT, BARTScore, or manual marking.

[0078] TER (Translation Edit Rate) is the minimum number of edits required to modify the generated text to perfectly match any reference text. This number is normalized using the average length of the reference text.

[0079] BLEURT (Bilingual Evaluation Understudy using Representations from Transformers) is a metric used to evaluate the quality of generated text.

[0080] BARTScore is a metric used to evaluate the quality of generated text. It treats the evaluation of generated text as a text generation problem. The core idea is that the higher the quality of the generated text, the higher the probability that the pre-trained model will convert the generated text into the reference output or source text.

[0081] Human labeling involves setting evaluation criteria and using a learning model to assess the threat level of AI-generated misinformation. These criteria typically include: fluency (e.g., whether the language used to generate the misinformation is natural and fluent, and whether there are any grammatical errors, spelling mistakes, or awkward phrasing); relevance (e.g., how relevant the generated text is to a given topic, question, or instruction); information completeness (e.g., whether the generated text contains key information); credibility (e.g., whether the generated text is logically sound and based on common sense); and diversity (e.g., measuring the diversity of expression in the generated text to avoid repetitive and monotonous statements).

[0082] Existing models for spreading misinformation often overlook the unique characteristics of misinformation and the background knowledge of internet users, resulting in a significant discrepancy between the model's propagation effect and real-world scenarios. Therefore, this invention categorizes individual users into four roles: ordinary users, debunking users, users spreading misinformation, and deceived users. As an embodiment of this invention, the state types of user nodes in the network structure include: ordinary user nodes, debunking user nodes, users spreading misinformation, and deceived user nodes.

[0083] It's important to clarify that user nodes spreading false information are those that initially publish AI-generated false information and then propagate it to their neighboring user nodes. Debunking user nodes, upon receiving AI-generated false information, become active and propagate accurate information to their neighboring user nodes to prevent them from being deceived by rumors. Deceived user nodes are those that have been misled by AI-generated false information and may continue to spread it. From the perspective of AI-generated false information propagation, ordinary user nodes face two possibilities after receiving false information: if they believe it, they become deceived user nodes; otherwise, they remain ordinary user nodes. From the perspective of accurate information propagation, ordinary user nodes become debunking user nodes upon receiving accurate information; while user nodes spreading false information and deceived user nodes, upon receiving accurate information, choose to believe the facts and become debunking user nodes.

[0084] As an example, at least one set of AI information is randomly extracted from the same or different threat level datasets as a training set, including: randomly extracting at least one set of AI information from the first threat level dataset multiple times to obtain multiple subsets, and generating multiple different large models of ordinary users accordingly; randomly extracting at least one set of AI information from the first threat level dataset, the second threat level dataset, and the third threat level dataset multiple times respectively to obtain multiple subsets, and generating multiple different large models of debunking users accordingly.

[0085] It should be noted that both the user nodes spreading false information and the deceived user nodes are directly deceived when they receive false information from the AI. Therefore, their node states do not need to be simulated using a large model.

[0086] When AI-generated misinformation spreads in cyberspace, in order to simulate the reaction of network user nodes to AI-generated misinformation as realistically as possible, this invention uses Qwen2.5-0.5B-Instruct as the large model base, and uses randomly extracted real information and corresponding AI-generated misinformation as the training set to train several large models of ordinary users and large models of debunking users respectively.

[0087] The general user model is used to determine whether ordinary network nodes believe and continue to spread AI-generated false information after receiving it; the debunking user model is used to determine whether debunking user network nodes will recognize and publish true information after receiving AI-generated false information.

[0088] This invention targets 34 nodes in Zachary's Karate Club network and generates 34 large user models using the LoRA fine-tuning method. Specifically, 60 pieces of AI-generated misinformation are labeled with threat levels to obtain low-threat, medium-threat, and high-threat datasets.

[0089] For the large-scale model for ordinary users, 50% of the AI ​​information from the low-threat dataset was randomly selected as the training set, and the remaining AI information was used as the validation set for fine-tuning the large-scale model. To understand the extent to which the large-scale model for ordinary users reacts to different AI false information, performance validation was performed on all large-scale models for ordinary users after training, and the results are shown in Table 3.

[0090] Table 3 Performance of Large Model for Ordinary Users (Unit: %)

[0091]

[0092] As can be seen, tests were conducted on different threat levels and the recognition rates of AI-generated misinformation and real information contained within them. For example, in the medium threat dataset, the average accuracy rate of the large model for 34 ordinary users in recognizing AI-generated misinformation was 67.65%, while the average accuracy rate in recognizing real information was 77.21%.

[0093] Although the large-scale model for ordinary users only used a subset of AI information from the low-threat dataset as training data, in the testing phase, the recognition rates for AI-generated fake information with low, medium, and high threat levels were 88.88%, 67.65%, and 60.74%, respectively, while the recognition rate for all AI-generated fake information was 76.67%. This demonstrates that using multiple subsets of AI information randomly extracted from the low-threat dataset as training data and employing a large model to simulate network users' reactions to AI-generated fake information has a certain degree of generalization and feasibility.

[0094] In scenarios involving the spread of AI-generated misinformation online, debunking users need to be more accurate than ordinary users in distinguishing between AI-generated misinformation and real information online. To gain a deeper and more accurate understanding of the training performance of the large-scale debunking user models, this invention uses three different training settings for testing. In each setting, 10 large-scale debunking user models are generated to test their accuracy in identifying AI-generated misinformation and real information. The training settings for each large-scale debunking user model are as follows:

[0095] For the debunking user large model A, only 90% of the AI ​​information in the low-threat dataset is randomly extracted as the training set, and the remaining data is used as the validation set.

[0096] For the debunking user large model B, 90% of all AI information in the AI ​​information set is randomly extracted as the training set, and the remaining data is used as the validation set.

[0097] For the debunking user large model C, 50% of all AI information in the AI ​​information set is randomly extracted as the training set, and the remaining data is used as the validation set.

[0098] By comparing the large-scale debunking user models A and B, we can understand the training effect of using AI-generated misinformation of varying threat levels as training data. Similarly, by comparing the large-scale debunking user models B and C, we can understand the training effect of using different amounts of AI-generated misinformation as training data. Specific experimental results are shown in Table 4.

[0099] Table 4 Performance of the large-scale user model for debunking rumors (unit: %)

[0100]

[0101] Table 4 shows that the debunking user model B performed best, achieving 100% accuracy in identifying all AI-generated misinformation. This is because the training of the debunking user model B not only expanded the data types with different threat levels but also used more AI-generated misinformation samples for training. Compared to the debunking user model B, the debunking user model A, which was not trained on AI-generated misinformation with medium and high threat levels, had a poorer ability to identify AI-generated misinformation in the "medium" and "high" threat levels. While the debunking user model C was trained using AI-generated misinformation with multiple threat levels, it used relatively fewer samples for training. Therefore, compared to the debunking user model A, its experimental results were worse in low-threat scenarios, but better in high-threat scenarios.

[0102] By comparing the performance of the large-scale model of debunking users (Type A) and the large-scale model of ordinary users, it can be found that, without expanding the threat level, increasing the training samples of AI-generated misinformation can greatly enhance the accuracy of the large-scale model in identifying AI-generated misinformation.

[0103] Step 103: Randomly place different user models onto user nodes in the network structure for simulation. Select any user node in the network structure to deliver AI false information. The AI ​​false information propagates along the edges in the network structure to surrounding user nodes. Use the corresponding user models on each user node to simulate different individual users. Analyze the state transitions of each user node in the network structure to realize the analysis of AI false information in the network propagation scenario.

[0104] In other words, based on the network dataset, this invention generates a large user model corresponding to each user node in the network structure. When AI misinformation is introduced into the network structure, the user nodes in the network structure will be transformed into different node states according to the identification results of the AI ​​misinformation by the large model, thereby describing the propagation scenario of AI misinformation in a complex network.

[0105] The network structure used in this invention can be the Zachary's Karate Club dataset. To analyze AI-driven disinformation network propagation scenarios, a total of 46,240 AI-driven disinformation propagation scenarios were simulated, including 1,360 scenarios without debunking users and 44,880 scenarios with debunking users.

[0106] Next, we will demonstrate and analyze the general and special dissemination scenarios under the two types of conditions. In all the diagrams, gray nodes represent ordinary users, red nodes represent users spreading false information and deceived users, and green nodes represent users debunking rumors.

[0107] In a propagation scenario without a large-scale user model for debunking misinformation, AI-generated misinformation will spread from the initial user node that spreads the misinformation until it reaches all nodes in the network, or until the node states in the network no longer change. To further investigate the propagation scenarios of AI-generated misinformation with different threat levels in the network, this invention will analyze the propagation of high-threat AI-generated misinformation and low-threat AI-generated misinformation in the network separately.

[0108] In a propagation scenario without a large-scale user model for debunking misinformation, when AI-generated misinformation poses a high threat and the initial user node spreading the misinformation is a super node, such as... Figure 8 As shown, AI-generated misinformation spreads extremely quickly, covering the entire network in just three transmissions. It's important to note that supernodes are nodes with special properties or play crucial roles in the network, typically characterized by high connectivity, significant influence, strong centrality, and abundant resources. On the other hand, when a node (especially a supernode) can identify AI-generated misinformation and ceases to spread it, such as... Figure 9 As shown, the impact of AI-generated misinformation on cyberspace will be reduced to some extent. Therefore, it is necessary to strengthen the AI-generated misinformation identification capabilities of supernodes in cyberspace. When supernodes are unaffected by AI-generated misinformation, the spread speed of AI-generated misinformation and its impact on cyberspace will be greatly reduced.

[0109] In a propagation scenario without a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is high, and the initial user node spreading the misinformation is a general node, such as... Figure 10 As shown, since the user nodes that initially spread false information are relatively close to the super nodes, even highly threatening AI false information can still spread rapidly through the super nodes.

[0110] In a propagation scenario without a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is high and the initial user node spreading the misinformation is a marginal node, such as... Figure 11 As shown, AI-generated misinformation spreads slowly at first, but its spread accelerates dramatically once it reaches supernodes. Conversely, when key nodes (such as supernodes) close to the user nodes initially spreading the misinformation can identify it, such as… Figure 12 As shown, the spread of AI-generated misinformation may be contained to a small area and prevented from spreading outwards.

[0111] In a propagation scenario without a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is low, and the initial user node spreading the misinformation is a super node, such as... Figure 13As shown, because the threat level of AI-generated misinformation is low, surrounding nodes are all capable of recognizing it, thus preventing its spread in cyberspace. On the other hand, when the threat level of AI-generated misinformation increases slightly, and there are nodes around the supernode that have been deceived by the misinformation, such as... Figure 14 As shown, AI-generated misinformation can spread in a small area and may propagate along a specific path.

[0112] In a propagation scenario with a large user model for debunking misinformation, when the threat level of AI-generated misinformation is high, the debunking user node is adjacent to the initial user node spreading the misinformation, and the initial user node spreading the misinformation is a super node, such as... Figure 15 Although debunking user nodes are adjacent to supernodes and can detect the spread of AI-generated misinformation immediately, the speed at which supernodes spread this misinformation means that by the time debunking user nodes release accurate information, the misinformation has already reached a considerable scale in cyberspace. Conversely, when debunking user nodes are located at the edge of cyberspace, such as... Figure 16 As shown, although debunking user nodes can detect AI-generated misinformation and disseminate accurate information immediately, the spread of accurate information lags far behind the spread of AI-generated misinformation. Therefore, it is necessary to deploy a certain number of debunking user nodes near supernodes. On the one hand, these nodes can easily detect AI-generated misinformation reaching supernodes and enhance the detection of information released by supernodes. On the other hand, once an outbreak of AI-generated misinformation occurs, these debunking user nodes can disseminate accurate information through supernodes, thereby reducing the impact of AI-generated misinformation on cyberspace.

[0113] In a propagation scenario with a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is high, and the debunking user nodes are far apart from the initial user nodes spreading the misinformation, such as... Figure 17 As shown, debunking user nodes require a long reaction time to detect AI-generated misinformation and disseminate accurate information, by which time the AI-generated misinformation has already had a significant impact on cyberspace. Furthermore, as... Figure 18 As shown, when the debunking user node is located at the edge of cyberspace, the debunking user node cannot detect the AI ​​false information because the key node 0 identifies that the AI ​​false information has not spread. Therefore, in this scenario, even if there are debunking user roles in cyberspace, the AI ​​false information cannot be eliminated.

[0114] In a propagation scenario with a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is high, the initial user node spreading the misinformation is located in a peripheral area, and the debunking user node is located on the critical path of AI-generated misinformation dissemination, such as... Figure 19As shown, AI-generated misinformation can be eliminated promptly when its impact is minimal. On the other hand, if key node 0 has already identified the AI-generated misinformation but has not spread it before the debunking user node discovers it, such as... Figure 20 As shown, this AI-generated misinformation will remain within a certain range and cannot be eliminated. When the threat level of AI-generated misinformation is low, its spread may be controlled within a small area, preventing it from spreading outwards. This results in AI-generated misinformation being undetectable and undebunked. Therefore, it is necessary to set up some debunking user nodes in the edge areas of cyberspace to mitigate the impact of AI-generated misinformation on cyberspace.

[0115] In a propagation scenario with a large model of debunking users, when the threat level of AI-generated misinformation is low, the debunking user node is adjacent to the initial user node spreading the misinformation, and the initial user node spreading the misinformation is a super node, such as... Figure 21 As shown, because the threat level of AI-generated misinformation is relatively low, its spread is very difficult, making it easy for debunking user nodes to detect and eliminate. However, when the threat level of AI-generated misinformation increases slightly, it will spread along certain paths, such as... Figure 22 As shown, although the impact of AI-generated misinformation is relatively limited at this point, debunking user nodes still need to disseminate AI-generated misinformation on a large scale to eliminate its influence on cyberspace. Therefore, debunking user nodes need to be positioned in important locations within cyberspace as much as possible. This will allow them to detect the spread of AI-generated misinformation more quickly, while also ensuring that accurate information spreads rapidly throughout cyberspace.

[0116] In a propagation scenario with a large-scale user model for debunking misinformation, when the threat level of AI-generated misinformation is low and the debunking user nodes are far apart from the initial user nodes spreading the misinformation, such as... Figure 23 As shown, because AI-generated misinformation poses a relatively low threat and has a limited impact on cyberspace, it is only possible to detect and disseminate accurate information when the debunking user node is located in the vicinity of the AI-generated misinformation's propagation path. Otherwise, as... Figure 24 As shown, debunking user nodes cannot detect the spread of AI-related misinformation, making it impossible to eliminate such misinformation in cyberspace. When the threat level of AI-related misinformation is low, its spread may be discretely distributed throughout cyberspace, making it difficult for debunking user nodes to detect its occurrence. Therefore, it is necessary to set up debunking user nodes in key locations within cyberspace to more comprehensively prevent the spread of AI-related misinformation.

[0117] It's important to note that key nodes in cyberspace, especially supernodes, need to enhance their ability to identify AI-generated misinformation, and debunking user nodes should be deployed around them to facilitate the dissemination of accurate information. Furthermore, the perceived low threat level of AI-generated misinformation is not necessarily a good thing. Because this misinformation is scattered throughout cyberspace and lacks clustering characteristics, it is difficult to detect, potentially allowing it to persist in cyberspace indefinitely. To address this, a more specific and targeted study of cyberspace topology is needed, deploying a certain number of debunking user nodes in different locations within cyberspace to completely eliminate the impact of AI-generated misinformation on cyberspace.

[0118] Secondly, this invention provides an AI-based system for analyzing the spread of misinformation on a large-scale network, comprising a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of any of the methods described above. The system's technical solution is consistent with the methods described above, and will not be elaborated further here.

[0119] It should be noted that, for the sake of simplicity, the foregoing embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, and some steps may be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0120] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed methods or systems can be implemented in other ways in the several embodiments provided in this application. For example, the embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0122] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A large model-based AI false information network propagation scene analysis method, characterized in that, The method comprises: generating an AI information set; the AI information set comprises multiple groups of AI information, each group of AI information comprising an AI false information generated by AI and a true information corresponding to the AI false information; randomly extracting at least one group of AI information in the AI information set multiple times to obtain multiple subsets, and separately training a large model using the multiple subsets as training sets, one subset corresponding to generating one user large model, thereby generating multiple different user large models; randomly placing different user large models on user nodes in a network structure for simulation, and placing an AI false information in any user node in the network structure, the AI false information propagating to surrounding user nodes along edges in the network structure, and simulating different user individuals by using corresponding user large models on each user node, and realizing analysis of AI false information in a network propagation scenario by analyzing state transitions of each user node in the network structure; The method further comprises: marking the threat level of the AI false information in each group of AI information, and dividing the AI information into different threat level data sets; the threat level data sets include threat level data sets arranged from low to high: a first threat level data set, a second threat level data set, and a third threat level data set; randomly extracting at least one group of AI information from the same or different threat level data sets as a training set, comprising: randomly extracting at least one group of AI information from the first threat level data set multiple times to obtain multiple subsets, corresponding to generating multiple different ordinary user large models; and randomly extracting at least one group of AI information from the first threat level data set, the second threat level data set, and the third threat level data set multiple times to obtain multiple subsets, corresponding to generating multiple different rumor-busting user large models.

2. The large model-based AI false information network propagation scenario analysis method according to claim 1, characterized in that, The state types of the user nodes in the network structure include: ordinary user nodes, rumor-busting user nodes, false information spreading user nodes, and deceived user nodes.

3. The method according to claim 1, wherein, The method for threat level marking comprises one or more of TER, BLEURT, BARTScore, or manual marking.

4. The method according to claim 1, wherein, The multiple subsets are separately used as training sets to individually train the large model using a LoRA fine-tuning method.

5. The method according to claim 1, wherein, The method for generating the AI false information comprises: obtaining network false information and corresponding true information from a scientific rumor-busting website, using a Qwen large model, inputting a false theme according to a prompt word framework, and generating AI false information corresponding to the false theme.

6. The large model-based AI false information network propagation scenario analysis method according to claim 5, characterized in that, The prompt word framework comprises an ICIO framework or a CRISPE framework.

7. A large model-based AI false information network propagation scene analysis system, comprising a memory, a processor and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Social network false information propagation detection method based on motif degree

    CN112819645A

  • Image detection model generation method, detection method, terminal and storage medium

    CN113807281A