Rumor detection method and system based on prompt learning and meta-learning in low-resource environment

By adopting a method based on prompt learning and meta-learning in a low-resource environment, using a two-way long and short-term memory network, multi-layer perceptron and mask language model, the rumor detection task is transformed into a mask prediction task, which solves the problem of difficult to improve the accuracy and accuracy of rumors detection in a low-resource environment, and realizes the efficient adaptation and accuracy of the detection model under limited labeled data.

CN119961456AActive Publication Date: 2025-05-09湖南工商大学
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510446008.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In a low-resource environment, there are fewer data samples for rumor detection, which makes it difficult to improve detection accuracy and accuracy.

Method used

Using a method based on prompt learning and meta-learning, by obtaining soft prompt parameters and detection models, a two-way long and short-term memory network, a multi-layer perceptron and mask language model is used to convert the rumor detection task into a mask prediction task, and feature learning and task adaptation are performed.

Benefits of technology

Under the limited labeled data conditions, the adaptability and detection accuracy of the detection model to rumor detection tasks are improved, and the detection accuracy and accuracy problems in low-resource environments are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961456A_ABST
    Figure CN119961456A_ABST
Patent Text Reader

Abstract

The invention discloses a rumor detection method and system based on prompt learning and meta-learning in a low-resource environment. The method comprises the steps of obtaining a to-be-detected input text, a soft prompt parameter for guiding a rumor detection task and a detection model; inputting the input text and the soft prompt parameter into the detection model to obtain an output result of the detection model; the detection model comprises a bidirectional long-short-term memory network used for carrying out context information modeling on the soft prompt parameters; the two-layer multi-layer perceptron is used for adding nonlinear characteristics to the soft prompt parameters after context modeling; the mask language model is used for executing a mask prediction task on a prompt text formed by the soft prompt parameters added with the nonlinear features and the input text and predicting the soft prompt parameters in the prompt text; the prototype expressor is used for mapping a prediction result of the mask language model into a corresponding category label; and determining a rumor detection result according to the output result. According to the scheme, the rumor detection precision in a low-resource environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rumor detection, and in particular to a rumor detection method and system based on prompt learning and meta-learning in a low-resource environment. Background Art

[0002] Rumor detection is crucial in today's era of information explosion. By identifying and curbing the spread of false information, it can effectively prevent social panic, public crises and collapse of trust, especially in emergencies (such as epidemics and disasters), to ensure the accuracy of key information; at the same time, it maintains the order of the digital space, reduces economic losses and social divisions caused by rumors, and promotes information equity through technical means (such as multilingual models, low-resource algorithms), building a defense barrier against rumors for different regions and cultural groups, and is the cornerstone of digital social governance and security.

[0003] However, for some specific fields, there are few data samples for rumor detection in low-resource environments. How to improve the detection accuracy and precision of rumors in low-resource environments is a technical problem that needs to be solved urgently. Summary of the invention

[0004] In order to solve the above technical problems, an embodiment of the present invention provides a rumor detection method and system based on prompt learning and meta-learning in a low-resource environment.

[0005] The technical solution of the embodiment of the present invention is achieved as follows: An embodiment of the present invention provides a rumor detection method based on prompt learning and meta-learning in a low-resource environment, the method comprising: Obtaining input text to be detected, soft prompt parameters and a detection model; the soft prompt parameters are used to guide the rumor detection task; Input the input text and the soft prompt parameters into the detection model to obtain the output result of the detection model; wherein the detection model includes a prompt model and a prototype expressor; the prompt model includes a bidirectional long short-term memory network, a two-layer multi-layer perceptron and a mask language model; the bidirectional long short-term memory network is used to model the context information of the soft prompt parameters; the two-layer multi-layer perceptron is used to add nonlinear features to the soft prompt parameters after context modeling; the mask language model is used to perform a mask prediction task on the prompt text composed of the soft prompt parameters with added nonlinear features and the input text, and predict the soft prompt parameters in the prompt text; the prototype expressor is used to map the prediction results of the mask language model to corresponding category labels; Determine the rumor detection result corresponding to the input text according to the category label output by the detection model.

[0006] In one embodiment, the prompt text is: in, For the prompt text, is a special token for the masked language model, x is the input text, is a fixed anchor mark. represents a trainable vector filled with soft hint parameters, Indicates a sentence separator or end-of-sequence marker.

[0007] In one embodiment, obtaining soft prompt parameters and a detection model includes: Randomly draw samples from the source data multiple times to construct the meta-prompt task set; For each meta-prompt task in the meta-prompt task set, obtaining a corresponding task dataset, wherein the task dataset includes a support set and a query set; Acquire an initial prompt model, and use the support set to perform a first-order gradient multi-step parameter update on initial soft prompt parameters and initial mask language model parameters in the initial prompt model to obtain updated first soft prompt parameters and first mask language model parameters; Inputting the first soft prompt parameter and the first mask language model parameter into the initial prompt model, and performing a second-order gradient multi-step parameter update on the first soft prompt parameter and the first mask language model parameter in the input initial prompt model using a query set to obtain updated second soft prompt parameters and second mask language model parameters; Inputting the second mask language model parameters into the initial prompt model, and obtaining a detection model according to the input initial prompt model and the prototype expressor; The second soft prompt parameter is used as the final soft prompt parameter.

[0008] In one embodiment, using the support set to perform first-order gradient multi-step parameter update on initial soft prompt parameters and initial mask language model parameters in the initial prompt model includes: The following calculation formula is used to perform first-order gradient multi-step parameter update: in, For the The masked language model parameters of the step, For the Soft hint parameters for the step, For the -1 step masked language model parameters; For the - Soft hint parameters for step 1, is the learning rate when updating the model parameters on the support set, symbol represents the gradient operation, is the loss function, Meta-prompt task The support set, To enter -1 step masked language model parameters and Hint model for soft hint parameters of -1 step.

[0009] In one embodiment, a second-order gradient multi-step parameter update is performed on a first soft prompt parameter and a first mask language model parameter in an input initial prompt model using a query set, including: The following calculation formula is used to perform second-order gradient multi-step parameter update: in, For soft prompt parameters, arrow Indicates assignment or update of value. is the learning rate when updating the model parameters on the query set, Soft prompt parameter The second derivative of represents the gradient operation, is the loss function, Meta-prompt task The query set, To prompt the model, Meta-cue task The corresponding soft prompt parameters, is the identity matrix, is the learning rate when updating the model parameters on the support set, Represents the parameters on the query set The gradient of the loss function is represents the loss function on the support set, Represents the query set about parameters The gradient of the loss function is Express Seeking guidance, Indicates the parameters Compute the Hessian matrix.

[0010] In one embodiment, the stop update condition of the second-order gradient multi-step parameter update is: in, Represent the model parameters θ and soft prompt parameters The optimal value of is the loss function, Meta prompt task The query set, Prompt model.

[0011] In one embodiment, mapping the prediction result of the masked language model to the corresponding category label includes: For each category label, construct the corresponding category prototype; Encoding the prediction result of the masked language model into a feature representation, and calculating the similarity between the feature representation and the category prototype; Determine a category label corresponding to the prediction result of the masked language model according to the similarity.

[0012] In one embodiment, for each category label, a corresponding category prototype is constructed, including: The initial prototype is obtained by calculating the embedding average of all sample sets for each category label in the support set: in, represents the initial prototype, Indicates that the support concentration category label is A sample set of Indicates that the category label is No. samples; Calculate the attention weight of each sample according to the initial prototype; in, represents the Euclidean distance; Indicates that the category label is No. samples; According to the attention weight of each sample, the optimized category prototype of each sample is calculated using the following formula: in, Represents a prototype.

[0013] In one embodiment, calculating the similarity between the feature representation and the category prototype includes: The similarity between the feature representation and the category prototype is calculated using the following formula: in, Represents a query sample Belongs to category The probability score of represents the encoded feature representation of the prediction result of the masked language model, represents the set of all category labels, Expressed in natural logarithm The exponential function with base , It is the cosine similarity calculation formula.

[0014] An embodiment of the present invention also provides a rumor detection system based on prompt learning and meta-learning in a low-resource environment, comprising: a processor and a memory for storing a computer program that can be run on the processor; wherein, when the processor is used to run the computer program, it executes the steps of the above-mentioned method.

[0015] The embodiments of the present invention have the following beneficial effects: The embodiment of the present invention utilizes prompt text with soft prompt parameters and a mask language model to convert the rumor detection task into a mask prediction task, so that the detection model can perform effective feature learning and task adaptation with the help of the mask language model in the detection model under the condition of limited labeled data, thereby improving the adaptability and detection accuracy of the detection model to the rumor detection task in the absence of a large amount of labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a rumor detection method based on prompt learning and meta-learning in a low-resource environment according to an embodiment of the present invention; Figure 2 1 is a diagram showing the internal structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments.

[0018] The embodiment of the present invention provides a rumor detection method based on prompt learning and meta-learning in a low-resource environment. Figure 1 As shown, the method includes: Step 101: obtaining input text to be detected, soft prompt parameters and a detection model; the soft prompt parameters are used to guide the rumor detection task; Step 102: input the input text and the soft prompt parameters into the detection model to obtain the output result of the detection model; wherein the detection model includes a prompt model and a prototype expressor; the prompt model includes a bidirectional long short-term memory network, a two-layer multilayer perceptron and a mask language model; the bidirectional long short-term memory network is used to model the context information of the soft prompt parameters; the two-layer multilayer perceptron is used to add nonlinear features to the soft prompt parameters after context modeling; the mask language model is used to perform a mask prediction task on the prompt text composed of the soft prompt parameters with added nonlinear features and the input text, and predict the soft prompt parameters in the prompt text; the prototype expressor is used to map the prediction results of the mask language model to corresponding category labels; Step 103: Determine the rumor detection result corresponding to the input text according to the category label output by the detection model.

[0019] The method in this embodiment combines prompt learning and meta-learning to solve the low resource problem faced by social media rumor detection.

[0020] Specifically, in the prompt learning part, this embodiment converts the rumor detection task into a pre-training task (such as masked language modeling) by designing a combination of soft and hard templates. This method constructs a task-specific prompt template so that the detection model can perform effective feature learning and task adaptation with the help of the masked language model in the detection model under limited annotated data conditions.

[0021] In the meta-learning part, this embodiment proposes a meta-prompt learning framework, which aims to improve the adaptability of the detection model to the rumor detection task in the absence of a large amount of labeled data through effective knowledge transfer and meta-learning mechanism (meta-learning includes meta-training and meta-testing).

[0022] Specifically, the method of this embodiment includes the following contents: Step 1: Design a combination of soft and hard prompt template 1.1 Template construction This embodiment proposes a soft prompt template with an anchor mark, and the prompt mark is composed of a soft prompt mark (indicating a soft prompt parameter ) and an anchor tag (which represents an embedding fixed to a specific word). The combination of hard and soft hint templates makes the model more flexible while retaining enough semantic information to trigger the masked language model to generate correct predictions.

[0023] Taking the news headline classification task and a commonly used masked language model BERT as an example, the text prompted by the soft prompt template is as follows: in, is the prompt text, [CLS] is a special marker for the mask language model, This post is a fixed anchor mark for input text, and the filling of the [MASK] position here is not completed by a fixed vocabulary, but by a trainable vector (soft prompt parameter ) to indicate the guidance of the task, and [SEP] indicates a sentence separator or sequence end marker.

[0024] Step 2: Use the mask hint model to predict tags and convert the rumor detection task into a pre-training task.

[0025] 2.1 Overall architecture of the detection model The detection model in this embodiment includes a prompt model and prototype expressors, where the prompt model Including: Bidirectional Long Short-Term Memory Network (BiLSTM), two-layer Multilayer Perceptron (MLP), Masked Language Model (BERT).

[0026] Prompt Model The learnable parameters of BERT are divided into two parts: the masked language model BERT parameters and soft prompt parameters (Parameters used to generate soft hint embeddings).

[0027] To enhance the hint representation, the soft hint parameter A two-layer bidirectional long short-term memory network (BiLSTM) is first used to model context information and capture the sequence dependency within the prompt vector. Then, the output of the BiLSTM is sent to a two-layer multi-layer perceptron (MLP) to further process and give the prompt vector richer nonlinear features, thereby obtaining an enhanced prompt vector. This transformation enables it to smoothly escape from the local minimum. Next, With the original input text Splice to form a complete prompt text , for low-resource rumor detection. Input to a masked language model (such as BERT), which will Perform language modeling tasks on it, such as predicting the filler word at the [MASK] position.

[0028] The prototype expressor is used to map the model output to specific category labels.

[0029] 2.2 Post Encoder In prompt learning, the masked language model (MLM) is based on the prompt text The predicted [MASK] placeholder is propagated layer by layer through the self-attention mechanism. Due to the soft hint parameter Contained in the prompt text In the process, the self-attention mechanism propagates layer by layer, so that the final [MASK] combines the information embedded in the input text and the anchor mark (hard prompt) and soft prompt. Therefore, the hidden state of the [MASK] mark in the last layer of the masked language model BERT is As input text Then, the final feature representation is obtained through linear transformation . Using a masked language model, the rumor detection task can be converted into a pre-training task (masked language modeling).

[0030] in, is the weight matrix.

[0031] This embodiment uses the idea of ​​prompt learning and designs a template to convert the classification task (i.e., judging whether a given text is a rumor or not) into a fill-in-the-blank task (such as predicting [MASK]). In this way, the classification task is aligned with the pre-training task, thereby narrowing the gap between the tasks and further stimulating the potential of the masked language model (BERT).

[0032] Since hint-based methods, especially those adopting soft hints, are very sensitive to parameter initialization, an optimization-based meta-learning approach is introduced to the hint method to find better initialization points for the hint model and further explore its capability in low-resource scenarios.

[0033] In the following steps 3 and 4, we will first introduce the hint tuning objective, then describe how to construct the meta-hint task, and finally, explain the parameter update strategy in detail.

[0034] Step 3: Set the hint tuning target (i.e. the tuning target for the meta-training task).

[0035] For any low-resource task given in the source data , get the task dataset corresponding to the task Represents the training sample. Tip: The optimization objective can be formulated as follows (principle: minimizing the loss function is equivalent to maximizing the conditional probability predicted by the model): in, To indicate the tuning target, is the loss function, is the prompt model, is the conditional probability, which is determined by the masked language model parameters and soft prompt parameters composition.

[0036] In low-resource scenarios, Contains little labeled data, so the model parameters and soft prompt parameters The initialization of is crucial to the performance of the model.

[0037] Step 4: Based on the prompt tuning objective, a low-resource meta-prompt task set is constructed to meta-learn the prompt model in order to find a better initialization point for the prompt-based model.

[0038] 4.1 Constructing the Meta-prompt Task In order to find a better initialization point, this embodiment uses posts from source data (crawled from social media) Randomly draw samples to construct the meta-prompt task set , where N is the number of meta-prompt tasks. For each meta-prompt task , the meta-prompt task Corresponding task dataset Further divided into two disjoint subsets: support set and queryset .

[0039] That is, each meta-prompt task The sampling is: In the low-resource setting, the meta-training task and the meta-testing task are sampled from different domains to prevent the model from simply memorizing the training samples.

[0040] 4.2 Applying Meta-Learning to Hint Models Meta-learning includes meta-training and meta-testing. Meta-training is used to perform first-order gradient multi-step parameter updates on the support set to determine the first parameters of the prompt model. and The meta-test is the first parameter and Based on this, a second-order gradient multi-step parameter update is performed, and the prompt tuning target determined above is used as the stop update condition, and the final parameter obtained by stopping the update is used as the optimal initial parameter and .

[0041] Meta-training In the meta-training part, this embodiment uses the first-order gradient to update the parameters on the support set.

[0042] For each given meta-prompt task , clone the parameters of the related model that has been trained and is similar to the meta-prompt task (by cloning some or all of the parameters, the model can be trained faster on a new but related task instead of training from scratch. This method is particularly suitable for situations where the amount of data for the target task is small), and obtain the initial parameters of the prompt model for this meta-training and . Reuse support set Update initial parameters and During the meta-training process, the initial parameters and There will be multiple steps of gradient parameter update, each step is updated as follows: in, For The learning rate when updating model parameters, symbol represents the gradient operation, Indicates the number of inner steps.

[0043] Here, the stopping condition for parameter updating can comprehensively consider the fixed step limit (most common), gradient change, loss reduction, and computing resource constraints, and can be adjusted according to the experimental situation.

[0044] The final parameters obtained after stopping the update are used as the initial parameters of the meta-test for subsequent meta-tests.

[0045] Meta testing In the meta-test part, this embodiment uses the second-order gradient optimization to find the best parameters on the query set.

[0046] make To hint the model in the queryset The learning rate when updating model parameters, H is the Hessian matrix. In the query set The second-order gradient calculated above is formalized as follows: in, Soft prompt parameter The second-order derivative of Meta-cue task Corresponding soft prompt parameter, arrow Indicates assignment / update value, Represents the parameters on the query set The gradient of the loss function is is the identity matrix, represents the inner loop learning rate on the support set, represents the loss function on the support set, Prompt the model in the query set during meta-training The learning rate when updating model parameters, Represents the query set about parameters The gradient of the loss function is Express Seeking guidance, Indicates the parameters Compute the Hessian matrix.

[0047] In the meta-test part, the prompt tuning target mentioned above is used as the stop update condition of the meta-test, that is, the update is stopped when the prompt tuning target is reached, and the parameters of the last step are obtained. and As the optimal initial parameters of the prompt model.

[0048] That is, stop updating the target: in, Represent the model parameters and soft hint embed The optimal value of .

[0049] Step 5: Build a prototype expressor to create a category prototype 5.1 Prototype Construction Strategy The Prototypical Verbalizer is the bridge between the output of the prompt model and the final task label. It is responsible for mapping the results predicted by the prompt model into specific category labels.

[0050] This embodiment proposes a method of using a prototype expressor. The prototypes of the corresponding categories are established using training data, and two contrastive learning losses are designed to guide the optimization process. Considering the influence of noise in the scenario of limited training data, a noise-resistant prototype construction strategy is also designed to automatically assign weights to samples to obtain more effective prototypes.

[0051] This embodiment decides to introduce the prototype into the language generator to solve the overfitting problem when the training data is limited, because the prototype network can alleviate the overfitting problem to a certain extent through a metric-based method.

[0052] The basic unit of meta-learning is meta-task. For each meta-task, we construct corresponding category prototypes for the two categories (rumors and non-rumors) of its support set. ,in .

[0053] The category prototype is a conceptual unit that summarizes the overall information of the category. Previous prototype network work mainly calculates the prototype by averaging the embeddings of all samples of a category. In scenarios with limited samples, noisy data will seriously affect the accuracy of the prototype. Therefore, this embodiment proposes a noise-resistant prototype construction strategy that automatically assigns different weights to each sample as an attention mechanism to improve the robustness of the model.

[0054] (1) First, we obtain the prototype by averaging the embeddings of posts in each category: in, Supports centralized categories A sample set of Indicates that it belongs to the category label The first (rumor or not) Sample posts.

[0055] (2) Calculate sample weight (attention allocation) Before building a noise-resistant prototype, we first calculate the weight of each sample , that is, the contribution of the sample to the final prototype. Uncertainty weighting based on distance. For a sample , if its distance from other similar samples is large, it means it may be a noise sample.

[0056] in, represents the Euclidean distance, Indicates that it belongs to the category label The first (rumor or not) Sample posts, Indicates that it belongs to the category label The first (rumor or not) Sample posts, Represents the aforementioned prototype; in, Supports centralized categories A sample set of Indicates that it belongs to the category label The first (rumor or not) The prior art directly calculates by averaging, which contains noise, while the calculation method proposed in this embodiment has the effect of resisting noise.

[0057] 5.2 Contrastive Learning Loss Guided Prototyping In order to better optimize the prototype construction, this embodiment extracts low-level and high-level information from the relationship between posts and between posts and prototypes to assist model optimization. Specifically, this embodiment implements a two-level contrastive learning loss to guide parameter optimization: (1) Instance-level contrastive learning: Posts belonging to the same category should have higher similarity. This example optimizes the objective by minimizing the following loss function: in, Represents the set of posts with the same label on the support set , Represents the set of posts with different labels on the support set , is a temperature hyperparameter used to control the model's discrimination against negative samples. is the cosine similarity. This loss function aims to increase the similarity between similar posts, reduce the similarity between heterogeneous posts, and ensure local smoothness, thereby assisting prototype construction.

[0058] (2) Prototype-level contrastive learning: Prototype Should be labeled The posts with higher similarity have higher similarity. This embodiment optimizes the objective by minimizing the following loss function: in, represents the support set, Indicates support for centralized labeling A sample set of Get all possible category indices (including the correct category and other categories), represents the true category to which the sample belongs, It's a prototype The estimated concentration ratio is The larger it is, the smaller the concentration is. represents the estimated value of concentration, represents the set of all categories, Represents the prototype of a class.

[0059] This embodiment estimates the : in, To support centralized labeling A collection of samples.

[0060] The centrality estimate will adjust the weights used in the similarity calculations between the post vector and the prototype. ) will reduce the similarity, making the embedding Closer to the original, and smaller Can avoid embedding Too close to the prototype. The model can be prompted to generate a prototype with a more balanced concentration. In addition, according to the model verification formula, this embodiment can find that when =1 (derived from the formula at the prototype level), The essence of is the cross entropy loss function. Compared with the contrast loss between instances, it can utilize prototype information, and this high-level information usually contains more extensive and stable semantic information. This parameter-free prototype classifier usually performs better than the linear classifier with parameters in the few-sample scenario. Finally, this embodiment combines these two losses to obtain the final optimization goal: Step 6: The prompt model encodes the new post into a feature representation and calculates its similarity with the category prototype to obtain a probability score.

[0061] 6.1 Reasoning In the inference stage, this embodiment obtains probability scores through a metric-based method. For each category (rumor and non-rumor), a corresponding category prototype is constructed during the training process. For a new post, the prompt model first encodes it into a feature representation and calculates the similarity between the post and each category prototype. The similarity is calculated using the cosine similarity metric. By calculating the similarity between the input post and each category prototype, the prototype expressor can evaluate the possibility that the post belongs to different categories. For example, if the representation of the post is closer to the prototype of the "rumor" category, the post is classified as a rumor; otherwise, it is classified as a non-rumor.

[0062] For query sample ,category The probability score is: This is the feature representation of the output of the masked language model post encoder , Represents the prototype mentioned above, Get all possible category indices (including the correct category and other categories), Represents the set of all categories.

[0063] Then make a prediction using the argmax function: Under the masked language model MLM, given the input ,category The posterior probability (PosteriorProbability).

[0064] In the inference stage, this embodiment proposes a probability scoring method based on metric learning. By calculating the similarity between the input sample and the prototype vector, the model can output the probability score of each category, thereby achieving accurate detection of rumors.

[0065] The embodiment of the present invention utilizes prompt text with soft prompt parameters and a mask language model to convert the rumor detection task into a mask prediction task, so that the detection model can perform effective feature learning and task adaptation with the help of the mask language model in the detection model under the condition of limited labeled data, thereby improving the adaptability and detection accuracy of the detection model to the rumor detection task in the absence of a large amount of labeled data.

[0066] In order to implement the method of an embodiment of the present invention, an embodiment of the present invention also provides a rumor detection system based on prompt learning and meta-learning in a low-resource environment, including: a processor and a memory for storing a computer program that can be run on the processor; wherein, when the processor is used to run the computer program, the steps of the above-mentioned method are executed.

[0067] The above-mentioned system provided in this embodiment belongs to the same concept as the above-mentioned method embodiment. The specific implementation process thereof is detailed in the method embodiment and will not be repeated here.

[0068] In order to implement the method of the embodiment of the present invention, the embodiment of the present invention also provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the above method.

[0069] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present invention, the embodiment of the present invention further provides an electronic device (computer device). Specifically, in one embodiment, the computer device may be a terminal, and its internal structure diagram may be as follows: Figure 2As shown. The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05 and a memory (not shown in the figure) connected through a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A06. The network interface A02 of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor A01, the method of any one of the above embodiments is implemented. The display screen A04 of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device A05 of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0070] Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0071] The device provided by the embodiment of the present invention includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the method of any one of the above embodiments is implemented.

[0072] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0073] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0074] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0075] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0076] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0077] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0078] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0079] It can be understood that the memory of the embodiment of the present invention can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.

[0080] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0081] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A rumor detection method based on prompt learning and meta-learning in a low-resource environment, characterized in that: The method comprises: Obtaining input text to be detected, soft prompt parameters and a detection model; the soft prompt parameters are used to guide the rumor detection task; Input the input text and the soft prompt parameters into the detection model to obtain the output result of the detection model; wherein the detection model includes a prompt model and a prototype expressor; the prompt model includes a bidirectional long short-term memory network, a two-layer multilayer perceptron and a mask language model; the bidirectional long short-term memory network is used to model the context information of the soft prompt parameters; the two-layer multilayer perceptron is used to add nonlinear features to the soft prompt parameters after context modeling; the mask language model is used to perform a mask prediction task on the prompt text composed of the soft prompt parameters with added nonlinear features and the input text, and predict the soft prompt parameters in the prompt text; the prototype expressor is used to map the prediction results of the mask language model to corresponding category labels; Determine the rumor detection result corresponding to the input text according to the category label output by the detection model.

2. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 1 is characterized in that: The prompt text is: in, is the prompt text, [CLS] is a special marker for the mask language model, To input text, is a fixed anchor mark. represents a trainable vector filled with soft hint parameters, Indicates a sentence separator or end-of-sequence marker.

3. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 1 is characterized in that: Get soft prompt parameters and detection model, including: Randomly draw samples from the source data multiple times to construct the meta-prompt task set; For each meta-prompt task in the meta-prompt task set, obtaining a corresponding task dataset, wherein the task dataset includes a support set and a query set; Acquire an initial prompt model, and use the support set to perform a first-order gradient multi-step parameter update on initial soft prompt parameters and initial mask language model parameters in the initial prompt model to obtain updated first soft prompt parameters and first mask language model parameters; Inputting the first soft prompt parameter and the first mask language model parameter into the initial prompt model, and performing a second-order gradient multi-step parameter update on the first soft prompt parameter and the first mask language model parameter in the input initial prompt model using a query set to obtain updated second soft prompt parameters and second mask language model parameters; Inputting the second mask language model parameters into the initial prompt model, and obtaining a detection model according to the input initial prompt model and the prototype expressor; The second soft prompt parameter is used as the final soft prompt parameter.

4. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 3 is characterized in that: Performing a first-order gradient multi-step parameter update on initial soft prompt parameters and initial mask language model parameters in the initial prompt model using the support set, including: The following calculation formula is used to perform first-order gradient multi-step parameter update: in, For the The masked language model parameters of the step, For the Soft hint parameters for the step, For the -1 step masked language model parameters; For the -1 step soft hint parameter, is the learning rate when updating the model parameters on the support set, symbol represents the gradient operation, is the loss function, Meta-prompt task The support set, To enter -1 step masked language model parameters and Hint model for soft hint parameters of -1 step.

5. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 3 is characterized in that: The query set is used to perform a second-order gradient multi-step parameter update on the first soft prompt parameter and the first mask language model parameter in the input initial prompt model, including: The following calculation formula is used to perform second-order gradient multi-step parameter update: in, For soft prompt parameters, arrow Indicates assignment or update of value. is the learning rate when updating the model parameters on the query set, Soft prompt parameter The second derivative of represents the gradient operation, is the loss function, Meta prompt task The query set, To prompt the model, Meta-cue task The corresponding soft prompt parameters, is the identity matrix, is the learning rate when updating the model parameters on the support set, Represents the parameters on the query set The gradient of the loss function is represents the loss function on the support set, Represents the query set about parameters The gradient of the loss function is Express Seeking guidance, Indicates the parameters Compute the Hessian matrix.

6. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 5 is characterized in that: The stop update condition for the second-order gradient multi-step parameter update is: in, Represent the model parameters and soft prompt parameters The optimal value of is the loss function, Meta-prompt task The query set, Prompt model.

7. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 1 is characterized in that: Map the prediction results of the masked language model to the corresponding category labels, including: For each category label, construct the corresponding category prototype; Encoding the prediction result of the masked language model into a feature representation, and calculating the similarity between the feature representation and the category prototype; Determine a category label corresponding to the prediction result of the masked language model according to the similarity.

8. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 7 is characterized in that: For each category label, construct the corresponding category prototype, including: The initial prototype is obtained by calculating the embedding average of all sample sets for each category label in the support set: in, represents the initial prototype, Indicates that the support concentration category label is A sample set of Indicates that the category label is No. samples; Calculate the attention weight of each sample according to the initial prototype; in, represents the Euclidean distance; Indicates that the category label is No. samples; According to the attention weight of each sample, the optimized category prototype of each sample is calculated using the following formula: in, Represents a prototype.

9. The rumor detection method based on prompt learning and meta-learning in a low-resource environment according to claim 7, characterized in that: Calculating the similarity between the feature representation and the category prototype includes: The similarity between the feature representation and the category prototype is calculated using the following formula: in, Represents a query sample Belongs to category The probability score of represents the encoded feature representation of the prediction result of the masked language model, represents the set of all category labels, Expressed in natural logarithm The exponential function with base , It is the cosine similarity calculation formula.

10. A rumor detection system based on prompt learning and meta-learning in a low-resource environment, characterized in that: include: A processor and a memory for storing a computer program that can be run on the processor; wherein, when the processor is used to run the computer program, the steps of the method according to any one of claims 1 to 9 are executed.

Citation Information

Patent Citations

  • Pre-training language model processing method based on comparative learning and intelligent question answering system

    CN114528383A

  • Station detection method, device and system based on combination of prompt learning and external knowledge

    CN116662807A

  • Social media rumor detection method based on active learning iteration

    CN117992651A

  • Cross-language false news detection method based on prompt learning

    CN118350358A

  • Low-resource environment rumor detection method and system

    CN119204021A