A false news detection method and device based on a dynamic social background
Patent Information
- Application Number
- CN202311392732.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-25
AI Technical Summary
然而在实际场景中,每条新闻都存在于一个动态变化的社会背景中,而以往的虚假信息检测忽略了新闻的动态社会背景变化,导致虚假新闻检测的真实性较低
[0086]本发明实施例提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN117520813B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fake news detection technology, and in particular to a method and apparatus for detecting fake news based on dynamic social context. Background Technology
[0002] In recent years, work on fake news detection has primarily focused on detecting the authenticity of news articles containing text, images, and social network data, and has designed deep learning models including pre-trained language models and graph neural networks. Typically, previous research on fake news detection has only studied the correlation between news content and authenticity labels. However, in real-world scenarios, each news article exists within a dynamically changing social context, and previous fake news detection methods ignored these dynamic changes in social context, resulting in low accuracy in detecting fake news. Summary of the Invention
[0003] This invention provides a method and apparatus for detecting fake news based on a dynamic social context. The technical solution is as follows:
[0004] On the one hand, a method for detecting fake news based on a dynamic social context is provided. This method is implemented by an electronic device and includes:
[0005] S1. Construct an initial fake news detection model;
[0006] S2. Obtain the news training set; the news training set includes sample news texts, the real values corresponding to the sample news texts, and the publication time of the news.
[0007] S3. Obtain the initial news dynamic background representation matrix;
[0008] S4. Based on the news training set and the initial news dynamic background representation matrix, pre-train the initial fake news detection model to obtain a trained fake news detection model.
[0009] S5. Divide the news training set into multiple subsets according to the publication time of the sample news texts, and train the initial news dynamic background representation model to obtain the trained news dynamic background representation model.
[0010] S6. Obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model.
[0011] S7. Obtain the news item to be identified and its publication time;
[0012] S8. Input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified;
[0013] S9. Based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model, determine the authenticity of the news to be identified.
[0014] Optionally, the fake news detection model includes a text embedding layer, a feature extractor, and a classifier;
[0015] S4 pre-trains the initial fake news detection model based on the news training set and the initial news dynamic background representation matrix to obtain a trained fake news detection model, including:
[0016] S41. Input the sample news texts from the news training set into the text embedding layer to generate word embeddings;
[0017] S42. The generated word embeddings and the initial news dynamic background representation matrix are jointly input into the feature extractor to obtain joint features;
[0018] S43. Input the joint features into the classifier to obtain the predicted value of the authenticity of the sample news text;
[0019] S45. Using the binary cross-entropy loss function, construct the objective function of the initial fake news detection model, optimize the gap between the predicted and actual values of the sample news texts, and obtain the trained fake news detection model.
[0020] Optionally, S5 divides the news training set into multiple subsets according to the publication time of the sample news texts, and trains the initial news dynamic background representation model to obtain a trained news dynamic background representation model, including:
[0021] S51. Freeze the parameters of the text embedding layer, the parameters of the feature extractor, and the parameter matrix of the classifier;
[0022] S52. Divide the news training set into multiple subsets based on the publication time of the sample news texts;
[0023] S53. Train the initial news dynamic background representation model using the divided subsets to obtain the trained news dynamic background representation model.
[0024] Optionally, the time series model includes a discrete-time time series model or a continuous-time time series model.
[0025] Optionally, the training process for a discrete-time time series model is as follows:
[0026] The Discrete Long Short-Term Memory (LSTM) network was used as the initial discrete-time time series model.
[0027] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial discrete-time time series model is then trained to obtain the trained discrete-time time series model.
[0028] Optionally, the training process for a continuous-time time series model is as follows:
[0029] The true value of the news dynamic background representation is obtained based on the trained news dynamic background representation model, and the time-based dynamic differential equation of the true value of the news dynamic background representation is obtained.
[0030] The time-based dynamic differential equations are used as the initial continuous-time time series model;
[0031] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial continuous-time time series model is then trained to obtain the trained continuous-time time series model.
[0032] Optionally, S9 determines the authenticity of the news to be identified based on the news to be identified, its dynamic background representation, and a trained fake news detection model, including:
[0033] S91. Embed the news input text to be recognized into the layer and generate word embeddings;
[0034] S92. Input the publication time of the news to be identified into the time series model to obtain the dynamic background representation of the news to be identified;
[0035] S93. Input the word embeddings of the news to be identified and the obtained dynamic background representation of the news into the feature extractor to obtain joint features;
[0036] S94. Input the joint features into the classifier to obtain the authenticity result of the news to be identified.
[0037] On the other hand, a fake news detection device based on a dynamic social context is provided. This device is applied to a fake news detection method based on a dynamic social context, and the device includes:
[0038] Building blocks are used to construct the initial fake news detection model;
[0039] The first acquisition unit is used to acquire the news training set; the news training set includes sample news texts, the real values corresponding to the sample news texts, and the publication time of the news.
[0040] The second acquisition unit is used to acquire the initial news dynamic background representation matrix;
[0041] The third acquisition unit is used to acquire the news to be identified and the publication time of the news to be identified;
[0042] The first training unit is used to pre-train the initial fake news detection model based on the news training set and the initial news dynamic background representation matrix, so as to obtain a trained fake news detection model.
[0043] The second training unit is used to divide the news training set into multiple subsets according to the news release time, train the initial news dynamic background representation model, and obtain the trained news dynamic background representation model.
[0044] The third training unit is used to obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and to train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model.
[0045] The first prediction unit is used to input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified.
[0046] The second prediction unit is used to determine the authenticity of the news to be identified based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model.
[0047] Optionally, the fake news detection model includes a text embedding layer, a feature extractor, and a classifier;
[0048] The first training unit is used for:
[0049] Input sample news texts from the news training set into the text embedding layer to generate word embeddings;
[0050] The generated word embeddings are jointly input into the feature extractor along with the initial news dynamic background matrix to obtain joint features;
[0051] The joint features are input into the classifier to obtain the predicted authenticity value of the sample news text;
[0052] Using the binary cross-entropy loss function, the objective function of the fake news detection model is constructed, and the gap between the predicted and actual values of the sample news texts is optimized to obtain the trained fake news detection model.
[0053] The second training unit is used for:
[0054] Freeze the parameters of the text embedding layer, the parameters of the feature extractor, and the parameter matrix of the classifier;
[0055] The news training set is divided into multiple subsets based on the publication time of the news;
[0056] The initial news dynamic background representation model is trained by dividing the data into multiple subsets to obtain a trained news dynamic background representation model.
[0057] The third training unit is used for:
[0058] Time series models include discrete-time time series models and continuous-time time series models;
[0059] The training process for a discrete-time time series model is as follows:
[0060] The Discrete Long Short-Term Memory (LSTM) network was used as the initial discrete-time time series model.
[0061] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial discrete-time time series model is then trained to obtain the trained discrete-time time series model.
[0062] The training process for a continuous-time time series model is as follows:
[0063] The true value of the news dynamic background representation is obtained based on the trained news dynamic background representation model, and the time-based dynamic differential equation of the true value of the news dynamic background representation is obtained.
[0064] The time-based dynamic differential equations are used as the initial continuous-time time series model;
[0065] Based on the obtained news dynamic background representation of the true value, the initial continuous time series model is trained to obtain a trained continuous time series model.
[0066] Optionally, the second prediction unit is used for:
[0067] The text to be recognized is input into the text embedding layer to generate word embeddings;
[0068] By inputting the publication time of the news to be identified into the time series model, a dynamic background representation of the news to be identified is obtained.
[0069] The word embeddings of the news to be identified and the obtained dynamic background representation of the news are jointly input into the feature extractor to obtain joint features;
[0070] The joint features are input into the classifier to obtain the authenticity result of the news to be identified.
[0071] Optionally, the time series model includes a discrete-time time series model or a continuous-time time series model.
[0072] Optionally, the training process for a discrete-time time series model is as follows:
[0073] The Discrete Long Short-Term Memory (LSTM) network was used as the initial discrete-time time series model.
[0074] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial discrete-time time series model is then trained to obtain the trained discrete-time time series model.
[0075] Optionally, the training process for a continuous-time time series model is as follows:
[0076] The true value of the news dynamic background representation is obtained based on the trained news dynamic background representation model, and the time-based dynamic differential equation of the true value of the news dynamic background representation is obtained.
[0077] The time-based dynamic differential equations are used as the initial continuous-time time series model;
[0078] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial continuous-time time series model is then trained to obtain the trained continuous-time time series model.
[0079] Optionally, S9 determines the authenticity of the news to be identified based on the news to be identified, its dynamic background representation, and a trained fake news detection model, including:
[0080] S91. Embed the news input text to be recognized into the layer and generate word embeddings;
[0081] S92. Input the publication time of the news to be identified into the time series model to obtain the dynamic background representation of the news to be identified;
[0082] S93. Input the word embeddings of the news to be identified and the obtained dynamic background representation of the news into the feature extractor to obtain joint features;
[0083] S93. Input the joint features into the classifier to obtain the authenticity result of the news to be identified.
[0084] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned fake news detection method based on dynamic social context.
[0085] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, which is loaded and executed by a processor to implement the above-mentioned fake news detection method based on dynamic social context.
[0086] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0087] This invention first pre-trains an initial fake news detection model using the acquired news training set and an initial news dynamic background representation matrix, obtaining a trained initial fake news detection model. Considering that each news item exists in a dynamically changing social context in real-world scenarios, the initial news dynamic background representation model is further trained using the news release time and the news training set, obtaining a news dynamic background representation model. Addressing the temporal variation characteristic of fake news, the initial time series model is trained using the real values of the news dynamic background, obtaining a time series model. The fake news detection method proposed in this invention considers both social and temporal changes, achieving higher prediction accuracy compared to existing fake news detection methods. Attached Figure Description
[0088] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0089] Figure 1 This is a flowchart of a fake news detection method based on a dynamic social background provided by an embodiment of the present invention;
[0090] Figure 2 This is a schematic diagram of the structure of a fake news detection model provided in an embodiment of the present invention;
[0091] Figure 3 This is a block diagram of a fake news detection device based on a dynamic social background provided in an embodiment of the present invention;
[0092] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0093] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0094] This invention provides a method for detecting fake news based on a dynamic social context. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart shown is for a fake news detection method based on a dynamic social context. The processing flow of this method may include the following steps:
[0095] S1. Obtain the news training set.
[0096] The news training set can include sample news texts, the real values corresponding to the sample news texts, and the publication time of the news.
[0097] In one feasible implementation, the news training set can be obtained by web scraping.
[0098] S2. Obtain the initial news dynamic background representation matrix.
[0099] The initial news background representation matrix can be an initialized matrix, represented as follows:
[0100] S3. Construct an initial fake news detection model.
[0101] Optionally, the fake news detection model includes a text embedding layer, a feature extractor, and a classifier.
[0102] Among them, the text embedding layer can be selected from the embedding layer of the pre-trained language model BERT; the feature extractor can be any fake news feature extractor; and the classifier can be selected from a realism classifier based on a feedforward neural network.
[0103] S4. Based on the news training set and the initial news dynamic background representation matrix, pre-train the initial fake news detection model to obtain a trained fake news detection model.
[0104] Optionally, the specific operation of S4 may include the following steps S41-S44:
[0105] S41. Input the sample news texts from the news training set into the text embedding layer to generate word embeddings.
[0106] In one feasible implementation, any sample news article from the news training set D can be used. Input text embedding layer of pre-trained language model BERT Get the word embeddings of this news article
[0107] S42. The generated word embeddings and the initial news dynamic background representation matrix are jointly input into the feature extractor to obtain joint features.
[0108] S43. Input the joint features into the classifier to obtain the predicted value of the authenticity of the sample news text.
[0109] In one feasible implementation, the authenticity prediction value of the sample news text can be obtained by the following formula (1).
[0110]
[0111] Where [·; ·] represents the concatenation operation, F θ (·) represents a classifier. The matrix representing the initial news background is called e. i Representative word embedding, P i W represents the predicted value of the news text's authenticity. C Let i represent a realism classifier based on a feedforward neural network, where i can take the values 1, 2, 3, ..., n, and n represents the number of samples in the news training set.
[0112] S44. Using the binary cross-entropy loss function, construct the objective function of the fake news detection model, optimize the gap between the predicted and actual values of the sample news texts, and obtain the trained fake news detection model.
[0113] In one feasible implementation, the objective function of the initial fake news detection model can be constructed using the following formula (2).
[0114]
[0115] Among them, l BCE (·,·) denotes the classic binary cross-entropy loss function, y i This represents the true label of the news, where N represents the number of samples in the training set, and P represents the true label of the news. i denoted as the predicted value of the news text's authenticity, where i can take the values 1, 2, 3, ..., n, and n represents the number of samples in the news training set.
[0116] S5. Divide the news training set into multiple subsets according to the publication time of the sample news texts, and train the initial news dynamic background representation model to obtain the trained news dynamic background representation model.
[0117] Optionally, the specific operation of S5 may include the following steps S51-S53:
[0118] S51. Combine the parameters Π of the text embedding layer, the parameters θ of the feature extractor, and the parameter matrix W of the classifier. C freeze.
[0119] In one feasible implementation, the parameters of the text embedding layer, the feature extractor, and the classifier are trained through the above steps. To accelerate the training of the entire model, the parameters of the trained text embedding layer, the feature extractor, and the classifier can be frozen during the initial training of the news dynamic background representation model, so that the parameters of the trained text embedding layer, the feature extractor, and the classifier will not change.
[0120] S52. The news training set D is divided according to the publication time t of the sample news texts. i Dividing into multiple subsets is represented as
[0121] S53. Train the initial news dynamic background representation model using the divided subsets to obtain the trained news dynamic background representation model.
[0122] In one feasible implementation, a subset D of Γ time periods can be obtained using the following formula (3). Γ Authenticity prediction.
[0123]
[0124] Among them, z Γ The background information represents the news and events. D represents the news subset of the Γth time period. Γ The true predicted value, D represents the news subset of the Γth time period. Γ Word embedding, W C F represents a realism classifier based on a feedforward neural network. θ (·) represents the classifier, where i can take the values 1, 2, 3, ..., n, and n represents the number of samples in the news training set.
[0125] Based on step S5 above, the objective function of the fake news detection model is optimized in stages, which can be expressed by the following formula (4).
[0126]
[0127] Among them, l BCE (·,·) denotes the classic binary cross-entropy loss function. D represents the news subset of the Γth time period. Γ The true predicted value, D represents the news subset of the Γth time period. Γ The authenticity label, where i can be 1, 2, 3, ..., N. Γ N Γ Represents D Γ The number of training samples in the dataset.
[0128] S6. Obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model.
[0129] Among them, the time series model uses f Φ(·,t) represents the context. Time series models can be used to obtain a dynamic background representation of news events in future time periods.
[0130] Optionally, the time series model includes discrete-time or continuous-time models. Compared to discrete-time models, changes in the social context of news are often continuous and sudden. Therefore, a continuous-time model is trained to capture this dynamic. During testing, a continuous-time model can be preferred. The training processes for discrete-time and continuous-time models are described below:
[0131] 1. The training process for discrete-time time series models is as follows:
[0132] (1) The Discrete Long Short-Term Memory Network (LSTM) is used as the initial discrete-time time series model.
[0133] (2) Obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and train the initial discrete time series model to obtain the trained discrete time series model.
[0134] In one feasible implementation, the discrete time series model can be represented by the following formula (5).
[0135]
[0136] Among them, f Φ (·,t) represents the time series model. 0 ,z 1 ,...,z Γ-1 The} symbol represents the background of a news update, indicating the actual value.
[0137] Based on the time series models in steps S1-S5 and S6, the overall objective function of the entire fake news detection model can be expressed by the following formula (6):
[0138]
[0139] Where ||·|| represents Euclidean distance.
[0140] 2. The training process for a continuous-time time series model is as follows:
[0141] (1) Obtain the true value of the news dynamic background representation based on the trained news dynamic representation model, and obtain the time-based dynamic differential equation of the true value of the news dynamic background representation.
[0142] The time-based dynamic differential equation can be expressed as:
[0143] (2) The time-based dynamic differential equation is used as the initial continuous-time time series model.
[0144] (3) Obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and train the initial continuous time series model to obtain the trained continuous time series model.
[0145] In one feasible implementation, a continuous-time time series model can be represented by the following formula (7).
[0146]
[0147] Among them, z 0 The background information representing the news event indicates the true value, f. Φ (·,t) represents the time-based differential of the news dynamic background representation.
[0148] To optimize the parameter Φ, the parameter Φ of the continuous-time time series model can be optimized using the following formula (8).
[0149]
[0150] in, and This indicates that two feedforward neural networks will z Γ Encoded as a hidden layer representation And decode it back. Among them, for The dynamic equation can be solved using the following formula (9).
[0151]
[0152] Here, ODESOLVER stands for Differential Equation Solver.
[0153] S7. Obtain the news item to be identified and its publication time.
[0154] S8. Input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified.
[0155] The time series model can be any one of those in step S6.
[0156] In one feasible implementation, the news dynamic background representation Z of the news release time Γ to be identified can be obtained according to formula (7) in step S6. Γ Alternatively, according to formula (5) in step S6, the dynamic background representation Z of the news release time Γ to be identified can be obtained. Γ .
[0157] S9. Based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model, determine the authenticity of the news to be identified.
[0158] Optionally, the specific operation of S9 may include the following steps S91-S94:
[0159] S91. Embed the news input text to be recognized into the layer and generate word embeddings;
[0160] S92. Input the publication time of the news to be identified into the time series model to obtain the dynamic background representation of the news to be identified;
[0161] S93. Input the word embeddings of the news to be identified and the obtained dynamic background representation of the news into the feature extractor to obtain joint features;
[0162] S94. Input the joint features into the classifier to obtain the authenticity result of the news to be identified.
[0163] This invention first pre-trains an initial fake news detection model using the acquired news training set and an initial news dynamic background representation matrix, obtaining a trained initial fake news detection model. Considering that each news item exists in a dynamically changing social context in real-world scenarios, the initial news dynamic background representation model is further trained using the news release time and the news training set, obtaining a news dynamic background representation model. Addressing the temporal variation characteristic of fake news, the initial time series model is trained using the real values of the news dynamic background, obtaining a time series model. The fake news detection method proposed in this invention considers both social and temporal changes, achieving higher prediction accuracy compared to existing fake news detection methods.
[0164] Figure 3 This is a block diagram illustrating a fake news detection device based on a dynamic social context, according to an exemplary embodiment. The device is used in a fake news detection method based on a dynamic social context. (Refer to...) Figure 3 The device includes a construction unit 310, a first acquisition unit 320, a second acquisition unit 330, a third acquisition unit 340, a first training unit 350, a second training unit 360, a third training unit 370, a first prediction unit 380, and a second prediction unit 390, wherein:
[0165] Building unit 310 is used to build the initial fake news detection model;
[0166] The first acquisition unit 320 is used to acquire a news training set; the news training set includes sample news texts, the real values corresponding to the sample news texts, and the publication time of the news.
[0167] The second acquisition unit 330 is used to acquire the initial news dynamic background representation matrix;
[0168] The third acquisition unit 340 is used to acquire the news to be identified and the publication time of the news to be identified;
[0169] The first training unit 350 is used to pre-train the initial fake news detection model based on the news training set and the initial news dynamic background representation matrix to obtain a trained fake news detection model.
[0170] The second training unit 360 is used to divide the news training set into multiple subsets according to the news release time, train the initial news dynamic background representation model, and obtain the trained news dynamic background representation model.
[0171] The third training unit 370 is used to obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and to train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model.
[0172] The first prediction unit 380 is used to input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified.
[0173] The second prediction unit 390 is used to determine the authenticity of the news to be identified based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model.
[0174] Optionally, the fake news detection model includes a text embedding layer, a feature extractor, and a classifier;
[0175] Optionally, the first training unit 350 is used for:
[0176] Input sample news texts from the news training set into the text embedding layer to generate word embeddings;
[0177] The generated word embeddings are jointly input into the feature extractor along with the initial news dynamic background matrix to obtain joint features;
[0178] The joint features are input into the classifier to obtain the predicted authenticity value of the sample news text.
[0179] Using the binary cross-entropy loss function, the objective function of the fake news detection model is constructed, and the gap between the predicted and actual values of the sample news texts is optimized to obtain the trained fake news detection model.
[0180] Optionally, the second training unit 360 is used for:
[0181] Freeze the parameters of the text embedding layer, the parameters of the feature extractor, and the parameter matrix of the classifier;
[0182] The news training set is divided into multiple subsets based on the time of news release;
[0183] The initial news dynamic background representation model is trained by dividing the data into multiple subsets to obtain a trained news dynamic background representation model.
[0184] Optionally, the third training unit 370 is used for:
[0185] Time series models include discrete-time time series models and continuous-time time series models;
[0186] The training process for a discrete-time time series model is as follows:
[0187] The Discrete Long Short-Term Memory (LSTM) network was used as the initial discrete-time time series model.
[0188] The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial discrete-time time series model is then trained to obtain the trained discrete-time time series model.
[0189] The training process for a continuous-time time series model is as follows:
[0190] The true value of the news dynamic background representation is obtained based on the trained news dynamic background representation model, and the time-based dynamic differential equation of the true value of the news dynamic background representation is obtained.
[0191] The time-based dynamic differential equations are used as the initial continuous-time time series model;
[0192] Based on the obtained news dynamic background representation of the true value, the initial continuous time series model is trained to obtain a trained continuous time series model.
[0193] Optionally, the second prediction unit 390 is used for:
[0194] The text to be recognized is input into the text embedding layer to generate word embeddings;
[0195] By inputting the publication time of the news to be identified into the time series model, a dynamic background representation of the news to be identified is obtained.
[0196] The word embeddings of the news to be identified and the obtained dynamic background representation of the news are jointly input into the feature extractor to obtain joint features;
[0197] The joint features are input into the classifier to obtain the authenticity result of the news to be identified.
[0198] This invention first pre-trains an initial fake news detection model using the acquired news training set and an initial news dynamic background representation matrix, obtaining a trained initial fake news detection model. Considering that each news item exists in a dynamically changing social context in real-world scenarios, the initial news dynamic background representation model is further trained using the news release time and the news training set, obtaining a news dynamic background representation model. Addressing the temporal variation characteristic of fake news, the initial time series model is trained using the real values of the news dynamic background, obtaining a time series model. The fake news detection method proposed in this invention considers both social and temporal changes, achieving higher prediction accuracy compared to existing fake news detection methods.
[0199] Figure 4 This is a schematic diagram of the structure of an electronic device 400 provided in an embodiment of the present invention. The electronic device 400 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 401 and one or more memories 402. The memory 402 stores at least one instruction, which is loaded and executed by the processor 401 to implement the steps of the Chinese text spelling check method described above.
[0200] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the Chinese text spelling check method described above. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0201] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0202] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for detecting fake news based on dynamic social context, characterized in that the method... include: S1. Construct an initial fake news detection model; S2, Obtain the news training set; The news training set includes sample news texts, the corresponding real values of the sample news texts, and the publication time of the news. S3. Obtain the initial news dynamic background representation matrix; S4. Based on the news training set and the initial news dynamic background representation matrix, pre-train the initial fake news detection model to obtain a trained fake news detection model. S5. Divide the news training set into multiple subsets according to the publication time of the sample news texts, and train the initial news dynamic background representation model to obtain the trained news dynamic background representation model. S6. Obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model. S7. Obtain the news item to be identified and its publication time; S8. Input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified; S9. Based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model, determine the authenticity of the news to be identified.
2. The method according to claim 1, characterized in that, The fake news detection model includes a text embedding layer, a feature extractor, and a classifier; S4 pre-trains the initial fake news detection model based on the news training set and the initial news dynamic background representation matrix to obtain a trained fake news detection model, including: S41. Input the sample news texts from the news training set into the text embedding layer to generate word embeddings; S42. The generated word embeddings and the initial news dynamic background representation matrix are jointly input into the feature extractor to obtain joint features; S43. Input the joint features into the classifier to obtain the predicted value of the authenticity of the sample news text; S45. Using the binary cross-entropy loss function, construct the objective function of the initial fake news detection model, optimize the gap between the predicted and actual values of the sample news texts, and obtain the trained fake news detection model.
3. The method according to claim 1, characterized in that, S5 divides the news training set into multiple subsets according to the publication time of the sample news texts, and trains the initial news dynamic background representation model to obtain a trained news dynamic background representation model, including: S51. Freeze the parameters of the text embedding layer, the parameters of the feature extractor, and the parameter matrix of the classifier; S52. Divide the news training set into multiple subsets based on the publication time of the sample news texts; S53. Train the initial news dynamic background representation model using the divided subsets to obtain the trained news dynamic background representation model.
4. The method according to claim 1, characterized in that, Time series models include discrete-time time series models or continuous-time time series models.
5. The method according to claim 4, characterized in that, The training process for a discrete-time time series model is as follows: The Discrete Long Short-Term Memory (LSTM) network was used as the initial discrete-time time series model. The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial discrete-time time series model is then trained to obtain the trained discrete-time time series model.
6. The method according to claim 4, characterized in that, The training process for a continuous-time time series model is as follows: The true value of the news dynamic background representation is obtained based on the trained news dynamic background representation model, and the time-based dynamic differential equation of the true value of the news dynamic background representation is obtained. The time-based dynamic differential equations are used as the initial continuous-time time series model; The true values of the news dynamic background representation are obtained from the trained news dynamic background representation model. The initial continuous-time time series model is then trained to obtain the trained continuous-time time series model.
7. The method according to claim 1, characterized in that, S9 determines the authenticity of the news to be identified based on the news to be identified, its dynamic background representation, and a trained fake news detection model, including: S91. Embed the news input text to be recognized into the layer and generate word embeddings; S92. Input the publication time of the news to be identified into the time series model to obtain the dynamic background representation of the news to be identified; S93. Input the word embeddings of the news to be identified and the obtained dynamic background representation of the news into the feature extractor to obtain joint features; S94. Input the joint features into the classifier to obtain the authenticity result of the news to be identified.
8. A fake news detection device based on dynamic social context, characterized in that, The device includes: Building blocks are used to construct the initial fake news detection model; The first acquisition unit is used to acquire the news training set; The news training set includes sample news texts, the corresponding real values of the sample news texts, and the publication time of the news. The second acquisition unit is used to acquire the initial news dynamic background representation matrix; The third acquisition unit is used to acquire the news to be identified and the publication time of the news to be identified; The first training unit is used to pre-train the initial fake news detection model based on the news training set and the initial news dynamic background representation matrix, so as to obtain a trained fake news detection model. The second training unit is used to divide the news training set into multiple subsets according to the news release time, train the initial news dynamic background representation model, and obtain the trained news dynamic background representation model. The third training unit is used to obtain the true value of the news dynamic background representation based on the trained news dynamic background representation model, and to train the initial time series model based on the true value of the news dynamic background representation to obtain the trained time series model. The first prediction unit is used to input the publication time of the news to be identified into the trained time series model to obtain the dynamic background representation of the news to be identified. The second prediction unit is used to determine the authenticity of the news to be identified based on the news to be identified, the dynamic background representation of the news to be identified, and the trained fake news detection model.
9. The apparatus according to claim 8, characterized in that, The fake news detection model includes a text embedding layer, a feature extractor, and a classifier; The first training unit is used for: Input sample news texts from the news training set into the text embedding layer to generate word embeddings; The generated word embeddings are jointly input into the feature extractor along with the initial news dynamic background matrix to obtain joint features; The joint features are input into the classifier to obtain the predicted authenticity value of the sample news text. Using the binary cross-entropy loss function, we construct the objective function of the fake news detection model, optimize the gap between the predicted and actual values of the sample news texts, and obtain the trained fake news detection model.
10. The apparatus according to claim 8, characterized in that, The second training unit is used for: Freeze the parameters of the text embedding layer, the parameters of the feature extractor, and the parameter matrix of the classifier; The news training set is divided into multiple subsets based on the publication time of the sample news texts; The initial news dynamic background representation model is trained using multiple predefined subsets to obtain a trained news dynamic background representation model.