Deep-wide multimodal network health rumor detection method and device integrating language style
The language style characteristics of social media posting content are extracted through deep-width multimodal network and Aristotle rhetorical theory, which solves the problem of poor health rumors detection in the existing technology, and achieves higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202411443862.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The existing rumor detection methods have poor detection of health rumors and cannot effectively identify mixed information of true and falsehood.
The deep-width multimodal network health rumor detection method is adopted with a fusion language style. The shallow and deep characteristics of social media posting content are extracted through the pre-trained width module and the depth module, and the appeal to logic, emotion and personality traits are extracted in combination with Aristotle's rhetorical theory, and multi-classification detection is performed through a classifier.
It improves the accuracy and robustness of health rumors identification, and can effectively distinguish between real information, false information and mixed information of true and false information.
Smart Images

Figure CN119577134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rumor detection, and in particular to a deep-wide multimodal network health rumor detection method and device integrating language style. Background Art
[0002] The spread of health rumors on the Internet poses a serious threat to public health and effective detection is urgently needed.
[0003] However, existing rumor detection research mainly focuses on areas such as fake news, fake reviews, and general social media rumors. Unlike fake news and other false information that requires exaggerated fabricated content to attract traffic, health rumors are mostly mixed truth and falsehood.
[0004] Therefore, conventional rumor detection methods targeting fake news are not effective in detecting health rumors. Summary of the Invention
[0005] The present invention provides a deep-wide multimodal network health rumor detection method and device that integrates language style, so as to make up for the defect of poor detection effect of health rumor in existing technology and realize a social media health rumor detection method with higher accuracy and robustness.
[0006] The present invention provides a deep-wide multimodal network health rumor detection method integrating language style, comprising:
[0007] Inputting the to-be-identified social media content and publisher features into a pre-trained width module for extraction, thereby obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features;
[0008] Inputting the social media content to be identified into a pre-trained deep module for extraction, and obtaining deep features representing deep semantics output by the deep module;
[0009] The shallow features and the deep features are spliced and input into a classifier to obtain the output of the classifier, and the social media content to be identified is detected as true information, false information, or a mixture of true and false information.
[0010] According to the present invention, a deep-wide multimodal network health rumor detection method integrating language style is provided.
[0011] The emotional appeal features include at least the number of exclamation marks, the number of emoticons, the number of emotional words corresponding to different emotional categories, and the emotional scores of the social media posts to be identified;
[0012] The emotion score corresponding to each emotion category is determined according to the occurrence frequency and emotion intensity of each emotion word of the corresponding emotion category, and the number of degree adverbs and negative words in the context of a preset length of each emotion word.
[0013] According to a deep, wide, multimodal network health rumor detection method that integrates language style provided by the present invention, the appeal to logic features include at least the proportion of cognitive process words, the proportion of causal words, the number of function words, the number of content words in the text of the social media content to be identified, and the image-text consistency features of the social media content to be identified.
[0014] According to the present invention, a deep-wide multimodal network health rumor detection method integrating language style is provided.
[0015] In the case where the social media content to be identified contains both images and text, each image and text content in the social media content to be identified is sequentially input into the Clip model to obtain an image-text consistency value for each image and text content output by the Clip model;
[0016] Calculating an average of the image-text consistency values of all images in the social media content to be identified, and using the average as the image-text consistency feature of the social media content to be identified;
[0017] In the case that the social media content to be identified does not contain an image, the average of the image-text consistency features of any multiple pieces of social media content containing images is used as the image-text consistency feature of the social media content to be identified.
[0018] According to a deep-wide multimodal network health rumor detection method that integrates language style provided by the present invention, the appeal to personality characteristics at least includes the publisher characteristics, the proportion of first-person pronouns and the number of medical terms of the social media content to be identified.
[0019] According to a method for detecting health rumors on a network using deep and wide multimodal methods that integrate language styles, the present invention further includes:
[0020] Collect multiple health rumor-debunking messages from various rumor-debunking platforms, and use the content theme of each health rumor-debunking message as a keyword to search and collect the content data corresponding to the keyword, wherein each piece of content data includes the ID, text, and image of the content, as well as the ID, gender, number of fans, number of followers, total number of interactions, and VIP status of the publisher of the content;
[0021] After the collected published content data is cleaned, a portion of the cleaned published content data is randomly extracted as a data set used in the training process.
[0022] The present invention also provides a deep-wide multimodal network health rumor detection device integrating language style, comprising:
[0023] A width module is used to input the social media content and publisher features to be identified into the pre-trained width module for extraction, thereby obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features;
[0024] A deep module is used to input the social media content to be identified into a pre-trained deep module for extraction, and obtain deep features representing deep semantics output by the deep module;
[0025] The classification module is used to splice the shallow features and the deep features and input them into a classifier to obtain the output of the classifier, and detect the social media content to be identified as true information, false information, or a mixture of true and false information.
[0026] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the deep-wide multimodal network health rumor detection method integrating language styles as described above.
[0027] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for detecting health rumors in a deep, wide, and multimodal network by integrating language styles as described above is implemented.
[0028] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned deep-wide multimodal network health rumor detection methods that integrate language styles.
[0029] The present invention provides a deep-wide multimodal network health rumor detection method and device that integrates language style. A social media health rumor detection model is constructed based on a deep-wide model. Based on Aristotle's rhetoric theory, the emotional appeal features, personality appeal features, and logical appeal features of the social media content to be identified are extracted as shallow features, and the deep features of the text information of the social media content to be identified are extracted. The deep features and the shallow features are spliced and input into a classifier to realize the prediction of whether the social media content to be identified is true, false, or a mixture of true and false. By introducing the language style features of health rumors, the recognition accuracy of health rumors is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 This is one of the flow charts of the deep-wide multimodal network health rumor detection method integrating language style provided by the present invention;
[0032] Figure 2 This is an architectural diagram of the corresponding model of the deep-wide multimodal network health rumor detection method integrating language style provided by the present invention;
[0033] Figure 3 This is the second flow chart of the deep-wide multimodal network health rumor detection method integrating language style provided by the present invention;
[0034] Figure 4 This is a schematic diagram of the structure of the deep-wide multimodal network health rumor detection device integrating language styles provided by the present invention;
[0035] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0036] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0037] The following combination Figures 1 to 3 The present invention introduces a deep-wide multimodal network health rumor detection method that integrates language style, such as Figure 1 Shown, including:
[0038] Step 101: Input the to-be-identified social media content and publisher features into a pre-trained width module for extraction, and obtain shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features;
[0039] Most current rumor detection models are data-driven and algorithm-driven, focusing primarily on the language content of rumors and lacking attention to language style. For health rumor detection, most current models still rely on traditional machine learning techniques, focusing on language statistics, network features, user characteristics, interaction features, and sentiment.
[0040] However, health rumors differ from typical social media rumors in their generation, content characteristics, and dissemination mechanisms. Specifically, their text needs to convince users of the accuracy of their content. Therefore, in identifying health rumors, features that characterize the linguistic style of text content can be introduced to improve the accuracy of health rumor recognition.
[0041] On this basis, the present invention is based on Aristotle's rhetoric theory to extract features that characterize the language style of health rumors.
[0042] Specifically, Aristotle's rhetorical theory categorizes persuasive styles into three modes: appeal to personality (Ethos), appeal to logic (Logos), and appeal to emotion (Pathos). He suggests that persuaders can utilize this rhetorical triad to achieve their persuasive goals. Appeal to personality, which builds trust by demonstrating specific personality traits such as credibility, reputation, moral integrity, and expertise, is considered the most effective persuasive tool. Highly professional or credible communicators can effectively influence audiences, changing their attitudes and behaviors regarding the issues, products, information, or individuals addressed in the message. Appeal to logic, a persuasive approach based on rationality and logic, primarily involves providing facts, data, statistics, and reasoned arguments to support a viewpoint or position, allowing the audience to evaluate the argument's validity. This is considered the primary persuasive factor. Appeal to emotion, which aims to influence audience opinion by appealing to their emotions, aims to stimulate feelings or emotions, such as fear, sympathy, anger, or joy, to enhance persuasive effectiveness.
[0043] Therefore, based on Aristotle's rhetorical theory, appeal to personality, appeal to logic and appeal to emotion features are extracted from the social media content to be identified and the publisher characteristics, so as to improve the accuracy of health rumor identification based on the above three types of features.
[0044] Optionally, the present invention realizes the extraction of emotional appeal features, personality appeal features and logical appeal features of social media content to be identified in the Chinese context based on a wide and deep model.
[0045] like Figure 2As shown in the figure, the deep-wide model includes a shallow part, namely the width module, which is mainly responsible for extracting shallow features in the data and memorizing the corresponding outputs in historical data; while the deep part, namely the depth module, learns the nonlinear relationship and high-order feature combination of features through neural networks.
[0046] Therefore, by pre-training a deep-wide model, the features of the social media posts and publishers to be identified are extracted through its wide module, and the appeal to emotion features, appeal to personality features and appeal to logic features are obtained as shallow features.
[0047] It can be understood that the social media content to be identified is a piece of health-related information published by the publisher on the social media platform.
[0048] Optionally, the social media platform may be any type of social media platform that is primarily text-based, primarily image-based, or primarily video-based.
[0049] Optionally, the social media content to be identified may be in the form of pure text, a mixture of text and images, or a mixture of text and video.
[0050] Optionally, the text in the social media content to be identified may be text published directly in the form of text, or may be text published in the form of an image or video.
[0051] Optionally, publisher features are used to characterize the publisher's identity portrait. The selected publisher features can be adjusted accordingly for different social media platforms, but can generally include the number of followers of the publisher, as well as historical forwarding, commenting and / or like data. The specific data category selected shall be based on the data that can be obtained by the corresponding social media platform.
[0052] Step 102: Input the social media content to be identified into a pre-trained deep module for extraction, and obtain deep features representing deep semantics output by the deep module;
[0053] The deep module in the deep-wide model is used to extract deep features that represent deep semantics from the input social media content to be identified.
[0054] In one feasible implementation, the deep module uses the BERT (Bidirectional Encoder Representations from Transformers) pre-trained model to extract deep features.
[0055] The BERT model is a deep learning model that has achieved significant breakthroughs in natural language processing and possesses powerful text understanding capabilities. Based on the Transformer architecture, this model uses a bidirectional training approach to fully capture contextual information in text, thereby generating richer word vector representations. As a result, it has been widely used in rumor detection tasks.
[0056] In actual use, the corresponding BERT model only inputs the text content of the social media posts to be identified into the deep module for extracting deep features.
[0057] Optionally, the text content may be text content directly published by the publisher.
[0058] Optionally, when the social media platform is a video-based platform, the text content may be the title of the video and the text content in the introduction.
[0059] Optionally, when the social media content to be identified is a mixed content of text and image published in the form of a long image, the text content in the long image is identified as text content.
[0060] In step 103, the shallow features and the deep features are combined and input into a classifier to obtain an output of the classifier, and the social media content to be identified is detected as true information, false information, or a mixture of true and false information.
[0061] Alternatively, as Figure 2 As shown, a prediction module is designed to jointly learn and predict labels for deep features and shallow features.
[0062] First, the shallow features and deep features obtained by the width module and the depth module are mapped to the same dimension through two fully connected layers, and then the two vectors of the same dimension are concatenated to obtain a more comprehensive feature representation.
[0063] Furthermore, a self-attention layer is set in the prediction module to receive the spliced features and dynamically adjust the importance of different features through a multi-head self-attention mechanism to capture the long-distance dependencies between features, so as to enhance the model's perception and representation capabilities of key information.
[0064] The core idea of the multi-head self-attention mechanism is to independently calculate the weighted correlation between features in different subspaces to capture richer feature interaction information. By transforming the concatenated feature vectors, we obtain three vectors: query, key, and value. The attention score is then used to weight the value vector and sum it to obtain a weighted representation of the features.
[0065] On this basis, since health rumors often present the characteristics of a mixture of true and false information, unlike the binary classification setting in conventional rumor detection models, in order to solve the possible limitations of the generalization performance of binary classification tasks for real tasks, the present invention expands the rumor detection task into a multi-classification task, that is, the output detection results are true information, false information, or a mixture of true and false information.
[0066] Therefore, in the prediction module, an MLP (Multilayer Perceptron) is trained as a classifier. The MLP learns the nonlinear relationships between features and maps the processed features to the final prediction space to generate the final prediction results. The MLP consists of three fully connected layers with 256, 64, and 3 neurons, respectively, reducing the feature dimension layer by layer. ReLU activation functions are introduced between each layer.
[0067] In general, MLP maps the feature representation of the self-attention output to a 256-dimensional vector space through a linear layer, then applies the ReLU activation function to the vector and maps it to 64 dimensions through a linear layer. Finally, a 3-dimensional linear layer outputs the final classification result, that is, the social media content to be identified as true information, false information, or a mixture of true and false information.
[0068] The present invention constructs a social media health rumor detection model based on a deep-wide model. Based on Aristotle's rhetoric theory, it extracts the emotional features, personality features and logical features of the social media content to be identified as shallow features, extracts the deep features of the text information of the social media content to be identified, and maps the deep features and the shallow features to the same dimension through two fully connected layers. The two vectors of the same dimension are then spliced and input into a classifier after splicing to realize the classification of the social media content to be identified as true, false or a mixture of true and false. By introducing the language style features of health rumors, the recognition accuracy of health rumors is improved.
[0069] In the deep-wide multimodal network health rumor detection method integrating language style of the present invention, the emotional appeal features include at least the number of exclamation marks, the number of emoticons, the number of emotional words corresponding to different emotional types, and the emotional scores;
[0070] The emotion score corresponding to each emotion category is determined according to the occurrence frequency and emotion intensity of each emotion word of the corresponding emotion category, and the number of degree adverbs and negative words in the context of a preset length of each emotion word.
[0071] Optionally, appeal to emotion features are extracted from the text of the social media content to be identified. The definition of appeal to emotion features is shown in Table 1:
[0072] Table 1
[0073]
[0074] The appeal to emotion feature characterizes the language style feature of the appeal to emotion, taking into account the number and emotion scores of emotion words of different emotion categories in the published content to be identified. In a feasible implementation, the emotion words are determined based on an existing emotion dictionary.
[0075] For example, the text of the social media post to be identified is segmented, and each segmented word is searched and matched against a sentiment dictionary to determine whether it belongs to a sentiment word and the corresponding sentiment category. Finally, the sentiment words corresponding to each sentiment category and the corresponding number of sentiment words are counted.
[0076] Since exclamation marks are usually used to express strong emotions or mood swings, the punctuation marks mainly extract the exclamation mark "!" and count the number of exclamation marks as one of the emotional features.
[0077] Emoticons are widely used as a medium for conveying emotions in online communication, becoming an explicit and fixed way to express emotions. Their number can reflect the intensity of the poster's emotional expression, so we count the number of emoticons in a text as one of its emotional features.
[0078] In addition, the sentiment score indicates the intensity of the publisher's sentiment level on each sentiment category.
[0079] In one possible implementation, the sentiment score is calculated as follows:
[0080] For a given text T Specific emotions in e , calculate the text according to formula (1) T The i words t i Sentiment score :
[0081]
[0082] in, is the context window size, for The negation value of for The degree adverb value of is shown in formula (2). For words If the word Not in the sentiment dictionary middle, Recorded as 0, otherwise Recorded as the corresponding sentiment intensity value in the sentiment dictionary, and calculated Negative word values in context and degree adverb values , based on which the final sentiment score of the word is calculated .
[0083] Adverbs of degree are calculated by referring to the adverb entries in WordNet, as shown in formula (3); negation words are calculated by referring to the negation word list in NLTK, as shown in formula (4).
[0084] Then, the text T Each word in Sentiment score Add up and you get the text T Specific emotions e Sentiment score , as shown in formula (5):
[0085] (5)
[0086] Symbolic sentiment features include the number of exclamation marks and emoticons. As key components of text, symbolic content such as punctuation and emoticons can complement text sentiment features. For example, exclamation marks ("!") are often used to express strong emotions, while emoticons, as graphical symbols of emotional expression widely used in online communication, can convey more subtle emotions, thereby achieving persuasive purposes.
[0087] In the deep-wide multimodal network health rumor detection method that integrates language style of the present invention, the appeal to logic features include at least the proportion of cognitive process words, the proportion of causal words, the proportion and number of functional words in the text of the social media content to be identified, and the image-text consistency features of the social media content to be identified.
[0088] Optionally, appeal to logic features are extracted from the text and image of the social media content to be identified. The definition of appeal to logic features is shown in Table 2:
[0089] Table 2
[0090] ;
[0091] The appeal to logic feature represents the characteristic dimension of the persuasive language style of appeal to logic. Optionally, the proportion of causal words, the number of functional words, the number of content words, the proportion of cognitive process words and the consistency of text and picture are used as the appeal to logic features.
[0092] Specifically, causal words such as "because" and "therefore" can reflect how the publisher constructs arguments, connects ideas, and demonstrates causal relationships, which helps to understand the relationships between sentences and paragraphs in the text and shows the logic and coherence of the text. Therefore, the frequency or proportion of different types of causal words in the text is usually used to measure the logical structure of the text.
[0093] Function words refer to words that do not carry specific semantics in a sentence, such as "is" and "of". Although function words themselves do not have strong semantics, their use affects the grammar structure, causal relationships, coherence, and fluency of the text, thus affecting the overall logic and rigor of the text.
[0094] Optionally, with the help of LIWC-2015 (Linguistic Inquiry and Word Count, a text analysis software), calculate the proportion and quantity of 8 common function words in the Chinese scenario.
[0095] Cognitive process words (such as causal words, difference words, etc.) can indicate that the author is expressing logical thoughts such as reasoning, analysis, and causal relationships. Therefore, their proportion is related to the logic of the published content.
[0096] Content words (such as nouns, verbs, etc.) often carry more semantic information and can reflect the content density of the published content. More content words may mean that the published content is rich in information, has a more compact structure, and is conducive to logical expression.
[0097] The quantity and proportion of words in the above categories are all obtained based on the text statistics of the social media posts to be identified.
[0098] In addition, since the consistency between pictures and text represents the consistency of the information content expressed by pictures and text, it is generally considered that the higher the consistency between pictures and text, the higher the clarity of information logic, which is more conducive to the effective dissemination of information and the audience's understanding of information.
[0099] Therefore, take the consistency between pictures and text as one of the logical features and calculate the consistency between the images and text of the social media posts to be identified.
[0100] Optionally, when there are multiple images in the social media posts to be identified, calculate the consistency between each image and all the text, and then take the average value of the consistency between pictures and text of all the text as the consistency between pictures and text of the social media posts to be identified.
[0101] Optionally, when there are multiple images in the social media posts to be identified and each image has clear corresponding text content, calculate the consistency between each image and its corresponding text content, and then take the average value of the consistency between pictures and text of all the text as the consistency between pictures and text of the social media posts to be identified.
[0102] Optionally, when the social media content to be identified includes video and text, several key frames of the video are captured as images of the social media content to be identified, for calculating the image-text consistency of the social media content to be identified.
[0103] On this basis, in a feasible implementation, the number and proportion of each function word, the proportion of causal words, the proportion of cognitive process words, the number of content words and the consistency of pictures and texts are used as the appeal to logic features obtained, that is, the proportion of causal words is counted separately. In other feasible implementations, similarly, the proportion of difference words in cognitive process words can also be counted separately, and / or the total proportion of all function words can be counted as one of the output appeal to logic features, which can be adjusted according to the training results.
[0104] It should be noted that, unlike conventional rumor-debunking models that directly extract images from published content to obtain image modality features, the present invention takes into account that the semantic features of images alone are difficult to achieve accurate rumor detection. Therefore, it does not use pre-trained models such as ResNet to extract its features. Instead, it chooses to calculate the image-text consistency features of the image and text content, and use this as the features in the width module to assist in the rumor detection task.
[0105] In the deep-wide multimodal network health rumor detection method that integrates language style of the present invention, when the social media content to be identified contains both images and text, each image and text content in the social media content to be identified is sequentially input into the Clip model to obtain the image-text consistency value of each image and text content output by the Clip model;
[0106] Calculating an average of the image-text consistency values of all images in the social media content to be identified, and using the average as the image-text consistency feature of the social media content to be identified;
[0107] Specifically, the Clip model is used to calculate image-text consistency and incorporate it into the model. Clip (Contrastive Language-Image Pre-training) is a multimodal model based on contrastive learning. It uses 400 million text-image pairs for pre-training. The Clip model has a shared embedding space, allowing images and text to be embedded in the same space for consistency calculation.
[0108] When the social media posts to be identified contain both images and text, the Clip model is used to calculate the image-text consistency value of each image and the complete text respectively. The image-text consistency values of all images are then averaged, and the obtained mean is used as the image-text consistency feature of the social media posts to be identified.
[0109] The above method can reduce the deviation caused by individual images and obtain a more stable and reliable similarity measurement.
[0110] In the case that the social media content to be identified does not contain an image, the average of the image-text consistency features of any multiple pieces of social media content containing images is used as the image-text consistency feature of the social media content to be identified.
[0111] When the social media post to be identified does not contain an image, the average of the image-text consistency features of multiple other social media posts containing images is used for filling in the gaps to ensure data integrity and improve the robustness of the model, so that it can be applied to detect the authenticity of social media posts with or without images.
[0112] In a specific embodiment, during the training phase of the model, social media posts containing images are first used for training to obtain the mean image-text consistency of all social media posts containing images used for training. This mean is then used to fill in the image-text consistency features of social media posts that do not contain images during the detection process.
[0113] It is understandable that the image-text consistency feature represented by the above mean is not biased, that is, the image-text consistency feature of social media content that does not contain images does not have a biased impact on the judgment of its authenticity.
[0114] In the deep-wide multimodal network health rumor detection method that integrates language style of the present invention, the appeal to personality characteristics at least includes the publisher characteristics, the proportion of first-person pronouns and the number of medical terms of the social media content to be identified.
[0115] Optionally, publisher characteristics, the proportion of first-person pronouns, and the number of medical terms are extracted from the input publisher characteristics and the text information of the social media content to be identified as appeal to personality characteristics. The definition of appeal to personality characteristics is shown in Table 3:
[0116] Table 3
[0117] ;
[0118] Appeal to personality traits characterizes the dimension of persuasive language style characteristics of appeal to personality, specifically considering the publisher's characteristics, the proportion of the first person in the text of the social media content to be identified, and the number of medical terms.
[0119] Optionally, publisher characteristics may include the number of fans, number of followers, gender, total number of interactions, authentication information, and VIP status of the publisher on the corresponding social media platform, which are specifically determined based on the information available on the corresponding social media platform.
[0120] These characteristics can reflect the publisher's status, social influence, and credibility on the platform. The number of followers of a publisher represents their influence to a certain extent. Users with greater influence tend to maintain their reputation and image, making their posts more likely to be authentic.
[0121] Similarly, people who follow fewer people tend to reduce the false interactions that come with following each other, and are therefore generally considered to have greater autonomy and independence, and the information they post is more reliable.
[0122] Publishers of different genders often have different tendencies and preferences when publishing and processing information. Moreover, the gender of the information publisher will directly affect the receiver's perception of the credibility and authority of the information, especially in the field of health.
[0123] In addition, if the content published by a publisher is often forwarded, commented on and liked by a large number of users, it means that the publisher has a certain degree of authority and influence in the field or community. Such publishers usually pay more attention to their own reputation and image, and the authenticity of their posts may be higher.
[0124] Publishers with certified information are generally considered to have higher authority and credibility in their professional fields, and their posts are less likely to be rumors.
[0125] VIP status refers to the user's subscription status, which to a certain extent reflects the publisher's cost investment in the platform. VIP users may have a higher marginal cost for spreading rumors and a lower possibility of spreading rumors.
[0126] In addition, the proportion of first-person pronouns and the number of medical terms are calculated from the text information of the social media posts to be identified as personality traits.
[0127] The proportion of first-person pronouns can reflect the credibility and authority that the publisher attempts to demonstrate. Therefore, a higher proportion of first-person pronouns usually indicates that the publisher pays more attention to expressing personal positions in information transmission and attempts to enhance the audience's trust through self-disclosure.
[0128] The number of medical terms describes the amount of health information contained in the published content, and to a certain extent can affect the audience's perception of the professionalism and authority of the information publisher.
[0129] The ability of the trustee will affect the trust level of the trustgiver. The more frequently medical terms are used, the more credible the publisher's ability tends to appear in the eyes of the audience, thereby enhancing the persuasiveness of health rumors.
[0130] On this basis, the number of fans, number of followers, gender, total number of interactions and authentication information of the publisher of the social media content to be identified, as well as the proportion of first-person pronouns and the number of medical terms are used as the appeal to personality characteristics.
[0131] In the deep-wide multimodal network health rumor detection method integrating language style of the present invention, before the step of inputting the social media content to be identified and the publisher features into the pre-trained wide module for extraction, the method further includes:
[0132] Collect multiple health rumor-debunking messages from various rumor-debunking platforms, and use the content theme of each health rumor-debunking message as a keyword to search and collect the content data corresponding to the keyword, wherein each piece of content data includes the ID, text, and image of the content, as well as one or more of the ID, number of fans, number of followers, gender, total number of interactions, authentication information, and VIP status of the publisher of the content;
[0133] Since existing data sets are difficult to be used for training the model corresponding to the method of the present invention, it is necessary to construct a corresponding data set.
[0134] Specifically, first, multiple health and wellness rumor-refuting information are manually collected from various Internet joint rumor-refuting platforms and / or each social platform’s own official rumor-refuting platform. In this implementation, a total of 300 pieces of information are collected, and the collected rumor-refuting information is used as the basis for manual labeling.
[0135] We then extracted the themes of each rumor-debunking message and used them as search keywords to construct a keyword list. We then searched social media platforms using each keyword in the keyword list as a search keyword, obtaining all published content related to the search keywords. A total of 79,409 pieces of data were obtained, forming the original dataset.
[0136] The data obtained for each published content includes the ID, text, and pictures of the published content, as well as one or more of the ID, number of fans, number of followers, gender, total number of interactions, authentication information, and VIP status of the publisher of the published content, whichever is available on the corresponding social media platform.
[0137] At the same time, the image data is saved and numbered, and an index is built to associate and map each post text with the image to ensure the correlation between different modal data.
[0138] After the collected published content data is cleaned, a portion of the cleaned published content data is randomly extracted as a data set used in the training process.
[0139] To improve data quality, the data was preprocessed by removing duplicate values, removing noise, and cleaning text. Then, a 10% random sample was taken from the original dataset to generate 7,620 data points as the initial experimental dataset, which was used for training the deep and wide models and classifiers.
[0140] Furthermore, the initial experimental dataset was annotated. To maximize the accuracy of the annotations and reduce annotation bias due to subjective factors, official rumor-busting information was used as the basis for annotation, and cross-validation was used for annotation.
[0141] First, five research assistants were fully trained to ensure they were familiar with the labeling rules and all rumor-debunking information. Then, 50 posts were randomly selected for a pre-test. The five annotators were asked to independently read each post and determine its category. During the labeling process, the annotators were unable to refer to others' judgments. This was done to assess the annotators' understanding and consistency, and to fully discuss and resolve potential issues.
[0142] Based on the pre-test, five annotators conducted formal annotation. After the independent annotation was completed, samples with consistent annotations were directly adopted as the final labeling of the sample. For samples with different annotation results, the five annotators fully discussed the results based on the medical big model and the opinions of domain experts to reach a consensus.
[0143] The final experimental dataset contained 2,546 pieces of two-sided health information, 1,830 pieces of health rumors, and 3,244 pieces of real health information. The processed dataset was randomly divided into two parts in an 8:2 ratio to form the training and test sets required for model training.
[0144] On this basis, the overall process of the present invention is as follows Figure 3 shown.
[0145] In order to evaluate and test the performance of the model corresponding to the method of the present invention, a comparative experiment was set up during the experiment. The baseline model selected in the comparative experiment is:
[0146] Traditional machine learning models: including two benchmark models, Random Forest (RF) and XGBoost, which mainly use shallow language style features for prediction.
[0147] Deep learning models: including two benchmark models, BERT and bi-LSTM, which mainly use deep language content for prediction.
[0148] LSTM-Attention: A rumor detection model that integrates content features and user features. This model uses an LSTM-based attention mechanism to extract deep semantic information from text and perform task prediction.
[0149] HMCAN: A hierarchical multimodal contextual attention network rumor detection model. This model uses BERT and ResNet models to extract deep semantic features of text and image features, respectively. It then fuses and concatenates the fused features through a multimodal contextual attention mechanism, and finally completes the classification task through a fully connected layer.
[0150] The above baseline model and model are tested on the dataset, and the results are shown in Table 4:
[0151] Table 4
[0152] ;
[0153] In the table, the MWDHRD-ART framework is the model corresponding to the method of the present invention. It can be found that the MWDHRD-ART framework proposed in the present invention outperforms the five baseline models in all performance indicators. First, compared with the traditional machine learning methods RF and XGBoost, MWDHRD-ART improves the F1 value by 5.5% and 4.26%, respectively. Second, compared with the deep learning methods BERT and bi-LSTM, MWDHRD-ART improves the F1 value by 6.16% and 5.78%, respectively. Finally, compared with two baseline models in existing research, LSTM-Attention and HMCAN, MWDHRD-ART improves the F1 value by 7.86% and 1.54%, respectively. Compared with LSTM-Attention and HMCAN, MWDHRD-ART simultaneously utilizes deep language content features and shallow language style features for prediction, especially considering the amount of medical information in the text, making it more suitable for rumor identification tasks in the health field. These results show that the MWDHRD-ART framework proposed in the present invention is effective for carrying out the task of detecting health rumors on social media.
[0154] It is worth mentioning that in the baseline model of the present invention, the F1 values of the two machine learning models that only use shallow language style features are slightly higher than the F1 values of the deep learning model, which to a certain extent shows that the language style features extracted in this paper are of high quality. Therefore, in the face of the real scenario of massive tasks to be predicted, only by reasonably extracting lower-dimensional language style features can it achieve a health rumor detection effect similar to that based on the deep language content feature model. It can be seen that language style features have high efficiency and practicality in the task of health rumor detection, especially when computing resources are limited, its application prospects are even broader.
[0155] The following describes the deep, wide, multimodal network health rumor detection device with integrated language style provided by the present invention. The deep, wide, multimodal network health rumor detection device with integrated language style described below and the deep, wide, multimodal network health rumor detection method with integrated language style described above can be referenced to each other.
[0156] like Figure 4 As shown, the deep-wide multimodal network health rumor detection device integrating language style includes a width module 401, a depth module 402 and a classification module 403:
[0157] The width module 401 is configured to input the social media content and publisher features to be identified into a pre-trained width module for extraction, thereby obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features;
[0158] Most current rumor detection models are data-driven and algorithm-driven, focusing primarily on the language content of rumors and lacking attention to language style. For health rumor detection, most current models still rely on traditional machine learning techniques, focusing on language statistics, network features, user characteristics, interaction features, and sentiment.
[0159] However, health rumors differ from typical social media rumors in their generation, content characteristics, and dissemination mechanisms. Specifically, their text needs to convince users of the accuracy of their content. Therefore, in identifying health rumors, features that characterize the linguistic style of text content can be introduced to improve the accuracy of health rumor recognition.
[0160] On this basis, the present invention is based on Aristotle's rhetoric theory to extract features that characterize the language style of health rumors.
[0161] Specifically, Aristotle's rhetorical theory categorizes persuasive styles into three modes: appeal to personality (Ethos), appeal to logic (Logos), and appeal to emotion (Pathos). He suggests that persuaders can utilize this rhetorical triad to achieve their persuasive goals. Appeal to personality, which builds trust by demonstrating specific personality traits such as credibility, reputation, moral integrity, and expertise, is considered the most effective persuasive tool. Highly professional or credible communicators can effectively influence audiences, changing their attitudes and behaviors regarding the issues, products, information, or individuals addressed in the message. Appeal to logic, a persuasive approach based on rationality and logic, primarily involves providing facts, data, statistics, and reasoned arguments to support a viewpoint or position, allowing the audience to evaluate the argument's validity. This is considered the primary persuasive factor. Appeal to emotion, which aims to influence audience opinion by appealing to their emotions, aims to stimulate feelings or emotions, such as fear, sympathy, anger, or joy, to enhance persuasive effectiveness.
[0162] Therefore, based on Aristotle's rhetoric theory, appeal to personality, appeal to logic and appeal to emotion features are extracted from the social media content to be identified and the publisher characteristics, so as to improve the accuracy of health rumor identification based on the above three types of features.
[0163] Optionally, the present invention realizes the extraction of emotional appeal features, personality appeal features and logical appeal features of social media content to be identified in the Chinese context based on a wide and deep model.
[0164] like Figure 2 As shown in the figure, the deep-wide model includes a shallow part, namely the width module, which is mainly responsible for extracting shallow features in the data and memorizing the corresponding outputs in historical data; while the deep part, namely the depth module, learns the nonlinear relationship and high-order feature combination of features through neural networks.
[0165] Therefore, by pre-training a deep-wide model, the features of the social media posts and publishers to be identified are extracted through its wide module, and the appeal to emotion features, appeal to personality features and appeal to logic features are obtained as shallow features.
[0166] It can be understood that the social media content to be identified is a piece of health-related information published by the publisher on the social media platform.
[0167] Optionally, the social media platform may be any type of social media platform that is primarily text-based, primarily image-based, or primarily video-based.
[0168] Optionally, the social media content to be identified may be in the form of pure text, a mixture of text and images, or a mixture of text and video.
[0169] Optionally, the text in the social media content to be identified may be text published directly in the form of text, or may be text published in the form of an image or video.
[0170] Optionally, publisher features are used to characterize the publisher's identity portrait. The selected publisher features can be adjusted accordingly for different social media platforms, but can generally include the number of followers of the publisher, as well as historical forwarding, commenting and / or like data. The specific data category selected shall be based on the data that can be obtained by the corresponding social media platform.
[0171] A depth module 402 is configured to input the social media content to be identified into a pre-trained depth module for extraction, and obtain deep features representing deep semantics output by the depth module;
[0172] The deep module in the deep-wide model is used to extract deep features that represent deep semantics from the input social media content to be identified.
[0173] In one feasible implementation, the deep module uses the BERT (Bidirectional Encoder Representations from Transformers) pre-trained model to extract deep features.
[0174] The BERT model is a deep learning model that has achieved significant breakthroughs in natural language processing and possesses powerful text understanding capabilities. Based on the Transformer architecture, this model uses a bidirectional training approach to fully capture contextual information in text, thereby generating richer word vector representations. As a result, it has been widely used in rumor detection tasks.
[0175] In actual use, the corresponding BERT model only inputs the text content of the social media posts to be identified into the deep module for extracting deep features.
[0176] Optionally, the text content may be text content directly published by the publisher.
[0177] Optionally, when the social media platform is a video-based platform, the text content may be the title of the video and the text content in the introduction.
[0178] Optionally, when the social media content to be identified is a mixed content of text and image published in the form of a long image, the text content in the long image is identified as text content.
[0179] The classification module 403 is used to splice the shallow features and the deep features and input them into a classifier to obtain the output of the classifier, and detect the social media content to be identified as true information, false information, or a mixture of true and false information.
[0180] Alternatively, as Figure 2 As shown, a prediction module is designed to jointly learn and predict labels for deep features and shallow features.
[0181] First, the shallow features and deep features obtained by the width module and the depth module are mapped to the same dimension through two fully connected layers, and then the two vectors of the same dimension are concatenated to obtain a more comprehensive feature representation.
[0182] Furthermore, a self-attention layer is set in the prediction module to receive the spliced features and dynamically adjust the importance of different features through a multi-head self-attention mechanism to capture the long-distance dependencies between features, so as to enhance the model's perception and representation capabilities of key information.
[0183] The core idea of the multi-head self-attention mechanism is to independently calculate the weighted correlation between features in different subspaces to capture richer feature interaction information. By transforming the concatenated feature vectors, we obtain three vectors: query, key, and value. The attention score is then used to weight the value vector and sum it to obtain a weighted representation of the features.
[0184] On this basis, since health rumors often exhibit the characteristics of a mixture of true and false information, unlike the binary classification setting in conventional rumor detection models, in order to solve the possible limitations of the generalization performance of binary classification tasks for real tasks, the present invention expands the rumor detection task into a multi-classification task, that is, the output detection results are true information, false information, or a mixture of true and false information.
[0185] Therefore, in the prediction module, an MLP (Multilayer Perceptron) is trained as a classifier. The MLP learns the nonlinear relationships between features and maps the processed features to the final prediction space to generate the final prediction results. The MLP consists of three fully connected layers with 256, 64, and 3 neurons, respectively, reducing the feature dimension layer by layer. ReLU activation functions are introduced between each layer.
[0186] In general, MLP maps the feature representation of the self-attention output to a 256-dimensional vector space through a linear layer, then applies the ReLU activation function to the vector and maps it to 64 dimensions through a linear layer. Finally, a 3-dimensional linear layer outputs the final classification result, that is, the social media content to be identified as true information, false information, or a mixture of true and false information.
[0187] The present invention constructs a social media health rumor detection model based on a deep-wide model. Based on Aristotle's rhetoric theory, it extracts the emotional features, personality features and logical features of the social media content to be identified as shallow features, extracts the deep features of the text information of the social media content to be identified, and maps the deep features and the shallow features to the same dimension through two fully connected layers. The two vectors of the same dimension are then spliced and input into a classifier after splicing to realize the detection of true, false or mixed true and false content of the social media content to be identified. By introducing the language style features of health rumors, the recognition accuracy of health rumors is improved.
[0188] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call the logic instructions in the memory 530 to execute a social media health rumor detection method based on Aristotle's rhetoric theory. The method includes: inputting the social media content to be identified and the publisher's features into a width module for extraction, obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features; inputting the social media content to be identified into a pre-trained deep module for extraction, obtaining deep features representing deep semantics output by the depth module; splicing the shallow features and the deep features and inputting them into a classifier to obtain the output of the classifier, and detecting the social media content to be identified as true information, false information, or a mixture of true and false information.
[0189] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0190] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the deep-wide multimodal network health rumor detection method that integrates language styles provided by the above methods. The method includes: inputting the social media content to be identified and the publisher characteristics into a pre-trained wide module for extraction, and obtaining shallow features output by the wide module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features; inputting the social media content to be identified into a pre-trained deep module for extraction, and obtaining deep features representing deep semantics output by the deep module; splicing the shallow features and the deep features and inputting them into a classifier to obtain the output of the classifier, and detecting the social media content to be identified as true information, false information, or a mixture of true and false information.
[0191] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the deep-wide multimodal network health rumor detection method that integrates language styles provided by the above-mentioned methods, the method comprising: inputting the social media content to be identified and the publisher's features into a pre-trained wide module for extraction, and obtaining shallow features output by the wide module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features; inputting the social media content to be identified into a pre-trained deep module for extraction, and obtaining deep features representing deep semantics output by the deep module; splicing the shallow features and the deep features and inputting them into a classifier to obtain the output of the classifier, and detecting the social media content to be identified as true information, false information, or a mixture of true and false information.
[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0193] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A deep-wide multimodal network health rumor detection method integrating language style, characterized by: include: Inputting the to-be-identified social media content and publisher features into a pre-trained width module for extraction, thereby obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features; Inputting the social media content to be identified into a pre-trained deep module for extraction, and obtaining deep features representing deep semantics output by the deep module; The shallow features and the deep features are combined and input into a classifier to obtain an output of the classifier, and the social media content to be identified is detected as true information, false information, or a mixture of true and false information; The emotional appeal features include at least the number of exclamation marks, the number of emoticons, the number of emotional words corresponding to different emotional categories, and the emotional scores of the social media content to be identified; The emotion score corresponding to each emotion category is determined according to the frequency of occurrence and emotion intensity of each emotion word of the corresponding emotion category, and the number of degree adverbs and negative words in the context of each emotion word of a preset length; The appeal to logic features include at least the proportion of cognitive process words, the proportion of causal words, the number of function words, the number of content words in the text of the social media content to be identified, and the consistency feature of the text and image of the social media content to be identified; In the case where the social media posting to be identified contains both images and text, each image and text content in the social media posting to be identified is sequentially input into the Clip model to obtain an image-text consistency value for each image and text content output by the Clip model; the average image-text consistency value of all images in the social media posting to be identified is calculated, and the average value is used as the image-text consistency feature of the social media posting to be identified; In the case where the social media content to be identified does not contain an image, the average of the image-text consistency features of any multiple pieces of social media content containing images is used as the image-text consistency feature of the social media content to be identified; The personality traits appealed to include at least the number of fans, number of followers, gender, total number of interactions and authentication information of the publisher of the social media content to be identified, as well as the proportion of first-person pronouns and the number of medical terms.
2. The deep-wide multimodal network health rumor detection method integrating language style according to claim 1 is characterized in that: Before the step of inputting the to-be-identified social media content and publisher features into the pre-trained width module for extraction, the method further includes: Collect multiple health rumor-debunking messages from various rumor-debunking platforms, and use the content theme of each health rumor-debunking message as a keyword to search and collect the content data corresponding to the keyword, wherein each piece of content data includes the ID, text, and image of the content, as well as the ID, gender, number of fans, number of followers, total number of interactions, and VIP status of the publisher of the content; After the collected published content data is cleaned, a portion of the cleaned published content data is randomly extracted as a data set used in the training process.
3. A deep-wide multimodal network health rumor detection device integrating language style, characterized by: include: A width module is used to input the social media content and publisher features to be identified into the pre-trained width module for extraction, thereby obtaining shallow features output by the width module, wherein the shallow features include appeal to logic features, appeal to emotion features, and appeal to personality features; A deep module is used to input the social media content to be identified into a pre-trained deep module for extraction, and obtain deep features representing deep semantics output by the deep module; A classification module, configured to combine the shallow features and the deep features and input the combined features into a classifier, obtain an output of the classifier, and detect the social media content to be identified as true information, false information, or a mixture of true and false information; The emotional appeal features include at least the number of exclamation marks, the number of emoticons, the number of emotional words corresponding to different emotional categories, and the emotional scores of the social media content to be identified; The emotion score corresponding to each emotion category is determined according to the frequency of occurrence and emotion intensity of each emotion word of the corresponding emotion category, and the number of degree adverbs and negative words in the context of each emotion word of a preset length; The appeal to logic features include at least the proportion of cognitive process words, the proportion of causal words, the number of function words, the number of content words in the text of the social media content to be identified, and the consistency feature of the text and image of the social media content to be identified; In the case where the social media posting to be identified contains both images and text, each image and text content in the social media posting to be identified is sequentially input into the Clip model to obtain an image-text consistency value for each image and text content output by the Clip model; the average image-text consistency value of all images in the social media posting to be identified is calculated, and the average value is used as the image-text consistency feature of the social media posting to be identified; In the case where the social media content to be identified does not contain an image, the average of the image-text consistency features of any multiple pieces of social media content containing images is used as the image-text consistency feature of the social media content to be identified; The personality traits appealed to include at least the number of fans, number of followers, gender, total number of interactions and authentication information of the publisher of the social media content to be identified, as well as the proportion of first-person pronouns and the number of medical terms.
4. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the deep-wide multimodal network health rumor detection method integrating language styles as described in claim 1 or 2 is implemented.
5. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the deep-wide multimodal network health rumor detection method integrating language styles as described in claim 1 or 2 is implemented.
6. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the deep-wide multimodal network health rumor detection method integrating language styles as described in claim 1 or 2 is implemented.