Multilingual text cognition rumor detection method and system, product and medium
By obtaining and comparing the cognitive response data of users to text in different language areas, calculating the information understanding deviation value and positioning the difference fragments, the problem of low accuracy of rumor detection in cross-language information dissemination is solved, and more accurate rumor recognition and early warning is achieved.
Patent Information
- Application Number
- CN202510138877.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately reflect the actual understanding and cognitive status of the user group of information in the dissemination of cross-language information, which affects the accuracy of rumor detection.
By obtaining the cognitive response data of the user to the text by the source and target language areas, calculating the information understanding deviation value, positioning the difference fragments and determining the location and degree of tampering. If the degree and number of tampering exceed the threshold, it is determined as a rumor text and an early warning is issued.
It improves the accuracy of rumor detection, effectively prevents rumor spread, and maintains the authenticity and stability of the information dissemination environment.
Smart Images

Figure CN120011518A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic digital data processing, and in particular to a multilingual text cognitive rumor detection method, system, product and medium. Background Art
[0002] In the era of globalization, information is increasingly transmitted between different language and cultural groups. When information is transmitted from one language region to another, there are often significant differences in the cognition and understanding of the same information by different groups. This cognitive difference may lead to distortion or misinterpretation of information, causing misunderstandings and conflicts between groups, and in serious cases, affecting social stability and international relations. Therefore, it is of great significance to effectively monitor cross-language information transmission.
[0003] In related technologies, monitoring can be carried out by comparing text features. First, texts in different languages are segmented, and then a word vector model is established to obtain semantic representation. Algorithms such as cosine similarity are used to calculate the similarity between texts. Finally, a similarity threshold is used to determine whether the information has changed significantly.
[0004] However, due to the differences in the meanings of words in different languages and cultural backgrounds, text feature comparison is difficult to accurately reflect the actual understanding and cognitive status of user groups towards information, resulting in a large deviation between the system's judgment of information changes and actual user cognition, affecting the accuracy of rumor detection. Summary of the invention
[0005] The present application provides a multilingual text recognition rumor detection method, system, product and medium for improving the accuracy of rumor detection.
[0006] In the first aspect, the present application provides a multilingual text cognitive rumor detection method, which is applied to a multilingual text cognitive rumor detection system, the method comprising: obtaining source cognitive response data of users in a source language area to a source text, the cognitive response data including an understanding index, an identification index and a dissemination willingness index of the source text, and the source text becomes a target text through an information dissemination process; obtaining target cognitive response data of users in a target language area to the target text, comparing the source cognitive response data with the target cognitive response data to obtain cognitive difference data, the cognitive difference data including an information understanding deviation value; when the information understanding deviation value exceeds a preset deviation threshold, locating the difference segment according to the semantic feature mapping relationship between the source text and the target text, and determining the tampering position and degree of tampering of the difference segment; if the tampering degree of the target text is greater than a preset degree threshold and the number of the tampering positions is greater than a preset number threshold, then determining that the target text is a rumor text; and sending rumor warning information to an information publishing platform.
[0007] By adopting the above technical solution, we first obtain the cognitive response data of the source language area users to the source text and the cognitive response data of the target language area users to the target text, among which the understanding index, identification index and dissemination willingness index can fully reflect the user's cognitive state of the text, and compare the two to obtain cognitive difference data. When the information understanding deviation value exceeds the preset threshold, the difference fragment is located according to the semantic feature mapping relationship between the source text and the target text, and the tampering position and degree are determined. If the tampering degree and number of positions of the target text exceed the threshold, it is determined to be a rumor text and an early warning is issued. By integrating various data and analysis, the accuracy of identifying rumor texts is improved, the spread of rumors is effectively prevented, and the authenticity and stability of the information dissemination environment are maintained.
[0008] In combination with some embodiments of the first aspect, in some embodiments, the target cognitive response data of users in the target language area to the target text is obtained, and the source cognitive response data is compared with the target cognitive response data to obtain cognitive difference data, specifically including: obtaining the target cognitive response data of users in the target language area to the target text, inputting the target cognitive response data into a cognitive analysis model to obtain a target cognitive feature, and the target cognitive feature includes a text fact consistency feature, a user emotional tendency feature, and a user cognitive bias feature; inputting the source cognitive response data into the cognitive analysis model to obtain a source cognitive feature, and the source cognitive feature includes a fact consistency feature, a user emotional tendency feature, and a user cognitive bias feature; performing comparative analysis on the target cognitive feature and the source cognitive feature, and calculating the fact cognition difference degree, the emotional tendency difference degree, and the cognitive bias degree; and calculating the cognitive difference data according to the fact cognition difference degree, the emotional tendency difference degree, and the cognitive bias degree.
[0009] By adopting the above technical solution, after obtaining the target cognitive response data of users in the target language area to the target text, it is input into the cognitive analysis model together with the source cognitive response data. The model can extract cognitive features containing key information such as text fact consistency features, user emotional tendency features and user cognitive bias features, and calculate the multi-dimensional difference degree by comparing and analyzing these features, and then obtain cognitive difference data. By using the mutual correlation and difference comparison of multi-dimensional features, it can deeply explore the essence of cognitive differences of users in different regions to information, making cognitive difference data more accurate and comprehensive, and providing a reliable basis for subsequent rumor judgment.
[0010] In combination with some embodiments of the first aspect, in some embodiments, before the step of obtaining target cognitive response data of users in the target language area to the target text, the method also includes: searching the source text in a cognitive database corresponding to the target language area to obtain related texts; extracting the degree of overlap of key information between each related text and the source text, and determining the related text whose information overlap is greater than a overlap threshold as the target text.
[0011] By adopting the above technical solution, before obtaining the target cognitive response data, the source text is retrieved from the cognitive database corresponding to the target language area to obtain relevant texts, and the target texts with key information overlap greater than the threshold are screened out. This step narrows the scope of analysis and focuses on the target texts that are closely related to the source texts. Subsequent cognitive analysis of these target texts avoids interference from irrelevant information and improves the efficiency and pertinence of the entire detection process. Ensure that the cognitive response data of the target text is obtained more accurately within limited resources and time.
[0012] In combination with some embodiments of the first aspect, in some embodiments, when the information understanding deviation value exceeds a preset deviation threshold, the step of locating the difference segment according to the semantic feature mapping relationship between the source text and the target text, and determining the tampering position and tampering degree of the difference segment specifically includes: calculating the semantic similarity of corresponding segments between the source text and the target text; marking the text segment whose semantic similarity is lower than the preset similarity threshold as a difference segment, analyzing the number of words that have been changed, deleted or added in the difference segment, and obtaining the tampering position; calculating the semantic change amplitude of the tampered words in the difference segment, and obtaining the tampering degree of the tampering position.
[0013] By adopting the above technical solution, when the information understanding deviation value exceeds the preset deviation threshold, the semantic similarity of the corresponding segments of the source text and the target text is first calculated to find the difference segments with low semantic similarity. Then, the change, deletion or addition of words in the difference segments is analyzed to determine the tampering location, and then the degree of tampering is determined by calculating the semantic change range of the tampered words. Based on the judgment standard of semantic similarity, the key location and degree of text tampering can be accurately locked, so that the text tampering situation can be clearly presented.
[0014] In combination with some embodiments of the first aspect, in some embodiments, after the step of issuing rumor warning information to the information release platform, the method also includes: extracting rumor refuting verification elements from the source text, the rumor refuting verification elements including the time of the event, the location of the event, information about people related to the event, and the event description content; determining the information understanding characteristics of each target language area based on the factual cognitive difference, emotional tendency difference, and cognitive bias degree in the cognitive difference data; matching corresponding language expression rules in a preset language expression rule library based on the information understanding characteristics; generating rumor refuting content based on the language expression rules and the rumor refuting verification elements, the rumor refuting content including multiple languages.
[0015] By adopting the above technical solution, after issuing rumor warning information, rumor-refuting verification elements containing key information of the event are extracted from the source text, and then the information understanding characteristics of the target language area are determined based on cognitive difference data, and then the rules are matched with the preset language expression rule library. Finally, multilingual rumor-refuting content is generated by combining verification elements and rules, which solves the problem of rumor-refuting content generation, enhances the effectiveness of rumor-refuting, enables the public to obtain accurate information, and maintains the normal order of information dissemination.
[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of sending rumor warning information to the information release platform, the method also includes: obtaining time zone information of each target language area, determining the optimal release time interval based on the time zone information; and publishing the rumor-refuting content on the information release platform according to the optimal release time interval.
[0017] By adopting the above technical solution, the time zone information of each target language area is obtained after the early warning information is issued, and the user activity scores of different time periods are calculated through statistical analysis of historical user behavior, so as to determine the best release time interval. The time zone information provides the regional time basis, and the user activity analysis clarifies the time period when the audience is most likely to receive information. In this way, the rumor-refuting content is released on the information release platform, which can make the rumor-refuting information reach the target audience at the most appropriate time, improve the dissemination effect of the rumor-refuting information, and enhance the timeliness and influence of the rumor-refuting work.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of automatically batch publishing the rumor-debunking content on the information publishing platform according to the optimal publishing time interval, the method also includes: collecting user response data of the users to the rumor-debunking content; when the user response data is lower than a preset response threshold within a preset time period, stopping publishing the rumor-debunking content.
[0019] By adopting the above technical solution, after publishing the rumor-debunking content according to the optimal publishing time interval, the system will collect users' response data to the rumor-debunking content, including direct interaction data such as the number of likes and comments, as well as indirect interaction data and sentiment tendency data such as dwell time. By setting a specific preset time period and response threshold, if the score is lower than the threshold for two consecutive periods, the release will be automatically stopped to avoid the spread of invalid information and ensure that the rumor-debunking work is efficient and accurate.
[0020] In a second aspect, an embodiment of the present application provides a multilingual text cognitive rumor detection system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the multilingual text cognitive rumor detection system to execute the method described in the first aspect and any possible implementation method of the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions. When the above-mentioned computer program product is run on a multilingual text cognitive rumor detection system, the above-mentioned multilingual text cognitive rumor detection system executes the method described in the first aspect and any possible implementation method of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions. When the instructions are executed on a multilingual text cognitive rumor detection system, the multilingual text cognitive rumor detection system executes the method described in the first aspect and any possible implementation method of the first aspect.
[0023] It is understandable that the multilingual text recognition rumor detection system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. This application first obtains the cognitive response data of the source language area users to the source text and the cognitive response data of the target language area users to the target text, among which the understanding index, identification index and dissemination willingness index can fully reflect the user's cognitive state of the text, and compares the two to obtain cognitive difference data. When the information understanding deviation value exceeds the preset threshold, the difference fragment is located based on the semantic feature mapping relationship between the source text and the target text, and the tampering position and degree are determined. If the tampering degree and number of positions of the target text exceed the threshold, it is determined to be a rumor text and an early warning is issued. Comprehensive multi-faceted data and analysis can improve the accuracy of identifying rumor texts, effectively prevent the spread of rumors, and maintain the authenticity and stability of the information dissemination environment.
[0025] 2. This application obtains the target cognitive response data of users in the target language area to the target text and inputs it into the cognitive analysis model together with the source cognitive response data. The model can extract cognitive features including key information such as text fact consistency features, user emotional tendency features, and user cognitive bias features. By comparing and analyzing these features, the multi-dimensional difference is calculated to obtain cognitive difference data. By using the mutual correlation and difference comparison of multi-dimensional features, the cognitive difference nature of users in different regions to information can be deeply explored, making the cognitive difference data more accurate and comprehensive, and providing a reliable basis for subsequent rumor judgment.
[0026] 3. This application calculates the semantic similarity of the corresponding segments of the source text and the target text when the information understanding deviation value exceeds the preset deviation threshold, so as to find the difference segments with low semantic similarity. Then, the change, deletion or addition of words in the difference segments is analyzed to determine the tampering location, and then the degree of tampering is determined by calculating the semantic change range of the tampered words. Based on the judgment standard of semantic similarity, the key location and degree of text tampering can be accurately locked, so that the text tampering situation can be clearly presented. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flow chart of a multilingual text recognition rumor detection method in an embodiment of the present application; Figure 2 It is another flowchart of the multilingual text recognition rumor detection method in the embodiment of the present application; Figure 3 It is a schematic diagram of the structure of a physical device of a multilingual text recognition rumor detection system in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification of the present application, the singular expressions "one", "a kind of", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more of the listed items.
[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.
[0030] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.
[0031] During an international sports event in 2023, a piece of news about a well-known athlete's post-match interview sparked heated discussions on social media. The athlete expressed his views on the rules of the game in English, saying "The rules need to be reviewed". When the news spread to other language regions, the word "review" was interpreted differently by different language media. Some translations interpreted it as "questioning and denying the existing rules", while the original meaning was "suggesting a routine review of the rules". This difference in interpretation led to serious differences in the attitudes of users in different language regions towards the athlete: users in the English-speaking region believed that this was a rational suggestion, while users in other language regions believed that the athlete was questioning the fairness of the event. This cognitive difference quickly evolved into an online controversy, affecting the reputation of the event and the atmosphere of international sports exchanges. Traditional public opinion monitoring systems can only detect the increase in topic popularity, but cannot promptly identify cognitive biases in this cross-language communication process, resulting in continued misunderstandings.
[0032] The following introduces a scenario of using the multilingual text cognitive rumor detection method in related technologies.
[0033] A multinational company uses an existing text similarity analysis system to monitor the dissemination of its product launch information in the global market. The system mainly uses a word vector model to analyze the similarity of news reports in different languages. At a new product launch conference, the company's CEO said in English that the product has a "revolutionary" technological breakthrough. The system shows through text similarity analysis that the similarity of each language version is over 90%, and believes that the information is accurately disseminated. However, the actual situation is that in some language areas, the translation of the word "revolutionary" carries negative implications of "disruptive" and "destructive", causing local consumers to worry about the product. Despite the high text similarity, users' actual cognition and emotional reactions are very different. The existing system is unable to capture this subtle semantic difference and the resulting cognitive bias, causing the company to miss the opportunity to adjust its communication strategy in a timely manner, which ultimately affects the acceptance of the product in some markets.
[0034] The following describes a scenario in which the multilingual text recognition rumor detection method in this application is used.
[0035] An international news organization used the cognitive data monitoring method of this application to track a news report on environmental protection. The original English report said that a certain company "committed to gradual carbon reduction". The system collected cognitive response data from users in the English-speaking region, showing that users generally understood that this was a positive long-term commitment (understanding level 90%, recognition level 85%, and willingness to spread 75%). When the news spread to other language regions, the system detected that the cognitive response data of users in the target language region showed significant deviations (understanding level 60%, recognition level 45%, and willingness to spread 30%). The system immediately located the semantic deviation caused by the translation of the word "gradual". In some languages, the translation of this word implies "delay" and "inaction". The system promptly issued an early warning to the news organization, providing the specific location and degree of tampering, so that the news organization could quickly adjust the reporting terms and supplement the explanations, effectively avoiding cognitive differences between users in different language regions.
[0036] A global social media platform further optimized and applied this solution, combining the cognitive monitoring method with a real-time feedback mechanism. In the report of a major international event in 2024, the platform used the system to monitor the cognitive responses of multilingual users in real time. When a sensitive term is detected to cause cognitive differences in different language regions, the system not only issues an early warning, but also automatically generates multilingual explanations and displays supplementary information next to the relevant content through algorithmic recommendations. For example, when it is detected that the understanding deviation of the word "diplomatic pressure" in different language regions exceeds the preset threshold, the system automatically adds a multilingual annotation box next to the content to explain the specific meaning and usage scenarios of the word in different cultural contexts. At the same time, the system continuously collects user feedback on the supplementary explanations and dynamically adjusts the display method and depth of the explanation content. This optimized application enables the platform to proactively prevent cognitive bias, provide more accurate cross-cultural communication services, and significantly improve users' information acquisition experience and cross-cultural understanding level.
[0037] For ease of understanding, the following describes the process of the method provided by this implementation in combination with the above scenario. Figure 1 , which is a flow chart of the multilingual text recognition rumor detection method in the embodiment of the present application.
[0038] S101. Acquire source cognitive response data of users in a source language region to a source text, wherein the cognitive response data includes an understanding index, an identification index, and a dissemination intention index of the source text. The source text is transformed into a target text through an information dissemination process.
[0039] Among them, the source language area refers to the language usage area where the information was originally generated and disseminated; the source text refers to the original text content first published in the source language area; the cognitive response data refers to the quantitative data of the user's cognitive reactions to the text, such as understanding, recognition and willingness to disseminate; the understanding index is used to indicate the user's understanding of the text content; the recognition index is used to indicate the user's recognition of the text content; and the dissemination intention index is used to indicate the user's tendency to forward and share the text.
[0040] This step is executed when the system starts the text cognition rumor detection task, and is used to obtain the user cognition benchmark data of the source text in the original language area. Specifically, the system first determines the release location and time of the source text, and then collects the user's reading behavior data, interaction data, and dissemination behavior data on the text in the language area. By analyzing these data, the system calculates the quantitative indicators of the three dimensions of understanding index, identification index, and dissemination willingness index, forming a complete source cognition response data set.
[0041] In some embodiments, the acquisition of source cognitive response data can be achieved in the following ways: Optionally, the system uses a questionnaire survey method to design a questionnaire containing comprehension test questions, identity score and dissemination willingness multiple-choice questions, randomly selects users in the source language area for survey, and statistically analyzes the questionnaire results to obtain various indexes; Optionally, the system collects user behavior data such as reading time, likes, comments, and forwarding of the source text in real time through the social media platform API interface, analyzes user comment content in combination with the text semantic understanding model, and calculates various indexes. It is understandable that other data collection and analysis methods can also be used to obtain source cognitive response data.
[0042] S102, obtaining target cognitive response data of users in the target language region to the target text, comparing the source cognitive response data with the target cognitive response data to obtain cognitive difference data, wherein the cognitive difference data includes an information understanding deviation value.
[0043] Among them, the target language area refers to the language area used by the target audience after the information is disseminated; the target cognitive response data refers to the cognitive reaction data of users in the target language area to the text after dissemination; the cognitive difference data is used to indicate the cognitive differences of users in different language areas on the same information; the information understanding deviation value refers to the degree of deviation in the degree of text understanding of users in different language areas.
[0044] This step is performed after obtaining the source cognitive response data, and is used to evaluate the cognitive changes of information during cross-language communication. Specifically, the system first collects users' cognitive response data on the target text in the target language area, and quantifies it using the same indicator system as the source cognitive response data. The two sets of data are then input into the comparative analysis model, and by calculating the difference values of the indicators in each dimension, cognitive difference data including understanding bias, recognition bias, and communication intention bias are generated.
[0045] In some embodiments, the acquisition of cognitive difference data can be achieved in the following ways: Optionally, the system uses a deep learning model to analyze the semantic representation of the source text and the target text, and combines the user cognitive response data to calculate the cognitive difference in the text understanding dimension; Optionally, the system uses sentiment analysis technology to compare and analyze the emotional tendencies of users in different language regions, and combines user behavior data to quantitatively evaluate the differences in recognition and dissemination willingness. It is understandable that other analysis methods can also be used to calculate cognitive difference data.
[0046] This step specifically includes: Retrieving the source text in a cognitive database corresponding to the target language area to obtain relevant text; Extracting the key information overlap between each of the related texts and the source text, and determining the related texts whose information overlap is greater than the overlap threshold as target texts; Obtain target cognitive response data of users in the target language region to the target text, input the target cognitive response data into a cognitive analysis model, and obtain target cognitive features, wherein the target cognitive features include text fact consistency features, user emotional tendency features, and user cognitive bias features; The source cognitive response data is input into the cognitive analysis model to obtain source cognitive features, wherein the source cognitive features include fact consistency features, user emotional tendency features, and user cognitive bias features; Compare and analyze the target cognitive features and the source cognitive features, and calculate the difference in factual cognition, the difference in emotional tendency and the degree of cognitive bias; The cognitive difference data is calculated based on the factual cognitive difference, the emotional tendency difference and the cognitive bias degree.
[0047] Among them, the cognitive database refers to a structured database that stores the cognitive data of users in a specific language area; related text refers to the target language text that has content association with the source text; the key information overlap refers to the degree of matching between two texts in core information elements; the overlap threshold is used to indicate the minimum standard for determining text relevance; the cognitive analysis model refers to a machine learning model used to analyze user cognitive characteristics; the text fact consistency feature indicates the degree of conformity between the text and objective facts; the emotional tendency feature refers to the user's emotional attitude characteristics towards the text content; the cognitive bias feature is used to indicate the degree of deviation between the user's understanding and the original intention.
[0048] This step is performed after obtaining the source cognitive response data. It is used to find and confirm the corresponding text in the target language area and perform cognitive difference analysis. Specifically, the system first retrieves the text content related to the source text in the cognitive database of the target language area, and then screens out closely related target texts through information overlap analysis. For the determined target text, the system collects user cognitive response data and inputs it into the cognitive analysis model to extract cognitive features. At the same time, the source cognitive response data is input into the same model to obtain the source cognitive features. Finally, by comparing and analyzing the two sets of features, multi-dimensional cognitive difference data is calculated.
[0049] In some embodiments, the determination of the target text and the analysis of cognitive differences can be achieved in a variety of ways: Optionally, the system first uses multilingual text retrieval technology to retrieve relevant texts in the database, then calculates the text relevance through keyword matching and topic models, then uses information extraction technology to extract core information for overlap calculation, and then uses a deep learning model to analyze user cognitive features, and finally calculates cognitive differences through feature vectors; Optionally, the system first uses a cross-language text similarity algorithm to screen candidate texts, then applies semantic role labeling technology to analyze information structure consistency, then extracts user cognitive features through sentiment analysis and cognitive modeling, then calculates the difference values of each dimension based on feature comparison, and finally comprehensively evaluates to obtain cognitive difference data. It is understandable that other methods can also be used to achieve the determination of relevant texts and cognitive difference analysis, which are not limited here.
[0050] It should be noted that the core function of the cognitive analysis model is to analyze the user's understanding of the text. Specifically, this model focuses on three aspects: first, whether the user accurately understands the basic facts described in the text, such as "where something happened", "when exactly", "who is involved" and other key information; second, the user's emotional reaction after reading the text, such as whether it is positive, negative or neutral, by analyzing the user's comments, whether they like or forward, and other behaviors; third, whether the user's understanding is biased, such as whether some important details are exaggerated or ignored.
[0051] The model works as follows: First, it collects various user responses to the text, including comments, text added when forwarding, interactive behaviors, etc. Then, it analyzes the specific information contained in these responses, such as whether the factual details mentioned in the user's comments are accurate, whether the emotional attitude expressed is positive or negative, and whether the understanding of the original text is complete and accurate. Finally, the model will give a comprehensive score that reflects the user group's overall understanding of the text.
[0052] The process of building a cognitive analysis model can be divided into three main stages: data preparation, model training, and optimization and adjustment: In the data preparation stage, we first collect the input data required for training. The input data includes: original text content (including key elements such as time, place, and people), user comments on the text, added content when forwarding, interactive behavior data such as likes and collections. At the same time, for each training sample, professionals will annotate the reference answer, including: a list of key factual elements in the text, emotional labels expressed in user comments (positive, negative, neutral), and a score of the degree of deviation between the user's understanding and the original text.
[0053] The supervised learning method is used in the model training stage. The input layer of the model receives text content and user response data, which are input into three functional submodules after preprocessing such as text segmentation and key information extraction. The fact understanding submodule is responsible for extracting specific facts mentioned in user responses and matching them with the original text; the sentiment analysis submodule processes the sentiment vocabulary and interactive behavior data in user responses; and the cognitive bias submodule compares the key information differences between the original text and the user's understanding. The model continuously adjusts parameters based on the reference answers annotated by experts to optimize the analysis accuracy of these three dimensions.
[0054] During the optimization and adjustment phase, the model performance is continuously improved through test data from actual application scenarios. The final output of the model includes: fact understanding accuracy (a score between 0-1, indicating the accuracy of user understanding), sentiment tendency value (a score between -1 and 1, negative values indicate negativity, positive values indicate positivity), and cognitive bias (a score between 0-1, the larger the value, the greater the bias). These three output scores together constitute the quantitative evaluation results of the user's cognitive situation.
[0055] S103: When the information understanding deviation value exceeds a preset deviation threshold, the difference segment is located according to the semantic feature mapping relationship between the source text and the target text, and the tampering position and tampering degree of the difference segment are determined.
[0056] Among them, the preset deviation threshold refers to the allowable upper limit of information understanding deviation set in advance by the system; the semantic feature mapping relationship represents the semantic correspondence between the source text and the target text at the word, phrase and sentence levels; the difference fragment refers to the text fragment with semantic differences between the source text and the target text; the tampering location is used to indicate the specific location where the text is changed, deleted or added; the degree of tampering refers to the severity of the change in text content, including the magnitude of semantic changes and the scope of impact.
[0057] This step is triggered when the system detects that the information understanding deviation value exceeds the preset threshold, and is used to accurately locate and evaluate text changes. Specifically, the system first establishes the semantic feature vectors of the source text and the target text, and constructs a semantic mapping relationship between the two language texts through a deep learning model. Then, based on the semantic similarity calculation, text segments with significant semantic differences are identified. For each difference segment, the system analyzes the addition, deletion, and modification of words in it to determine the specific location of the tampering. At the same time, by calculating the semantic distance and contextual influence range of the words before and after the tampering, the severity of each tampering is evaluated. The system integrates these analysis results to form a complete text tampering assessment report.
[0058] In some embodiments, the location of the difference segment and the evaluation of the degree of tampering can be achieved in a variety of ways: Optionally, the system first performs word segmentation and syntactic analysis on the source text and the target text, constructs a semantic dependency tree, and then identifies structural differences through a tree structure matching algorithm, then uses word vectors to calculate the semantic similarity of words, and finally integrates syntactic and semantic information to determine the location and degree of the difference segment; Optionally, the system first uses bilingual alignment technology to establish text correspondence, then applies a text comparison algorithm to detect content changes, then uses semantic role labeling technology to analyze semantic changes, and finally combines contextual information to evaluate the impact of tampering. It is understandable that other text analysis and comparison methods can also be used to achieve difference detection and evaluation, which are not limited here.
[0059] This step specifically includes: Calculate the semantic similarity between corresponding segments of the source text and the target text; Mark the text segment whose semantic similarity is lower than the preset similarity threshold as a difference segment, analyze the number of words changed, deleted or added in the difference segment, and obtain the tampering position; The semantic change range of the tampered words in the difference segment is calculated to obtain the tampering degree of the tampered position.
[0060] First, the system will segment and align the source text and the target text according to sentences or semantic units. For each pair of aligned text segments, the semantic similarity between them is calculated. The calculation method is to convert the text segments into semantic vectors and then calculate the cosine similarity between the vectors (the value range is 0-1). For example, the semantic similarity between "It rained today" and "Today is a rainy day" may be 0.9, while the semantic similarity between "It rained today" and "Today is sunny" may be only 0.3.
[0061] Next, the system sets a preset similarity threshold (for example, 0.7). When the semantic similarity of a pair of text segments is lower than this threshold, it is marked as a difference segment. In each difference segment, the system finds the specific change location through word-level comparative analysis. For example, if "Today's temperature in Beijing is 25 degrees" is changed to "Today's temperature in Beijing is 35 degrees", the system will locate the specific location where "25" is changed to "35". By counting such change locations, the system can obtain all the tampering locations in the text.
[0062] Finally, the system will evaluate the degree of tampering at each tampered position. The specific method is to calculate the extent of the semantic change before and after the tampered word. For example, if the semantically opposite modification is changed from "slight" to "serious", the degree of tampering will be judged as high; while the modification with similar semantics, such as changing "today" to "now", will be judged as low. The system quantifies the degree of tampering by calculating the semantic vector distance between the original word and the new word, and finally obtains a value between 0-1. The larger the value, the more serious the tampering.
[0063] S104: If the tampering degree of the target text is greater than a preset degree threshold and the number of tampering locations is greater than a preset number threshold, the target text is determined to be a rumor text.
[0064] Among them, the preset degree threshold refers to the judgment standard of the degree of text tampering pre-defined by the system; the preset quantity threshold represents the upper limit of the number of text tampering positions set by the system; the degree of tampering refers to the severity score of the change in text content; the number of tampering positions is used to indicate the total number of tampered fragments in the text; rumor text refers to false information text that has been maliciously tampered with or distorted.
[0065] This step is performed after the text tampering analysis is completed, and is used to determine whether the text constitutes a rumor based on multi-dimensional thresholds. Specifically, the system first compares the detected degree of text tampering with the preset degree threshold, and at the same time counts the total number of tampering locations and compares it with the quantity threshold. When both indicators exceed their respective thresholds at the same time, it indicates that the text content has undergone significant and multiple changes, and the system will determine that the target text has the nature of a rumor. This dual-threshold judgment mechanism can effectively reduce the false positive rate and identify rumor texts with systematic tampering characteristics.
[0066] In some embodiments, rumor determination can be achieved in a variety of ways: Optionally, the system first calculates the semantic change score of each tampering, then performs a weighted summation of the scores of all tampered locations to obtain the overall tampering degree, then counts the number of tampered locations and evaluates their distribution density, and finally compares each indicator with a preset threshold to complete the determination; Optionally, the system first constructs a text tampering feature vector, then inputs a pre-trained rumor detection model for prediction, then verifies the prediction result in combination with rule judgment, and finally determines whether it is a rumor based on the comprehensive score. It is understandable that other analysis methods can also be used to achieve rumor identification and determination, which are not limited here.
[0067] S105. Send rumor warning information to the information release platform.
[0068] Among them, information release platform refers to online platforms such as social media and news websites that have information dissemination functions; rumor warning information refers to rumor identification and reminder information issued by the system; warning information contains key information such as the source of the rumor text, the scope of dissemination, and the degree of harm.
[0069] This step is executed immediately after the system confirms that the target text is a rumor, and is used to timely warn and control the spread of rumors. Specifically, the system generates a warning information package containing rumor identification results, tampering evidence, impact assessment, and disposal suggestions based on the characteristics and spread of the rumor text. The system pushes the warning information to the security management system of the relevant information release platform through the API interface, and sends a warning notification to the relevant regulatory authorities to prompt the platform to take timely control measures.
[0070] In some embodiments, the release of early warning information can be achieved in a variety of ways: optionally, the system first organizes the key data of rumor detection to form an early warning report, then encapsulates the early warning information according to the interface specifications of different platforms, then sends the early warning data through an encrypted channel, and finally tracks and confirms the receipt and processing status of the early warning information; optionally, the system first establishes an early warning information push queue, then sends the early warning information one by one in order of priority, then monitors the response of each platform, and finally records the early warning process log information. It is understandable that other methods can also be used to achieve the release of rumor early warning information, which is not limited here.
[0071] For example, a piece of news about "a multinational company entering some countries to carry out digital inclusion plans" was initially reported in a neutral and objective manner on Chinese social platforms, emphasizing that the project will promote local development through technological innovation. System analysis found that the cognitive response of Chinese users to this news reflected their concern for opportunities in scientific and technological development. Comments mostly focused on how digital technology can drive employment and technological progress. Users' understanding index, recognition index and willingness to spread all showed a positive trend. However, when this news spread to other language regions, the narrative perspective changed significantly. By analyzing the cognitive response data of users in the target language region, the system found that users' focus shifted to data security risks and market competition issues. Comments reflected concerns about technological dependence and economic control. Cognitive difference data showed that the information understanding deviation value reached 0.6. The system further compared the semantic features of the source text and the target text and found that the core differences were mainly reflected in the description of corporate motivations and project impacts. This difference was not a simple tampering of facts, but was due to differences in cognitive positions on the same business behavior in different cultural backgrounds. For example, "promoting the popularization of technology" was interpreted as "market penetration" during the dissemination process, and "creating employment opportunities" was understood as "controlling the local economy." Although this cognitive difference leads to a higher tampering index, the system identifies that this is information translation based on cultural cognitive differences rather than malicious rumor spreading. Therefore, the warning suggestions issued by the system focus on the integrity and background supplement of information dissemination to help users from different cultural backgrounds establish a more comprehensive understanding, rather than simply marking it as a rumor for processing.
[0072] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , is another flow chart of the multilingual text recognition rumor detection method in the embodiment of the present application.
[0073] S201. Send rumor warning information to the information release platform.
[0074] It can be understood that this step is similar to step S105 and will not be described in detail here.
[0075] S202: extract rumor-refuting verification elements from the source text, where the rumor-refuting verification elements include the time of the event, the location of the event, information about people involved in the event, and a description of the event.
[0076] The rumor-refuting verification factors include the time, location, personnel information and description content of the event, which are key data for verifying the authenticity of the information. The system uses multiple entity recognition and information extraction technology to extract factors during processing.
[0077] In the specific implementation process, the system first performs word segmentation and preprocessing on the source text. For the time element, the time expression recognition module is used to extract and convert various time expressions in the text (including exact time, relative time, time period, etc.) into a standard time format. The system determines the time reference point through context analysis, parses relative time expressions (such as "yesterday", "next week", etc.), and finally outputs standardized time information.
[0078] Location element extraction uses a geographic information recognition module and combines it with a geographic location knowledge base for processing. The system identifies place name entities in the text, establishes a hierarchical relationship of geographic locations (such as province-city-district), and normalizes the abbreviations and aliases of place names. For ambiguous location descriptions, the system completes and clarifies them through contextual information.
[0079] Personnel information extraction includes identifying information such as name, identity, position, and organization. The system uses the name recognition module to annotate name entities and extracts identity description information related to the person in combination with dependency syntax analysis. For organizational names, the system uses special organizational identification rules to extract and associate people with their organizations.
[0080] The extraction of event description content is based on semantic analysis technology. The system constructs the relationship between event elements through syntactic analysis and extracts key information such as the cause, process, and result of the event. For complex events, the system identifies the logical relationship between multiple event fragments and reconstructs the complete event chain. All extracted elements are stored and managed according to a unified data structure.
[0081] S203: Determine information comprehension features of each target language region based on the fact cognition difference degree, emotional tendency difference degree and cognitive bias degree in the cognitive difference data.
[0082] The system determines the information comprehension characteristics of the target language area based on cognitive difference data. This process includes three main steps: data standardization, feature calculation, and feature vector construction.
[0083] In the data standardization stage, the system normalizes the three indicators of factual cognition difference, emotional tendency difference and cognitive bias. By calculating the maximum and minimum values of each indicator, the values are mapped to a unified interval to eliminate dimensional differences. The standardized data is more suitable for subsequent feature calculations.
[0084] During the feature calculation process, the system first calculates the comprehensive cognitive difference value. By setting the weights of different indicators (0.4 for factual cognition, 0.3 for emotional tendency, and 0.3 for cognitive bias), the standardized indicators are weighted and summed. The weight values are determined based on expert experience and historical data analysis and can be adjusted according to actual application scenarios.
[0085] During the feature vector construction phase, the system generates a 48-dimensional feature vector for each language region. The understanding accuracy feature includes indicators such as text consistency and key information retention rate; the understanding depth feature includes indicators such as context relevance and logical reasoning ability; the understanding preference feature includes indicators such as expression preference and information acceptance tendency. The system forms a complete feature vector by quantitatively calculating various indicators.
[0086] S204: According to the information understanding feature, a corresponding language expression rule is matched in a preset language expression rule library.
[0087] Language expression rule matching is the process of matching information understanding features with the rules in the preset rule library. The system adopts a multi-level screening strategy to achieve accurate matching.
[0088] In the process of rule matching, the system first performs preliminary screening based on the dimensional characteristics of the feature vector. For each dimension, the system sets a corresponding matching threshold, such as 0.8 for understanding accuracy and 0.7 for understanding depth. Only rules that meet the threshold requirements of each dimension will enter the next round of screening.
[0089] The system uses the cosine similarity algorithm to calculate the similarity between the feature vector and the rule condition vector. The direction and size of the vector are taken into account during the calculation. The larger the similarity value, the higher the matching degree. The system sets the similarity threshold to 0.8 and selects the rules whose similarity exceeds the threshold as candidate rules.
[0090] For the selected candidate rules, the system optimizes and combines them. First, it checks the conflict relationship between the rules and deletes the conflicting rules. Then, it analyzes the complementarity of the rules and combines the complementary rules to form a rule sequence. Finally, it selects the top 5 rules with the highest matching degree as the execution rules.
[0091] The system applies the selected rules to the subsequent rumor-refuting content generation process and records the rule execution results for continuous optimization of the rule base. The rule selection and application process fully considers the language characteristics and expression habits of the target language area to ensure that the generated content meets the understanding characteristics of local users.
[0092] S205: Generate rumor-refuting content according to the language expression rule and the rumor-refuting verification element.
[0093] Language expression rules are a set of norms that define the expression methods, grammatical structures and rhetoric of different languages. The rumor-refuting verification elements include core information such as the time, location, personnel and description of the event. The rumor-refuting content is a multilingual rumor-refuting text generated based on the verification elements and language rules. Multilingual includes different language versions used by the target audience.
[0094] When the system generates rumor-debunking content, it first deconstructs the rumor-debunking verification elements into basic information units. For each target language, the system constructs a content template according to the corresponding language expression rules. The template contains sentence structure rules (such as the subject-verb-object structure in Chinese and the SOV structure in Japanese), rhetoric rules (such as positive arguments and rhetorical questions), and tone rules (such as rigorous statements and euphemistic explanations). The system reorganizes the information units according to the template rules to generate the initial text. Then, grammatical normalization is performed, including tense adjustment, addition of modal particles, optimization of conjunctions, etc. Finally, multiple rounds of text optimization are performed to ensure that the language expression is natural and fluent. For each language version, the system executes a complete generation process to ensure that the rumor-debunking content conforms to the expression characteristics of each language.
[0095] S206: Obtain time zone information of each target language region, and determine the best publishing time interval according to the time zone information.
[0096] Time zone information refers to the standard time zone data of the target language area, including the time difference from UTC and daylight saving time rules. The best publishing time interval refers to the time period when users in a specific area are most active. The target language area refers to a collection of geographical areas that use the same language.
[0097] The system first obtains the time zone information of each target language area from the time zone database. For language areas that span multiple time zones (such as English-speaking areas spanning the United States and the United Kingdom), the system obtains time zone data for major areas. Then it analyzes the user activity data for each area, which comes from historical user behavior statistics and includes user online rates, interaction rates, and content dissemination rates at different times of the day. The system classifies this data by hourly period and calculates the user activity score for each period. By setting an activity threshold (such as 80% of the peak), the optimal publishing time interval is determined. For areas with multiple time zones, the system selects a publishing interval that covers prime time in major time zones.
[0098] S207. Release the rumor-refuting content on the information release platform according to the optimal release time interval.
[0099] Information publishing platforms are social media or news platforms that support multilingual content publishing. The publishing process refers to the operational process of pushing rumor-refuting content to a designated platform at a scheduled time.
[0100] When the system executes a publishing task, it first establishes a publishing interface connection with each target platform to perform identity authentication and permission verification. Then the rumor-refuting content is organized into publishing task packages by language. Each task package contains information such as the publishing content, target platform, and publishing time. The system assigns the publishing task to the corresponding time point based on the optimal publishing time interval. When executing the release, the system encapsulates the text according to the platform's content format specifications, and adds necessary tags, classifications, and related information. The system submits the content to each platform through the API interface, monitors the publishing status at the same time, and retries failed publishing tasks. The system records the release time, platform, and status information of each version for subsequent effect tracking and analysis.
[0101] S208. Collect user response data to the rumor-refuting content.
[0102] User response data includes direct interaction data (such as likes, comments, and reposts) of users to the rumor-refuting content, indirect interaction data (such as dwell time, browsing depth, and bounce rate), and sentiment tendency data (such as comment sentiment analysis results and emoji response types). The collection process refers to the process by which the system continuously obtains this data through the platform API interface and data crawlers.
[0103] The system performs data collection in the following ways: First, a data collection task pool is established to create an independent data tracking task for each rumor-refuting content released. The system sets a collection interval of 15 minutes and regularly obtains the latest data from each platform. Direct interaction data directly obtains specific values through the platform API; indirect interaction data collects user behavior data through the monitoring code embedded in the content page; sentiment tendency data is obtained through real-time text analysis of user comments and responses. The collected raw data is cleaned (removing duplicate data and outliers) and formatted, and then stored in a unified data structure. The system summarizes the collected data in real time, calculates the current value and change trend of each indicator, and uses the analysis results to evaluate the effect of rumor refuting.
[0104] S209: When the user response data is lower than a preset response threshold within a preset time period, stop publishing the rumor-refuting content.
[0105] The preset time period refers to the fixed time period during which the system evaluates user response data, usually set to 4 hours. The preset response threshold refers to the minimum standard that user response data needs to meet within the time period, including the weighted comprehensive value of different types of response data. Stop publishing means that the system stops the continuous push of the rumor-refuting content on the relevant platform.
[0106] The specific process of the system's response evaluation and release control is as follows: The system calculates the user response comprehensive score every 4 hours. The calculation formula is: response score = 0.4 × direct interaction score + 0.3 × indirect interaction score + 0.3 × emotional tendency score. The direct interaction score is obtained by dividing the number of likes, comments, and reposts by the benchmark value (set according to the average level of the platform) and then summing them up; the indirect interaction score is obtained by weighted summing after standardizing indicators such as average dwell time and browsing completion rate; the emotional tendency score is calculated by the proportion of positive emotions. When the response score is lower than the preset threshold (set to 0.6) for two consecutive cycles, the system automatically stops the release operation: remove the content from the release queue, cancel the subsequent scheduled release tasks, and save the relevant data for subsequent analysis. For the published content, the system does not perform the removal operation, but only stops new release behavior.
[0107] The following describes the multilingual text recognition rumor detection system in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , which is a schematic diagram of the physical device structure of the multilingual text recognition rumor detection system in the embodiment of the present application.
[0108] It should be noted that Figure 3 The structure of the multilingual text recognition rumor detection system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0109] like Figure 3 As shown, the multilingual text recognition rumor detection system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 to the random access memory (RAM) 303, such as executing the method described in the above embodiment. In RAM 303, various programs and data required for system operation are also stored. CPU 301, ROM 302 and RAM 303 are connected to each other through bus 304. Input / output (I / O) interface 305 is also connected to bus 304.
[0110] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD) and an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.
[0111] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 309, and / or installed from a removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present invention are performed.
[0112] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.
[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings.
[0114] Specifically, the multilingual text cognitive rumor detection system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the multilingual text cognitive rumor detection method provided in the above embodiment is implemented.
[0115] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the multilingual text cognitive rumor detection system described in the above embodiment; or may exist independently without being assembled into the multilingual text cognitive rumor detection system. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of the multilingual text cognitive rumor detection system, the multilingual text cognitive rumor detection system implements the multilingual text cognitive rumor detection method provided in the above embodiment.
[0116] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0117] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.
[0118] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.
Claims
1. A multilingual text cognitive rumor detection method, characterized in that: Applied to a multilingual text cognitive rumor detection system, the method comprises: Acquiring source cognitive response data of users in a source language region to a source text, wherein the cognitive response data includes an understanding index, an identification index, and a dissemination willingness index of the source text, and the source text is transformed into a target text through an information dissemination process; Obtain target cognitive response data of users in a target language region to the target text, compare the source cognitive response data with the target cognitive response data to obtain cognitive difference data, wherein the cognitive difference data includes an information understanding deviation value; When the information understanding deviation value exceeds a preset deviation threshold, locating the difference segment according to the semantic feature mapping relationship between the source text and the target text, and determining the tampering position and tampering degree of the difference segment; If the tampering degree of the target text is greater than a preset degree threshold and the number of tampering positions is greater than a preset number threshold, the target text is determined to be a rumor text; Send rumor warning information to information release platforms.
2. The method according to claim 1, characterized in that The obtaining of target cognitive response data of users in the target language region to the target text, and comparing the source cognitive response data with the target cognitive response data to obtain cognitive difference data specifically includes: Obtain target cognitive response data of users in a target language region to the target text, input the target cognitive response data into a cognitive analysis model, and obtain target cognitive features, wherein the target cognitive features include text fact consistency features, user emotional tendency features, and user cognitive bias features; Inputting the source cognitive response data into the cognitive analysis model to obtain source cognitive features, wherein the source cognitive features include fact consistency features, user emotional tendency features, and user cognitive bias features; Comparative analysis is performed on the target cognitive features and the source cognitive features to calculate the difference in factual cognition, the difference in emotional tendency and the degree of cognitive bias; The cognitive difference data is calculated based on the factual cognitive difference, the emotional tendency difference and the cognitive bias degree.
3. The method according to claim 2, characterized in that Before the step of obtaining target cognitive response data of users in the target language area to the target text, the method further includes: Retrieving the source text in a cognitive database corresponding to the target language area to obtain relevant text; The key information overlap between each of the related texts and the source text is extracted, and the related texts whose information overlap is greater than an overlap threshold are determined as target texts.
4. The method according to claim 1, characterized in that When the information understanding deviation value exceeds a preset deviation threshold, the step of locating the difference segment according to the semantic feature mapping relationship between the source text and the target text, and determining the tampering position and tampering degree of the difference segment specifically includes: Calculate the semantic similarity between corresponding segments of the source text and the target text; Marking the text segments whose semantic similarity is lower than a preset similarity threshold as difference segments, analyzing the number of words that are changed, deleted or added in the difference segments, and obtaining the tampering position; The semantic change range of the tampered words in the difference segment is calculated to obtain the tampering degree of the tampered position.
5. The method according to claim 1, characterized in that After the step of sending rumor warning information to the information publishing platform, the method further includes: Extract rumor-refuting verification elements from the source text, wherein the rumor-refuting verification elements include the time of the event, the location of the event, information about people related to the event, and the event description content; Determining information comprehension characteristics of each target language region according to the fact cognition difference degree, emotional tendency difference degree and cognitive bias degree in the cognitive difference data; According to the information understanding feature, matching the corresponding language expression rules in the preset language expression rule library; The rumor-refuting content is generated according to the language expression rules and the rumor-refuting verification elements, and the rumor-refuting content includes multiple languages.
6. The method according to claim 5, characterized in that After the step of sending rumor warning information to the information publishing platform, the method further includes: Obtaining time zone information of each target language region, and determining an optimal publishing time interval according to the time zone information; The rumor-refuting content is published on the information publishing platform according to the optimal publishing time interval.
7. The method according to claim 6, characterized in that After the step of automatically publishing the rumor-refuting content in batches on the information publishing platform according to the optimal publishing time interval, the method further includes: Collecting user response data to the rumor-refuting content; When the user response data is lower than a preset response threshold within a preset time period, the rumor-refuting content is stopped from being published.
8. A multilingual text recognition rumor detection system, characterized in that: The multilingual text cognitive rumor detection system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the multilingual text cognitive rumor detection system to execute the method described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on the multilingual text-cognitive rumor detection system, the multilingual text-cognitive rumor detection system executes the method as described in any one of claims 1-7.
10. A computer program product, characterized in that When the computer program product runs on a multilingual text-cognitive rumor detection system, the multilingual text-cognitive rumor detection system executes the method as described in any one of claims 1-7.
Citation Information
Cited By
Communication clue analysis reminding method and system
CN121034319A