Method for generating news think tank data reports based on different languages

CN122528830APending Publication Date: 2026-08-07SHANDONG WISDOM YIBAI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG WISDOM YIBAI INFORMATION TECH CO LTD
Filing Date
2026-07-07
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

现有方法难以将候选隐含事件转换为内部监测对象,也难以根据后续公开新闻中实体、行为和对象的响应情况计算探针响应特征量,从而无法对候选隐含事件的可信度进行动态调整,也不便形成具有时间顺序和语种来源支撑的响应轨迹,降低了智库数据报告的可追溯性和研判可靠性

Benefits of technology

[0012]相较于现有技术,本发明的有益效果如下:(1)本发明通过目标事件类型匹配常识核对规则库,生成预期追问问题集合,并将其与各语种新闻中的实际追问问题进行结构化比对,能够识别各语种新闻均未覆盖的共同失问项。由此,智库数据报告不再仅依赖已报道的显性内容,还能够发现目标事件中可能被共同忽略的原因、责任对象或潜在风险线索,为后续候选隐含事件生成提供依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528830A_ABST
    Figure CN122528830A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of natural language processing and data mining, and relates to a news think tank data report generation method based on different languages. The present application obtains news of each language of a target event, identifies the target event type, matches a common sense checking rule library to generate a set of expected follow-up questions, extracts actual follow-up questions, and compares to obtain cross-language common missing items. The cross-language distribution differences of different entities are combined to generate candidate implied events and convert them into internal monitoring probe sequences. In the current time window, public news of each language is obtained, probe response characteristics are calculated, the credibility of the candidate implied events is adjusted according to the trigger criterion, and high-risk events to be output are marked, and then a think tank data report is generated. The present application solves the problem that the existing multilingual news report is difficult to find common missing problems and implied risk events, and achieves the effect of improving the report research depth and traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of natural language processing and data mining technology, and relates to a method for generating news think tank data reports based on different languages. Background Technology

[0002] With the increasing speed of dissemination of international events, regional security incidents, industrial policy events, and public emergencies, news texts have become a crucial data source for think tank analysis and situation assessment. Because the same event is often reported in different languages ​​by different countries, regions, or media outlets, news in different languages ​​often differs in reporting angle, focus, questioning style, and information disclosure pace. To generate data reports from multilingual news that can be used for decision-making, existing technologies typically employ machine translation, keyword retrieval, and topic clustering to produce event overviews, opinion summaries, or public opinion reports.

[0003] However, existing methods for generating multilingual news reports typically focus on extracting explicit, already reported content, such as the subject of the report, the sequence of events, emotional tone, and keyword popularity. They struggle to identify common questions missing across different language news reports on the same target event. For think tank analysis, questions not addressed in news reports across all languages ​​often correspond to the cause of the event, responsible parties, hidden stakeholders, or potential risks. Simply translating and summarizing existing content easily overlooks this shared cross-language omission, resulting in think tank data reports that remain superficial summaries and fail to effectively uncover potential hidden risks behind the target event.

[0004] Furthermore, existing technologies typically generate analytical conclusions directly upon detecting discrepancies in reports, lacking a mechanism for continuous verification based on subsequent publicly available news. Especially in scenarios with continuously updated multilingual news, determining whether a candidate risk event has truly escalated often requires considering the responses to publicly available news in different languages ​​within a subsequent time window. Existing methods struggle to convert candidate implicit events into internal monitoring targets and to calculate probe response characteristics based on the responses of entities, behaviors, and objects in subsequent publicly available news. This makes it impossible to dynamically adjust the credibility of candidate implicit events and to generate response trajectories supported by time sequence and language source, thus reducing the traceability and reliability of think tank data reports. Summary of the Invention

[0005] In view of this, in order to solve the problems mentioned in the background technology above, a method for generating news think tank data reports based on different languages ​​is proposed.

[0006] The objective of this invention can be achieved through the following technical solution: a method for generating news think tank data reports based on different languages, including: S1, acquiring news of a target event in various languages, identifying the type of the target event, matching it with a common sense check rule base, and generating a set of expected follow-up questions.

[0007] S2. Extract the actual follow-up questions from news articles in each language, compare the actual follow-up questions with the expected follow-up question set, and identify common missing questions that are missing across all languages.

[0008] S3. Extract the cross-language distribution differences of different entities from news articles in various languages, combine them with common missing items, generate candidate latent events, and convert them into internal monitoring probe sequences.

[0009] S4. Within the current time window, obtain publicly available news in various languages, execute the internal monitoring probe sequence, and calculate the probe response characteristic quantity.

[0010] S5. If the probe response feature quantity meets the trigger criterion, the confidence level of the corresponding candidate hidden event is increased and it is marked as a high-risk event to be output; if the probe response feature quantity does not meet the trigger criterion, the confidence level is decreased and monitoring continues.

[0011] S6. If the current time window ends and there are high-risk events to be output, extract common missing items, high-risk events to be output, and response trajectories to generate a think tank data report; if there are no high-risk events to be output, proceed to the next time window and re-execute S4.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention generates a set of expected follow-up questions by matching the target event type with a common sense check rule base, and compares it with the actual follow-up questions in news in various languages ​​in a structured manner, which can identify common missing questions that are not covered by news in various languages. As a result, the think tank data report no longer relies solely on the reported explicit content, but can also discover the reasons, responsible parties or potential risk clues that may be commonly overlooked in the target event, providing a basis for the generation of subsequent candidate implicit events.

[0013] (2) This invention combines common missing items with the cross-language distribution differences of entity objects to generate candidate latent events, and converts the candidate latent events into an internal monitoring probe sequence composed of entity probes, behavior probes, and object probes. By matching publicly available news in each language within a subsequent time window, calculating probe response feature quantities, and dynamically adjusting the credibility of candidate latent events, a think tank data report containing common missing items, high-risk events to be output, and response trajectories can be generated, improving the traceability and reliability of the report. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of the method for generating news think tank data reports based on different languages ​​in this invention; Figure 2 This is a diagram showing the relationship between the response features of candidate latent event probes and the triggering criteria in this invention. Figure 3 This is a diagram showing the correspondence between the high-risk event response fields to be output in this invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0018] The specific solution of the news think tank data report generation method based on different languages ​​provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0019] Please see Figure 1 As shown, the implementation of this invention includes S1 to S6: During the continuous dissemination of multilingual news, the reporting focus, questioning angle, and information disclosure pace of the same target event are not entirely consistent in different language news reports. Existing news think tank reports usually focus on translating, clustering, and summarizing existing content, easily overlooking common issues that are not followed up in different language news reports. These commonly missing issues may correspond to the causes of the event, the responsible parties, implicit related parties, or subsequent risk points. Therefore, it is necessary to establish a benchmark of follow-up questions adapted to the type of target event before generating the report, in order to determine which key issues in different language news reports have not been jointly addressed.

[0020] This invention first acquires news articles in various languages ​​corresponding to a target event and identifies the type of the target event. Then, it uses a common sense check rule base to generate a set of expected follow-up questions that match the type of the target event. Subsequently, it compares the actual follow-up questions in the news articles in various languages ​​with the set of expected follow-up questions to identify common missing items that are missing across languages. Furthermore, it combines the differences in cross-language distribution of different entities in the news articles in various languages ​​to generate candidate latent events and converts them into internal monitoring probe sequences. Finally, within a subsequent time window, it calculates the probe response feature based on the response to public news, adjusts the credibility of the candidate latent events, and generates a think tank data report.

[0021] S1. Obtain news about the target event in various languages, identify the type of the target event, match it with the common sense check rule base, and generate a set of expected follow-up questions.

[0022] Because the reporting angles, questioning methods, and information disclosure focuses of the same target event may differ in news reports across different languages, directly summarizing or statistically analyzing news articles in each language only yields explicit content that has already been reported or questioned. It's difficult to determine which questions should be raised at the common-sense level for a given event. Therefore, before identifying common unanswered questions across languages, it's necessary to first determine the event type based on its semantic content and then match the standard questioning text structure corresponding to that event type from a common-sense check rule base. This generates a set of expected follow-up questions, providing a unified benchmark for subsequent comparisons of actual follow-up questions.

[0023] First, the headlines and body text of news articles in each language are segmented into sentences. The headlines are treated as independent news sentences, and the body text is further segmented according to the positions of periods, question marks, exclamation marks, semicolons, and paragraph breaks to obtain multiple news sentences. For news sentences containing multiple predicate verbs after segmentation, they are further split into clauses according to the positions of commas, pauses, or conjunctions, and each clause is treated as a news sentence to be analyzed.

[0024] Subsequently, syntactic dependency analysis is performed on each news sentence to be analyzed using a natural language processing toolkit. The toolkit can be either spaCy or NLTK. spaCy performs word segmentation, part-of-speech tagging, and dependency parsing on the news sentences, while NLTK assists in word segmentation and part-of-speech tagging. After word segmentation and part-of-speech tagging, the word nodes, part-of-speech tagging results, dependency relation labels, and governing word nodes from the dependency parsing results are read. Words with dependency relation labels indicating a core predicate are identified as core predicate verbs. If the governing word node of a word node is a core predicate verb, and the dependency relation label between that word node and the core predicate verb is a subject-predicate relation label, then the word corresponding to that word node is identified as the subject. If the governing word node of a word node is a core predicate verb, and the dependency relation label between that word node and the core predicate verb is a verb-object relation label, then the word corresponding to that word node is identified as the object. Finally, the subject, predicate verb, and object of the same news statement to be analyzed are grouped into event triplets in sequence.

[0025] It is important to note that the subject represents the participants in the target event, the predicate verb represents the action in the target event, and the object represents the target object. Furthermore, if the same target event corresponds to multiple news articles in different languages, the event triples for each news article are extracted separately, and the subject, predicate verb, and object in the event triples are subjected to synonym normalization. Synonym normalization includes: first, converting the subject, predicate verb, and object from different languages ​​into the target language text; then, based on a thesaurus, a cross-language translation dictionary, and an entity alias table consisting of the full names, abbreviations, translations, and aliases of entities in the news articles from each language, determining whether different words represent the same subject, the same action, or the same target object. If they represent the same meaning, they are replaced with synonym normalized words. Event triples with identical three components after normalization are grouped and statistically analyzed; event triples whose frequency is greater than or equal to the average frequency of event triples are selected as the features to be matched for the target event.

[0026] Secondly, event triples are used as features to be matched and input into a standard event classification tree. Feature word comparison is performed in the order from parent node to child node. The standard event classification tree is built based on event type records in the common sense check rule base. Event type records include parent and child event types. The standard event classification tree uses parent event types as parent nodes and child event types as child nodes. Parent event types are obtained by combining subject feature words and behavior feature words, while child event types are obtained by combining subject feature words, behavior feature words, and object feature words.

[0027] During the comparison, the subject of the event triple is first compared with the subject feature words in the parent event type, and the predicate verb of the event triple is compared with the behavior feature words in the parent event type. If both the subject feature words and behavior feature words in the parent node are the same, the parent node is considered a successful match. Continuing under that parent node, the subject, predicate verb, and object of the event triple are compared with the corresponding subject feature words, behavior feature words, and object feature words in the lower event type, respectively, to obtain the child node matching results. The ratio between the number of matched feature words and the total number of feature words participating in the comparison is used as the node matching degree, and the child node category with the highest matching degree is taken as the target event type. If no successfully matched child node exists, the parent node category with the highest matching degree is taken as the target event type.

[0028] Next, the parent node category corresponding to the target event type is compared with the parent event type in the common sense check rule base, or the child node category corresponding to the target event type is compared with the child event type in the common sense check rule base. If the comparison results are the same, the event type record containing the parent or child event type is taken as the matching record, and the standard question text structure bound to the matching record is read as the standard question text structure that matches the target event type.

[0029] The common sense verification rule base is generated from historical news text records. First, statements representing the intent to inquire are extracted from these records as the statements to be processed. Then, a text vectorization tool converts these statements into semantic vectors. Semantic clustering is performed based on the cosine similarity between the semantic vectors. The mean of each cluster is calculated to extract the cluster center sentences whose sample size exceeds the mean. The subject and object of each cluster center sentence are replaced with placeholders to generate a standard question text structure. This standard question text structure includes interrogative elements, predicate verbs, subject placeholders, and object placeholders, used to represent common question patterns for the same event type. For example, if the cluster center sentence is "Who should be held responsible for Company A's acquisition of Company B," the interrogative element "who" is used as the subject placeholder, and "Company B" is used as the object placeholder. The generated standard question text structure is "The subject placeholder should be held responsible for Company A's acquisition of the object placeholder."

[0030] Subsequently, the frequency of the subject, predicate verb, and object in the news records corresponding to each standard question text structure was statistically analyzed. Words with frequencies higher than the corresponding average frequency were extracted as subject feature words, behavior feature words, and object feature words, respectively. Finally, the subject feature words and behavior feature words were combined into a superior event type, and the subject feature words, behavior feature words, and object feature words were combined into a subordinate event type. The superior event type and subordinate event type were combined into an event type record, and the event type record was bound to the standard question text structure to generate a common sense verification rule base.

[0031] Furthermore, statements representing the intent to inquire about elements are determined through semantic role labeling results. Specifically, semantic role labeling is performed on the sentence-segmented text of historical news records, extracting the subject, object, time component, location component, cause component, and responsible object component corresponding to the predicate verb. Subsequently, term detection is performed on the above components. If any component contains an interrogative pronoun, interrogative adverb, causal inquiry word, or responsibility inquiry word, it indicates that the text is used to inquire about event elements, and the text is identified as a statement representing the intent to inquire about elements, and is included as a statement to be processed in the commonsense check rule base construction process. If none of the above components contain an interrogative pronoun, interrogative adverb, causal inquiry word, or responsibility inquiry word, the text is not included as a statement to be processed.

[0032] Finally, based on the event triples corresponding to the target event, the subject and object of the event triples are read. The subject is filled into the subject placeholder in the standard question text structure, and the object is filled into the object placeholder in the standard question text structure, thus obtaining the expected follow-up questions for the current target event. If the same target event type corresponds to multiple standard question text structures, the placeholders are replaced separately, and the expected follow-up questions obtained after the replacement are summarized to generate an expected follow-up question set. Thus, the expected follow-up question set represents the questions that should be asked under the common sense check rule base for the current target event. It is used to compare with the actual follow-up questions in news articles of various languages ​​to identify common missing questions that are absent across languages.

[0033] S2. Extract the actual follow-up questions from news articles in each language, compare the actual follow-up questions with the expected follow-up question set, and identify common missing questions that are missing across all languages.

[0034] After obtaining the expected set of follow-up questions, this set only represents the questions that should be asked about the current target event according to common sense verification rules; it does not yet indicate whether these questions have actually appeared in news articles in various languages. Since news articles in different languages ​​may use different expression orders, interrogative words, and sentence structures, even if the follow-up questions are the same, there may be differences in the surface text. Therefore, it is necessary to first convert the actual follow-up questions in news articles in various languages ​​into a unified question text structure, and then compare it with the standard question text structure corresponding to the expected set of follow-up questions. This will filter out follow-up content not covered in all languages, serving as common missing questions for subsequently generating candidate implicit events.

[0035] First, the headlines and body texts of news articles in each language are segmented into sentences. Following the method for identifying sentences that represent the intent to inquire, semantic role labeling is performed on each segmented news sentence, extracting the subject, object, time component, location component, cause component, and responsible party component. Then, term detection is performed on these components. If any component contains an interrogative pronoun, interrogative adverb, causal inquiry word, or responsibility inquiry word, the news sentence is identified as a statement used to inquire about the subject, object, time, place, cause, or responsible party of the event, and is thus considered the actual inquiry question for each language news article.

[0036] The actual follow-up questions are derived from news articles in various languages ​​corresponding to the current target event, not from historical news text records. Each actual follow-up question retains at least the question text, source language, news record, and publication time. For multiple news articles in the same source language, the actual follow-up questions are identified separately and grouped according to the source language to facilitate subsequent determination of whether a particular question text structure appears in that language.

[0037] Secondly, semantic role labeling is performed on the actual follow-up questions, extracting the subject, object, and predicate verb. The subject and object are then replaced with placeholders identical to the standard question text structure, resulting in the actual question text structure. Replacing the subject and object with placeholders eliminates the impact of differences in specific entity names across different language news articles on the comparison of question structures; the predicate verb is retained to reflect the event or action addressed by the follow-up question. If the same actual follow-up question contains multiple subjects or multiple objects, corresponding actual question text structures are generated for each and associated with the source language of the actual follow-up question.

[0038] Subsequently, the standard question text structure corresponding to the expected set of follow-up questions is read, and the actual question text structure is compared with the standard question text structure. During the comparison, the predicate verbs in both the actual and standard question text structures are first normalized using synonymy, and then the question component types, predicate verb normalization results, and placeholder positions are compared. If the question component types are the same, the predicate verb normalization results are the same, and the subject and object placeholders correspond correctly in the question structure, then the actual question text structure is determined to match the standard question text structure; if any of the above is inconsistent, then no match is determined.

[0039] Finally, using the standard question text structure as the comparison object, the matching results of the standard question text structure in the corresponding actual question text structures of each source language are statistically analyzed. If a standard question text structure has a matching result in at least one source language's corresponding actual question text structure, it indicates that the news in that source language has actually pursued a follow-up question around that standard question text structure, and the standard question text structure is not identified as a common missing question item. If a standard question text structure has no matching result in any source language's corresponding actual question text structures, it indicates that the news in each language has not actually pursued a follow-up question around that standard question text structure, and the standard question text structure is identified as a common missing question item. Thus, common missing question items represent the follow-up question content that is missing across languages ​​in the current target event, and are used to combine with the cross-language distribution differences of different entities to generate candidate latent events.

[0040] S3. Extract the cross-language distribution differences of different entities from news articles in various languages, combine them with common missing items, generate candidate latent events, and convert them into internal monitoring probe sequences.

[0041] After identifying common missing questions, these terms only represent questions not pursued in news articles from different languages; they do not directly indicate which specific entity or event the missing content might point to. Therefore, it is necessary to further extract entity objects from the news articles in each language and compare the frequency of the same entity object across different source languages. If a particular entity object appears frequently in some language news articles but less frequently or not at all in others, it indicates an uneven distribution of that entity object across cross-language reports. Combining this unevenly distributed entity object with the common missing questions can generate candidate implicit events that require further verification.

[0042] First, entity objects are extracted from the headlines and body texts of news articles in various languages. The frequency of the same entity object in each language's news is then statistically analyzed according to its source language to obtain the frequency differences. Here, the source language refers to the text language of each news article, used to distinguish the reporting sources of the same target event in different language news articles. Entity objects, derived from the headlines and body texts of each language's news, refer to textual units that can serve as participants in an event, objects affected by an event, or related to the occurrence of an event, including names of people, organizations, locations, or events.

[0043] When extracting entity objects, the headlines and body texts of news articles in each language are first segmented into sentences, and then the segmented news sentences are segmented and tagged with parts of speech. If the part-of-speech tagging of consecutive words results in a person's name, organization's name, place name, or proper noun, then the consecutive words are extracted as entity objects. Similarly, if the subject or object of a news sentence is a person's name, organization's name, place name, or event name, it is also extracted as an entity object. For entity objects expressing the same meaning in different languages, synonym normalization is first performed, and entity objects with the same normalization result are considered the same entity object. Then, the frequency of the same entity object in the news articles of each source language is counted separately. If an entity object does not appear in a certain source language, its frequency in that source language is recorded as zero. Finally, the highest and lowest frequencies of the same entity object in each source language are read, and the difference between the highest and lowest frequencies is taken as the frequency difference of the entity object.

[0044] Secondly, the frequency differences of each entity object are read separately. The sum of these frequency differences is then divided by the number of entity objects to obtain the mean frequency difference. If the frequency difference of a particular entity object is higher than the mean, it indicates that the difference in the report distribution of that entity object across different source languages ​​exceeds the average difference level of entity objects in the current target event, and this entity object is identified as an anomalous entity object. Subsequently, placeholder markers in common missing questions are read, and the filling position is determined based on the grammatical components of the anomalous entity object in the original news sentence: if the anomalous entity object appears as the subject in the original news sentence, it is filled into the subject placeholder marker; if the anomalous entity object appears as the object or proper noun in the original news sentence, it is filled into the object placeholder marker. After filling, the interrogative words or follow-up questions in the common missing questions are transformed into declarative statements to obtain candidate implicit events composed of anomalous entity objects, event behaviors, and event objects. If the same common missing question corresponds to multiple anomalous entity objects, placeholder markers are filled separately to generate multiple candidate implicit events.

[0045] Finally, syntactic dependency analysis is performed on the candidate latent events to locate the core predicate verbs. Words with subject-verb dependencies on the core predicate verb are designated as subjects, and words with verb-object dependencies are designated as objects. The subjects are used as entity probes to check if related entities continue to appear in public news articles; the predicate verbs are used as behavior probes to determine if corresponding event behaviors appear in public news articles; and the objects are used as object probes to determine if the event behavior points to the same object. The entity probes, behavior probes, and object probes corresponding to the same candidate latent event are bound into a set of probe records and arranged according to the generation order of the candidate latent events, forming an internal monitoring probe sequence. This internal monitoring probe sequence is then used to match public news articles in various languages ​​within the current time window and to calculate probe response feature quantities.

[0046] Therefore, through the aforementioned steps, cross-language missing inquiry content has been identified from news articles in various languages. Candidate latent events have been generated by combining this with cross-language distribution differences of entities and objects. These candidate latent events are then decomposed into entity probes, behavior probes, and object probes. Since these candidate latent events are still derived from common missing inquiries and distribution differences, they cannot be directly used as report conclusions. Therefore, further verification using publicly available news articles in various languages ​​is required within a subsequent time window. Based on this, subsequent steps utilize the common responses of entities, behaviors, and objects in publicly available news to calculate probe response features. The credibility of the candidate latent events is then adjusted based on these features to determine whether to generate a think tank data report.

[0047] S4. Within the current time window, obtain publicly available news in various languages, execute the internal monitoring probe sequence, and calculate the probe response characteristic quantity.

[0048] After generating the internal monitoring probe sequence, the candidate latent events are still events to be verified, derived from common missing items and anomalous entity objects. It is necessary to determine whether they have received an external response by checking whether the same entities, behaviors, and objects appear in subsequent publicly available news. Therefore, within the current time window, publicly available news in various languages ​​is continuously acquired, and the internal monitoring probe sequence is used for consistent matching to obtain probe response features that reflect the subsequent manifestation of the candidate latent events.

[0049] First, publicly available news articles published within the current time window in various languages ​​are retrieved, and the titles and texts of these articles are segmented into sentences to obtain the news sentences to be matched. The current time window refers to the news collection time range corresponding to this round of monitoring, and each publicly available news article includes at least the news text, the source language, and the publication time. The news sentences to be matched are derived from publicly available news articles whose publication time falls within the current time window and are used for consistency matching with the internal monitoring probe sequence.

[0050] Secondly, probe records in the internal monitoring probe sequence are read one by one. These records include entity probes, behavior probes, and object probes. Syntactic dependency analysis is performed on each news statement to be matched, locating the core predicate verb and extracting the subject with a subject-verb dependency relationship, the object with a verb-object dependency relationship, and proper nouns within the news statement. Entity probes are matched against the subjects or proper nouns in the news statement; behavior probes are matched against the core predicate verb; and object probes are matched against the objects. If a news statement matches all three probes (entity, behavior, and object), it is considered a valid response statement corresponding to that probe record; otherwise, it is not considered a valid response statement.

[0051] Finally, for each probe record, the total number of valid response statements matched within the current time window is counted, along with the total number of source languages ​​containing the valid response statements. The total number of valid response statements represents the scale of the candidate implicit event's response in public news, while the total number of source languages ​​represents the spread of the response across different language news articles. Multiplying the total number of valid response statements by the total number of source languages ​​yields the probe response feature of the candidate implicit event corresponding to that probe record. This probe response feature is then used to subsequently determine whether the candidate implicit event meets the triggering criteria and adjust the credibility of the candidate implicit event accordingly.

[0052] like Figure 2 The diagram illustrates the correspondence between the total number of valid response statements, the total number of source languages, and the probe response feature values ​​for each candidate hidden event. The horizontal axis represents the total number of valid response statements, the vertical axis represents the total number of source languages ​​containing valid response statements, and the bubble size represents the probe response feature value, which is obtained by multiplying the total number of valid response statements by the total number of source languages. The dashed line in the diagram represents the trigger criterion boundary, which is determined based on the average value of the probe response feature values ​​corresponding to each candidate hidden event within the current time window. When the bubble corresponding to a candidate hidden event is above the trigger criterion boundary, it indicates that the probe response feature value of that candidate hidden event is greater than or equal to the average response feature value, and it can be determined to meet the trigger criterion in subsequent steps, and used to increase the confidence level of the corresponding candidate hidden event.

[0053] S5. If the probe response feature quantity meets the trigger criterion, the confidence level of the corresponding candidate hidden event is increased and it is marked as a high-risk event to be output; if the probe response feature quantity does not meet the trigger criterion, the confidence level is decreased and monitoring continues.

[0054] After obtaining the probe response feature values ​​corresponding to each candidate latent event, it is necessary to determine whether the public news response level of the candidate latent event within the current time window is higher than the average response level of all candidate latent events. If the response level reaches or exceeds the average response level, it indicates that the candidate latent event has a strong indication of being displayed in subsequent public news, and its credibility should be increased and included in the output objects; if the response level is lower than the average response level, it indicates that the candidate latent event has not yet generated a sufficient response within the current time window, and its credibility should be reduced and continued observation should continue.

[0055] First, within the current time window, the probe response feature values ​​corresponding to each candidate hidden event are read, and the number of candidate hidden events is counted. The sum of each probe response feature value is divided by the number of candidate hidden events to obtain the average response feature value, which is then used as the trigger criterion within the current time window. The probe response feature value corresponding to each candidate hidden event is obtained by matching the corresponding probe record, and the number of candidate hidden events is the total number of candidate hidden events participating in confidence adjustment within the current time window. If the number of candidate hidden events is zero, or if the probe response feature values ​​corresponding to each candidate hidden event are all zero, it is determined that there are no candidate hidden events satisfying the trigger criterion within the current time window, and monitoring continues in the next time window.

[0056] Secondly, if the probe response feature value corresponding to any candidate latent event is greater than or equal to the mean response feature value, then the candidate latent event is determined to meet the triggering criterion. Subsequently, the difference between the probe response feature value corresponding to the candidate latent event and the mean response feature value is calculated, and the ratio of the difference to the mean response feature value is used as the credibility increase ratio. The credibility of the candidate latent event is increased according to the credibility increase ratio, and the candidate latent event with increased credibility is marked as a high-risk event to be output. The credibility increase ratio indicates the extent to which the public news response level of the candidate latent event exceeds the current average response level; the high-risk events to be output are used to generate a think tank data report at the end of the current time window.

[0057] Finally, if the probe response feature value corresponding to any candidate latent event is less than the mean of the response feature values, the candidate latent event is determined not to meet the triggering criterion. Subsequently, the difference between the mean of the response feature values ​​and the probe response feature value corresponding to the candidate latent event is calculated, and the ratio of the difference between the mean of the response feature values ​​and the probe response feature value to the mean of the response feature values ​​is used as the credibility decay ratio. The credibility of the candidate latent event is reduced according to the credibility decay ratio, and the candidate latent event and its corresponding internal monitoring probe sequence are retained so that public news matching can continue in the next time window. Thus, the credibility of the candidate latent event can be dynamically adjusted according to the public news response situation in each time window, and provides high-risk events to be output for subsequent think tank data report generation.

[0058] S6. If the current time window ends and there are high-risk events to be output, extract common missing items, high-risk events to be output, and response trajectories to generate a think tank data report; if there are no high-risk events to be output, proceed to the next time window and re-execute S4.

[0059] After adjusting the credibility of each candidate implicit event within the current time window, it is necessary to determine whether to generate a think tank data report based on whether the current time window has ended and whether there are any high-risk events to be output. If there are high-risk events to be output, it means that the corresponding candidate implicit event has already received a public news response within the current time window, and its common source of failure, response statements, and time-series changes need to be compiled into the report content; if there are no high-risk events to be output, it means that no output-worthy high-risk events have yet formed within the current time window, and it is necessary to continue monitoring in the next consecutive time window.

[0060] First, the start time and length of the current time window are read, and the start time and length are added together to obtain the end time. The current acquisition time is compared with the end time; if the current acquisition time reaches or exceeds the end time, the current time window is considered to have ended; if the current acquisition time has not yet reached the end time, public news in various languages ​​continues to be acquired within the current time window, and internal monitoring probe sequence matching and probe response feature calculation are performed.

[0061] Secondly, if the current time window has ended and there are still high-risk events to be output, then extract the common missing questions, valid response statements, source language, and publication time corresponding to the high-risk events to be output. The common missing questions represent the original missing inquiry content corresponding to the high-risk event to be output; valid response statements are derived from news statements that simultaneously hit entity probes, behavior probes, and object probes; the source language and publication time are derived from publicly available news records where the valid response statements appear, representing the language range and chronological order of the response statements.

[0062] like Figure 3The diagram illustrates the correspondence between valid response statements, source languages, publication times, and the cumulative number of valid response statements for each high-risk event to be output. The horizontal axis represents the publication time of the valid response statement, the vertical axis represents the source language, each bubble represents a valid response statement, and the size of the bubble indicates the cumulative number of valid response statements up to that publication time. Figure 3 This allows us to determine which source languages ​​produce valid responses to the same high-risk event within the current time window, and the temporal order of these valid response statements. Therefore, Figure 3 The basis for generating response trajectory nodes is to combine each valid response statement, its corresponding source language, publication time, and the cumulative number of valid response statements up to that publication time into a response trajectory node.

[0063] Subsequently, valid response statements are arranged in chronological order of publication time, and a response trajectory node is generated for each valid response statement. Each response trajectory node includes the valid response statement, the corresponding source language, the publication time, and the cumulative number of valid response statements up to that publication time. The cumulative number of valid response statements is obtained by counting the number of valid response statements published within the current time window that are no later than the publication time of that response trajectory node. All response trajectory nodes are arranged in chronological order to generate the response trajectories corresponding to the high-risk events to be output.

[0064] Finally, common missing questions, high-risk events to be output, and response trajectories are written into the report fields to generate a think tank data report. Common missing questions describe which cross-language missing content triggered the high-risk event; high-risk events to be output describe the implicit risk objects that this report needs to highlight; and response trajectories describe the temporal changes in the public news response to the high-risk event within the current time window. If the current time window ends and there are no high-risk events to be output, the next consecutive time window is updated to the current time window, and the internal monitoring probe sequence is re-executed to continue matching and verifying subsequent public news in various languages ​​using the internal monitoring probe sequence.

[0065] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0066] Those skilled in the art will recognize that the algorithmic steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0067] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0068] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

[0069] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating news think tank data reports based on different languages, characterized in that: include: S1. Obtain news about the target event in various languages, identify the type of the target event, match it with the common sense check rule base, and generate a set of expected follow-up questions; S2. Extract the actual follow-up questions from news articles in each language, compare the actual follow-up questions with the expected follow-up question set, and identify common missing questions that are missing across all languages; S3. Extract the cross-language distribution differences of different entities from news articles in various languages, combine them with common missing items, generate candidate latent events, and convert them into internal monitoring probe sequences; S4. Obtain publicly available news in various languages ​​within the current time window, execute the internal monitoring probe sequence, and calculate the probe response characteristic quantity; S5. If the probe response feature quantity satisfies the trigger criterion, then increase the confidence level of the corresponding candidate hidden event and mark it as a high-risk event to be output. If the probe response feature does not meet the triggering criterion, the confidence level is lowered and monitoring continues. S6. If the current time window ends and there are high-risk events to be output, extract common missing items, high-risk events to be output, and response trajectories to generate a think tank data report; if there are no high-risk events to be output, proceed to the next time window and re-execute S4.

2. The method for generating news think tank data reports based on different languages ​​as described in claim 1, characterized in that, S1 specifically includes: Syntactic dependency analysis is performed on news articles in various languages ​​to extract the subject, predicate verb, and object, forming event triples; The event triples are used as features to be matched. Feature words are compared level by level in the standard event classification tree, and the node category with the highest matching degree is extracted as the target event type. Based on common sense, check the rule base to obtain the standard question text structure that matches the target event type; Extract the subject and object from the event triples, replace the placeholders in the standard question text structure, and generate the expected set of follow-up questions.

3. The method for generating news think tank data reports based on different languages ​​as described in claim 2, characterized in that, The method for obtaining the common sense verification rule base is as follows: Obtain historical news text records and extract statements representing key elements and intent as statements to be processed; Semantic clustering is performed on the statements to be processed, the cluster center sentences with a sample size greater than the mean are extracted, and the subject and object of the center sentences are replaced with placeholders to generate a standard question text structure. The frequency of the subject, predicate verb, and object in the news records corresponding to the text structure of each standard question was statistically analyzed, and words with a frequency higher than the corresponding average frequency were extracted as subject feature words, behavior feature words, and object feature words, respectively. The main feature words and behavior feature words are combined to form the superior event type. The main feature words, behavior feature words and object feature words are combined to form the subordinate event type. The superior event type and subordinate event type are combined to form the event type record. The event type record is bound to the standard question text structure to generate a common sense verification rule base.

4. The method for generating news think tank data reports based on different languages ​​as described in claim 2, characterized in that, The step of obtaining a standard question text structure that matches the target event type by checking the rule base based on common sense specifically includes: Read the event type records in the common sense verification rule base, where the event type records include parent event types and child event types; By using the higher-level event type as the parent node and the lower-level event type as the child node, a tree-like node structure is built according to the correspondence in the records of the same event type, forming a standard event classification tree.

5. The method for generating news think tank data reports based on different languages ​​as described in claim 4, characterized in that, The method for determining the intent of the statement representing the element is as follows: Semantic role labeling is performed on the text to identify the subject, object, time component, location component, cause component, and responsible party component in the sentences; If the subject, object, time element, place element, cause element, or object of responsibility element contains interrogative pronouns, interrogative adverbs, causal inquiry words, or responsibility inquiry words, then the text is identified as a statement representing the intent to inquire about the element.

6. The method for generating news think tank data reports based on different languages ​​as described in claim 5, characterized in that, S2 specifically includes: Sentence segmentation is performed on news articles in various languages ​​to identify sentences that represent the intent to follow up on key information, which are then used as the actual questions to be asked in each language's news articles. Replace the subject and object in the actual follow-up questions with placeholders to obtain the actual question text structure, and compare the actual question text structure with the standard question text structure in the expected follow-up question set; If the standard question text structure has no matching result in the actual question text structure in all languages, then the standard question text structure is identified as a common missing question item.

7. The method for generating news think tank data reports based on different languages ​​as described in claim 1, characterized in that, S3 specifically includes: Entity objects are extracted from news articles in various languages. These entity objects include names of people, organizations, places, or events. The frequency of the same entity object in news articles from different source languages ​​is counted, and the difference between the highest and lowest frequency of occurrence is taken as the frequency difference. Entities whose frequency difference is higher than the average frequency difference are identified as abnormal entities, and abnormal entities are filled into placeholders in common missing items to generate candidate hidden events. Extract the subject, predicate verb, and object from the candidate hidden events, and use the subject, predicate verb, and object as entity probes, behavior probes, and object probes respectively to form an internal monitoring probe sequence.

8. The method for generating news think tank data reports based on different languages ​​as described in claim 1, characterized in that, S4 specifically includes: By utilizing internal monitoring probe sequences, consistent matching is performed on publicly available news in various languages ​​within the current time window, and news statements that simultaneously contain entity probes, behavior probes, and object probes are extracted as valid response statements. The total number of valid response statements is counted, and the total number of source languages ​​containing valid response statements is counted. The total number is multiplied by the total number of source languages ​​to obtain the probe response feature.

9. The method for generating news think tank data reports based on different languages ​​as described in claim 1, characterized in that, S5 specifically includes: Within the current time window, read the probe response feature values ​​corresponding to each candidate hidden event. Sum the probe response feature values ​​and divide by the number of candidate hidden events to obtain the average response feature value, which is used as the triggering criterion. If the probe response feature value corresponding to any candidate hidden event is greater than or equal to the mean of the response feature value, the triggering criterion is met. The difference between the probe response feature value and the mean of the response feature value is calculated, and the ratio of the difference between the probe response feature value and the mean of the response feature value to the mean of the response feature value is used as the confidence increase ratio. The confidence of the corresponding candidate hidden event is increased according to the confidence increase ratio and marked as a high-risk event to be output. If the probe response feature value corresponding to any candidate hidden event is less than the mean of the response feature value, it is determined that the trigger criterion is not met. The difference between the mean of the response feature value and the probe response feature value is calculated, and the ratio of the difference between the mean of the response feature value and the probe response feature value to the mean of the response feature value is used as the confidence decay ratio. The confidence of the corresponding candidate hidden event is reduced according to the confidence decay ratio and monitoring continues.

10. The method for generating news think tank data reports based on different languages ​​as described in claim 1, characterized in that, S6 specifically includes: Read the start time and length of the current time window, add the start time and length to get the end time, and determine the end time of the current time window if the current acquisition time reaches or exceeds the end time. If the current time window has ended and there are still high-risk events to be output, then extract the common failure items, valid response statements, source language, and publication time corresponding to the high-risk events to be output; Valid response statements are arranged in chronological order of publication time. Response trajectory nodes are formed by combining the valid response statements, their corresponding source languages, publication times, and the cumulative number of valid response statements up to the publication time. Response trajectory nodes are then arranged in order to generate a response trajectory. Write common missing items, high-risk events to be output, and response trajectories into the report fields to generate a think tank data report; If the current time window ends and there are no high-risk events to be output, the next consecutive time window will be updated to the current time window, and S4 will be re-executed.