News topic generation method
By generating event clusters and dividing them into levels, calculating credibility and dynamic influence, and constructing cross-level relationship signals, this technology solves the problem of lack of depth and innovation in news topics in existing technologies, and realizes the scientific generation and screening of high-value topics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to uncover high-value news angles, relying heavily on shallow data for statistical sorting, which lacks depth and innovation.
Event clusters are generated by acquiring multi-source data, candidate opinion sentences are identified and structured data is generated and divided into multiple levels. Based on the long-term behavioral data of the opinion subject and the indicator information associated with multi-source data, the credibility and dynamic influence of the event are calculated, cross-level relationship signals are constructed, multiple candidate perspectives are generated and scored and ranked.
It enables in-depth exploration of high-value news topics, ensuring the objectivity and completeness of the topics, improving the quality and efficiency of news topic selection, and providing scientific decision-making support.
Smart Images

Figure CN121615634B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of news media technology, and in particular to methods for generating news topics. Background Technology
[0002] As the core starting point of news production, television news topic selection directly determines the value orientation, dissemination effect, and audience acceptance of the report. In today's era of rapid development of new media technology and increasingly diversified information dissemination, news leads have expanded from traditional media channels to multiple online platforms, including government announcements, social media, and industry reports, forming a massive and heterogeneous information ecosystem. However, many related technologies focus on statistically ranking superficial data such as event popularity, topic traffic, or public sentiment, making it difficult to delve into high-value topic angles.
[0003] Currently, no effective solution has been proposed to address the difficulty in deeply exploring high-value topic angles in related technologies. Summary of the Invention
[0004] This application provides a method for generating news topics, which at least solves the problem in related technologies that it is difficult to deeply explore high-value topic angles.
[0005] This application provides a method for generating news topics, the method comprising:
[0006] Acquire multi-source data and generate event clusters based on the multi-source data; the event cluster is an aggregation of information about the same event;
[0007] Candidate opinion sentences are identified from the document set of the event cluster, and structured data is generated; the structured data includes the opinion subject, opinion content, opinion stance, and opinion source information;
[0008] Based on the identity characteristics and behavioral patterns in the opinion source information, as well as the opinion subject, the structured data is divided into multiple levels; for each level, credibility is obtained based on the long-term behavioral data of the opinion subject; and the dynamic influence of the event is obtained based on the indicator information associated with the multi-source data.
[0009] Based on the representative text content of the event cluster, the event propagation attributes are identified, and a scene vector is constructed; based on the scene vector, the key level missing judgment result is obtained; based on the key level missing judgment result, the penalty factor is calculated.
[0010] For each level, the level value contribution is calculated based on the credibility and the dynamic influence of the event; based on the level value contribution of each level and the penalty factor, the comprehensive hierarchical value is obtained.
[0011] Based on the event clusters and the structured data, a cross-layer relationship signal is constructed; based on the cross-layer relationship signal, the hierarchical value contribution, and the structured data, multiple candidate perspectives are generated;
[0012] Based on the comprehensive hierarchical value, the candidate angles are scored and ranked, and the top N candidate angles with the highest scores are selected as the final candidate angles; N is a positive integer.
[0013] In some embodiments, the hierarchy includes an authority layer, a professional layer, a social layer, and a grassroots layer; the construction of cross-layer relationship signals based on the event clusters and the structured data includes:
[0014] Based on the viewpoints and positions in the structured data corresponding to each level, conflict signals are obtained;
[0015] Based on the event cluster, the growth rate of social layer discussions and the intensity of content updates at the authority layer are calculated to obtain gap signals;
[0016] Based on the representative viewpoints in the structured data corresponding to the professional layer and the representative viewpoints in the structured data corresponding to the basic layer, a mutual verification signal is obtained;
[0017] Based on the indicator information associated with the multi-source data, the original heat value is calculated; based on the original heat value and the smoothed heat value of the previous moment, the smoothed heat value of the current moment is obtained; based on the smoothed heat value of the current moment, the smoothed heat value of the previous moment, and the trend sensitivity threshold, a trend signal is obtained.
[0018] Based on the conflict signal, the gap signal, the mutual verification signal, and the trend signal, a cross-layer relationship signal is constructed.
[0019] In some embodiments, obtaining conflict signals based on viewpoints and positions in the structured data corresponding to each level includes:
[0020] Based on the viewpoints and positions in the structured data corresponding to each level, a position distribution of viewpoints at each level is formed;
[0021] Based on the position distribution, calculate the position difference degree between any two levels;
[0022] Based on the aforementioned degree of difference in positions, a conflict signal is obtained.
[0023] In some embodiments, obtaining the mutual verification signal based on representative viewpoints in the structured data corresponding to the professional layer and representative viewpoints in the structured data corresponding to the grassroots layer includes:
[0024] Semantic similarity is calculated based on representative viewpoints in the structured data corresponding to the professional layer and representative viewpoints in the structured data corresponding to the basic layer.
[0025] Based on the semantic similarity and the preset credibility weight, a mutual verification signal is obtained.
[0026] In some embodiments, the cross-layer relationship signal includes conflict signals, gap signals, mutual verification signals, and trend signals; the generation of multiple candidate perspectives based on the cross-layer relationship signal, the hierarchical value contribution, and the structured data includes:
[0027] Based on the cross-layer relationship signal and the hierarchical value contribution, the signal trigger intensity corresponding to the conflict signal, the gap signal, the mutual verification signal and the trend signal are calculated respectively; when the signal trigger intensity is greater than a preset threshold, multiple topic selection angle signals are triggered;
[0028] Based on the multiple topic selection angle signals, the structured data, and the hierarchical value contribution, multiple candidate angles are generated; the multiple candidate angles correspond to different types of cross-level relationship signals.
[0029] In some embodiments, the step of scoring the candidate angles based on the comprehensive hierarchical value and selecting the top N candidate angles as the final candidate angles includes:
[0030] Based on the candidate perspectives, an innovation score is generated;
[0031] Based on the comprehensive hierarchical value and the innovation score, a comprehensive score is given to multiple candidate angles, and the top N candidate angles with the highest comprehensive scores are selected as the final candidate angles.
[0032] In some embodiments, generating an innovation score based on the candidate angles includes:
[0033] Based on the candidate angles and historical topic selection angles, a content novelty score is obtained;
[0034] Based on the candidate angles, calculate the information entropy; based on the information entropy, obtain a unique perspective score.
[0035] Based on the candidate perspectives and the preset narrative template library, a score for formal innovation is generated;
[0036] An overall innovation score is obtained based on the content novelty score, the perspective uniqueness score, and the form innovation score.
[0037] In some embodiments, the process of obtaining credibility for each level based on long-term behavioral data of opinion subjects in the multi-source data includes:
[0038] For each level, a historical score is calculated based on the long-term behavioral data of the opinion subjects in the multi-source data; the credibility is obtained based on the preset static credibility and the historical score.
[0039] And / or, obtaining the dynamic impact of the event based on the indicator information associated with the multi-source data includes:
[0040] Based on the indicator information associated with the multi-source data, the original event impact is calculated;
[0041] Based on the propagation content and propagation behavior data of the event cluster, the emotional polarization index, text homogenization, and short-term abnormal outbreak characteristics are extracted to obtain the emotional noise suppression factor.
[0042] Based on the emotional noise suppression factor and the original event influence, the dynamic influence of the event is obtained.
[0043] In some embodiments, identifying candidate opinion sentences from the document set of the event cluster and generating structured data includes:
[0044] Identify opinion sentences from the document set of the event cluster to obtain candidate opinion sentences; input the candidate opinion sentences into a preset large language model to obtain initial structured data;
[0045] Based on the viewpoints contained in the initial structured data, representative viewpoints and supporting evidence are obtained;
[0046] The structured data is generated based on the representative viewpoints, supporting evidence, viewpoint source information associated with the multi-source data, and the initial structured data.
[0047] In some embodiments, after selecting the top N candidate angles with the highest scores as the final candidate angles, the method further includes:
[0048] The final candidate angles are subjected to entity recognition to extract place name and institution entities;
[0049] The place name institutions are mapped to a localized geographic knowledge graph, and the graph hop distance and semantic relevance between the place name institutions and the local core nodes in the knowledge graph are calculated. The knowledge graph includes local administrative division nodes, pillar industry related nodes, key livelihood project nodes and place name alias mapping relationships, forming a structured knowledge network covering local core related elements.
[0050] Based on the hop distance and semantic relevance of the graph, the association level between the final candidate angle and the local core node is determined;
[0051] Based on the aforementioned association level, localized candidate angles are generated;
[0052] The localized candidate angles are input into a preset feasibility scoring model to evaluate spatial accessibility, resource matching degree, time feasibility and news value, and an execution suggestion card is generated.
[0053] Compared to related technologies, the news topic generation method provided in this application obtains multi-source data and generates event clusters based on the multi-source data; an event cluster is an information aggregation of the same event; candidate opinion sentences are identified from the document set of the event cluster, and structured data is generated; the structured data includes opinion subject, opinion content, opinion stance, and opinion source information; based on the identity characteristics and behavioral patterns in the opinion source information, as well as the opinion subject, the structured data is divided into multiple levels; for each level, credibility is obtained based on the long-term behavioral data of the opinion subject; the dynamic influence of the event is obtained based on the indicator information associated with the multi-source data; and the event dissemination is identified based on the representative text content of the event cluster. The method involves: identifying attributes and constructing scene vectors; obtaining key level missing judgment results based on scene vectors; calculating penalty factors based on key level missing judgment results; calculating the level value contribution for each level based on credibility and event dynamic influence; obtaining a comprehensive hierarchical value based on the level value contribution of each level and the penalty factors; constructing cross-level relationship signals based on event clusters and structured data; generating multiple candidate angles based on cross-level relationship signals, level value contributions, and structured data; scoring and ranking the candidate angles based on the comprehensive hierarchical value, and selecting the top N candidate angles with the highest scores as the final candidate angles. This method solves the problem of difficulty in deeply mining high-value topic angles in related technologies.
[0054] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0055] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0056] Figure 1 This is a hardware structure block diagram of a terminal for a news topic generation method according to an embodiment of this application;
[0057] Figure 2 This is a flowchart of a news topic generation method according to an embodiment of this application;
[0058] Figure 3 This is a schematic diagram illustrating the calculation principle of the gap signal according to an embodiment of this application;
[0059] Figure 4 This is a schematic diagram illustrating the principle of conflict signal calculation according to an embodiment of this application;
[0060] Figure 5 This is a schematic diagram of the structured viewpoint matrix generation process according to an embodiment of this application;
[0061] Figure 6 This is a schematic diagram of the overall process of the news topic generation method according to an embodiment of this application;
[0062] Figure 7 This is a schematic diagram of the framework of a news topic generation system according to an embodiment of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0064] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0065] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application means two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The terms “first,” “second,” “third,” etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0066] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure block diagram of a terminal for a news topic generation method according to an embodiment of this application. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. Optionally, the terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0067] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the news topic generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0068] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0069] This embodiment provides a method for generating news topics. Figure 2 This is a flowchart of a news topic generation method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0070] Step S201: Obtain multi-source data and generate event clusters based on the multi-source data; an event cluster is an aggregation of information about the same event;
[0071] Specifically, the system first connects to multiple information sources, including news websites, government information platforms, social media, video platforms, and industry databases, through a multi-source collector consisting of an API adapter, a crawler engine, a stream processor, and a data cleaner. Then, the preprocessed data undergoes standardization, including text segmentation, language detection, and encoding normalization. Entity disambiguation is performed by constructing an entity alias mapping table and a semantic matching model, such as mapping different references to the same entity, like "municipal government" or "city government," to a standard entity. Next, a dual similarity calculation method is used. First, a 64-bit text fingerprint is generated using the SimHash algorithm for rapid initial screening. Then, semantic similarity is calculated based on a deep learning model, and the final similarity is obtained by linearly combining configurable weight parameters. When this similarity exceeds a preset threshold, similar documents are aggregated into event clusters using the HDBSCAN algorithm. Each event cluster generates a unique event ID and contains structured information such as a document list, time span, and keyword / topic set, forming an information aggregate of the same event.
[0072] Step S202: Identify candidate opinion sentences from the document set of the event cluster and generate structured data; the structured data includes opinion subject, opinion content, opinion stance and opinion source information.
[0073] First, a sentence-level in-depth analysis was conducted on the event cluster document set formed by the fusion of multi-source data. Opinion sentences were screened based on three core judgment criteria: subjectivity, stance, and structureability. The subjectivity criterion focused on sentences containing attitudinal, evaluative, or emotional words. The stance criterion targeted expressions that reflected clear attitudes such as support, opposition, or concern. The structureability criterion required that sentences could extract the core structure of "subject-object-stance". At the same time, it was clarified that objective fact sentences and pure data sentences were not considered opinion sentences, and sentences containing stance components in quotations or paraphrases were still treated as opinion sentences, thereby accurately identifying candidate opinion sentences.
[0074] Subsequently, candidate opinion sentences are input into a pre-defined large language model. Under the constraints of the news element layer, they are uniformly transcribed into structured triplets containing opinion subjects (individuals, organizations, institutions, media, etc.), opinion content (specific evaluations or attitudes expressed around the topic / object), and opinion stance (support, opposition, neutrality, pending, etc.). Stance labels and corresponding confidence scores (used to characterize the clarity and stability of stance determination) are output. For opinions with highly similar semantics, deduplication and merging are performed, and the sentence with the highest cluster center or weight is selected as the representative opinion. The remaining sentences in the cluster are used as supporting evidence, retaining the original text, source platform, and timestamp. Then, the opinion source information (including platform type, original document ID, publication time, capture timestamp, and other metadata) is automatically extracted and bound throughout the process by the collector in the multi-source data access stage. Finally, the data is integrated to generate structured data containing opinion subjects, opinion content, opinion stance, and opinion source information. This structured data has a complete source index, providing a standardized foundation for subsequent hierarchical division, credibility calculation, and other steps.
[0075] Step S203: Based on the identity characteristics and behavioral patterns in the opinion source information, as well as the opinion subject, the structured data is divided into multiple levels; for each level, the credibility is obtained based on the long-term behavioral data of the opinion subject; and the dynamic influence of the event is obtained based on the indicator information associated with multi-source data.
[0076] Specifically, we first extract identity features from the source information of viewpoints in the structured data, such as platform type, account type, authentication information, and attributes of the publishing entity. Then, we combine these with the behavioral patterns of the viewpoint entity in historical data, such as the frequency of publication, content type, and stability, and associate them with the attributes of the viewpoint entity itself. This allows us to accurately divide the structured data into four predefined levels: the authoritative level (government announcements, official media, etc.), the professional level (experts, scholars, industry analysis, etc.), the social level (social media discussions, forum opinions, etc.), and the grassroots level (personal statements, on-site witnesses, etc.).
[0077] For each level, a historical score (covering dimensions such as factual consistency, debunking rate, content quality and stability) is calculated based on the account entity of the information source on the specific publishing platform, taking its long-term behavioral data as the unit. This score is then combined with the pre-set static credibility of each level to obtain the final credibility of each level. At the same time, the original event influence is calculated based on the related indicator information such as the number of reads, reposts, and comment growth rates collected from multi-source data. Then, features such as emotional polarization index, text homogenization, and short-term abnormal outbreaks are extracted from the dissemination content and dissemination behavior data of the event cluster to generate emotional noise suppression factors. These factors are used to denoise and correct the original event influence, ultimately obtaining the dynamic influence of the event that can reflect the true dissemination capability of the event.
[0078] Step S204: Based on the representative text content of the event cluster, identify the event propagation attributes and construct a scene vector; based on the scene vector, obtain the key level missing judgment result; based on the key level missing judgment result, calculate the penalty factor.
[0079] Specifically, representative text content such as news articles, authoritative release summaries, and representative viewpoints that have been aggregated in the event cluster are extracted, and after unified preprocessing, they are input into a multi-label classification model to identify the degree of association of events in three communication scenarios: politics (P), people's livelihood and public concerns (L), and industry technology or in-depth analysis (D). The corresponding scenario probability values are output and a normalized scenario vector Scene=(P,L,D) is constructed (satisfying P+L+D=1).
[0080] Subsequently, based on the scenario vector, the core attributes of the current event are identified, thereby determining the key levels supporting the authenticity of the event and the completeness of the report. Political events correspond to the authoritative level, livelihood events correspond to the grassroots level, and in-depth industry events correspond to the professional level. By verifying whether there is valid viewpoint data at each key level, the system obtains the judgment result of the missing key level. If it is determined that there is a missing key level, the system will calculate the corresponding quantitative penalty factor based on the importance of the missing level, the event type, and the data completeness requirements. This penalty factor will be used for subsequent deduction calculation of the comprehensive layered value to ensure that the generated topics comply with the principle of integrity in news reporting.
[0081] Step S205: For each level, calculate the level value contribution based on credibility and event dynamic influence; based on the level value contribution and penalty factor of each level, obtain the comprehensive hierarchical value.
[0082] First, based on the event-driven scene vector Scene=(P,L,D) (politics / people's livelihood / in-depth analysis), the credibility importance coefficient α_scene and the influence importance coefficient β_scene under the current scene are calculated through a preset mapping function. Then, for the four levels of authority, professional, social, and grassroots, the final credibility (C) of each level is calculated. final (Weighted by static credibility and historical ratings) and event dynamic influence (I) after correction by emotional noise suppression factor. final As the core input, it is substituted into the layer value contribution calculation formula layer_value(i)=α_scene·C final (i)+β_scene·I final (i) Calculate the comprehensive value contribution of each level in the current event, where i is the level index; C final (i) represents the final credibility of level i; I final(i) represents the dynamic influence of the event at level i; α_scene represents the credibility importance coefficient; β_scene represents the influence importance coefficient; layer_value(i) represents the value contribution of the level, indicating the overall contribution of the level in the current event.
[0083] Subsequently, by combining the semantic relevance (i) of each level of viewpoint with the target column, the value contribution of the corresponding level is weighted and adjusted. Then, the weighted value contributions of all levels are summarized, and finally, the quantitative penalty factor (Penalty, triggered when the authoritative level is missing in political events, the grassroots level is missing in livelihood events, and the professional level is missing in in-depth industry events) calculated based on the judgment of the missing key levels is subtracted. The final result is a comprehensive hierarchical value that can comprehensively reflect the news value and reporting completeness of the event. The specific calculation formula is as follows: ;
[0084] Among them, Score is the comprehensive layered value, reflecting the news value and reporting completeness of the event; i is the layer index, representing the four levels of the source of the viewpoint (authoritative layer, professional layer, social layer, grassroots layer); layer_value(i) is the layer value contribution, representing the comprehensive contribution of the i-th layer to the news value of the event; Relevance(i) is the semantic fit, representing the degree of matching between the viewpoint of the i-th layer and the target column; Penalty is the penalty factor, representing the quantitative deduction for insufficient reporting completeness due to the absence of key layers.
[0085] Step S206: Construct cross-layer relationship signals based on event clusters and structured data; generate multiple candidate perspectives based on cross-layer relationship signals, hierarchical value contributions, and structured data.
[0086] Among them, four types of cross-level relationship signals are constructed based on event clusters and structured data (including extended triples, source levels, position labels, etc.). First, the position distribution of viewpoints at each level (authority level, professional level, social level, grassroots level) is used to obtain conflict signals (CI) that reflect the degree of position opposition between levels.
[0087] Then, based on dissemination indicators such as the number of reads and comments of relevant content at the social level, the growth rate of social level discussions is calculated. Combined with the release frequency and update time of authoritative level content, the update intensity of authoritative level content is obtained, thus obtaining the gap signal (GI) that represents the asymmetry between public opinion and official response.
[0088] Subsequently, representative viewpoint vectors from the professional and grassroots levels are extracted, their semantic similarity is calculated, and the minimum value of the static credibility of the two levels is combined to generate a mutual verification signal (VS) that reflects the relationship of mutual verification of facts.
[0089] Simultaneously, based on the propagation indicators in multi-source data, the original popularity value within the current time window is statistically analyzed. The smoothed popularity value of the current time is obtained by combining the smoothed popularity value of the previous moment with the exponential smoothing method. After comparing it with the trend sensitivity threshold (dynamically set based on the standard deviation of the topic's historical popularity over 7 days), a trend signal (TI) that characterizes the change in propagation potential energy is obtained, and then integrated to form a cross-layer relationship signal vector.
[0090] Next, based on the signal vector, the layer value contribution of each level (layer_value), and the scene weight factor mapped by the scene vector (Scene=(P,L,D)), the trigger intensity of the four types of signals is calculated respectively. When the trigger intensity exceeds the corresponding preset threshold, the topic angle signals such as opposing viewpoint type, hot spot gap type, in-depth mutual verification type, and trend tracking type are triggered.
[0091] Finally, the matching angle template is invoked to extract key supporting content (including representative viewpoints at each level, source information, position distribution, etc.) from the structured data. The value weight of different levels in the topic selection is clarified by combining the value contribution of each level, and multiple structured candidate angles are generated, including trigger signal summaries, level contribution, scene feature vectors, angle entry descriptions, and representative viewpoints and source paths.
[0092] Step S207: Based on the comprehensive hierarchical value, score and sort the candidate angles, and select the top N candidate angles with the highest scores as the final candidate angles; N is a positive integer.
[0093] Specifically, an innovative comprehensive evaluation is first conducted for each candidate angle. Then, the comprehensive hierarchical value (reflecting the news value and reporting completeness of the event) and the innovative comprehensive score are normalized respectively. Next, the comprehensive score of each candidate angle is calculated. Finally, all candidate angles are sorted in descending order based on the comprehensive score results, and the top N candidate angles with the highest scores (N is a positive integer and can be flexibly set according to actual news gathering and editing needs) are selected as the final candidate angles, providing core input for subsequent localization conversion and feasibility verification.
[0094] Steps S201 to S207 above involve aggregating the same event into an event cluster through multi-source data fusion, accurately identifying opinion statements and generating structured data with complete source information. After classification by source level, credibility and dynamic influence of the event are quantified by combining long-term behavioral data of the subject with multi-source indicators. Simultaneously, a scenario vector is constructed based on the event's propagation attributes to determine key level deficiencies and calculate penalty factors. Then, the level value contribution is calculated through credibility and dynamic influence, and a comprehensive stratified value is obtained by combining the penalty factor. Subsequently, cross-level relationship signals are constructed based on event clusters and structured data, and multiple candidate angles are generated by combining level value contributions. Finally, high-value candidate angles are selected by ranking and sorting through comprehensive stratified value scores. This process not only solves the problems of traditional topic selection relying on shallow data and incomplete opinion mining, but also ensures the objectivity and completeness of topic selection through hierarchical division and dynamic quantitative evaluation. It achieves differentiation and innovation in topic angles through cross-level signal mining, effectively improving the quality, depth, and screening efficiency of news topics, and providing scientific and systematic decision support for news gathering and editing.
[0095] In some embodiments, the hierarchy includes an authority layer, a professional layer, a social layer, and a grassroots layer; the construction of cross-layer relationship signals based on the event clusters and the structured data includes:
[0096] Based on the viewpoints and positions in the structured data corresponding to each level, conflict signals are obtained;
[0097] Based on the event cluster, the growth rate of social layer discussions and the intensity of content updates at the authority layer are calculated to obtain gap signals;
[0098] Based on the representative viewpoints in the structured data corresponding to the professional layer and the representative viewpoints in the structured data corresponding to the basic layer, a mutual verification signal is obtained;
[0099] Based on the indicator information associated with the multi-source data, the original heat value is calculated; based on the original heat value and the smoothed heat value of the previous moment, the smoothed heat value of the current moment is obtained; based on the smoothed heat value of the current moment, the smoothed heat value of the previous moment, and the trend sensitivity threshold, a trend signal is obtained.
[0100] Based on the conflict signal, the gap signal, the mutual verification signal, and the trend signal, a cross-layer relationship signal is constructed.
[0101] The aforementioned cross-level relationship signals are derived from comparative analysis of structured viewpoint data from different source levels, used to characterize the relationship features of different source levels in terms of stance, intensity, or changing trends. Cross-level relationship signals include conflict signals, gap signals, mutual corroboration signals, and trend signals. Among these, the core of the conflict signal CI is to capture the opposing differences in viewpoint stances between different levels, uncovering high-value angles such as "opinion clashes" for topic selection. Specifically, based on the structured data corresponding to each level, viewpoint stance information (support, opposition, neutral, pending, etc.) is extracted, and the proportion of each type of stance in that level is statistically analyzed, forming the stance distribution characteristics specific to each level. Subsequently, by calculating the degree of stance difference between any two levels, i.e., comparing the difference in the proportion of core stances such as support and opposition in different levels, the degree of viewpoint divergence between levels is quantified. Finally, the largest degree of stance difference among all level combinations is used as the core quantitative value of the conflict signal. The larger this value, the more significant the cross-level viewpoint conflict, and the greater the news reporting value from an opposing perspective.
[0102] The Gap Signal GI focuses on the asymmetry between public opinion intensity and official response, accurately identifying topic opportunities in "information vacuums." Based on the full document data contained in the event cluster, it calculates the growth rate of social discussion (social intensity) Trend. social This involves quantifying changes in public attention to an event by statistically analyzing the growth rates of metrics such as discussion volume, comment volume, and repost volume on social media platforms and forums within a given timeframe; and calculating the intensity of authoritative content updates (authoritative response). authority This involves tracking the frequency of content releases, information updates, and iterations of core viewpoints from authoritative channels such as government announcements and official media to measure the strength of official responses to events. The difference between the growth rate of public discussion and the intensity of updates from authoritative sources is calculated. social -Response authority A gap signal is obtained. When the signal is significantly positive, it means that public opinion is heating up but the official response is lagging behind, forming a typical information gap, which is the core entry point for news topics. Figure 3 This is a schematic diagram of the gap signal calculation principle according to an embodiment of this application. The horizontal axis represents time (t), and the vertical axis represents intensity (i.e., heat / response intensity). By comparing the changing trends of the social layer heat curve and the authoritative layer response curve, the "asymmetry between public opinion heat and official response" is presented intuitively. When the social layer discussion heat continues to rise while the authoritative layer content update intensity lags behind, the difference area formed between the two curves is the "information vacuum zone". The size of this area directly corresponds to the intensity of the gap signal and is the core basis for identifying high-value hot topics.
[0103] The Mutual Evidence Signal VS aims to uncover in-depth reporting angles where "professional interpretations and grassroots feedback corroborate each other," enhancing the credibility and persuasiveness of selected topics. First, representative viewpoints, after clustering and deduplication, are extracted from the structured data corresponding to the professional level—the core viewpoints that are semantically most representative and supported by the most substantial evidence at that level. Simultaneously, representative viewpoints that best reflect the actual situation, such as statements from parties involved and on-site eyewitness feedback, are selected from the structured data corresponding to the grassroots level. Natural language processing technology is used to calculate the semantic similarity between the two types of representative viewpoints, judging their degree of alignment in factual description and core conclusions. Then, combined with preset static credibility weights for the professional and grassroots levels, the semantic similarity results are weighted and calibrated (taking the minimum value of the two static credibility levels as the weight to avoid misjudgments caused by insufficient credibility at a single level). Finally, a mutual evidence signal is generated. The higher the signal strength, the higher the alignment between the professional viewpoint and the actual situation at the grassroots level, indicating value for in-depth analysis.
[0104] Trend signals (TI) are used to capture changes in the propagation potential of an event over time, identifying topics that are "rising hot topics." First, based on propagation metrics information correlated from multiple data sources (such as readership, reposts, and comment growth rates on various platforms), the original popularity value X of the event is calculated. t This directly reflects the surface heat intensity during the current period. To avoid interference from short-term fluctuations in trend judgment, an exponential smoothing method is used to process the original heat value: combining it with the smoothed heat value S from the previous moment. t-1 By adjusting the weights of the original heat value and the current heat value by setting a smoothing coefficient α, the smoothed heat value S at the current moment is obtained. t This allows for a smooth depiction of the transmission trend. The specific calculation formula is S. t =αX t +(1-α)S t-1 Then, the difference S between the smoothed heat value at the current time and the previous time is calculated. t -S t-1 The difference is compared with the trend sensitivity threshold (dynamically set based on k times the standard deviation of the topic's historical 7-day popularity). When the difference exceeds the threshold, it is determined that the spread of the event has a significant upward trend, generating a trend signal to provide data support for timely tracking of hot topics and seizing the opportunity to report.
[0105] Cross-layer relationship signals are a comprehensive value indicator system integrating the above four types of single signals, providing multi-dimensional triggering basis for generating topic selection angles. By structurally integrating conflict signals, gap signals, mutual corroboration signals, and trend signals, a unified cross-layer relationship signal vector is formed. This vector retains the core quantitative characteristics of each type of single signal while achieving a comprehensive portrayal of the news value of an event through multi-dimensional complementarity. Conflict signals reflect the tension of viewpoints, gap signals reflect the gap between information supply and demand, mutual corroboration signals ensure the credibility of content, and trend signals capture the timeliness of dissemination. Subsequently, based on this signal vector, combined with hierarchical value contribution and scenario adaptability, different types of topic selection angles can be triggered, ensuring that the generated topics not only align with the core characteristics of the event but also possess differentiation and depth.
[0106] Through the above steps, a systematic and quantitative capture and integration of multi-dimensional core value signals behind news events is achieved. This involves uncovering the focus of viewpoint conflicts through differences in hierarchical positions, identifying information gaps by comparing the dissemination and response of the social and authoritative levels, strengthening the credibility of the topic selection by using the semantic fit of the viewpoints represented by the professional and grassroots levels, accurately grasping the upward trend of dissemination through the smoothing of heat data and comparison of trend thresholds, and integrating the four types of signals into a comprehensive cross-level relationship signal. This breaks through the limitations of traditional topic selection that only relies on shallow heat data, and provides quantifiable and multi-dimensional core evidence for generating differentiated, in-depth news topic angles that are both credible and timely. This effectively improves the accuracy and value density of topic selection.
[0107] In some embodiments, obtaining conflict signals based on viewpoints and positions in the structured data corresponding to each level includes:
[0108] Based on the viewpoints and positions in the structured data corresponding to each level, a position distribution of viewpoints at each level is formed;
[0109] Based on the position distribution, calculate the position difference degree between any two levels;
[0110] Based on the aforementioned degree of difference in positions, a conflict signal is obtained.
[0111] Specifically, based on the viewpoints and positions in the structured data corresponding to each level, the proportion of supportive and opposing viewpoints within each level is statistically analyzed to form the position distribution of viewpoints at each level, which is then standardized. Based on the position distribution, the net position ratio r of each level p is calculated. p =P support -P oppose (i.e., the proportion P of the viewpoint supported at this level) support The proportion of opposing views P oppose The difference, r p(∈[-1,1], the closer the value is to 1, the more supporters there are; the closer it is to -1, the more opponents there are.) Then, by calculating the absolute difference between the net position ratios of any two different levels, the degree of position difference between any two levels is obtained. Based on the degree of position difference, the maximum value is selected as the consensus index / conflict signal CI (the larger the CI, the stronger the position divergence between levels; the smaller the CI, the more consistent the positions). When CI exceeds the preset conflict threshold θ, C (If the value is 0.5, it can be dynamically adjusted based on actual business needs or historical data statistical characteristics.) When this value is reached, a significant cross-layer conflict is determined, ultimately resulting in a conflict signal. The formula for calculating the consensus index / conflict signal CI is as follows:
[0112] ;
[0113] Where p and q represent any two different source levels (such as the authority level and the social level); r p and r q These are the net position ratios for levels p and q (i.e., the percentage of supporting opinions minus the percentage of opposing opinions for that level); |r p -r q | represents the absolute difference between the net position ratios of the two levels, indicating the degree of positional difference between the two levels; max (p,q) This is to take the maximum value of the positional difference among all combinations of levels. Figure 4 This is a schematic diagram of the conflict signal calculation principle according to an embodiment of this application. The diagram uses the net position ratio as an indicator to show the position tendencies of four groups: the authority level, the grassroots level, the professional level, and the social level. The authority level (-0.8) has a negative position, the grassroots level (+0.7) and the social level (+0.5) have a positive position, and the professional level (-0.2) is close to neutral. The diagram uses arrows to mark the significant opposition between the positions of the authority level and the grassroots level. This difference reflects the greatest conflict between the different groups.
[0114] The above steps first extract viewpoints and positions from structured data at each level and form a standardized position distribution, then quantify the degree of position difference between any two levels, and finally generate conflict signals that accurately reflect the degree of cross-level viewpoint opposition. This achieves a systematic and quantifiable capture of the positional differences among multiple groups behind news events, effectively breaking through the limitations of traditional topic selection, which relies on manual and subjective identification of viewpoint conflicts. It not only helps to avoid reporting bias caused by a single perspective, but also accurately identifies topic entry points with "viewpoint clash" value, providing objective data support for the subsequent generation of compelling and multi-dimensional news topics, while improving the efficiency and accuracy of topic conflict analysis.
[0115] In some embodiments, obtaining the mutual verification signal based on representative viewpoints in the structured data corresponding to the professional layer and representative viewpoints in the structured data corresponding to the grassroots layer includes:
[0116] Semantic similarity is calculated based on representative viewpoints in the structured data corresponding to the professional layer and representative viewpoints in the structured data corresponding to the basic layer.
[0117] Based on the semantic similarity and the preset credibility weight, a mutual verification signal is obtained.
[0118] The aforementioned mutual corroboration signals are generated based on the degree of semantic content consistency of structured viewpoint data from different source levels, and are used to characterize the mutual corroboration relationship between different source levels in terms of facts or viewpoints. Specifically, the most representative viewpoints after semantic clustering are selected from the structured data corresponding to the professional level and the grassroots level, respectively, and the two types of representative viewpoints are transformed into semantic vectors (i.e., claims). pro With claim grass The semantic similarity calculation model is used to calculate the degree of matching between the two at the content level, resulting in sim(claim). pro ,claim grass Then obtain the static credibility weight C of the professional layer. base (pro) and the static credibility weight C of the basic layer base (grass), taking the minimum of the two as the weighted constraint coefficient (to ensure that when the basic credibility of one layer is low, the mutual verification signal will not be amplified due to the high credibility of a single layer, thereby improving the robustness of the mutual verification judgment); finally, the semantic similarity result is multiplied by the weighted constraint coefficient to obtain the mutual verification signal VS. When VS exceeds the preset threshold θ V When the value is 0.6, the two viewpoints are considered to form a "positive mutual corroboration," possessing value for in-depth reporting. The formula for calculating VS is as follows:
[0119] ;
[0120] Among them, sim(claim) pro ,claim grass Professional representatives' viewpoints (claims) pro ) and grassroots representatives' views (claims) grass The semantic similarity between the two reflects the degree of consistency in their content; C base (pro) represents the static credibility weight of the professional layer, reflecting the basic reliability of the information source at the professional layer; C base (grass) represents the static credibility weight of the basic layer.
[0121] The above steps calculate the semantic similarity between the viewpoints represented by professionals and those at the grassroots level, and combine this with a weighted constraint based on two preset static credibility weights to generate quantifiable mutual verification signals. This achieves a systematic and quantifiable identification of whether professional interpretations and genuine feedback from the grassroots form positive corroboration. It not only locks in factual consistency at the content level through semantic matching, but also filters out misjudgments caused by unreliable single-level sources through credibility weights. This effectively improves the scientific rigor and robustness of mutual verification judgments, breaking through the limitations of traditional manual judgments that rely on subjective experience. At the same time, it accurately selects in-depth reporting materials that combine professionalism and realistic evidence, providing objective data support for the subsequent generation of highly credible and persuasive news topics, and further ensuring the authenticity and credibility of the reported content.
[0122] In some embodiments, the cross-layer relationship signal includes conflict signals, gap signals, mutual verification signals, and trend signals; the generation of multiple candidate perspectives based on the cross-layer relationship signal, the hierarchical value contribution, and the structured data includes:
[0123] Based on the cross-layer relationship signal and the hierarchical value contribution, the signal trigger intensity corresponding to the conflict signal, the gap signal, the mutual verification signal and the trend signal are calculated respectively; when the signal trigger intensity is greater than a preset threshold, multiple topic selection angle signals are triggered;
[0124] Based on the multiple topic selection angle signals, the structured data, and the hierarchical value contribution, multiple candidate angles are generated; the multiple candidate angles correspond to different types of cross-level relationship signals.
[0125] The cross-layer relationship signals include conflict signals (CI), gap signals (GI), mutual verification signals (VS), and trend signals (TI). Specifically, a cross-layer relationship signal vector V=(CI,GI,VS,TI) is first constructed. For each signal in the vector, its corresponding signal index value S is retrieved. i (i.e., the specific quantitative values of CI / GI / VS / TI), and the hierarchical value contribution L obtained based on previous dynamic evaluation results. i And the scene weight factor Scene obtained by mapping the scene vector Scene=(P,L,D). i (Reflecting the applicability and importance of this signal in current political, livelihood, or industry-specific scenarios), triggered by the intensity function Trigger. i =S i ·L i Scene i Calculate the signal trigger strength for each of the four signal types; when the trigger strength of a certain signal is Trigger... i Greater than its corresponding preset threshold θ i(e.g., the threshold θ for conflicting signals) C Notch signal threshold θ G Mutual verification signal θ V Trend signal θ T When a signal is triggered (e.g., conflict signals trigger opposing viewpoints, gap signals trigger hot topic gaps, mutual verification signals trigger in-depth mutual verification, and trend signals trigger trend-following angles), it is considered a valid topic selection angle signal.
[0126] Finally, the system calls the dedicated angle templates matched with each signal, extracts key supporting content from the structured data, and generates structured candidate angles containing five core elements according to a unified format: trigger signal summary (clearly identifying the CI / GI / VS / TI signals that contributed to the angle generation and their corresponding strengths), hierarchical contribution (layer_value weights for each level, such as the authority layer and professional layer, calculated based on a dynamic hierarchical evaluation model), scene feature vector (reflecting the attribute distribution of the event in three scenarios: political (P), livelihood (L), and in-depth industry (D), explaining the logic of weight selection), angle entry description (summarizing the core entry point in no more than 40 words), and representative viewpoints and tracing paths (extracting key viewpoints at each level and tracing information such as source level, time, and evidence location). The system closely integrates the model calculation results with viewpoint evidence usable by editors, ensuring that the generated candidate angles are both clear and interpretable, and provide editors with directly usable interview guidance, guaranteeing the practicality and credibility of the topic selection. Among these, multiple candidate angles correspond to different types of cross-layer relationship signals.
[0127] The above steps combine four types of cross-level relationship signals—conflict, gap, mutual corroboration, and trend—with hierarchical value contributions. First, the trigger strength of each signal is quantified and effective topic selection angle signals are screened. Then, structured data and hierarchical value contributions are integrated to generate multiple candidate angles. This realizes a data-driven topic selection angle mining model that replaces traditional experience-based judgment. It ensures the differentiation and innovation of topic selection angles through multi-dimensional signal triggers (covering opposing viewpoints, hot gaps, in-depth mutual corroboration, trend tracking, etc.) and makes candidate angles both interpretable and feasible with the support of hierarchical value contributions and structured data. This effectively breaks through the limitations of traditional topic selection angles being homogeneous and relying on manual screening for low efficiency. At the same time, it provides editors with clear value basis and evidence support, greatly improving the accuracy, diversity, and efficiency of news topic selection and editing.
[0128] In some embodiments, the step of scoring the candidate angles based on the comprehensive hierarchical value and selecting the top N candidate angles as the final candidate angles includes:
[0129] Based on the candidate perspectives, an innovation score is generated;
[0130] Based on the comprehensive hierarchical value and the innovation score, a comprehensive score is given to multiple candidate angles, and the top N candidate angles with the highest comprehensive scores are selected as the final candidate angles.
[0131] In the process of scoring candidate angles based on comprehensive hierarchical value and selecting the top N final candidate angles, the novelty of the content is first calculated for each candidate angle by combining historical topic data, the uniqueness of the perspective is obtained through information entropy analysis, and the formal innovation is evaluated by comparing with a preset narrative template library. This process is then integrated to obtain an innovation score. Next, the comprehensive hierarchical value and the innovation score are normalized separately, weighted and summed according to preset weights to obtain a comprehensive score. Finally, the top N candidate angles with the highest comprehensive scores are selected as the final candidate angles.
[0132] The above steps, by combining innovation scoring with comprehensive hierarchical value, not only ensure that the final candidate angles align with the core news value of the event and the integrity of the report, but also effectively avoid the problem of topic homogenization, significantly improving the differentiation and innovation of candidate angles. At the same time, through quantitative scoring and ranking mechanisms, the topic selection is made more objective and scientific, greatly reducing the subjective bias and workload of manual selection, and providing news gathering and editing with higher quality and more diverse topic directions.
[0133] In some embodiments, generating an innovation score based on the candidate angles includes:
[0134] Based on the candidate angles and historical topic selection angles, a content novelty score is obtained;
[0135] Based on the candidate angles, calculate the information entropy; based on the information entropy, obtain a unique perspective score.
[0136] Based on the candidate perspectives and the preset narrative template library, a score for formal innovation is generated;
[0137] An overall innovation score is obtained based on the content novelty score, the perspective uniqueness score, and the form innovation score.
[0138] First, the system obtains the content novelty score. It accesses a historical topic database, extracting past topic angles within a certain period (e.g., the last 3 months or 6 months) to construct a historical topic feature library. The system then compares the core semantics, keywords, and event dimensions of the current candidate angle with those of historical topic angles, quantifying the difference between the current and historical angles by calculating semantic distance and feature overlap. A higher degree of difference indicates lower content repetition and stronger novelty, resulting in a higher content novelty score. This avoids topic homogenization and ensures that new angles possess unique reporting value.
[0139] Secondly, there's the generation of the unique perspective score. The core of unique perspective lies in the diversity of stances and entry points. The system achieves quantitative evaluation by calculating information entropy. Specifically, based on structured data and other information, it statistically analyzes the distribution ratio of viewpoints associated with the candidate perspective across different stances such as support, opposition, and neutrality. Then, according to the information entropy calculation formula, it combines the balance and richness of the stance distribution to derive the information entropy value. Higher information entropy indicates a more diverse range of stances and a more comprehensive perspective, breaking through the limitations of a single stance and providing a more three-dimensional presentation for the report, thus resulting in a higher unique perspective score.
[0140] Furthermore, the system determines the score for formal innovation. It has a built-in library of pre-set narrative templates, including standard templates for various news presentation formats such as in-depth analysis, interviews, investigative reporting, and data visualization. The core content and expression logic of candidate angles are matched with various templates in the library to assess the degree of difference between their presentation and traditional templates. If a candidate angle can adapt to a new narrative structure, integrate diverse presentation techniques, or break through the framework limitations of existing templates, it is judged to have a high degree of formal innovation, resulting in a higher score and helping to enhance the appeal of the report.
[0141] Finally, the system integrates the comprehensive innovation score. Preset weights are assigned to the content novelty score, unique perspective score, and format innovation score (e.g., content novelty 0.4, unique perspective 0.35, format innovation 0.25, which can be adaptively adjusted according to the column's positioning or business needs). Through a weighted summation, the scores from the three dimensions are integrated into a unified comprehensive innovation score. This comprehensive score fully reflects the overall innovation level of the candidate angles in terms of content, perspective, and form, providing a crucial basis for the subsequent comprehensive ranking of candidate angles.
[0142] Through the above steps, a multi-dimensional and quantitative comprehensive evaluation of the innovation of candidate angles was achieved. This not only ensured the novelty of the content by comparing with historical topic angles and avoiding topic homogenization, but also captured the uniqueness of the perspective by using information entropy calculation and explored the reporting value of combinations of diverse positions. Furthermore, it enhanced the innovation of the presentation by matching with a pre-set narrative template library and strengthened the appeal of the report. The final integrated innovation score provides an objective and comprehensive basis for the selection and ranking of candidate angles, effectively supporting the generation of differentiated and high-quality news topics.
[0143] In some embodiments, the process of obtaining credibility for each level based on long-term behavioral data of opinion subjects in the multi-source data includes:
[0144] For each level, a historical score is calculated based on the long-term behavioral data of the opinion subjects in the multi-source data; the credibility is obtained based on the preset static credibility and the historical score.
[0145] And / or, obtaining the dynamic impact of the event based on the indicator information associated with the multi-source data includes:
[0146] Based on the indicator information associated with the multi-source data, the original event impact is calculated;
[0147] Based on the propagation content and propagation behavior data of the event cluster, the emotional polarization index, text homogenization, and short-term abnormal outbreak characteristics are extracted to obtain the emotional noise suppression factor.
[0148] Based on the emotional noise suppression factor and the original event influence, the dynamic influence of the event is obtained.
[0149] The credibility calculation logic is as follows: for each level (authoritative level, professional level, social level, grassroots level), credibility acquisition requires combining preset standards with dynamic behavioral data to achieve objective and accurate source assessment. First, the system extracts long-term behavioral data of the opinion subject from multi-source data, taking the account entity of the opinion subject on a specific publishing platform as the unit. This data covers factual consistency (the consistency between historical content and subsequent authoritative information), debunking rate (the proportion of historical content that is corrected or refuted), content quality stability (consistency of opinion, noise ratio, language standardization, etc.), and optional dimensions such as cross-platform performance and user interaction quality. Through rule models, statistical models, or machine learning regression models, a historical score (C) reflecting the subject's long-term performance is comprehensively calculated. history Based on this, and combined with the pre-set static credibility baselines (C) at each level... static For example, the authority level ≈ 0.95, the professional level ≈ 0.85, the grassroots level ≈ 0.70, and the social level ≈ 0.40, using the weighted formula λC static +(1-λ)C history The final credibility C is obtained by fusion. final The weight λ∈[0,1] can be dynamically adjusted according to the event type, source distribution, and data quality. This not only preserves the inherent credibility attributes of different levels but also effectively avoids the problem of low-quality information sources such as "pseudo-experts" and "marketing accounts" gaining excessive influence due to the inherent weight of the level through dynamic calibration of historical behavioral data.
[0150] Acquiring the dynamic influence of an event requires balancing the initial dissemination momentum with filtering for true value, avoiding being misled by unnatural dissemination behaviors. The first step involves calculating the initial event influence (I) based on dissemination metrics from multiple data sources, such as readership, reposts, comments, and discussion growth rates across various platforms, through normalization and weighted summation. This calculates the I representing the event's surface-level popularity.raw The second step, to eliminate interference from non-genuine dissemination behaviors such as emotional manipulation and paid online trolls, involves the system extracting three core features from the content and behavior data of the event clusters to generate emotional noise suppression factors: first, emotional polarization index, which measures the concentration of extreme emotions in the disseminated content; second, text homogenization, which assesses the presence of mass-produced paid online troll characteristics through semantic similarity and repetition rate; and third, short-term abnormal outbreaks, identifying instantaneous dissemination peaks without a reasonable event driver. The system comprehensively evaluates and normalizes these three features, generating a noise coefficient between 0 and 1. The closer the coefficient is to 1, the more obvious the non-genuine dissemination behavior. Finally, through formula I... raw (1-Noise) modifies the original event impact to obtain the event dynamic impact I, which truly reflects the level of social attention and dissemination value of the event. final This provides a reliable basis for calculating the value contribution of subsequent levels.
[0151] Through the above steps, precise quantification and dynamic optimization of hierarchical credibility and the dynamic influence of events are achieved. This is achieved by weighted fusion of historical scores generated from static credibility and long-term behavioral data of opinion subjects, effectively avoiding interference from low-quality sources such as "pseudo-experts" and "marketing accounts," thus ensuring the objectivity and accuracy of source evaluation. Furthermore, by extracting emotional polarization, text homogenization, and short-term abnormal outbreak characteristics to generate emotional noise suppression factors, the original event influence is denoised and corrected, eliminating misleading information from inaccurate dissemination behaviors such as paid online commentary and malicious hype. This truly reflects the degree of social attention and dissemination value of the event, providing reliable data support for subsequent hierarchical value contribution calculations and candidate angle generation, significantly improving the scientific rigor and accuracy of news topic selection.
[0152] In some embodiments, identifying candidate opinion sentences from the document set of the event cluster and generating structured data includes:
[0153] Identify opinion sentences from the document set of the event cluster to obtain candidate opinion sentences; input the candidate opinion sentences into a preset large language model to obtain initial structured data;
[0154] Based on the viewpoints contained in the initial structured data, representative viewpoints and supporting evidence are obtained;
[0155] The structured data is generated based on the representative viewpoints, supporting evidence, viewpoint source information associated with the multi-source data, and the initial structured data.
[0156] First, the system accurately identifies candidate opinion sentences. It performs deep sentence-level analysis on the document set within the event cluster, filtering opinion sentences based on three core criteria: subjectivity, stance, and structureability. The subjectivity criterion focuses on sentences containing attitudinal, evaluative, or emotional words; the stance criterion targets expressions that clearly demonstrate support, opposition, or concern; and the structureability criterion requires sentences to have a core structure of "subject – topic / object – stance." Simultaneously, it clarifies that objective factual sentences and purely data-driven sentences are not considered opinion sentences, while quotations or paraphrases containing stance components are still treated as opinion sentences. Through this filtering logic, candidate opinion sentences with reporting value are accurately identified. The subject is the actor expressing the opinion, which can be an individual (citizen, expert), organization (municipal government, property management company), institution, or media; the topic / object is the issue, event, measure, or policy addressed by the opinion, such as "nighttime odor," "emissions," or "road construction"; and the stance is the subject's attitude towards the object, including support, opposition, neutrality, skepticism, and concern.
[0157] Next, initial structured data is generated. The selected candidate opinion sentences are input into a pre-defined large language model. Under the constraints of the news element layer, the large language model performs semantic understanding and standardized transcription of the candidate opinion sentences, outputting initial structured data containing the opinion subject (individuals, organizations, institutions, media, etc.), opinion content (specific evaluations or attitudes around the topic / object), and opinion stance (support, opposition, neutrality, pending, etc.). Simultaneously, the confidence score corresponding to the stance label is attached. This confidence score is used to characterize the clarity and stability of the stance determination, providing a reference for subsequent data processing.
[0158] Next comes the extraction of representative viewpoints and supporting evidence. For the viewpoint content in the initial structured data, the system uses BERT sentence vector generation technology to transform each viewpoint into a semantic vector. Then, based on semantic similarity, cluster analysis is performed to merge viewpoints with high semantic similarity and expressing the same stance into the same viewpoint cluster, thus achieving viewpoint deduplication. Within each viewpoint cluster, the sentence with the most representative semantics and the highest weight is selected as the representative viewpoint, while the remaining viewpoints within the cluster are retained as supporting evidence. Simultaneously, the original text, source platform, and timestamp of the supporting evidence are recorded, forming a combination of "representative viewpoint + multi-source supporting evidence," which both simplifies the data and preserves complete evidence.
[0159] Finally, the structured data is integrated and generated. The representative viewpoints and supporting evidence obtained above are fully integrated with the viewpoint source information (including platform type, original document ID, publication time, capture timestamp, and other metadata) automatically extracted and bound throughout the multi-source data access phase, as well as the initial structured data. The resulting structured data not only contains core information such as the viewpoint subject, content, and stance, but also integrates supporting evidence and complete source tracing information, forming standardized data with an immutable source tracing index. This provides a solid foundation for subsequent steps such as hierarchical division, credibility calculation, and cross-level relationship signal construction.
[0160] Through the above steps, precise screening, standardized structuring, and end-to-end source integration of textual viewpoints within event clusters are achieved. This involves selecting candidate viewpoints with reporting value based on criteria of "subjectivity, stance, and structurability," transforming them into standardized initial structured data using a large language model, then deduplicating viewpoints through viewpoint clustering to extract representative viewpoints and supporting evidence, and finally integrating information from multiple sources to generate complete structured data. This approach ensures the standardization, completeness, and diversity of viewpoint data, providing high-quality data support for subsequent hierarchical division and cross-level signal construction. It also achieves end-to-end traceability of topic selection viewpoints, conforming to the core "verifiable" standard of the news industry, while avoiding viewpoint redundancy and homogenization, thus improving the efficiency and credibility of news topic generation.
[0161] In some embodiments, after selecting the top N candidate angles with the highest scores as the final candidate angles, the method further includes:
[0162] The final candidate angles are subjected to entity recognition to extract place name and institution entities;
[0163] The place name institutions are mapped to a localized geographic knowledge graph, and the graph hop distance and semantic relevance between the place name institutions and the local core nodes in the knowledge graph are calculated. The knowledge graph includes local administrative division nodes, pillar industry related nodes, key livelihood project nodes and place name alias mapping relationships, forming a structured knowledge network covering local core related elements.
[0164] Based on the hop distance and semantic relevance of the graph, the association level between the final candidate angle and the local core node is determined;
[0165] Based on the aforementioned association level, localized candidate angles are generated;
[0166] The localized candidate angles are input into a preset feasibility scoring model to evaluate spatial accessibility, resource matching degree, time feasibility and news value, and an execution suggestion card is generated.
[0167] Specifically, after selecting the top N candidate angles with the highest scores as the final candidate angles, further localization adaptation and feasibility verification will be carried out to ensure that the selected topics can be implemented. First, entity recognition (NER) is performed on the final candidate angles to accurately extract key entities such as place names and institutions. Then, these place name and institution entities are mapped to the built-in localized geographic knowledge graph. This graph covers local administrative division nodes, pillar industry related nodes, key livelihood project nodes, and place name alias mapping relationships, forming a structured knowledge network covering local core related elements. At the same time, the graph hop distance (reflecting the closeness of the relationship between the entity and the core node) and semantic relevance (quantifying the logical fit between the entity and the core node) are calculated between the entity and the local core node in the knowledge graph.
[0168] Subsequently, by combining the results of the graph hop distance and semantic relevance, the association level between the final candidate angle and the local core node is determined (e.g., high, medium, low association). Based on this association level, the candidate angle is localized to generate localized candidate angles. For example, national hot topics are combined with local industry characteristics and people's livelihood needs, and the topic selection and expression are adjusted to better suit the local audience's concerns. Then, based on the column style template library, the topic selection is automatically adjusted (e.g., "In-depth analysis" is converted to the "street interview" style of the people's livelihood column), and timeliness filtering is performed according to the production cycle constraints.
[0169] Finally, the localized candidate angles are input into the preset feasibility scoring model, which comprehensively evaluates them from four dimensions: spatial accessibility (assessing the transportation convenience and feasibility of reaching the interview location), resource matching degree (considering the suitability of resources such as reporters, equipment, and time slots), time feasibility (determining whether the collection and production can be completed within the program's broadcast window), and news value score (referencing the previous comprehensive score). The final result is an execution suggestion card containing key information such as recommended interviewees and priorities, suitable shooting locations, and estimated working hours, providing direct and actionable operational guidance for the news gathering and editing work.
[0170] Through the above steps, a closed-loop transformation from high-value candidate angles to localized and implementable news gathering and editing solutions is achieved. Firstly, by using entity recognition and mapping with localized geographic knowledge graphs, the implicit connections between national topics and local core elements are accurately uncovered, generating differentiated topics that resonate with local audiences and solving the problem of homogenization. Secondly, an executability scoring model evaluates from multiple dimensions—space, resources, time, and value—filtering out unfeasible angles. Finally, an execution suggestion card containing practical information such as specific interviewees and shooting locations is generated. This ensures the localization and news value of the topics while significantly reducing news gathering and editing decision-making costs and improving execution efficiency, achieving a leap from "theoretically feasible" to "practically usable" intelligent topic selection.
[0171] In some embodiments, feedback learning and dynamic optimization steps are also included.
[0172] The feedback learning and dynamic optimization mechanism includes two parts: negative sampling and adversarial optimization, and adaptive evolution of parameters and thresholds. Each module (such as the collector and the computing engine) can be implemented as a software program running on a computer device with a processor and memory. The system records the topic selection and adoption behavior and builds a negative sample queue. When an editor "rejects" a topic, the cross-level signal features (CI / GI / VS / TI numerical set) of the topic are marked as negative examples. When the model is updated, the recommendation weight of topics with similar signal features (such as "high conflict but low credibility") is reduced through the contrastive learning mechanism, and "pseudo-hot spots" are automatically filtered out. At the same time, the system will regularly update parameters based on the operation logs of the editors to optimize the strategy, including signal thresholds and weights such as conflict / gap, historical calibration of credibility at each level, weight of innovation score and column style template, etc., to realize the evolution from rule-based cold start to data-driven personalized recommendation.
[0173] In some embodiments, after constructing the cross-layer relationship signal, the data can be integrated to generate a unified structure including a structured viewpoint matrix, dynamic hierarchical evaluation results, scene vectors and penalty terms, and the cross-layer relationship signal, providing complete input for perspective generation, innovation scoring, and subsequent executability verification.
[0174] Specifically, it comprises three core levels: First, a structured opinion matrix M(level, stance, t), constructed based on structured data and combined with information on the source hierarchy of the opinions. Opinions are aggregated according to the source hierarchy (authoritative, professional, social, grassroots), the types of supportive, opposing, and neutral stances, and the time dimension (hourly, daily, or adaptive). Matrix elements cover the distribution of stances at each level, representative opinions and their confidence levels, hierarchical source and tracing information, and time-series behavioral indicators, serving as the foundational input for subsequent conflict, gap, mutual verification, and trend modeling. Here, level represents the source hierarchy; stance represents the opinion stance type; and t represents the time dimension, used to describe the changes in opinions within different time windows (e.g., by hour, day, or adaptive time period).
[0175] Second, the dynamic hierarchical evaluation results are presented in array form, showing the final credibility (C_final), dynamic influence (I_final), scene weights (α and β), and layer value contribution (layer_value) for each of the four levels. These are used for angle scoring and weight allocation in step three, with the following structure:
[0176] {
[0177] "layer": "Authority", / / Authority layer
[0178] "C_final": 0.92,
[0179] "I_final": 0.21,
[0180] "scene_weight": {α:0.83, β:0.17},
[0181] "layer_value": 0.79 / / = α·C_final + β·I_final
[0182] }
[0183] All levels form an array:
[0184] layer_evaluation = [
[0185] {...Authority...},
[0186] {...Expert...},
[0187] {...Social...},
[0188] {...Grassroot...} ]
[0190] It can also include scene vectors and penalty terms. The scene vector Scene=(P,L,D) specifies the political, social, and depth attributes of the event, and outputs a penalty term for missing levels. The penalty term Penalty marks the key level missing situations and the corresponding deduction value. The structure is as follows:
[0191] Penalty = {
[0192] "missing": ["Authority"],
[0193] "value": 0.35
[0194] }
[0195] Third, the cross-layer relationship signal structure includes conflict signals (CI), gap signals (GI), mutual verification signals (VS), trend signals (TI), and signal strength, which together support the accurate generation of topic selection angles. The structure is as follows:
[0196] {
[0197] "CI": ...,
[0198] "GI": ...,
[0199] "VS": ...,
[0200] "TI": ...,
[0201] "signal_strength": ...
[0202] }
[0203] The final output is a unified structure:
[0204] {
[0205] "opinion_matrix": M,
[0206] "layer_evaluation": [...],
[0207] "scene_vector": Scene,
[0208] "penalty": Penalty,
[0209] "cross_layer_signals": {CI,GI,VS,TI}
[0210] }
[0211] Figure 5 This is a schematic diagram of the structured opinion matrix generation process according to an embodiment of this application. It shows the complete processing flow of taking event cluster text as input, first extracting and structuring opinion, and then, based on a four-layer information source hierarchical model (authority layer, professional layer, social layer, and grassroots layer), performing credibility calculation, influence calculation, hierarchical value and cross-layer signal calculation (combined with scene weight) on each layer in sequence, and finally outputting a structured opinion matrix.
[0212] Through the above steps, a systematic integration and standardized output of news event-related data were achieved. This involved comprehensively aggregating viewpoints and source information from different levels, positions, and time dimensions through a structured viewpoint matrix, providing foundational data for subsequent modeling; clarifying core indicators such as credibility and influence at each level through dynamic hierarchical evaluation results, providing a quantitative basis for angle scoring; defining the impact of missing event attributes and levels through scene vectors and penalty items; and capturing high-value news characteristics using cross-level relationship signals. The resulting unified structure provides complete, coherent, and standardized input support for topic angle generation, innovation scoring, and feasibility verification, significantly improving the processing efficiency and accuracy of subsequent stages while ensuring data traceability and logical consistency.
[0213] In some embodiments, acquiring multi-source data and generating event clusters based on the multi-source data includes:
[0214] Acquire multi-source data, and use the simhash algorithm and a preset deep learning model to calculate the similarity of the multi-source data and filter out similar documents;
[0215] The hierarchical density clustering algorithm is used to aggregate the similar documents to obtain multiple event clusters.
[0216] The first step is the comprehensive collection and preprocessing of multi-source data. The system uses a multi-source collector to collect and integrate cross-platform data, accessing rich information sources such as news websites, government information platforms, social media, video platforms, and industry databases. The collector incorporates four core components: an API adapter, a crawler engine, a stream processor, and a data cleaner. The API adapter supports multiple protocols such as REST, GraphQL, and WebSocket, ensuring compatibility with data interfaces on different platforms. The crawler engine efficiently collects web page content based on a distributed framework. The stream processor uses message queues to process dynamic data streams in real time. The data cleaner uses regular expressions and natural language processing technologies to perform preprocessing tasks such as encoding standardization, HTML noise reduction, duplicate text filtering, and format correction, improving data quality.
[0217] Subsequently, data standardization and entity disambiguation were performed. Data standardization converted the time format to the ISO 8601 standard format (YYYY-MM-DDTHH:mm:ss±hh:mm), while also performing text segmentation, language detection, and encoding standardization to achieve data format uniformity. Entity disambiguation addressed the issue of inconsistent names for the same real entity across different data sources. By constructing an entity alias mapping table and a semantic matching model, different designations such as "municipal government," "city government," and "government department" were uniformly mapped to the same standard entity, preventing subsequent clustering errors caused by naming differences.
[0218] Next, document similarity calculation is performed. Event aggregation employs a dual similarity calculation method. First, the SimHash algorithm is used to generate 64-bit text fingerprints for rapid initial screening of preprocessed documents, efficiently narrowing the similarity comparison range. Then, based on a pre-defined deep learning model, semantic similarity of the text is calculated to accurately capture the degree of fit of the core meaning of the documents. The final similarity is calculated using the linear combination formula Sim... final =α·Sim hash +β·Sim semantic The dataset is obtained (in practical applications, this can be replaced by a logistic regression or neural network model). When Simfinal is greater than 0.85 (this threshold can be dynamically configured according to data characteristics and clustering accuracy requirements), the document is considered similar. finalThe final similarity score is given by α, which is the weight of the Simhash algorithm calculation result, used to balance the efficiency and accuracy of the initial screening; β is the weight of the deep learning model calculation result, reflecting the core importance of semantic similarity, and (α+β=1); Sim semantic The core function of semantic similarity calculated based on deep learning models (such as BERT) is to "accurately capture the core meaning of the text".
[0219] Finally, event clusters are generated through event aggregation. Hierarchical density clustering (HDBSCAN) can be used to aggregate the selected similar documents, automatically grouping highly similar documents into event clusters. For example, a minimum cluster size of 5 and a time window of 24 hours are selected, while noise points are automatically filtered to avoid interference from irrelevant documents. Each event cluster generates a unique event ID, containing structured information such as a document list, time span, and keyword / topic set, providing standardized input data for subsequent opinion and stance analysis.
[0220] Through the above steps, a systematic integration and precise event aggregation of multi-source heterogeneous data were achieved. This involved comprehensively covering various information sources and completing data preprocessing through a multi-component collector, resolving inconsistencies in data format and entity names through standardization and entity disambiguation, balancing document matching efficiency and semantic accuracy through dual similarity calculation, and finally efficiently aggregating similar documents to generate structured event clusters using a hierarchical density clustering algorithm. This effectively overcomes the limitations of traditional data integration, such as semantic fragmentation, inefficient matching, and chaotic clustering, providing comprehensive, standardized, and high-quality basic data support for subsequent stages such as opinion mining and stance analysis, and significantly improving the efficiency and accuracy of preliminary data processing for news topic generation.
[0221] Figure 6 This is a schematic diagram of the overall process of the news topic generation method according to the embodiments of this application. The process starts with "start" and proceeds sequentially through step one (multi-source data fusion and event aggregation, outputting aggregated event clusters), step two (viewpoint extraction, source stratification and dynamic evaluation, outputting a stratified viewpoint matrix and signal vector in combination with core processing steps), step three (topic angle generation and innovation evaluation, outputting a set of candidate topics), step four (localization conversion and executability verification, outputting a structured execution suggestion card in combination with executability processing steps), and finally step five (feedback learning and dynamic optimization) to complete parameter updates, forming a complete process from data processing to topic implementation and iterative optimization, and finally ending with "end".
[0222] Figure 7This is a schematic diagram of the framework of a news topic generation system according to an embodiment of this application. Starting from the user terminal / editorial terminal's interactive interface, the system first completes data acquisition and cleaning through a distributed crawler engine and interface adapter of a multi-source data perception layer. Then, the viewpoint structuring calculation engine realizes semantic vectorization, viewpoint extraction and layering, and generates a layered viewpoint matrix by combining a dynamic evaluation model. Subsequently, the intelligent topic generation server completes cross-layer signal recognition, candidate angle generation, and local graph mapping. At the same time, heterogeneous databases / knowledge graphs provide data support. The closed-loop feedback optimization module updates the evaluation model and weight parameters through operation log recording and parameter adaptive evolution. Finally, the topic suggestion card is fed back to the interactive interface, forming a complete system closed loop of "data acquisition-processing-topic generation-feedback optimization".
[0223] Furthermore, in conjunction with the news topic generation methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the news topic generation methods in the above embodiments.
[0224] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0225] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0226] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0227] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for generating news topics, characterized in that, include: Acquire multi-source data and generate event clusters based on the multi-source data; The event cluster is an aggregation of information about the same event; Candidate opinion sentences are identified from the document set of the event cluster, and structured data is generated; The structured data includes the subject of the viewpoint, the content of the viewpoint, the stance of the viewpoint, and information on the source of the viewpoint; Based on the identity characteristics and behavioral patterns in the source information of the viewpoints, as well as the viewpoint subject, the structured data is divided into multiple levels; For each level, credibility is obtained based on the long-term behavioral data of the opinion subject; and the dynamic influence of the event is obtained based on the indicator information associated with the multi-source data. The hierarchy includes the authority level, the professional level, the social level, and the grassroots level; The representative text content of the event cluster is input into a multi-label classification model to generate scene vectors; based on the scene vectors, the key levels are determined. By verifying whether valid viewpoint data exists in the key levels, a key level missing judgment result is obtained; if the key level missing judgment result is missing, a penalty factor is calculated based on the missing key level; wherein, the representative text content includes news articles, authoritative release summaries, and representative viewpoint content; the dimensions of the scenario vector include current affairs, people's livelihood and public concerns, and industry technology or in-depth analysis; the key level is one or more of the levels mentioned above; For each level, the level value contribution is calculated based on the credibility and the dynamic influence of the event; based on the level value contribution of each level and the penalty factor, the comprehensive hierarchical value is obtained. Based on the event clusters and the structured data, a cross-layer relationship signal is constructed; based on the cross-layer relationship signal, the hierarchical value contribution, and the structured data, multiple candidate perspectives are generated; Based on the comprehensive hierarchical value, the candidate angles are scored and ranked, and the top N candidate angles with the highest scores are selected as the final candidate angles; N is a positive integer. Specifically, based on the event clusters and the structured data, a cross-layer relationship signal is constructed, including: Based on the viewpoints and positions in the structured data corresponding to each level, conflict signals are obtained; Based on the event cluster, the growth rate of social layer discussions and the intensity of content updates at the authority layer are calculated to obtain gap signals; Based on the representative viewpoints in the structured data corresponding to the professional layer and the representative viewpoints in the structured data corresponding to the basic layer, a mutual verification signal is obtained; Based on the indicator information associated with the multi-source data, the original heat value is calculated; based on the original heat value and the smoothed heat value of the previous moment, the smoothed heat value of the current moment is obtained; based on the smoothed heat value of the current moment, the smoothed heat value of the previous moment, and the trend sensitivity threshold, a trend signal is obtained. Based on the conflicting signals, the gap signals, the mutual verification signals, and the trend signals, a cross-layer relationship signal is constructed; The process of generating multiple candidate perspectives based on the cross-layer relationship signal, the hierarchical value contribution, and the structured data includes: Based on the cross-layer relationship signal and the hierarchical value contribution, the signal trigger intensity corresponding to the conflict signal, the gap signal, the mutual verification signal and the trend signal are calculated respectively; when the signal trigger intensity is greater than a preset threshold, multiple topic selection angle signals are triggered; Based on the multiple topic selection angle signals, the structured data, and the hierarchical value contribution, multiple candidate angles are generated; the multiple candidate angles correspond to different types of cross-level relationship signals.
2. The news topic generation method according to claim 1, characterized in that, The conflict signals are obtained based on the viewpoints and positions in the structured data corresponding to each level, including: Based on the viewpoints and positions in the structured data corresponding to each level, a position distribution of viewpoints at each level is formed; Based on the position distribution, calculate the position difference degree between any two levels; Based on the aforementioned degree of difference in positions, a conflict signal is obtained.
3. The news topic generation method according to claim 1, characterized in that, The process of obtaining mutual verification signals based on representative viewpoints in the structured data corresponding to the professional level and representative viewpoints in the structured data corresponding to the grassroots level includes: Semantic similarity is calculated based on representative viewpoints in the structured data corresponding to the professional layer and representative viewpoints in the structured data corresponding to the basic layer. Based on the semantic similarity and the preset credibility weight, a mutual verification signal is obtained.
4. The news topic generation method according to claim 1, characterized in that, The process of scoring the candidate angles based on the comprehensive hierarchical value and selecting the top N candidate angles as the final candidate angles includes: Based on the candidate perspectives, an innovation score is generated; Based on the comprehensive hierarchical value and the innovation score, a comprehensive score is given to multiple candidate angles, and the top N candidate angles with the highest comprehensive scores are selected as the final candidate angles.
5. The news topic generation method according to claim 4, characterized in that, The generation of an innovation score based on the candidate perspectives includes: Based on the candidate angles and historical topic selection angles, a content novelty score is obtained; Based on the candidate angles, calculate the information entropy; based on the information entropy, obtain a unique perspective score. Based on the candidate perspectives and the preset narrative template library, a score for formal innovation is generated; An overall innovation score is obtained based on the content novelty score, the perspective uniqueness score, and the form innovation score.
6. The news topic generation method according to claim 1, characterized in that, For each level, credibility is obtained based on the long-term behavioral data of the opinion subjects in the multi-source data, including: For each level, a historical score is calculated based on the long-term behavioral data of the opinion subjects in the multi-source data; the credibility is obtained based on the preset static credibility and the historical score. And / or, obtaining the dynamic impact of the event based on the indicator information associated with the multi-source data includes: Based on the indicator information associated with the multi-source data, the original event impact is calculated; Based on the propagation content and propagation behavior data of the event cluster, the emotional polarization index, text homogenization, and short-term abnormal outbreak characteristics are extracted to obtain the emotional noise suppression factor. Based on the emotional noise suppression factor and the original event influence, the dynamic influence of the event is obtained.
7. The news topic generation method according to claim 1, characterized in that, The step of identifying candidate opinion sentences from the document set of the event cluster and generating structured data includes: Identify opinion sentences from the document set of the event cluster to obtain candidate opinion sentences; input the candidate opinion sentences into a preset large language model to obtain initial structured data; Based on the viewpoints contained in the initial structured data, representative viewpoints and supporting evidence are obtained; The structured data is generated based on the representative viewpoints, supporting evidence, viewpoint source information associated with the multi-source data, and the initial structured data.
8. The news topic generation method according to any one of claims 1 to 7, characterized in that, After selecting the top N candidate angles with the highest scores as the final candidate angles, the process also includes: The final candidate angles are subjected to entity recognition to extract place name and institution entities; The place name institutions are mapped to a localized geographic knowledge graph, and the graph hop distance and semantic relevance between the place name institutions and the local core nodes in the knowledge graph are calculated. The knowledge graph includes local administrative division nodes, pillar industry related nodes, key livelihood project nodes and place name alias mapping relationships, forming a structured knowledge network covering local core related elements. Based on the hop distance and semantic relevance of the graph, the association level between the final candidate angle and the local core node is determined; Based on the aforementioned association level, localized candidate angles are generated; The localized candidate angles are input into a preset feasibility scoring model to evaluate spatial accessibility, resource matching degree, time feasibility and news value, and an execution suggestion card is generated.
Citation Information
Patent Citations
News topic analysis method and device
CN106934049A
Multi-platform fusion intelligent topic selection inspiration generation method and topic selection inspiration engine
CN118780280A