Intelligent news data management method and system based on big data

Through the intelligent management method of news data based on big data, the problem of insufficient analysis of information authenticity and communication potential in news data processing is solved, efficient information collection and personalized recommendation are achieved, and the timeliness and user experience of news dissemination is improved.

CN120179834AInactive Publication Date: 2025-06-20江苏桃李阁企业管理发展(集团)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510245711.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks an effective information authenticity and accuracy verification mechanism in news data processing, and the comprehensive analysis of communication potential is insufficient, mainly relying on surface indicators such as view count, forwarding count and like count.

Method used

Adopt the intelligent management method of news data based on big data, collect and filter source messages through the template library, generate structured abstracts, and automatically generate news drafts using preset templates and entity framework structures to conduct content review and personalized recommendations.

Benefits of technology

It improves the efficiency of information collection and processing, ensures the authority and credibility of the source, automatically generates news content, realizes personalized recommendations, and improves the timeliness and user experience of news dissemination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005295431580000021
    Figure BDA0005295431580000021
  • Figure BDA0005295431580000031
    Figure BDA0005295431580000031
  • Figure BDA0005295431580000041
    Figure BDA0005295431580000041
Patent Text Reader

Abstract

The invention discloses a news data intelligent management method and system based on big data, and belongs to the technical field of news data processing. The system comprises a data acquisition module, a media analysis module, a news editing module and an intelligent pushing module. The data acquisition module is used for acquiring a template library, acquiring messages from different information sources and extracting metadata; the media analysis module is used for screening out reference information sources from all the information sources, and extracting message content of each reference information source to generate a structured abstract; the news editing module is used for filling data according to a preset template in a template library, generating a news manuscript, auditing and editing the news manuscript and then publishing the news And the intelligent pushing module is used for automatically distributing the news to different platforms and recommending different users in a personalized manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of news data processing, and particularly to an intelligent management method and system for news data based on big data. Background Art

[0002] With the rapid development of information technology and the wide application of the Internet, the way of news dissemination has undergone a major change. A vast amount of information has flooded into news platforms, which not only greatly enriches the choices of users but also profoundly changes the way people obtain information. In this era of information explosion, news clients have gradually become one of the main channels for people to obtain information.

[0003] Currently, many media and news agencies adopt an automated way to collect, edit, and broadcast news. This way improves the speed and efficiency of news dissemination, enabling users to obtain information faster. However, multiple challenges are faced in this process. On the one hand, existing technologies mostly rely on manual review or rule-based keyword filtering. Although they can analyze the format and structure of news content, there is still a lack of an effective verification mechanism for the authenticity and accuracy of information. On the other hand, in information retrieval and recommendation systems, although existing research has adopted machine learning and natural language processing technologies for news content sorting and recommendation, there is still insufficient comprehensive analysis of communication potential. Existing algorithms mostly focus on surface indicators such as view counts, repost counts, and like counts, and fail to deeply explore the underlying communication mechanism. Therefore, at this stage, a more intelligent and efficient news data processing technical solution is needed to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent management method and system for news data based on big data to solve the problems raised in the above background art.

[0005] To solve the above technical problems, the present invention provides an intelligent management method for news data based on big data, including:

[0006] S100. Collect a template library, collect messages from different information sources, and extract metadata.

[0007] S200. Screen out reference information sources from all information sources, and extract the message content of each reference information source to generate a structured summary.

[0008] S300. Fill in data according to the preset templates in the template library, generate a news draft, and publish it after review and editing.

[0009] S400. Automatically distribute the news to different platforms and make personalized recommendations to different users.

[0010] In S100, the template library contains preset templates for different types of news, and the preset templates contain feature indexes and a fixed entity frame structure, and different entities are filled in each entity frame.

[0011] The news types are set in advance by managers, usually in a hierarchical structure with a three-level system of "major categories - subcategories - sub-fields" as the basic framework.

[0012] The major categories are divided based on the core attributes of news, usually covering the common standards of mainstream media organizations around the world, including current affairs, finance, society, science and technology, sports, culture and entertainment, environment and health, military security, etc.

[0013] Subcategories are further divided under the main categories according to event types or domain characteristics.

[0014] The sub-sectors are non-optional levels extended based on actual business needs and are used for in-depth exploration of vertical scenarios.

[0015] The Entity Framework structure is used to define the core entities and attribute slots that must be included in this type of news, including mandatory entities and optional entities. Mandatory entities cannot be missing, and optional entities can be empty.

[0016] The feature index refers to a set of iconic words or phrases that are used to quickly match news types in the original text and activate the corresponding preset templates.

[0017] The source refers to the source of the message. Metadata includes the release time, forwarding path diagram and publisher characteristics. The forwarding path diagram refers to a tree diagram used to describe the propagation path and hierarchical relationship of messages in the information platform, including root nodes and propagation nodes. Each propagation node includes the forwarding time and the number of forwarding times. The publisher characteristics include subject type and historical falsification rate. Subject types include authoritative media, self-media and individual users. The historical falsification rate refers to the proportion of messages released in the past that have been falsified.

[0018] S200 includes:

[0019] S201. Obtain a forwarding path diagram of the message Q in the source SOI, analyze the forwarding relationship between different propagation nodes, and take the forwarding party as the child node and the forwarded party as the parent node.

[0020] S202, analyze the forwarding time of all nodes and calculate the time range T max , count the number of direct forwarding times j of all parent nodes, and obtain the maximum number of direct forwarding times j among all parent nodes max .

[0021] S203, calculate the influence coefficient YX of each parent node according to the formula; set the threshold YX thr , the marker influence coefficient is greater than YX thrThe parent node. The formula is as follows:

[0022]

[0023] In the formula, α is the time decay coefficient, and T i is the forwarding time difference between the parent node and its i-th child node.

[0024] In the calculation formula of the influence coefficient, is used to measure the direct propagation ability of the parent node, is used to measure the influence degree of the time decay effect.

[0025] The value of α is adaptively set according to the scenario. In short-term hot events, a high α value is set to strengthen real-time performance; while in long-term propagation events, a low α value is set to weaken the time penalty.

[0026] The influence coefficient is determined by the product of the propagation times ratio and the time decay weight. By associating the two through multiplication, both the core propagation ability can be reflected, and the parent nodes with serious delays but inflated forwarding volumes can be suppressed.

[0027] S204. Mark the longest node chain DL where each marked parent node is located in the forwarding path graph. Analyze the number of propagation nodes JS sum of DL and the number of propagation nodes JS n separating the corresponding marked parent node from the root node.

[0028] S205. Count the number m of all marked parent nodes, and set different weight coefficients for each subject type. Substitute into the formula to calculate the weight index WC of the information source SOI SOI :

[0029]

[0030] In the formula, W SOI is the weight coefficient corresponding to the subject type of the information source SOI, Z SOI is the historical falsification rate of the information source SOI, and YX n is the influence coefficient of the n-th marked parent node.

[0031] Different weight coefficients are set for each subject type, which are specifically set by the management according to actual needs. Usually, the weight coefficient of authoritative media is greater than that of self-media, and the weight coefficient of self-media is greater than that of individual users.

[0032] The weight index of the information source is jointly determined by the weight coefficient, the historical falsification rate, and the marked parent node.

[0033] The weight coefficient and the historical falsification rate represent the comprehensive credibility of the information source. The more it belongs to authoritative media and the lower the historical falsification rate, the greater the comprehensive credibility.

[0034] Marking the parent node represents the degree of the message propagation potential. When more parent nodes with large influence coefficients are at the end position of the propagation node chain, it indicates that the message has great propagation potential and there is a possibility of further increasing the news popularity.

[0035] S206. Calculate the weight index of each information source respectively and set the threshold WC. thr , and use the information sources with weight index greater than WC thr as the reference information sources. Extract the content Q of the messages of each reference information source by using natural language processing technology and generate a structured summary.

[0036] S300 includes:

[0037] S301. Analyze the structured summaries of all reference information sources, extract keywords and match them with the feature indexes in the template library to obtain the preset template MB corresponding to the type of news with successful matching.

[0038] S302. Obtain the entity framework structure of the preset template MB, extract the entities in the structured summary of each reference information source, and fill them into different entity boxes in the entity framework structure respectively.

[0039] The mapping between the news type and the preset template is a standardized generation framework centered on entity filling + feature triggering. Each news type corresponds to an exclusive preset template, and the entity elements that must be included in this type of news are preset in the template. The original data is parsed through an automated process and adapted to the template to generate compliant text.

[0040] Analyze the entity type distribution matrix and event verb syntactic pattern of the structured summary, combine the domain feature constraint conditions and the dynamic weight model, and match the predefined feature indexes in the template library to achieve accurate news type recognition.

[0041] The core elements include feature extraction, knowledge mapping and decision-making mechanism. Feature extraction: calculation of entity distribution entropy value (quantifying information focus degree) + parsing of SVO (subject-verb-object) structure; Knowledge mapping: geographical coding rules (ISO-3166) in the template library, role classification system (such as political / commercial figure identification); Decision-making mechanism: three-layer funnel architecture of rough classification (event domain judgment) → fine classification (subtype recognition) → rule correction (forced type assertion).

[0042] S303. Use the static word vector algorithm to calculate the semantic similarity between every two entities in each entity box respectively, set the threshold C, and mark the entity boxes with semantic similarity less than C.

[0043] S304. Classify all entities with semantic similarity greater than C in the marked entity boxes into the same category, and calculate the release duration TP according to the release time of the message Q in the information source corresponding to each entity. f, set the standard duration TP s , substitute into the formula to calculate the credibility coefficient KXD of each category:

[0044]

[0045] In the formula, v is the number of all entities under the category, WC f is the weight index of the information source corresponding to the f-th entity.

[0046] S305. Delete all categories with non-highest credibility coefficients within each marked entity box. According to the text information of all filled entities in the entity framework structure, use natural language generation technology to generate a news draft according to the requirements of the corresponding news type.

[0047] Through the preset typed templates, it is possible to improve the content generation efficiency and standardization degree on the premise of ensuring the integrity of the core news facts. This framework is especially suitable for scenarios that require rapid production of structured content, such as news flashes, earnings reports, and emergencies. At the same time, it allows for in-depth customization in vertical fields by adding sub-category templates.

[0048] S306. Conduct content review and automatic editing and replacement for the news draft. If it meets the preset release conditions, it will be published as a news release on the information platform.

[0049] Automatic editing includes title generation, content summary, format adjustment, etc. The edited news content is reviewed, and possible problem content, such as sensitive topics and false information, is automatically identified and marked, and public opinion analysis and keyword review are carried out.

[0050] In S400, the news is automatically synchronized and published to different platforms and social media, and is personalized and pushed to the user terminals that usually browse the same news type, and interactive functions such as user comments, likes, and shares are provided.

[0051] Provide customized news content according to the user's reading history and preferences.

[0052] The intelligent management system for news data based on big data includes a data collection module, a media analysis module, a news editing module, and an intelligent push module.

[0053] The data collection module is used to collect the template library, collect messages from different information sources, and extract metadata.

[0054] The media analysis module is used to screen out reference information sources from all information sources and extract the message content of each reference information source to generate a structured summary.

[0055] The news editing module is used to fill in data according to the preset templates in the template library, generate a news draft, and publish it after review and editing.

[0056] The intelligent push module is used to automatically distribute news to different platforms and provide personalized recommendations to different users.

[0057] The data collection module includes a template information collection unit and a message content collection unit.

[0058] The template information collection unit is used to collect the template library. The template library contains preset templates for different types of news, and the preset templates include feature indexes and fixed entity framework structures.

[0059] The message content collection unit is used to collect messages from different information sources and extract metadata. An information source refers to the source of a message. Metadata includes the release time, the forwarding path diagram, and the characteristics of the publisher. The forwarding path diagram is a tree diagram used to describe the propagation path and hierarchical relationship of messages in an information platform, including a root node and propagation nodes. Each propagation node includes the forwarding time and the number of times it has been forwarded. The characteristics of the publisher include the subject type and the historical falsification rate.

[0060] The media analysis module includes a quality assessment unit and a summary extraction unit.

[0061] The quality assessment unit is used to evaluate the propagation quality of each information source and screen out reference information sources.

[0062] First, obtain the forwarding path diagram of message Q in information source SOI, with the forwarding party as the child node and the forwarded party as the parent node. Analyze the forwarding time of all nodes and calculate the time range T max , count the direct forwarding times j of all parent nodes, and obtain the maximum direct forwarding times j among all parent nodes max .

[0063] Second, according to the formula: calculate the influence coefficient of each parent node. Set the threshold YX thr , and mark the parent nodes with an influence coefficient greater than YX thr . Among them, α is the time decay coefficient, and T i is the forwarding time difference between the parent node and its i-th child node.

[0064] Then, mark the longest node chain DL where each marked parent node is located, and analyze the number of propagation nodes JS of DL sum and the number of propagation nodes JS between the corresponding marked parent node and the root node n . Count the number m of all marked parent nodes, and set different weight coefficients for each subject type.

[0065] Finally, according to the formula: calculate the weight index of information source SOI. Calculate the weight index of each information source separately, set the threshold WC thr , and mark the information sources with a weight index greater than WC thrThe information source is used as a reference information source. Among them, W SOI is the weight coefficient corresponding to the main body type of the information source SOI, and Z SOI is the historical falsification rate of the information source SOI, and YX n is the influence coefficient of the nth marked parent node.

[0066] The abstract extraction unit uses natural language processing technology to extract the message Q content of the reference information source and generate a structured abstract.

[0067] The news editing module includes a template filling unit and a review and editing unit.

[0068] The template filling unit is used to fill data according to the preset templates in the template library and generate a news draft.

[0069] First, analyze the structured abstracts of all reference information sources, extract keywords and match them with the feature indexes in the template library to obtain the preset template MB corresponding to the successfully matched news type and its entity framework structure. Extract the entities in the structured abstracts of each reference information source and fill them into different entity boxes in the entity framework structure respectively.

[0070] Second, use the static word vector algorithm to calculate the semantic similarity between every two entities in each entity box respectively, set a threshold C, and mark the entity boxes with semantic similarity less than C. Classify all entities with semantic similarity greater than C in the marked entity boxes into the same category, and calculate the release duration TP according to the message Q release time in the information source corresponding to each entity f .

[0071] Finally, set the standard duration TP s , and according to the formula: Calculate the credibility coefficients of each category. Delete all categories with non - highest credibility coefficients in each marked entity box, and generate a news draft according to the text information of all filled entities in the entity framework structure by using natural language generation technology according to the requirements of the corresponding news type.

[0072] Among them, v is the number of all entities in the category, and WC f is the weight index of the information source corresponding to the fth entity.

[0073] The review and editing unit conducts content review and automatic editing and replacement on the news draft, and if it meets the preset release conditions, it is published as a news release draft on the information platform.

[0074] The intelligent push module includes a platform release unit, a news recommendation unit and a user interaction unit.

[0075] The platform release unit is used to synchronously publish news to different platforms and social media.

[0076] The news recommendation unit is used to push personalized news to user terminals that routinely browse the same type of news.

[0077] The user interaction unit is used to provide interactive functions such as user comments, likes, and sharing.

[0078] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0079] Efficient information collection and processing: By quickly collecting news information from multiple information sources and extracting metadata such as release time and publisher characteristics, the efficiency and accuracy of information collection are significantly improved.

[0080] Intelligent information source screening: By analyzing the forwarding path diagram and the influence of nodes, and using the influence coefficient and historical falsification rate to calculate the weight of information sources, it is ensured that the selected information sources are more authoritative and credible. This data analysis-based method is more effective than traditional manual information source screening.

[0081] Automated news generation: Using the structured summaries of reference information sources and preset templates, this solution can automatically generate text that meets news requirements, reducing the workload of manual editing and improving the speed and consistency of content generation.

[0082] Personalized recommendation system: By analyzing users' reading habits, the system can provide personalized recommendations for news content, enabling users to receive news that better matches their interests and improving user participation and satisfaction.

[0083] Real-time dissemination and multi-platform publishing: This technical solution allows news to be instantly distributed to multiple platforms and personalized pushed after generation, greatly enhancing the timeliness of news dissemination and adapting to the rapidly changing media environment.

[0084] Intelligent content review: Before publication, the system can automatically identify and mark possible problematic content such as sensitive topics and false information, conduct public opinion analysis and keyword review to ensure the compliance and security of the published content.

[0085] Through the above advantages, this technical solution can provide a more intelligent, efficient, and personalized news data processing solution in the context of the information age, meet the rapidly developing information dissemination needs, and improve the user experience. Description of the Drawings

[0086] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0087] Figure 1 is a schematic flow diagram of the intelligent management method for news data based on big data of the present invention;

[0088] Figure 2 It is a schematic structural diagram of the news data intelligent management system based on big data of the present invention. Detailed implementation manners

[0089] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0090] Please refer to Figure 1 , the present invention provides a news data intelligent management method based on big data, including:

[0091] S100. Collect a template library, collect messages from different information sources and extract metadata.

[0092] S200. Screen out reference information sources among all information sources, and extract the message content of each reference information source to generate a structured summary.

[0093] S300. Fill in data according to the preset templates in the template library, generate a news draft, and publish it after review and editing.

[0094] S400. Automatically distribute the news to different platforms and make personalized recommendations to different users.

[0095] In S100, the template library contains preset templates for different types of news. The preset templates include feature indexes and fixed entity framework structures, and different entities are filled in each entity box.

[0096] The news type is set in advance by the management personnel, usually adopting a hierarchical structure, with a three-level system of "major category - subcategory - subdivision field" as the basic framework.

[0097] The major category is divided based on the core attributes of the news, usually covering the general standards of global mainstream media organizations, including current politics, finance, society, technology, sports, culture and entertainment, environmental health, military security, etc.

[0098] The subcategory is further split under the major category according to the event type or field characteristics. For example, under the finance major category, there are macroeconomics, financial markets, and company dynamics; under the technology major category, there are artificial intelligence, hardware technology, and network security, etc.

[0099] The subdivision field is an optional level extended according to actual business needs and is used for in-depth mining of vertical scenarios. For example, sports (major category) → football (subcategory) → Premier League (subdivision).

[0100] The Entity Framework structure is used to define the core entities and attribute slots that must be included in this type of news, including mandatory entities and optional entities. Mandatory entities cannot be missing, and optional entities can be empty.

[0101] The feature index refers to a set of iconic words or phrases that are used to quickly match news types in the original text and activate the corresponding preset templates.

[0102] The source refers to the origin of the message. Metadata includes the release time, forwarding path diagram and publisher characteristics. The forwarding path diagram refers to a tree diagram used to describe the propagation path and hierarchical relationship of messages in the information platform, including root nodes and propagation nodes. Each propagation node includes the forwarding time and the number of forwarding times.

[0103] The publisher characteristics include subject type and historical falsification rate. Subject types include authoritative media, self-media and individual users. The historical falsification rate refers to the proportion of messages published in the past that have been falsified.

[0104] S200 includes:

[0105] S201. Obtain a forwarding path diagram of the message Q in the source SOI, analyze the forwarding relationship between different propagation nodes, and take the forwarding party as the child node and the forwarded party as the parent node.

[0106] S202, analyze the forwarding time of all nodes and calculate the time range T max , count the number of direct forwarding times j of all parent nodes, and obtain the maximum number of direct forwarding times j among all parent nodes max .

[0107] S203, calculate the influence coefficient YX of each parent node according to the formula; set the threshold YX thr , the marker influence coefficient is greater than YX thr The parent node of . The formula is as follows:

[0108]

[0109] Where α is the time attenuation coefficient, T i is the forwarding time difference between the parent node and its i-th child node.

[0110] In the calculation formula of influence coefficient, Used to measure the direct propagation ability of the parent node. Used to measure the impact of time decay effect.

[0111] The α value is set adaptively according to the scenario. In short-term hot events, a high α value is set to enhance real-time performance; in long-term propagation events, a low α value is set to weaken the time penalty.

[0112] The influence coefficient is determined by the product of the propagation times ratio and the time decay weight. By associating the two through multiplication, both the core propagation ability can be reflected and the parent nodes with serious delays but inflated forwarding volumes can be suppressed.

[0113] S204. Label the longest node chain DL where each marked parent node is located in the forwarding path graph. Analyze the number of propagation nodes JS of DL sum and the number of propagation nodes JS corresponding to the interval between the marked parent node and the root node n .

[0114] S205. Count the number m of all marked parent nodes, and set different weight coefficients for each subject type. Substitute into the formula to calculate the weight index WC of the information source SOI SOI :

[0115]

[0116] In the formula, W SOI is the weight coefficient corresponding to the subject type of the information source SOI, Z SOI is the historical falsification rate of the information source SOI, YX n is the influence coefficient of the nth marked parent node.

[0117] Different weight coefficients are set for each subject type, which are specifically set by the management according to actual needs. Usually, the weight coefficient of authoritative media is greater than that of self-media, and the weight coefficient of self-media is greater than that of individual users.

[0118] The weight index of the information source is jointly determined by the weight coefficient, the historical falsification rate, and the marked parent node.

[0119] The weight coefficient and the historical falsification rate represent the comprehensive credibility of the information source. The more it belongs to authoritative media and the lower the historical falsification rate, the greater the comprehensive credibility.

[0120] The marked parent node represents the degree of the message's propagation potential. When more parent nodes with large influence coefficients are at the end position of the propagation node chain, it indicates that the message has great propagation potential and the possibility of further increasing the news heat.

[0121] S206. Calculate the weight index of each information source respectively, and set the threshold WC thr , and use the information sources with weight indexes greater than WC thr as reference information sources. Use natural language processing technology to extract the message Q content of each reference information source and generate a structured summary.

[0122] S300 includes:

[0123] S301. Analyze the structured abstracts of all reference sources, extract keywords and match them with the feature indexes in the template library to obtain the preset template MB corresponding to the successfully matched news type.

[0124] S302. Obtain the entity framework structure of the preset template MB, extract the entities in the structured abstracts of each reference source, and fill them into different entity boxes in the entity framework structure respectively.

[0125] The mapping between news types and preset templates is a standardized generation framework centered on entity filling + feature triggering. Each news type corresponds to a dedicated preset template, and the entity elements that must be included in this type of news are preset in the template. The original data is parsed through an automated process and adapted to the template to generate compliant text.

[0126] Analyze the entity type distribution matrix and event verb syntactic patterns of the structured abstract, combine the domain feature constraint conditions and the dynamic weight model, and match the predefined feature indexes in the template library to achieve accurate news type recognition.

[0127] The core elements include feature extraction, knowledge mapping, and decision-making mechanisms. Feature extraction: calculation of entity distribution entropy value (quantifying information focus) + parsing of SVO (subject-verb-object) structure; Knowledge mapping: geographical coding rules (ISO-3166) in the template library, role classification system (such as political / business person identification); Decision-making mechanism: a three-layer funnel architecture of rough classification (event domain judgment) → fine classification (subtype recognition) → rule correction (forced type assertion).

[0128] Example logic chain: The structured abstract appears with high-frequency verbs such as "disaster level code: EH-Ⅲ" + "rescue" + "transfer" → the classifier activates the template matching path of "environmental health - natural disaster" → lock the final type based on the GIS geographical attribute verification result.

[0129] S303. Use the static word vector algorithm to calculate the semantic similarity between every two entities in each entity box respectively, set the threshold C, and mark the entity boxes with semantic similarity less than C.

[0130] S304. Classify all entities with semantic similarity greater than C in the marked entity boxes into the same category, calculate the release duration TP according to the message Q release time in the source corresponding to each entity f , set the standard duration TP s , substitute it into the formula to calculate the credibility coefficient KXD of each category:

[0131]

[0132] In the formula, v is the total number of all entities in the category, WC f is the weight index of the source corresponding to the f-th entity.

[0133] S305, deleting all classes whose credibility coefficients are not the highest in each marked entity frame, and generating news manuscripts according to the text information of all filled entities in the entity frame structure using natural language generation technology in accordance with the requirements of the corresponding news type.

[0134] By presetting typed templates, we can improve the efficiency and standardization of content generation while ensuring the integrity of the core facts of the news. This framework is particularly suitable for scenarios such as news flashes, financial report interpretations, and emergencies that require rapid production of structured content, while allowing for deep customization in vertical fields by adding sub-category templates.

[0135] S306. Conduct content review and automatic editing and replacement on the news draft. If it meets the preset release conditions, it will be released as a news release on the information platform.

[0136] Automatic editing includes title generation, content summary, format adjustment, etc. The edited news content is reviewed, and possible problematic content such as sensitive topics and false information are automatically identified and marked, and public opinion analysis and keyword review are performed.

[0137] In S400, news is automatically published to different platforms and social media, and is pushed to user terminals that browse the same type of news on a daily basis in a personalized manner, and interactive functions such as user comments, likes and sharing are provided.

[0138] Provide users with customized news content based on their reading history and preferences.

[0139] See also Figure 2 The present invention provides an intelligent news data management system based on big data, including a data acquisition module, a media analysis module, a news editing module and an intelligent push module.

[0140] The data collection module is used to collect template libraries, collect messages from different sources and extract metadata.

[0141] The media analysis module is used to filter out reference sources from all sources, extract the message content of each reference source and generate a structured summary.

[0142] The news editing module is used to fill in data according to the preset templates in the template library, generate news drafts, and publish them after review and editing.

[0143] The intelligent push module is used to automatically distribute news to different platforms and provide personalized recommendations to different users.

[0144] The data collection module includes a template information collection unit and a message content collection unit.

[0145] The template information collection unit is used to collect the template library. The template library contains preset templates for different types of news, and the preset templates include feature indexes and fixed entity framework structures.

[0146] The message content collection unit is used to collect messages from different information sources and extract metadata. The information source refers to the source of the message. The metadata includes the release time, the forwarding path diagram, and the characteristics of the publisher. The forwarding path diagram is a tree diagram used to describe the propagation path and hierarchical relationship of messages in the information platform, including the root node and propagation nodes. Each propagation node includes the forwarding time and the number of times it has been forwarded. The characteristics of the publisher include the subject type and the historical falsification rate.

[0147] The media analysis module includes a quality assessment unit and a summary extraction unit.

[0148] The quality assessment unit is used to evaluate the propagation quality of each information source and screen out reference information sources.

[0149] First, obtain the forwarding path diagram of message Q in information source SOI, with the forwarding party as the child node and the forwarded party as the parent node. Analyze the forwarding time of all nodes and calculate the time range T max , count the direct forwarding times j of all parent nodes, and obtain the maximum direct forwarding times j among all parent nodes max .

[0150] Second, according to the formula: calculate the influence coefficient of each parent node. Set the threshold YX thr , and mark the parent nodes with an influence coefficient greater than YX thr . Among them, α is the time decay coefficient, and T i is the forwarding time difference between the parent node and its i-th child node.

[0151] Then, for each marked parent node, mark the longest node chain DL where it is located, and analyze the number of propagation nodes JS of DL sum and the number of propagation nodes JS between the corresponding marked parent node and the root node n . Count the number m of all marked parent nodes, and set different weight coefficients for each subject type.

[0152] Finally, according to the formula: calculate the weight index of information source SOI. Calculate the weight index of each information source respectively, and set the threshold WC thr , and use the information sources with a weight index greater than WC thr as reference information sources. Among them, W SOI is the weight coefficient corresponding to the subject type of information source SOI, Z SOI is the historical falsification rate of information source SOI, and YX n is the influence coefficient of the n-th marked parent node.

[0153] The abstract extraction unit uses natural language processing technology to extract the content of message Q from the reference information source and generate a structured abstract.

[0154] The news editing module includes a template filling unit and a review and editing unit.

[0155] The template filling unit is used to fill data according to the preset templates in the template library and generate a news draft.

[0156] First, analyze the structured abstracts of all reference information sources, extract keywords and match them with the feature indexes in the template library to obtain the preset template MB corresponding to the news type with successful matching and its entity framework structure. Extract the entities in the structured abstract of each reference information source and fill them into different entity boxes in the entity framework structure respectively.

[0157] Second, use the static word vector algorithm to calculate the semantic similarity between every two entities in each entity box respectively, set a threshold C, and mark the entity boxes where the semantic similarity is less than C. Classify all entities with a semantic similarity greater than C in the marked entity boxes into the same category, and calculate the release duration TP according to the release time of message Q in the corresponding information sources of each entity. f 。

[0158] Finally, set the standard duration TP s , according to the formula: Calculate the credibility coefficients of each category. Delete all categories with non-highest credibility coefficients in each marked entity box, and generate a news draft according to the text information of all filled entities in the entity framework structure by using natural language generation technology according to the requirements of the corresponding news type.

[0159] Among them, v is the number of all entities in the category, and WC f is the weight index of the information source corresponding to the f-th entity.

[0160] The review and editing unit conducts content review and automatic editing and replacement on the news draft, and if it meets the preset release conditions, it is published as a news release draft on the information platform.

[0161] The intelligent push module includes a platform release unit, a news recommendation unit, and a user interaction unit.

[0162] The platform release unit is used to synchronously release news to different platforms and social media.

[0163] The news recommendation unit is used to push news personalized to the user terminals that usually browse the same news type.

[0164] The user interaction unit is used to provide interactive functions such as user comments, likes, and shares.

[0165] Example 1:

[0166] Assume that the maximum difference in forwarding time among all nodes is 210 s, the number of direct forwardings of the parent node A1 is 3 times, and the maximum number of direct forwardings among all parent nodes is 9 times;

[0167] When there are three child nodes under the parent node A1, and the forwarding time differences are 20 s, 30 s, and 40 s respectively, and the time decay coefficient α is 2, substitute into the formula to calculate the influence coefficient of the parent node A1:

[0168] Influence coefficient:

[0169] Then the influence coefficient of the parent node A1 is 0.24.

[0170] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0171] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent management method for news data based on big data, characterized by: The method includes: S100, a collection template library, collecting messages from different sources and extracting metadata; S200, selecting reference information sources from all information sources, extracting the message content of each reference information source and generating a structured summary; S300, filling in data according to the preset template in the template library, generating a news draft, reviewing and editing it, and then publishing it; S400, automatically distribute news to different platforms and provide personalized recommendations to different users.

2. The intelligent management method of news data based on big data according to claim 1 is characterized by: In S100, the template library contains preset templates for different types of news, and the preset templates contain feature indexes and a fixed entity frame structure, and different entities are filled in each entity frame; The source refers to the origin of the message; metadata includes release time, forwarding path diagram and publisher characteristics; the forwarding path diagram refers to a tree diagram used to describe the propagation path and hierarchical relationship of messages in the information platform, including root nodes and propagation nodes, and each propagation node includes the forwarding time and the number of forwarding times; publisher characteristics include subject type and historical falsification rate; subject types include authoritative media, self-media and individual users; the historical falsification rate refers to the proportion of messages published in the past that have been falsified.

3. The intelligent management method of news data based on big data according to claim 2 is characterized by: S200 includes: S201, obtaining a forwarding path diagram of the message Q in the source SOI, analyzing the forwarding relationship between different propagation nodes, taking the forwarding party as the child node and the forwarded party as the parent node; S202, analyze the forwarding time of all nodes and calculate the time range T max , count the number of direct forwarding times j of all parent nodes, and obtain the maximum number of direct forwarding times j among all parent nodes max ; S203, calculate the influence coefficient YX of each parent node according to the formula; set the threshold YX thr , the marker influence coefficient is greater than YX thr The parent node of ; the formula is as follows: Where α is the time attenuation coefficient, T i is the forwarding time difference between the parent node and its i-th child node; S204, in the forwarding path graph, mark the longest node chain DL for each marked parent node; analyze the number of propagation nodes JS of DL sum And the corresponding number of propagation nodes JS between the parent node and the root node n ; S205. Count the number of all marked parent nodes m, set different weight coefficients for each subject type; substitute into the formula to calculate the weight index WC of the source SOI SOI : Where W SOI is the weight coefficient corresponding to the source SOI main body type, Z SOI is the historical falsification rate of the source SOI, YX n is the influence coefficient of the nth marked parent node; S206, calculate the weight index of each information source respectively, and set the threshold WC thr , the weight index is greater than WC thr The sources are used as reference sources; natural language processing technology is used to extract the message Q content of each reference source and generate a structured summary.

4. The intelligent management method of news data based on big data according to claim 3 is characterized by: S300 includes: S301, analyzing the structured summaries of all reference sources, extracting keywords and matching them with feature indexes in the template library, and obtaining a preset template MB corresponding to the type of news that has been successfully matched; S302, obtaining an entity framework structure of a preset template MB, extracting entities in the structured summary of each reference source, and filling them into different entity boxes in the entity framework structure respectively; S303, using a static word vector algorithm to calculate the semantic similarity between every two entities in each entity box, setting a threshold C, and marking the entity boxes with a semantic similarity less than C; S304: Classify all entities in the marked entity box whose semantic similarity is greater than C as the same category, and calculate the release time TP according to the release time of the message Q in the corresponding source of each entity. f , set the standard time TP s , substitute into the formula to calculate the credibility coefficient KXD of each category: In the formula, v is the number of entities in the class, WC f is the weight index of the source corresponding to the f-th entity; S305, deleting all classes with non-highest credibility coefficients in each marked entity frame, and generating news manuscripts according to the text information of all filled entities in the entity frame structure using natural language generation technology in accordance with the requirements of the corresponding news type; S306. Conduct content review and automatic editing and replacement on the news draft. If it meets the preset release conditions, it will be released as a news release on the information platform.

5. The intelligent management method of news data based on big data according to claim 4 is characterized in that: In S400, news is automatically and synchronously published to different platforms and social media, and is pushed to user terminals that browse the same type of news on a daily basis in a personalized manner, and interactive functions such as user comments, likes and sharing are provided.

6. The intelligent management system for news data based on big data is characterized by: The system includes data collection module, media analysis module, news editing module and intelligent push module; The data collection module is used to collect template libraries, collect messages from different sources and extract metadata; The media analysis module is used to select reference sources from all sources, extract the message content of each reference source and generate a structured summary; The news editing module is used to fill in data according to the preset templates in the template library, generate news drafts, and publish them after review and editing; The intelligent push module is used to automatically distribute news to different platforms and provide personalized recommendations to different users.

7. The intelligent news data management system based on big data according to claim 6 is characterized by: The data collection module includes a template information collection unit and a message content collection unit; The template information collection unit is used to collect template libraries; the template library contains preset templates for different types of news, and the preset templates contain feature indexes and fixed entity framework structures; The message content collection unit is used to collect messages from different sources and extract metadata; the source refers to the source of the message; the metadata includes the release time, forwarding path map and publisher characteristics; The forwarding path diagram refers to a tree diagram used to describe the propagation path and hierarchical relationship of messages in the information platform, including root nodes and propagation nodes. Each propagation node includes the forwarding time and the number of forwarding times. Issuer characteristics include subject type and historical falsification rate.

8. The intelligent news data management system based on big data according to claim 7 is characterized by: The media analysis module includes a quality assessment unit and a summary extraction unit; The quality assessment unit is used to assess the transmission quality of each information source and select reference information sources; First, obtain the forwarding path diagram of message Q in the source SOI, take the forwarding party as the child node and the forwarded party as the parent node; analyze the forwarding time of all nodes and calculate the time range T max , count the number of direct forwarding times j of all parent nodes, and obtain the maximum number of direct forwarding times j among all parent nodes max ; Secondly, according to the formula: Calculate the influence coefficient of each parent node; set the threshold YX thr , the marker influence coefficient is greater than YX thr The parent node of ; where α is the time decay coefficient, T i is the forwarding time difference between the parent node and its i-th child node; Then, each marked parent node is marked with the longest node chain DL, and the number of propagation nodes JS of DL is analyzed sum And the corresponding number of propagation nodes JS between the parent node and the root node n ; Count the number of all marked parent nodes m, and set different weight coefficients for each subject type; Finally, according to the formula: Calculate the weight index of the source SOI; calculate the weight index of each source separately and set the threshold WC thr , the weight index is greater than WC thr The source is used as the reference source; among them, W SOI is the weight coefficient corresponding to the source SOI main body type, Z SOI is the historical falsification rate of the source SOI, YX n is the influence coefficient of the nth marked parent node; The summary extraction unit uses natural language processing technology to extract the message Q content of the reference source and generate a structured summary.

9. The intelligent news data management system based on big data according to claim 8 is characterized by: The news editing module includes a template filling unit and a review editing unit; The template filling unit is used to fill in data according to the preset template in the template library and generate a news manuscript; First, the structured summaries of all reference sources are analyzed, keywords are extracted and matched with feature indexes in the template library, and the preset template MB and its entity framework structure corresponding to the successfully matched type of news are obtained; entities in the structured summaries of each reference source are extracted and filled into different entity boxes in the entity framework structure respectively; Secondly, the static word vector algorithm is used to calculate the semantic similarity between every two entities in each entity box, set a threshold C, and mark the entity boxes with semantic similarity less than C; all entities with semantic similarity greater than C in the marked entity box are classified into the same category, and the release time TP is calculated according to the release time of the message Q in the corresponding source of each entity f ; Finally, set the standard time TP s , according to the formula: Calculate the credibility coefficient of each category; delete all categories with non-highest credibility coefficients in each marked entity frame, and generate news manuscripts according to the corresponding news type requirements using natural language generation technology based on the text information of all filled entities in the entity frame structure; Among them, v is the number of all entities under the class, WC f is the weight index of the source corresponding to the f-th entity; The review and editing unit conducts content review and automatic editing and replacement on news drafts. If they meet the preset release conditions, they will be published as news releases on the information platform.

10. The intelligent news data management system based on big data according to claim 9 is characterized in that: The intelligent push module includes a platform publishing unit, a news recommendation unit, and a user interaction unit; The platform publishing unit is used to publish news to different platforms and social media simultaneously; The news recommendation unit is used to push personalized news to user terminals that browse the same type of news on a daily basis; The user interaction unit is used to provide interactive functions such as user comments, likes and sharing.