A media fusion content clue information intelligent analysis method
By analyzing news leads through a file upload page and pre-configured rules, combined with a keyword database and image recognition model, the problem of the large number and diverse forms of news leads has been solved. This has enabled efficient and personalized news lead processing and recommendation, improving processing efficiency and content quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-03-31
AI Technical Summary
The current news lead analysis process suffers from a large number of news leads, irregular and chaotic data, which makes manual screening time-consuming and labor-intensive, unable to track hot topics in real time, and unable to provide personalized recommendations based on editors' expertise, resulting in low efficiency and poor content quality.
News leads and their attributes are collected through the lead file upload page. The data is then analyzed using pre-configured lead extraction rules to generate alert configurations and news editing tasks. By combining a keyword database and an image recognition model, structured and personalized recommendations for news leads are achieved.
It has improved the efficiency and timeliness of news processing, enhanced the efficiency and quality of news gathering and editing, and enabled the priority processing and personalized recommendation of high-value leads.
Smart Images

Figure CN120705417B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent information analysis, specifically to a method for intelligent analysis of content clues in converged media. Background Technology
[0002] With the continuous development and rapid popularization of the Internet and artificial intelligence technologies, the habits of collecting news leads in power news publicity have undergone tremendous changes. In the digital network environment, a large amount of news information brings convenience to journalists, but also brings problems such as a huge number of news leads, irregular data, and chaos.
[0003] The following technical problems often exist in the existing news lead analysis process:
[0004] First, news leads are generated at an alarming rate. Relying on manual screening of valuable news leads from massive amounts of data and publishing news is not only time-consuming and labor-intensive, but also makes it impossible to track current news hotspots in real time, which can easily lead to the problem of untimely release of major news.
[0005] Second, with the popularization of media such as text, images, and short videos, news leads are taking on more and more forms. Faced with a massive amount of news leads in various forms, how to efficiently extract the most newsworthy news leads has become an urgent problem to be solved.
[0006] Third, different editors have expertise in different areas, but traditional systems cannot make personalized recommendations based on editors' expertise and preferences, resulting in low editors' work enthusiasm and adoption rate, which in turn leads to low efficiency in news gathering and editing and poor content quality. Summary of the Invention
[0007] The summary section of this invention provides a brief overview of the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0008] This invention proposes an intelligent analysis method for converged media content clue information to solve one or more of the technical problems mentioned in the background section above.
[0009] This invention provides an intelligent analysis method for converged media content clue information, including:
[0010] The news tip file and news tip attribute information submitted by users are collected through the tip file upload page;
[0011] Based on the pre-configured clue extraction rules, the news clue files are extracted to obtain news clue information and update the news clue information list.
[0012] Based on the news lead file and news lead attribute information, generate a reminder configuration and add the reminder configuration to the news lead information. The reminder configuration includes the reminder target, reminder time, and reminder method.
[0013] Based on the top N news leads in the news lead information list according to their overall news value, generate news editing tasks and send them to the corresponding news editing terminals.
[0014] Optionally, the news lead file can be a text file; and
[0015] Extract news leads from news lead files, obtain news lead information, and update the news lead information list, including:
[0016] The news lead file is matched against a pre-configured keyword database to obtain the matched keyword groups, the keyword category of each keyword, and the number of times each keyword is matched. Different keyword categories are configured with different statistical weights.
[0017] Based on keyword groups, keyword categories for each keyword, hit counts, and statistical weights for different keyword categories, the isolated news value of news lead files is generated.
[0018] Based on the cumulative hit count of each keyword in the keyword library within the target time period, a global news value score is generated for each news lead file. Based on the isolated news value score and the global news value score, an overall news value score for each news lead file is generated. The overall news value score, keyword group, and news lead attribute information of each news lead file constitute the corresponding news lead information.
[0019] The news leads that rank in the top N in terms of overall news value are marked in the news lead list.
[0020] Optionally, news lead attribute information includes the news source organization information; and
[0021] Based on the news lead file and news lead attribute information, generate an alert configuration and add the alert configuration to the news lead information, including;
[0022] Based on the information of the news source organization and the pre-configured organizational structure information, determine the basic reminder recipients;
[0023] Based on the overall news value and the pre-configured list of alert scopes, an expanded alert object is generated. The basic alert object and the expanded alert object together constitute the alert object.
[0024] Receive the reminder time and reminder method configured by the user for news tip files;
[0025] Configure the reminder by specifying the recipient, time, and method, and then add the reminder configuration to the news lead information.
[0026] Optionally, news lead files can be image files; and
[0027] Extract news leads from news lead files, obtain news lead information, and update the news lead information list, including:
[0028] Input an image file into an image content recognition model to obtain the content information displayed in the image file;
[0029] The content information is treated as a text file, and a first news value score corresponding to the content information is generated.
[0030] Input the image file into the news value extraction model to obtain the second news value score corresponding to the image file;
[0031] The overall news value of the image file is obtained by combining the primary news value score and the secondary news value score.
[0032] Optionally, the keyword library includes a configured keyword set and an online keyword set. Both the configured and online keyword sets are configured with corresponding keyword category weights. The keyword library is updated through the following steps:
[0033] For each keyword in the keyword library, determine the number of times the keyword was adopted in the news within the target time period, sort the keywords in the keyword library according to the number of news adoptions, and remove the M keywords at the bottom of the list from the keyword library.
[0034] Get the current set of trending words and select M trending words from it to add to the keyword library.
[0035] Optionally, news lead attribute information includes the time of the event; and the above methods also include:
[0036] For each news lead that ranks in the top N in terms of overall news value, query related news lead groups based on the matched keyword groups.
[0037] The related news leads in the related news lead information group are sorted according to the chronological order of the events to obtain the related news lead information sequence;
[0038] The sequence of related news leads is simultaneously distributed to the news editing terminal along with the news editing task. When a user clicks on the sequence of related news leads, it is displayed in the form of a timeline.
[0039] Optionally, generating a resizing reminder object also includes the following steps:
[0040] The editorial response efficiency is matched with historically similar clues based on the overall news value;
[0041] If the editing response efficiency of similar historical clues is lower than a preset threshold, add a high-priority editing group to the expanded reminder list; configure a tiered reminder strategy for the expanded reminder objects.
[0042] Optionally, the intelligent analysis method for converged media content clue information of the present invention further includes:
[0043] After the news editing terminal completes the editing task, it collects data on the dissemination effect of the final published news article.
[0044] The keyword category weights in the keyword library are updated based on the dissemination effect data; if the dissemination effect exceeds the benchmark value, the corresponding keyword category weight is increased; if the dissemination effect continues to be lower than the benchmark value, a warning is triggered to downgrade the weight of the corresponding keyword category in the keyword library.
[0045] Optionally, the intelligent analysis method for converged media content clue information of the present invention further includes:
[0046] Sentiment analysis is performed on the content information in news lead files to obtain the sentiment bias of the content information;
[0047] Based on the overall news value of the news lead file, combined with sentiment bias and pre-configured public opinion risk thresholds, a public opinion risk assessment is conducted on the news lead information to obtain the assessment results. The assessment results indicate whether there are potential public opinion risks.
[0048] If the assessment results indicate potential public opinion risks, the news leads will be marked and the corresponding news editing tasks will be distributed to the pre-set public opinion handling task force terminal.
[0049] Configure reminder strategies for the public opinion management task force's terminals.
[0050] The present invention has the following beneficial effects:
[0051] 1. Improved processing efficiency and timeliness. Specifically, the system receives user-submitted news lead files and attribute information via the lead file upload page and stores them in a designated data structure, ensuring data integrity and traceability for subsequent processing. Based on preset lead extraction rules, the system performs content parsing and keyword matching on the news lead files to extract news lead information, which is then automatically updated to the news lead information list, achieving structured and standardized lead information and reducing manual processing burden. Based on the news lead files and their attribute information, the system automatically generates reminder configurations, including the reminder recipient, reminder time, and reminder method, and attaches them to the news lead information, achieving automation and personalization of the notification process and improving the timeliness and accuracy of notifications. Furthermore, based on the overall news value ranking in the news lead information list, the system selects the top N high-value news leads, automatically generates news editing tasks, and distributes them to the corresponding editing terminals, prioritizing high-value leads and improving processing efficiency and timeliness.
[0052] 2. Improved news processing efficiency. Specifically, by integrating heterogeneous news lead files such as text and images through a unified platform, and combining semantic analysis (keyword database matching) and image recognition (content recognition model + value extraction model) technologies, the system can automatically extract structured lead information and present the evolution of events through a timeline visualization engine.
[0053] 3. Improve the efficiency and quality of news gathering and editing. Specifically, the system continuously records the historical behavioral data of each editor. When new news leads enter the system, it analyzes the characteristics of the lead itself (such as keywords, news value, and sentiment), and then compares these characteristics with each editor's user profile. The system will only send highly relevant leads as customized recommendations to the editors most likely to be interested. It not only pushes leads but also tracks the adoption rate of recommended leads and the final dissemination effect of the articles, thus making future recommendations more accurate. Attached Figure Description
[0054] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0055] Figure 1 This is a flowchart of an intelligent analysis method for converged media content clue information according to the present invention. Detailed Implementation
[0056] The invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the drawings and embodiments of the invention are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0057] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0058] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0059] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0060] The names of messages or information exchanged between the various devices of this invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0061] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0062] like Figure 1 The diagram shows a flowchart of an intelligent analysis method for converged media content clue information according to the present invention; specifically, it includes the following steps:
[0063] Step 101: Collect news lead files and news lead attribute information submitted by users through the lead file upload page;
[0064] In some embodiments, the implementing entity of the intelligent analysis method for converged media content clue information of the present invention is a server that has deployed an intelligent analysis system for converged media content clue information.
[0065] In practice, the news tip file upload page is a user interface where users can submit news tip files and related attribute information. Users can be individuals or teams submitting news tips, such as news gatherers from different departments (e.g., project centers). A news tip file refers to the original material submitted by the user, which can be a text file or an image file. The news tip attribute information describes the metadata of the news tip, including but not limited to user information, material category, material size, creation time, status information, news source organization information, and event occurrence time. Based on this, this invention receives information through an interface called the "News Tip File Upload Page," which allows users to upload news tip files to the system and simultaneously fill in the associated metadata (i.e., news tip attribute information).
[0066] Step 102: Extract clues from the news clue file according to the pre-configured clue extraction rules, obtain news clue information, and update the news clue information list.
[0067] In some embodiments, pre-configured lead extraction rules refer to a series of matching and extraction logics for analyzing news lead files, set by the administrator or during the algorithm training process before system deployment or operation. Based on this, structured data is extracted and organized from the news lead files, including basic file information, extracted keywords, event type, occurrence time, location, and other indicators. The extracted key information is used to form news lead information. The newly generated "news lead information" is then added to or modified in the "news lead information list." Here, news lead information refers to structured information obtained after intelligent analysis, including overall news value, matched keyword groups, and news lead attribute information. The news lead information list is a list used to store all news lead information.
[0068] Step 103: Based on the news lead file and news lead attribute information, generate a reminder configuration and add the reminder configuration to the news lead information. The reminder configuration includes the reminder target, reminder time, and reminder method.
[0069] In some embodiments, news lead attribute information includes news source organization information. In this case, based on the news lead file and news lead attribute information, generating a reminder configuration and adding the reminder configuration to the news lead information includes: determining the basic reminder object based on the news source organization information and pre-configured organizational structure information; generating expanded reminder objects based on the overall news value and a pre-configured reminder scope list, with the basic reminder object and expanded reminder objects constituting the reminder object; receiving the reminder time and reminder method configured by the user for the news lead file; and combining the reminder object, reminder time, and reminder method into a reminder configuration and adding the reminder configuration to the news lead information. Here, news source organization information refers to structured information that identifies the organization or unit that provided the news lead, used to determine the source of the lead and support the subsequent determination of the reminder object. The reminder configuration refers to the configuration information used to send prompts to relevant personnel during subsequent editing, review, or publishing processes. It includes the reminder object, reminder time, and reminder method. The reminder object refers to the personnel or department that needs to receive the news lead prompt, which can be represented by a specific account, role identifier, or organizational structure node. The reminder time refers to the time at which the reminder is sent. The reminder method refers to the channel through which the reminder information is delivered, such as client pop-ups, SMS, email, APP push, instant messaging tools, etc.
[0070] Step 104: Based on the top N news leads in the news lead information list according to their overall news value, generate a news editing task and send the news editing task to the corresponding news editing terminal.
[0071] In some embodiments, the overall news value of each news lead is read from the news lead information list and sorted from high to low value. Then, the top N records of the sorted results are extracted. For each news lead in the top N records, elements such as title, event time, location, and keywords are extracted, combined to generate a task description text, appended with the access path or attachment identifier of the news lead file, and the task priority is determined based on the overall news value. The generated task data is encapsulated as structured records and stored in a task management database table. Subsequently, news editors or teams (task recipients) are determined based on news lead attribute information and preset task allocation rules, and the terminal information bound to the task recipient (such as device ID and client account) is read. The task is then sent to the corresponding news editing terminal via push service, which may include notification via SMS, email, etc. After receiving the task, the news editing terminal displays the task in the task list interface and supports viewing details, accepting the order, and providing feedback. Here, a news editing task refers to a task order generated by the system based on the selected news lead information, used to guide news editors in operations such as writing, processing, and reviewing. Task content includes, but is not limited to: task title (which can be generated from news lead titles or keywords); task description (including an overview of the news event, keyword phrases, value score, event time, location, etc.); attachments or material links (original news lead files, images, videos, etc.); completion deadline, priority, and other task management information. A news editing terminal refers to the hardware device or software client used to receive and process news editing tasks. This can be a computer-based news gathering and editing system, a mobile news gathering and editing app, etc. It is linked to the editor's account information and can receive tasks through push services, task list interfaces, etc. Overall news value is a comprehensive scoring indicator used to measure the importance, timeliness, and attention received by news leads.
[0072] These embodiments improve processing efficiency and timeliness. Specifically, the system receives user-submitted news lead files and attribute information via a lead file upload page and stores them in a designated data structure to ensure data integrity and traceability in subsequent processing. Based on preset lead extraction rules, the system performs content parsing and keyword matching on the news lead files to extract news lead information, which is then automatically updated to the news lead information list, achieving structured and standardized lead information and reducing manual processing burden. Based on the news lead files and their attribute information, a reminder configuration is automatically generated, including the reminder recipient, reminder time, and reminder method, and appended to the news lead information, achieving automation and personalization of the notification process and improving the timeliness and accuracy of notifications. Furthermore, based on the overall news value ranking in the news lead information list, the top N high-value news leads are selected, news editing tasks are automatically generated, and these tasks are distributed to the corresponding editing terminals, achieving priority processing of high-value leads and improving processing efficiency and timeliness.
[0073] In some embodiments, to further address the second technical problem described in the background section, namely, "With the popularization of media such as text, images, and short videos, news leads are taking on increasingly diverse forms. Faced with a massive amount of news leads in various forms, how to efficiently extract the most newsworthy news leads has become an urgent problem to be solved," in some embodiments of the present invention, the news lead file is a text file; and
[0074] Extract news leads from news lead files, obtain news lead information, and update the news lead information list, including:
[0075] Step 1: Match the news lead file against a pre-configured keyword database to obtain the matched keyword groups, the keyword category of each keyword, and the number of times each keyword is matched. Different keyword categories are configured with different statistical weights.
[0076] In some embodiments, the value of a single piece of content is quantitatively assessed by matching keywords in a text file with a keyword library and calculating isolated news value by combining keyword categories and statistical weights. The global news value is calculated using the cumulative number of keyword hits over a target time period and combined with the isolated news value to obtain a more comprehensive overall news value, improving the accuracy of the evaluation. Clues with high overall news value are marked to facilitate rapid identification and processing of high-value information, improving content filtering and decision-making efficiency. Specifically, the text file includes news headlines, body text, or other textual descriptions related to the event. The pre-configured keyword library refers to a pre-configured set of words used to identify the core elements of news clues. Each keyword has a unique identifier in the keyword library and is categorized into a keyword category. Based on this, the content of the news clue file is converted into a unified code and segmented. The segmentation results are then compared one by one with the entries in the keyword library. If they match, a hit is determined, and a "hit list" is generated, containing the hit keywords, their categories, and the number of hits. All hit keywords are extracted from the hit list and deduplicated to form keyword groups. Each keyword carries keyword category and hit count information. The system then reads the weight values corresponding to keyword categories from the keyword database. Here, a keyword group refers to the set of all keywords successfully matched in the news lead file. Keyword category is the classification label to which a keyword belongs, used to distinguish the importance of different types of information in news value assessment. For example, keyword categories can be numerical keywords (e.g., "100%)", location keywords (e.g., "a certain city"), organizational structure keywords (e.g., "human resources department"), and event keywords (e.g., "Prasang"). The hit count refers to the number of times a keyword appears in the news lead file.
[0077] Step two: Based on the keyword group, the keyword category of each keyword, the number of hits, and the statistical weight of different keyword categories, generate the isolated news value of the news lead file;
[0078] In some embodiments, based on the keyword group, the keyword category, the number of times the keyword was hit, and the statistical weight corresponding to the category for each keyword are read sequentially. Then, the number of hits for the keyword is multiplied by its statistical weight to obtain the weighted contribution value of the keyword to the news value. The weighted contribution values of all keywords in the keyword group are summed to obtain the total score of the news lead file. Finally, the total score is normalized to obtain the isolated news value score of the news lead file. As an example, suppose a news lead file undergoes keyword matching, yielding the following keyword groups and their attributes: the keyword "earthquake" corresponds to the event keyword category, with a statistical weight of 1.5 and a hit count of 2; the keyword "City A" corresponds to the location keyword category, with a statistical weight of 1.2 and a hit count of 1; the keyword "80%" corresponds to the numerical keyword category, with a statistical weight of 1.6 and a hit count of 3. The calculation process is as follows: First, calculate the weighted contribution value of each keyword: "earthquake": 2 (hit count) multiplied by 1.5 (statistical weight), resulting in 3.0; "City A": 1 multiplied by 1.2, resulting in 1.2; "80%": 3 multiplied by 1.6, resulting in 4.8. Second, sum all keyword contribution values to obtain the total score: 3.0 plus 1.2, plus 4.8, resulting in 9.0. Finally, normalization (assuming the maximum possible total score is 15 and the minimum is 0): dividing 9.0 by 15 gives an isolated news value of 0.6.
[0079] Step 3: Based on the cumulative hit count of each keyword in the keyword library within the target time period, generate the global news value of the news lead file, and based on the isolated news value and the global news value, generate the overall news value of the news lead file; the overall news value, keyword group, and news lead attribute information of each news lead file constitute the corresponding news lead information.
[0080] In some embodiments, the target time period refers to a preset time window, such as the most recent 1 day, 3 days, or 7 days, used to count the frequency of keyword occurrences. Based on this, the cumulative hit count for each keyword in the keyword library within the target time period is calculated. The cumulative hit count refers to the total number of times a keyword is matched in all news lead files within the target time period. Based on the cumulative hit count and the keyword category weight, the hot topic contribution value of each keyword is calculated (similar to the calculation of isolated news value). The hot topic contribution values corresponding to keywords in the news lead files are weighted and summed to obtain the global news value score for that file. The isolated news value and the global news value are weighted and averaged according to a preset weight ratio (e.g., 0.6 and 0.4) to obtain the overall news value. As an example, suppose that within the target time period (e.g., the most recent 7 days), the cumulative hit count for each keyword in the keyword library is as follows: the cumulative hit count for earthquake is 100, the cumulative hit count for City A is 80, and the cumulative hit count for 80% is 50. Based on this, the hot topic contribution value for each keyword is calculated using the following formula:
[0081]
[0082] in, Contribution value to trending topics For statistical weighting, The cumulative hit count is calculated using the above formula. The hot topic contribution value for the keyword "earthquake" is 3.0065; the hot topic contribution value for the keyword "City A" is 2.2902; and the contribution value for 80% of the hot topics is 2.7321. Then, the hot topic contribution values for each keyword are summed to obtain a global news value score of 8.0288. The global news value score is then normalized (assuming the maximum possible value is 10, normalization is 8.0288 divided by 10, resulting in 0.80288). Based on this, the weighting coefficient for isolated news value is set to 0.6, and the weighting coefficient for global news value is 1 minus 0.6, resulting in 0.4. Therefore, the overall news value score is: first, multiply 0.6 by 0.6 to get 0.36; then multiply 0.4 by 0.80288 to get 0.321152; finally, add 0.36 and 0.321152 to obtain an overall news value score of 0.681152. Based on this, the calculated overall news value is integrated with the corresponding keyword groups and news lead attribute information (such as the news source organization, event time, etc.) to form complete news lead information. The overall news value measures the importance of keywords in the news lead file within the overall public opinion and trending environment, reflecting the news content's popularity or attention level relative to the entire news ecosystem. The overall news value is a comprehensive score derived by considering both isolated news value (the importance of the news lead itself) and overall news value (the trending nature of the keywords within the target time period).
[0083] Step four: Mark the news leads that rank in the top N in terms of overall news value in the news lead list.
[0084] In some embodiments, the news leads are sorted from the news lead information list according to their overall news value, and the top N news leads are extracted. A flag is added to the top N news leads in the memory data structure and marked as "high" when displayed on the interface.
[0085] Among them, news tip files are image files; and
[0086] Extract news leads from news lead files, obtain news lead information, and update the news lead information list, including:
[0087] Step 1: Input the image file into the image content recognition model to obtain the content information displayed in the image file;
[0088] Step two: Save the content information as a text file and generate the first news value score corresponding to the content information;
[0089] In some embodiments, content recognition is performed on image files to extract textual information, and a first news value score is calculated based on the textual information, achieving semantic analysis of the image. The image is then directly input into a news value extraction model to calculate a second news value score, achieving a comprehensive value assessment of the image's visual features, scene, objects, and other elements. The first and second news value scores are then integrated, comprehensively considering text and image features to improve the comprehensiveness and accuracy of the image's news value assessment. Specifically, news lead files can also be image files, where image files refer to image data serving as news leads, such as news scene photos, surveillance screenshots, posters, table screenshots, etc. Based on this, the executing entity reads the image file of the news lead; inputs the image pixel data into the image content recognition model; and the model outputs readable content information of the image (e.g., "earthquake disaster area, building collapse, City A"). The image content recognition model can be an algorithm model based on computer vision technology, such as an open-source OCR (Optical Character Recognition) model (e.g., PaddleOCR). Image content recognition models are used to identify describable textual content (OCR), object categories (such as "earthquake scene," "conference room"), and scene information (such as "city street," "indoor") from images. The output is structured or semi-structured "content information" (text description). Based on this, the obtained content information is treated as a text file; the isolated news value is calculated using the same method as for news clue files, i.e., keyword matching, followed by weighted statistics and normalization to obtain a first news value score, for example, 0.62. This first news value score is derived from the text information output by the image content recognition model, and then calculated using the same method as for text-based clues. It reflects the news importance of the information contained in the image itself.
[0090] Step 3: Input the image file into the news value extraction model to obtain the second news value score corresponding to the image file;
[0091] Step four: Combine the first news value score and the second news value score to obtain the overall news value score of the image file.
[0092] In some embodiments, the same image file is directly input into a pre-trained news value extraction model. The model outputs a second news value score, for example, 0.75. The pre-trained news value extraction model can be a deep learning-based image analysis model that comprehensively analyzes the visual features of the image, the identified scenes and objects, and the accompanying textual information to output a value score for the image in a news report. Based on this, the first and second news value scores are weighted and averaged according to a preset weight ratio (e.g., 0.4 and 0.6) to obtain the overall news value score of the image file. For example, first, 0.4 is multiplied by 0.62 to get 0.248, then 0.6 is multiplied by 0.75 to get 0.45, and finally, 0.45 and 0.248 are added together to obtain an overall news value score of 0.698 for the image file.
[0093] The keyword database includes a configured keyword set and an online keyword set. Each set is configured with corresponding keyword category weights. The keyword database is updated through the following steps:
[0094] For each keyword in the keyword library, determine the number of times the keyword was adopted in the news within the target time period, sort the keywords in the keyword library according to the number of news adoptions, and remove the M keywords at the bottom of the list from the keyword library.
[0095] Get the current set of trending words and select M trending words from it to add to the keyword library.
[0096] In some embodiments, by counting the number of times keywords are adopted in news during a target time period, low-frequency keywords are eliminated and hot words are introduced to achieve the dynamic optimization and update of the keyword library. Configure the keyword category weights so that keyword matching reflects the importance of different categories when calculating the value degree, improving the rationality and flexibility of the calculation. By continuously introducing hot keywords, ensure that the keyword library can quickly respond to current trends and hot events, enhancing the timeliness and adaptability of the system. Specifically, configuring the keyword set refers to a keyword set pre-manually set, usually covering news elements that are stable and long-term concerned (such as "earthquake", "typhoon", etc.). The online keyword set refers to a keyword set dynamically obtained from the Internet during operation, usually from real-time monitored news hotspots (such as "Paris Olympics", "Prassang", etc.). The keyword category weight means that each keyword belongs to a certain category (event keyword, location keyword, person keyword, numerical keyword, etc.), and the execution entity configures weights for each category for weighted calculation when calculating the news value degree. On this basis, the keyword library is updated through the following steps: First, read the news content of the news editing tasks generated within the target time period (such as the most recent 7 days) from the news release database. For each keyword in the keyword library, count the number of times it appears in the adopted news (news adoption times). Then, sort all keywords in descending order according to the news adoption times. Remove the last M keywords (the M keywords with the lowest popularity) from the keyword library. On this basis, obtain recent hot words through the Internet platform. Specifically, perform word segmentation on the captured news titles, abstracts, social media hot posts, etc.; count the word frequencies within a certain time window (such as the past 24 hours or the past 7 days); remove stop words (such as "de", "le", "shi") and irrelevant words to obtain the words with the top popularity rankings; a threshold (word frequency, heat index, etc.) can be set to filter out online hot words. And within the set target time period (such as the past 7 days), count the hit times of each configured keyword; regard the words with hit times higher than a certain threshold as internal hot words. Combine the online hot words with the internal hot words of the configured keyword set and remove duplicates. If the same word appears in both sources, its heat value can be combined (such as weighted average). The finally obtained is the "current hot word set", which can be used to update the keyword library. On this basis, select M high-popularity keywords that do not appear in the keyword library from the hot word set and add them to the keyword library. And assign corresponding keyword categories and their category weights to the newly added keywords.
[0097] Among them, the news clue attribute information includes the news source unit information; and
[0098] According to the news clue file and the news clue attribute information, generate a reminder configuration and add the reminder configuration to the news clue information, including;
[0099] Step 1: Determine the basic notification recipients based on the information of the news source organization and the pre-configured organizational structure information;
[0100] Step 2: Based on the overall news value and the pre-configured list of alert scopes, generate expanded alert objects. The basic alert objects and the expanded alert objects together constitute the alert objects.
[0101] Step 3: Receive the reminder time and reminder method configured by the user for the news lead file;
[0102] Step 4: Configure the reminder by specifying the recipient, time, and method, and then add the reminder configuration to the news lead information.
[0103] In some embodiments, basic reminder targets are automatically generated by reading news source unit information and matching organizational structure information, achieving automation and precise positioning of notification targets. Based on the overall news value, expanded personnel are determined from the reminder scope list, enabling dynamic expansion of the notification scope for high-value content and ensuring the breadth and importance of information coverage. Reminder targets are combined with user-defined reminder times and methods to form reminder configurations, achieving personalization and automation of the reminder process and improving response timeliness and operational convenience. Specifically, organizational structure information refers to pre-configured organizational structure data of the organization, including the hierarchical relationships, responsible persons, and contact information of branches, business units, and functional departments. This is the basis for locating basic reminder targets. Based on this, the news source unit information in the news lead attribute information is first read. The department and corresponding personnel of that unit are then matched in the organizational structure information. The matched personnel list is used as the basic reminder targets. Then, the overall news value corresponding to the news lead file is read; based on the overall news value, the scope of reminder personnel to be expanded is determined from the pre-configured reminder scope list, where the pre-configured reminder scope list includes an overall news value range and the expanded reminder personnel range corresponding to each range. The identified personnel to be alerted are added to the alert recipient list, forming the expanded alert recipient group. The basic alert recipient group and the expanded alert recipient group are then merged (duplicated) to form the final alert recipient group. The alert scope list is a pre-defined rule table used to determine the scope of personnel to be notified based on the importance of the news (overall news value). The alert recipient group refers to the set of personnel receiving news tip alerts. The basic alert recipient group consists of relevant responsible persons and departments located within the organizational structure based on the news source unit information. The expanded alert recipient group consists of additional recipients added based on the overall news value and the alert scope list, such as adding superior departments or broadcasting to all employees. As an example, suppose the group's organizational structure information includes the following: News source unit: Branch B; Department of Branch B: News and Publicity Department of Branch B; Relevant personnel in the News and Publicity Department of Branch B: Xiao Li (Publicity Manager), Xiao Zhang (News Specialist); The pre-configured list of alert scopes is as follows: Overall news value range of 0.9 to 1.0, corresponding to the expanded alert scope of the headquarters leadership (President, Executive Vice President) and all personnel of the Group's Publicity Department; Overall news value range of 0.7 to 0.9, corresponding to the expanded alert scope of the Group's Publicity Department; Overall news value range of 0.5 to 0.7, corresponding to the expanded alert scope of the regional news liaison; Overall news value range of 0.0 to 0.5, corresponding to no expanded alert scope.Based on this, the source unit is retrieved: Branch Company B. The organizational structure is matched, and Xiao Li and Xiao Zhang are identified as the initial notification recipients. The overall news value of this news item is found to be 0.82. The notification range list is checked, and it is determined that 0.82 falls within the 0.7 to 0.9 range. Therefore, the corresponding expanded notification recipient range is all personnel in the Group's Publicity Department (assuming they are: Lao Wang, Lao Liu, and Lao Chen). The final notification recipients are: Xiao Li, Xiao Zhang, Lao Wang, Lao Liu, and Lao Chen.
[0104] In some embodiments, users can set the reminder time and method for the news lead in the system interface (e.g., immediate + SMS push, next morning + email). If not manually set, the default reminder strategy is used. Based on this, the reminder recipient, reminder time, and reminder method are encapsulated into a reminder configuration record. This reminder configuration is added to the corresponding news lead information and saved in the news lead information list. When the reminder time arrives, the system will notify the target personnel of the news lead according to the set method. Here, the reminder configuration refers to the reminder rule record stored by the executing entity, consisting of the reminder recipient, reminder time, and reminder method, and is bound to a specific news lead information.
[0105] The news lead attribute information includes the time of the event; at this point, the above method also includes the following steps:
[0106] Step 1: For each news lead in the top N news value rankings, query related news lead groups based on the matched keyword groups.
[0107] Step two: Sort the related news leads in the related news lead information group according to the chronological order of the events to obtain the related news lead information sequence;
[0108] Step 3: The associated news lead information sequence and news editing task are simultaneously sent to the news editing terminal. When the user clicks on the associated news lead information sequence, it is displayed in the form of a timeline.
[0109] In some embodiments, based on keyword groups of high-value news leads, related news leads are queried and a sequence sorted by event occurrence time is generated, enabling timeline tracking of the development of related events. The timeline is distributed synchronously with news editing tasks, allowing editors to quickly obtain the context and development of related events while processing tasks, improving the coherence and completeness of reporting. Related news leads are displayed in timeline form on the terminal, enhancing the visualization and comprehensibility of information and facilitating rapid understanding of event progress. Specifically, news leads are sorted in descending order based on "overall news value" from the news lead information list. The top N leads are selected as key leads for correlation analysis. Keyword groups of key leads are extracted and matched against the news lead database, using keyword Boolean matching (e.g., at least one keyword must be matched). News leads with a hit rate reaching a threshold are then collected into related news lead information groups. The "event occurrence time" field is read from the related news lead information. These are sorted in ascending order by time, forming a time-series event sequence. The sorting results are saved as a transmittable data structure, including: timestamp, news title, summary or content summary, and source information. The target news lead and its corresponding related news sequence are then packaged into a task package and pushed to the news editing terminal. The front end parses the related news sequence into a timeline. Users can click on a specific time point to expand the detailed content or jump to the full text page of the news. The matched keyword group refers to all keywords matched in the keyword database for the news lead content. These keywords are used as "topic tags" to find other potentially related news leads. The related news lead information group refers to a set of other news leads highly related to the target news lead's theme, filtered through keyword group matching. The related news lead information sequence is an ordered list obtained by sorting the related news lead groups by event occurrence time. The news lead database refers to dedicated data used to store, manage, and retrieve information related to news leads. It typically contains the following types of information: the original news lead file, news lead attribute information, keyword matching results, news value information (isolated news value, global news value, overall news value), alert configuration, and relationship information. The event occurrence time refers to the actual time or time interval of the event described in a news lead.
[0110] The process of generating a capacity expansion reminder object also includes the following steps:
[0111] Step 1: Match the editorial response efficiency with historically similar clues based on the overall news value;
[0112] Step 2: If the editing response efficiency of similar historical clues is lower than the preset threshold, add a high-priority editing group to the expanded reminder object from the reminder scope list; configure a hierarchical reminder strategy for the expanded reminder object.
[0113] In some embodiments, the importance of a lead is determined based on its overall news value, combined with historical data to assess the potential risks of processing such leads. When historical records show that content similar to a current high-value lead was processed inefficiently in the past, the lead is deemed at risk of being "delayed." Once a risk is identified, the system expands the alert recipients from the regular editorial group to the high-priority editorial group. To ensure timely receipt and processing of information, the system also configures a tiered alert strategy for these high-priority recipients to maximize the effective reach of leads. Specifically, historically similar leads refer to all news leads that are highly similar to the current news lead in terms of content, theme, or keywords within a past period. The executing entity analyzes newly submitted leads using natural language processing technology and retrieves similar old leads from the database. Editorial response efficiency measures the speed and effectiveness with which the editorial team processes similar leads. For example, the percentage of leads processed within a specified time can be statistically analyzed. A preset threshold is a pre-defined value or standard. In the example of "editorial response efficiency," this threshold might be a time value (e.g., "average response time exceeds 30 minutes") or an efficiency value (e.g., "adoption rate is below 50%)." When the actual response efficiency falls below this threshold, the system determines that there is a problem with the current processing flow. Based on this, it first searches and identifies historically similar leads in the historical database according to the value and content of the lead. Then, it retrieves the editing response efficiency data of these historically similar leads for statistical analysis and comparison. When the matching and comparison results show that the editing response efficiency of historically similar leads is below the preset threshold, a predetermined high-priority editing group is selected from the alert range list. The identifier of this high-priority editing group is added to the alert target list for the current news lead. Once the high-priority editing group is added to the alert target list, the system sets up a special notification mechanism for them according to the rules of the tiered alert strategy. For example, if the overall news value of the lead is at the highest level, the system may directly configure the highest-level alert strategy, that is, simultaneously notify via system messages, emails, and telephones to maximize information reach. A high-priority editing group refers to an editing team in the alert range list that is given higher processing authority and responsibility. When ordinary editing groups cannot effectively handle important leads, the system will escalate the task to this group to ensure that the lead receives timely attention. Tiered alert strategy: a set of alert rules for different levels of importance and urgency. It's not just a single system notification; it can include a combination of various methods, such as: Level 1 alert: an in-system pop-up or message notification. Level 2 alert: simultaneously sending an email or SMS. Level 3 alert: in emergency situations, automatically dialing the high-priority editing group for a voice alert.
[0114] The present invention provides an intelligent analysis method for converged media content clue information, which further includes:
[0115] Step 1: After the news editing terminal completes the editing task, collect data on the dissemination effect of the final published news article;
[0116] Step 2: Update the keyword category weights in the keyword library based on the dissemination effect data; if the dissemination effect exceeds the benchmark value, the corresponding keyword category weight will be increased; if the dissemination effect continues to be lower than the benchmark value, a warning will be triggered to downgrade the weight of the corresponding keyword category in the keyword library.
[0117] In some embodiments, the front-end analysis of leads is integrated with the back-end performance of articles, forming a complete feedback loop. This not only evaluates and recommends leads but also tracks their actual performance after they become news, thus verifying the accuracy of the initial "news value" assessment. By comparing dissemination effect data with preset benchmarks, it determines which types of leads are truly popular with users. If an article corresponding to a certain keyword has a good dissemination effect, the system will increase its weight in value assessment, giving it higher priority in future lead selection. This ensures that the keyword library is no longer a static, unchanging list. When a hot topic or news type gradually loses its appeal, its dissemination effect will continue to decline, and the system will trigger a downgrade warning, reminding managers to reassess its weight. This ensures that the system can flexibly adapt to constantly changing public interests and market demands, avoiding missing truly valuable news due to reliance on outdated standards. Specifically, the final published press release refers to content that has been edited, reviewed, and modified, and finally publicly released on converged media platforms (such as news websites, apps, social media accounts, etc.). Dissemination effect data: measures the response and interaction among the audience after the press release is published. It includes, but is not limited to: Views / Browses: the number of times the article is opened and read. Shares / Forwards: the number of times the article is shared to social media platforms by users. Comments / Interactions: the number of comments or likes from users on the article. Duration: the average time users spend on the article page. The benchmark is a pre-set performance standard. In the example of dissemination effect, it can be a fixed number (such as "more than 100,000 views") or a dynamic average (such as "dissemination effect is higher than the average level of similar articles"). Based on this, after the press release is published, the system will automatically retrieve relevant dissemination effect data from the publishing platform through a preset API (Application Programming Interface) and store it in the effect database. The system will periodically or in real time compare the dissemination effect data with the benchmark. If the dissemination effect of articles corresponding to a certain keyword category (such as "Human Resources Department") generally exceeds the benchmark, the system will trigger an algorithm to increase the weight value of that category. If the effect continues to be lower than the benchmark, a weight downgrade warning will be triggered, and the weight of that category will be marked or reduced in the database. A weight downgrade warning is an alert or notification issued when the dissemination effect of a certain keyword category continues to fall below the benchmark value.
[0118] The present invention provides an intelligent analysis method for converged media content clue information, which further includes:
[0119] Step 1: Perform sentiment analysis on the content information in the news lead file to obtain the sentiment tendency of the content information;
[0120] Step 2: Based on the overall news value of the news lead file, combined with sentiment bias and pre-configured public opinion risk thresholds, conduct a public opinion risk assessment on the news lead information to obtain the assessment results. The assessment results indicate whether there are potential public opinion risks.
[0121] Step 3: If the assessment results indicate potential public opinion risks, the news leads will be marked and the corresponding news editing tasks will be distributed to the pre-set public opinion handling task force terminal.
[0122] Step 4: Configure reminder strategies for the public opinion management task force's terminals.
[0123] In some embodiments, the approach goes beyond simply assessing the value of the news itself; it delves deeper into the potential emotional undercurrents behind the leads through sentiment analysis. This method moves the risk assessment process upstream, determining from the source of the lead whether it might trigger negative public opinion, rather than waiting until after the news is published to take remedial action. Once the system determines, through public opinion risk assessment, that a high-value lead simultaneously possesses negative sentiment or controversy, it immediately activates an escalation mechanism for task distribution. This means the lead is not distributed to ordinary editors but is directly routed to a specialized public opinion handling team with professional experience, ensuring the task is handled by the most suitable personnel. To ensure that urgent and sensitive leads are not overlooked, the system configures an alert strategy for the specialized team. This strategy, through multi-channel and strong alerts, ensures that relevant personnel receive and process tasks immediately, thereby minimizing the risk of public opinion spiraling out of control due to delayed response. Specifically, the text content of the news lead file (such as the title, body text, and text identified from images) is used as input. A sentiment analysis model is invoked, which is a deep learning model based on a recurrent neural network. This model is trained on a large corpus and can understand the emotional tone and context of words. The model outputs one or more sentiment bias values or labels. The overall news value and sentiment bias are compared with a preset public opinion risk threshold. For example, if a lead's overall news value exceeds 0.8 and its sentiment bias is strongly negative (e.g., -0.8), the lead is considered to have potential public opinion risk. When the public opinion risk assessment result is "potential public opinion risk exists," a preset rule is triggered. This rule removes the task corresponding to this news lead from the regular editing task list and automatically assigns it to the task queue of the public opinion handling task force terminal. Once the task is assigned to the task force, the system immediately initiates the corresponding notification process according to a preset reminder strategy. Sentiment analysis: This is a natural language processing technique used to identify and extract the emotions expressed in text. It categorizes content information as positive, negative, or neutral, or gives a specific sentiment score. For example, "A touching rescue event occurred in a certain place" would be judged as positive sentiment, while "Consumers expressed strong dissatisfaction with a certain product" would be judged as negative sentiment. Sentiment bias: This is the output of sentiment analysis, representing the emotional color of the content information. It is typically a numerical value (such as a score ranging from -1 to +1). The public opinion risk threshold is a pre-set parameter used to define which sentiment trends or combinations of sentiment scores might trigger public opinion risk. For example, a risk warning could be triggered "when the news value is high and the sentiment trend is strongly negative." This threshold can be dynamically adjusted based on historical data and the experience of public opinion experts. The public opinion handling task force terminal refers to an editorial team specifically responsible for handling sensitive, high-risk, or potentially negative news leads.Their terminal systems are configured with specific functions and permissions to handle these tasks more efficiently. Alert strategy: A customized notification scheme for high-risk clues. It is typically more urgent and prominent than ordinary task alerts, for example, sending alerts through multiple channels such as system pop-ups, SMS, email, and instant messaging tools. The alert message highlights phrases like "Public Opinion Risk" or "Urgent Handling." A continuous alert mechanism is configured to resend notifications if the issue is not addressed within a certain timeframe.
[0124] These embodiments improve news processing efficiency. Specifically, by integrating heterogeneous news lead files such as text and images through a unified platform, and combining semantic analysis (keyword database matching) and image recognition (content recognition model + value extraction model) technologies, the automatic extraction of structured lead information is achieved, and the evolution of events is presented through a timeline visualization engine.
[0125] In some embodiments, to further address the third technical problem described in the background section, namely, "different editors specialize in different fields, but traditional systems cannot make personalized recommendations based on editors' expertise and preferences, resulting in low editors' work enthusiasm and adoption rate, thereby leading to low efficiency in news gathering and editing and poor content quality," some embodiments of the present invention further include:
[0126] Step 1: Obtain historical editing behavior data; Based on the historical editing behavior data of the news editing terminal, establish a user profile for the editor;
[0127] Step two: Using user profiles, combined with data such as keyword groups, overall news value, and sentiment bias of news leads, match news leads.
[0128] Step 3: The top K news leads with the highest matching degree are sent as customized lead recommendations to the corresponding news editing terminals;
[0129] Step four: Dynamically adjust the user profile based on the adoption rate of customized leads and the corresponding dissemination effect data.
[0130] In some embodiments, the system records every action of each editor (such as clicking, viewing, accepting, and abandoning). By analyzing these action logs and extracting relevant data from the database, comprehensive historical editing behavior data is constructed. Machine learning algorithms (such as clustering and collaborative filtering) are used to analyze this historical editing behavior data. The algorithms identify the editor's interests, preferences, and style, and structure this information to form the editor's user profile. For a news lead, features such as keyword phrases, news value, and sentiment are extracted. These lead features are then compared with the target editor's user profile features to calculate a similarity score. After the matching score is calculated, the system selects the top K leads with the highest scores. Then, through a preset interface, this customized lead list is pushed to the corresponding editor's news editing terminal, for example, displaying a "Recommended for You" section on the backend homepage. The system continuously tracks the acceptance rate and final dissemination effect data of customized leads. If an editor accepts a recommended lead and the lead's dissemination effect is good, the system will increase the weight of features related to that lead in the editor's user profile based on this positive feedback. Conversely, if the adoption rate of recommended leads is low, or the dissemination effect after adoption is poor, the system will reduce the weight of relevant features. Historical editorial behavior data refers to all quantifiable behavioral records of editors in their past work, such as: Adopted lead categories: Does the editor most frequently handle social news, financial news, or breaking news? Popular topics: Which keywords or events are the editor particularly interested in? Editor style preferences: For example, do they prefer in-depth reporting, breaking news, or special feature planning? Adoption time: During which time periods are editors typically most active? An editor's user profile is a structured data model used to comprehensively describe an editor's interests, preferences, and behavioral habits. It is not a simple label, but a multi-dimensional, dynamically adjustable "digital profile," the foundation for personalized recommendations. Customized lead recommendation refers to filtering leads from a massive amount of news leads based on a specific editor's user profile, selecting those that best match their interests and work habits, and presenting them in the form of a personalized list or customized push notification.
[0131] In these embodiments, the efficiency and quality of news gathering and editing are improved. Specifically, the system continuously records the historical behavioral data of each editor. When a new news lead enters the system, it analyzes the characteristics of the lead itself (such as keywords, news value, and sentiment), and then compares these characteristics with each editor's user profile. The system only sends leads with high matching scores as customized recommendations to the editors most likely to be interested. It not only pushes leads but also tracks the adoption rate of recommended leads and the dissemination effect of the final articles, thus making future recommendations more accurate.
[0132] The above description is merely a selection of preferred embodiments of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this invention.
Claims
1. A media fusion content clue information intelligent analysis method, characterized in that, The method comprises the following steps: collecting news clue files and news clue attribute information; extracting clues from the news clue files to obtain news clue information and update the news clue information to a news clue information list; matching keywords in the text file in a keyword library, multiplying the hit number of the keyword by the statistical weight to obtain the weighted contribution value of the keyword to the news value degree, accumulating the weighted contribution values of all keywords in the keyword group to obtain the isolated news value degree score, calculating the global news value degree using the cumulative hit number of the keywords in the target time period, combining the global news value degree with the isolated news value degree to obtain the overall news value degree, extracting the text information in the picture file to calculate the first news value degree, inputting the picture into a news value extraction model to calculate the second news value degree, and fusing the first news value degree and the second news value degree to obtain the overall news value degree of the picture file; generating a reminder configuration according to the news clue files and the news clue attribute information and adding the reminder configuration to the news clue information, including: determining a basic reminder object according to news source unit information and organizational structure information, matching the overall news value degree with the editing response efficiency of historical similar clues, adding a high-priority editing group to the expanded reminder object from a reminder range list if the editing response efficiency is lower than a preset threshold, and configuring a hierarchical reminder strategy; generating a news editing task according to news clue information in the news clue information list with an overall news value degree in the top N positions and sending the news editing task to a news editing terminal; collecting propagation effect data of a final published manuscript, updating the keyword category weight in the keyword library according to the propagation effect data, counting the number of times a keyword is adopted in the target time period, eliminating low-frequency keywords and introducing hot keywords to update the keyword library; for a news clue, extracting a keyword group, a news value degree, and an emotional tendency feature, comparing the features with user portrait features of a target editor to calculate a similarity score, selecting the top K clues with the highest scores as customized clues and pushing the customized clues to the news editing terminal of the corresponding editor, tracking the adoption rate and propagation effect data of the customized clues, and adjusting the user portrait, including: if a recommended clue is adopted, comparing the propagation effect data with a preset baseline value to determine the type of clue that is popular with users and enhancing the weight of the feature in the editor's user portrait; if the adoption rate of the recommended clue is low or the propagation effect after adoption is poor, the weight of the feature in the editor's user portrait is reduced. 2.The method of claim 1, wherein, The news clue file is a text file. The method further comprises the following steps: matching the news clue file in a pre-configured keyword library to obtain a hit keyword group, the keyword category of each keyword, and the hit number, different keyword categories being configured with different statistical weights; generating the isolated news value degree of the news clue file according to the keyword group, the keyword category of each keyword, the hit number, and the statistical weights of different keyword categories. According to the cumulative hit number of each keyword in the keyword library within a target time period, a global news value degree of the news clue file is generated, and according to the isolated news value degree and the global news value degree, an overall news value degree of the news clue file is generated; the overall news value degree, the keyword group and the news clue attribute information of each news clue file form corresponding news clue information; News clue information whose overall news value degree is ranked in the top N is marked in the news clue information list. 3.The method of claim 2, wherein, The news clue attribute information includes news source unit information; and The generating of the reminding configuration and the adding of the reminding configuration to the news clue information according to the news clue file and the news clue attribute information includes; According to the news source unit information and pre-configured organizational structure information, a basic reminding object is determined; According to the overall news value degree and a pre-configured reminding range list, an expanded reminding object is generated, and the basic reminding object and the expanded reminding object form a reminding object; A reminding time and a reminding manner configured by a user for the news clue file are received; The reminding object, the reminding time and the reminding manner form a reminding configuration, and the reminding configuration is added to the news clue information. 4.The method of claim 3, wherein, The news clue file is a picture file; And The clue extraction on the news clue file, the obtaining of news clue information and the updating of the news clue information to a news clue information list include: The picture file is input into a picture content recognition model to obtain content information displayed by the picture file; The content information is taken as a text file, and a first news value degree corresponding to the content information is generated; The picture file is input into a news value extraction model to obtain a second news value degree corresponding to the picture file; The first news value degree and the second news value degree are fused to obtain an overall news value degree of the picture file. 5.The method of claim 2, wherein, The keyword library includes a configuration keyword set and an online keyword set, and the configuration keyword set and the online keyword set are respectively configured with corresponding keyword category weights, and the keyword library is updated through the following steps: For each keyword in the keyword library, the number of news adoptions of the keyword within a target time period is determined, and each keyword in the keyword library is sorted according to the number of news adoptions, and M keywords sorted at the end are removed from the keyword library; A current hot word set is obtained, and M hot words are selected from the current hot word set and added to the keyword library. 6.The method of claim 5, wherein, The news clue attribute information includes event occurrence time; And the method further includes: For each news clue information in the news clue information whose overall news value degree is ranked in the top N, associated news clue information groups are queried according to the hit keyword group; Each associated news clue information in the associated news clue information groups is sorted according to the chronological order of the event occurrence time to obtain an associated news clue information sequence; The associated news clue information sequence and a news editing task are synchronously issued to a news editing terminal, and when a user clicks the associated news clue information sequence, the associated news clue information sequence is displayed in the form of a timeline. 7.The method of claim 6, wherein, The generating the expansion reminding object further comprises the following steps: According to the overall news value degree, matching the editing response efficiency of the historical similar clues; If the editing response efficiency of the historical similar clues is lower than the preset threshold, adding a high-priority editing group from the reminding range list to the expansion reminding object; configuring a hierarchical reminding strategy for the expansion reminding object. 8.The method of claim 7, wherein, Further comprising: After the news editor terminal completes the editing task, collecting the propagation effect data of the finally published news article; According to the propagation effect data, updating the keyword category weight in the keyword library; If the propagation effect exceeds the benchmark value, the corresponding keyword category weight is improved; if the propagation effect continues to be lower than the benchmark value, the weight degradation warning of the corresponding keyword category in the keyword library is triggered. 9.The method of claim 8, wherein, Further comprising: Performing sentiment analysis on the content information in the news clue file to obtain the sentiment tendency of the content information; According to the overall news value degree of the news clue file, combining the sentiment tendency and the pre-configured public opinion risk threshold, performing public opinion risk assessment on the news clue information to obtain an assessment result, wherein the assessment result represents whether there is a potential public opinion risk; If the assessment result represents that there is a potential public opinion risk, the news clue information is marked, and the corresponding news editing task is distributed to the preset public opinion processing special team terminal; Configuring a reminding strategy for the public opinion processing special team terminal.
Citation Information
Patent Citations
News task planning and distribution method and system based on big data and storage medium
CN114969196A
News production system and method based on big data and artificial intelligence
CN120218019A