Target information processing method and device, electronic equipment and computer program product

By constructing a historical semantic baseline pool and a multi-dimensional scoring model, the problems of information fragmentation and noise interference were solved, enabling a deep understanding of unstructured content and accurate identification of high-value information. This improved the continuity and accuracy of information processing and provided information services that better meet user needs.

CN121997933APending Publication Date: 2026-05-08ZHENGZHOU APUS DIGITAL CLOUD INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHENGZHOU APUS DIGITAL CLOUD INFORMATION TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing information aggregation and news processing systems suffer from problems such as information fragmentation, noise interference, and insufficient adaptability. They are unable to effectively capture unstructured pages or dynamically loaded content, and their evaluation mechanisms are simplistic, making it difficult to distinguish between promotional content and substantive technological progress.

Method used

By constructing a historical semantic baseline pool, using a multi-dimensional scoring model to perform incremental value judgment and structured processing of target information, employing multiple intelligent models to collaboratively perform semantic parsing and scoring, generating a comprehensive weight value, and adjusting model parameters based on user feedback data.

Benefits of technology

It enables continuous information processing, improves the depth and quality of information collection, accurately distinguishes high-value content from low-value noise, enhances the purity and accuracy of information filtering, and provides information services that better meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997933A_ABST
    Figure CN121997933A_ABST
Patent Text Reader

Abstract

The invention provides a target information processing method and device, electronic equipment and a computer program product, belongs to the field of data processing, and aims to solve the problems of information fragmentation, noise interference, insufficient adaptive capacity and the like in the existing method. The method comprises the following steps: acquiring historical information and current target information, obtaining a historical semantic baseline pool according to the historical information, and determining an increment value of the current target information based on the historical semantic baseline pool; in response to the increment value of the current target information, processing the current target information to obtain structured information of the current target information; based on the structured information, determining a comprehensive weight value of the current target information through a multi-dimensional scoring model; and generating target content of the current target information according to the comprehensive weight value, and adjusting parameters of the multi-dimensional scoring model according to feedback data of the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more particularly to a method, apparatus, electronic device, and computer program product for processing target information. Background Technology

[0002] Existing information aggregation and news processing systems suffer from the following limitations. First, traditional methods typically process information in isolation on a daily basis, lacking the ability to model the continuity and development of events over time. This results in fragmented information, making it difficult for users to grasp the evolutionary context and incremental value of the information. Second, information collection largely relies on web crawlers based on fixed rules, lacking adaptability and generalization capabilities, and failing to effectively crawl unstructured pages or dynamically loaded content. Finally, information evaluation and filtering mechanisms often employ a single dimension, failing to distinguish between promotional content and substantive technological advancements, and struggling to differentiate weighting based on different sources or entities, leading to information overload and the suppression of high-quality information. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, and computer program product for processing target information, which can solve the problems of information fragmentation, noise interference, and insufficient adaptive capability in existing methods.

[0004] In a first aspect, embodiments of this application provide a method for processing target information, the method comprising the following steps: Acquire historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool; In response to the fact that the current target information has incremental value, the current target information is processed to obtain structured information of the current target information; Based on the structured information, a multidimensional scoring model is used to determine the comprehensive weight value of the current target information; The target content of the current target information is generated based on the comprehensive weight value, and the parameters of the multidimensional scoring model are adjusted based on the feedback data of the target content.

[0005] Secondly, embodiments of this application provide a target information processing apparatus, which includes the following: The acquisition module is used to acquire historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool. The processing module is configured to process the current target information in response to the fact that the current target information has incremental value, and obtain structured information of the current target information; The scoring module is used to determine the comprehensive weight value of the current target information based on the structured information using a multi-dimensional scoring model; and, The generation module is used to generate target content of the current target information based on the comprehensive weight value, and adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

[0006] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the target information processing method as described in the first aspect.

[0007] Fourthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, the program instructions being executed by a computer to implement the steps of the target information processing method as described in the first aspect.

[0008] This application's embodiments determine the incremental value of current target information by constructing and utilizing a historical semantic baseline pool. This enables the identification of the information's evolution over time, effectively filtering duplicate or redundant content and ensuring the true continuity of the output information, thus solving the problems of isolated and fragmented information in traditional information aggregation. Multiple intelligent models collaborate to perform semantic parsing and structuring processing on identified high-value information, replacing traditional fixed-template-based crawling methods. This allows for deep understanding and information extraction of unstructured, dynamically loaded content, helping to expand the scope and quality of effective information collection. A multi-dimensional scoring model quantitatively evaluates target information from multiple perspectives, generating a comprehensive weight value. This overcomes the limitations of relying solely on single indicators such as popularity, accurately distinguishing between substantive content and low-value noise, significantly improving the purity and accuracy of information filtering. Target content is generated based on the comprehensive weight value, and the parameters of the multi-dimensional scoring model are dynamically adjusted based on user behavior feedback data. This allows for continuous learning of user preferences and adaptive optimization of the evaluation strategy, providing information services that better meet user needs and improving user experience. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1This is a flowchart illustrating a method for processing target information provided in an embodiment of this application; Figure 2 This is a flowchart illustrating another method for processing target information provided in an embodiment of this application; Figure 3 This is a flowchart illustrating another method for processing target information provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a target information processing device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0012] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0013] Against the backdrop of the rapid development of artificial intelligence (AI) technology, existing information aggregation solutions suffer from serious problems of "technology silos" and "noise overload." These problems greatly affect the effectiveness of information aggregation and the user experience of obtaining effective information. Specifically, they can be manifested in aspects such as temporal discontinuity, blind spots in data collection, and a single evaluation dimension.

[0014] Temporal discontinuity refers to the tendency of traditional systems to treat daily information as isolated events, lacking effective processing of the temporal relationships between information. For example, when OpenAI releases an update to Sora, if the system lacks prior release context (inheritance mechanism), it cannot accurately determine the incremental value of that information. This makes it difficult for users to fully understand the development context and importance of information when acquiring it.

[0015] The existence of blind spots in crawling stems from the fact that traditional web crawlers rely on fixed templates, a method with significant limitations that struggles to adapt to diverse online environments. This is particularly true for unstructured, dynamically loaded technical communities (such as Discord or personal technical blogs), where traditional crawlers often fail to effectively extract information, resulting in the loss of a large amount of valuable data.

[0016] The evaluation criteria are too simplistic because existing methods rely solely on keyword popularity ranking. This single metric fails to comprehensively and accurately assess the value of information. It cannot distinguish between "PR press releases" and "deep technological breakthroughs," nor can it dynamically weight information for specific high-value entities (such as Google and Anthropic). This leads to truly valuable information being overlooked, while low-quality information may receive excessive attention.

[0017] Therefore, there is an urgent need for a method and device for processing target information, especially for processing information such as news, to solve the problems of information fragmentation, noise interference, and insufficient adaptive capability in existing methods.

[0018] The target information processing method and apparatus provided in this application adopt semantic inheritance and multi-dimensional intelligent evaluation, which can be used for AI news automation processing. It aims to realize the automation of news processing and provide users with more efficient and accurate news services. It integrates key technologies such as semantic inheritance and multi-dimensional intelligent evaluation, and has unique innovation and practicality in the field of news processing.

[0019] The following is in conjunction with the appendix Figures 1 to 5 The present application provides a detailed description of a target information processing method, apparatus, electronic device, and computer program product through specific embodiments and application scenarios.

[0020] Figure 1 This application illustrates an embodiment of a method for processing target information. This method can be executed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be executed by software or hardware installed on the server and / or terminal device, and includes the following steps: Step 110: Obtain historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool.

[0021] The historical semantic baseline pool addresses issues like temporal discontinuity and lack of context in traditional information processing. Based on each processing cycle (e.g., daily), high-value information (e.g., information with a comprehensive weight value exceeding a set threshold from the previous day) that has undergone preliminary screening and verification is selected as candidates. A pre-trained natural language processing model transforms the core content of each piece of information (e.g., title, summary, key paragraphs, etc.) into a high-dimensional semantic feature vector to mathematically represent the deep semantic information of the information. These semantic vector features are associated with the information's metadata and stored in a dedicated vector database or index structure, thus forming the historical semantic baseline pool. This pool can be updated dynamically according to a strategy, such as retaining high-value vectors from the most recent cycles, ensuring that the historical semantic baseline pool reflects recent focus while also possessing historical depth.

[0022] The target information can refer to valuable, structured text streams, such as news, technical documents, academic papers, forum discussions, and market reports. In terms of timeliness, it can be either highly time-sensitive information (e.g., breaking news) or less time-sensitive but high-value information (e.g., classic papers, historical analyses). Its source can be any publicly available or authorized text source, including media, institutions, communities, individuals, and databases.

[0023] The incremental value is achieved by using the same model as the historical semantic baseline pool to transform the current target information into a semantic feature vector. The similarity between this semantic feature vector and all or some representative vectors in the historical semantic baseline pool is calculated. Cosine similarity can be used for calculation, with a value range of [-1, 1]. The closer the value is to 1, the more similar the semantics are.

[0024] The determination is made based on the calculated highest or average similarity, combined with a preset threshold. A specific example is as follows: If the similarity is higher than the preset high threshold (e.g., 0.9), it indicates that the current information is highly semantically repetitive with historical known information and there is no significant new content. It can be marked as redundant information, downweighted or filtered directly to avoid information overload.

[0025] If the similarity is lower than the preset low threshold (e.g., 0.5), it indicates that the current information has a weak correlation with the historical baseline and may involve entirely new topics, technologies, or events. It can be marked as high incremental value information, given a higher basic inheritance score, and then enter the subsequent in-depth processing process.

[0026] If the similarity score is between high and low thresholds (e.g., 0.5 to 0.9), it indicates that the current information is related to the historical context but is not a simple repetition. In this case, further analysis using entity recognition techniques can be conducted to determine if the vector contains new technical parameters, performance indicators, partners, or other key entities. Entity recognition technology can be Named Entity Recognition (NER), a key technology in natural language processing designed to identify and classify entities with specific meanings from text, such as names, locations, organizations, and time expressions. If new and important entities are included, it is considered "evolutionary information," signifying the development and updating of existing events or technologies, and is assigned a higher inheritance score. If no substantial new entities are found, it may be classified as a restatement or commentary on known information, and given a lower weight.

[0027] Step 120: In response to the fact that the current target information has incremental value, the current target information is processed to obtain the structured information of the current target information.

[0028] The processing of the current target information can be achieved through the in-depth mining and standardized transformation of information using multiple dedicated models working in tandem. These dedicated models can include a link discovery model and a content parsing model. The link discovery model starts with the current target information and is responsible for broad exploration, analyzing page structure, semantic relationships, and external references. It proactively discovers and filters extended links related to the core content (e.g., technical documents, community discussions, official announcements), constructing a network of related information to be crawled. The content parsing model is responsible for deep understanding. It receives the links output by the link discovery model and performs semantic parsing on the acquired raw content (e.g., unstructured HTML, dynamically loaded content) to understand the content's main idea, identify key entities (e.g., companies, technologies, parameters), and determine the content type, rather than simply extracting text.

[0029] The collected raw text content undergoes standardized preprocessing and structured transformation. Cleaning and deduplication remove noisy content such as advertisements and navigation bars, and merge highly repetitive information. Entity recognition and relation extraction utilize technologies such as Named Entity Recognition (NER) to automatically extract key fields, such as company, technology stack, performance metrics, people, and time. Generating structured objects involves encapsulating the cleaned text and extracted entity information into structured data objects according to predefined templates.

[0030] Step 130: Based on the structured information, determine the comprehensive weight value of the current target information through a multidimensional scoring model.

[0031] The multidimensional scoring model calculates a comprehensive weight value for the current target information based on multiple dimensions. These dimensions can include content and source dimensions. The content dimension can further include technological innovation, market influence, and information density. Technological innovation is used to assess the innovativeness of the information's content, while market influence is used to assess the information's potential impact on industry structure and market dynamics. Information density is used to assess the abundance and effectiveness of the information itself. The source dimension can further include source compensation weights, serving as a bias term for the basic confidence level and importance.

[0032] The scores from the above multiple dimensions are synthesized using a preset quantification formula to obtain the final comprehensive weight value.

[0033] Step 140: Generate the target content of the current target information based on the comprehensive weight value, and adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

[0034] The target content can be intelligently generated based on the comprehensive weight value. Multiple current target information processed in the same period are sorted in descending order according to the comprehensive weight value. Information with high relevance and belonging to the same event context is grouped according to the fields in the structured information. Then, the sorted and grouped structured information and its historical background (e.g., the relevant vector information summary extracted from the historical semantic benchmark pool) are input into the fine-tuned Large Language Model (LLM). The model can generate the final target content, such as a structured summary, according to the preset template and style.

[0035] The feedback data can be obtained by collecting various user interaction information, such as reading depth metrics, interaction behavior metrics, and subsequent exploration behavior. The feedback data is periodically aggregated and analyzed to adjust the parameters of the multidimensional scoring model in step 130, such as adjusting the weight coefficients of the content dimension or source dimension and revising the scoring rules.

[0036] In this embodiment, by constructing and utilizing a historical semantic baseline pool to determine the incremental value of current target information, the evolution of information over time can be identified. This effectively filters out repetitive or redundant content, ensuring the true continuity of the output information and solving the problems of isolated and fragmented information in traditional information aggregation. Multiple intelligent models collaborate to perform semantic parsing and structuring processing on identified high-value information, replacing the traditional fixed-template-based crawling method. This enables deep understanding and information extraction of unstructured, dynamically loaded content, helping to expand the scope and quality of effective information collection. A multi-dimensional scoring model is used to quantitatively evaluate target information from multiple perspectives, generating a comprehensive weight value. This overcomes the limitations of relying solely on single indicators such as popularity, accurately distinguishing between substantive content and low-value noise, significantly improving the purity and accuracy of information filtering. Target content is generated based on the comprehensive weight value, and the parameters of the multi-dimensional scoring model are dynamically adjusted based on user behavior feedback data. This allows for continuous learning of user preferences and adaptive optimization of the evaluation strategy, providing information services that better meet user needs and improving user experience.

[0037] Figure 2 This diagram illustrates a flowchart of another method for processing target information provided by an embodiment of this application. This method can be executed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be executed by software or hardware installed on the server and / or terminal device, and includes the following steps: Step 210: Based on step 110 of the above embodiment, historical information and current target information are obtained, a historical semantic baseline pool is obtained based on the historical information, and the incremental value of the current target information is determined based on the historical semantic baseline pool. The method of this embodiment may further include the following specific steps: Extract semantic feature vectors from each historical information to form a historical semantic baseline pool; extract the semantic feature vector of the current target information, calculate the similarity between the semantic feature vector of the current target information and the vectors in the historical semantic baseline pool, and determine the incremental value of the current target information relative to the historical information based on the similarity.

[0038] This embodiment uses news as the target information for specific example illustration.

[0039] This step involves constructing a temporal semantic inheritance and baseline pool. It's not simply about data storage, but rather establishing a mathematical framework that quantifies the evolution of information. This framework allows for more scientific and accurate analysis and processing of information, providing a solid foundation for subsequent news processing.

[0040] Specifically, each day, the selected high-weight news items are transformed into an N-dimensional feature vector V_hist and stored in a vector database to form a historical semantic baseline pool. The purpose of this is to establish a stable information reference standard, facilitating the subsequent evaluation of the value of new information.

[0041] When new information flows in, calculate the cosine similarity between the new information vector V_new and the vector V_hist in the baseline pool.

[0042] Current target information vector V_new: A newly received news article is transformed into a high-dimensional (e.g., 768-dimensional) semantic feature vector using a model (e.g., BERT, SentenceTransformer, etc.). Historical baseline pool vector V_hist: A set of semantic feature vectors that have been stored for historical high-value information.

[0043] Calculate the cosine similarity between V_new and each (or the most relevant) V_hist in the baseline pool. Take the highest similarity value or the average similarity value with the relevant topic clusters as the criterion.

[0044] If the similarity is extremely high but the core parameters remain unchanged, it is considered redundant; if the similarity is moderate and contains new entities, it is marked as "evolutionary information" and assigned a higher inheritance score. In this way, valuable new information can be effectively filtered out, avoiding interference from redundant information.

[0045] In addition, structured data can be synchronously retrieved from pre-set authoritative sites (such as arXiv, TechCrunch, etc.) as daily verification points for the authenticity of information. This helps ensure the authenticity and reliability of information, providing users with high-quality news content.

[0046] Step 220: In response to the fact that the current target information has incremental value, the current target information is processed to obtain the structured information of the current target information. Step 230: Based on the structured information, determine the comprehensive weight value of the current target information through a multidimensional scoring model.

[0047] Step 240: Generate the target content of the current target information based on the comprehensive weight value, and adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

[0048] Steps 220-240 can be found above. Figure 1 The specific descriptions of steps 120-140 in the illustrated embodiment are provided, and the same technical effects can be achieved. To avoid repetition, they will not be repeated here.

[0049] In this embodiment, by quantifying historical high-value information into feature vectors and constructing a baseline pool, discrete and isolated historical information is transformed into structured and computable data. When new information flows in, semantic similarity comparison with the baseline pool automatically identifies the connections and evolution between information, determining whether it is redundant, repetitive, an important evolution, or a completely new theme. This solves the problem of temporal discontinuity caused by treating daily information as isolated events in traditional information processing, giving the output information stream a continuous historical background and development trajectory, thus enhancing the intuitive understanding of events. By setting a similarity threshold and combining it with entity change analysis, high-value incremental information and low-value noise information can be efficiently and automatically distinguished, improving the automation and accuracy of information filtering, providing high-quality information for subsequent deep processing steps, and avoiding resource waste on invalid information.

[0050] In yet another exemplary embodiment, the method further includes the following steps: Semantic feature vectors entering the historical semantic baseline pool are managed according to their timeliness classification. The semantic feature vectors that are closer to the current time are stored in the pool with finer granularity and have higher comparison priority. In response to the inflow of new current target information, similarity comparison is preferentially performed with the nearest historical vector with the finest storage granularity according to the timeliness classification. When determining the incremental value of the current target information based on the similarity comparison, a time-based decay factor is introduced to correct the similarity calculation results. The longer the time interval between the historical vector and the current target information, the more significant the decay of its corresponding similarity weight.

[0051] The timeliness-based hierarchical management involves using hierarchical time windows and semantic centroid compression for the historical semantic baseline pool, and employing different processing strategies for data from different time periods. Specifically, the historical semantic baseline pool is divided as follows: Hot Window: stores the full semantic feature vectors within the last N days (e.g., 7 days), and uses HNSW (Hierarchical Navigable Small World) index for high-frequency comparison; Warm Window: stores information within N to M days (e.g., 30 days), and uses K-means clustering algorithm to compress similar information into "semantic centroids," comparing only the centroids to reduce computation; ColdArchive: performs dimensionality reduction processing on information from very long periods.

[0052] In addition, a "semantic decay coefficient" is introduced. When calculating incremental value, the weight of historical information decreases exponentially over time. The specific formula can be expressed as: S_final = S_similarity

[0053] Where S_final is the final similarity value after attenuation correction, and S_similarity is the original similarity value. For time difference.

[0054] This mechanism ensures that it can identify long-term technological trends without being interfered with by outdated and redundant information in real time.

[0055] In yet another exemplary embodiment, the method for determining the incremental value of the current target information relative to historical information based on the similarity may further include the following specific steps: If the similarity is higher than a first threshold, the current target information is determined to be redundant information; if the similarity is between the first threshold and the second threshold, and the current target information includes new entity information, the current target information is determined to be evolved information.

[0056] Specifically, a high similarity (e.g., greater than 0.9) indicates that the new information and historically known information are almost identical in semantic space, and can be initially judged as redundant or slightly updated. If no new entities / parameters are included, low weight or filtering is applied. A medium similarity (e.g., between 0.5 and 0.9) indicates that the new information is strongly related to historical information, but there is a certain deviation in direction, which may represent an important evolution. The entity recognition results will be further combined to determine whether new entities have appeared, thereby confirming its incremental value. A low similarity (e.g., less than 0.5) indicates that the new information is weakly related to the historical baseline, and may be a completely new topic or domain. It can be marked as new information with high potential value and enter the deep processing flow.

[0057] In this embodiment, by setting a first threshold and a second threshold, the continuous value of similarity is transformed into discrete value levels such as redundancy and evolution. This simulates the intuitive judgment logic of information value, thereby enabling rapid and accurate segmentation of massive amounts of information. Redundant information is assigned low weight or directly filtered to avoid consuming computational resources. Evolved information is marked and assigned a high inheritance score, triggering standard or enhanced deep processing procedures to uncover its incremental content.

[0058] Figure 3 This diagram illustrates a flowchart of another method for processing target information provided by an embodiment of this application. This method can be executed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be executed by software or hardware installed on the server and / or terminal device, and includes the following steps: Step 310: Obtain historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool.

[0059] Step 310 can be found above. Figure 1 The specific description of step 110 in the illustrated embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.

[0060] Step 320: Based on step 120 of the above embodiment, in response to the current target information having incremental value, the current target information is processed to obtain structured information of the current target information. The method of this embodiment may further include the following specific steps: The link discovery model is used to obtain one or more links associated with the current target information; the content parsing model is used to perform semantic parsing on the content pointed to by the links; the results of the semantic parsing are cleaned, deduplicated, and entity recognized to generate structured information with predefined fields.

[0061] The processing of current target information can be carried out through multiple intelligent models or agents working together.

[0062] First, instead of simply processing the initial content of the current target information (e.g., the body of a news report), a proactive exploration process is initiated through a link discovery model (or "discovery agent"). This model uses the current target information as a starting point (seed), analyzes its internal and external links, referencing relationships, and related network structures, and uses its reasoning capabilities to proactively filter and recommend a series of links highly relevant to the core topic and potentially containing deeper information. These links might point to official technical documentation, API descriptions for related products, in-depth discussion threads in technical communities, industry analysis and commentary articles, official announcements from relevant companies, etc.

[0063] The link set output by the link discovery model is then deeply processed by a content parsing model (or "extraction agent"). This model accesses the links, retrieves their original content (including handling dynamically loaded content from JavaScript), and performs key operations such as semantic understanding, main idea summarization, and relationship identification. Semantic understanding includes analyzing the overall page structure, identifying and understanding the core content of the main text, and distinguishing the main text from noisy modules such as advertisements and navigation. Main idea summarization involves extracting the core viewpoints and key information of the page content. Relationship identification involves preliminary analysis of the relationships between entities involved in the content (e.g., a company has released a product). This model transforms unstructured HTML or text into an intermediate representation containing core semantic information by simulating how humans read and understand web pages.

[0064] The raw results obtained from semantic parsing undergo a standardized post-processing workflow to generate output in a uniform format. Cleaning and deduplication include removing HTML tags, irrelevant scripts, and formatting characters, and merging and deduplicating highly repetitive text from different links to ensure information uniqueness. Entity recognition utilizes Network Entity Recognition (NER) technology to automatically identify and classify key entities from the cleaned text. Finally, the cleaned text summary, the identified entity list and their relationships, original links, and other metadata are encapsulated according to a predefined data pattern to form a structured data object.

[0065] Based on the implementation approach of the processing models, the "link discovery" and "content parsing" functions can be implemented as two independent, parameterized processing models (e.g., fine-tuned deep learning models). The link discovery model is a link ranking and classification model, whose input is seed information and a set of candidate links, and whose output is a ranked list of high-value links. The content parsing model is an optimized information extraction and text understanding model, whose input is the original web page content, and whose output is an intermediate representation of structured data.

[0066] The agent-based implementation can use software agents with autonomous reasoning and action capabilities to achieve the same functionality. The discovery agent is used to find relevant, high-quality links, autonomously planning actions (e.g., browsing pages, clicking links, analyzing anchor text), and utilizing internal or external tools (e.g., a link prediction model) to assist decision-making. The extraction agent is used to understand and extract core page information, autonomously deciding how to parse the page layout and which NER service to invoke. The agent-based implementation emphasizes the goal-driven nature of modules, their autonomous decision-making capabilities, and dynamic behavior.

[0067] This system utilizes a multi-agent collaborative approach to achieve dynamic, deep crawling of target information. It replaces traditional rule-based crawlers with the reasoning capabilities of agents, enabling "understanding-based crawling" that more intelligently understands webpage content, improving the efficiency and accuracy of information crawling. A "discovery agent" is deployed to traverse massive amounts of website links, while an "extraction agent" performs semantic parsing of unstructured HTML. Through reasonable task decomposition, different agents can focus on different tasks, improving the overall system performance. The crawled text is cleaned and deduplicated, and entity recognition is performed, converting unstructured text into structured objects containing fields such as "company," "technology stack," and "performance metrics." This facilitates further analysis and processing of the information.

[0068] Step 330: Based on the structured information, determine the comprehensive weight value of the current target information through a multidimensional scoring model.

[0069] Step 340: Generate the target content of the current target information based on the comprehensive weight value, and adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

[0070] Steps 330 and 340 can be found above. Figure 1 The specific descriptions of steps 130 and 140 in the illustrated embodiment are provided, and they achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0071] In this embodiment, by deploying a link discovery model / agent, a network of related information centered on the current target information is proactively constructed. This overcomes the limitations of single reports, intelligently tracking and acquiring diverse and in-depth related content such as official documents, technical discussions, and industry analyses. This ensures that the acquired information is not isolated points but a rich tapestry of knowledge, providing comprehensive and multi-dimensional raw materials for subsequent in-depth analysis and addressing the blind spots and superficiality in information acquisition. Employing a content parsing model / agent, deep semantic understanding can recognize the main text, extract the main idea, and analyze entity relationships like a human, rather than mechanically matching tags. This effectively handles unstructured, dynamically loaded complex web pages, accurately stripping away noise and extracting truly valuable core semantic information. The final output structured information object can provide a high-quality data foundation for downstream quantitative evaluation.

[0072] In yet another exemplary embodiment, a policy controller dynamically regulates the collaborative operation of the link discovery model and the content parsing model. This dynamic regulation includes: dynamically adjusting the access rate of the link discovery model to the target site in response to site feedback information acquired by the link discovery model during operation; dynamically adjusting the collection depth of content associated with a high-value entity identified by the content parsing model during parsing; and performing compliance verification based on the target site's access control protocol before the link discovery model initiates an access request.

[0073] Specifically, to ensure the stability of collaborative crawling, the system employs a policy controller between the link discovery model and the content parsing model. This controller executes the following logic: Adaptive rate limiting: Based on the token bucket algorithm, the number of concurrent crawlers is dynamically adjusted according to the target site's response time (Round-Trip Time, RTT) and HTTP 429 status code feedback.

[0074] Context-aware depth control: In response to the deep crawling logic triggered by "high-value entities", the policy controller dynamically allocates the crawling depth (Depth-Limit) based on the entity's "semantic centripetal value". If the semantic relevance of subsequent pages to the core entity is lower than the preset threshold, the crawling link is automatically cut off to prevent it from falling into a "crawler black hole".

[0075] Compliance Filtering: Before initiating a request, the system automatically parses the target site's robots.txt protocol and combines it with a pre-defined blacklist / whitelist filtering mechanism to ensure that data collection activities operate within legal and compliant boundaries. The robots.txt protocol is a non-mandatory, text-based industry standard used by website administrators to communicate with web crawlers (such as search engine spiders and data collection programs).

[0076] In this embodiment, adaptive rate limiting is implemented through a policy controller. Taking real-time feedback from the site as input, the concurrent request rate is dynamically adjusted using algorithms such as token bucket. This allows the system to sense and adapt to the load capacity and tolerance of the target site, achieving the best balance between user-friendly access and efficient data collection. This helps reduce the task failure rate and the risk of being blocked, ensuring high stability and success rate for large-scale, long-cycle data collection tasks.

[0077] In yet another exemplary embodiment, during the semantic parsing process, if a preset high-value entity is identified in the content, the priority of the current acquisition task is automatically increased, and the content acquisition scope is expanded to acquire extended content associated with the high-value entity.

[0078] High-value entities can be identified by introducing a specific entity perception mechanism. When the content parsing model identifies a pre-defined high-value entity in the text, it will automatically increase the priority of the current data collection task, trigger the deep mining mode, and instruct the link discovery model to expand the crawling scope around the entity (such as its associated project documents, social media comments, industry analysis reports, etc.), ensuring comprehensive coverage of high-value clues.

[0079] For example, when the system identifies specific company names such as Google, OpenAI, and NVIDIA in the text, it automatically raises the priority of the task and triggers a site-wide deep mining mode (i.e., crawling related API documents, Twitter comments, etc.), which ensures a comprehensive and in-depth mining of relevant information about important companies.

[0080] In this embodiment, the processing strategy is adjusted in real time according to the value of the information content. When a preset high-value entity is identified, the task priority is automatically increased and a deep mining mode is triggered. This allows limited computing resources, network bandwidth, and storage space to be concentrated on mining and tracking truly high-value information sources, which helps to improve the input-output ratio and operating efficiency of information content processing.

[0081] In yet another exemplary embodiment, based on step 130 of the above embodiment, based on the structured information, the comprehensive weight of the current target information is determined through a multidimensional scoring model. The method of this embodiment may further include the following specific steps: The multidimensional scoring model analyzes the structured information to evaluate the current target information from the content dimension and the source dimension, obtaining a content dimension score and a source dimension score; the content dimension score and the source dimension score are weighted and summed to obtain the comprehensive weight value.

[0082] The multidimensional scoring model can primarily evaluate structured information from two dimensions: Content-based assessment primarily evaluates the intrinsic value of information itself. It quantifies and scores content by analyzing fields such as entity lists, text summaries, and core relationships within structured information objects. It typically includes the following sub-dimensions (which can be configurable and combined): technological innovation, market influence, and information density.

[0083] Technological innovation assesses the novelty, breakthrough, or optimization of the technology involved in the information. Scoring criteria include whether the identified technological entities belong to cutting-edge fields, whether performance indicators have been significantly improved, and whether new algorithms or architectures are mentioned. Market influence assesses the potential impact of the information on industry structure, competitive landscape, or market dynamics. Scoring criteria include the industry position of the companies involved, and whether significant collaborations, investments, mergers and acquisitions, or policy changes are mentioned. Information density assesses the abundance and effectiveness of the information itself. Scoring criteria include the ratio of valid entities to text length, the logical coherence of the argument, and the completeness of data support.

[0084] Based on preset rules or a lightweight scoring model, calculations are performed on each sub-dimension, and finally aggregated (e.g., by taking the average) to generate a comprehensive content dimension score.

[0085] Source dimension assessment is primarily used to evaluate the credibility and initial importance bias of information. Based on a pre-defined, dynamically maintainable source authority mapping table, a source compensation weight is directly assigned to the source of the current target information. This mapping table can be graded according to the source's historical accuracy, domain expertise, institutional credibility, and other long-term performance. For example, top academic journals, authoritative official channels, and well-known industry media have higher weights (e.g., 0.8-1.0), while ordinary blogs and personal social media accounts have lower weights (e.g., 0.1-0.3).

[0086] After obtaining the content dimension score and source compensation weight, they are synthesized through a preset quantitative formula to calculate the final comprehensive weight value.

[0087] In this embodiment, the evaluation of information value is transformed from relying on subjective experience or a single popularity indicator to an automated calculation process based on clearly defined dimensions, rules, and mathematical formulas. The final weight of each piece of information is traceable and can be explained through its content dimension score and source compensation weight, thereby improving the objectivity, transparency, and interpretability of the evaluation process. By outputting a single, comparable comprehensive weight value, a direct and accurate decision-making basis is provided for subsequent information processing, making automated priority management of massive amounts of information possible. This ensures that the highest-value information is presented and processed first, significantly improving information flow efficiency.

[0088] In yet another exemplary embodiment, the content dimension includes at least two of the following: content innovation, market influence, and information density.

[0089] This embodiment uses news as the target information for specific illustration. By establishing a multi-dimensional evaluation matrix, the qualitative concept of "good news" is transformed into quantitative values. Based on this quantitative approach, the value of news can be evaluated more objectively and accurately.

[0090] First, four core scoring dimensions are set, with values ​​typically ranging from [0.0 to 1.0]. These dimensions evaluate the value of the news from different perspectives, and can comprehensively and accurately reflect the quality of the news.

[0091] Technological innovation (S_tech) is used to determine whether it includes algorithm optimization, model release, or hardware breakthroughs. This dimension can measure the technological innovation and advancement of news.

[0092] Market impact (S_market) is used to determine its influence on the industry landscape. This dimension reflects the news's influence and importance on the market.

[0093] Information density (S_density) is the ratio of effective information content to the number of words in a text. This dimension can be used to assess the richness and effectiveness of news content.

[0094] The source compensation weight (W_source) is initially biased based on the source's authority and a pre-defined "high-value company list." For example, the initial W_source value for an official OpenAI announcement could be set to 0.95, while the initial W_source value for a regular blog post might be set to 0.3. This weight reflects the differences in credibility and importance of news from different sources.

[0095] Then, the formula for calculating the overall weight value is set as follows: W_final = α* (S_tech + S_market + S_density) / 3 + β* W_source Where α and β are dynamically adjusted weighting coefficients.

[0096] This formula allows for a comprehensive consideration of factors across various dimensions to arrive at the final overall weight value for the news item.

[0097] In this embodiment, by specifically defining the dimensional composition, quantification methods, and synthesis formulas of the multidimensional scoring model, four core dimensions are clearly defined: technological innovation, market influence, information density, and source authority. For each dimension, calculable scoring rules are defined (e.g., S_tech determines technological breakthroughs, and S_density calculates the effective information ratio). This makes "good news" no longer a vague concept, but a quantifiable object characterized by multiple objective indicators. News can be accurately scored from different perspectives, achieving standardization, objectification, and automation of value assessment. The evaluation framework employs content and source dimensions. The content dimension delves into the intrinsic quality of the information itself, identifying its technological value and depth of insight. The source dimension assesses its external credibility and basic authority. Finally, a weighted formula synthesizes the two, ensuring that the evaluation results neither blindly favor news reports with empty content due to authoritative sources nor bury excellent in-depth analysis due to ordinary sources. This allows for a comprehensive and balanced assessment of the composite value of information, with stronger resistance to interference.

[0098] As a further improvement to the above embodiment, the weight coefficients α and β can be constrained to α+β=1, making the calculation result of the comprehensive weight value more stable.

[0099] In this embodiment, the strategy can be switched in different scenarios by dynamically adjusting the weight coefficients α and β. For example, in the scenario of tracking the forefront of technology, α (content weight) can be increased to focus more on the technological breakthrough itself. Or, in the scenario of major decision-making reference, α can be decreased (or β increased) to rely more on information from authoritative sources. Constraining α and β to α+β=1 ensures that no matter how α and β are adjusted, the output range of the comprehensive weight value remains stable (e.g., falling within the range of 0 to 1), making the weight values ​​calculated under different periods and strategies directly comparable. This provides a good mathematical foundation for subsequent long-term ranking, threshold filtering, and historical comparison, and helps to solve the score drift problem in model iteration.

[0100] In yet another exemplary embodiment, based on step 140 of the above embodiment, the target content of the current target information is generated according to the comprehensive weight value, and the parameters of the multidimensional scoring model are adjusted according to the feedback data of the target content. The method of this embodiment may further include the following specific steps: Multiple current target information items are sorted based on their respective comprehensive weight values; according to the sorting results, a structured summary is generated through a large language model as the target content.

[0101] The target content can be a structured summary or outline. Using a finely tuned LLM, news items are sorted in descending order based on W_final and then structured and summarized according to "Industry Trends," "Technological Frontiers," and "In-depth Commentary." This structured summary allows users to quickly grasp the main content and key information of the news.

[0102] In this embodiment, after sorting based on comprehensive weight values, the large language model is driven to generate a structured summary or overview. This proactively integrates, summarizes, and re-creates high-value information, outputting a clearly themed, chapter-based report. This eliminates the need for manual filtering and piecing together from massive amounts of articles, directly providing a structured knowledge briefing. This improves information acquisition and enhances the efficiency of information digestion and decision support. The generated structured summary or overview, as the final target content presented to the user, serves as the most direct and effective vehicle for collecting user feedback. User behavior data, such as reading time and interaction depth for the overall report and each chapter, provides a basis for adjusting the parameters of the multidimensional scoring model and optimizing the LLM generation prompts. This ensures continuous alignment of content generation quality with user preferences, contributing to a complete reinforcement loop.

[0103] In yet another exemplary embodiment, based on step 140 of the above embodiment, the target content of the current target information is generated according to the comprehensive weight value, and the parameters of the multidimensional scoring model are adjusted according to the feedback data of the target content. The method of this embodiment may further include the following specific steps: Collect feedback data on the target content, including click behavior data or reading time data; adjust the weight coefficients used by the multidimensional scoring model in determining the comprehensive weight value based on the feedback data.

[0104] The feedback data can include user feedback such as reading time and click behavior on the report. If a user frequently clicks on a certain type of news that the system has rated low, the weighting coefficients used in the multi-dimensional scoring model in step 130 can be automatically adjusted, forming a continuously iterative closed-loop system. Through this feedback mechanism, the system's performance can be continuously optimized, improving user satisfaction.

[0105] In yet another exemplary embodiment, adjusting the weight coefficients used by the multidimensional scoring model in determining the comprehensive weight value based on the feedback data includes the following steps: Obtain user interaction feedback data regarding the target information; based on the interaction feedback data, determine the feedback adjustment amount for the weight coefficients of the multidimensional scoring model; process the feedback adjustment amount using a damped feedback condition algorithm to generate a smoothed weight adjustment amount; update the weight coefficients of the multidimensional scoring model based on the smoothed weight adjustment amount.

[0106] The damping smoothing algorithm is used for feedback smoothing and weight stability mechanisms in multidimensional scoring models. Specifically, when adjusting the parameters of the multidimensional scoring model based on feedback data, a damping feedback adjustment algorithm is introduced. Instead of directly modifying the weights based on a single feedback iteration, it calculates the exponential moving average (EMA) of the feedback vector. The update formula is shown below: ω_new = (1 - α) · ω_old + α · ω_feedback Where ω_new is the updated weight coefficient value, ω_old is the current weight coefficient value, ω_feedback is the calculated value of the feedback data, and α is the learning rate (damping coefficient), with a value range of (0, 0.1).

[0107] This approach ensures that the evaluation criteria capture the long-term evolution of user preferences while filtering out accidental and extreme user ratings, thus guaranteeing the smoothness and consistency of the generated AI news summaries in terms of quality.

[0108] Corresponding to the target information processing method provided in the above embodiments, based on the same technical concept, this application also provides a target information processing apparatus. See also Figure 4 The device 400 includes an acquisition module 410, a processing module 420, a scoring module 430, and a generation module 440.

[0109] The acquisition module 410 is used to acquire historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool; the processing module 420 is used to process the current target information in response to the incremental value of the current target information to obtain the structured information of the current target information; the scoring module 430 is used to determine the comprehensive weight value of the current target information based on the structured information through a multidimensional scoring model; and the generation module 440 is used to generate the target content of the current target information based on the comprehensive weight value, and adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

[0110] It should be noted that the target information processing device and the target information processing method provided in the embodiments of this application are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned target information processing method, and the repeated parts will not be described again.

[0111] The target information processing method and apparatus provided in this application can be applied to AI news automation processing systems, solving the fragmentation problem of information aggregation systems and giving the output summaries profound background reference value. This allows users to better understand the ins and outs of information when obtaining news, forming a more coherent understanding. Through multi-dimensional scoring, industry noise can be filtered out, making it easier for users to obtain valuable information and avoiding interference from a large amount of noisy information. The system does not rely on fixed rules and can automatically adjust its focus according to the rise and fall of AI industry entities, enabling the system to better adapt to industry changes and provide users with news content that better meets their needs.

[0112] This application relates to the fields of internet information gathering, natural language processing (NLP), and large language model (LLM) applications. Specifically, it relates to a news aggregation and distribution system that combines a temporal semantic inheritance mechanism, multi-agent collaborative crawling, and a high-dimensional weighted quantitative evaluation model. Based on these advanced technologies, this system aims to solve problems related to information integration, filtering, and distribution in the news processing process, thereby improving the efficiency and quality of news processing.

[0113] Corresponding to the target information processing method provided in the above embodiments, based on the same technical concept, this application also provides an electronic device for executing the above method. Figure 5 To illustrate the structure of an electronic device according to various embodiments of this application, as shown in the following diagrams... Figure 5 As shown. Electronic device 500 can vary considerably due to differences in configuration or performance, and may include one or more processors 510 and memory 520. Memory 520 may store one or more application programs or data. Memory 520 may be temporary or persistent storage. The application programs stored in memory 520 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device. Furthermore, processor 510 may be configured to communicate with memory 520 and execute the series of computer-executable instructions stored in memory 520 on the electronic device.

[0114] This application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, implement the steps of the target information processing method described above.

[0115] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, apparatus, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0116] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems, devices), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0119] In a typical configuration, an electronic device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0120] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0122] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0123] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0124] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for processing target information, characterized in that, The method includes the following steps: Acquire historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool; In response to the fact that the current target information has incremental value, the current target information is processed to obtain structured information of the current target information; Based on the structured information, a multidimensional scoring model is used to determine the comprehensive weight value of the current target information; The target content of the current target information is generated based on the comprehensive weight value, and the parameters of the multidimensional scoring model are adjusted based on the feedback data of the target content.

2. The method according to claim 1, characterized in that, The process of acquiring historical information and current target information, obtaining a historical semantic baseline pool based on the historical information, and determining the incremental value of the current target information based on the historical semantic baseline pool includes the following steps: Extract semantic feature vectors from each historical information to form a historical semantic baseline pool; Extract the semantic feature vector of the current target information, calculate the similarity between the semantic feature vector of the current target information and the vector in the historical semantic baseline pool, and determine the incremental value of the current target information relative to the historical information based on the similarity.

3. The method according to claim 2, characterized in that, Determining the incremental value of the current target information relative to historical information based on the similarity includes the following steps: If the similarity is higher than a first threshold, the current target information is determined to be redundant information; If the similarity is between the first threshold and the second threshold, and the current target information includes new entity information, the current target information is determined to be evolution information.

4. The method according to claim 1, characterized in that, The step of processing the current target information to obtain structured information in response to the incremental value of the current target information includes the following steps: The link discovery model retrieves one or more links associated with the current target information; The content pointed to by the links is semantically parsed using a content parsing model; The semantic parsing results are cleaned, deduplicated, and entity recognized to generate structured information with predefined fields.

5. The method according to claim 1, characterized in that, The step of determining the comprehensive weight value of the current target information based on the structured information through a multidimensional scoring model includes the following steps: The multidimensional scoring model analyzes the structured information and evaluates the current target information from the content dimension and the source dimension to obtain the content dimension score and the source dimension score. The content dimension score and the source dimension score are weighted and summed to obtain the comprehensive weight value.

6. The method according to claim 1, characterized in that, The step of generating the target content of the current target information based on the comprehensive weight value, and adjusting the parameters of the multidimensional scoring model based on the feedback data of the target content, includes the following steps: Sort multiple current target information based on their respective comprehensive weight values; Based on the sorting results, a structured summary is generated using a large language model, which serves as the target content.

7. The method according to claim 1, characterized in that, The step of generating the target content of the current target information based on the comprehensive weight value, and adjusting the parameters of the multidimensional scoring model based on the feedback data of the target content, includes the following steps: Collect feedback data on the target content, including click behavior data or reading time data; Based on the feedback data, the weight coefficients used by the multidimensional scoring model in determining the comprehensive weight value are adjusted.

8. A target information processing apparatus, characterized in that, The device includes the following: The acquisition module is used to acquire historical information and current target information, obtain a historical semantic baseline pool based on the historical information, and determine the incremental value of the current target information based on the historical semantic baseline pool. The processing module is configured to process the current target information in response to the fact that the current target information has incremental value, and obtain structured information of the current target information; The scoring module is used to determine the comprehensive weight value of the current target information based on the structured information through a multi-dimensional scoring model; as well as, The generation module is used to generate target content of the current target information based on the comprehensive weight value, and to adjust the parameters of the multidimensional scoring model based on the feedback data of the target content.

9. An electronic device, characterized in that, It includes a processor, a memory, a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method for processing target information as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, implement the steps of the method for processing target information as described in any one of claims 1 to 7.