Cross-source information processing method and device based on meta-learning, equipment and storage medium
Through the cross-original information processing method based on meta-learning, the internal loop optimization specific information source strategies and the general mechanism of external loop learning are solved, and efficient and low-cost cross-original information understanding and integration are achieved.
Patent Information
- Application Number
- CN202510914378.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multi-source heterogeneous information, the prior art has problems such as insufficient generalization capability and poor adaptability to new data types, which leads to high computing costs and difficulty in achieving unified processing and integration of cross-origin information.
The cross-original information processing method based on meta-learning is adopted, and the target large language model and two-layer optimization architecture (internal loop and external loop) are used for information processing, the internal loop optimizes the processing strategies of specific information sources, and the external loop learns a general processing mechanism to generate a target processing framework to uniformly process multi-source information.
It realizes the powerful generalization ability and self-evolution characteristics of the cross-original information processing system, can efficiently process heterogeneous information sources, have the ability to accurately process known information sources, and quickly adapt to new information sources, reducing system development and maintenance costs.
Smart Images

Figure CN120409555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing, and particularly to a cross-source information processing method, apparatus, device, and storage medium based on meta-learning. Background Art
[0002] With the rapid development of Internet technology and the continuous improvement of the degree of informatization, various types of information sources have shown an explosive growth, including news websites, academic databases, social media, professional forums, enterprise reports, and other heterogeneous information sources. These information sources have significant differences in data structure, content format, language style, information quality, etc., bringing huge challenges to the unified processing and effective utilization of information. In recent years, although some information processing technologies based on machine learning have been developed, these methods still have problems such as insufficient generalization ability and poor adaptability to new data types. Especially when dealing with multi-source heterogeneous information, existing technologies often need to train models separately for each type of information source, which not only increases the computational cost but also makes it difficult to achieve true cross-source information understanding and integration.
[0003] Therefore, there is an urgent need for an intelligent information processing technology that can uniformly process heterogeneous information sources, has strong generalization ability and the ability to quickly adapt to new information sources, in order to meet the growing demand for cross-source information aggregation. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a cross-source information processing method, apparatus, device, and storage medium based on meta-learning, which can meet the growing demand for cross-source information aggregation. The specific solutions are as follows:
[0005] In a first aspect, the present application discloses a cross-source information processing method based on meta-learning, which is applied to a cross-source information processing system and includes:
[0006] Obtain the cases to be optimized corresponding to each target information source, and use the target large language model to analyze the target information source to determine the information processing characteristics and information processing challenges corresponding to each target information source;
[0007] Use the inner-loop optimization architecture of the target double-layer optimization architecture to optimize the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determine the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies;
[0008] Use the outer-loop optimization architecture of the target double-layer optimization architecture to perform aggregation processing on the information processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources, and use the target large language model to determine the corresponding target general strategy based on the target common challenges;
[0009] Determine the target processing strategy corresponding to each target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, and determine the target processing framework based on all the target processing strategies, so as to process the cross-source information obtained from all the target information sources by using the target processing framework;
[0010] Among them, the target two-layer optimization architecture is a two-layer optimization architecture implemented based on meta-learning.
[0011] Optionally, before obtaining the cases to be optimized corresponding to each target information source, it further includes:
[0012] Determine each target information source;
[0013] Use the target application programming interface to dock the cross-source information processing system with each target information source;
[0014] Obtain the cross-source information to be processed from the target information source based on the target subscription method and the target information acquisition frequency.
[0015] Optionally, the cross-source information processing method based on meta-learning further includes:
[0016] Evaluate the information quality of the cross-source information obtained from each target information source based on a preset information scoring algorithm;
[0017] Among them, the scoring dimensions of the preset information scoring algorithm include information timeliness, information authority, and information integrity; the score corresponding to information timeliness is the score determined based on the difference between the information release time of the cross-source information and the current time, the information authority is the score determined based on the authority level of the target information source corresponding to the cross-source information, and the information integrity is the score determined based on the text length and structural integrity of the cross-source information.
[0018] Optionally, optimizing the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determining the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies, includes:
[0019] Use the cross-source information processing system to construct a first analysis prompt template;
[0020] Batch input the case to be optimized into the target large language model based on the first analysis prompt template to obtain the corresponding target analysis result;
[0021] Optimize the processing strategy corresponding to the case to be optimized based on the target analysis result to determine the optimized processing strategy corresponding to the case to be optimized;
[0022] Use a preset automated testing mechanism to score all the optimized post - processing strategies to determine the target optimized strategy corresponding to the case to be optimized.
[0023] Optionally, the aggregating the information - processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources includes:
[0024] Aggregate the information - processing challenges corresponding to all the target information sources based on the challenge type to establish a target cross - source challenge knowledge base corresponding to each challenge type;
[0025] Use the cross - source information - processing system to construct a second analysis prompt template;
[0026] Based on the second analysis prompt template, use the target large - language model to analyze the cross - source challenge data in each target cross - source challenge knowledge base to determine the target common challenges corresponding to all the target information sources;
[0027] Wherein, the challenge types include structured data extraction challenges, unstructured text understanding challenges, and multi - modal content processing challenges.
[0028] Optionally, the cross - source information - processing method based on meta - learning further includes:
[0029] If there is a new target information source, use a target few - shot learning algorithm to construct the target processing strategy corresponding to the new target information source based on the target general strategy.
[0030] Optionally, the processing the cross - source information obtained from all the target information sources by using the target processing framework includes:
[0031] Use the target processing framework to uniformly extract the cross - source information obtained from all the target information sources to obtain a target extraction result;
[0032] Use the target double - layer integration framework of the cross - source information - processing system to call the target large - language model to generate an information - processing result corresponding to the cross - source information based on the target extraction result;
[0033] Wherein, the information - processing result includes fact summary, view analysis, and trend prediction.
[0034] In a second aspect, the present application discloses a cross - source information - processing device based on meta - learning, which is applied to a cross - source information - processing system and includes:
[0035] A case acquisition module, configured to acquire cases to be optimized corresponding to each target information source, and analyze the target information sources using a target large language model to determine the information processing characteristics and information processing challenges corresponding to each of the target information sources;
[0036] An inner loop optimization module, configured to optimize the processing strategies corresponding to the cases to be optimized using the inner loop optimization architecture of a target two-layer optimization architecture to determine the optimized processing strategies corresponding to the cases to be optimized, and determine the target optimized strategies corresponding to the cases to be optimized based on all the optimized processing strategies;
[0037] An outer loop optimization module, configured to perform an aggregation process on the information processing challenges corresponding to all the target information sources using the outer loop optimization architecture of the target two-layer optimization architecture to determine the target common challenges corresponding to all the target information sources, and determine the corresponding target general strategies based on the target common challenges using the target large language model;
[0038] An information processing module, configured to determine the target processing strategies corresponding to the target information sources based on the target optimized strategies and the target general strategies corresponding to each of the target information sources, and determine a target processing framework based on all the target processing strategies, so as to process cross-source information obtained from all the target information sources using the target processing framework;
[0039] Wherein, the target two-layer optimization architecture is a two-layer optimization architecture implemented based on meta-learning.
[0040] In a third aspect, the present application discloses an electronic device, including:
[0041] A memory, configured to store a computer program;
[0042] A processor, configured to execute the computer program to implement the aforementioned cross-source information processing method based on meta-learning.
[0043] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the aforementioned cross-source information processing method based on meta-learning.
[0044] In this application, when processing cross-source information, the cross-source information processing system obtains the cases to be optimized corresponding to each target information source, and uses the target large language model to analyze the target information source to determine the information processing characteristics and information processing challenges corresponding to each target information source; uses the inner loop optimization architecture of the target two-layer optimization architecture to optimize the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determines the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies; uses the outer loop optimization architecture of the target two-layer optimization architecture to perform aggregation processing on the information processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources, and uses the target large language model to determine the corresponding target general strategy based on the target common challenges; determines the target processing strategy corresponding to the target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, and determines the target processing framework based on all the target processing strategies, so as to use the target processing framework to process the cross-source information obtained from all the target information sources; wherein, the target two-layer optimization architecture is a two-layer optimization architecture implemented based on meta-learning. It can be seen that the cross-source information of multiple target information sources obtained in this application provides a rich training data basis for meta-learning. Meta-learning is performed using the target two-layer optimization architecture including an inner loop architecture and an outer loop architecture, realizing the autonomous learning and continuous optimization of the cross-source information processing ability. Among them, the inner loop architecture focuses on improving the processing ability of specific information source types, and uses the target large language model to optimize the strategy corresponding to the case to be optimized to obtain the target optimized strategy, ensuring that the system can achieve the best processing effect on each known information source type, and laying a foundation for building a high-quality information aggregation system; the outer loop architecture focuses on learning the general processing mechanism across information source types, and identifies the core bottlenecks and key elements of cross-source information processing by analyzing the common processing problems existing in different information sources, so as to generate a general processing strategy (i.e., the target general strategy) that can improve the cross-source processing ability. The outer loop is the core manifestation of meta-learning, enabling the cross-source information processing system to have strong generalization ability. The inner loop ensures the processing quality of the system on each known information source and provides high-quality learning data for the outer loop; the general mechanism learned by the outer loop guides the inner loop to better process specific information sources, enabling the cross-source information processing system to have the precise processing ability for known information sources, and realizing the intelligent understanding and unified processing of heterogeneous information sources. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0046] Figure 1 Flowchart of a cross-source information processing method based on meta-learning disclosed in this application;
[0047] Figure 2 Schematic structural diagram of a cross-source information processing device based on meta-learning disclosed in this application;
[0048] Figure 3 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0050] With the rapid development of Internet technology and the continuous improvement of the degree of informatization, various types of information sources have shown an explosive growth, including news websites, academic databases, social media, professional forums, enterprise reports and other heterogeneous information sources. These information sources have significant differences in data structure, content format, language style, information quality, etc., bringing huge challenges to the unified processing and effective utilization of information. In recent years, although some information processing technologies based on machine learning have been developed, these methods still have problems such as insufficient generalization ability and poor adaptability to new data types. Especially when dealing with multi-source heterogeneous information, the existing technologies often need to train models separately for each type of information source, which not only increases the computational cost, but also makes it difficult to achieve true cross-source information understanding and integration. To solve the above technical problems, this application discloses a cross-source information processing method based on meta-learning, which can meet the growing demand for cross-source information aggregation.
[0051] See Figure 1 As shown, the embodiments of the present invention disclose a cross-source information processing method based on meta-learning, which is applied to a cross-source information processing system and includes:
[0052] Step S11: Obtain the cases to be optimized corresponding to each target information source, and use the target large language model to analyze the target information source to determine the information processing characteristics and information processing challenges corresponding to each target information source.
[0053] In this embodiment, the cross-source information processing system is a cross-source information intelligent collection and integration system based on meta-learning. This system constructs an adaptive cross-source information processing framework through the meta-learning algorithm to achieve intelligent understanding and unified processing of heterogeneous information sources, thereby creating an information aggregation system with strong generalization ability and self-evolution characteristics, avoiding the complexity of the traditional method that requires developing separate strategies for each information source. The entire system is deployed using a microservices architecture, and each module communicates through RESTful API (Application Programming Interface). The data storage uses a MySQL database to store structured data, MongoDB to store unstructured content, and Redis as a cache layer to improve the system response speed. The system supports horizontal scaling and can dynamically increase or decrease service instances according to the processing volume requirements.
[0054] In this embodiment, before obtaining the cases to be optimized corresponding to each target information source, it is also necessary to determine each target information source; use the target application programming interface to dock the cross-source information processing system with each target information source; and obtain the cross-source information to be processed from the target information source based on the target subscription method and the target information acquisition frequency. Taking the financial investment decision-making scenario as an example, this system needs to collect information from multiple heterogeneous information sources such as news websites, brokerage research reports, social media, and regulatory announcements, and integrate and generate investment analysis reports. Therefore, in the system initialization stage, an information source registry is first established through manual configuration, including major financial news websites, brokerage research report platforms, social media platforms, regulatory information disclosure websites, etc. The system uses standard APIs and RSS (Really Simple Syndication) subscription methods to regularly obtain the content of each information source, and configures technical parameters such as access frequency control, data format parsing rules, and authentication parameters. By identifying, classifying, and managing multiple heterogeneous information sources, a rich training data foundation can be provided for meta-learning.
[0055] It can be understood that since the quality of information may vary due to information sources, timeliness, and content integrity, cross-source information obtained from each target information source can be evaluated for information quality based on a preset information scoring algorithm. Among them, the scoring dimensions of the preset information scoring algorithm include information timeliness, information authority, and information integrity. The score corresponding to information timeliness is a score determined based on the difference between the information release time of the cross-source information and the current time. The information authority is a score determined based on the authority level of the target information source corresponding to the cross-source information. The information integrity is a score determined based on the text length and structural integrity of the cross-source information. Specifically, the preset information scoring algorithm can adopt a weighted average algorithm. For example, the total score of the information = 0.4 × timeliness score + 0.4 × authority score + 0.2 × integrity score.
[0056] In this embodiment, after the system is initialized and each target information source is connected to the system, the optimization cases corresponding to each target information source can be obtained, and the target large language model can be used to analyze the target information source to determine the information processing characteristics and information processing challenges corresponding to each target information source. Taking the processing of brokerage research reports as an example, the system first establishes an optimization case database to record specific instances where errors occur in the extraction of research report information or the extraction results do not meet expectations, including optimization types such as PDF parsing errors, incomplete extraction of table data, and incorrect identification of investment ratings. The information processing challenge is a set of common problems obtained by abstracting and summarizing the optimization cases of all target information sources using a large model. Specifically, it includes, but is not limited to, the optimization cases of various information sources collected in the inner loop, technical difficulties identified during the processing, such as structured data extraction and multi-modal content understanding, and common problems found during cross-source processing, such as timeliness differences and inconsistent formats.
[0057] Step S12: Use the inner loop optimization architecture of the target double-layer optimization architecture to optimize the processing strategy corresponding to the optimization case to determine the optimized processing strategy corresponding to the optimization case, and determine the target optimized strategy corresponding to the optimization case based on all the optimized processing strategies.
[0058] In this embodiment, the target double-layer optimization architecture is a double-layer optimization architecture implemented based on meta-learning, including an inner-loop architecture and an outer-loop architecture. The inner-loop optimization is performed separately for each type of information source (academic papers, news reports, social media, professional forums, etc.), collecting the failure cases of the current processing framework on this type of information source, analyzing the common features and failure reasons of these failure cases through a large language model, generating improvement solutions for this type of information source based on the analysis results, and selecting the optimal processing strategy through performance evaluation. The inner-loop ensures that the system can achieve the best processing effect on each known type of information source, laying a foundation for building a high-quality information aggregation system. Specifically, optimizing the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determining the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies may include: constructing a first analysis prompt template using a cross-source information processing system; batch inputting the case to be optimized into the target large language model based on the first analysis prompt template to obtain the corresponding target analysis result; optimizing the processing strategy corresponding to the case to be optimized based on the target analysis result to determine the optimized processing strategy corresponding to the case to be optimized; using a preset automated test mechanism to score all the optimized processing strategies to determine the target optimized strategy corresponding to the case to be optimized.
[0059] In a specific implementation manner, the analysis of the case to be optimized is performed in batches using the target large language model. The system constructs a first analysis prompt template such as "Please analyze the following failure cases of brokerage research report processing: [list of failure cases]. Please identify the common failure patterns, analyze the root causes of the failures, and propose improvement suggestions. Output format: 1. Failure pattern: [specific pattern] 2. Root cause: [analysis result] 3. Improvement suggestion: [specific suggestion]". The large language model analyzes the failure cases based on this prompt template and outputs a structured target analysis result. Based on the target analysis result, the system generates multiple improvement solutions for the processing strategy (i.e., the optimized processing strategy corresponding to the case to be optimized). The improvement solutions exist in the form of optimized prompt templates. For example, the optimized prompt template for brokerage research reports can be: "You are a professional financial analyst. Please extract the following key information from the following brokerage research reports: 1. Underlying stock code and name 2. Investment rating (buy / hold / sell) 3. Target price 4. Core investment logic 5. Risk warning. Please ensure that the extracted values are accurate and the rating expressions are standardized."
[0060] In this embodiment, the performance evaluation of the inner loop adopts a targeted evaluation strategy, and different types of information sources adopt different combinations of evaluation indicators, including: (1) For structured information extraction tasks (such as investment ratings and target prices in brokerage research reports), the F1 score of exact matching is adopted; (2) For unstructured text understanding tasks (such as news sentiment analysis), semantic similarity scoring is adopted; (3) For information sources with high real-time requirements (such as social media), processing latency and information freshness indicators are added. The evaluation process is carried out on the dedicated test sets of various information sources to ensure that the evaluation results can truly reflect the actual effects of the processing strategies. That is to say, when using the preset automated test mechanism to score all the optimized post-processing strategies to determine the target optimized strategy corresponding to the case to be optimized, the corresponding evaluation index sets can be adopted according to different information source types. For example, for structured information extraction, accuracy, recall rate, and F1 score are adopted; for unstructured text understanding, semantic similarity and key information coverage rate are adopted; for information sources with high timeliness requirements, a time delay index is added. In a specific implementation manner, the extraction accuracy of each optimized post-processing strategy is tested on the verification dataset of brokerage research reports. The verification dataset contains 100 labeled brokerage research report samples, and the labeled content includes key information such as the investment rating and target price of the standard answer. The system calculates the F1 score of each optimized post-processing strategy and selects the solution with the highest score as the target optimized strategy for brokerage research reports.
[0061] Step S13: Use the outer loop optimization architecture of the target double-layer optimization architecture to aggregate the information processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources, and use the target large language model to determine the corresponding target general strategy based on the target common challenges.
[0062] In this embodiment, the outer loop optimization focuses on learning a general processing mechanism across information source types, aggregating the processing challenges of all information source types, analyzing the common processing problems existing between different information sources, identifying the core bottlenecks and key elements of cross-source information processing. Based on these common problems, a general processing strategy that can improve cross-source processing capabilities is generated through a large language model, and the comprehensive performance of the framework is evaluated on all information source types. Through iterative optimization, a general processing mechanism that can adapt to various information source types is learned. The outer loop is the core manifestation of meta-learning, enabling the information aggregation system to have strong generalization capabilities. Specifically, aggregating the information processing challenges corresponding to all target information sources to determine the target common challenges corresponding to all target information sources may include: aggregating the information processing challenges corresponding to all target information sources based on the challenge type to establish a target cross-source challenge knowledge base corresponding to each challenge type; constructing a second analysis prompt template using a cross-source information processing system; analyzing the cross-source challenge data in each target cross-source challenge knowledge base using a target large language model based on the second analysis prompt template to determine the target common challenges corresponding to all target information sources; where the challenge types include structured data extraction challenges, unstructured text understanding challenges, and multimodal content processing challenges.
[0063] In a specific implementation, the system aggregates the processing challenge data of all information source types (news, research reports, social media, announcements) to establish a cross-source challenge knowledge base. The knowledge base is classified by challenge type, including structured data extraction challenges (table and chart processing), unstructured text understanding challenges (sentiment analysis, event extraction), multimodal content processing challenges (graphic-text mixed content), etc. Then, the system uses the target large language model to perform cross-source commonality analysis based on this cross-source challenge knowledge base to determine the target general strategy based on the target common challenges using the target large model. Among them, the second analysis prompt template can be: "Please analyze the following processing challenges from different information sources: [cross-source challenge data]. Please identify the common factors affecting the processing effects of multiple information sources, analyze the key bottlenecks, and propose general solutions. Output format: 1. Common factors: [specific factors] 2. Key bottlenecks: [bottleneck analysis] 3. General strategy: [strategy plan]". Based on the above process of cross-source commonality analysis, the system generates a candidate solution for the general processing mechanism (i.e., the target general strategy). The general mechanism exists in the form of a system-level prompt template, for example: "You are a professional information analysis expert. Please follow the following general principles to process various information sources: 1. Prioritize extracting key factual information 2. Identify the time, source, and credibility of the information 3. Standardize the output format 4. Mark uncertain or missing information."
[0064] Step S14: Determine the target processing strategy corresponding to each target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, and determine the target processing framework based on all the target processing strategies, so as to process the cross-source information obtained from all the target information sources by using the target processing framework.
[0065] In this embodiment, when determining the target processing strategy corresponding to each target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, a weighted fusion algorithm can be used to combine the target general strategy optimized by the outer loop and the dedicated strategies (i.e., the target optimized strategies) corresponding to each target information source optimized by the inner loop, so as to obtain the target processing strategy corresponding to each target information source, thereby realizing the collaborative mechanism of the inner and outer loops. The specific implementation is as follows: Final prompt = × General mechanism prompt + × Dedicated strategy prompt, where and are weight parameters, which are optimized and determined on the validation set through a grid search algorithm.
[0066] It should be noted that the collaborative mechanism of the inner and outer loops realizes the unity of specialization and generalization capabilities. The inner loop ensures the processing quality of the system on each known information source and provides high-quality learning data for the outer loop; the general mechanism learned by the outer loop guides the inner loop to better process specific information sources and provides a strong knowledge base for processing new information sources. This two-layer structure enables the information aggregation system to have both the precise processing ability for known information sources and the fast adaptation ability for new information sources.
[0067] In this embodiment, if there is a new target information source, the target few-shot learning algorithm is used to construct the target processing strategy corresponding to the new target information source based on the target general strategy. Specifically, when the system encounters a new type of information source (such as an emerging investment community website), the general processing mechanism learned by the outer loop serves as a strong knowledge base, greatly reducing the learning cost of adapting to the new information source. The system adopts few-shot learning technology and only needs a small number of new information source examples to quickly establish an effective processing strategy through the inner loop. By calculating the feature similarity through a large language model, the new information source is mapped to a similar known type, drawing on similar processing experiences, and establishing an incremental learning mechanism to continuously optimize the processing effect. This fast adaptation ability ensures the scalability and continuous effectiveness of the information aggregation system.
[0068] In a specific implementation, when the system encounters a new information source, the feature similarity calculation uses the cosine similarity algorithm to compare the feature vectors of the new information source with those of the known information sources to find the most similar known type. Few-shot learning adopts the few-shot prompting technique, and only 5-10 new information source examples are required to establish a preliminary processing strategy. The incremental learning mechanism continuously collects the processing effect data of new information sources during actual use and periodically updates the processing strategy. The fast adaptation mechanism built based on the target general strategy obtained from the outer loop can achieve efficient processing of information of new information source types. The double-layer optimization architecture not only ensures the processing quality of specific information sources but also realizes the fast adaptation to new information sources, greatly reducing the system development and maintenance costs. The meta-learning framework ensures the continuous learning and self-evolution ability of the system, ensuring the long-term competitive advantage of the information aggregation system.
[0069] In this embodiment, the target processing framework is an optimized processing framework implemented based on the target general strategy obtained through outer loop optimization and all the dedicated strategies obtained through inner loop optimization. The cross-source information obtained from all target information sources is processed using the target processing framework, including: uniformly extracting the cross-source information obtained from all target information sources using the target processing framework to obtain a target extraction result; invoking the target large language model using the target double-layer integration framework of the cross-source information processing system to generate an information processing result corresponding to the cross-source information based on the target extraction result; where the information processing result includes fact summary, opinion analysis, and trend prediction.
[0070] In a specific implementation, information extraction performed based on the optimized target processing framework may include: intelligent classification using a large language model. The system constructs a classification prompt template to guide the large language model to analyze the technical features (such as API response format, data field structure), content features (such as theme distribution, language style, writing characteristics), and quality features (such as authority, timeliness, integrity) of the information source, and judges the exact type of the information source based on these comprehensive features. Dynamic policy selection is based on a policy mapping table. The mapping table is stored in a dictionary data structure, with the information source type as the key and the corresponding policy combination parameters as the value. The policy selection algorithm is: input information source type → query mapping table → obtain policy parameters → construct the final prompt template. Information extraction execution adopts a pipeline processing architecture. In the content acquisition stage, the original content of the information source is obtained through standard API calls, RSS parsing, etc., and the configuration includes authentication keys, request headers, timeout parameters, etc. In the preprocessing stage, a JSON / XML parsing library is used for data parsing, and regular expressions are used for text cleaning to remove irrelevant content such as advertisements and navigation bars.
[0071] When using large language models to perform actual information extraction tasks, an extraction prompt template can be dynamically constructed according to a policy combination to guide the large language model to extract key information from the original content. When the system calls the large language model, corresponding parameters are configured to ensure output stability and accuracy, including setting a lower randomness parameter to ensure result consistency, dynamically adjusting the output limit according to the input content length, and setting a specific end flag to ensure output integrity. The prompt template adopts a structured design, including system role definition, specific task description, input content to be processed, and expected output format requirements, to ensure the accuracy and consistency of the extraction results. In the post-processing stage, a multi-layer verification mechanism can be adopted. Format verification uses JSONschema to verify the correctness of the output format; logical verification checks the reasonableness of the numerical range (e.g., stock prices cannot be negative); consistency verification compares the consistency of the same information in different processing results, and when the difference exceeds the threshold, manual review is triggered.
[0072] In this embodiment, the target double-layer integration framework includes an information fusion layer and a knowledge generation layer, which can be used to achieve deep fusion and intelligent analysis of multi-source information, and integrate multi-source information into an organic knowledge system. The information fusion layer is responsible for processing the original information from different information sources, performing semantic understanding and content standardization through a large language model, identifying and resolving information conflicts, supplementing missing information, and converting heterogeneous information into a unified knowledge representation format. Based on the fused information, the knowledge generation layer performs in-depth analysis and reasoning through a large language model, identifies potential associations between information, mines implicit knowledge patterns, and generates structured knowledge outputs with insightful value, including various knowledge products such as fact summaries, opinion analyses, and trend predictions.
[0073] In this embodiment, the target double-layer integration framework of the cross-source information processing system is used to call the target large language model to generate an information processing result corresponding to the cross-source information based on the target extraction result. Specifically, it may include: The information fusion layer first performs semantic understanding processing. A dedicated large language model with a small number of parameters (such as a language model with a parameter scale of 0.8B) is used for named entity recognition and relationship extraction. Such models have faster inference speed and lower computational cost while ensuring the processing effect, and are more suitable for large-scale information processing scenarios. The system constructs a dedicated entity recognition and relationship extraction prompt template to guide the small-parameter model to identify key entities such as person names, organization names, and place names from the text, and identify the semantic relationships between entities, including subordinate relationships, cooperation relationships, competition relationships, etc. Finally, a structured knowledge graph is constructed to store this semantic information. The knowledge generation layer performs in-depth analysis and knowledge output. Association analysis uses a graph neural network algorithm to identify implicit relationships between entities. Pattern mining uses time series analysis algorithms to identify trend patterns, and anomaly detection algorithms to identify abnormal events. The final knowledge output can adopt a templated generation method. Taking an investment analysis report as an example, the output template includes standard chapters such as executive summary, market overview, individual stock analysis, risk warning, and investment advice. The large language model fills the template based on the integrated knowledge data to generate a structured investment analysis report. The multi-level knowledge integration architecture realizes true cross-source information understanding, enabling the information aggregation system to provide knowledge services with in-depth insight value.
[0074] In a specific implementation manner, the conflict detection algorithm is based on information similarity calculation. For numerical information, absolute difference comparison is used; for text information, semantic similarity calculation (using vector models such as bge-m3) is used. Information with a similarity exceeding the threshold but inconsistent content is marked as a conflict. Conflict resolution adopts a combination of a rule engine and a large language model. The rule engine processes simple conflicts (such as preferentially using information sources with higher authority); complex conflicts are submitted to the large language model for analysis, and the prompt template is: "The following information conflicts: [conflicting information]. Please make a judgment based on the authority, timeliness, and logical rationality of the information source and give the most likely correct information." Information supplementation adopts an active query mechanism. The system uses the large language model to identify missing key information, and the prompt template is: "Please analyze the integrity of the following information: [existing information]. Identify the missing key information items and evaluate the importance of the missing information." Based on the importance ranking of the missing information, the system automatically triggers supplementary queries to relevant information sources.
[0075] It can be seen that the cross-source information of multiple target information sources obtained in this application provides a rich training data basis for meta-learning. By using the target double-layer optimization architecture including the inner-loop architecture and the outer-loop architecture for meta-learning, the autonomous learning and continuous optimization of cross-source information processing capabilities are achieved. Among them, the inner-loop architecture focuses on improving the processing capabilities of specific information source types, and uses the target large language model to optimize the strategy corresponding to the case to be optimized to obtain the target optimized strategy, ensuring that the system can achieve the best processing effect on each known information source type and laying a foundation for building a high-quality information aggregation system; the outer-loop architecture focuses on learning the general processing mechanism across information source types, and identifies the core bottlenecks and key elements of cross-source information processing by analyzing the common processing problems existing in different information sources, so as to generate a general processing strategy (i.e., the target general strategy) that can improve cross-source processing capabilities. The outer loop is the core embodiment of meta-learning, enabling the cross-source information processing system to have strong generalization capabilities. The inner loop ensures the processing quality of the system on each known information source and provides high-quality learning data for the outer loop; the general mechanism learned by the outer loop guides the inner loop to better process specific information sources, enabling the cross-source information processing system to have the precise processing capabilities for known information sources, realizing the intelligent understanding and unified processing of heterogeneous information sources.
[0076] See Figure 2 As shown, this application discloses a cross-source information processing device based on meta-learning, which is applied to a cross-source information processing system and includes:
[0077] A case acquisition module 11, configured to acquire cases to be optimized corresponding to each target information source, and analyze the target information source by using a target large language model to determine the information processing characteristics and information processing challenges corresponding to each of the target information sources;
[0078] An inner-loop optimization module 12, configured to optimize the processing strategy corresponding to the case to be optimized by using the inner-loop optimization architecture of the target double-layer optimization architecture to determine the optimized processing strategy corresponding to the case to be optimized, and determine the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies;
[0079] An outer-loop optimization module 13, configured to perform aggregation processing on the information processing challenges corresponding to all the target information sources by using the outer-loop optimization architecture of the target double-layer optimization architecture to determine the target common challenges corresponding to all the target information sources, and determine the corresponding target general strategy based on the target common challenges by using the target large language model;
[0080] An information processing module 14 is configured to determine a target processing strategy corresponding to each target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, and determine a target processing framework based on all the target processing strategies, so as to process cross-source information obtained from all the target information sources by using the target processing framework;
[0081] Wherein, the target double-layer optimization architecture is a double-layer optimization architecture implemented based on meta-learning.
[0082] It can be seen that the cross-source information of multiple target information sources obtained in this application provides a rich training data basis for meta-learning. Meta-learning is performed by using a target double-layer optimization architecture including an inner loop architecture and an outer loop architecture, realizing the autonomous learning and continuous optimization of cross-source information processing capabilities. Among them, the inner loop architecture focuses on improving the processing capabilities of specific information source types, and uses a target large language model to optimize the strategy corresponding to the case to be optimized to obtain a target optimized strategy, ensuring that the system can achieve the best processing effect on each known information source type, laying a foundation for building a high-quality information aggregation system; the outer loop architecture focuses on learning general processing mechanisms across information source types, identifying the core bottlenecks and key elements of cross-source information processing by analyzing common processing problems existing in different information sources, so as to generate a general processing strategy (i.e., the target general strategy) that can improve cross-source processing capabilities. The outer loop is the core manifestation of meta-learning, enabling the cross-source information processing system to have strong generalization capabilities. The inner loop ensures the processing quality of the system on each known information source, providing high-quality learning data for the outer loop; the general mechanism learned by the outer loop guides the inner loop to better process specific information sources, enabling the cross-source information processing system to have the precise processing capabilities for known information sources, realizing the intelligent understanding and unified processing of heterogeneous information sources.
[0083] In a specific implementation manner, the device may further include:
[0084] An information source determination module, configured to determine each target information source;
[0085] An information source docking module, configured to dock the cross-source information processing system with each target information source by using a target application programming interface;
[0086] An information acquisition module, configured to acquire cross-source information to be processed from the target information source based on a target subscription method and a target information acquisition frequency.
[0087] In a specific implementation manner, the device may further include:
[0088] An information evaluation module, configured to perform information quality evaluation on the cross-source information obtained from each target information source based on a preset information scoring algorithm;
[0089] Among them, the scoring dimensions of the preset information scoring algorithm include information timeliness, information authority, and information integrity; the score corresponding to the information timeliness is a score determined based on the difference between the information release time of the cross-source information and the current time, the information authority is a score determined based on the authority level of the target information source corresponding to the cross-source information, and the information integrity is a score determined based on the text length and structural integrity of the cross-source information.
[0090] In a specific implementation manner, the inner loop optimization module 12 may specifically include:
[0091] A first template construction unit, configured to construct a first analysis prompt template by using the cross-source information processing system;
[0092] An analysis result acquisition unit, configured to batch input the cases to be optimized into the target large language model based on the first analysis prompt template to obtain corresponding target analysis results;
[0093] A strategy optimization unit, configured to optimize the processing strategy corresponding to the case to be optimized based on the target analysis result to determine the optimized processing strategy corresponding to the case to be optimized;
[0094] An optimization strategy determination unit, configured to score all the optimized processing strategies by using a preset automated test mechanism to determine the target optimized strategy corresponding to the case to be optimized.
[0095] In a specific implementation manner, the outer loop optimization module 13 may specifically include:
[0096] A knowledge challenge library establishment unit, configured to perform aggregation processing on the information processing challenges corresponding to all the target information sources based on the challenge types to establish a target cross-source challenge knowledge library corresponding to each challenge type;
[0097] A second template construction unit, configured to construct a second analysis prompt template by using the cross-source information processing system;
[0098] A common challenge determination unit, configured to analyze the cross-source challenge data in each target cross-source challenge knowledge library by using the target large language model based on the second analysis prompt template to determine the target common challenges corresponding to all the target information sources;
[0099] Among them, the challenge types include structured data extraction challenges, unstructured text understanding challenges, and multimodal content processing challenges.
[0100] In a specific implementation manner, the device may further include:
[0101] A processing strategy construction module, which is used to, if there is a new target information source, use a target few-shot learning algorithm to construct a corresponding target processing strategy for the new target information source based on the target general strategy.
[0102] In a specific embodiment, the information processing module 14 may specifically include:
[0103] An information extraction unit, which is used to uniformly extract cross-source information obtained from all the target information sources by using the target processing framework to obtain a target extraction result;
[0104] A processing result generation unit, which is used to call the target large language model by using the target double-layer integration framework of the cross-source information processing system to generate an information processing result corresponding to the cross-source information based on the target extraction result;
[0105] Among them, the information processing result includes fact summary, opinion analysis, and trend prediction.
[0106] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 3 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation to the scope of use of the present application.
[0107] Figure 3 It is a structural schematic diagram of an electronic device 20 provided by the embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the cross-source information processing method based on meta-learning disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0108] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0109] In addition, as a carrier for storing resources, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be transient storage or permanent storage.
[0110] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the meta-learning-based cross-source information processing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0111] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the foregoing disclosed meta-learning-based cross-source information processing method. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0113] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0114] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0115] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element.
[0116] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A cross-source information processing method based on meta-learning, characterized in that Applied to a cross - source information processing system, including: Obtain the cases to be optimized corresponding to each target information source, and use a target large - language model to analyze the target information source to determine the information processing characteristics and information processing challenges corresponding to each of the target information sources; Use the inner - loop optimization architecture of the target two - layer optimization architecture to optimize the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determine the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies; Use the outer - loop optimization architecture of the target two - layer optimization architecture to aggregate the information processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources, and use the target large - language model to determine the corresponding target general strategy based on the target common challenges; Determine the target processing strategy corresponding to the target information source based on the target optimized strategy and the target general strategy corresponding to each of the target information sources, and determine the target processing framework based on all the target processing strategies, so as to use the target processing framework to process the cross - source information obtained from all the target information sources; Wherein, the target two - layer optimization architecture is a two - layer optimization architecture implemented based on meta - learning.
2. The cross-source information processing method based on meta-learning according to claim 1, wherein Before obtaining the cases to be optimized corresponding to each target information source, it further includes: Determine each target information source; Use a target application programming interface to dock the cross - source information processing system with each of the target information sources; Obtain the cross - source information to be processed from the target information source based on a target subscription method and a target information acquisition frequency.
3. The cross-source information processing method based on meta-learning according to claim 1, wherein, It further includes: Evaluate the information quality of the cross - source information obtained from each of the target information sources based on a preset information scoring algorithm; Wherein, the scoring dimensions of the preset information scoring algorithm include information timeliness, information authority, and information integrity; the score corresponding to information timeliness is a score determined based on the difference between the information release time of the cross - source information and the current time, the information authority is a score determined based on the authority level of the target information source corresponding to the cross - source information, and the information integrity is a score determined based on the text length and structural integrity of the cross - source information.
4. The cross-source information processing method based on meta-learning according to claim 1, wherein, The step of optimizing the processing strategy corresponding to the case to be optimized to determine the optimized processing strategy corresponding to the case to be optimized, and determining the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies, includes: Use the cross - source information processing system to construct a first analysis prompt template; Batch - input the case to be optimized into the target large - language model based on the first analysis prompt template to obtain the corresponding target analysis result; Optimize the processing strategy corresponding to the case to be optimized based on the target analysis result to determine the optimized processing strategy corresponding to the case to be optimized; Use a preset automated test mechanism to score all the optimized processing strategies to determine the target optimized strategy corresponding to the case to be optimized.
5. The cross-source information processing method based on meta-learning according to claim 1, wherein Aggregating the information processing challenges corresponding to all the target information sources to determine the target common challenges corresponding to all the target information sources, including: Aggregating the information processing challenges corresponding to all the target information sources based on the challenge type to establish a target cross-source challenge knowledge base corresponding to each challenge type; Using the cross-source information processing system to construct a second analysis prompt template; Analyzing the cross-source challenge data in each target cross-source challenge knowledge base using the target large language model based on the second analysis prompt template to determine the target common challenges corresponding to all the target information sources; Wherein, the challenge types include structured data extraction challenges, unstructured text understanding challenges, and multi-modal content processing challenges.
6. The cross-source information processing method based on meta-learning according to claim 1, wherein It also includes: If there is a new target information source, using the target few-shot learning algorithm to construct the target processing strategy corresponding to the new target information source based on the target general strategy.
7. The cross-source information processing method based on meta-learning according to any one of claims 1 to 6, characterized in that The processing of the cross-source information obtained from all the target information sources using the target processing framework includes: Uniformly extracting the cross-source information obtained from all the target information sources using the target processing framework to obtain a target extraction result; Invoking the target large language model using the target double-layer integration framework of the cross-source information processing system to generate an information processing result corresponding to the cross-source information based on the target extraction result; Wherein, the information processing result includes fact summary, opinion analysis, and trend prediction.
8. A cross-source information processing device based on meta-learning, characterized in that, Applied to a cross-source information processing system, including: A case acquisition module, configured to acquire the cases to be optimized corresponding to each target information source, and analyze the target information sources using the target large language model to determine the information processing characteristics and information processing challenges corresponding to each target information source; An inner loop optimization module, configured to optimize the processing strategy corresponding to the case to be optimized using the inner loop optimization architecture of the target double-layer optimization architecture to determine the optimized processing strategy corresponding to the case to be optimized, and determine the target optimized strategy corresponding to the case to be optimized based on all the optimized processing strategies; An outer loop optimization module, configured to aggregate the information processing challenges corresponding to all the target information sources using the outer loop optimization architecture of the target double-layer optimization architecture to determine the target common challenges corresponding to all the target information sources, and determine the corresponding target general strategy using the target large language model based on the target common challenges; An information processing module, configured to determine the target processing strategy corresponding to the target information source based on the target optimized strategy and the target general strategy corresponding to each target information source, and determine a target processing framework based on all the target processing strategies, so as to process the cross-source information obtained from all the target information sources using the target processing framework; Wherein, the target double-layer optimization architecture is a double-layer optimization architecture implemented based on meta-learning.
9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the cross-source information processing method based on meta-learning according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein when the computer program is executed by a processor, it implements the cross-source information processing method based on meta-learning according to any one of claims 1 to 7.