A method, system, device and medium for strengthening the connection of heterogeneous A / B data sources
By performing semantic analysis and vectorization on heterogeneous A/B data sources, and combining them with large model technology for information retrieval and association analysis, the integration challenges of heterogeneous data sources are resolved, efficient data fusion and in-depth interpretation are achieved, and accurate trend forecasts and insights are provided.
Patent Information
- Application Number
- CN202510846200.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing technologies make it difficult to efficiently integrate and deeply mine heterogeneous A/B data sources, resulting in information dispersion, inefficient fusion, lack of in-depth interpretation and prediction capabilities, and difficulty in supporting cross-industry analysis.
By receiving natural language queries, performing semantic parsing and vectorization processing, combining large model technology for information retrieval and similarity matching, establishing cross-source data associations, and using artificial intelligence models for in-depth analysis, we can predict technology trends and business opportunities.
It achieves deep integration across heterogeneous data sources, provides accurate technical insights and forward-looking trend predictions, breaks down data silos, and improves the efficiency and depth of information retrieval and analysis.
Smart Images

Figure CN120353835B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer data processing technology, and in particular to an artificial intelligence-based method for information retrieval, analysis, association, and trend prediction of heterogeneous data sources, specifically a method, system, device, and medium for strengthening the connection of heterogeneous A / B data sources. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, data-driven decision-making and innovation are becoming increasingly important across industries. In this context, effectively integrating and leveraging heterogeneous data sources from diverse sources and structures—for example, Class A data sources (such as structured technical literature databases and internal case libraries) and Class B data sources (such as unstructured industry reports and external market dynamics data)—has become crucial for enhancing the depth and breadth of analysis. However, traditional methods for acquiring and processing information from these heterogeneous A / B data sources have numerous shortcomings in achieving efficient integration and deep fusion, specifically:
[0003] 1. Barriers and integration challenges between A / B data sources: Traditional methods often focus on the independent application of a single data source (A or B), resulting in information fragmentation between Category A and Category B data. This makes it difficult to achieve effective cross-source connectivity and comprehensive integration, creating de facto "data silos" and limiting the potential for comprehensive analysis.
[0004] 2. Inefficient heterogeneous A / B data fusion processing: Existing technical tools rely on static matching or shallow association for Class A and Class B data sources with different structures, formats, and semantics. This results in inefficiency in cross-source data alignment, conversion, and fusion, making it difficult to deeply explore the complex internal connections and complementary value between A and B data, hindering their value-added utilization after integration.
[0005] 3. Lack of in-depth interpretation and prediction capabilities based on integrated data: Due to the inability to effectively connect and integrate heterogeneous A / B data sources, existing technologies often lack comprehensive insights based on AI deep learning and big data analysis. They are unable to conduct in-depth interpretation of integrated complex industry cases and multi-source technical data, and it is even more difficult to make accurate trend forecasts.
[0006] 4. Insufficient interactivity and personalization in post-connection data exploration: Even after initial A / B data aggregation is achieved, traditional platforms often provide static, preset query and analysis modes, making it difficult to support users in conducting instant, customized exploratory analysis and interactive feedback based on the post-connection A / B data, thus reducing the efficiency of data value discovery.
[0007] 5. Lack of cross-industry A / B data integration and analysis capabilities: When Category A and Category B data sources belong to different industries or fields, due to the lack of an effective integration mechanism and support for cross-domain knowledge graphs, existing platforms find it difficult to integrate and analyze cross-industry technology development trends and potential cross-industry opportunities.
[0008] Therefore, there is an urgent need to develop a method for strengthening the integration of heterogeneous A / B data sources, aiming to break down the barriers between different data sources, improve the efficiency and depth of heterogeneous data fusion processing, and enable in-depth interpretation and intelligent prediction based on multi-source data, so as to solve the defects and shortcomings of the above-mentioned existing technologies in integrating and utilizing heterogeneous A / B data sources. Summary of the Invention
[0009] Based on this, the present invention aims to solve the many shortcomings of traditional information acquisition and processing methods for heterogeneous A / B data sources, as well as the many deficiencies in achieving efficient connection and deep integration, and to provide a heterogeneous A / B data source connection enhancement method, system, equipment and medium.
[0010] In a first aspect, the present invention provides a method for strengthening the connectivity of heterogeneous A / B data sources, comprising the following steps:
[0011] Receive a natural language query request input by a user, and perform semantic parsing and vectorization processing on the natural language query request;
[0012] Based on the results of the semantic analysis and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results;
[0013] Performing data source association analysis on the information from different heterogeneous data sources in the search results to establish a traceable association relationship between cross-source data;
[0014] Use artificial intelligence models to conduct in-depth analysis of the data after the correlation analysis to extract key information and identify potential technology trends and business opportunities;
[0015] Based on the results of the in-depth analysis, we will develop trends and predict technology development directions and business opportunities;
[0016] The output includes structured analysis results of original search results, association relationships and trend predictions.
[0017] As an optional implementation of the first aspect of the present application, in the step of receiving a natural language query request input by a user and performing semantic parsing and vectorization processing on the query request, the semantic parsing and vectorization processing on the natural language query request includes: using a pre-trained language model to semantically parse the query text and extract the query intent; and converting the query text into a semantic vector through an embedding model.
[0018] As an optional implementation of the first aspect of the present application, according to the results of the semantic parsing and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources. In the step of obtaining the retrieval results, the information retrieval and similarity matching include: parallel execution of keyword retrieval based on the inverted index and approximate nearest neighbor search of the vector database based on the semantic vector; and dynamic weighted fusion processing of the matching score of the keyword retrieval and the vector semantic similarity score to generate the retrieval result with comprehensive ranking.
[0019] As an optional implementation of the first aspect of the present application, data source association analysis is performed on the information from different heterogeneous data sources in the search results, and in the step of establishing a traceable association relationship between cross-source data, the data source association analysis includes: extracting technical features through named entity recognition, constructing a technology-case association graph; and implementing classification system mapping based on a predefined rule base, and establishing a bidirectional traceable association network between industry cases, technology nodes and technical documents.
[0020] As an optional implementation of the first aspect of the present application, it also includes: when querying a certain industry case, automatically expanding the technical information related to the industry case according to the bidirectional traceable association network; when querying certain technical information, automatically expanding the industry case related to the technical information according to the bidirectional traceable association network.
[0021] As an optional implementation method of the first aspect of this application, in the step of using an artificial intelligence model to conduct in-depth analysis on the data after the association analysis, extract key information and identify potential technological trends and business opportunities, the in-depth analysis using the artificial intelligence model includes: extracting technical entities and business elements through a large language model; analyzing the mapping relationship between the technology evolution path and the market application scenario; and generating a structured analysis report including technology maturity assessment, commercialization path planning, title, future outlook, solutions and technology selection.
[0022] As an optional implementation manner of the first aspect of the present application, in the step of performing trend expansion based on the results of the in-depth analysis and predicting the direction of technological development and business opportunities, the trend expansion includes: adopting a three-level transfer learning framework, the three-level transfer learning framework including: near-level migration based on similarity matching of cases in the same industry, middle-level migration based on the co-occurrence network of technical terms to discover cross-sub-industry associations, and far-level migration based on the attention mechanism to capture cross-domain technical features; and based on the analysis results of the three-level transfer learning framework, forming a three-stage trend forecast report including short-term, medium-term and long-term.
[0023] In a second aspect, an embodiment of the present application provides a heterogeneous A / B data source interoperability enhancement system, including:
[0024] An intelligent retrieval module is configured to receive natural language query requests input by users, perform semantic parsing and vectorization processing on the natural language query requests, and perform information retrieval and similarity matching from multiple heterogeneous data sources based on the results of the semantic parsing and vectorization processing to obtain retrieval results;
[0025] A data source association module is used to perform data source association analysis on the information from different heterogeneous data sources in the search results and establish a traceable association relationship between cross-source data;
[0026] An AI deep analysis module is used to use an artificial intelligence model to perform in-depth analysis on the data after the correlation analysis, extract key information, and identify potential technology trends and business opportunities;
[0027] The trend expansion module is used to expand trends based on the results of the in-depth analysis, predict technology development directions and business opportunities, and output structured analysis results including original search results, association relationships and trend predictions.
[0028] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0029] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0030] Compared with the existing technology, the present invention proposes a method for strengthening the connection of heterogeneous A / B data sources, which has the following beneficial effects:
[0031] (1) Improve search accuracy and depth: By combining natural language processing and big model technology, it can accurately understand user query intentions, conduct multi-dimensional search and information fusion, deeply explore potential connections in technical information, overcome barriers between different data sources, and obtain more accurate and relevant technical solutions and industry cases.
[0032] (2) Breaking down data silos: Through the data source association module, information from different heterogeneous data sources can be accurately matched and integrated, achieving deep integration between data sources and providing users with technical insights from a global perspective.
[0033] (3) Provide forward-looking trend forecasts: Use large models to conduct in-depth data analysis, combine historical industry data with existing trends, accurately identify technology development trends, market demand changes and potential opportunities, and provide users with forward-looking industry trend forecasts.
[0034] (4) Realize mutual reinforcement of information: By building bidirectional associations between heterogeneous data sources, querying one type of data (such as industry cases) can inversely lead to related data of another type (such as technical literature), and vice versa, thus achieving mutual verification and value enhancement of information. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a structural diagram of a heterogeneous A / B data source connection and mutual reinforcement system of the present invention.
[0036] Figure 2 This is a flow chart of a method for connecting and mutually reinforcing heterogeneous A / B data sources according to the present invention. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the application can be implemented in a sequence other than those illustrated or described here. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated before and after are in a kind of "or" relationship. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically limited.
[0039] It should be noted that the "module" described in the embodiments of the present invention can be a software unit, a hardware unit, or a combination of the two. For a software unit, it can be stored in a computer-readable storage medium and realize the corresponding function when executed.
[0040] Example 1: System-based implementation
[0041] Reference Figure 1 The present invention provides a heterogeneous A / B data source integration and enhancement system for implementing the method of Example 2 below. The system may include: an intelligent retrieval module 101, an AI in-depth analysis module 102, a data source association module 103, and a trend expansion module 104. These modules work together to implement a complete process from user query to trend prediction.
[0042] Intelligent Retrieval Module 101:
[0043] This module is responsible for receiving natural language queries input by users (for example, “How to improve the conversion efficiency of solar cells?”).
[0044] First, the query question is vectorized and parsed using a pre-trained large-scale language model (for example, calling the text-embedding-3-small model API: response = await client.embeddings.create(input=text, model='text-embedding-3-small')) to convert the text into a vectorized response.
[0045] Then, based on the vectorized semantic analysis results, hybrid retrieval technology is used to perform information retrieval in multiple heterogeneous databases (such as data source A: industry case library, data source B: technical literature library). Hybrid retrieval includes:
[0046] 1. Keyword search: Extract keywords (such as "solar cell", "conversion efficiency", "improve") from the query and match them in the case database and literature database.
[0047] 2. Semantic Search: Based on the vectorized result of the query, response (i.e., goal_vec), search for semantically similar content in the vector database. Matching is performed by calculating the cosine similarity between the query vector and the stored vector stored_vec in the database: similarity = np.dot(response, stored_vec) / (np.linalg.norm(goal_vec)np.linalg.norm(stored_vec)).
[0048] Finally, the module integrates the retrieved information, sorting it based on a weighted fusion of keyword matching scores and semantic similarity scores, and returns it to the user. (For example, the top result might be an academic paper on the application of novel perovskite materials in solar cells.) The second-ranked result might be a patent describing a technology for improving solar cell efficiency through multi-layer anti-reflective coatings. The third-ranked result might be an industry case study that describes a company improving solar cell performance by optimizing its manufacturing process.
[0049] Intelligent retrieval module 101 is responsible for receiving natural language queries entered by users through the user layer and performing vectorized parsing of the queries using a pre-trained large-scale language model to extract the query intent. The specific implementation steps are as follows: The pre-trained language model is used to convert the query text into a high-dimensional vector representation. Based on the semantic understanding results, retrieval technology is used to perform similarity matching within the vector database to retrieve relevant industry cases and technical information. A sorting algorithm (comparing vector similarities) is used to optimize the order of the search results to ensure the most relevant information is returned.
[0050] AI Deep Analysis Module 102:
[0051] This module receives the retrieved industry cases and technical information output by the intelligent retrieval module 101 and passes it to the customized "Industry and Technology Analysis Agent" of this invention. It extracts key information and identifies potential technology trends and business opportunities. It also conducts a comprehensive evaluation of the case, analyzing its market prospects, technical feasibility, and potential application areas, providing users with comprehensive decision support. The following is relevant information about the agent:
[0052] Role: A seasoned case analyst, skilled at extracting key insights from complex cases, crafting engaging headlines, future perspectives, and proposing innovative solutions. Conduct in-depth analysis based on the context and details of specific cases, and accurately select relevant AI technologies.
[0053] Skill 1: Information Extraction: Carefully read the case study, capture the core information, and ensure the title is attractive and expressive to spark reader interest.
[0054] Skill 2: Future Outlook. Thoroughly analyze the case content and technologies used, and make detailed predictions about future development trends, including timeframes, potential market impacts, and specific manifestations of technological advancement.
[0055] Skill 3: Solution-Based Approach. Thoroughly analyze the problems and challenges in the case, provide a comprehensive root cause analysis, and propose practical solutions based on the actual situation. Detail the implementation steps and use successful case studies to enhance persuasiveness.
[0056] Skill 4: Technology Selection: Select 3 to 5 technologies from the technology list that match the case. Ensure that each technology is selected with sufficient reasons and explain in detail its applicability and how it is applied to the case.
[0057] Output format: Please organize the output in JSON format.
[0058] For example, suppose an industry case is retrieved: "A company has developed a new type of solar cell based on perovskite material with a conversion efficiency of 28% and plans to commercialize it within the next two years."
[0059] The workflow of the AI in-depth analysis module 102 includes:
[0060] 1. Automated analysis and key information extraction
[0061] The AI deep analysis module 102 uses an agent to analyze the case and extract the following key information: 1. Core technology: perovskite materials, solar cells, 28% conversion efficiency. 2. Commercialization plan: Commercialization within the next two years. 3. Potential advantages: high efficiency, low cost, and environmental friendliness.
[0062] 2. Identification of technology trends and business opportunities
[0063] The module then uses semantic analysis and large-scale model networking technology to identify the following trends and opportunities: 1. Technology Trend: The application of perovskite materials in the photovoltaic field is rapidly developing and may become the mainstream technology for next-generation solar cells. 2. Business Opportunity: High-efficiency solar cells have broad application prospects in distributed energy, electric vehicle charging stations, and other fields.
[0064] 3. Comprehensive assessment and decision support
[0065] The final module provides a comprehensive evaluation of the case study, including: 1. Market Prospects: The global solar cell market is projected to grow at an average annual rate of 15% over the next five years, with perovskite cells poised to capture a significant share. 2. Technical Feasibility: Perovskite materials have a high level of technical maturity, but stability and large-scale production challenges remain. 3. Application Areas: Suitable for rooftop photovoltaics, portable charging devices, space energy, and other fields.
[0066] 4. Generate headlines, future outlook, solutions and technology options
[0067] Based on the analysis results, the module generates the following output in JSON:
[0068] {"title": "Perovskite solar cells: A dual breakthrough in high efficiency and commercial prospects",
[0069] "future_view": "Perovskite solar cells are expected to achieve large-scale commercialization within the next five years and are expected to capture 20% of the global photovoltaic market by 2028. Their high efficiency (over 30%) and low cost will drive the rapid development of distributed energy and electric vehicle charging infrastructure. Technological progress will focus on material stability and production process optimization."
[0070] To accelerate the commercialization of perovskite solar cells, we recommend the following steps: 1. Strengthen material stability research and develop new packaging technologies through collaboration with universities and research institutions. 2. Optimize production processes and introduce automated production lines to reduce costs. 3. Expand application scenarios, such as collaborating with electric vehicle manufacturers to develop integrated photovoltaic roofs. Success Case Study: A company increased the lifespan of perovskite cells from 1,000 hours to 5,000 hours by improving packaging technology, laying the foundation for commercialization.
[0071] "technologies": "Knowledge Graph, Transformer Infrastructure, Semantic Web, End-to-End Learning, Image Processing"}
[0072] In AI Deep Analysis Module 102, large-scale model analysis is used to automatically analyze retrieved industry cases and technical information, extracting key information and identifying potential technological trends and business opportunities. This module then conducts a comprehensive assessment of the case, analyzing its market prospects, technical feasibility, and potential application areas, providing comprehensive decision support for users. The specific implementation steps are as follows: Large-scale models (such as GPT-4 and DeepSeek) are used to deeply analyze the text information in the search results, extracting key technical elements, application scenarios, and industry impact, and analyzing future development trends and technical solutions.
[0073] Data source association module 103:
[0074] This module uses a dual-channel dynamic association mechanism to accurately match information from different heterogeneous data sources (such as technical literature, industrial cases, etc.). It uses the intelligent extraction channel of technical features: context-aware technology recognition based on the large language model (LLM), and the case feature mapping channel: multi-dimensional analysis of structured industrial cases to achieve bidirectional traceable association between technical literature and industrial cases, solving the data silo problem of technology industrialization analysis in traditional methods.
[0075] The implementation process includes:
[0076] 1. Feature extraction of heterogeneous data:
[0077] For the industrial case (data source A), accurate extraction of technical features is achieved through multiple rounds of prompt engineering. The pseudo code is as follows:
[0078] - Title: {title}
[0079] - Solution: {solution}
[0080] - Future Outlook: {future_view}”
[0081] Output example: "Artificial neural network, convolutional neural network, deep learning".
[0082] 2. For technical documents (data source B, such as patents), document parsing and vector retrieval are combined to vectorize the technical documents. The pseudo code is as follows:
[0083] soup = BeautifulSoup(description, 'html.parser')
[0084] text = str()
[0085] for p in soup.find_all('p'):
[0086] text += p.text
[0087] searches = await vector_retrieval(text)
[0088] Where text is "technical document subject: {title} technical document classification: {classification} content: {text}"; searches is the vectorized result of the technical document.
[0089] 3. Dynamic association construction:
[0090] Association rule engine: Achieve precise matching of technology names. Technology classification tree mapping (through the classification field) can accurately assign technology categories and use semantic similarity to assist in technical document matching (through vectorized results).
[0091] Relational database structure:
[0092] A[Industry Case] -- applies --> B((Technology Node)),
[0093] B -- improves -->C[patent document],
[0094] C -- cited_by -->A
[0095] 4. Two-way query service: Through the above steps, when querying a specific industry case, the module will automatically expand the relevant technical information of the industry case based on the above associations. When querying a specific technical information, the module will also automatically expand the relevant industry cases of the technical information based on the above associations.
[0096] For example, assume the following heterogeneous data sources: an "industry case": a company successfully developed and commercialized perovskite solar cells; and a "patent document": a patent on "perovskite material packaging technology."
[0097] First, a relational database was constructed to extract keywords from industrial cases: perovskite, solar energy, commercialization, and from patent documents: packaging technology, perovskite, production process.
[0098] Secondly, precise matching and association: The module uses the above technology to accurately match information from different data sources: associating "perovskite" in patent documents with "perovskite" in industrial cases.
[0099] Finally, the module uses data fusion technology to integrate data from various sources into a unified knowledge base, eliminating information silos and ensuring users receive comprehensive and multi-dimensional technical information. For example, when a user searches for "perovskite solar cells," the module returns the following related information: Industry Case Studies: Successful Commercial Applications; Patent Documents: Innovations in Perovskite Encapsulation Technology.
[0100] In the data source association module 103, the module constructs a relational database to accurately match information from different heterogeneous data sources (such as technical literature and industry cases), achieving deep connections between technologies and application cases. Data fusion technology is used to integrate data from different sources, eliminating information silos and ensuring that users receive comprehensive and multi-dimensional technical information. The specific implementation steps are as follows: Key information is extracted from multiple heterogeneous data sources to construct a database containing multi-level relationships such as technical information and industry cases. Large model technology is used to discover potential connections between technologies and cases and optimize the matching degree. Data fusion technology is used to integrate data from different sources, eliminating information silos and ensuring that users receive comprehensive and multi-dimensional technical information. The innovations of this module are reflected in: 1. Dynamic feature extraction mechanism: LLM prompt engineering is used to achieve context-sensitive technology identification; 2. Bidirectional association indexing: A three-level traceability network is established from case → technology → patent.
[0101] Trend Expansion Module 104:
[0102] This module pioneers a "three-level migration prediction model," quantifying the results from Data Source Association Module 103 through technology correlation metrics. This model uses dynamic weighting based on the correlation between industry cases and technology information. A cross-industry migration engine predicts technology migration across three industry scenarios: near-, mid-, and long-term. Multi-dimensional trend verification leverages industry data from similar peers and cross-industry cases to analyze their performance and predict future technology trends, helping users identify potential industry innovation opportunities. By combining multi-dimensional analysis of historical technology cases and current industry performance with market dynamics, users are provided with long- and short-term technology trend forecasts to assist in enterprise strategic planning.
[0103] The implementation process includes:
[0104] Phase 1: Industry Case Feature Extraction: Extract the core technology label "technology" from the target case and determine its position in the industry knowledge graph based on industry classification.
[0105] Phase 2: Level 3 case migration.
[0106] Near migration: Look for cases applying similar technologies (such as dynamic pricing on food delivery platforms) within the same sub-category (for example, if the target case is logistics route optimization, then in the field of terminal delivery).
[0107] Mid-term migration: Look for cases where similar technologies are applied within different subcategories of the same industry (for example, within the logistics industry but in different links) (e.g., air cargo loading optimization).
[0108] Far transfer: Find use cases for applying similar technologies across industries (e.g., healthcare) (e.g., hospital operating room scheduling).
[0109] Phase 3: Intelligent Trend Generation. Leveraging a large model (e.g., a carefully designed system_prompt = "As a trend analysis expert, please generate a report based on the following dimensions: - Industry Status Analysis - Success Case Lessons Learned... - Cross-Industry Collaboration") combined with the results of the three-level migration analysis, a trend forecast report is generated that includes the current state of the industry, lessons learned from success cases, innovation strategy recommendations, potential risks, future prospects, and cross-industry collaboration recommendations.
[0110] For example:
[0111] 1. Enter the case study: "A logistics company used an AI-powered route optimization algorithm (technology: reinforcement learning + spatiotemporal prediction) to achieve a 22% improvement in delivery efficiency."
[0112] Industry classification: Logistics → Terminal delivery → Route planning
[0113] 2. Level 3 Migration Case Matching Results
[0114]
[0115] 3. Generate key paragraphs for trend reports
[0116] (1) Analysis of the current status of the industry:
[0117] Currently, the application of artificial intelligence in the logistics industry has entered a period of rapid growth. According to a 2023 Deloitte report, 85% of leading logistics companies have deployed AI route optimization systems, and the industry's average delivery efficiency has increased by 17.5% (your case is 22% better than the average).
[0118] (2) Lessons from successful cases
[0119] graph LR
[0120] A [Logistics Path Optimization] --> | Near Migration | B [Dynamic Pricing for Takeout]
[0121] A -->|Mid-term Migration| C[Air Loading Optimization]
[0122] A -->|Far Migration| D[Operating Room Scheduling]
[0123] Recent migration case: Dynamic pricing system of a food delivery platform
[0124] Key to success: Reusing 85% of the underlying logic of the path optimization algorithm
[0125] Replicable Strategy: Add Price Sensitivity Parameters to Your System
[0126] Remote Migration Case: A Hospital Operating Room Scheduling System
[0127] Lessons learned: Medical scenarios verify the robustness of algorithms in complex scenarios
[0128] Adaptation suggestion: Develop a delivery priority rule configuration panel
[0129] (3) Innovation strategy recommendations
[0130] Short-term goal (within 1 year): Establish a "path optimization algorithm supermarket" and encapsulate 3 preset modes: cost priority mode, time priority mode
[0131] Balanced mode.
[0132] Long-term goal (3-5 years): Develop a cross-industry technology adaptation engine to support: logistics → medical scenario conversion (requires a 120-day development cycle)
[0133] Logistics → retail scenario conversion (requires a 90-day development cycle).
[0134] (4) Potential risks and responses:
[0135] Risk Matrix:
[0136]
[0137] Key Warning: Urban low-speed delivery scenarios require special attention to data collection on non-motorized vehicle lanes (referring to the Shanghai pilot experience)
[0138] (5) Outlook for future trends
[0139] Technology integration direction: Edge computing devices will optimize paths in real time (reducing cloud dependence), and digital twin technology will enable full-scenario simulation testing.
[0140] Sustainable development recommendations: Develop a "green path model," prioritizing routes around energy vehicle charging stations and low-emission traffic control areas.
[0141] (6) Cross-industry collaboration recommendations
[0142] List of collaboration opportunities:
[0143]
[0144] In the trend expansion module 104, this module combines industry data from similar peers and cross-industry technology cases, analyzes their performance, predicts future technology development trends, and helps users identify potential industry innovation opportunities. The specific implementation steps are as follows: The collected multi-source data is normalized and feature selection methods are used to extract features valuable for trend prediction. Industry cases and technology information from different fields are integrated, and cross-industry data integration and comparative analysis are conducted to identify potential cross-border applications and innovation opportunities for technologies. Based on the output of the large model, the system automatically generates a report on future industry development trends, covering technology development directions, changes in market demand, and potential business opportunities.
[0145] Example 2: Method Flow
[0146] Reference Figure 2 The present invention provides a method for strengthening the connection of heterogeneous A / B data sources, which specifically includes the following steps:
[0147] Step S201: receiving a natural language query request input by a user, and performing semantic parsing and vectorization processing on the natural language query request.
[0148] The system receives a natural language query (e.g., "How can I improve the conversion efficiency of solar cells?") entered by the user through the user interface. It uses a pre-trained language model (such as BERT or GPT architecture) to perform semantic parsing on the query text to understand the user's intent. Subsequently, an embedding model (such as text-embedding-3-small) is used to convert the query text into a 768-dimensional semantic vector. This process can be implemented by calling the corresponding embedding API (such as OpenAI's Embeddings API). For example, if the user enters text = "How can I improve the conversion efficiency of solar cells?", calling client.embeddings.create(input=text, model='text-embedding-3-small') will generate the vectorized response.
[0149] Step S202: Based on the results of the semantic analysis and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results.
[0150] The system performs two retrieval processes in parallel:
[0151] Keyword search based on inverted index: Extract keywords (such as "solar cell", "conversion efficiency", "improve") from the query request, search for matching documents in a pre-built index containing industry cases (data source A) and technical literature (data source B), and calculate the matching score between the terms and documents.
[0152] Semantic retrieval based on a vector database: Using the query semantic vector generated in step S201, perform an approximate nearest neighbor (ANN) search within a pre-built vector database (which stores semantic vectors of industry cases and technical literature). Calculate the similarity score between the query vector and the stored vectors in the database using cosine similarity. For example, similarity = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec)np.linalg.norm(stored_vec)).
[0153] Search results are processed using a dynamic weighted fusion algorithm. For example, the comprehensive score = α * semantic similarity score + (1-α) keyword matching score, where α is an adjustable parameter (e.g., 0.6 by default). Search results are sorted based on the comprehensive score to generate a comprehensive ranking list.
[0154] Step S203: performing data source association analysis on the information from different heterogeneous data sources in the search results, and establishing a traceable association relationship between cross-source data.
[0155] Using bidirectional association engine technology, the heterogeneous data in the search results obtained in step S202 are associated:
[0156] Construction of technology-case association graph: Extract key technical features (such as "perovskite" and "packaging technology") from industry cases and technical literature through named entity recognition (NER).
[0157] Classification system mapping and association network construction: Based on a predefined rule base (e.g., a technology classification tree), extracted technical features are mapped to corresponding technology nodes. This establishes a traceable association between industry cases (A), technology nodes (B), and patents / technical documents (C), e.g., A applies --> B, B improves --> C, C cited by --> A.
[0158] Through this step, if an industrial case "A company successfully developed and commercialized perovskite solar cells" is retrieved, the system can associate it with the patent document "A patent on 'perovskite material packaging technology'" through the common technology node "perovskite".
[0159] Step S204: Use an artificial intelligence model to conduct in-depth analysis on the data after the association analysis, extract key information and identify potential technology trends and business opportunities.
[0160] Deploy a multi-task analysis model (usually based on a large language model such as GPT-4 or a custom agent) to analyze the associated data:
[0161] 1. Entity and element extraction: Extract core technology entities and business elements (such as market size, competitors, and business models) from industry cases and technical literature.
[0162] 2. Evolution and Application Mapping: Analyze the mapping relationship between the evolution path of technology (such as from laboratory research to commercial application) and market application scenarios (such as rooftop photovoltaics and portable devices).
[0163] 3. Report Generation: A structured analysis report is ultimately generated, which may include technology maturity assessment, commercialization path planning, an engaging title, future outlook, solution recommendations, and relevant technology selections. The output format is JSON, as described in the AI in-depth analysis module 102 in Example 1.
[0164] Step S205: Based on the results of the in-depth analysis, trend expansion is performed to predict technology development directions and business opportunities.
[0165] A three-level transfer learning framework is used to expand the analysis results of step S204 and predict technology trends:
[0166] 1. Near-layer transfer: Based on similarity matching (e.g., cosine similarity > 0.85) of cases in the same industry (e.g., solar cell industry), analyze the application and performance of similar technologies in similar scenarios.
[0167] Mid-level migration: Discover technology associations across sub-industries (e.g., from photovoltaic power generation to energy storage technology) through the co-occurrence network of technology terms, and analyze the migration potential of technology in related sub-industries.
[0168] 2. Long-Layer Transfer: Leverage deep learning models such as attention mechanisms to capture similarities in technical features across domains (e.g., from energy materials to biomedical materials) and explore long-range technology transfer opportunities.
[0169] 3. Based on this three-level migration analysis, combined with market data and expert knowledge, a three-stage trend forecast report is generated, covering short-term (1-3 years), medium-term (3-5 years), and long-term (5+ years). The report content includes the "Industry Status Analysis" and "Success Case Studies" sections described in the trend expansion module 104 of Example 1.
[0170] Step S206: Outputting structured analysis results including original search results, association relationships, and trend predictions.
[0171] The system integrates the processing results and returns a structured result set to the user. This result set contains three main dimensions:
[0172] 1. Original search result list (sorted).
[0173] 2. Associated knowledge graph (displays the association relationship between data, with better visualization).
[0174] 3. In-depth analysis and trend forecast report (including the output of AI in-depth analysis module 102 and trend expansion module 104).
[0175] In this way, a complete service closed loop from basic information retrieval to advanced decision support is achieved.
[0176] In the embodiments of the present application, a heterogeneous A / B data source interoperability enhancement system can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, network attached storage (NAS), personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.
[0177] In the embodiments of the present application, a heterogeneous A / B data source interoperability enhancement system may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0178] The embodiment of the present application provides a heterogeneous A / B data source connection and strengthening system that can achieve Figure 1 In the method embodiment, each process of implementing a heterogeneous A / B data source penetration enhancement method is not described here to avoid repetition.
[0179] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned embodiment of the heterogeneous A / B data source connection enhancement method is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0180] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned heterogeneous A / B data source connection enhancement method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0181] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.
[0182] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0183] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.
[0184] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A method for strengthening the connection of heterogeneous A / B data sources, characterized in that: The following steps are involved: Receive a natural language query request input by a user, and perform semantic parsing and vectorization processing on the natural language query request; Based on the results of the semantic analysis and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results; Performing data source association analysis on the information from different heterogeneous data sources in the search results to establish traceable association relationships across source data; specifically, extracting technical features through named entity recognition to construct a technology-case association map; and implementing classification system mapping based on a predefined rule base to establish a bidirectional traceable association network between industry cases, technology nodes, and technical documents; Use artificial intelligence models to conduct in-depth analysis of the data after the correlation analysis to extract key information and identify potential technology trends and business opportunities; Based on the results of the in-depth analysis, trend expansion is conducted to predict technology development directions and business opportunities. Specifically, this includes: using a three-level transfer learning framework, which includes: near-level transfer based on similarity matching of cases in the same industry, mid-level transfer using a technology word co-occurrence network to discover cross-sub-industry connections, and far-level transfer using an attention mechanism to capture cross-domain technology features; and based on the analysis results of the three-level transfer learning framework, a three-stage trend forecast report is generated, covering short-term, medium-term, and long-term trends. The output includes structured analysis results of original search results, association relationships and trend predictions.
2. The method according to claim 1, characterized in that In the step of receiving a natural language query request input by a user and performing semantic parsing and vectorization processing on the natural language query request, performing semantic parsing and vectorization processing on the query request includes: Use pre-trained language models to perform semantic analysis on query text and extract query intent; and The query text is converted into a semantic vector through an embedding model.
3. The method according to claim 2, characterized in that According to the results of the semantic analysis and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results. The information retrieval and similarity matching include: executing keyword retrieval based on the inverted index and approximate nearest neighbor search in a vector database based on the semantic vector in parallel; and The matching score of the keyword search and the vector semantic similarity score are dynamically weighted and fused to generate the search results with comprehensive ranking.
4. The method according to claim 1, wherein Also includes: When a certain industry case is queried, technical information related to the industry case is automatically expanded based on the bidirectional traceability association network; When a certain technical information is queried, the industrial cases related to the technical information are automatically expanded according to the bidirectional traceable association network.
5. The method according to claim 1, wherein In the step of using an artificial intelligence model to conduct in-depth analysis on the data after the correlation analysis, extract key information and identify potential technology trends and business opportunities, the in-depth analysis using the artificial intelligence model includes: Extract technical entities and business elements through large language models; Analyze the mapping relationship between technology evolution paths and market application scenarios; and Generate a structured analysis report that includes technology maturity assessment, commercialization path planning, title, future outlook, solutions and technology selection.
6. A heterogeneous A / B data source connection and enhancement system, characterized in that: include: An intelligent retrieval module is used to receive natural language query requests input by users and perform semantic parsing and vectorization processing on the natural language query requests; Based on the results of the semantic analysis and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results; A data source association module is used to perform data source association analysis on the information from different heterogeneous data sources in the search results and establish a traceable association relationship between cross-source data. Specifically, it includes: extracting technical features through named entity recognition to build a technology-case association map; and implementing classification system mapping based on a predefined rule base to establish a bidirectional traceable association network between industry cases, technology nodes and technical documents; An AI deep analysis module is used to use an artificial intelligence model to perform in-depth analysis on the data after the correlation analysis, extract key information, and identify potential technology trends and business opportunities; The trend expansion module is used to expand trends based on the results of the in-depth analysis and predict technology development directions and business opportunities; output structured analysis results including original search results, association relationships and trend predictions; specifically includes: adopting a three-level transfer learning framework, the three-level transfer learning framework includes: near-level transfer based on similarity matching of cases in the same industry, middle-level transfer based on the discovery of cross-sub-industry associations through the technical word co-occurrence network, and far-level transfer based on the attention mechanism to capture cross-domain technical features; and based on the analysis results of the three-level transfer learning framework, forming a three-stage trend forecast report including short-term, medium-term and long-term.
7. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the method implements the steps of a heterogeneous A / B data source connection enhancement method as described in any one of claims 1 to 5.
8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the heterogeneous A / B data source connection enhancement method as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Multi-source heterogeneous data association query method and system
CN110837585A
Graphene product retrieval method and system based on large model
CN118820545A
Graphene industry application discovery method based on large model and knowledge graph analysis
CN119168423A