Heterogeneous A / B data source run-through strengthening method, system, equipment and medium
By receiving natural language queries and performing semantic analysis and vectorization processing, cross-origin data associations are established, and artificial intelligence models are used for in-depth analysis, the integration problem of heterogeneous A/B data sources is solved, and efficient data fusion and forward-looking trend prediction are achieved.
Patent Information
- Application Number
- CN202510846200.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
It is difficult for existing technology to efficiently integrate and deeply interpret heterogeneous A/B data sources, resulting in information dispersion and inefficient fusion, lack of in-depth interpretation and prediction capabilities, and it is difficult to support cross-industry analysis and personalized exploration.
By receiving natural language queries, performing semantic analysis and vectorization processing, performing information retrieval and matching similarity, establishing cross-origin data associations, using artificial intelligence models for in-depth analysis, and predicting technological trends and business opportunities.
It has achieved deep integration of cross-origin data, provided accurate technical insights and forward-looking trend predictions, broken data silos, and supported instant customized analysis and interactive feedback.
Smart Images

Figure CN120353835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer data processing, and in particular to a method for information retrieval, analysis, association and trend prediction of heterogeneous data sources based on artificial intelligence. Specifically, it is a method, system, device and medium for strengthening the connection between heterogeneous A / B data sources. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, data-driven decision-making and innovation have become increasingly important in various industries. In this context, how to effectively integrate and utilize heterogeneous data sources from different sources with different structures, such as type A data sources (such as structured technical literature databases, internal case bases) and type B data sources (such as unstructured industry reports, external market dynamics data), has become the key to enhancing the depth and breadth of analysis. However, traditional methods for obtaining and processing such heterogeneous A / B data sources have many deficiencies in achieving efficient connection and deep integration, specifically manifested as follows: 1. Barriers and integration problems between A / B data sources: Traditional methods often focus on the independent application of a single data source (A or B), resulting in scattered information between type A data and type B data, making it difficult to achieve effective cross-source connection and comprehensive integration, forming de facto "data islands", which limits the potential of comprehensive analysis.
[0003] 2. Low efficiency of heterogeneous A / B data fusion processing: For type A and type B data sources with different structures, formats and semantics, existing technical tools mostly rely on static matching or shallow association, and are inefficient in cross-source data alignment, conversion and fusion, making it difficult to deeply explore the complex internal connections and complementary values between A and B data, which hinders the value-added utilization after connection.
[0004] 3. Lack of in-depth interpretation and prediction ability based on connected data: Due to the failure to effectively connect and integrate heterogeneous A / B data sources, existing technologies often lack comprehensive insights based on AI deep learning and big data analysis, and are unable to deeply interpret the integrated complex industrial cases and multi-source technical data, let alone make accurate trend predictions.
[0005] 4. Insufficient interactivity and personalization in data exploration after connection: Even if the initial collection of A / B data is achieved, traditional platforms mostly provide static and preset query and analysis modes, making it difficult to support users to conduct instant and customized exploratory analysis and interactive feedback based on the connected A / B data, reducing the efficiency of data value discovery.
[0006] 5. Lack of cross - industry A / B data integration and analysis capabilities: When data sources of type A and type B belong to different industries or fields, due to the lack of effective integration mechanisms and support from cross - domain knowledge graphs, existing platforms have difficulty in integrating and analyzing the technological development trends and potential cross - over opportunities across industries.
[0007] Therefore, there is an urgent need to develop a method for strengthening the integration of heterogeneous A / B data sources, aiming to break down the barriers between different data sources, improve the efficiency and depth of heterogeneous data fusion processing, empower in - depth interpretation and intelligent prediction based on multi - source data, and solve the defects and deficiencies of the above - mentioned existing technologies in integrating and utilizing heterogeneous A / B data sources. Summary of the Invention
[0008] Based on this, the present invention aims to solve the many deficiencies in traditional information acquisition and processing methods for heterogeneous A / B data sources, as well as in achieving efficient integration and deep fusion, and provides a method, system, device, and medium for strengthening the integration of heterogeneous A / B data sources.
[0009] In the first aspect, the present invention provides a method for strengthening the integration of heterogeneous A / B data sources, including the following steps: Receiving a natural - language query request input by a user, and performing semantic parsing and vectorization processing on the natural - language query request; According to the results of the semantic parsing and vectorization processing, retrieving information and performing similarity matching from multiple heterogeneous data sources to obtain retrieval results; Performing data - source correlation analysis on the information from different heterogeneous data sources in the retrieval results, and establishing a traceable correlation relationship for cross - source data; Using an artificial - intelligence model to perform in - depth analysis on the data after the correlation analysis, extracting key information and identifying potential technological trends and business opportunities; Based on the results after the in - depth analysis, performing trend expansion to predict the technological development direction and business opportunities; Outputting a structured analysis result including the original retrieval results, correlation relationships, and trend predictions.
[0010] As an optional implementation manner of the first aspect of the present application, in the step of receiving a natural - language query request input by a user and performing semantic parsing and vectorization processing on the query request, performing semantic parsing and vectorization processing on the natural - language query request includes: using a pre - trained language model to perform semantic parsing on the query text to extract the query intent; and converting the query text into a semantic vector through an embedding model.
[0011] As an alternative implementation of the first aspect of the present application, in the step of retrieving information and performing similarity matching from multiple heterogeneous data sources according to the results of the semantic parsing and vectorization processing, the information retrieval and similarity matching include: parallelly executing keyword retrieval based on an inverted index and approximate nearest neighbor search in a vector database based on the semantic vectors; and dynamically weighted fusion processing of the matching degree scores of the keyword retrieval and the vector semantic similarity scores to generate the retrieved results with comprehensive sorting.
[0012] As an alternative implementation of the first aspect of the present application, in the step of performing data source association analysis on the information from different heterogeneous data sources in the retrieved results and establishing a traceable association relationship for cross-source data, the data source association analysis includes: extracting technical features through named entity recognition to construct a technology-case association graph; and realizing classification system mapping based on a predefined rule library to establish a two-way traceable association network among industrial cases, technical nodes, and technical documents.
[0013] As an alternative implementation of the first aspect of the present application, it further includes: when querying a certain industrial case, automatically expanding the technical information related to the industrial case according to the two-way traceable association network; when querying a certain technical information, automatically expanding the industrial cases related to the technical information according to the two-way traceable association network.
[0014] As an alternative implementation of the first aspect of the present application, in the step of using an artificial intelligence model to deeply analyze the data after the association analysis, extract key information, and identify potential technical trends and business opportunities, the deep analysis using the artificial intelligence model includes: extracting technical entities and business elements through a large language model; analyzing the mapping relationship between the technical evolution path and the market application scenario; and generating a structured analysis report including technical maturity assessment, commercialization path planning, title, future outlook, solution, and technology selection.
[0015] As an alternative implementation of the first aspect of the present application, in the step of performing trend expansion based on the results of the deep analysis to predict the technical development direction and business opportunities, the trend expansion includes: adopting a three-level transfer learning framework, where the three-level transfer learning framework includes: near-layer transfer based on similarity matching of cases in the same industry, middle-layer transfer for discovering cross-sub-industry associations through a technical word co-occurrence network, and far-layer transfer for capturing cross-domain technical features using an attention mechanism; and forming a three-stage trend prediction report including short-term, medium-term, and long-term based on the analysis results of the three-level transfer learning framework.
[0016] In a second aspect, an embodiment of the present application provides a heterogeneous A / B data source penetration and enhancement system, including: An intelligent retrieval module, configured to receive a natural language query request input by a user, perform semantic parsing and vectorization processing on the natural language query request; retrieve information and perform similarity matching from multiple heterogeneous data sources according to the results of the semantic parsing and vectorization processing to obtain retrieval results; A data source association module, configured to perform data source association analysis on the information from different heterogeneous data sources in the retrieval results, and establish a traceable association relationship for cross-source data; An AI in-depth analysis module, configured to use an artificial intelligence model to perform in-depth analysis on the data after the association analysis, extract key information, and identify potential technical trends and business opportunities; A trend expansion module, configured to expand trends based on the results after the in-depth analysis, predict the technical development direction and business opportunities; output a structured analysis result including the original retrieval results, association relationships, and trend predictions.
[0017] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0019] Compared with the prior art, the beneficial effects of the heterogeneous A / B data source penetration and enhancement method proposed by the present invention are as follows: (1) Improve retrieval accuracy and depth: By combining natural language processing and large model technology, it can accurately understand the user's query intention, perform multi-dimensional retrieval and information fusion, deeply explore the potential associations in technical information, cross the barriers of different data sources, and obtain more accurate and relevant technical solutions and industrial cases.
[0020] (2) Break data silos: Through the data source association module, accurately match and integrate information from different heterogeneous data sources, achieve deep integration between data sources, and provide users with a global perspective of technical insights.
[0021] (3) Provide forward-looking trend prediction: Use a large model to perform in-depth analysis on the data, combine industry historical data and existing trends, accurately identify technical development trends, market demand changes, and potential opportunities, and provide users with forward-looking industry trend predictions.
[0022] (4) Realize mutual reinforcement of information: By constructing two-way associations between heterogeneous data sources, querying one type of data (such as industrial cases) can reverse-extract relevant another type of data (such as technical literature), and vice versa, achieving mutual verification and value improvement of information. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a schematic structural diagram of a system for integrating and mutually reinforcing heterogeneous A / B data sources according to the present invention.
[0024] Figure 2 FIG. is a flowchart of a method for integrating and mutually reinforcing heterogeneous A / B data sources according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the associated objects before and after are in an "or" relationship. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0027] It should be noted that the "module" described in the embodiments of the present invention can be a software unit, a hardware unit, or a combination of both. For a software unit, it can be stored in a computer-readable storage medium and implement corresponding functions when executed.
[0028] Embodiment 1: Implementation Based on the System Referring to Figure 1 , a system for integrating and strengthening heterogeneous A / B data sources provided by the present invention is used to implement the method in Embodiment 2 below. The system may include: an intelligent retrieval module 101, an AI in-depth analysis module 102, a data source association module 103, and a trend expansion module 104. These modules work together to implement a complete process from user query to trend prediction.
[0029] Intelligent Retrieval Module 101: This module is responsible for receiving natural language queries input by users (e.g., "How to improve the conversion efficiency of solar cells?").
[0030] First, the query question is vectorized and parsed through a pre-trained large language model (e.g., calling the API of the text-embedding-3-small model: response = await client.embeddings.create(input=text, model='text-embedding-3-small')), converting the text text into a vectorized result response.
[0031] Then, according to the vectorized semantic parsing result, a hybrid retrieval technique is called to perform information retrieval in multiple heterogeneous databases (such as data source A: industry case library, data source B: technical literature library). Hybrid retrieval includes: 1. Keyword retrieval: Extract keywords (such as "solar cell", "conversion efficiency", "improve") from the query and match them in the case database and the literature database.
[0032] 2. Semantic retrieval: Based on the vectorized result response of the query (i.e., goal_vec), find content with similar semantics in the vector database. Matching is performed by calculating the cosine similarity between the query vector and the stored vector stored_vec in the database: similarity = np.dot(response, stored_vec) / (np.linalg.norm(goal_vec) * np.linalg.norm(stored_vec)).
[0033] Finally, the module integrates the retrieved information, sorts it according to the weighted fusion result of the keyword matching score and the semantic similarity score, and returns it to the user. (For example: The top-ranked result may be an academic paper on "The application of new perovskite materials in solar cells". The second-ranked result may be a patent describing a technology of "improving the efficiency of solar cells through multi-layer antireflection coatings". The third-ranked result may be an industry case introducing a company that has improved the battery performance by optimizing the manufacturing process).
[0034] In the intelligent retrieval module 101, this module is responsible for receiving the natural language queries input by the user through the user layer, and vectorizing and parsing the queries through a pre-trained large-scale language model to extract the query intent. The specific implementation steps are as follows: Use the pre-trained language model to convert the query text into a high-dimensional vector representation. According to the semantic understanding results, call the retrieval technology to perform similarity matching in the vector database to retrieve relevant industrial cases and technical information. Use a sorting algorithm (comparing the magnitudes of vector similarities) to optimize the sorting of the retrieval results to ensure the return of the most relevant information.
[0035] AI Deep Analysis Module 102: This module receives the retrieved industrial cases and technical information output by the intelligent retrieval module 101 and delivers them to the "Industry and Technology Analysis Intelligent Agent" customized by the present invention to extract key information, identify potential technical trends and business opportunities. And comprehensively evaluate this case, analyze its market prospects, technical feasibility, and potential application fields, providing comprehensive decision-making support for users. The following is the relevant information about the intelligent agent: Role: A senior case analyst, good at extracting key information from complex cases, writing attractive titles, future outlooks, and proposing innovative solutions. Please conduct in-depth analysis based on the background and details of specific cases and accurately select relevant artificial intelligence technologies.
[0036] Skill 1: Information extraction. Read the case carefully, capture the core information, and ensure that the title is attractive and expressive enough to arouse the interest of readers.
[0037] Skill 2: Future outlook. Deeply analyze the case content and the technologies used, and detailedly predict future development trends, including the time frame, potential market impact, specific manifestations of technological progress, etc.
[0038] Skill 3: Solution. Thoroughly analyze the problems and challenges in the case, provide a comprehensive root cause analysis, and propose practical solutions in combination with the actual situation, detailing the implementation steps and citing successful cases to enhance the persuasion.
[0039] Skill 4: Technology selection. Select 3 to 5 technologies that match the case from the technology name list, ensure that each technology selection has sufficient reasons, and detailedly explain its applicability and how it is applied to the case.
[0040] Output format: Please organize the output in JSON format.
[0041] Exemplarily, assume that an industrial case is retrieved: "A company has developed a new type of perovskite-based solar cell with a conversion efficiency of 28% and plans to commercialize it within the next two years." The working process of the AI Deep Analysis Module 102 includes: 1. Automated Analysis and Key Information Extraction The AI in-depth analysis module 102 uses agents to analyze cases and extract the following key information: 1. Technical core: perovskite materials, solar cells, conversion efficiency of 28%. 2. Commercialization plan: achieve commercialization within the next two years. 3. Potential advantages: high efficiency, low cost, environmentally friendly.
[0042] 2. Technical Trends and Business Opportunity Identification Then the module uses semantic analysis and large model networking technology to identify the following trends and opportunities: 1. Technical trend: The application of perovskite materials in the photovoltaic field is developing rapidly and may become the mainstream technology for the next generation of solar cells. 2. Business opportunity: High-efficiency solar cells have broad application prospects in fields such as distributed energy and electric vehicle charging stations.
[0043] 3. Comprehensive Evaluation and Decision Support Finally, the module conducts a comprehensive evaluation of the case, including: 1. Market prospect: The global solar cell market is expected to have an average annual growth rate of 15% in the next five years, and perovskite cells are expected to occupy an important share. 2. Technical feasibility: The technical maturity of perovskite materials is relatively high, but stability and large-scale production issues still need to be addressed. 3. Application fields: Suitable for rooftop photovoltaics, portable charging devices, space energy, etc.
[0044] 4. Generate Title, Future Outlook, Solutions and Technology Options Based on the analysis results, the module generates the following JSON output: {"title": "Perovskite Solar Cells: Dual Breakthroughs in High Efficiency and Commercialization Prospects", "future_view": "In the next five years, perovskite solar cells are expected to achieve large-scale commercialization, and it is estimated to occupy 20% of the global photovoltaic market by 2028. Their high efficiency (exceeding 30%) and low-cost characteristics will drive the rapid development of distributed energy and electric vehicle charging infrastructure. Technological progress will focus on material stability and production process optimization.", "solution": "To accelerate the commercialization of perovskite solar cells, the following steps are recommended: 1. Strengthen research on material stability, and develop new encapsulation technologies through cooperation with universities and research institutions. 2. Optimize the production process and introduce automated production lines to reduce costs. 3. Expand application scenarios, such as cooperating with electric vehicle manufacturers to develop integrated photovoltaic roofs. Reference to successful cases: A company improved the encapsulation technology and increased the lifespan of perovskite cells from 1000 hours to 5000 hours, laying the foundation for commercialization.", "technologies": "Knowledge Graph, Transformer Infrastructure, Semantic Web, End-to-End Learning, Image Processing"} In the AI Deep Analysis Module 102, the module uses large model parsing to automatically analyze the retrieved industrial cases and technical information, extract key information, identify potential technical trends and business opportunities. And comprehensively evaluate the case, analyze its market prospects, technical feasibility and potential application fields, providing users with comprehensive decision-making support. The specific implementation steps are as follows: Use large models (such as GPT-4, DeepSeek, etc.) to deeply analyze the text information in the retrieval results, extract key technical elements, application scenarios and industry impacts, and analyze future development trends and technical solutions.
[0045] Data Source Association Module 103: Through a dual-channel dynamic association mechanism, this module precisely matches information (such as technical literature, industrial cases, etc.) in different heterogeneous data sources. It uses a technical feature intelligent extraction channel: context-aware technology recognition based on large language models (LLMs), and a case feature mapping channel: multi-dimensional analysis of structured industrial cases, to achieve two-way traceable association between technical literature and industrial cases, solving the problem of data silos in traditional technical industrialization analysis methods.
[0046] The implementation process includes: 1. Heterogeneous data feature extraction: For industrial cases (Data Source A), precise extraction of technical features is achieved through multiple rounds of Prompt engineering. The pseudo-code is as follows: - Title: {title} - Solution: {solution} - Future Outlook: {future_view}” Output example: "Artificial Neural Network, Convolutional Neural Network, Deep Learning".
[0047] 2. For technical documents (Data Source B, such as patents), a combination of document parsing and vector retrieval is used to vectorize the technical documents. The pseudo-code is as follows: soup = BeautifulSoup(description, 'html.parser') text = str() for p in soup.find_all('p'): text += p.text searches = await vector_retrieval(text) where text is "Technical document theme: {title}, Technical document classification: {classification}, Content: {text}"; searches are the vectorized results of technical documents.
[0048] 3. Dynamic association construction: Association rule engine: Achieve exact matching of technology names. The technology classification tree mapping (through the classification field) can accurately assign the categories of technologies, and use semantic similarity to assist in technical document matching (through the vectorized results).
[0049] Association database structure: A [Industrial case] -- applies --> B ((Technical node)), B -- improves --> C [Patent document], C -- cited_by --> A 4. Two-way query service: Through the above steps, when querying a certain industrial case, the module will automatically expand the relevant technical information of the industrial case according to the above associations. At the same time, when querying a certain technical information, the module will also automatically expand the relevant industrial cases of the technical information according to the above associations.
[0050] Exemplarily, assume there are the following heterogeneous data sources: "Industrial case": A company successfully developed and commercialized perovskite solar cells. And "Patent document": A patent on "Perovskite material encapsulation technology".
[0051] First, construct an association database, extract keywords from the industrial case: perovskite, solar energy, commercialization, and extract keywords from the patent document: encapsulation technology, perovskite, production process.
[0052] Second, precise matching and association: The module uses the above technologies to precisely match the information in different data sources: associate "perovskite" in the patent document with "perovskite" in the industrial case.
[0053] Finally, the module integrates the data from different sources into a unified knowledge base through data fusion technology: eliminate information silos and ensure that the technical information obtained by users is extensive and multi-dimensional. For example, when a user queries "perovskite solar cells", the module returns the following associated information: Industrial case: Successful experience in commercial applications. Patent document: Innovation points of perovskite encapsulation technology.
[0054] In the data source association module 103, the module builds an association database to precisely match information (such as technical literature, industrial cases, etc.) from different heterogeneous data sources, achieving a deep association between technologies and application cases. Through data fusion technology, data from different sources are integrated to eliminate information silos, ensuring that the technical information obtained by users has comprehensiveness and a multi-dimensional perspective. The specific implementation steps are as follows: Extract key information from multiple heterogeneous data sources and build a database containing multi-level relationships such as technical information and industrial cases. Utilize large model technology to discover potential association relationships between technologies and cases and optimize the matching degree. Through data fusion technology, integrate data from different sources to eliminate information silos, ensuring that the technical information obtained by users has comprehensiveness and a multi-dimensional perspective. The innovation of this module lies in: 1. Dynamic feature extraction mechanism: Achieve context-related technology identification through LLM prompt engineering; 2. Bidirectional association index: Establish a three-layer traceable relationship network of case → technology → patent.
[0055] Trend expansion module 104: This module pioneered the "three-level migration prediction model", which quantitatively analyzes the results obtained from the data source association module 103 through technical association degree: Based on the dynamic weight calculation of industrial case - technical information association, cross-industry migration engine: Realize the technical migration prediction of near / medium / far three-level industry scenarios, multi-dimensional trend verification: Utilize industry data of the same type and peers and cross-industry cases, and predict future technology development trends by analyzing their performances to help users identify potential industry innovation opportunities. Through multi-dimensional analysis of historical technical cases and current industry performances, combined with market dynamics, provide long-term and short-term technology trend predictions for users to assist enterprise strategic planning.
[0056] The implementation process includes: Phase 1: Industrial case feature extraction. Extract the core technology label technology from the target case and determine its position in the industry knowledge graph according to industry classification.
[0057] Phase 2: Three-level case migration.
[0058] Near migration: Search for cases applying similar technologies (such as dynamic pricing of food delivery platforms) in the same subclass industry (for example, if the target case is logistics route optimization, then in the last-mile delivery field).
[0059] Medium migration: Search for cases applying similar technologies (such as air cargo loading optimization) in different subclasses of the same major category industry (for example, in different links of the logistics industry).
[0060] Far migration: Search for cases applying similar technologies (such as hospital operating room scheduling) in cross-industry fields (such as the medical industry).
[0061] Phase 3: Intelligent Trend Generation. Leverage large models (e.g., through a carefully designed system_prompt = "As a trend analysis expert, please generate a report based on the following dimensions: - Industry Status Analysis - Insights from Successful Cases... - Cross-Industry Collaboration") combined with the results of three-level migration analysis to generate a trend prediction report that includes industry status, insights from successful cases, innovative strategic recommendations, potential risks, future outlook, and cross-industry collaboration suggestions.
[0062] Exemplarily: 1. Input case: "A logistics company uses an AI path optimization algorithm (Technology: Reinforcement Learning + Spatiotemporal Prediction) to achieve a 22% improvement in delivery efficiency." Industry Classification: Logistics Industry → Last-Mile Delivery → Route Planning 2. Results of three-level migration case matching 3. Key paragraphs of the generated trend report (1) Industry Status Analysis: The application of artificial intelligence in the current logistics industry has entered a rapid growth period. According to a 2023 report by Deloitte, 85% of leading logistics companies have deployed AI path optimization systems, and the industry's average delivery efficiency has increased by 17.5% (your case of 22% is better than the average). (2) Insights from Successful Cases graph LR A[Logistics Path Optimization] -->|Near Migration| B[Food Delivery Dynamic Pricing] A -->|Medium Migration| C[Aircraft Loading Optimization] A -->|Far Migration| D[Operating Room Scheduling] Near migration case: A food delivery platform's dynamic pricing system Key to success: Reuse 85% of the underlying logic of the path optimization algorithm Replicable strategy: Add a price sensitivity parameter module to your system Far migration case: A hospital operating room scheduling system Lessons learned: The medical scenario verifies the robustness of the algorithm in complex scenarios Adaptation suggestion: Develop a delivery priority rule configuration panel (3) Innovative Strategic Recommendations Short-term goal (within 1 year): Establish a "Path Optimization Algorithm Supermarket" and encapsulate 3 preset modes: Cost Priority Mode, Timeliness Priority Mode Balance Mode.
[0063] Long-term goal (3 - 5 years): Develop a cross-industry technology adaptation engine to support: Logistics → Medical Scenario Conversion (requires a 120-day development cycle) Logistics → Retail scenario conversion (90-day development cycle required).
[0064] (4) Potential risks and countermeasures: Risk matrix: Key warning: Special attention should be paid to the data collection of non-motor vehicle lanes in the urban low-speed distribution scenario (refer to the Shanghai pilot experience). (5) Outlook on future trends Technical integration direction: Edge computing devices will optimize paths in real time (reduce dependence on the cloud), and digital twin technology will enable full-scenario simulation testing Sustainable development suggestions: Develop a "green path mode", and give priority to: routes around energy vehicle charging stations, low-emission traffic control areas (6) Suggestions for cross-industry collaboration List of cooperation opportunities: In the trend expansion module 104, this module combines the industry data of the same kind of peers and the technical cases of cross-industries, analyzes their performances, predicts the future technical development trends, and helps users identify potential industry innovation opportunities. The specific implementation steps are as follows: Normalize the multi-source data collected, and use feature selection methods to extract features valuable for trend prediction. Integrate industrial cases and technical information from different fields, conduct cross-industry data integration and comparative analysis, and identify potential cross-border applications and innovation opportunities of technologies. The system automatically generates a report on the future development trends of the industry according to the output of the large model, covering the technical development direction, market demand changes, and potential business opportunities.
[0065] Example 2: Method flow Refer to Figure 2 , a heterogeneous A / B data source penetration and strengthening method provided by the present invention specifically includes the following steps: Step S201: Receive a natural language query request input by a user, and perform semantic parsing and vectorization processing on the natural language query request.
[0066] The system receives a natural language query request input by the user through the interface (e.g., "How to improve the conversion efficiency of solar cells?"). It uses a pre-trained language model (such as BERT, GPT architecture) to perform semantic parsing on the query text to understand the user's intention. Subsequently, the query text is transformed into a 768-dimensional semantic vector through an embedding model (such as text-embedding-3-small). This process can be achieved by calling the corresponding embedding API (such as OpenAI's Embeddings API). For example, when the user inputs text = "How to improve the conversion efficiency of solar cells?", calling client.embeddings.create(input=text, model='text-embedding-3-small') obtains the vectorization result response.
[0067] Step S202: According to the results of the semantic parsing and vectorization processing, information retrieval and similarity matching are performed from multiple heterogeneous data sources to obtain retrieval results.
[0068] The system executes two retrieval processes in parallel: Keyword retrieval based on inverted index: Extract keywords (such as "solar cell", "conversion efficiency", "improve") from the query request, search for matching documents in the pre-constructed index containing industrial cases (data source A) and technical literature (data source B), and calculate the matching score between the term and the document.
[0069] Semantic retrieval based on vector database: Use the query semantic vector generated in step S201 to perform approximate nearest neighbor (ANN) search in the pre-constructed vector database (which stores the semantic vectors of industrial cases and technical literature). Calculate the similarity score between the query vector and the stored vectors in the database through cosine similarity. For example, similarity = np.dot(query_vec, stored_vec) / (np.linalg.norm(query_vec) * np.linalg.norm(stored_vec)).
[0070] The retrieval results are processed by a dynamic weighted fusion algorithm. For example, the comprehensive score = α * semantic similarity score + (1 - α) * keyword matching score, where α is an adjustable parameter (such as the default 0.6). The retrieval results are sorted according to the comprehensive score to generate a comprehensive sorted list.
[0071] Step S203: Perform data source association analysis on the information from different heterogeneous data sources in the retrieval results to establish a traceable association relationship for cross-source data.
[0072] Using a two-way association engine technology, associate the heterogeneous data in the retrieval results obtained in step S202: Construction of technology-case association graph: Extract key technology features (such as "perovskite", "encapsulation technology") from industrial cases and technical literatures through named entity recognition (NER).
[0073] Classification system mapping and association network construction: Map the extracted technology features to corresponding technology nodes based on a predefined rule base (such as a technology classification tree). Establish a traceable association relationship among industrial cases (A), technology nodes (B), and patents / technical documents (C), for example, A -- applies --> B, B -- improves --> C, C -- cited_by --> A.
[0074] Through this step, if an industrial case "a company successfully developed perovskite solar cells and achieved commercialization" is retrieved, the system can associate it with a patent document "a patent on 'perovskite material encapsulation technology'" through the common technology node "perovskite".
[0075] Step S204: Use an artificial intelligence model to deeply analyze the data after the association analysis, extract key information, and identify potential technology trends and business opportunities.
[0076] Deploy a multi-task analysis model (usually based on large language models, such as GPT-4 or customized agents) to analyze the associated data: 1. Entity and element extraction: Extract core technology entities and business elements (such as market size, competitors, business models) from industrial cases and technical literatures.
[0077] 2. Evolution and application mapping: Analyze the mapping relationship between the evolution path of technology (such as from laboratory research to commercial application) and market application scenarios (such as rooftop photovoltaic, portable devices).
[0078] 3. Report generation: Finally, generate a structured analysis report, the content of which can include technology maturity assessment, commercialization path planning, engaging titles, future outlook, solution suggestions, relevant technology selection, etc. The output format is JSON, as described in AI deep analysis module 102 in Embodiment 1.
[0079] Step S205: Based on the results of the in-depth analysis, conduct trend expansion to predict the technology development direction and business opportunities.
[0080] Adopt a three-level transfer learning framework to expand the analysis results of step S204 and predict technology trends: 1. Near-layer migration: Based on the similarity matching of cases within the same industry (e.g., within the solar cell industry) (such as cosine similarity > 0.85), analyze the application and performance of similar technologies in similar scenarios.
[0081] Mid-layer migration: Discover the technical associations across sub-industries (e.g., from photovoltaic power generation to energy storage technologies) through the co-occurrence network of technical terms, and analyze the migration potential of technologies in related sub-industries.
[0082] 2. Far-layer migration: Use deep learning models such as the attention mechanism to capture the technical feature similarities across domains (e.g., from energy materials to biomedical materials), and explore long-distance technology migration opportunities.
[0083] 3. Based on this three-level migration analysis, combined with market data and expert knowledge, form a three-stage trend prediction report covering the short term (1 - 3 years), medium term (3 - 5 years), and long term (more than 5 years). The report content is such as the "Industry Status Analysis", "Inspiration from Successful Cases", etc. described in the trend expansion module 104 of Example 1.
[0084] Step S206: Output the structured analysis results including the original retrieval results, association relationships, and trend predictions.
[0085] The system integrates the processing results and returns a structured result set to the user. This result set includes three main dimensions: 1. The list of original retrieval results (sorted).
[0086] 2. The associated knowledge graph (showing the association relationships between data, and visualization is better).
[0087] 3. The in-depth analysis and trend prediction report (including the outputs of the AI in-depth analysis module 102 and the trend expansion module 104).
[0088] In this way, a complete service closed-loop from basic information retrieval to advanced decision support is achieved.
[0089] A heterogeneous A / B data source penetration enhancement system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a Personal Computer (PC), etc., which are not specifically limited in the embodiments of the present application.
[0090] A heterogeneous A / B data source penetration enhancement system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0091] A heterogeneous A / B data source penetration enhancement system provided in an embodiment of the present application can implement Figure 1 each process implemented by a heterogeneous A / B data source penetration enhancement method in the method embodiment. To avoid repetition, it will not be elaborated here.
[0092] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned heterogeneous A / B data source penetration enhancement method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0093] An embodiment of the present application further provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned heterogeneous A / B data source penetration enhancement method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0094] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0095] It should be noted that in this text, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0097] The embodiments of the present application have been described above with reference to the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A method for strengthening the connection between heterogeneous A / B data sources, characterized in that, It includes the following steps: Receive a natural language query request input by the user, and perform semantic parsing and vectorization processing on the natural language query request; According to the results of the semantic parsing and vectorization processing, perform information retrieval and similarity matching from multiple heterogeneous data sources to obtain retrieval results; Perform data source association analysis on the information from different heterogeneous data sources in the retrieval results, and establish a traceable association relationship for cross-source data; Use an artificial intelligence model to deeply analyze the data after the association analysis, extract key information, and identify potential technological trends and business opportunities; Based on the results after the deep analysis, perform trend expansion to predict the technological development direction and business opportunities; Output a structured analysis result including the original retrieval results, association relationships, and trend predictions.
2. The method according to claim 1, wherein In the step of receiving a natural language query request input by the user and performing semantic parsing and vectorization processing on the natural language query request, performing semantic parsing and vectorization processing on the query request includes: Adopt a pre-trained language model to perform semantic parsing on the query text and extract the query intent; and Convert the query text into a semantic vector through an embedding model.
3. The method according to claim 2, wherein In the step of performing information retrieval and similarity matching from multiple heterogeneous data sources according to the results of the semantic parsing and vectorization processing to obtain retrieval results, performing information retrieval and similarity matching includes: Parallelly execute keyword retrieval based on an inverted index and approximate nearest neighbor search in a vector database based on the semantic vector; and Perform dynamic weighted fusion processing on the matching degree score of keyword retrieval and the vector semantic similarity score to generate the retrieval results with comprehensive sorting.
4. The method according to claim 1, wherein In the step of performing data source association analysis on the information from different heterogeneous data sources in the retrieval results and establishing a traceable association relationship for cross-source data, performing data source association analysis includes: Extract technical features through named entity recognition and construct a technology-case association graph; and Based on a predefined rule library, implement classification system mapping and establish a two-way traceable association network among industrial cases, technical nodes, and technical documents.
5. The method according to claim 4, characterized in that, It also includes: When querying a certain industrial case, automatically expand the technical information related to the industrial case according to the two-way traceable association network; When querying a certain technical information, automatically expand the industrial cases related to the technical information according to the two-way traceable association network.
6. The method according to claim 1, wherein In the step of using an artificial intelligence model to deeply analyze the data after the association analysis, extract key information, and identify potential technological trends and business opportunities, using an artificial intelligence model for deep analysis includes: Extract technical entities and business elements through a large language model; Analyze the mapping relationship between the technical evolution path and the market application scenario; and Generate a structured analysis report including technical maturity assessment, commercialization path planning, title, future outlook, solutions, and technology selection.
7. The method according to claim 1, characterized in that In the step of performing trend expansion based on the results after the deep analysis to predict the technological development direction and business opportunities, performing trend expansion includes: Adopt a three - level transfer learning framework, and the three - level transfer learning framework includes: near - layer transfer based on the similarity matching of cases in the same industry, middle - layer transfer for discovering cross - sub - industry associations through a technical word co - occurrence network, and far - layer transfer for capturing cross - domain technical features using an attention mechanism; and Based on the analysis results of the three - level transfer learning framework, form a three - stage trend prediction report including short - term, medium - term, and long - term trends.
8. A heterogeneous A / B data source connection and enhancement system, characterized in that Comprising: An intelligent retrieval module, configured to receive a natural language query request input by a user, and perform semantic parsing and vectorization processing on the natural language query request; According to the results of the semantic parsing and vectorization processing, perform information retrieval and similarity matching from multiple heterogeneous data sources to obtain retrieval results; A data source association module, configured to perform data source association analysis on the information from different heterogeneous data sources in the retrieval results, and establish a traceable association relationship for cross - source data; An AI deep - analysis module, configured to use an artificial intelligence model to deeply analyze the data after the association analysis, extract key information, and identify potential technical trends and business opportunities; A trend expansion module, configured to expand trends based on the results after the deep analysis, and predict the technical development direction and business opportunities; Output a structured analysis result including the original retrieval results, association relationships, and trend predictions.
9. An electronic device, characterized in that, Comprising a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for strengthening the connection between heterogeneous A / B data sources as described in any one of claims 1 - 7 are implemented.
10. A readable storage medium, characterized in that, The program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of a method for strengthening the connection between heterogeneous A / B data sources as described in any one of claims 1 - 7 are implemented.
Citation Information
Patent Citations
Fusion query method and device of heterogeneous multi-source data
CN108090154A
Multi-source heterogeneous data association query method and system
CN110837585A
Transform-based subject information retrieval system and method
CN117332782A
Occupational development prediction system and method based on big data analysis
CN117371625A
Graphene product retrieval method and system based on large model
CN118820545A