Sci-tech project matching method and system based on multi-dimensional data fusion

By using a multi-dimensional data fusion system, the problems of data heterogeneity, privacy leakage, low matching accuracy, and poor scenario adaptability in matching technology projects have been solved. It has achieved cross-domain data fusion and accurate matching, and improved the system's scalability and ease of use.

CN122196568APending Publication Date: 2026-06-12SHIJIAZHUANG TIEDA KEXIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHIJIAZHUANG TIEDA KEXIAN INFORMATION TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing methods for matching science and technology projects suffer from problems such as strong heterogeneity of multi-source data, inconsistent quality, easy leakage of sensitive information, lack of interpretability, low matching accuracy, and poor adaptability to different scenarios.

Method used

A multi-dimensional data fusion system is adopted, including multi-source data adaptive access, intelligent data cleaning, multi-dimensional feature fusion, dynamic scenario-based matching, and feedback closed-loop mechanism. Through technologies such as federated learning, differential privacy, hybrid fusion algorithms, and knowledge graphs, cross-domain data fusion and accurate matching are achieved.

Benefits of technology

It resolves the conflict between data barriers and privacy protection, improves matching accuracy and scenario adaptability, achieves high scalability and ease of use, and meets the personalized needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122196568A_ABST
    Figure CN122196568A_ABST
Patent Text Reader

Abstract

The application discloses a technology project matching method and system based on multidimensional data fusion, belongs to the technical field of data fusion, and aims at solving the problems of strong heterogeneity, uneven quality, easy leakage of sensitive information, lack of interpretability, low matching precision and poor scene adaptability of traditional multi-source data.The technology project matching system based on multidimensional data fusion comprises a data layer, a data management layer, a fusion algorithm layer, a matching logic layer, an application layer and a feedback loop module.The application realizes deep fusion of cross-node multi-source data, guarantees that original sensitive data is not leaked, solves the contradiction between data barriers and privacy protection, and can simultaneously realize dynamic noise adjustment and automatic desensitization processing, and balance data availability and compliance.In addition, the application solves the problem of an algorithm black box, and realizes adaptive optimization of the algorithm by combining migration learning and light-weight technology, and balances matching precision and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data fusion technology, specifically to a method and system for matching science and technology projects based on multi-dimensional data fusion. Background Technology

[0002] With the rapid development of the technology industry, the number and types of technology projects are becoming increasingly diverse, and the demand for cross-domain and cross-regional project cooperation continues to grow. However, existing technology project matching methods and systems have many shortcomings, as follows: At the data level, multi-source data suffers from strong heterogeneity, inconsistent quality, and easy leakage of sensitive information, forming data barriers. At the algorithm level, traditional fusion algorithms are mostly single-structured, lack interpretability, and are difficult to adapt to changes in data distribution and scenarios, resulting in an "algorithm black box" problem. At the matching logic level, most use static and single matching rules, relying solely on explicit keyword matching, failing to uncover implicit cross-domain connections, leading to low matching accuracy and poor scenario adaptability. At the application level, the system is complex to operate and lacks scalability, making it difficult to meet the personalized needs of different users, thus hindering its practical application.

[0003] To address the above issues, a method and system for matching science and technology projects based on multi-dimensional data fusion is proposed. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for matching science and technology projects based on multi-dimensional data fusion. By using this invention, the problems of strong heterogeneity, inconsistent quality, easy leakage of sensitive information, lack of interpretability, low matching accuracy and poor scenario adaptability of traditional multi-source data in the above-mentioned background are solved.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a technology project matching system based on multi-dimensional data fusion, comprising a data layer, a data governance layer, a fusion algorithm layer, a matching logic layer, an application layer, and a feedback closed-loop module, as detailed below: Data layer: Configured with a multi-source data adaptive access module and a privacy-compliant data processing module, enabling compliant access to multi-source heterogeneous data; Data governance layer: Configured with intelligent data cleaning module, heterogeneous data standardization module and multi-dimensional feature engineering module to complete data preprocessing; Fusion Algorithm Layer: Configures a hybrid fusion algorithm module, an algorithm interpretability enhancement module, and an adaptive algorithm optimization module to implement feature fusion; Matching logic layer: Configures a multi-dimensional matching indicator module, a dynamic weight adjustment module, and an implicit association mining module to achieve scenario-based matching; Application layer: Configure role-based operation modules, modular extension modules, and full-process interaction modules to implement practical applications; Feedback loop module: Collects user feedback data and optimizes system parameters and rules at each level through reinforcement learning.

[0006] Furthermore, the role-based operation module at the application layer is designed with personalized interfaces and functional modules for enterprise R&D personnel, government administrators, and university researchers; The modular extension module adopts a microservice architecture, which can separate data access, data governance, algorithm fusion and matching logic into independent microservices, and reserves interfaces for emerging technology fields; The full-process interactive module allows users to rate the matching results with stars and annotate the reasons, and synchronizes the feedback data to the feedback closed-loop module in real time.

[0007] This invention also proposes another technical solution: a technology project matching method based on multi-dimensional data fusion, comprising the following steps: S1: Multi-source heterogeneous data access, adopting multi-source adaptation and privacy compliance mechanisms, accessing structured, semi-structured and unstructured multi-type data, achieving cross-node data feature fusion through the fusion of federated learning and differential privacy technologies, and simultaneously de-identifying sensitive information; S2: Data governance preprocessing, through cleaning, standardization and feature extraction operations, performs intelligent cleaning, heterogeneous data standardization transformation and static and dynamic feature extraction on the data accessed in S1 to generate standardized feature vectors; S3: Multi-dimensional feature fusion, which adopts shallow feature fusion and deep semantic fusion, combines improved traditional machine learning algorithms with general artificial intelligence techniques to build an interpretable adaptive fusion algorithm model for fusion processing of features extracted in S2; S4: Dynamic scenario-based matching, based on a three-dimensional indicator system of technology, value and collaboration dimensions, achieves accurate matching of technology projects in multiple scenarios through dynamic weight adjustment and implicit correlation mining; S5: Matching result feedback optimization. Collect user feedback on matching results and combine reinforcement learning techniques to optimize algorithm parameters, matching weights, and data standards to form a closed-loop iteration throughout the entire process.

[0008] Furthermore, the multi-source data adaptive access module can adapt to the data update frequency of different data sources. In the privacy-compliant data processing mechanism, each data source node keeps the data locally stored and only transmits the feature fusion results. The intensity of differential privacy noise addition is dynamically adjusted according to the data sensitivity level.

[0009] Furthermore, the intelligent data cleaning module adopts a fusion strategy of rule engine and machine learning, as follows: The rule engine builds basic cleaning rules based on the experience of domain experts, and machine learning identifies abnormal data through the isolated forest algorithm and removes duplicate data based on the semantic hash deduplication algorithm. The heterogeneous data standardization module is used to build a standard dictionary for unified data in the field of science and technology, and to transform semi-structured or unstructured data into standardized feature vectors through natural language processing technology.

[0010] Furthermore, the static features extracted by the multi-dimensional feature engineering module include the technical field, funding scale, and quantitative indicators of partner qualifications; while the dynamic features are extracted through time series analysis and trend prediction, specifically including the potential for technology industrialization, patent citation growth rate, and market demand change trends.

[0011] Furthermore, in the first stage, the hybrid fusion algorithm module employs an improved weighted Bayesian algorithm for structured data and dynamically assigns weights based on data quality, while extracting deep semantic features from unstructured data using a BERT pre-trained model. In the second stage, a deep learning model with an attention mechanism is introduced to fuse multi-dimensional features, which is used to automatically focus on core features.

[0012] Furthermore, the algorithm interpretability enhancement module uses an attention weight visualization mechanism to display the weight allocation and association rules in the feature fusion process in the form of heatmaps and logical link diagrams. The adaptive algorithm optimization module reuses existing model parameters through transfer learning techniques and introduces a lightweight model structure to optimize inference speed in order to meet the requirements of real-time matching.

[0013] Furthermore, in S4, the technical dimensions of the three-dimensional indicator system include technological similarity, technological complementarity, and technological maturity; the value dimension includes industrialization potential, market prospects, and input-output ratio; and the synergy dimension includes partner suitability, resource complementarity, and regional synergy. The dynamic weight adjustment module can preset typical scenario weight templates and automatically optimize weight allocation by combining reinforcement learning algorithms with historical matching data.

[0014] Furthermore, in S4, the implicit association mining module integrates knowledge graph and semantic reasoning technologies to construct a science and technology knowledge graph covering technology, patents, enterprises, universities, and industry chain nodes. It mines cross-domain implicit associations through graph neural networks and achieves implicit association item matching based on semantic similarity reasoning.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention adopts a dual-technology fusion scheme of federated learning and differential privacy to construct a decentralized data fusion architecture, which not only achieves deep fusion of multi-source data across nodes, but also ensures that the original sensitive data is not leaked. It not only solves the contradiction between data barriers and privacy protection, but also takes into account the usability and compliance of data through dynamic noise adjustment and automatic desensitization processing.

[0016] 2. This invention integrates traditional machine learning and general artificial intelligence techniques, constructs a two-stage fusion model of shallow and deep layers, and introduces the idea of ​​ensemble learning to improve stability; it also achieves algorithm interpretability through attention weight visualization and result traceability, solving the "algorithm black box" problem, and combines transfer learning and lightweight technology to achieve adaptive optimization of the algorithm, balancing matching accuracy and efficiency.

[0017] 3. This invention constructs a three-dimensional indicator system of technology, value, and collaboration, which can comprehensively cover and match core needs. Furthermore, based on the dynamic weight adjustment mechanism of reinforcement learning, it can adapt to different application scenarios. In addition, this invention also breaks through the limitations of traditional keyword matching by integrating knowledge graph and graph neural network technologies to mine implicit cross-domain relationships, thereby improving matching accuracy and coverage.

[0018] 4. This invention adopts a microservice modular design, which realizes the system's high scalability and low barrier to entry ease of use; and it constructs a closed-loop mechanism for the entire process, which drives the continuous iteration of data, algorithms and matching logic through user feedback, thereby improving the system's long-term applicability and implementability. Attached Figure Description

[0019] Figure 1 This is a system flowchart of the present invention; Figure 2 This is a diagram illustrating the method steps of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] To address the technical challenges of traditional multi-source data, such as strong heterogeneity, inconsistent quality, easy leakage of sensitive information, lack of interpretability, low matching accuracy, and poor scenario adaptability, etc. Figure 1 As shown, the following preferred technical solutions are provided: The technology project matching system based on multi-dimensional data fusion includes a data layer, a data governance layer, a fusion algorithm layer, a matching logic layer, an application layer, and a feedback closed-loop module, as detailed below: Data layer: Configured with a multi-source data adaptive access module and a privacy-compliant data processing module, enabling compliant access to multi-source heterogeneous data; Data governance layer: Configured with intelligent data cleaning module, heterogeneous data standardization module and multi-dimensional feature engineering module to complete data preprocessing; Fusion Algorithm Layer: Configures a hybrid fusion algorithm module, an algorithm interpretability enhancement module, and an adaptive algorithm optimization module to implement feature fusion; Matching logic layer: Configures a multi-dimensional matching indicator module, a dynamic weight adjustment module, and an implicit association mining module to achieve scenario-based matching; Application layer: Configure role-based operation modules, modular extension modules, and full-process interaction modules to implement practical applications; Feedback loop module: Collects user feedback data and optimizes system parameters and rules at each level through reinforcement learning.

[0022] like Figure 2 As shown, to better explain the above embodiments, the present invention also discloses another implementation method: a technology project matching method based on multi-dimensional data fusion, comprising the following steps: S1: Multi-source heterogeneous data access, adopting multi-source adaptation and privacy compliance mechanisms, accessing structured, semi-structured and unstructured multi-type data, achieving cross-node data feature fusion through the fusion of federated learning and differential privacy technologies, and simultaneously de-identifying sensitive information; S2: Data governance preprocessing, through cleaning, standardization and feature extraction operations, performs intelligent cleaning, heterogeneous data standardization transformation and static and dynamic feature extraction on the data accessed in S1 to generate standardized feature vectors; S3: Multi-dimensional feature fusion, which adopts shallow feature fusion and deep semantic fusion, combines improved traditional machine learning algorithms with general artificial intelligence techniques to build an interpretable adaptive fusion algorithm model for fusion processing of features extracted in S2; S4: Dynamic scenario-based matching, based on a three-dimensional indicator system of technology, value and collaboration dimensions, achieves accurate matching of technology projects in multiple scenarios through dynamic weight adjustment and implicit correlation mining; S5: Matching result feedback optimization. Collect user feedback on matching results and combine reinforcement learning techniques to optimize algorithm parameters, matching weights, and data standards to form a closed-loop iteration throughout the entire process.

[0023] The core objective of the data layer is to address three major issues: poor data quality, high heterogeneity, and privacy breaches. Therefore, a multi-source adaptability and privacy-compliant data access mechanism is constructed to provide high-quality, compliant raw data support for subsequent processing, as detailed below: The multi-source data adaptive access module supports access to all types of structured, semi-structured, and unstructured data, covering various core data related to science and technology projects. Structured data includes directly quantifiable data such as project funding, technical parameters, and partner qualification certificates; semi-structured data includes patent abstracts, project applications, and technical feasibility reports; and unstructured data includes technical documents, academic papers, and video scripts for demonstrating research results.

[0024] To achieve efficient access to multi-source data, the multi-source data adaptive access module has designed a standardized interface adaptation system. For different data sources such as enterprise R&D systems, university patent databases, government project management platforms, and industry databases, the module pre-sets adaptation interface templates and supports both real-time data synchronization and batch import modes. Specifically: For scenarios with high real-time requirements, the real-time synchronization mode is used, capturing data update events from the data source through an interface listening mechanism and triggering the data synchronization process; for scenarios with large data volumes and lower real-time requirements, the batch import mode is used, and it can adapt to the data update frequency of different data sources, automatically adjusting the synchronization cycle.

[0025] To balance data fusion depth and privacy security, the privacy-compliant data processing module adopts a dual-technology fusion scheme of federated learning and differential privacy to construct a decentralized data fusion architecture, as detailed below: Decentralized data storage: Data from each data source node is stored locally and raw sensitive data is not transmitted externally. Feature extraction and preliminary processing are performed locally to avoid the risk of raw data leakage.

[0026] Federated learning cross-node fusion: Based on the federated learning framework, each node joins the federated learning network as a participant and shares the intermediate parameters required for feature fusion through an encrypted transmission protocol, rather than the original data. During the fusion process, a horizontal federated learning mode is used to aggregate features of data of the same type, and a vertical federated learning mode is used to complement features of data of different types, thereby achieving deep fusion of cross-node data features while ensuring data privacy.

[0027] Differential privacy noise addition: Differential privacy noise is dynamically added to the intermediate feature data during the federated learning fusion process according to the data sensitivity level. The sensitivity level is divided into three levels (high sensitivity: corporate trade secrets, unpublished scientific research results; medium sensitivity: project funding details, technical parameter details; low sensitivity: publicly available patent information, basic project introduction). The noise intensity added to high-sensitivity data is higher than that added to medium and low-sensitivity data. This not only ensures that attackers cannot reverse-engineer the original sensitive data through the fusion result, but also controls the impact of noise on data usability.

[0028] Automatic data anonymization: An embedded data anonymization module automatically identifies sensitive information through a combination of natural language processing technology and a rule engine, as detailed below: For text-based data, methods such as replacement, deletion, and masking are used for desensitization; for numerical data, methods such as perturbation and normalization are used for desensitization to ensure zero leakage of sensitive information.

[0029] The core objective of the data governance layer is to improve data quality and solve the challenge of heterogeneous data fusion. Therefore, it is necessary to establish a full-process data governance mechanism that includes cleaning, standardization, and feature extraction to transform raw data into standardized, high-value feature data, as detailed below: The intelligent data cleaning module employs a cleaning strategy that integrates rule engines and machine learning to achieve precise data cleaning, removing redundant, abnormal, and missing data to ensure data accuracy.

[0030] Rule Engine Cleaning: A basic cleaning rule base is built based on domain expert experience, covering dimensions such as data missing rate thresholds, logical consistency checks, and format standardization checks. For example, samples with a data missing rate > 30% are directly removed; the logical relationship between "project start time" and "project completion time" is checked, and abnormal data with a completion time earlier than the start time is removed; the format standardization of technical parameters is checked, and non-numerical parameters are converted to numerical values ​​or marked as invalid data.

[0031] Machine Learning Anomaly Detection: An anomaly detection model is constructed by introducing the Isolation Forest algorithm. This model belongs to the traditional machine learning algorithm and has the characteristics of fast training speed and strong adaptability to high-dimensional data. By training on historical normal data, the anomaly detection model can automatically identify outliers that deviate from the normal data distribution, such as exaggerated technology maturity indicators by enterprises or lagging industry parameters. The anomaly data identified by the anomaly detection model will enter the manual review process, and will be removed or corrected after confirmation by domain experts.

[0032] Semantic hashing deduplication: To address the problem of duplicate data, a deduplication algorithm based on semantic hashing is used. This algorithm extracts semantic features from the data and generates a fixed-length hash value. When the similarity of the hash values ​​of two data is ≥95%, they are determined to be duplicate data. In this case, the data with higher quality is retained to avoid the interference of redundant data on subsequent fusion and matching.

[0033] Due to the different data formats from different data sources, semantic definitions vary, requiring standardization to achieve cross-domain semantic alignment. This invention constructs a unified data standard dictionary for the scientific and technological field through a heterogeneous data standardization module, covering core dimensions such as technical terms, evaluation indicators, and data formats, and supports dynamic updates.

[0034] Structured data standardization: For structured data, field mapping and type conversion are performed according to a unified data standard dictionary. For example, fields such as "project budget" and "total funding" from different data sources are uniformly mapped to the "funding scale" field, and string-type monetary data is converted to numeric types, with the unit uniformly set to "ten thousand yuan".

[0035] Standardization of semi-structured / unstructured data: Extract key information (such as technical field, core technology, and patent applicant) from semi-structured data (such as patent abstracts) using algorithms such as word segmentation, part-of-speech tagging, and entity recognition in natural language processing.

[0036] For unstructured data (such as academic papers), deep semantic analysis is performed using a BERT pre-trained model to extract the core content of the paper, such as the research topic, technical solutions, and innovations. The extracted key information is then transformed into standardized feature vectors, which not only preserves the details of technical innovations but also achieves format consistency with structured data.

[0037] Feature engineering is key to improving matching accuracy. The multi-dimensional feature engineering module extracts features from both static and dynamic dimensions to build a comprehensive feature system, as detailed below: Static feature extraction: Static features are fixed indicators that can be directly quantified, including technical field, funding scale, partner qualifications, number of patents, and project cycle.

[0038] Dynamic feature extraction: Dynamic features are extracted through time series analysis and trend prediction techniques, which can reflect the changing trends and development potential of data. By using time series analysis algorithms (such as the ARIMA model) to analyze historical data, future trends can be predicted, including the potential for technology industrialization, the growth rate of patent citations, the changing trend of market demand, and the growth rate of R&D investment by partners. The introduction of dynamic features allows the matching process to consider not only the current state but also future development potential, thus improving the foresight of the matching.

[0039] The core objective of the fusion algorithm layer is to overcome the limitations of traditional algorithms. Therefore, a lightweight, interpretable, and adaptive fusion algorithm model was constructed, which can balance matching accuracy and efficiency and solve the problem of the "algorithm black box," as detailed below: The hybrid fusion algorithm model adopts a two-stage fusion strategy of shallow feature fusion and deep semantic fusion, and combines traditional machine learning and artificial intelligence general techniques to achieve efficient fusion of multi-dimensional features.

[0040] The first stage is shallow feature fusion. For structured data, an improved weighted Bayesian algorithm is used. The weights of traditional weighted Bayesian algorithms are mostly set manually, which is highly subjective. However, this invention dynamically allocates weights based on data quality. Data quality is quantified and scored through three dimensions: data completeness, accuracy, and timeliness. The higher the data quality score, the greater the weight. The calculation formula is: Weight of a feature = Data quality score of that feature / Sum of data quality scores of all features, which is used to avoid the subjective bias of manual setting.

[0041] For unstructured data, deep semantic feature extraction is performed using the BERT pre-trained model. The BERT model can capture the contextual semantic relationships in the text, and is especially suitable for semantic understanding of cross-domain technical documents. For example, it can identify the semantic relationship between "artificial intelligence" and "machine learning" and "deep learning" and extract more representative semantic feature vectors.

[0042] The second stage involves deep feature fusion. A deep learning model with an attention mechanism is introduced to fuse the shallow features of structured data extracted in the previous stage with the semantic features of unstructured data. The attention mechanism can automatically focus on core features and dynamically adjust the feature attention according to the needs of different application scenarios. For example, in the scenario of matching industry-academia-research projects, it automatically focuses on core features such as technological complementarity and resource complementarity; in the scenario of enterprise technology introduction, it automatically focuses on core features such as technology maturity and industrialization potential, thereby improving the fusion accuracy. At the same time, the model incorporates the idea of ​​ensemble learning, which can reduce the generalization error of a single model by voting on the prediction results of multiple sub-models, further improving the stability of the fusion.

[0043] To address the "black box" problem of deep learning models and meet the traceability requirements of scenarios such as government project approval and corporate decision-making, an algorithm interpretability enhancement mechanism embeds an attention weight visualization module into the deep learning model. The specific implementation is as follows: Weight allocation visualization: The attention weight of each feature during the feature fusion process is displayed in the form of a heatmap. The color intensity of the heatmap represents the weight magnitude, allowing users to intuitively see which features play a dominant role in the matching results.

[0044] Association rule visualization: The relationship between features is shown through a logical link diagram, such as the association path between "patent technology A" and "enterprise demand B", the causal relationship between "technology maturity" and "industrialization potential", etc., so as to clarify the core basis of the matching results.

[0045] Results traceability function: Users can click on any indicator in the matching results to view the calculation process, data source and weight allocation basis of the indicator, realizing full-link traceability from the matching results to the original data, and meeting the compliance requirements of scenarios such as government project review and enterprise cooperation decision-making.

[0046] The adaptive algorithm optimization module dynamically adjusts algorithm parameters based on data distribution and scene changes to ensure the model's adaptability in different data environments and application scenarios, while also optimizing inference speed to meet the needs of real-time matching.

[0047] Transfer learning parameter reuse: When a new data source or technical field is added, traditional models need to re-label a large amount of data for training, which is time-consuming and labor-intensive; while the adaptive algorithm optimization module reuses the parameters of the existing model through transfer learning technology, and only needs to label a small amount of new data to quickly adapt to the new scenario, which greatly shortens the training time.

[0048] Lightweight model optimization: Lightweight techniques such as model pruning and quantization are used to remove redundant parameters in the model and compress the model size while ensuring model accuracy. The optimized model has a significantly improved inference speed, which can meet the requirements of real-time matching.

[0049] Dynamic parameter adjustment: By monitoring the distribution changes of the data, if the data feature distribution offset is greater than the preset threshold, the parameter adjustment process is automatically triggered. The gradient descent algorithm is used to iteratively optimize the model parameters to ensure that the model always maintains a high fusion accuracy.

[0050] The core objective of the matching logic layer is to break through the static, single matching logic. Therefore, a matching mechanism combining multi-dimensional, dynamic weighting and implicit association mining is constructed to improve scenario adaptability and matching coverage, as detailed below: The multi-dimensional matching indicator module constructs a three-dimensional indicator system encompassing technology, value, and collaboration dimensions, comprehensively covering the core needs of matching technology projects, as detailed below: The technology dimension includes three secondary indicators: technology similarity, technology complementarity, and technology maturity. Technology similarity is calculated using the cosine similarity algorithm, based on standardized technology feature vectors to calculate similarity values ​​(range 0-1, with higher values ​​indicating higher similarity). Technology complementarity is calculated through the association degree of technology nodes in the knowledge graph. If the technology nodes of project A and project B are complementary in the knowledge graph, the complementarity score is high. Technology maturity adopts the Technology Readiness Level (TRL) grading standard, divided into 1-9 levels, with higher levels indicating higher maturity.

[0051] Value dimensions include three secondary indicators: industrialization potential, market prospects, and input-output ratio. Industrialization potential is calculated based on a weighted average of factors such as technological maturity, production process feasibility, and policy support. Market prospects are predicted using data such as market demand growth rate, competitive landscape, and consumer group size. Input-output ratio is calculated as the ratio of expected project returns to invested funds.

[0052] Collaboration dimension: This includes three secondary indicators: partner suitability, resource complementarity, and regional synergy. Partner suitability is assessed based on factors such as the partner's industry position, cooperation history, and credit rating; resource complementarity assesses the degree of complementarity between partners in terms of resources such as capital, technology, talent, and equipment; and regional synergy considers factors such as the geographical location of both parties, the distribution of industrial clusters, and the compatibility with regional policies.

[0053] The dynamic weight adjustment module adaptively adjusts the indicator weights based on the scenario, balancing preset templates and custom requirements to ensure that the weight settings fit the actual application scenario, as detailed below: Preset scenario weight templates: Preset weight templates are provided for typical scenarios such as government-industry-academia-research cooperation, enterprise technology introduction, and cross-border project cooperation.

[0054] Custom weight adjustment: Users can fine-tune the weights according to their personalized needs. Users can adjust the weight ratio of each indicator through the slider, and the system will calculate the adjusted matching results in real time to meet the matching needs of special scenarios. Automatic optimization through reinforcement learning: Combining historical matching data and user feedback data, the weight allocation is automatically optimized through reinforcement learning algorithms; user satisfaction with the matching results is used as a reward signal. When a user rates the matching result as "excellent", a positive reward is given to strengthen the current weight combination; when the rating is "unsatisfactory", a negative reward is given to adjust the weight combination, so that the weight setting gradually conforms to the actual application needs and continuously improves the matching accuracy.

[0055] The implicit association mining module integrates knowledge graph and semantic reasoning technologies, breaking through the limitations of traditional keyword matching. It can uncover implicit associations across domains and broaden the matching coverage, as detailed below: Construction of a Knowledge Graph in the Technology Field: The construction process of a knowledge graph encompasses three core stages: knowledge extraction, knowledge processing, and knowledge fusion. In the knowledge extraction stage, natural language processing technology is used to extract technical entities and relational entities (such as "Company A - Cooperation - University B" and "Technology C - Application - Industry D") from patent documents, academic papers, and technical documents. In the knowledge processing stage, the extracted entities are disambiguated, categorized, and their attributes are completed. For example, "AI" and "Artificial Intelligence" are unified into a single entity, and attributes such as the industry classification of enterprises and the authorization time of patents are completed. In the knowledge fusion stage, knowledge extracted from multiple sources is integrated to eliminate data conflicts and construct a unified knowledge graph in the technology field, covering multiple types of nodes such as technology, patents, enterprises, universities, and industry chains, as well as multiple types of relationships such as cooperation, application, and citation. Implicit Association Mining: This method utilizes Graph Neural Networks (GNNs) to deeply analyze the relationships between nodes in a knowledge graph, uncovering implicit cross-domain associations. For example, it can identify connections between artificial intelligence technology and drug molecule screening technology in the biomedical field, or between blockchain technology and fund tracking technology in supply chain finance. Then, based on semantic similarity reasoning, it achieves precise matching of implicitly related items. For instance, if a company needs "intelligent detection technology," the system not only matches items directly containing that keyword but also items involving semantically related technologies such as "machine vision detection" and "intelligent analysis of sensor data," thus broadening the matching coverage.

[0056] The core objective of the application layer is to improve the system's usability and scalability. Therefore, a modular, low-barrier, and fully interactive system platform has been built to implement the technical solutions, as detailed below: The role-based operation module can design personalized operation interfaces and functional modules to meet the different needs of different user groups, simplifying the operation process and allowing users to get started without professional technical background.

[0057] Enterprise-side features include project retrieval, requirement posting, matching result comparison, and collaboration management. Project retrieval supports multi-condition filtering, fuzzy search, and quick filtering using popular tags. The requirement posting function allows enterprises to fill in information such as technical requirements, cooperation expectations, and budget range, and the system automatically recommends suitable projects. The matching result comparison function displays the scores of various indicators of multiple candidate projects in multiple formats such as tables and radar charts to help enterprises make quick decisions.

[0058] Government-side functions: Provides functions such as project review, matching basis traceability, data statistical analysis, and policy adaptation recommendation; among them, the project review function supports batch review of project application materials and is the core basis for viewing matching results; the data statistical analysis function generates statistical reports such as the cooperation popularity of regional science and technology projects, the distribution of technical fields, and the effectiveness of industry-university-research cooperation; and the policy adaptation recommendation function can recommend suitable government support policies based on the characteristics of the project.

[0059] University-side functions: Provides functions such as scientific research achievement display, cooperation project matching, and achievement transformation tracking; among them, the scientific research achievement display supports uploading achievement materials such as patents, papers and technology prototypes, and can generate visual display pages; the cooperation project matching function can automatically match enterprises or government projects with cooperation needs and push matching information; and the achievement transformation tracking function can record data such as the progress of scientific research achievement transformation and economic benefits.

[0060] Modular extension modules, employing a microservice modular design, enhance system scalability and maintainability, adapting to the evolving technological needs of the technology sector, as detailed below: Microservice module decomposition: Core functions such as data access, data governance, algorithm fusion, and matching logic are decomposed into independent microservice modules, and each module communicates through API interfaces; for example, the data access microservice is responsible for accessing multi-source data and handling privacy, and the data governance microservice is responsible for data cleaning, standardization, and feature extraction. Each module can be deployed, upgraded, and maintained independently.

[0061] Flexible scalability: When adding a new data source, only the standardized interface of the data access microservice needs to be connected, without modifying other modules; when adding a new matching scenario, the scenario template of the dynamic weight adjustment module and the indicator system of the matching logic module can be configured to quickly adapt to the new scenario without the need for overall secondary development.

[0062] Emerging technology adaptation: Reserved interfaces support data access and algorithm adaptation in emerging technology fields. When new technical terms and evaluation indicators emerge in emerging technology fields, the need for rapid adaptation to technology iteration can be met by updating the unified data standard dictionary in the science and technology field and adjusting the feature extraction algorithm.

[0063] The end-to-end interactive module establishes a closed-loop mechanism for matching, evaluation, and optimization, driving continuous iterative optimization of the system through user feedback, as detailed below: Matching result evaluation: Users can rate the matching results with stars and indicate the reason for the rating. The system also supports text input of custom rating reasons. Feedback data collection: The system automatically collects user evaluation data and operation behavior data, and stores them in the feedback database as the core data support for system optimization; Closed-loop optimization execution: The feedback closed-loop module periodically analyzes the feedback data and optimizes the algorithm parameters and matching weights through reinforcement learning technology. It also updates the unified data standard dictionary and knowledge graph in the science and technology field and synchronizes the optimization results to modules at all levels to achieve continuous iterative improvement of system performance.

[0064] The feedback loop module is the core hub connecting the application layer with other layers. By collecting user feedback data, it drives continuous optimization of the data layer, data governance layer, fusion algorithm layer, and matching logic layer, forming a closed-loop iteration throughout the entire process, as detailed below: Data layer optimization: Based on user feedback regarding issues such as "inaccurate data" and "missing data," we adjusted the data access verification rules and optimized the desensitization strategy for privacy compliance processing mechanisms.

[0065] Data governance layer optimization: Based on user feedback regarding issues such as "unreasonable matching metrics" and "incomplete feature extraction," the data standard dictionary is updated, and the extraction dimensions and calculation methods of feature engineering are adjusted.

[0066] Fusion algorithm layer optimization: Based on user evaluation of the matching results, the weight allocation strategy of the fusion algorithm parameters is adjusted to optimize the model structure, thereby improving fusion accuracy and interpretability.

[0067] Matching logic layer optimization: Based on user feedback regarding issues such as "poor scenario adaptability" and "undiscovered implicit associations," the scenario weight template is updated, and the entity and relationship types of the knowledge graph are expanded to improve the rationality and coverage of the matching logic.

[0068] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0069] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A technology project matching system based on multi-dimensional data fusion, characterized in that, It includes a data layer, a data governance layer, a fusion algorithm layer, a matching logic layer, an application layer, and a feedback loop module, as detailed below: Data layer: Configured with a multi-source data adaptive access module and a privacy-compliant data processing module, enabling compliant access to multi-source heterogeneous data; Data governance layer: Configured with intelligent data cleaning module, heterogeneous data standardization module and multi-dimensional feature engineering module to complete data preprocessing; Fusion Algorithm Layer: Configures a hybrid fusion algorithm module, an algorithm interpretability enhancement module, and an adaptive algorithm optimization module to implement feature fusion; Matching logic layer: Configures a multi-dimensional matching indicator module, a dynamic weight adjustment module, and an implicit association mining module to achieve scenario-based matching; Application layer: Configure role-based operation modules, modular extension modules, and full-process interaction modules to implement practical applications; Feedback loop module: Collects user feedback data and optimizes system parameters and rules at each level through reinforcement learning.

2. The technology project matching system based on multi-dimensional data fusion according to claim 1, characterized in that: The role-based operation module at the application layer is designed with personalized interfaces and functional modules for enterprise R&D personnel, government administrators, and university researchers. The modular extension module adopts a microservice architecture, which can separate data access, data governance, algorithm fusion and matching logic into independent microservices, and reserves interfaces for emerging technology fields; The full-process interactive module allows users to rate the matching results with stars and annotate the reasons, and synchronizes the feedback data to the feedback closed-loop module in real time.

3. A technology project matching method based on multi-dimensional data fusion, applied to the technology project matching system based on multi-dimensional data fusion as described in claims 1-2, characterized in that, Includes the following steps: S1: Multi-source heterogeneous data access, adopting multi-source adaptation and privacy compliance mechanisms, accessing structured, semi-structured and unstructured multi-type data, achieving cross-node data feature fusion through the fusion of federated learning and differential privacy technologies, and simultaneously de-identifying sensitive information; S2: Data governance preprocessing, through cleaning, standardization and feature extraction operations, performs intelligent cleaning, heterogeneous data standardization transformation and static and dynamic feature extraction on the data accessed in S1 to generate standardized feature vectors; S3: Multi-dimensional feature fusion, which adopts shallow feature fusion and deep semantic fusion, combines improved traditional machine learning algorithms with general artificial intelligence techniques to build an interpretable adaptive fusion algorithm model for fusion processing of features extracted in S2; S4: Dynamic scenario-based matching, based on a three-dimensional indicator system of technology, value and collaboration dimensions, achieves accurate matching of technology projects in multiple scenarios through dynamic weight adjustment and implicit correlation mining; S5: Matching result feedback optimization. Collect user feedback on matching results and combine reinforcement learning techniques to optimize algorithm parameters, matching weights, and data standards to form a closed-loop iteration throughout the entire process.

4. The technology project matching method based on multi-dimensional data fusion according to claim 3, characterized in that: The multi-source data adaptive access module can adapt to the data update frequency of different data sources. In the privacy-compliant data processing mechanism, each data source node keeps the data locally stored and only transmits the feature fusion results. The intensity of differential privacy noise addition is dynamically adjusted according to the data sensitivity level.

5. The technology project matching method based on multi-dimensional data fusion according to claim 4, characterized in that: The intelligent data cleaning module adopts a fusion strategy of rule engine and machine learning, as detailed below: The rule engine builds basic cleaning rules based on the experience of domain experts, and machine learning identifies abnormal data through the isolated forest algorithm and removes duplicate data based on the semantic hash deduplication algorithm. The heterogeneous data standardization module is used to build a standard dictionary for unified data in the field of science and technology, and to transform semi-structured or unstructured data into standardized feature vectors through natural language processing technology.

6. The technology project matching method based on multi-dimensional data fusion according to claim 5, characterized in that: The static features extracted by the multi-dimensional feature engineering module include the technical field, funding scale, and quantitative indicators of the qualifications of partners; the dynamic features are extracted through time series analysis and trend prediction, specifically including the potential for technology industrialization, the patent citation growth rate, and the trend of market demand changes.

7. The technology project matching method based on multi-dimensional data fusion according to claim 6, characterized in that: In the first stage, the hybrid fusion algorithm module uses an improved weighted Bayes algorithm for structured data and dynamically assigns weights based on data quality. For unstructured data, deep semantic features are extracted using a BERT pre-trained model. The second stage involves introducing a deep learning model with an attention mechanism to fuse multi-dimensional features, which is used to automatically focus on core features.

8. The technology project matching method based on multi-dimensional data fusion according to claim 7, characterized in that: The algorithm interpretability enhancement module uses an attention weight visualization mechanism to display the weight allocation and association rules in the feature fusion process in the form of heatmaps and logical link diagrams. The adaptive algorithm optimization module reuses existing model parameters through transfer learning techniques and introduces a lightweight model structure to optimize inference speed in order to meet the requirements of real-time matching.

9. The technology project matching method based on multi-dimensional data fusion according to claim 8, characterized in that: In S4, the three-dimensional indicator system includes the following dimensions: technology similarity, technology complementarity, and technology maturity; value dimension including industrialization potential, market prospects, and input-output ratio; and collaboration dimension including partner compatibility, resource complementarity, and regional collaboration. The dynamic weight adjustment module can preset typical scenario weight templates and automatically optimize weight allocation by combining reinforcement learning algorithms with historical matching data.

10. The technology project matching method based on multi-dimensional data fusion according to claim 9, characterized in that: In S4, the implicit association mining module integrates knowledge graph and semantic reasoning technologies to construct a science and technology knowledge graph covering technology, patents, enterprises, universities, and industry chain nodes. It mines cross-domain implicit associations through graph neural networks and achieves implicit association item matching based on semantic similarity reasoning.