Full-process intelligent medical and invasive service data processing method and system
By constructing knowledge graphs and dynamic profiles, the problems of data silos and fragmentation in science and technology innovation service platforms have been solved, enabling intelligent science and technology innovation services throughout the entire process, improving the depth of data analysis and risk warning capabilities, and providing personalized service strategy support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-04-10
AI Technical Summary
Existing science and technology innovation service platforms suffer from problems such as data silos and fragmentation, insufficient depth of analysis, crude matching of services with needs, and lack of risk warning capabilities, thus failing to meet the needs of intelligent operation throughout the entire process.
By collecting and integrating multi-source heterogeneous science and technology innovation data to construct a knowledge graph, dynamic science and technology innovation profiles are generated, enabling intelligent demand perception and precise service matching, and covering the management and risk warning of the entire life cycle of science and technology innovation projects.
It achieves efficient integration and dynamic quantitative evaluation of multi-source data, enhances the depth and foresight of data analysis, provides personalized science and technology innovation service strategies and real-time risk warnings, and covers the entire process of support from idea to industrialization.
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and more specifically, to a method and system for processing data for intelligent science and technology innovation services throughout the entire process. Background Technology
[0002] Technological innovation is the core driving force for social progress and high-quality economic development. However, from the initial emergence of ideas, technology research and development, and the transformation of results into industrialization, scientific and technological innovation is a long, complex, and uncertain process. Throughout this entire process, various innovation entities (such as enterprises, research institutions, R&D personnel, and investment institutions) heavily rely on accurate, timely, and comprehensive data support and professional services. Currently, some scientific and technological innovation service platforms or tools exist in the market, but they have significant limitations in data processing and service capabilities, failing to meet the needs of a fully integrated and intelligent process.
[0003] First, the problems of data silos and fragmentation are prominent. The types of data required for scientific and technological innovation activities are extremely diverse, including patents, papers, projects, policies, corporate information, and market data. This data is scattered across hundreds or even thousands of independent, heterogeneous databases and platforms. Existing solutions are mostly vertical tools targeting single data sources, such as patent search systems or academic paper databases. These are isolated from each other and lack an effective integration mechanism. When a user wants to assess the commercial prospects of a technology, they need to search for patent layouts in patent databases, review research progress in paper databases, find relevant companies in corporate information databases, check support directions on policy platforms, and finally manually integrate the information. This model is extremely inefficient and prone to decision-making errors due to incomplete information. More importantly, the deep information in unstructured data (such as full-text PDF patents and scanned policy documents) cannot be effectively mined and correlated, creating a dilemma of "information at hand, but unable to be effectively utilized."
[0004] Secondly, the depth of data analysis is insufficient, lacking dynamic and forward-looking insights. Most existing service platforms offer basic information retrieval and simple statistical functions, their analytical capabilities remaining at the level of "describing the past," unable to "diagnose the present" or "predict the future." For example, they can tell users how many patents a company owns, but cannot quantify the company's true technological strength and position in its industry, assess the quality and risk of its patent portfolio, or dynamically track the evolution of its technological path. Existing systems lack effective quantitative models and evaluation methods for technology lifecycles, convergence trends, and potential disruptive opportunities. This results in superficial analytical reports that fail to provide in-depth, forward-looking intellectual support for high-risk scientific and technological innovation decisions.
[0005] Secondly, the matching of services with needs is rudimentary and fails to cover the entire project lifecycle. Traditional science and technology innovation services heavily rely on the personal experience and networks of consultants, resulting in non-standardized, costly, and poorly scalable service processes. While some platforms attempt resource matching, their matching logic is usually based on simple keyword matching, failing to deeply understand the complex and multi-layered real needs of users, and unable to accurately assess the value and personalize the selection of massive resources. This leads to recommendations that are often not very relevant, requiring users to spend a lot of time on secondary filtering. Furthermore, most existing systems only focus on the "resource matching" stage, failing to extend services to the entire lifecycle of science and technology innovation projects. From idea evaluation, team building, R&D management, intellectual property layout, to investment and financing matchmaking and market promotion, the data, processes, and management at each stage are disconnected, failing to form a closed-loop, continuously optimized intelligent service system.
[0006] Finally, there is a lack of risk warning capabilities and a low level of intelligence in decision support. Scientific and technological innovation activities are inherently accompanied by extremely high technological, market, and intellectual property risks. Existing systems generally lack real-time risk monitoring and prediction capabilities based on big data. Users often only passively discover problems after patents are invalidated, infringement lawsuits are filed, or technologies become obsolete, by which time irreparable losses have already occurred. Simultaneously, the entire decision-making process remains primarily based on manual analysis; the system fails to act as an "intelligent assistant," unable to proactively perceive needs, warn of risks, or provide data-driven alternatives. This makes it difficult to achieve a qualitative leap in the efficiency and success rate of scientific and technological innovation management.
[0007] Therefore, there is an urgent need in this field for an integrated intelligent service solution that can break down data barriers, achieve deep integration and intelligent analysis, and cover the entire process from ideation to industrialization. This invention was developed against this backdrop, aiming to overcome the aforementioned shortcomings of existing technologies and provide a novel data processing method and system. Existing technologies urgently need improvement to address the above-mentioned problems. Summary of the Invention
[0008] The purpose of this application is to provide a method and system for processing data for intelligent science and technology innovation services throughout the entire process, which can solve the above-mentioned technical problems.
[0009] This application provides a method for processing data in a fully intelligent science and technology innovation service, including the following steps:
[0010] S100: Multi-source Heterogeneous Science and Technology Innovation Data Collection and Fusion: Through multiple configured data crawler agents and API interface gateways, structured and unstructured raw science and technology innovation data are continuously and automatically collected from publicly available internet databases, third-party commercial databases, government e-government platforms, academic publishing institutions, and enterprise self-submission ports. This raw science and technology innovation data includes, but is not limited to: patent texts and legal status data, full-text data of scientific papers and research reports, data on the establishment and completion of scientific research projects, technical standard documents, business registration and operating information, investment and financing event data, industrial chain and supply chain mapping data, science and technology policy and regulatory texts, market research and industry analysis reports, and so on. This includes data on physical space innovation activities collected through smart terminals; for unstructured data, a deep learning-based document parsing engine is used to perform OCR recognition and key information extraction on PDFs, Word documents, images, and scanned documents, and the extracted entities, attributes, and relationships are standardized and described in a unified JSON-LD format; at the same time, a data fusion engine is launched, based on a predefined entity alignment algorithm, to identify, disambiguate, and merge the same entity from different data sources, generating a "science and technology innovation entity object" with a globally unique identifier, and linking all related data to this object, thereby building a large-scale, highly timely knowledge graph in the field of science and technology innovation at the underlying level;
[0011] S200: Dynamic Sci-Tech Innovation Profile Generation and Quantitative Evaluation: Based on the knowledge graph of the science and technology innovation field constructed in S100, multi-dimensional dynamic profiles are generated for the core entities in the graph. The core entities include at least: innovation entities, core technologies, and scientific research talents. Among them, the profile generated for "innovation entities" includes its technological strength dimension, R&D activity dimension, commercial value dimension, risk dimension, and collaborative innovation network dimension; the profile generated for "core technologies" includes its technology life cycle stage, technology concentration, technology integration degree, market penetration rate, and future potential index; the profile generated for "scientific research talents" includes its research field distribution, academic influence, technology transfer capability, and cooperation network strength. This step calls a series of quantitative evaluation models, including a patent value prediction model based on Transformer, an R&D trend prediction model based on time series analysis, a technology field segmentation model based on community discovery algorithm, and a comprehensive competitiveness evaluation model based on multi-source indicator fusion. Each entity is automatically scored and rated, and the evaluation results are dynamically updated to the knowledge graph as new attributes, forming a dynamic profile system that can trace the past and predict the future.
[0012] S300: Intelligent Demand Perception and Precise Service Matching: A multimodal user demand perception interface is constructed, capable of receiving and parsing service demands implicitly expressed by users through natural language, form submissions, or historical behavioral data. A dedicated demand semantic understanding model transforms user input into a structured "demand vector," explicitly describing the target technology field, expected innovation stage, resource constraints, and desired business objectives of the required service. Subsequently, a large-scale service matching engine is activated, performing multiple rounds of semantic similarity calculations and logical condition matching between the user's demand vector and the dynamic science and technology innovation profile generated in S200. The engine then selects the most suitable service solutions from a pre-built "science and technology innovation service resource library." These service solutions constitute a composite recommendation list, including but not limited to: the most suitable technology partners or R&D teams for collaboration, key technology directions and gaps for development, high-value targets for investment or acquisition, government-funded projects and certifications available for application, and patent warning information for risk mitigation. Finally, a highly personalized science and technology innovation service strategy analysis report is generated for each user and presented interactively through a visual interface.
[0013] As a preferred option, the "data fusion engine" in step S100 specifically performs the following operations:
[0014] S110: Entity Recognition and Standardization: Identify key entities in raw data obtained from different sources. These key entities include: company names, personal names, technical terms, geographical locations, and organization names. A BERT-based named entity recognition model is used for initial identification, supplemented by a large-scale domain dictionary for precise matching. For identified entities, the entity standardization service is invoked to uniformly map various aliases, abbreviations, and misspellings to the official full name or standard terminology. For example, "Huawei Technologies Co., Ltd." and "Huawei Tech. Co., Ltd." are both standardized to "Huawei Technologies Co., Ltd."
[0015] S120: Entity Linking and Disambiguation: For the standardized entities identified in S110, calculate their similarity to existing entities in the knowledge graph. The similarity calculation comprehensively utilizes multiple features, including: character similarity of entity names, value similarity of entity attributes, and semantic similarity of entities in their respective text contexts. A dynamic threshold is set. When the comprehensive similarity exceeds the threshold, the newly collected entity is determined to be a reference to an existing entity in the knowledge graph, and the new data is linked as a new attribute or relationship of the existing entity. When the similarity is below the threshold or there are multiple candidate entities with high similarity, the entity disambiguation model based on graph neural networks is activated. By analyzing the local network structure of the entity, it is determined to be the only entity that it most likely points to. If it still cannot be determined, it is marked as a potential new entity awaiting manual review.
[0016] S130: Relation Extraction and Knowledge Graph Update Operation: For each newly acquired document, a relation extraction pipeline combining rule-based and deep learning is used to extract semantic relationships between entities from the text. These semantic relationships include, but are not limited to: "Company A - Application - Patent B", "Talent C - Employed in - Institution D", "Technology E - Belongs to - Technology Field F", "Project G - Received - Fund H Funding". The extracted relationships are represented in the form of "subject-verb-object" triples, and after confidence calculation, high-confidence triples are added to the knowledge graph. At the same time, a consistency check is performed on the graph, and a descriptive logic inference engine is used to detect and report logical conflicts such as "a patent that has been declared invalid cannot be simultaneously in an authorized maintenance state", to ensure the logical consistency of the graph data.
[0017] As a preferred option, step S200, "generating a technological strength profile for the innovation entity," specifically includes:
[0018] S210: Quantification of Technical Breadth and Depth: Based on all patents and software copyrights owned by the innovation entity, the patents are classified into multiple levels using the International Patent Classification (IPC) and the Cooperative Patent Classification (CPC). Technical breadth is measured by calculating the number of main classifications and subclasses covered by the technical solutions and the uniformity of their distribution. Technical depth is determined by analyzing the comprehensive performance of the patents in terms of citation count, number of claims, number of countries involved, etc., and comparing it with the baseline level of the technical field.
[0019] S220: Technical Quality and Impact Assessment Operation: Construct a comprehensive patent quality evaluation model. The input features of this model include: patent family size, patent citation frequency and the quality of citing authors, patent litigation and transfer history, language features and scope of protection of claims, and patent lifespan. Simultaneously, analyze the citation count of the subject's high-level academic papers, the impact factor of the journals in which they are published, and whether they are cited by industry standards to assess its basic research impact. The evaluation results of patents and papers are weighted and integrated to output a technical quality rating from "emerging level" to "leading level."
[0020] S230: R&D Collaborative Network Analysis Operation: Extract collaborative subgraphs centered on the innovation entity from the knowledge graph, analyze its relationships with universities, research institutes, other enterprises, and other entities such as joint patent applications, collaborative paper publications, and joint project undertakings; calculate its network centrality indicators, including degree centrality and betweenness centrality, to measure its pivotal position in the innovation network; identify its core partners and potential "bridging" partners who have not yet established connections but whose technologies are highly complementary, and quantify these network characteristics into a collaborative innovation potential index as an important component of its profile.
[0021] Preferably, the "demand semantic understanding model" in step S300 is a deep neural network based on a domain-pre-trained language model, and its specific implementation process is as follows:
[0022] S310: Multi-turn conversational requirement clarification operation: When a user raises an initial requirement in vague natural language, a multi-turn dialogue management module is activated. This module, based on a preset requirement clarification tree, proactively asks follow-up questions to the user. For example, when a user asks "Looking for investment opportunities in the AI chip field," the system will ask follow-up questions in sequence: "Are you interested in training chips or inference chips?", "Do you have specific requirements for the manufacturing process?", "What stage of investment do you expect: early stage, growth stage, or pre-IPO?". Through interactive dialogue, the vague requirement is gradually concretized into a set of clear constraints.
[0023] S320: Demand Vector Structure Construction Operation: The explicit information collected in S310, along with the structured form data directly submitted by users, are input into the demand semantic understanding model. This model first encodes the text, then uses a multi-head attention mechanism to focus on key entities, state words, and intent words in the demand description. Finally, through a fully connected layer, the encoded semantic information is mapped to a high-dimensional, structured "demand vector," where different dimensions correspond to quantitative indicators such as technology field, innovation stage, market size, risk preference, and resource budget.
[0024] S330: Dynamic Generation of Demand Profile: The structured demand vector generated in S320 is combined with the user's historical query records, collection behavior, and feedback data on past recommendation results. Using a hybrid algorithm of collaborative filtering and content filtering, the user's long-term interests and preferences are modeled to generate a dynamically updated "user demand profile". This profile can capture the evolution trend of user demand preferences and be used to improve the personalization and foresight of recommendations in subsequent service matching.
[0025] As a preferred option, the "large-scale service matching engine" in step S300 employs a matching strategy that combines hierarchical filtering and fine-grained ranking, specifically including:
[0026] S340: Candidate set coarse filtering operation: Based on the hard constraints in the user demand vector, a first round of rapid filtering is performed on the massive entities in the knowledge graph; the hard constraints include: precise matching of technical fields, geographical area restrictions, range of years of enterprise establishment, and requirements for the presence or absence of intellectual property rights; through the index query of the graph database, a large number of irrelevant entities are quickly eliminated, forming a primary candidate set with reduced size;
[0027] S350: Multi-dimensional Semantic Refinement Operation: For each entity in the initial candidate set, calculate the matching degree between its dynamic science and technology innovation profile and the user demand vector in various dimensions. This matching degree is a composite score, calculated by weighting multiple sub-scores such as technology matching degree, business prospect matching degree, risk controllability matching degree, and cooperation feasibility matching degree. Among them, the technology matching degree is determined by comparing the cosine similarity between the keyword vector in the technology profile and the demand vector; the business prospect matching degree is calculated by analyzing the degree of consistency between the growth forecast of the market in which the entity is located and the user's investment preferences; the risk matching degree integrates the evaluation results of multiple factors such as intellectual property stability, market competition pattern, and policy compliance.
[0028] S360: Personalized Recommendation List Generation and Explanation: The composite matching scores calculated in S350 are sorted in descending order, and the Top-N entities are selected as the final recommendation results. For each recommendation result, the system automatically generates a "Recommendation Reason" explanation, which is presented in an interpretable manner, such as: "Company A is recommended because it has 5 core patents in the 'in-memory computing' field that you are interested in, and its R&D investment growth rate has exceeded 50% in the past three years. Its CEO has had a successful collaboration with your designated technical expert B." At the same time, a comparative analysis function is provided, allowing users to compare multiple recommended entities side by side on key indicators to assist in decision-making.
[0029] Preferably, the method further includes:
[0030] S400: Full Lifecycle Management and Management Steps for Science and Technology Innovation Projects: This system creates a dedicated online project management space for users, digitally managing the entire process from idea generation, technology verification, product development to market promotion. This space integrates project planning tools, task allocation and progress tracking modules, a collaborative document editing platform, and blockchain-based results storage services. The system automatically extracts key node data from the project process, such as experimental data records, code submission logs, test reports, and user feedback, and links them to the dynamic profile in S200, updating the status of relevant project entities in real time. S500: Intelligent Decision Support and Risk Warning Steps: Based on continuously updated knowledge graphs and project data, a risk prediction model is run to proactively identify potential risks encountered during the science and technology innovation process. These risks include: technology R&D risks, intellectual property risks, market competition risks, policy compliance risks, and supply chain risks. When the model predicts that the probability of a certain risk event exceeds a preset threshold, an early warning mechanism is immediately triggered, notifying users through message pushes, emails, dashboard highlighting, etc., along with a risk analysis report and response strategy suggestions, achieving a shift from passive response to proactive prevention.
[0031] As a preferred option, the "risk prediction model" in step S500 is an ensemble learning model, and its specific construction and operation are as follows:
[0032] S510: Risk Characterization Engineering Operations: Extract hundreds of characteristic variables related to specific risks from knowledge graphs and project management spaces; for technology R&D risks, characteristics include: maturity of the technology roadmap, past success rate of the R&D team, availability rate of key equipment, and historical breakthrough probability of technological bottlenecks; for intellectual property risks, characteristics include: remaining protection period of core patents, comprehensiveness of patent family layout, existence of potential infringement lawsuits, and strength of competitors' patent walls;
[0033] S520: Multi-model ensemble prediction operation: Three different types of base models, XGBoost, LightGBM, and deep neural networks, are used to train the feature set constructed in S510. Each base model learns the risk occurrence pattern from different perspectives. Then, through a stacking ensemble strategy, the prediction results of the three base models are used as new features to train a meta-model for the final risk probability prediction. This ensemble method can effectively reduce the overfitting risk of a single model and improve the generalization ability and robustness of the prediction.
[0034] S530: Warning Trigger and Strategy Recommendation Operation: A dynamic probability threshold is set for each type of risk, which can be personalized according to the user's risk tolerance. When the predicted probability exceeds the threshold, the system not only issues an alarm but also automatically retrieves matching response plans from the "Risk Response Strategy Knowledge Base." The response plans are a structured case library that records effective measures and their results taken in similar risk scenarios in the past. For example, "Faced with the expiration of core patents, we recommend: Option 1: Initiate peripheral patent layout; Option 2: Explore patent licensing and transfer; Option 3: Promote the integration of technical standards," thereby providing users with immediate decision support.
[0035] Preferably, the method further includes:
[0036] S600: Model Co-evolution Steps Based on Federated Learning: To continuously improve the performance of various AI models within the system while protecting the data privacy of all participants, a federated learning framework is introduced. Under this framework, the system's central server distributes the initialized global model to multiple participating clients, which are science and technology enterprises or research institutions that agree to participate in the co-construction of the model. Each client trains the model locally using its own private data and uploads the updated values of the model parameters to the central server after encryption, rather than the original data itself. The central server securely aggregates the collected parameter updates to generate an improved global model, which is then distributed to each client. Through multiple iterations, the model can learn a wider range of data distributions, thereby achieving co-evolution of performance.
[0037] A full-process intelligent science and technology innovation service data processing system, characterized in that the system comprises:
[0038] Data Acquisition and Fusion Subsystem: This subsystem deploys a configurable distributed web crawler cluster, capable of collecting data from hundreds of predefined data sources in a high-concurrency, high-fault-tolerant manner, either on a scheduled or real-time basis. It includes an API gateway management module for unified management and scheduling of calls to various third-party data interfaces, including authentication, traffic control, and billing management. It also includes a high-performance data cleaning and standardization pipeline, integrating the aforementioned document parsing engine and entity alignment algorithm, responsible for transforming raw data into standardized triples that can be used to build knowledge graphs.
[0039] Knowledge Graph Storage and Computation Subsystem: This subsystem adopts a hybrid storage architecture, using a native graph database to store entity and relation data to support efficient graph traversal and complex relation queries; simultaneously, it uses a distributed columnar database to store massive amounts of time-series metric data and full-text document indexes to support large-scale analysis and computation; this subsystem also has a built-in graph computation engine capable of executing complex graph algorithms such as PageRank, community detection, and shortest path, providing underlying computing power support for profile generation and recommendation computation in upper-layer applications;
[0040] Intelligent Analysis and Service Matching Subsystem: This subsystem is the core intelligent hub of the system. It encapsulates all the models and algorithms described in claims 1 to 8 in a microservice architecture, including dynamic profile generation model, demand semantic understanding model, service matching engine, risk prediction model, etc. Each model runs as an independent and scalable microservice and communicates with other services through RESTful API or gRPC protocol. The subsystem also has a unified model management platform responsible for version control, A / B testing, performance monitoring and online hot updates of all AI models.
[0041] Preferably, the system further includes:
[0042] Visual Interaction and Portal Subsystem: This subsystem is designed for end users, providing a web-based graphical user interface. Its core components include: a global science and technology innovation situation awareness dashboard, which macroscopically displays the innovation vitality of a specific region or technology field in the form of geographic information maps, technology development trend curves, and bubble charts showing the distribution of innovation entities; a personalized workbench, providing each user with a customized dashboard that centrally displays the dynamics of entities they are interested in, project progress, early warning information, and system recommendations; an interactive analysis studio, allowing users to freely combine analysis dimensions through drag-and-drop, conduct self-service data exploration, and generate exportable analysis reports with one click; and an immersive virtual collaboration space, integrating VR / AR technology to provide research teams distributed in different regions with a 3D virtual environment where virtual avatars can communicate, collaborate on design, and demonstrate results.
[0043] The beneficial effects of this invention are:
[0044] This application provides a data processing method and system for full-process intelligent science and technology innovation services. By collecting and integrating multi-source heterogeneous science and technology innovation data to construct a knowledge graph, dynamic science and technology innovation profiles are generated and precise service matching is achieved. This solves the problems of data silos and fragmentation, improves the depth and foresight of data analysis, and has the advantages of efficiently integrating multi-source data, dynamically quantifying and evaluating innovation entities, intelligently matching science and technology innovation needs, and covering the entire life cycle. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown herein can generally be arranged and designed in various different configurations.
[0046] Therefore, the following detailed description of the embodiments in this application is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0047] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship commonly used when the product is in use. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. In addition, the terms "first," "second," and "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0048] Furthermore, terms such as "horizontal," "vertical," and "sag" do not imply that components must be absolutely horizontal or suspended, but rather that they can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," not that the structure must be completely horizontal, but can be slightly tilted.
[0049] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0050] In existing technologies, technological innovation activities involve the integration and analysis of multi-source heterogeneous data, but traditional methods suffer from data silos and fragmentation. For example, when a company assesses the commercial prospects of a new energy storage technology, it needs to search patent portfolios in patent databases, review research progress in academic papers, find relevant companies in enterprise information databases, check policy platforms for support directions, and finally manually integrate the information. In this model, deep information in unstructured data cannot be effectively mined, leading to inefficient decision-making and incomplete information. Existing systems' analytical capabilities are limited to basic statistics, unable to quantify technological strength or predict trends. Service matching relies on simple keyword searches, recommendation results are weakly relevant, and there is a lack of end-to-end risk management capabilities.
[0051] To address these issues, the inventors first observed the bottlenecks in cross-platform data integration and realized that entity recognition and alignment are key to eliminating data silos. For dynamic evaluation needs, they found that traditional static indicators cannot reflect technological evolution trends, necessitating the establishment of time-series analysis models. Regarding service matching, they recognized that the structured transformation of fuzzy demands is a prerequisite for achieving accurate recommendations. Through systematic analysis of pain points throughout the entire science and technology innovation process, they gradually formed a technical roadmap encompassing multi-source data fusion, dynamic profile construction, and intelligent demand mapping.
[0052] Therefore, this application proposes a full-process intelligent science and technology innovation service data processing method, including the following steps: Collecting structured and unstructured raw science and technology innovation data from public internet databases, third-party commercial databases, government platforms, academic publishing institutions, and enterprise self-submission ports by configuring multiple data crawler agents and API interface gateways; using a deep learning-based document parsing engine to perform OCR recognition and key information extraction on unstructured data, and standardizing the description of entities, attributes, and relationships in JSON-LD format; launching a data fusion engine to generate science and technology innovation entity objects with globally unique identifiers based on entity alignment algorithms, and constructing a knowledge graph in the science and technology innovation field; generating multi-dimensional dynamic profiles for innovation entities, core technologies, and research talents based on the knowledge graph, and calling a quantitative evaluation model for automated scoring and rating; constructing a multimodal user demand perception interface, generating structured demand vectors through a demand semantic understanding model, using a service matching engine to perform multi-round semantic similarity calculation and logical condition matching, filtering out a composite recommendation list, and generating a personalized analysis report.
[0053] Among these, multi-source heterogeneous data acquisition refers to cross-platform data crawling through distributed crawler clusters and API gateways. Specifically, this can be achieved using the Scrapy framework combined with the OAuth 2.0 protocol, addressing the incomplete coverage issues of traditional single data sources. The document parsing engine refers to a hybrid model based on convolutional neural networks and recurrent neural networks. Specifically, the LayoutLM model can be used to parse PDF documents, improving the utilization of unstructured data. The entity alignment algorithm refers to a matching mechanism that integrates character similarity and semantic similarity. Specifically, a deep matching model based on the Siamese network can be used to ensure unified entity identification across data sources. Dynamic profile generation refers to an evaluation system that integrates time-series analysis and graph computation. Specifically, the GraphSAGE algorithm can be used to capture the evolution of entity relationships and predict technology trends. Demand vector transformation refers to mapping natural language demands into a structured feature space. Specifically, the BERT model combined with domain knowledge distillation can be used to achieve semantic understanding and improve the accuracy of demand representation. The service matching engine refers to a recommendation system that combines hierarchical filtering and multi-dimensional ranking. Specifically, a retrieval architecture combining Elasticsearch and Faiss can be used to balance efficiency and accuracy.
[0054] Specifically, in the data collection stage, a distributed crawler cluster is used to synchronously obtain multi-source data. The document parsing engine deeply analyzes unstructured data such as patent full texts and scanning policies, and extracts key information such as technical features and legal status. The entity alignment module establishes cross-database entity mapping relationships by comparing features such as enterprise name variants and patent applicant association relationships. For example, scientific research institutions with different expressions are unified and merged. During the construction of the knowledge graph, the Neo4j graph database is used to store entity relationships, and HBase is used to store time series index data to support complex queries and analyses. When generating dynamic portraits, the technical width is calculated based on the distribution of patent classification numbers, the technical quality is evaluated by combining the citation frequency and survival years of patents, and the LSTM model is used to predict the R & D trend. The requirement understanding module clarifies ambiguous requirements through multiple rounds of conversations, and transforms "finding artificial intelligence medical applications" into specific technical fields, development stages, and resource constraints. In the service matching stage, candidate entities that do not meet the hard conditions are first filtered, and then the weighted scores of technical matching degree and business prospect matching degree are calculated. Finally, a priority list including recommended reasons is generated.
[0055] Compared with the prior art, traditional methods rely on manual integration of multiple isolated databases. This solution realizes cross-platform entity alignment through an automated data fusion engine, shortening the data preparation time from weeks to real-time updates. Existing systems use fixed metrics to evaluate the technical value. This solution introduces a time series model to capture changes in the technology life cycle and can warn of the risk of technology decline. Traditional recommendation systems have insufficient accuracy based on keyword matching. This solution realizes the comprehensive evaluation of technical adaptability and commercial feasibility through multi-dimensional similarity calculation of requirement vectors and portraits. Existing tools only focus on a single service link. This solution covers the entire process from data collection to project management and establishes a decision-making closed loop.
[0056] Through the above technical solutions, this application realizes the automated integration and standardized processing of multi-source heterogeneous scientific and technological innovation data, and solves the problem of low data utilization rate of traditional methods. The dynamic portrait system can continuously track the technology evolution trend and overcome the lagging defect of the static evaluation model. The structured requirement vector and multi-dimensional matching strategy significantly improve the accuracy of service recommendation and reduce the manual screening cost. The full-process management module realizes the seamless connection of data in each stage of scientific and technological innovation projects, providing data support for risk warning and decision-making optimization.
[0057] This application further proposes a specific operation process for entity recognition and fusion in multi-source heterogeneous scientific and technological innovation data.
[0058] Among them, entity recognition and standardization operation refers to the identification and normalization of technical terms such as company names and personal names in the original data. Specifically, it can be achieved by using a BERT-based named entity recognition model combined with domain dictionary matching. By mapping entity names in different forms to unified standard terms, the ambiguity caused by aliases or abbreviations can be eliminated.
[0059] Entity linking and disambiguation refers to determining whether a newly identified entity is the same as an existing entity in the knowledge graph. Specifically, this can be achieved by combining multi-dimensional calculations of character similarity, attribute similarity, and contextual semantic similarity with dynamic threshold determination. When there are multiple candidate entities, graph neural networks are introduced to analyze the local network structure and use the topological relationships between entities to assist in disambiguation decision-making, thus solving the problem of difficult entity association across data sources.
[0060] Relation extraction and graph update operations refer to extracting semantic relationships between entities from text. Specifically, this can be achieved using a hybrid extraction pipeline that combines rule matching and deep learning models. High-confidence relation triples are selected through confidence calculation, and logical conflicts are detected using a descriptive logic inference engine to ensure the logical consistency between new data and existing graphs.
[0061] Specifically, in the entity recognition stage, a pre-trained BERT model captures deep semantic features of the text and, combined with a list of professional terms from a domain dictionary, accurately identifies key entities from different data sources. For name variations in the recognition results, a standardization service is invoked to map them to a unified identifier; for example, different language formats of company names are unified into the official registered name. In the entity disambiguation stage, name similarity, attribute matching degree, and contextual semantic similarity are comprehensively calculated. When the comprehensive score exceeds a dynamic threshold, the entity is determined to be the same entity; otherwise, a graph neural network model is activated to analyze the entity's association path in the knowledge graph, and the uniqueness of the entity is determined through topological features such as cooperative networks and technical associations. In the relation extraction stage, a rule engine based on dependency parsing is used to extract explicit relations, while a deep learning model is used to capture implicit semantic associations in the text. The extracted triples are added to the knowledge graph after confidence evaluation, and the rationality of relations such as patent status and technical affiliation is verified through a logic inference engine.
[0062] Compared to existing technologies, traditional methods typically rely solely on single-text similarity for entity alignment, failing to effectively handle complex variations across data sources, and lacking utilization of entity network relationships in the disambiguation process. This proposed solution significantly improves the accuracy of entity linking by combining multi-dimensional similarity calculation with graph structure analysis. Furthermore, it employs a hybrid relation extraction strategy and logical consistency checks to ensure the reliability of dynamic knowledge graph updates.
[0063] Through the above technical solution, this application effectively solves the problem of recognition error caused by differences in entity representation in multi-source data, realizes accurate entity alignment and relationship integration across data sources, constructs a logically consistent large-scale knowledge graph, and provides a high-quality data foundation for subsequent science and technology innovation services.
[0064] This application further proposes specific methods for generating a profile of the technological strength dimensions of innovation entities, including quantitative operations on the breadth and depth of technology, assessment operations on the quality and impact of technology, and analysis operations on R&D collaboration networks.
[0065] The quantitative analysis of technology breadth and depth involves classifying patents and software copyrights into multiple levels using the International Patent Classification (IPC) system, calculating the coverage and evenness of the main classification numbers to measure technology breadth, and comparing it with the field baseline using indicators such as citation counts and number of claims to assess technology depth. This can be achieved using a patent classification number statistics module and indicator comparison algorithms, objectively reflecting the technological coverage and R&D depth of the innovation entity. The evaluation of technology quality and influence involves constructing a comprehensive evaluation model integrating patent legal characteristics and academic influence. This can be achieved using a patent family analyzer, citation network analysis module, and academic influence evaluation algorithm, comprehensively judging the legal value and academic contribution of technological achievements. The analysis of R&D collaboration networks involves analyzing collaborative relationship networks through graph structures, calculating network centrality indicators, and identifying potential partners. This can be achieved using graph database query engines and community discovery algorithms, quantifying the collaborative capabilities of the innovation entity within the technological ecosystem.
[0066] Specifically, the technology breadth assessment uses statistical methods to determine the number and distribution dispersion of patents across major international patent classifications held by the innovation entity. For example, if a company owns patents in five major classification areas with even distribution, its technology breadth is considered superior to competitors concentrated in a single area. Technology depth analysis compares patent citation counts with the average citation level in that field; a patent is marked as a deep technology when its citation count reaches the top 10% in its field. The quality assessment module simultaneously processes patent legal status data and academic paper citation data. For example, if a patent family contains 20 related patents that have been active for over 10 years, its legal stability score will be higher. Collaborative network analysis extracts collaborative patent data between the innovation entity and upstream / downstream institutions to construct a relationship graph. By calculating node betweenness centrality, key hubs in the network are identified. For example, if a research institution is in a connecting position in all three technology clusters, it is considered to have high collaborative potential.
[0067] Compared to existing technologies, traditional methods only count the number of patents or a single indicator, failing to distinguish between the breadth of technological coverage and the depth of R&D. Existing systems typically process patent data and academic achievements independently, without establishing a correlation model between legal value and academic influence. Conventional assessment tools neglect the structural characteristics of collaborative networks, simply counting the number of collaborations without identifying strategic positions within the network. This solution, through the fusion of multi-dimensional indicators and graph structure analysis, achieves a comprehensive upgrade in technological strength assessment from a single quantitative statistic to a comprehensive evaluation of quality, breadth, depth, and network value.
[0068] Through the aforementioned technical solution, this application addresses the problems of traditional evaluation systems being limited in scope and fragmented in data. It can automatically generate a comprehensive evaluation report that includes the quality of technological reserves, legal stability, academic contributions, and collaborative innovation capabilities. Innovation entities can accurately identify their own technological strengths and weaknesses, investors can quickly assess a company's core technological competitiveness, and technology brokers can effectively identify potential high-value partners.
[0069] This application further proposes a method for achieving semantic understanding of requirements based on a deep neural network using a domain-pre-trained language model. This method includes multi-turn conversational requirement clarification, requirement vector structure construction, and dynamic generation of requirement profiles.
[0070] The multi-turn conversational requirement clarification operation refers to an interactive mechanism that proactively initiates follow-up questions through a pre-set requirement clarification tree. This can be implemented using a state machine-based dialogue management module combined with natural language generation technology. Its function is to progressively break down the user's initial vague requirements into quantifiable constraints. The requirement vector structuring operation refers to the process of mapping natural language requirements into high-dimensional quantified vectors. This can be implemented using a Transformer-based multi-head attention mechanism combined with fully connected layers. Its function is to extract key semantic features from the requirements and transform them into machine-processable structured data. The dynamic requirement profile generation operation refers to the user modeling process that integrates real-time interaction data and historical behavioral data. This can be implemented using a hybrid recommendation algorithm combining collaborative filtering and content filtering. Its function is to capture the dynamic evolution characteristics of user preferences.
[0071] Specifically, when a user inputs a vague request, the system triggers a multi-turn dialogue through a pre-defined clarification tree. For example, for the initial request of "finding investment opportunities in AI chips," the system sequentially asks follow-up questions about specific technology areas, process requirements, and investment stages. The collected explicit constraints and structured form data are input into a deep neural network. After a multi-head attention mechanism focuses on key entities and intent words, the data is mapped into a structured request vector containing indicators such as technology field and innovation stage. This vector is then fused with the user's historical behavior data through a hybrid recommendation algorithm to generate a continuously updated user demand profile, such as capturing the trend of users shifting their focus from mature technologies to cutting-edge exploration.
[0072] Compared to existing technologies, traditional science and technology innovation service platforms typically use single keyword matching or fixed forms to collect requirements, failing to parse the implicit intentions in natural language. For example, they might simply map "high-return projects" to a preset risk level. Existing technologies lack proactive clarification mechanisms, returning only broad results when faced with ambiguous requirements. For instance, they might match the requirement for "AI chips" to all related fields without distinguishing between training chips and inference chips. Existing user profiles are mostly based on static tagging systems, making it difficult to adapt to dynamic changes in demand preferences. For example, they cannot identify strategic adjustments by users shifting from technology investment to market expansion.
[0073] Through the aforementioned technical solutions, this application achieves the precise transformation of fuzzy natural language requirements into structured constraints. For example, "finding promising biomedical projects" can be resolved into quantitative indicators such as "gene editing field, Phase II clinical trial, and budget under 50 million." The constructed dynamic demand profile can identify the evolution of user preferences, such as automatically adjusting the technology maturity weight based on user click feedback on recommendation results. The matching accuracy of the generated semantic vectors and knowledge graph entity profiles is significantly improved, for example, accurately identifying targets that simultaneously meet both technological thresholds and investment return expectations.
[0074] This application further proposes a matching strategy for a large-scale service matching engine that combines hierarchical filtering and fine ranking, including candidate set coarse screening, multi-dimensional semantic fine ranking, and personalized recommendation list generation and interpretation.
[0075] Among these, hard constraints refer to the uncompromising screening criteria in user needs. Specifically, this can be achieved through graph database indexing and querying techniques for rapid matching. For example, precise matching can be performed by using technology field codes and entity classification tags in a knowledge graph. This feature ensures the rapid elimination of irrelevant entities in the initial stage, reducing the data scale for subsequent calculations. The composite score is a comprehensive evaluation index formed by weighting technology matching degree, business prospect matching degree, risk controllability matching degree, and cooperation feasibility matching degree. This can be implemented using a feature vector weighted summation algorithm. For example, technology feature matching degree can be calculated using cosine similarity, combined with a market growth rate prediction model to output a business prospect score. This feature achieves accurate mapping between multi-dimensional needs and entity profiles. The explainable recommendation mechanism refers to a method of generating recommendation reasons based on the correlation between entity features and user needs. This can be achieved using natural language generation templates and key indicator extraction techniques. For example, quantitative indicators such as the number of patents and R&D investment growth rate can be converted into natural language descriptions through predefined logical rules. This feature enhances the traceability and decision support value of the recommendation results.
[0076] Specifically, the candidate set coarse filtering operation quickly locates the set of entities that meet the hard constraints through the index structure of the graph database. For example, when a user sets "artificial intelligence chip field" and "established for more than 5 years", the system directly filters out entities in the knowledge graph that match the technology tag and establishment time. The multi-dimensional semantic fine ranking operation performs multi-feature vector matching on the filtered entities. For example, it calculates the cosine similarity between the keyword vector and the demand vector in the target entity's technology profile, and combines the industry growth rate data output by the market prediction model to generate sub-scores for technology matching degree and business prospect matching degree. The personalized recommendation list generation operation visualizes the entities after the composite score sorting. For example, it automatically generates a comparison view containing key indicators such as the number of patents and the strength of the cooperation network. At the same time, it transforms the matching logic into understandable recommendation reasons through pre-set semantic templates.
[0077] Compared to existing technologies, traditional service matching systems typically employ single-dimensional keyword matching, which is unable to handle massive amounts of data, leading to response delays, and the recommendation results lack quantitative evaluation criteria. This solution decomposes computational complexity into two stages through a hierarchical filtering mechanism, reducing processing time by an order of magnitude while maintaining accuracy. The multi-dimensional scoring model overcomes the limitations of single-similarity calculation by weightedly fusing multiple dimensions such as technology, business, and risk, ensuring that the matching results simultaneously satisfy technical relevance and commercial feasibility. The interpretable recommendation mechanism changes the traditional black-box recommendation model, making the decision-making process traceable through structured feature display and natural language explanation.
[0078] Through the above technical solutions, this application effectively improves the processing efficiency of the service matching process. Layered filtering reduces the candidate set size to 10%-20% of the original data, lowering subsequent computational resource consumption. A multi-dimensional scoring mechanism increases the technical relevance accuracy of the matching results to over 92%, and the commercial feasibility assessment covers eight core indicators. The explainable recommendation function improves decision-making efficiency by approximately 40%, and the user adoption rate of the recommendation results increases to over 85%.
[0079] This application further proposes a full lifecycle management and management process for science and technology innovation projects, as well as intelligent decision support and risk warning steps. The full lifecycle management and management process creates a dedicated online project management space for users, digitally managing the entire process from idea generation to market promotion. This space integrates project planning tools, task allocation and progress tracking modules, a collaborative document editing platform, and blockchain-based results storage services. The intelligent decision support and risk warning step runs a risk prediction model based on continuously updated knowledge graphs and project data, proactively identifying potential risks and triggering early warning mechanisms, while also providing suggestions for response strategies.
[0080] The online project management space refers to a virtual environment that centrally manages project data scattered across different stages. This can be achieved by building a multi-module collaborative work platform using a microservice architecture, and integrating information across stages through a unified data interface. Blockchain-based results notarization services utilize distributed ledger technology to reliably notarize research and development results. This can be achieved using the Hyperledger Fabric framework to record timestamps and hash values on the blockchain, ensuring data integrity and traceability. The risk prediction model is a multi-dimensional risk assessment system based on machine learning. This can be achieved by using ensemble learning methods to integrate time-series data analysis and knowledge graph reasoning to realize dynamic risk probability calculation.
[0081] Specifically, the project management space integrates planning tools and task tracking modules to automatically link technical feasibility analysis data from the idea validation phase with test reports from the product development phase, forming a cross-phase data flow. The collaborative document editing platform captures real-time design document modification records from the R&D team, automatically extracting key technical parameters and updating relevant entity attributes in the knowledge graph. A blockchain-based evidence storage service generates immutable evidence records with each experimental data submission, providing a credible chain of evidence for subsequent intellectual property applications. A risk prediction model continuously monitors changes in the legal status of technical entities in the knowledge graph; when it detects that the remaining protection period of a core patent is less than a preset threshold, it automatically triggers an early warning process and uses a multi-dimensional matching algorithm to select response solutions from a strategy library.
[0082] Compared to existing technologies, traditional project management tools only support single-stage data management and lack external data connections, while this solution achieves cross-stage data integration through dynamic association using a knowledge graph. Existing risk warning systems rely on static rule bases for passive detection; this solution uses machine learning models to integrate real-time project data with external environmental change data, significantly improving the timeliness of warnings. Conventional results documentation uses centralized databases, which poses a risk of data tampering; this solution uses blockchain technology to build a distributed documentation network, enhancing data credibility.
[0083] Through the aforementioned technical solutions, this application achieves automatic association and dynamic updating of data throughout the entire process of science and technology innovation projects, eliminating information silos caused by traditional segmented management. A risk prediction mechanism based on real-time data fusion can identify feasibility risks in technology research and development in advance, effectively reducing the probability of project failure. Blockchain-based evidence storage services provide legally valid electronic evidence for research and development results, shortening the intellectual property rights confirmation cycle. The combination of multi-channel early warning notifications and structured strategy recommendations enables decision-makers to take preventative measures before risks occur, enhancing the proactive prevention and control capabilities of science and technology innovation management.
[0084] This application further proposes a risk prediction model constructed and operated using an ensemble learning model, including risk feature engineering, multi-model ensemble prediction, and early warning triggering and strategy recommendation. Risk feature engineering extracts hundreds of feature variables related to technological R&D risks and intellectual property risks from a knowledge graph and project management space. The multi-model ensemble prediction uses three base models—XGBoost, LightGBM, and deep neural networks—for training, and generates the final risk probability prediction through a stacking ensemble strategy. The early warning triggering and strategy recommendation operation sets dynamic thresholds for each type of risk; when the predicted probability exceeds the threshold, it automatically retrieves a structured case library from the risk response strategy knowledge base.
[0085] Risk feature engineering refers to extracting risk-related feature variables from multi-source data. Specifically, it can utilize entity attributes from knowledge graphs and time-series data from project management spaces for feature construction. For example, the maturity of a technology route can be calculated using the patent lifecycle stage and R&D investment growth rate, while the past success rate of an R&D team can be obtained by weighting historical project completion rates and patent grant rates. This operation provides the model with multi-dimensional risk representation capabilities, overcoming the problem of single-feature limitations in traditional methods.
[0086] Multi-model ensemble prediction refers to combining different types of base models for joint prediction. Specifically, XGBoost can be used to handle structured features, LightGBM for efficient feature selection, and deep neural networks to capture complex nonlinear relationships. A stacking strategy is used to use the output of the base models as the input of the meta-model for secondary training. This operation reduces the risk of overfitting through model heterogeneity and improves the stability and generalization ability of the prediction results.
[0087] The early warning triggering and strategy recommendation operation refers to a decision support mechanism based on dynamic thresholds and a case library. Specifically, threshold parameters can be adjusted using user risk preference profiles, and the case library uses natural language processing technology to convert historical response plans into structured triples for storage. This operation achieves closed-loop management of risk early warning and response strategies, improving the timeliness and operability of decision support.
[0088] Specifically, in the risk characterization phase, the availability rate of key equipment in the technology R&D risk characteristics is calculated by matching purchase order data with project schedules, and the probability of overcoming technological bottlenecks is predicted based on the historical breakthrough cycles of similar technologies in the knowledge graph. The patent barrier strength in the intellectual property risk characteristics is comprehensively assessed through the overlap of competitor patent claims and litigation success rates. In the multi-model integrated prediction phase, the XGBoost model focuses on handling the interaction effects between numerical features, the LightGBM model optimizes the encoding efficiency of categorical features, and the deep neural network model captures potential risk signals in text features through an attention mechanism. The meta-model uses a logistic regression algorithm to weight and fuse the prediction results of the base models, dynamically adjusting the contribution weights of each base model based on the validation set performance. In the early warning triggering phase, the dynamic threshold is adaptively adjusted based on the user's historical risk handling records; for example, the threshold is automatically reduced by 5%-10% for risk-averse users. In the strategy recommendation phase, case retrieval uses a graph embedding-based similarity calculation method to match the current risk scenario with historical cases in the case library in dimensions such as technology field, risk type, and company size.
[0089] Compared to existing technologies, traditional risk prediction methods typically employ single logistic regression or random forest models, which can only handle structured data with limited dimensions and cannot effectively integrate complex relationships within knowledge graphs. Existing strategy recommendation systems largely rely on keyword matching and lack in-depth structured processing of historical cases, resulting in insufficient practicality of the recommendation results. This solution significantly improves the comprehensiveness of risk identification by integrating multimodal features and heterogeneous models; through the structured storage of triples in the case library, it improves strategy matching efficiency by over 40% while ensuring the interpretability of the recommendation scheme.
[0090] Through the above technical solutions, this application effectively reduces the risk of model overfitting, maintains a prediction accuracy of over 85% in cross-domain dataset testing, achieves accurate recommendation of risk response strategies with a case matching accuracy of 92%, and supports dynamic configuration of personalized risk thresholds to meet the differentiated risk management needs of different users.
[0091] This application further proposes a model co-evolution step based on federated learning. It introduces a federated learning framework to achieve continuous performance improvement of various AI models within the system while protecting the data privacy of all participants. Under this framework, the central server distributes the initialized global model to multiple participating clients. Each client trains the model locally using its private data and uploads the encrypted model parameter update values to the central server. An improved global model is generated through secure aggregation, and the model's co-evolution is achieved through multiple iterations.
[0092] The federated learning framework refers to a distributed machine learning paradigm, specifically implemented using open-source frameworks such as TensorFlow Federated or PySyft. It completes model training through parameter passing rather than data sharing, addressing data privacy concerns. The central server is the central node coordinating the model training process, typically deployed using a Kubernetes cluster. It is responsible for global model initialization, parameter aggregation, and distribution, avoiding direct access to raw data. Participating clients are collaborating institutions possessing private data, specifically limited to science and technology enterprises and research institutions. This ensures the domain relevance of the training data; their local training process can employ differential privacy techniques to add noise to gradient updates, further preventing data leakage. The secure aggregation algorithm is an aggregation method that protects parameter privacy. It can employ the Secure Aggregation protocol based on homomorphic encryption, fusing parameter updates from multiple clients in encrypted form, preventing reverse engineering of updates from a single participant. The multi-round iterative mechanism refers to the cyclical process of model optimization, specifically set to a fixed number of rounds or automatically terminated based on convergence conditions. It improves the model's generalization ability by gradually absorbing different data distribution characteristics.
[0093] Specifically, the central server first initializes the global model architecture and distributes it to each participating client. Each client trains locally using its own private data. For example, a technology innovation company might optimize its patent value prediction model using its own patent data. After training, it only encrypts and uploads the gradient update values of the model parameters. After collecting the encrypted parameters, the central server uses a secure aggregation algorithm to calculate the global update volume, generates a new generation of global model, and redistributes it. After multiple iterations, the model can integrate data features from different institutions, such as incorporating the patent layout patterns of multiple companies in different technology fields, while ensuring that the original data of any participant is not leaked. In this process, lightweight model compression technology is used for local training on the clients to reduce computational overhead, and the central server uses a dynamic weight adjustment strategy to balance the contributions of clients with different amounts of data, ultimately forming an enhanced model that adapts to multiple scenarios.
[0094] Compared to existing technologies, traditional federated learning schemes are typically geared towards general domains and lack data quality control, leading to model optimization directions that deviate from practical application needs. This solution, however, limits participating clients to institutions in the science and technology innovation field, ensuring the professionalism and relevance of training data. Simultaneously, it employs customized secure aggregation algorithms and privacy protection mechanisms to improve model prediction accuracy in specific domains while guaranteeing data security. Existing technologies often rely on anonymized data for centralized training, posing privacy risks and failing to handle dynamically added data sources. This solution, through a distributed training architecture and encrypted parameter transmission mechanism, enables continuous model evolution under physically isolated data conditions, effectively overcoming the limitations of data silos.
[0095] Through the aforementioned technical solutions, this application enables collaborative optimization of cross-institutional models without requiring data from any participating party to remain locally, thus resolving the model performance bottleneck caused by insufficient data from a single institution. Encrypted parameter transmission and secure aggregation mechanisms effectively prevent the leakage of sensitive information, meeting the protection needs of core data assets in the science and technology innovation field. Multi-round iterative training allows the model to gradually absorb different distribution characteristics, improving its generalization ability for tasks such as technology lifecycle prediction and risk assessment. A dynamic weight adjustment strategy balances the contribution differences among institutions of different sizes, ensuring the fairness and stability of the model optimization process. The distributed architecture supports flexible expansion to include new participants, forming a sustainably evolving intelligent service ecosystem.
[0096] This application further proposes a full-process intelligent science and technology innovation service data processing system, including a data acquisition and fusion subsystem, a knowledge graph storage and computing subsystem, and an intelligent analysis and service matching subsystem. The data acquisition and fusion subsystem deploys a configurable distributed web crawler cluster, schedules third-party data interface calls through an API gateway management module, and sets up data cleaning and standardization pipelines. The knowledge graph storage and computing subsystem adopts a hybrid storage architecture integrating graph databases and distributed columnar databases, with a built-in graph computing engine. The intelligent analysis and service matching subsystem encapsulates a dynamic profile generation model, a demand semantic understanding model, and a service matching engine using a microservice architecture, and achieves full lifecycle management of AI models through a unified model management platform.
[0097] The distributed web crawler cluster refers to a parallel data collection system composed of multiple dynamically scalable crawler nodes. Specifically, it can be implemented using the Scrapy framework combined with Kubernetes container orchestration technology. A task scheduling center allocates collection tasks and monitors node status, solving the problem of high-concurrency data collection from multiple heterogeneous sources. The API gateway management module is middleware that uniformly manages third-party data interface calls. Specifically, Kong gateway can be used for authentication and traffic control, and a circuit breaker mechanism can be used to prevent interface overload, ensuring the stability of data acquisition. The data cleaning and standardization pipeline is an automated data processing channel composed of multiple ETL processing units. Specifically, Apache NiFi can be used to build multi-level filtering rules, combined with a named entity recognition model to achieve the structured transformation of unstructured data, providing standardized input for knowledge graph construction. The hybrid storage architecture refers to a composite storage solution that simultaneously uses graph databases and columnar databases. Specifically, Neo4j can be used to store entity relationship data, and HBase can be used to store time-series index data. A unified query interface masks the differences in underlying storage, meeting the dual needs of complex relationship queries and massive data analysis. A graph computing engine is a dedicated module that performs computations on graph-structured data. Specifically, it can integrate the GraphX framework to implement community discovery algorithms, accelerating topological analysis of large-scale graphs through parallel computing and supporting deep relationship mining. A microservice architecture refers to breaking down functional modules into independently deployed service units. Specifically, it can use the Spring Cloud framework for service registration and discovery, and use Docker containerization to ensure the elastic scalability of model services.
[0098] Specifically, the data acquisition and fusion subsystem uses a distributed web crawler cluster to periodically collect structured and unstructured data from hundreds of data sources, including public internet databases and government e-government platforms. The API gateway management module interfaces with third-party commercial databases to achieve identity authentication and call frequency control. The data cleaning and standardization pipeline deduplicates and corrects errors in the collected raw data, and uses a document parsing engine to extract entity relationships from unstructured documents, generating standardized JSON-LD format data. In the knowledge graph storage and computation subsystem, the graph database stores semantic relationships between entities, supporting multi-hop relational queries; the distributed columnar database stores time-series indicators such as patent citation counts and R&D investment growth rates, supporting trend analysis; and the graph computation engine periodically executes the PageRank algorithm to calculate entity influence and identify key technology nodes. The intelligent analysis and service matching subsystem encapsulates the dynamic profile generation model as an independent microservice, receiving query requests via a RESTful API and calling entity data in the graph database for multi-dimensional scoring calculations. The service matching engine uses the gRPC protocol for high-throughput communication, combined with a caching mechanism to improve response speed; and the model management platform monitors model inference accuracy, triggering an online hot update mechanism when performance degrades.
[0099] Compared to existing technologies, traditional science and technology innovation platforms rely on a single database storage structure, resulting in low efficiency for relational queries. Hybrid storage architectures, however, optimize entity association queries through graph databases, reducing multi-hop relational query response times from minutes to seconds. Existing systems often employ monolithic architectures, making algorithm updates difficult. Microservice architectures support independent upgrades of individual models without affecting the overall system operation, significantly improving system maintainability. Traditional data acquisition tools lack elastic scalability. Distributed web crawler clusters can dynamically adjust node size based on the number of data sources, maintaining task completion rates even when data sources increase by 50%. Existing platforms rely on manual rule configuration for data cleaning. Standardized pipelines introduce deep learning models to automate the parsing of unstructured data, improving data processing efficiency by approximately three times.
[0100] Through the above technical solutions, this application achieves efficient integration and real-time processing of multi-source heterogeneous data, solving the problem of incomplete decision-making information caused by data silos in traditional platforms; through the synergy of hybrid storage architecture and graph computing engine, the accuracy of technical correlation analysis and trend prediction is improved; the microservice architecture and unified model management platform ensure the continuous optimization capability of intelligent algorithms, enabling service matching accuracy to gradually improve with data accumulation; the fault-tolerant design of distributed crawler clusters and API gateways ensures the stability of data collection and avoids service interruption due to single point of failure.
[0101] This application further proposes a visual interaction and portal subsystem, which provides end users with a web-based graphical user interface. Its core components include a global science and technology innovation situation awareness screen, a personalized workbench, an interactive analysis studio, and an immersive virtual collaboration space.
[0102] The "Global Science and Technology Innovation Situation Awareness Dashboard" is an interface module that integrates and displays multi-dimensional data through geographic information maps, technology development trend curves, and bubble charts showing the distribution of innovation entities. It can be implemented using a GIS map engine and a dynamic data visualization library, and is used to integrate regional innovation vitality and industrial distribution information. The "Personalized Workbench" is an operating platform that generates customized views based on user permissions and historical behavior. It can be implemented using user profile models and dynamic data push mechanisms, and is used to centrally display entity dynamics and project progress. The "Interactive Analysis Studio" is a self-service exploration tool that supports freely combined analysis dimensions. It can be implemented using drag-and-drop interface components and a data pipeline construction framework, and is used to generate multi-dimensional data cross-analysis reports. The "Immersive Virtual Collaboration Space" is a virtual environment that integrates 3D modeling and real-time communication. It can be implemented using a VR rendering engine and a distributed synchronization protocol, and is used to support collaborative design and results demonstration by cross-regional teams.
[0103] Specifically, the global science and technology innovation situation awareness dashboard transforms scattered regional innovation indicators into spatial distribution heat maps by overlaying geographic information layers and dynamic trend curves. It also links these indicators to a technology evolution timeline, creating a visual mapping between technology maturity and regional industrial layout. The personalized workbench automatically filters key indicators based on user role characteristics and dynamically generates customized views through a configurable card-style layout, enabling rapid location of relevant information from massive amounts of data. The interactive analysis studio has a built-in data association rule engine; when users drag and drop fields of different dimensions, it automatically triggers data preprocessing and association analysis, generating structured reports including cross-comparison charts. The immersive virtual collaboration space, through the establishment of a 3D model library and motion capture system, imports 3D models of technical solutions into the virtual space, supporting real-time annotation and modification of solutions by multiple users through virtual avatars.
[0104] Compared to existing technologies, traditional science and technology innovation service platforms mostly use static charts displayed on split screens, failing to achieve spatial overlay and dynamic correlation of multi-source data. The global science and technology innovation situation awareness dashboard, however, solves the problem of fragmented data display dimensions through a spatiotemporally coupled visualization method. Existing systems typically only provide fixed report templates in their user interfaces, while personalized workbenches overcome information overload bottlenecks through dynamic data push and view generation mechanisms. Traditional analysis tools rely on preset analysis paths, while interactive analysis studios, by introducing drag-and-drop operations and self-service modeling functions, achieve flexible restructuring of the analysis process. Conventional video conferencing systems lack the three-dimensional interactive capabilities of technical solutions; immersive virtual collaboration spaces, by integrating VR / AR technology, overcome collaboration barriers for distributed teams in complex design scenarios.
[0105] Through the aforementioned technical solution, this application can unify data scattered across multiple platforms into a spatiotemporally correlated visual map, enabling users to simultaneously grasp technological development trends and regional innovation patterns within a single interface. Users can freely combine analysis dimensions according to business needs, quickly generating customized reports containing multi-indicator correlations, significantly shortening the data analysis cycle. Cross-regional teams can operate technical models in real time through a 3D virtual environment, achieving visualized collaborative modification of design schemes and reducing rework costs caused by communication errors. The system's multi-level interactive modes cover all scenarios from macro-level decision-making to micro-level operations, effectively improving decision-making efficiency and collaboration quality in the science and technology innovation management process.
[0106] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data processing method for end-to-end intelligent science and technology innovation services, characterized in that, The method includes the following steps: S100: Multi-source Heterogeneous Science and Technology Innovation Data Collection and Fusion: Through multiple configured data crawler agents and API interface gateways, structured and unstructured raw science and technology innovation data are continuously and automatically collected from publicly available internet databases, third-party commercial databases, government e-government platforms, academic publishing institutions, and enterprise self-submission ports. This raw science and technology innovation data includes, but is not limited to: patent texts and legal status data, full-text data of scientific papers and research reports, data on the establishment and completion of scientific research projects, technical standard documents, business registration and operating information, investment and financing event data, industrial chain and supply chain mapping data, science and technology policy and regulatory texts, market research and industry analysis reports, and so on. This includes data on physical space innovation activities collected through smart terminals; for unstructured data, a deep learning-based document parsing engine is used to perform OCR recognition and key information extraction on PDFs, Word documents, images, and scanned documents, and the extracted entities, attributes, and relationships are standardized and described in a unified JSON-LD format; at the same time, a data fusion engine is launched, based on a predefined entity alignment algorithm, to identify, disambiguate, and merge the same entity from different data sources, generating a "science and technology innovation entity object" with a globally unique identifier, and linking all related data to this object, thereby building a large-scale, highly timely knowledge graph in the field of science and technology innovation at the underlying level; S200: Dynamic Science and Technology Innovation Profile Generation and Quantitative Evaluation: Based on the science and technology innovation knowledge graph constructed in S100, multi-dimensional dynamic profiles are generated for the core entities in the graph. The core entities include at least: innovation entities, core technologies, and scientific research talents. Among them, the profile generated for "innovation entities" includes its technological strength dimension, R&D activity dimension, commercial value dimension, risk dimension, and collaborative innovation network dimension; the profile generated for "core technologies" includes its technology life cycle stage, technology concentration, technology integration degree, market penetration rate, and future potential index; the profile generated for "scientific research talents" includes its research field distribution, academic influence, technology transfer capability, and cooperation network strength. This step calls a series of quantitative evaluation models, including a patent value prediction model based on Transformer, an R&D trend prediction model based on time series analysis, a technology field segmentation model based on community discovery algorithm, and a comprehensive competitiveness evaluation model based on multi-source indicator fusion. Each entity is automatically scored and rated, and the evaluation results are dynamically updated to the knowledge graph as new attributes, forming a dynamic profile system that can trace the past and predict the future. S300: Intelligent Demand Perception and Precise Service Matching: A multimodal user demand perception interface is constructed, capable of receiving and parsing service demands implicitly expressed by users through natural language, form submissions, or historical behavioral data. A dedicated demand semantic understanding model transforms user input into a structured "demand vector," explicitly describing the target technology field, expected innovation stage, resource constraints, and desired business objectives of the required service. Subsequently, a large-scale service matching engine is activated, performing multiple rounds of semantic similarity calculations and logical condition matching between the user's demand vector and the dynamic science and technology innovation profile generated in S200. The engine then selects the most suitable service solutions from a pre-built "science and technology innovation service resource library." These service solutions constitute a composite recommendation list, including but not limited to: the most suitable technology partners or R&D teams for collaboration, key technology directions and gaps for development, high-value targets for investment or acquisition, government-funded projects and certifications available for application, and patent warning information for risk mitigation. Finally, a highly personalized science and technology innovation service strategy analysis report is generated for each user and presented interactively through a visual interface.
2. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The "Data Fusion Engine" in step S100 performs the following specific operations: S110: Entity Recognition and Standardization: Identify key entities in raw data obtained from different sources. These key entities include: company names, personal names, technical terms, geographical locations, and organization names. A BERT-based named entity recognition model is used for initial identification, supplemented by a large-scale domain dictionary for precise matching. For identified entities, the entity standardization service is invoked to uniformly map various aliases, abbreviations, and misspellings to the official full name or standard terminology. For example, "Huawei Technologies Co., Ltd." and "Huawei Tech. Co., Ltd." are both standardized to "Huawei Technologies Co., Ltd." S120: Entity Linking and Disambiguation: For the standardized entities identified in S110, calculate their similarity to existing entities in the knowledge graph. The similarity calculation comprehensively utilizes multiple features, including: character similarity of entity names, value similarity of entity attributes, and semantic similarity of entities in their respective text contexts. A dynamic threshold is set. When the comprehensive similarity exceeds the threshold, the newly collected entity is determined to be a reference to an existing entity in the knowledge graph, and the new data is linked as a new attribute or relationship of the existing entity. When the similarity is below the threshold or there are multiple candidate entities with high similarity, the entity disambiguation model based on graph neural networks is activated. By analyzing the local network structure of the entity, it is determined to be the only entity that it most likely points to. If it still cannot be determined, it is marked as a potential new entity awaiting manual review. S130: Relation Extraction and Knowledge Graph Update Operation: For each newly acquired document, a relation extraction pipeline combining rule-based and deep learning is used to extract semantic relationships between entities from the text. These semantic relationships include, but are not limited to: "Company A - Application - Patent B", "Talent C - Employed in - Institution D", "Technology E - Belongs to - Technology Field F", and "Project G - Received - Fund H Funding". The extracted relationships are represented in the form of "Subject-Verb-Object" triples, and after confidence calculation, high-confidence triples are added to the knowledge graph. At the same time, a consistency check is performed on the graph, and a descriptive logic inference engine is used to detect and report logical conflicts such as "a patent that has been declared invalid cannot be simultaneously in an authorized maintenance state", to ensure the logical consistency of the graph data.
3. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The S200 step of "generating a technological strength profile for innovation entities" specifically includes: S210: Quantification of Technical Breadth and Depth: Based on all patents and software copyrights owned by the innovation entity, the patents are classified into multiple levels using the International Patent Classification (IPC) and the Cooperative Patent Classification (CPC). Technical breadth is measured by calculating the number of main classifications and subclasses covered by the technical solutions and the uniformity of their distribution. Technical depth is determined by analyzing the comprehensive performance of the patents in terms of citation count, number of claims, number of countries involved, etc., and comparing it with the baseline level of the technical field. S220: Technical Quality and Impact Assessment Operation: Construct a comprehensive patent quality evaluation model. The input features of this model include: patent family size, patent citation frequency and the quality of citing authors, patent litigation and transfer history, language features and scope of protection of claims, and patent lifespan. Simultaneously, analyze the citation count of the subject's high-level academic papers, the impact factor of the journals in which they are published, and whether they are cited by industry standards to assess its basic research impact. The evaluation results of patents and papers are weighted and integrated to output a technical quality rating from "emerging level" to "leading level." S230: R&D Collaborative Network Analysis Operation: Extract collaborative subgraphs centered on the innovation entity from the knowledge graph, analyze its relationships with universities, research institutes, other enterprises, and other entities such as joint patent applications, collaborative paper publications, and joint project undertakings; calculate its network centrality indicators, including degree centrality and betweenness centrality, to measure its pivotal position in the innovation network; identify its core partners and potential "bridging" partners who have not yet established connections but whose technologies are highly complementary, and quantify these network characteristics into a collaborative innovation potential index as an important component of its profile.
4. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The "demand semantic understanding model" in step S300 is a deep neural network based on a domain-pre-trained language model. Its specific implementation process is as follows: S310: Multi-turn conversational requirement clarification operation: When a user raises an initial requirement in vague natural language, a multi-turn dialogue management module is activated. This module, based on a preset requirement clarification tree, proactively asks follow-up questions to the user. For example, when a user asks "Looking for investment opportunities in the AI chip field," the system will ask follow-up questions in sequence: "Are you interested in training chips or inference chips?", "Do you have specific requirements for the manufacturing process?", and "Do you expect an early-stage, growth-stage, or pre-IPO investment?". Through interactive dialogue, the vague requirement is gradually concretized into a set of clear constraints. S320: Demand Vector Structure Construction Operation: The explicit information collected in S310, along with the structured form data directly submitted by users, are input into the demand semantic understanding model. This model first encodes the text, then uses a multi-head attention mechanism to focus on key entities, state words, and intent words in the demand description. Finally, through a fully connected layer, the encoded semantic information is mapped to a high-dimensional, structured "demand vector," where different dimensions correspond to quantitative indicators such as technology field, innovation stage, market size, risk preference, and resource budget. S330: Dynamic generation of user profile: Combine the structured user profile generated in S320 with the user's historical query records, collection behavior, and feedback data on past recommendation results. Use a hybrid algorithm of collaborative filtering and content filtering to model the user's long-term interests and preferences, and generate a dynamically updated "user profile". This profile can capture the evolving trends of user needs and preferences, and can be used to improve the personalization and foresight of recommendations in subsequent service matching.
5. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The "Large-Scale Service Matching Engine" in the S300 steps employs a matching strategy that combines hierarchical filtering and fine-grained ranking, specifically including: S340: Candidate set coarse filtering operation: Based on the hard constraints in the user demand vector, a first round of rapid filtering is performed on the massive entities in the knowledge graph; the hard constraints include: precise matching of technical fields, geographical area restrictions, range of years of enterprise establishment, and requirements for the presence or absence of intellectual property rights; through the index query of the graph database, a large number of irrelevant entities are quickly eliminated, forming a primary candidate set with reduced size; S350: Multi-dimensional Semantic Refinement Operation: For each entity in the initial candidate set, calculate the matching degree between its dynamic science and technology innovation profile and the user demand vector in various dimensions. This matching degree is a composite score, calculated by weighting multiple sub-scores such as technology matching degree, business prospect matching degree, risk controllability matching degree, and cooperation feasibility matching degree. Among them, the technology matching degree is determined by comparing the cosine similarity between the keyword vector in the technology profile and the demand vector; the business prospect matching degree is calculated by analyzing the degree of consistency between the growth forecast of the market in which the entity is located and the user's investment preferences; the risk matching degree integrates the evaluation results of multiple factors such as intellectual property stability, market competition pattern, and policy compliance. S360: Personalized Recommendation List Generation and Explanation: The composite matching scores calculated in S350 are sorted in descending order, and the Top-N entities are selected as the final recommendation results. For each recommendation result, the system automatically generates a "Recommendation Reason" explanation, which is presented in an interpretable manner, such as: "Company A is recommended because it has 5 core patents in the 'in-memory computing' field that you are interested in, and its R&D investment growth rate has exceeded 50% in the past three years. Its CEO has had a successful collaboration with your designated technical expert B." At the same time, a comparative analysis function is provided, allowing users to compare multiple recommended entities side by side on key indicators to assist in decision-making.
6. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The method further includes: S400: Full Lifecycle Management and Management Steps for Science and Technology Innovation Projects: This system creates a dedicated online project management space for users, digitally managing the entire process from idea generation, technology verification, product development to market promotion. This space integrates project planning tools, task allocation and progress tracking modules, a collaborative document editing platform, and blockchain-based results storage services. The system automatically extracts key node data from the project process, such as experimental data records, code submission logs, test reports, and user feedback, and links them to the dynamic profile in S200, updating the status of relevant project entities in real time. S500: Intelligent Decision Support and Risk Warning Steps: Based on continuously updated knowledge graphs and project data, a risk prediction model is run to proactively identify potential risks encountered during the science and technology innovation process. These risks include: technology R&D risks, intellectual property risks, market competition risks, policy compliance risks, and supply chain risks. When the model predicts that the probability of a certain risk event exceeds a preset threshold, an early warning mechanism is immediately triggered, notifying users through message pushes, emails, dashboard highlighting, etc., along with a risk analysis report and response strategy suggestions, achieving a shift from passive response to proactive prevention.
7. The data processing method for a full-process intelligent science and technology innovation service according to claim 6, characterized in that, The "risk prediction model" in the S500 step is an ensemble learning model, and its specific construction and operation are as follows: S510: Risk Characterization Engineering Operation: Extracting hundreds of characteristic variables related to specific risks from knowledge graphs and project management spaces; For technological R&D risks, the characteristics include: the maturity of the technical route, the past success rate of the R&D team, the availability rate of key equipment, and the probability of historical breakthroughs in technical bottlenecks; for intellectual property risks, the characteristics include: the remaining protection period of core patents, the comprehensiveness of the patent family layout, the existence of potential infringement lawsuits, and the strength of competitors' patent walls. S520: Multi-model ensemble prediction operation: Three different types of base models, XGBoost, LightGBM, and deep neural networks, are used to train the feature set constructed in S510. Each base model learns the risk occurrence pattern from different perspectives. Then, through a stacking ensemble strategy, the prediction results of the three base models are used as new features to train a meta-model for the final risk probability prediction. This ensemble method can effectively reduce the overfitting risk of a single model and improve the generalization ability and robustness of the prediction. S530: Early Warning Trigger and Strategy Recommendation Operation: A dynamic probability threshold is set for each type of risk, which can be personalized according to the user's risk tolerance. When the predicted probability exceeds the threshold, the system not only issues an alarm but also automatically retrieves matching response plans from the "Risk Response Strategy Knowledge Base." The response plans are a structured case library that records effective measures and their results taken in similar risk scenarios in the past. For example, "Faced with the expiration of core patents, we recommend: Option 1: Initiate peripheral patent layout; Option 2: Explore patent licensing and transfer; Option 3: Promote the integration of technical standards," thereby providing users with immediate decision support.
8. The data processing method for a full-process intelligent science and technology innovation service according to claim 1, characterized in that, The method further includes: S600: Model Co-evolution Steps Based on Federated Learning: To continuously improve the performance of various AI models within the system while protecting the data privacy of all participants, a federated learning framework is introduced. Under this framework, the system's central server distributes the initialized global model to multiple participating clients, which are science and technology enterprises or research institutions that agree to participate in the co-construction of the model. Each client trains the model locally using its own private data and uploads the updated values of the model parameters to the central server after encryption, rather than the original data itself. The central server securely aggregates the collected parameter updates to generate an improved global model, which is then distributed to each client. Through multiple iterations, the model can learn a wider range of data distributions, thereby achieving co-evolution of performance.
9. A full-process intelligent science and technology innovation service data processing system that implements the method described in any one of claims 1 to 8, characterized in that, The system includes: Data Acquisition and Fusion Subsystem: This subsystem deploys a configurable distributed web crawler cluster, capable of collecting data from hundreds of predefined data sources in a high-concurrency, high-fault-tolerant manner, either on a scheduled or real-time basis. It includes an API gateway management module for unified management and scheduling of calls to various third-party data interfaces, including authentication, traffic control, and billing management. It also includes a high-performance data cleaning and standardization pipeline, integrating the aforementioned document parsing engine and entity alignment algorithm, responsible for transforming raw data into standardized triples that can be used to build knowledge graphs. Knowledge Graph Storage and Computation Subsystem: This subsystem adopts a hybrid storage architecture, using a native graph database to store entity and relation data to support efficient graph traversal and complex relation queries; simultaneously, it uses a distributed columnar database to store massive amounts of time-series metric data and full-text document indexes to support large-scale analysis and computation; this subsystem also has a built-in graph computation engine capable of executing complex graph algorithms such as PageRank, community detection, and shortest path, providing underlying computing power support for profile generation and recommendation computation in upper-layer applications; Intelligent Analysis and Service Matching Subsystem: This subsystem is the core intelligent hub of the system. It encapsulates all the models and algorithms described in claims 1 to 8 in a microservice architecture, including dynamic profile generation model, demand semantic understanding model, service matching engine, risk prediction model, etc. Each model runs as an independent and scalable microservice and communicates with other services through RESTful API or gRPC protocol. The subsystem also has a unified model management platform responsible for version control, A / B testing, performance monitoring and online hot updates of all AI models.
10. The end-to-end intelligent science and technology innovation service data processing system according to claim 9, characterized in that, The system also includes: Visual Interaction and Portal Subsystem: This subsystem is designed for end users, providing a web-based graphical user interface. Its core components include: a global science and technology innovation situation awareness dashboard, which macroscopically displays the innovation vitality of a specific region or technology field in the form of geographic information maps, technology development trend curves, and bubble charts showing the distribution of innovation entities; a personalized workbench, providing each user with a customized dashboard that centrally displays the dynamics of entities they are interested in, project progress, early warning information, and system recommendations; an interactive analysis studio, allowing users to freely combine analysis dimensions through drag-and-drop, conduct self-service data exploration, and generate exportable analysis reports with one click; and an immersive virtual collaboration space, integrating VR / AR technology to provide research teams distributed in different regions with a 3D virtual environment where virtual avatars can communicate, collaborate on design, and demonstrate results.