Omnibearing data comprehensive management system and method oriented to enterprise research and development strength
By constructing a comprehensive data management system, integrating multi-source data and adopting a two-tier architecture evaluation model, the system addresses the limitations and lag of existing evaluation methods, enabling a comprehensive, dynamic evaluation and forward-looking prediction of the company's R&D capabilities, thereby enhancing the scientific and intelligent nature of management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for assessing corporate R&D capabilities focus on outcome outputs, neglecting process efficiency, resource effectiveness, and innovation quality. This results in assessments that fail to accurately and comprehensively reflect a company's overall R&D capabilities. Furthermore, they lack a dynamic, longitudinal, and temporal perspective, making forward-looking predictions impossible and leading to lagging management decisions.
We will build a comprehensive data management system, integrate multi-source data through automated collection strategies and interfaces, construct a wide table of R&D facts, define key quantitative indicators, and adopt a two-layer architecture of horizontal fusion evaluation and vertical time series prediction to build a comprehensive evaluation model of enterprise R&D strength and conduct statistical performance verification.
It enables a multi-faceted, quantifiable, and comparable accurate assessment of a company's R&D capabilities, overcoming the one-sidedness of traditional assessments, providing forward-looking decision support, and improving the scientific and intelligent level of management.
Smart Images

Figure CN121787844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of enterprise R&D capability assessment, specifically to a comprehensive data management system and method for assessing enterprise R&D capability. Background Technology
[0002] In an era driven by innovation, a company's R&D strength is crucial to its core competitiveness and sustainable development. A scientific, accurate, and dynamic assessment of a company's R&D strength has become an important basis for internal management optimization, external investment decisions, and policy support. Traditionally, the understanding of R&D strength has largely remained at the level of qualitative descriptions or single financial indicators (such as total R&D investment and number of patents), lacking a quantitative evaluation system that comprehensively reflects the input, process, output, and benefits of R&D activities. Meanwhile, with the deepening of enterprise informatization, massive amounts of multi-source, heterogeneous data have been generated from various aspects of R&D activities (such as project management, human resources, finance, assets, and intellectual property). This data should be the objective basis for assessing R&D strength, but in reality, it is often scattered across different "data silos," with inconsistent standards and varying quality, making it impossible to effectively integrate and utilize.
[0003] Existing methods often focus on outcome outputs (such as the number of patents and papers) while neglecting key dimensions such as process efficiency (such as delivery cycle and code quality), resource effectiveness (such as input-output ratio and talent retention) and innovation quality (such as high-value patent density). This results in evaluation results that fail to truly and comprehensively reflect the company's overall R&D strength. Furthermore, most evaluation models are merely static and horizontal comparisons of historical data (such as industry rankings), lacking a longitudinal and temporal perspective. They cannot reveal the dynamic evolution trend of R&D strength, let alone make forward-looking predictions, leading to a lag in management decisions. Summary of the Invention
[0004] To address the aforementioned technical issues, this paper provides a comprehensive data management system and method for enterprises' R&D capabilities. This technical solution resolves the problem mentioned in the background that focuses on outcome-based outputs while neglecting key dimensions such as process efficiency, resource effectiveness, and innovation quality. This results in evaluation results that fail to accurately and comprehensively reflect the enterprise's overall R&D capabilities. Furthermore, most evaluation models only provide static and horizontal comparisons of historical data, lacking a longitudinal and temporal perspective. They cannot reveal the dynamic evolution trend of R&D capabilities, let alone make forward-looking predictions, leading to a lag in management decisions.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A comprehensive data management approach tailored to enterprise R&D capabilities, comprising: Identify the types of multi-source data throughout the enterprise's R&D process, set up automated data collection strategies and interfaces, preprocess the data, and construct a wide table of R&D facts; Based on a wide table of R&D facts, calculate the key quantitative indicators needed to assess R&D capabilities; Based on key quantitative indicators, a two-tier architecture of horizontal fusion assessment and vertical time-series prediction is adopted to construct a comprehensive assessment model of enterprise R&D strength. Based on the output of the comprehensive evaluation model of enterprise R&D capabilities, the reliability of the model's prediction results is confirmed through statistical performance verification.
[0006] Preferably, the steps of determining the multi-source data types throughout the enterprise's R&D process, setting automated data collection strategies and interfaces, preprocessing the data, and constructing a wide table of R&D facts specifically include: Identify the multi-source data types involved in the entire enterprise R&D process; The multi-source data types include R&D project data, human resource data, intellectual property data, funding data, equipment asset data, and technology transfer data; Construct an enterprise R&D data asset map. This map takes the R&D process as the main line, maps the data types, data source systems, data formats and update frequencies generated at each stage, and forms a top-level design blueprint for data collection. Based on the data asset map, we design automated collection strategies and interfaces for each type of data. The automated data collection strategy is divided into three categories based on the openness of the data source system and the degree of data structuring: fully automated interface collection, semi-automatic file import, and manual confirmation and entry. The collected multi-source data were standardized and preprocessed, and a wide table of R&D facts was constructed. The preprocessing includes at least: data cleaning, format standardization, outlier handling, missing value imputation, text feature extraction, category feature encoding, and data association and fusion.
[0007] Preferably, the key quantitative indicators required for calculating and evaluating R&D capabilities based on the R&D fact wide table specifically include: Define the key quantitative indicator types needed to assess R&D capabilities, including: efficiency indicators, quality indicators, innovation and output indicators, and resource and efficiency indicators; Efficiency metrics include: average delivery time of requirements, code deployment frequency, and team iteration rate; Quality metrics include: defect rate per thousand lines of code, average time to resolve defects, and test case pass rate; Innovation and output indicators include: high-value patent density, revenue share of new products, number of registered technological achievements, and technology debt ratio; Resource and efficiency indicators include: R&D input-output ratio, key talent retention rate, per capita R&D output, and equipment utilization rate; Based on the data in the R&D fact table, the key quantitative indicators required to assess R&D capabilities are calculated, and a list of key indicators is generated.
[0008] Preferably, the construction of a comprehensive evaluation model for enterprise R&D strength based on key quantitative indicators, using a two-tier architecture of horizontal fusion assessment and vertical time-series prediction, specifically includes: A two-layer architecture of horizontal fusion assessment and vertical time series prediction is adopted. The horizontal fusion assessment layer outputs the current R&D strength comprehensive index and assessment confidence level, while the vertical time series prediction layer outputs the predicted value of the R&D strength index for future multiple periods and its confidence interval. Based on the enterprise's R&D facts wide table, a list of key indicators is generated, and the values of key quantitative indicators over T historical periods are extracted to construct an indicator matrix. The last row of the indicator matrix is used as the input to the horizontal fusion evaluation layer, and the entire matrix is used as the input to the vertical time series prediction layer. The index matrix is subjected to standardized preprocessing, which includes at least: normalization based on index direction, scaling based on historical extreme values, and time series-based imputation for missing values. Based on the horizontal fusion evaluation layer, the weights of the indicators are determined by a combination of subjective and objective weighting methods. The combined weighting includes at least: calculating objective weights based on the entropy weight method, determining subjective weights based on the expert scoring method, and performing linear combination through the harmonic coefficient; The current comprehensive index is calculated using linear weighting based on the combined weights. The comprehensive confidence level is calculated through data integrity and volatility, and the current comprehensive index of R&D strength is output. Based on the vertical time series prediction layer, indicators that have a leading influence on the composite index are selected as prediction inputs through mutual information analysis based on the indicator matrix. The longitudinal time series prediction layer adopts a temporal convolutional network model framework, taking the comprehensive index and its leading influence indicators as inputs, and outputting the predicted value of the R&D strength index for future multi-cycle periods and its confidence interval. The trained two-layer architecture of horizontal fusion evaluation and vertical time series prediction is defined as a comprehensive evaluation model of enterprise R&D strength.
[0009] Furthermore, this solution proposes a comprehensive data management system tailored to enterprise R&D capabilities, used to implement the aforementioned comprehensive data management method for enterprise R&D capabilities, including: The data acquisition module is used to determine the multi-source data types in the entire enterprise R&D process, set automated acquisition strategies and interfaces, preprocess the data, and construct a R&D fact wide table. The model building and validation module is used to calculate key quantitative indicators required to evaluate R&D strength based on a wide table of R&D facts; based on the key quantitative indicators, a two-layer architecture of horizontal fusion evaluation and vertical time series prediction is adopted to build a comprehensive evaluation model of enterprise R&D strength; based on the output results of the comprehensive evaluation model of enterprise R&D strength, the reliability of the model prediction results is confirmed through statistical performance verification.
[0010] Preferably, the model building and verification module includes: A quantitative indicator unit is used to calculate key quantitative indicators required to evaluate R&D capabilities based on a wide table of R&D facts. The model building unit is used to construct a comprehensive evaluation model of enterprise R&D strength based on key quantitative indicators and using a two-layer architecture of horizontal fusion evaluation and vertical time series prediction. The model validation unit is used to comprehensively evaluate the output results of the model based on the enterprise's R&D capabilities, and to confirm the credibility of the model's prediction results through statistical performance verification.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a comprehensive data management solution for enterprise R&D strength. By constructing a data asset map covering the entire R&D process and implementing an automated data collection strategy, it integrates scattered and heterogeneous multi-source data into a unified R&D fact table, providing an accurate, complete, and traceable data foundation for quantitative assessment. Based on this, the solution defines and calculates a key quantitative indicator system covering four dimensions: efficiency, quality, innovation, and resources. This breaks through the one-sidedness of traditional assessments that only focus on the number of patents or total investment, achieving a multi-faceted, quantifiable, and comparable accuracy of enterprise R&D strength. Furthermore, the solution adopts a two-layer architecture of horizontal fusion assessment and vertical time-series prediction to construct a comprehensive assessment model. At the horizontal level, it calculates the current R&D strength comprehensive index and outputs confidence levels through a subjective and objective weighting method (entropy weighting combined with expert scoring), overcoming the biases caused by single subjective weighting or pure data-driven approaches. This ensures that the static assessment results are both objective and strategically oriented. At the vertical level, it introduces mutual information analysis to screen for factors that influence the comprehensive index. This approach utilizes leading indicators with nonlinear leading effects and incorporates a Temporal Convolutional Network (TCN) model for multi-period prediction, outputting predicted values and their confidence intervals. This overcomes the limitations of existing evaluation methods, which can only perform static and post-hoc analyses, achieving a leap from current status assessment to trend prediction. It provides data intelligence support for forward-looking corporate decision-making. Finally, through triple statistical performance verification—weight sensitivity analysis, logical boundary testing, and time series cross-validation—the reliability of the model's output is rigorously confirmed, ensuring the robustness and reliability of the evaluation and prediction results in management applications. This forms a complete technical closed loop encompassing data governance, indicator construction, model evaluation / prediction, and result verification. This fundamentally transforms corporate R&D capability management from a traditional, fragmented, subjective, and static model to an integrated, objective, and dynamic intelligent decision-making model. It effectively addresses the three core pain points that have long plagued the industry: weak data foundation, incomplete evaluation dimensions, and lack of forward-looking insights, significantly improving the refinement, scientific rigor, and intelligence of R&D management. Attached Figure Description
[0012] Figure 1 This is a flowchart of a comprehensive data management method for enterprises' R&D capabilities, based on the present invention. Figure 2 To establish an automated data collection strategy and interface for this invention, the data is preprocessed, and a flowchart of the R&D fact wide table is constructed. Figure 3 The flowchart of the comprehensive evaluation model for enterprise R&D strength is presented in this invention, which employs a two-layer architecture of horizontal fusion evaluation and vertical time-series prediction. Detailed Implementation
[0013] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0014] Reference Figure 1 As shown, a comprehensive data management method for enterprise R&D capabilities includes: Identify the types of multi-source data throughout the enterprise's R&D process, set up automated data collection strategies and interfaces, preprocess the data, and construct a wide table of R&D facts; Based on a wide table of R&D facts, calculate the key quantitative indicators needed to assess R&D capabilities; Based on key quantitative indicators, a two-tier architecture of horizontal fusion assessment and vertical time-series prediction is adopted to construct a comprehensive assessment model of enterprise R&D strength. Based on the output of the comprehensive evaluation model of enterprise R&D capabilities, the reliability of the model's prediction results is confirmed through statistical performance verification.
[0015] It can be explained that comprehensive data management is the prerequisite and cornerstone for realizing the assessment of enterprise R&D capabilities from subjective, one-sided, and static to objective, comprehensive, and dynamic. Enterprise R&D capability assessment, on the other hand, is the value outlet and goal orientation of comprehensive data management. Therefore, this solution focuses on assessing enterprise R&D capabilities. Specifically, firstly, it is necessary to systematically sort out all kinds of key data generated throughout the entire lifecycle of enterprise R&D, from concept proposal, project approval, design and development, testing and verification to results transformation and iterative optimization. This involves clearly defining six core data sources: R&D project data, human resource data, intellectual property data, funding input data, equipment asset data, and results transformation data, and then drawing a diagram based on the R&D process. A data asset map clearly maps the data types, source systems, formats, and update frequencies at each stage. Using this map as a top-level design blueprint ensures the comprehensiveness and systematic nature of data collection, fundamentally avoiding problems such as data silos, inconsistent standards, and omissions of key information. After clarifying the data blueprint, the next step is to address the efficient and reliable acquisition of data. This solution employs a layered automated collection strategy, dividing the collection methods based on the openness and structure level of the source systems: fully automated interface collection (e.g., connecting to project management systems and code repository APIs), semi-automatic file import (e.g., periodically parsing structured reports exported from HR and financial systems), and manual confirmation and entry (used to supplement external benchmarking information). This strategy, which categorizes data into three types (including qualitative descriptions), maximizes the use of the enterprise's existing IT infrastructure. While significantly improving data collection efficiency, it also ensures the objectivity and timeliness of the data. However, when multi-source data is aggregated, the raw data often suffers from problems such as formatting issues, recording errors, and missing standards, making it unsuitable for direct analysis. Therefore, a standardized data preprocessing workflow is needed. Through a series of technical means such as data cleaning, format standardization, outlier handling, missing value imputation, text feature extraction, and category feature coding, the raw data is transformed into usable data. Then, using key fields such as project ID and employee ID, data association and fusion technology is used to connect information scattered across different data tables, constructing a comprehensive and consistent database. A traceable R&D fact table, stored in the enterprise data lake, provides a high-quality, single, and reliable data source for all subsequent analyses. After laying the data foundation, it is necessary to transform the data into measurable insights. Based on the R&D fact table, this solution defines and calculates key quantitative indicators covering four dimensions: efficiency (such as average delivery cycle of requirements and code deployment frequency), quality (such as defect rate per thousand lines of code), innovation and output (such as high-value patent density and revenue share of new products), and resources and effectiveness (such as R&D input-output ratio and key talent retention rate). This enables multi-dimensional and quantifiable measurement of the speed, reliability, value creation, and resource utilization efficiency of R&D activities, elevating management perception from vague experience to precise data.To further extract data value and empower decision-making, this solution adopts a two-layer architecture of horizontal fusion assessment and vertical time-series forecasting to construct a comprehensive assessment model of enterprise R&D strength. The horizontal fusion assessment layer uses a weighted method combining subjective and objective factors (entropy weighting combined with expert scoring) to weight and fuse key indicators for the current period, outputting a comprehensive R&D strength index and confidence level for the current period, achieving an objective and comprehensive assessment of the static status quo. The vertical time-series forecasting layer introduces mutual information analysis to automatically filter leading indicators with non-linear leading influence on the comprehensive index from historical indicator data. Then, a temporal convolutional network (TCN) model is used, taking the historical comprehensive index sequence and leading indicator sequence as input, to output the predicted R&D strength index value and its confidence level for future multiple periods. This approach, by overcoming the limitations of traditional methods that can only perform post-hoc analysis, achieves a leap from current status assessment to trend prediction. Finally, to ensure the reliability and credibility of the model's output assessment and prediction results and their ability to support key management decisions, this solution verifies the model's robustness through weight sensitivity analysis, validates the correctness of the model's mathematical logic through logical boundary testing, and evaluates the accuracy and trend judgment ability of the prediction model through time series cross-validation. This validation ensures the reliability and usability of the entire method, from data to model to output results, ultimately forming a complete, closed-loop, and implementable comprehensive data management solution for enterprise R&D capabilities, encompassing data governance, indicator construction, intelligent assessment / prediction, and result verification.
[0016] Reference Figure 2 As shown, the steps of setting up automated data collection strategies and interfaces, preprocessing the data, and constructing a wide table of R&D facts specifically include: Identify the multi-source data types involved in the entire enterprise R&D process; The multi-source data types include R&D project data, human resource data, intellectual property data, funding data, equipment asset data, and technology transfer data; Construct an enterprise R&D data asset map. This map takes the R&D process as the main line, maps the data types, data source systems, data formats and update frequencies generated at each stage, and forms a top-level design blueprint for data collection. Based on the data asset map, we design automated collection strategies and interfaces for each type of data. The automated data collection strategy is divided into three categories based on the openness of the data source system and the degree of data structuring: fully automated interface collection, semi-automatic file import, and manual confirmation and entry. The collected multi-source data were standardized and preprocessed, and a wide table of R&D facts was constructed. The preprocessing includes at least: data cleaning, format standardization, outlier handling, missing value imputation, text feature extraction, category feature encoding, and data association and fusion.
[0017] To achieve quantitative assessment and comprehensive data management of R&D capabilities, the primary task is to build a comprehensive, accurate, and timely data foundation. Therefore, the first step is to systematically review all key data categories generated throughout the entire R&D process, clearly define six core data categories, and construct a data asset map for global management. This map, as a core tool for data governance, clearly defines "where (process / system), what data, and how to obtain it," ensuring that data collection is systematic and without major omissions. After clarifying the data blueprint, the second step is to address "how to efficiently and reliably obtain data." This solution abandons the traditional model of relying solely on manual data entry and instead designs a hierarchical automated data collection strategy. This aims to maximize the use of existing IT systems, automatically acquiring data from the source, improving data objectivity and accuracy while ensuring efficiency. Specifically, fully automated interface collection is suitable for connecting to data sources with standard APIs, such as project management systems, code repositories, and financial systems; semi-automatic file import is suitable for structured reports regularly exported from HR systems and asset management systems; and human... The confirmation and reporting process serves as a supplement, used to collect external benchmarking information that cannot be automatically obtained and necessary qualitative descriptions. This process can be simplified by pre-filling data in the system and having the responsible person verify it. After the data is collected, the raw data often suffers from inconsistent formats, recording errors, and missing standards, making it unsuitable for direct analysis. Therefore, the third step implements a standardized data preprocessing workflow, using a series of technical means to transform the "raw data" into "usable data." Data cleaning aims to remove duplicate and erroneous records; format unification ensures that similar data have the same fields and units of measurement; outlier handling uses statistical methods to identify and reasonably correct values that significantly deviate from the normal range; missing value imputation uses interpolation, mean imputation, or prediction based on association rules to ensure data integrity; finally, through data association and fusion, information scattered across different data tables is connected using project ID, employee ID, and financial subject as key fields to form a comprehensive wide table reflecting the input, process, output, and environment of R&D activities. This table is stored in the enterprise's R&D data lake, providing high-quality data raw materials for subsequent indicator calculations, model analysis, and visualization. In specific implementation cases, the determination of multi-source data types involved in the entire enterprise R&D process includes: based on the entire lifecycle of enterprise R&D from concept proposal, project approval, design and development, testing and verification to results transformation and iterative optimization, the system identifies key data generated at each stage; wherein, the R&D project data originates from the project management system and includes, but is not limited to: basic project information, task progress, resource consumption, code submission records, defect tracking records, test reports, and project documents; the human resources data originates from the human resources management system and includes, but is not limited to: R&D personnel roster, organizational structure, skill matrix, performance records, training records, and personnel turnover; the intellectual property data originates from the intellectual property management system or the legal department. This includes, but is not limited to: patent application and authorization lists, software copyright registration information, trade secret registration ledgers, and patent maintenance status; the funding input data comes from financial systems or project accounting systems, including but not limited to: general ledgers and detailed ledgers of R&D expenses, budgets and actual expenditures by project / department / technology field, equipment procurement costs, and external cooperation costs; the equipment asset data comes from laboratory management systems, including but not limited to: R&D instrument and equipment ledgers, usage status records, calibration and maintenance logs, and shared reservation information; the technology transfer data comes from product data management, customer relationship management, and financial systems, including but not limited to: new product launch lists, technology transfer / licensing contracts, new product sales revenue, and technology transfer benefit analysis reports; In specific implementation cases, the design of automated data collection strategies and interfaces for each type of data specifically includes: For R&D project data, code data, and some test data: By calling the open API interfaces of the project management system, code repository, and continuous integration / continuous deployment platform, and writing scheduled task scripts, fully automated, real-time, or near-real-time synchronous collection of project status, commit logs, build results, and defect data can be achieved; For human resources data, financial expenditure data, and equipment asset data: In consultation with the human resources system, enterprise resource planning system, and asset management system, data reports in standard formats (such as CSV and Excel) are exported from the source system periodically (e.g., daily or weekly), and ETL tools are developed to achieve semi-automatic scheduled file capture, parsing, and storage. After storage, triggering... A notification is sent to the data administrator for consistency verification. For intellectual property data, technology transfer data, and external industry benchmarking data: a structured data entry template is designed in the system. For data that can be partially obtained automatically (such as the deadline for paying official intellectual property fees), the system pre-fills the data. For data that requires external querying or subjective judgment (such as the market impact of achievements or competitors' technological trends), designated data managers (such as intellectual property managers, product managers, and strategic researchers) regularly supplement and confirm the data. The system records the person filling out the form and the time of filling out the form, and supports uploading attachments as evidence. All data collected through automated and semi-automated methods is recorded and stored in the enterprise data lake, including the data source system, collection time, collection method, and original values, to ensure data traceability. In specific implementation cases, for text feature extraction (such as project document summaries, detailed defect descriptions, etc.), for technical documents, the TF-IDF algorithm is used to extract the word frequency weights of the top 50 key terms as feature vectors; for problem description text, a pre-trained all-MiniLM-L6-v2 model is used to generate 384-dimensional semantic embedding vectors as features. For category feature encoding (such as project status, patent legal status, etc.), one-hot encoding is used for unordered fields with few values (such as "project status"), and label encoding is used for ordered fields (such as "defect priority") (e.g., mapping low, medium, and high to [1,2,3]). A star schema is used to define a core fact table (whose primary key is the project number and year / month, containing aggregateable numerical indicators such as code submissions, valid defects, R&D man-hours, and expenses) and a related dimension table (whose primary key is the project number, containing project name, technology, etc.). The system uses Apache Spark SQL to perform a left join between the fact table and the project dimension table (whose primary key is the employee ID, which includes attributes such as department, skill tags, and performance level) and the personnel dimension table (whose primary key is the employee ID, which includes attributes such as department, skill tags, and performance level). For many-to-many relationships (such as multiple people participating in a project), the fact table is pre-expanded into multiple rows. The resulting full wide table, containing all original fields and derived feature columns, is written to a specified path in the enterprise data lake in Parquet columnar storage format. The schema version number and data time range of this version of the wide table are recorded. The left join refers to using each row of the left table (main table) as a basis to search for matching rows in the right table (lookup table). Regardless of whether a matching item is found in the right table, all rows in the left table will definitely appear in the final result.
[0018] The key quantitative indicators required for calculating and evaluating R&D capabilities based on the R&D fact wide table specifically include: Define the key quantitative indicator types needed to assess R&D capabilities, including: efficiency indicators, quality indicators, innovation and output indicators, and resource and efficiency indicators; Efficiency metrics include: average delivery time of requirements, code deployment frequency, and team iteration rate; Quality metrics include: defect rate per thousand lines of code, average time to resolve defects, and test case pass rate; Innovation and output indicators include: high-value patent density, revenue share of new products, number of registered technological achievements, and technology debt ratio; Resource and efficiency indicators include: R&D input-output ratio, key talent retention rate, per capita R&D output, and equipment utilization rate; Based on the data in the R&D fact table, the key quantitative indicators required to assess R&D capabilities are calculated, and a list of key indicators is generated.
[0019] To achieve accurate assessment and scientific management of R&D capabilities, after building a high-quality data foundation, the core task is to transform these "data raw materials" into directly measurable, comparable, and traceable "metrics." Therefore, this solution, based on a broad table of R&D facts, calculates key quantitative indicators needed to assess R&D capabilities. These indicators, as a core tool for value extraction, aim to answer key management questions such as R&D efficiency, innovation level, and effective resource utilization, ensuring that the assessment moves from vague perception to precise measurement. Specifically, efficiency indicators measure the speed and throughput of R&D activities, reflecting the agility in transforming ideas into deliverables; quality indicators assess the reliability and stability of R&D outputs, revealing the level of process control and technology debt management; innovation and output indicators quantify the value creation and knowledge accumulation of R&D activities, measuring technological leadership and market influence; and resource and efficiency indicators aim to understand the economic benefits and rationality of R&D investment, assessing the intensive and sustainable use of resources. It should be further explained that the specific quantitative relationships of each indicator are as follows: Average delivery time for a requirement = Σ(requirement closure time - requirement creation time) / total number of requirements; Code deployment frequency = the number of times a code is successfully deployed to the production environment per unit of time; Team iteration rate = number of story points or tasks completed per unit of time; Defect rate per thousand lines of code = (number of defects discovered after release / number of lines of code) × 1000; Average defect resolution time = Σ(defect closure time - defect creation time) / total number of defects; Test case pass rate = (number of passed test cases / total number of test cases) × 100%; High-value patent density = (number of high-value patents assessed / total number of R&D personnel). High-value patents need to be determined by combining the type of invention patent, overseas layout, patent citation rate and industrialization potential. New product revenue share = (New product sales revenue / Total company revenue) × 100%; Number of registered technological achievements = Total number of technological achievements registered through official recognition during the reporting period; Technical debt ratio = (critical code defects to be fixed + number of outdated components to be updated + unreasonable architectural items + number of missing core documents) / total number of system modules; R&D input-output ratio = (gross profit of new products + technology licensing revenue) / total R&D expenses; Key talent retention rate = (1 - number of key R&D personnel leaving / total number of key R&D personnel at the beginning of the period) × 100%; R&D output per employee = (gross profit from new products + patent licensing revenue + technical service revenue) / average number of R&D personnel; Equipment utilization rate = (effective equipment usage hours / total available equipment hours) × 100%.
[0020] Reference Figure 3 As shown, the comprehensive evaluation model for enterprise R&D strength constructed using a two-tier architecture of horizontal fusion assessment and vertical time-series prediction specifically includes: A two-layer architecture of horizontal fusion assessment and vertical time series prediction is adopted. The horizontal fusion assessment layer outputs the current R&D strength comprehensive index and assessment confidence level, while the vertical time series prediction layer outputs the predicted value of the R&D strength index for future multiple periods and its confidence interval. Based on the enterprise's R&D facts wide table, a list of key indicators is generated, and the values of key quantitative indicators over T historical periods are extracted to construct an indicator matrix. The last row of the indicator matrix is used as the input to the horizontal fusion evaluation layer, and the entire matrix is used as the input to the vertical time series prediction layer. The index matrix is subjected to standardized preprocessing, which includes at least: normalization based on index direction, scaling based on historical extreme values, and time series-based imputation for missing values. Based on the horizontal fusion evaluation layer, the weights of the indicators are determined by a combination of subjective and objective weighting methods. The combined weighting includes at least: calculating objective weights based on the entropy weight method, determining subjective weights based on the expert scoring method, and performing linear combination through the harmonic coefficient; The current comprehensive index is calculated using linear weighting based on the combined weights. The comprehensive confidence level is calculated through data integrity and volatility, and the current comprehensive index of R&D strength is output. Based on the vertical time series prediction layer, indicators that have a leading influence on the composite index are selected as prediction inputs through mutual information analysis based on the indicator matrix. The longitudinal time series prediction layer adopts a temporal convolutional network model framework, taking the comprehensive index and its leading influence indicators as inputs, and outputting the predicted value of the R&D strength index for future multi-cycle periods and its confidence interval. The trained two-layer architecture of horizontal fusion evaluation and vertical time series prediction is defined as a comprehensive evaluation model of enterprise R&D strength.
[0021] It can be explained that, after establishing a broad table of standard data collection for R&D facts and a specific system of key quantitative indicators, constructing an intelligent model that can comprehensively assess and output the confidence level of an enterprise's R&D strength and future trends is a crucial step in realizing the value of the entire system and empowering decision-making. Because the dimensions and directions of the indicators differ (some are better the larger they are, such as the pass rate; others are better the smaller they are, such as the defect rate), standardization is necessary. Therefore, the normalization based on the indicator direction specifically includes: For positive metrics (the higher the better, such as income share, pass rate): ; For negative metrics (the smaller the better, such as defect rate and delivery time): ; in, For the first The sample at the th The original values of the indicators The normalized index value, and For the first The indicator's actual minimum and maximum values in historical data result in all ; The calculation of objective weights based on the entropy weight method specifically includes: Calculate the first Entropy value of the indicator: ,in, , ; Calculate the coefficient of difference: ; Calculate the weights: ; Using the entropy weight method, based on the dispersion of the indicator data itself, the greater the variation of the indicator data (the more information it provides), the higher the weight. The method of determining subjective weights based on expert scoring specifically refers to inviting relevant experts to conduct pairwise comparisons and scoring of four major categories of indicators (efficiency, quality, innovation, and resources) and their specific internal indicators to form a judgment matrix and calculate the subjective weight vector. The harmonic coefficient is the degree of influence of objective entropy weight (data regularity) and subjective weight (strategic orientation). Based on the consistency test optimization method, the historical T-period data of the enterprise containing the comprehensive evaluation index and the real R&D performance label are extracted. The harmonic coefficient is traversed in the [0,1] interval with a step size of 0.1 and the corresponding comprehensive index is calculated. The harmonic coefficient that makes the evaluation result most consistent with the real R&D strength (minimum error) is found as the theoretical value through the Pearson correlation coefficient or the coefficient of determination. Furthermore, the current comprehensive index is calculated using linear weighting based on combined weights. The comprehensive confidence level is calculated based on data integrity and volatility. Specifically, the current comprehensive R&D strength index is obtained by summing the products of the combined weights and the indicator data in the last row of the corresponding indicator matrix. The index is then normalized to obtain its percentage value. Simultaneously, the comprehensive confidence level is calculated. The data integrity confidence level is the ratio of the number of non-empty indicators to the total number of indicators. The data volatility confidence level is the ratio of the number of indicators with controllable volatility (e.g., an absolute Z-score less than 2) to the total number of indicators. Both are weighted and summed according to a set weight (e.g., both set to 0.5) to obtain the comprehensive confidence level. Finally, the R&D strength comprehensive index and the comprehensive confidence level are output simultaneously.
[0022] The specific indicators that have a leading influence on the composite index through mutual information analysis include: From the enterprise R&D fact table, the target evaluation sequence and candidate indicator sequence are extracted and preprocessed. Define a set of positive integers as the set of leading periods to be examined; For each candidate indicator sequence and each set leading period, calculate the mutual information value between the value of the candidate indicator at the specified leading period and the value of the target evaluation sequence in the current period. For each mutual information value, a permutation test is used to perform a statistical significance test, and a mutual information reference distribution under the assumption of no association is constructed. For each candidate indicator, among all the leading periods that pass the significance test, select the leading period with the largest mutual information value and determine it as the best leading period of that indicator relative to the comprehensive strength index. Set a global mutual information threshold, and compare the maximum mutual information calculated for each candidate indicator under its best lead period with this threshold. Only candidate indicators whose maximum mutual information exceeds the threshold and pass the significance test are retained and identified as leading indicators with predictive value. All selected leading indicators are sorted in descending order of their maximum mutual information to generate a list of leading indicators. The identifier of each leading indicator, its optimal leading period, and the corresponding mutual information strength value are recorded.
[0023] This can be explained by the fact that when predicting the future trend of the comprehensive index of corporate R&D strength, many key quantitative indicators change before the comprehensive index and can affect the subsequent fluctuations of the comprehensive index (for example, some indicators rise first, and the comprehensive index will rise after a period of time). Mutual information analysis can accurately capture this "early correlation" between different indicators and the comprehensive index. Regardless of whether the correlation is linear or non-linear, it can judge the correlation strength through quantitative values, eliminate false correlations caused by random fluctuations, and find the indicators that truly have a leading impact on the comprehensive index. Adding these indicators to the prediction model can allow the model to capture the signals driving changes in R&D strength earlier, just like receiving an "early warning", thereby greatly improving the accuracy and reliability of future trend prediction and making the prediction results more valuable. When constructing a longitudinal time series prediction model, based on this list, the sequence of each leading indicator after shifting according to its best leading period is used as an auxiliary input feature and input into the prediction model along with the historical comprehensive strength index sequence for training and prediction. It should be noted that the target evaluation sequence is a sequence formed by sorting the comprehensive R&D strength index of each historical period by time, and the candidate indicator sequence is a series of independent sequences formed by sorting all the key quantitative indicators to be screened by time. Time point alignment is performed on all sequences to ensure that the same index position corresponds to the same statistical period. For sequences with missing data, time series interpolation is used to fill in the missing data to ensure that each sequence is continuous and complete. The mutual information value is used to quantify the nonlinear statistical dependency between two variables. The larger the value, the stronger the lead-lag correlation. Before calculation, continuous values need to be discretized and preprocessed. The set of leading periods to be examined is a set consisting of multiple preset positive integers, where each element represents a possible time lead that needs to be examined, in order to systematically analyze the leading correlation strength between the value of the candidate indicator and the current comprehensive strength index when it is several statistical periods in advance. To eliminate spurious associations caused by random fluctuations, the method of using permutation test to perform statistical significance testing specifically includes: constructing a mutual information reference distribution under the assumption of no association by randomly rearranging the target sequence multiple times and recalculating the mutual information; comparing the actual observed mutual information value with the reference distribution; if the actual value exceeds the preset significance threshold, it is determined that the association between the candidate indicator and the target evaluation sequence under the corresponding leading period is statistically significant. It should be further noted that the specific values of the global mutual information threshold and the preset significance threshold can be preset by the percentile method based on the distribution of historical data or the p-value method based on statistical hypothesis testing.
[0024] The temporal convolutional network model framework specifically includes: A temporal convolutional network model is constructed. The input of the model is a multi-dimensional time window containing the historical composite index sequence and the leading indicator sequence identified by mutual information analysis. The output is the composite index prediction sequence for multiple future periods and the confidence interval for each prediction point. The network structure of the model should include at least: A one-dimensional causal convolution kernel with an exponentially increasing dilation coefficient based on 2 is used. Four layers are stacked, with 64 filters in each layer and a kernel size of 3. Each TCN block contains a convolutional layer, weight normalization, ReLU activation function and random deactivation layer, and the block input and convolution output are added through residual connections; The output sequence of the last TCN layer is expanded in the time dimension and mapped through a fully connected layer; Quantile regression is used, and multiple independent fully connected layers are connected in parallel after the last layer as output heads to output the predicted values of the specified quantiles for multiple future periods. The specific steps of model training and prediction include: Using a historical time window of length T as the sample input feature and the real sequence of the composite index for multiple future periods as the sample label, a large number of training sample pairs are generated from the complete time series data through a sliding window. The mean squared error loss function is used to calculate the overall error between the model-predicted future multiple periodic sequences and the actual sequences. Multiple quantile output heads are trained simultaneously by minimizing the quantile loss function; The latest historical data window of length T is input into the trained TCN model, and the model directly outputs the composite index prediction sequence for multiple future periods and the quantile of each prediction point, and then synthesizes the confidence interval.
[0025] This can be explained by the fact that the TCN model, due to its causal convolutional structure, can ensure that the prediction does not depend on future information. At the same time, its dilated convolution design enables it to efficiently capture long-term historical patterns. Compared with recurrent neural networks, it is more stable in training and has a higher degree of parallelism. By integrating leading indicators, the model can more sensitively capture the leading signals that drive changes in comprehensive strength. Quantile regression provides a measure of uncertainty for point prediction, and the output confidence interval provides a quantitative basis for managers to judge the reliability of the prediction and formulate risk response strategies. Thus, it realizes the evolution from single-point prediction to probabilistic prediction and enhances the decision support value of the prediction results. In a specific implementation, the historical window length T of the model is 12 months, the input includes one comprehensive index sequence and K leading indicator sequences, the input dimension is [12, 1+K], the network depth is 4 layers, the initial inflation coefficient is initially set to 1, and then doubles to 2, 4, and 8 layer by layer. The quantile regression outputs the 10th, 50th, and 90th quantiles, and the 10th and 90th quantiles form an 80% confidence interval.
[0026] The verification of the reliability of the model's prediction results based on the output of the comprehensive evaluation model of enterprise R&D strength, through statistical performance validation, specifically includes: The output results of the comprehensive evaluation model of enterprise R&D strength were obtained and verified from three aspects: weight sensitivity analysis, logical boundary testing and time series cross-validation. The weight sensitivity analysis refers to applying a small-amplitude random perturbation to the weight of each indicator to generate multiple weight combinations, and calculating the comprehensive index and ranking respectively. If the standard deviation of the ranking change is less than the preset threshold and the Spearman rank correlation coefficient is greater than 0.95, then the model is not sensitive to the weight perturbation. The logical boundary test refers to constructing test datasets of extremely excellent and extremely poor performance and inputting them into the evaluation model. The output index of the extremely excellent data is required to be close to the full score, and the output index of the extremely poor data is required to be close to the minimum score, in order to verify the correctness of the mathematical logic and the ability to distinguish between different data. The time series cross-validation refers to using a rolling time window to train and predict multiple times, comparing the predicted values with the actual values, and calculating the mean absolute percentage error (MAPE) and the prediction direction accuracy. The requirement is that the MAPE is less than 25% and the direction accuracy is greater than 70%. A verification report is output based on the verification results. The verification report includes: Results of weighted sensitivity analysis, standard deviation of ranking change, and Spearman rank correlation coefficient; Output index of extremely excellent data and output index of extremely poor data; Mean absolute percentage error and prediction direction accuracy.
[0027] Explained by this, after constructing the comprehensive evaluation model for enterprise R&D capabilities, the reliability of its output results must be confirmed through systematic statistical verification to provide reliable support for management decisions. Therefore, robustness verification of the evaluation model is necessary. Weight sensitivity analysis simulates small changes in weights to test the stability of the model's output, avoiding drastic ranking fluctuations caused by the uncertainty of subjective weighting. Logical boundary testing mathematically verifies whether the model's behavior in extreme cases meets expectations, ensuring that its basic calculation logic is correct. Secondly, the accuracy of the prediction model needs to be verified. Time series cross-validation is the gold standard for evaluating time series prediction performance. It simulates the model's prediction process in real business when facing new data. Its output MAPE and directional accuracy directly reflect the model's prediction accuracy and trend judgment ability. Finally, all verification results are summarized into a rapid verification report to present the model's statistical performance in a clear and quantitative manner, and based on this, a conclusion is given on whether the conditions for pilot application are met, thus completing the complete technical closed loop from model construction to reliable performance confirmation. As an feasible case, based on the original weights of each indicator, a uniform random perturbation within the range of [-0.1, +0.1] is applied to generate multiple sets (e.g., 100 sets) of weight combinations, and the comprehensive index and the ranking of each evaluation object are calculated respectively. It is required that the standard deviation of the ranking change of all objects is less than a preset threshold (e.g., 2.0), and the Spearman rank correlation coefficient between the original ranking and the average ranking after perturbation is greater than 0.95, then the model is determined to be insensitive to the weight perturbation. Two test datasets are constructed: one extremely excellent (all positive indicators take the historical maximum value, and negative indicators take the historical minimum value) and the other extremely poor (all positive indicators take the historical minimum value, and negative indicators take the historical maximum value), and input into the evaluation model. The output index of the extremely excellent data is required to be ≥95 (out of 100), and the output index of the extremely poor data is required to be ≤5 (out of 100) to verify the correctness of the mathematical logic and the discrimination ability of the model.
[0028] Furthermore, based on the same inventive concept as the aforementioned comprehensive data management method for enterprise R&D capabilities, this solution proposes a comprehensive data management system for enterprise R&D capabilities, comprising: The data acquisition module is used to determine the multi-source data types in the entire enterprise R&D process, set automated acquisition strategies and interfaces, preprocess the data, and construct a R&D fact wide table. The model building and validation module is used to calculate key quantitative indicators required to evaluate R&D strength based on a wide table of R&D facts; based on the key quantitative indicators, a two-layer architecture of horizontal fusion evaluation and vertical time series prediction is adopted to build a comprehensive evaluation model of enterprise R&D strength; based on the output results of the comprehensive evaluation model of enterprise R&D strength, the reliability of the model prediction results is confirmed through statistical performance verification. The model building and verification module includes: A quantitative indicator unit is used to calculate key quantitative indicators required to evaluate R&D capabilities based on a wide table of R&D facts. The model building unit is used to construct a comprehensive evaluation model of enterprise R&D strength based on key quantitative indicators and using a two-layer architecture of horizontal fusion evaluation and vertical time series prediction. The model validation unit is used to comprehensively evaluate the output results of the model based on the enterprise's R&D capabilities, and to confirm the credibility of the model's prediction results through statistical performance verification.
[0029] In summary, the advantages of this invention are: it constructs quantitative indicators based on full-process data, and achieves dynamic evaluation and forward-looking prediction of R&D capabilities through a two-layer model that integrates mutual information analysis and temporal convolutional networks.
[0030] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A comprehensive data management method for enterprises' R&D capabilities, characterized in that, include: Identify the types of multi-source data throughout the enterprise's R&D process, set up automated data collection strategies and interfaces, preprocess the data, and construct a wide table of R&D facts; Based on a wide table of R&D facts, calculate the key quantitative indicators needed to assess R&D capabilities; Based on key quantitative indicators, a two-tier architecture of horizontal fusion assessment and vertical time-series prediction is adopted to construct a comprehensive assessment model of enterprise R&D strength. Based on the output of the comprehensive evaluation model of enterprise R&D capabilities, the reliability of the model's prediction results is confirmed through statistical performance verification.
2. The comprehensive data management method for enterprise R&D capabilities according to claim 1, characterized in that, The process of determining the multi-source data types throughout the enterprise's R&D process, setting automated data collection strategies and interfaces, preprocessing the data, and constructing a wide table of R&D facts specifically includes: Identify the multi-source data types involved in the entire enterprise R&D process; The multi-source data types include R&D project data, human resource data, intellectual property data, funding data, equipment asset data, and technology transfer data; Construct an enterprise R&D data asset map. This map takes the R&D process as the main line, maps the data types, data source systems, data formats and update frequencies generated at each stage, and forms a top-level design blueprint for data collection. Based on the data asset map, we design automated collection strategies and interfaces for each type of data. The automated data collection strategy is divided into three categories based on the openness of the data source system and the degree of data structuring: fully automated interface collection, semi-automatic file import, and manual confirmation and entry. The collected multi-source data were standardized and preprocessed, and a wide table of R&D facts was constructed. The preprocessing includes at least: data cleaning, format standardization, outlier handling, missing value imputation, text feature extraction, category feature encoding, and data association and fusion.
3. The comprehensive data management method for enterprise R&D capabilities according to claim 2, characterized in that, The key quantitative indicators required for calculating and evaluating R&D capabilities based on the R&D fact wide table specifically include: Define the key quantitative indicator types needed to assess R&D capabilities, including: efficiency indicators, quality indicators, innovation and output indicators, and resource and efficiency indicators; Efficiency metrics include: average delivery time of requirements, code deployment frequency, and team iteration rate; Quality metrics include: defect rate per thousand lines of code, average time to resolve defects, and test case pass rate; Innovation and output indicators include: high-value patent density, revenue share of new products, number of registered technological achievements, and technology debt ratio; Resource and efficiency indicators include: R&D input-output ratio, key talent retention rate, per capita R&D output, and equipment utilization rate; Based on the data in the R&D fact table, the key quantitative indicators required to assess R&D capabilities are calculated, and a list of key indicators is generated.
4. The comprehensive data management method for enterprise R&D capabilities according to claim 3, characterized in that, The construction of a comprehensive evaluation model for enterprise R&D strength based on key quantitative indicators and employing a two-tier architecture of horizontal fusion assessment and vertical time-series prediction specifically includes: A two-layer architecture of horizontal fusion assessment and vertical time series prediction is adopted. The horizontal fusion assessment layer outputs the current R&D strength comprehensive index and assessment confidence level, while the vertical time series prediction layer outputs the predicted value of the R&D strength index for future multiple periods and its confidence interval. Based on the enterprise's R&D facts wide table, a list of key indicators is generated, and the values of key quantitative indicators over T historical periods are extracted to construct an indicator matrix. The last row of the indicator matrix is used as the input to the horizontal fusion evaluation layer, and the entire matrix is used as the input to the vertical time series prediction layer. The index matrix is subjected to standardized preprocessing, which includes at least: normalization based on index direction, scaling based on historical extreme values, and time series-based imputation for missing values. Based on the horizontal fusion evaluation layer, the weights of the indicators are determined by a combination of subjective and objective weighting methods. The combined weighting includes at least: calculating objective weights based on the entropy weight method, determining subjective weights based on the expert scoring method, and performing linear combination through the harmonic coefficient; The current comprehensive index is calculated using linear weighting based on the combined weights. The comprehensive confidence level is calculated through data integrity and volatility, and the current comprehensive index of R&D strength is output. Based on the vertical time series prediction layer, indicators that have a leading influence on the composite index are selected as prediction inputs through mutual information analysis based on the indicator matrix. The longitudinal time series prediction layer adopts a temporal convolutional network model framework, taking the comprehensive index and its leading influence indicators as inputs, and outputting the predicted value of the R&D strength index for future multi-cycle periods and its confidence interval. The trained two-layer architecture of horizontal fusion evaluation and vertical time series prediction is defined as a comprehensive evaluation model of enterprise R&D strength.
5. The comprehensive data management method for enterprise R&D capabilities according to claim 4, characterized in that, The specific indicators that have a leading influence on the composite index through mutual information analysis include: From the enterprise R&D fact table, the target evaluation sequence and candidate indicator sequence are extracted and preprocessed. Define a set of positive integers as the set of leading periods to be examined; For each candidate indicator sequence and each set leading period, calculate the mutual information value between the value of the candidate indicator at the specified leading period and the value of the target evaluation sequence in the current period. For each mutual information value, a permutation test is used to perform a statistical significance test, and a mutual information reference distribution under the assumption of no association is constructed. For each candidate indicator, among all the leading periods that pass the significance test, select the leading period with the largest mutual information value and determine it as the best leading period of that indicator relative to the comprehensive strength index. Set a global mutual information threshold, and compare the maximum mutual information calculated for each candidate indicator under its best lead period with this threshold. Only candidate indicators whose maximum mutual information exceeds the threshold and pass the significance test are retained and identified as leading indicators with predictive value. All selected leading indicators are sorted in descending order of their maximum mutual information to generate a list of leading indicators. The identifier of each leading indicator, its optimal leading period, and the corresponding mutual information strength value are recorded.
6. The comprehensive data management method for enterprise R&D capabilities according to claim 5, characterized in that, The temporal convolutional network model framework specifically includes: A temporal convolutional network model is constructed. The input of the model is a multi-dimensional time window containing the historical composite index sequence and the leading indicator sequence identified by mutual information analysis. The output is the composite index prediction sequence for multiple future periods and the confidence interval for each prediction point. The network structure of the model should include at least: A one-dimensional causal convolution kernel with an exponentially increasing dilation coefficient based on 2 is used. Four layers are stacked, with 64 filters in each layer and a kernel size of 3. Each TCN block contains a convolutional layer, weight normalization, ReLU activation function and random deactivation layer, and the block input and convolution output are added through residual connections; The output sequence of the last TCN layer is expanded in the time dimension and mapped through a fully connected layer; Quantile regression is used, and multiple independent fully connected layers are connected in parallel after the last layer as output heads to output the predicted values of the specified quantiles for multiple future periods. The specific steps of model training and prediction include: Using a historical time window of length T as the sample input feature and the real sequence of the composite index for multiple future periods as the sample label, a large number of training sample pairs are generated from the complete time series data through a sliding window. The mean squared error loss function is used to calculate the overall error between the model-predicted future multiple periodic sequences and the actual sequences. Multiple quantile output heads are trained simultaneously by minimizing the quantile loss function; The latest historical data window of length T is input into the trained TCN model, and the model directly outputs the composite index prediction sequence for multiple future periods and the quantile of each prediction point, and then synthesizes the confidence interval.
7. The comprehensive data management method for enterprise R&D capabilities according to claim 6, characterized in that, The verification of the reliability of the model's prediction results based on the output of the comprehensive evaluation model of enterprise R&D strength, through statistical performance validation, specifically includes: The output results of the comprehensive evaluation model of enterprise R&D strength were obtained and verified from three aspects: weight sensitivity analysis, logical boundary testing and time series cross-validation. The weight sensitivity analysis refers to applying a small-amplitude random perturbation to the weight of each indicator to generate multiple weight combinations, and calculating the comprehensive index and ranking respectively. If the standard deviation of the ranking change is less than the preset threshold and the Spearman rank correlation coefficient is greater than 0.95, then the model is not sensitive to the weight perturbation. The logical boundary test refers to constructing test datasets of extremely excellent and extremely poor performance and inputting them into the evaluation model. The output index of the extremely excellent data is required to be close to the full score, and the output index of the extremely poor data is required to be close to the minimum score, in order to verify the correctness of the mathematical logic and the ability to distinguish between different data. The time series cross-validation refers to using a rolling time window to train and predict multiple times, comparing the predicted values with the actual values, and calculating the mean absolute percentage error (MAPE) and the prediction direction accuracy. The requirement is that the MAPE is less than 25% and the direction accuracy is greater than 70%. A verification report is output based on the verification results. The verification report includes: Results of weighted sensitivity analysis, standard deviation of ranking change, and Spearman rank correlation coefficient; Output index of extremely excellent data and output index of extremely poor data; Mean absolute percentage error and prediction direction accuracy.
8. A comprehensive data management system for enterprises' R&D capabilities, characterized in that: A comprehensive data management method for enterprise R&D capabilities as described in any one of claims 1-7 includes: The data acquisition module is used to determine the multi-source data types in the entire enterprise R&D process, set automated acquisition strategies and interfaces, preprocess the data, and construct a R&D fact wide table. The model building and validation module is used to calculate key quantitative indicators required to evaluate R&D strength based on a wide table of R&D facts; based on the key quantitative indicators, a two-layer architecture of horizontal fusion evaluation and vertical time series prediction is adopted to build a comprehensive evaluation model of enterprise R&D strength; based on the output results of the comprehensive evaluation model of enterprise R&D strength, the reliability of the model prediction results is confirmed through statistical performance verification.
9. A comprehensive data management system for enterprise R&D capabilities according to claim 8, characterized in that, The model building and verification module includes: A quantitative indicator unit is used to calculate key quantitative indicators required to evaluate R&D capabilities based on a wide table of R&D facts. The model building unit is used to construct a comprehensive evaluation model of enterprise R&D strength based on key quantitative indicators and using a two-layer architecture of horizontal fusion evaluation and vertical time series prediction. The model validation unit is used to comprehensively evaluate the output results of the model based on the enterprise's R&D capabilities, and to confirm the credibility of the model's prediction results through statistical performance verification.