Multi-dimensional data analysis-based medical and invasive power prediction method
By constructing a dynamic enterprise resource database and a Transformer model, combined with interpretive AI, the problems of data lag and subjectivity in traditional evaluation methods are solved, enabling efficient and intelligent evaluation and improvement guidance of enterprise innovation capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-03-13
AI Technical Summary
Traditional methods for assessing corporate innovation suffer from problems such as data lag, insufficient dimensional coverage, and strong subjectivity. They lack a systematic approach with a closed-loop process and are therefore unable to provide effective guidance for corporate decision-making.
We construct a dynamic enterprise resource database, use Python scripts and APIs to call data from multiple platforms, combine the Transformer model for feature learning, generate a science and technology innovation index through data cleaning, preprocessing and deep learning, and provide improvement suggestions using interpretive AI.
It enables objective, comprehensive, and dynamic evaluation of enterprise innovation capabilities, provides a unified and efficient evaluation system, enhances the timeliness and coverage of data, improves the intelligence and adaptability of evaluation, and supports comparative analysis across industries and time periods.
Smart Images

Figure CN121660147A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scientific and technological innovation capability assessment of research institutions, and more specifically, to a method for predicting scientific and technological innovation capability based on multi-dimensional data analysis. Background Technology
[0002] With the intensification of global technological competition, a company's technological innovation capability (referred to as "scientific and technological innovation power") has become an important indicator for measuring its sustainable development capability and core competitiveness. Currently, traditional assessments of corporate innovation power largely rely on static financial statements, qualitative scoring, or questionnaires, which suffer from problems such as strong data lag, insufficient dimensional coverage, and high subjectivity. These methods fail to comprehensively reflect a company's integrated innovation performance across multiple dimensions, including technological innovation, market responsiveness, internal management, and social responsibility. Furthermore, existing technologies for interpreting assessment results mostly remain at the score output stage, lacking effective explanatory mechanisms and improvement suggestion reasoning modules, making it difficult to provide directional guidance for corporate decision-making.
[0003] In recent years, the rapid development of artificial intelligence technology, especially the rise of deep learning and explained artificial intelligence (XAI), has provided new ideas for assessing corporate innovation capabilities. By constructing a multi-dimensional corporate resource database and combining it with advanced models such as Transformer to learn features and abstractly represent complex corporate behavioral data, more objective, comprehensive, and dynamic innovation capability modeling can be achieved. However, a unified and efficient systematic approach is still lacking in achieving a closed-loop process from data collection, preprocessing, modeling and evaluation to interpretation and recommendations. Summary of the Invention
[0004] In view of this, the present invention proposes a method for predicting scientific and technological innovation based on multi-dimensional data analysis, aiming to solve the problem of the lack of scientific and technological innovation processes in the current technology.
[0005] This invention proposes a method for predicting scientific and technological innovation based on multi-dimensional data analysis, comprising: Construct a dynamic enterprise resource database, and obtain enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information based on the dynamic enterprise resource database; The enterprise's technological innovation information, market performance information, operation and management information, and corporate social responsibility information are transmitted to the data processing module for preprocessing to obtain preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed corporate social responsibility information. The preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information are input into a deep learning model based on the Transformer architecture to obtain the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index. The technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index are transmitted to the data fusion module to obtain the enterprise's scientific and technological innovation index; By inputting the enterprise's technological innovation index into an interpretive AI, we can obtain directions for future improvement for the enterprise.
[0006] Furthermore, the specific content of constructing the dynamic enterprise resource database is as follows: using Python scripts and API calls to obtain enterprise information in real time from the State Intellectual Property Office, industrial and commercial and financial platforms, open paper databases, news and social media, and form a dynamic enterprise resource database.
[0007] Furthermore, the enterprise's technological innovation information includes its growth potential and strength; The company's market performance information includes market share, market growth rate, and revenue share of core products; The enterprise operation and management information includes the capabilities of the enterprise team; The corporate social responsibility information includes ESG ratings, energy conservation and emission reduction achievements, corporate credit, and corporate risks.
[0008] Furthermore, the specific content of transmitting enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information to the data processing module for preprocessing is as follows: Data cleaning algorithms are used to identify and remove duplicates, erroneous values, and incomplete records from the enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information. Then, NLP technology is applied to extract keywords, themes, and sentiment tendencies. Subsequently, outlier data is identified and processed using statistical rules and machine learning algorithms. Finally, normalization processing is performed to obtain preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information.
[0009] Furthermore, the specific content of using data cleaning algorithms to identify and remove duplicates, erroneous values, and incomplete records for enterprise technological innovation information, market performance information, operation and management information, and corporate social responsibility information is as follows: Hash matching, unique key merging, and field similarity calculation algorithms are used to compare enterprise data from multiple sources. Records with identical field content or similarity exceeding a preset threshold are identified as duplicates, automatically marked, and only one master record with the highest confidence is retained. The remaining duplicate information is deleted. Preliminary screening is performed using interval rules and statistical distribution methods, combined with logical rules for error judgment. For identified erroneous items, interpolation estimation or replacement correction is performed based on historical trends and industry averages. If estimation is not possible, the record is discarded. Records with more than 30% of the total number of missing fields are directly discarded; records with less than 30% of the total missing information are completed using multiple imputation methods.
[0010] Furthermore, the specific content of identifying and processing outlier data based on statistical rules and machine learning algorithms is as follows: outlier identification is performed using box plots and standard deviation rules. If data deviation is detected but can be corrected by fitting historical trend lines, time interpolation or industry averages are used for correction. If misrecording or structural errors are found, the record is directly removed. If the label field is found to have an unexpected classification, it is remapped to other classes or missing classes.
[0011] Furthermore, the preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information are input into a deep learning model based on the Transformer architecture to obtain the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index, expressed as: ; in, , Expressed as an index for evaluating technological innovation capabilities, (*) represents the first processing function. This is expressed as pre-processed enterprise technological innovation information; ; in, Expressed as a market competitiveness index, (*) is expressed as the second processing function. This is expressed as pre-processed information on the company's market performance. ; in, Expressed as a management efficiency index, (*) is expressed as the third processing function. This is expressed as pre-processed enterprise operation and management information; ; in, Expressed as a social responsibility index, (*) is expressed as the fourth processing function. This is expressed as pre-processed corporate social responsibility information.
[0012] Furthermore, the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index are transmitted to the data fusion module to obtain the enterprise's scientific and technological innovation index, which is expressed as: ; in, As an index of enterprise scientific and technological innovation, As the first weight, As the second weight, As the third weight, It is the fourth weight.
[0013] Furthermore, the specific content of inputting the enterprise's scientific and technological innovation index into the interpretive AI to obtain the enterprise's future improvement direction is as follows: the enterprise's scientific and technological innovation index and its constituent sub-indicators are submitted as input feature vectors to the interpretive AI analysis module. The interpretive AI analysis module uses the feature attribution method of attention mechanism to perform interpretive analysis on the input indicators, and outputs suggestions for future improvement directions based on the interpretive analysis results.
[0014] Furthermore, the proposed directions for future improvement include increasing R&D investment, optimizing patent portfolio, improving market strategies, enhancing organizational collaboration efficiency, or strengthening social responsibility practices.
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a multi-dimensional data analysis-based method for predicting scientific and technological innovation. Through Python scripts and multi-platform API interfaces, it collects data from multiple sources, including the State Intellectual Property Office, business and financial platforms, academic paper databases, news and social media, constructing a comprehensive information database covering enterprise technology, market, management, and responsibility. This database supports real-time updates, enhancing the timeliness and coverage of the data. Furthermore, this invention introduces cleaning algorithms for multi-source heterogeneous data, including hash matching, field similarity, logical rule recognition, historical trend interpolation, multiple imputation of missing values, and outlier identification. This significantly improves the integrity, consistency, and usability of the data, laying a solid foundation for subsequent modeling. Finally, this invention utilizes a deep learning-based Transformer model to analyze an enterprise's technological innovation, market competition, and operations. This invention utilizes feature learning and index output across management and social responsibility dimensions, leveraging NLP and time-series modeling capabilities to enhance the intelligence and adaptability of evaluation indicators. It inputs the four dimensions of indices into a weighted fusion model to generate a unified enterprise innovation index, providing a highly comparable and stable unified evaluation indicator that facilitates cross-industry and cross-time-period comparative analysis. Through an interpretable AI module based on an attention mechanism, the invention analyzes the contribution of each sub-indicator to the overall enterprise innovation capability, providing customized improvement suggestions (such as optimizing patent strategies and improving collaborative efficiency), significantly enhancing the application value and management support functions of the evaluation system. This invention is applicable to the monitoring and evaluation of technological innovation capabilities in enterprises of different industries and sizes, and can also serve as an important tool for governments, investment institutions, and research institutions in enterprise selection, policy formulation, and risk warning. Attached Figure Description
[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating a method for predicting scientific and technological innovation based on multi-dimensional data analysis, as described in an embodiment of the present invention. Detailed Implementation
[0017] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] See Figure 1 As shown, this embodiment of the invention provides a method for predicting scientific and technological innovation based on multi-dimensional data analysis, including: S1: Construct a dynamic enterprise resource database and obtain enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information based on the dynamic enterprise resource database; S2: Transmit the enterprise's technological innovation information, market performance information, operation and management information, and corporate social responsibility information to the data processing module for preprocessing to obtain preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed corporate social responsibility information; S3: Input the preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information into a deep learning model based on the Transformer architecture to obtain the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index. S4: Transmit the technology innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index to the data fusion module to obtain the enterprise's science and technology innovation index; S5: Input the enterprise's technological innovation index into the interpretive AI to obtain the enterprise's future improvement direction.
[0019] Furthermore, the specific content of constructing the dynamic enterprise resource database is as follows: using Python scripts and API calls to obtain enterprise information in real time from the State Intellectual Property Office, industrial and commercial and financial platforms, open paper databases, news and social media, and form a dynamic enterprise resource database.
[0020] Furthermore, the enterprise's technological innovation information includes its growth potential and strength; The company's market performance information includes market share, market growth rate, and revenue share of core products; The enterprise operation and management information includes the capabilities of the enterprise team; The corporate social responsibility information includes ESG ratings, energy conservation and emission reduction achievements, corporate credit, and corporate risks.
[0021] Furthermore, the enterprise's growth potential includes technological quality (authorized invention patents), technological competitiveness (comprehensive algorithm for technological competitiveness), and technological industrialization (recent patents, university-collaboration patents, and financing). The strength of the enterprise includes its size (registered capital, registration time, paid-in capital), growth rate (positive growth rate of social security employees over the past three years), technology certification (national, provincial, and municipal level technology certified enterprises), and innovation platform (national, provincial, and municipal level innovation platforms). The capabilities of the enterprise team include the technical level of senior executives (the number of patents held by directors, supervisors and senior executives) and the stability of the R&D team (the average tenure of patent inventors). The corporate credit includes quality credit (China Quality Award, nomination award | Provincial Governor's Quality Award, nomination award) and tax credit (taxpayer). The enterprise risks mentioned include operational risks (abnormal operations, administrative penalties, shell company index), financing risks (equity pledge, movable property mortgage, land mortgage), and judicial risks (enforcement judgment debtors, dishonest enforcement judgment debtors, major tax violations, serious violations of trust).
[0022] Specifically, the data collection period is the most recent three full fiscal years and the most recent twelve-month rolling period. The field set includes: authorized invention patents, recent patents, university cooperation patents, financing information, registered capital, paid-in capital, registration time, monthly sequence of social security employees, science and technology certification level, innovation platform level, list of patents held by directors, supervisors and senior executives, start and end dates of patent inventors' employment, market share, market growth rate, revenue share of core products, ESG rating, annual report data on energy conservation and emission reduction, quality credit and tax credit, list of operational / financing / judicial risk events, unified unique identifier for enterprises, and removal of duplicate names and codes and duplicate branch offices. Missing values are imputed by the median of the quantiles of the same size in the same industry, and extreme values are truncated by the P1 / P99 quantiles. For monetary values, the natural logarithm is taken and then standardized. For percentage and rate values, the range is limited to 0-100%. The industry benchmark is grouped by industry code + revenue size, and P10 and P90 are calculated as robust ranges. Sub-indicator calculation: Quantitative score, 0–100 points, using the default "quantile normalization" rule: score = 100 × (value - P10) / (P90 - P10), and truncated within 0–100. A: Enterprise growth potential A1: Technical quality: The number of invention patents authorized in the last five years is counted and scored according to the "per 100 R&D personnel" standard; 5 points are added when the proportion of patents maintained for ≥3 years is ≥60%, with a maximum of 100 points.
[0023] A2: Technological Competitiveness (Comprehensive Algorithm, Four Items with Equal Weights): IPC Coverage (Number of Different Subclasses / Industry P90) Normalized Score, Family Geographic Coverage (Number of Published Countries) Normalized Score, Average of Independent Claims Normalized Score, Effective Survival Rate = Number of Claims Under Examination or Granted and Not Terminated / Total Number of Claims Reverse Loss Rate Score, All four items are averaged after being taken from 0 to 100. A3: Technology Industrialization: Score of the proportion of newly authorized inventions in the past three years (weight 0.4), score of the proportion of patents in cooperation with universities (0.2), and score of the peer percentile of the most recent round of financing amount / valuation (0.4). Growth potential score = 0.35×A1 + 0.40×A2 + 0.25×A3, constraint: if any of A1, A2, or A3 is less than 30, then the upper limit of growth potential is 60.
[0024] B: Corporate Strength B1: Company Size: Both registered capital and paid-in capital are worth points. If the paid-in / registered capital ratio is less than 20%, 20 points will be deducted from this item. The minimum is not less than 0. B2: Growth rate: The number of times the number of people covered by social security has increased year-on-year is scored (0 times = 0 points, 1 time = 40 points, 2 times = 70 points, 3 times = 100 points). B3: Science and Technology Recognition: National level = 100, Provincial level = 80, Municipal level = 60. If multiple levels coexist, the highest score will be used and 10 points will be added, with a maximum of 100.
[0025] B4 Innovation Platform: National level = 100, Provincial level = 80, Municipal level = 60, and the same bonus rules apply. Enterprise strength score = 0.30×B1 + 0.25×B2 + 0.25×B3 + 0.20×B4, constraint: if the paid-in capital / registered capital ratio is less than 10%, the maximum enterprise strength score is 50.
[0026] C: Corporate Team Capabilities C1: Technical level of senior executives: The percentage of patents held and owned by directors, supervisors and senior executives to the total number of patents of the company is scored according to percentile. If the percentage is ≥30%, add 5 points. C2: R&D team stability: Average tenure of patent inventors (months), 30 points for 12 months or less, 100 points for 36 months or more; Team capability score = 0.5 × C1 + 0.5 × C2.
[0027] D: Corporate Market Performance D1: Market share percentile score.
[0028] D2: Market growth rate percentile score.
[0029] D3: Core product revenue share: 60%–85% is the optimal range. If it falls into the range, it scores 100 points. If it is below 40% or above 95%, it scores 60 points. The rest are scored by linear interpolation.
[0030] Market performance score = 0.4 × D1 + 0.4 × D2 + 0.2 × D3.
[0031] E: Corporate Social Responsibility E1: ESG rating mapping: AAA / AA / A / BBB / BB / B / CCC correspond to 100 / 90 / 80 / 70 / 60 / 50 / 30 points; E2: Energy conservation and emission reduction results: percentile score of the percentage decrease in energy consumption or carbon intensity per unit of revenue in the past three years, with 5 points added for three consecutive years of decline.
[0032] E3: Enterprise Credit: 100 points for winning the China Quality Award, 90 points for a nomination, 100 points for winning the Provincial Governor's Quality Award, and 90 points for a nomination. If both categories exist, the highest score is +10 points, with a maximum of 100 points. E4: Enterprise Risk (Deduction Items): -10 points for each abnormal operation, -10 points for each administrative penalty, -20 points for a high shell company index (above P90 in the same industry); -5 points for each equity pledge, -5 points for each movable property mortgage, -5 points for each land mortgage; -40 points for being an enforcement debtor, -60 points for being a dishonest enforcement debtor, -60 points for major tax violations, -80 points for serious violations of trust. The lowest risk is -100 points, and E4 is a negative value; Social responsibility score = 0.35×E1 + 0.25×E2 + 0.20×E3 + E4 (cut off between 0 and 100).
[0033] Furthermore, the specific content of transmitting enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information to the data processing module for preprocessing is as follows: Data cleaning algorithms are used to identify and remove duplicates, erroneous values, and incomplete records from the enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information. Then, NLP technology is applied to extract keywords, themes, and sentiment tendencies. Subsequently, outlier data is identified and processed using statistical rules and machine learning algorithms. Finally, normalization processing is performed to obtain preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information.
[0034] Furthermore, the specific content of using data cleaning algorithms to identify and remove duplicates, erroneous values, and incomplete records for enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information is as follows: Hash matching, unique key merging, and field similarity calculation algorithms (such as Levenshtein distance) are used to compare enterprise data from multiple sources. Records with identical field content or similarity exceeding a preset threshold (such as 95%) are identified as duplicates, automatically marked, and only one primary record with the highest confidence is retained, while other duplicate information is deleted. For numerical data (such as R&D investment, number of personnel, carbon emission values, etc.), interval rules and statistical distribution methods (such as mean ± 3 standard deviations) are used for preliminary screening, while logical rules (such as "R&D investment cannot exceed total enterprise revenue" and "the number of patents cannot be negative") are combined to make error judgments. For identified errors, interpolation estimation or replacement correction is performed by combining historical trends and industry averages. If estimation is not possible, the record is removed. Records with more than 30% of the total number of missing fields are directly removed. For records with a small amount of missing information (such as 1-2 fields), multiple imputation or KNN nearest neighbor imputation algorithm is used to complete the data to ensure data integrity.
[0035] Specifically, for enterprise technological innovation information, market performance information, operation and management information, and corporate social responsibility information, multi-dimensional verification rules and hash comparison algorithms are constructed to automatically identify and remove duplicate items, logically conflicting data (such as serious discrepancies between the number of enterprise patents and publicly available records), format errors (such as misaligned time fields or non-numerical monetary fields), and records with an excessively high proportion of missing fields, ensuring that the data structure is standardized and the fields are complete. Subsequently, Natural Language Processing (NLP) technology is applied to perform semantic analysis on unstructured text data (such as enterprise news reports, social media comments, and academic paper abstracts). This includes: extracting text keywords, topic tags, and technical terms based on BERT or similar pre-trained models to construct an enterprise text profile; using sentiment analysis models to assess the positive and negative impact of external public opinion on the enterprise; identifying the semantic center and domain trends of the text through topic modeling methods (such as LDA); and finally, using a method combining statistical rules and machine learning models to identify and process outlier data. Specifically, this includes: using IQR (Interquartile Range) and Z-score methods to identify distribution anomalies in structured indicator data (such as R&D investment ratio, profit margin, etc.); combining algorithms such as Isolation Forest and Local Outlier Factor (LOF) to determine, correct, or remove outliers in multidimensional space, improving the stability and representativeness of data distribution; and finally, normalizing the data. For numerical features, Min-Max normalization or Z-score normalization is used, while for textual features, TF-IDF vectorization or embedding representation methods are employed to ensure that data from different sources and scales are comparable and have a uniform scale in the model. This results in a preprocessed dataset with a unified structure, clear semantics, and no noise interference, namely, preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed corporate social responsibility information.
[0036] It should be noted that the above preprocessing steps not only significantly improve the quality and consistency of the data, but also greatly enhance the model's ability to recognize semantic information and complex features, effectively avoiding evaluation biases caused by redundancy, noise, and outliers in the original data. Simultaneously, this module provides a structured, semantically rich, and highly reliable input foundation for subsequent feature extraction, index calculation, and enterprise innovation index prediction, thereby improving the overall prediction accuracy and stability of the system and supporting more interpretable and action-oriented evaluation outputs.
[0037] Furthermore, the specific content of identifying and processing outlier data based on statistical rules and machine learning algorithms is as follows: outlier identification is performed using box plots and standard deviation rules. If data deviation is detected but can be corrected by fitting historical trend lines, time interpolation or industry averages are used for correction. If misrecording or structural errors are found, the record is directly removed. If the label field is found to have an unexpected classification, it is remapped to other classes or missing classes.
[0038] It should be noted that using statistical rules and machine learning algorithms to identify and process outlier data effectively reduces model bias and instability caused by extreme values or misrecording; trend interpolation and industry mean repair methods preserve structural changes in real economic behavior and improve data integrity; standard rules and automatic mapping mechanisms are used to handle label anomalies, ensuring the consistency of classification model inputs and the interpretability of prediction logic; and full-process log recording and removal traceability mechanisms are supported, improving system controllability and the efficiency of manual intervention.
[0039] Furthermore, the preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information are input into a deep learning model based on the Transformer architecture to obtain the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index, expressed as: ; in, , Expressed as an index for evaluating technological innovation capabilities, (*) represents the first processing function. This is expressed as pre-processed enterprise technological innovation information; ; in, Expressed as a market competitiveness index, (*) is expressed as the second processing function. This is expressed as pre-processed information on the company's market performance. ; in, Expressed as a management efficiency index, (*) is expressed as the third processing function. This is expressed as pre-processed enterprise operation and management information; ; in, Expressed as a social responsibility index, (*) is expressed as the fourth processing function. This is expressed as pre-processed corporate social responsibility information.
[0040] Specifically, the processing function performs feature interaction modeling and context relationship modeling on enterprise information to generate a global feature vector, and the output layer function maps the global feature function into a single numerical exponential result.
[0041] Furthermore, the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index are transmitted to the data fusion module to obtain the enterprise's scientific and technological innovation index, which is expressed as: ; in, As an index of enterprise scientific and technological innovation, As the first weight, As the second weight, As the third weight, It is the fourth weight.
[0042] The weighting coefficients are used to evaluate the dispersion of each index in the sample by standardizing the four indices to a uniform range. The greater the dispersion, the stronger the discrimination and the higher the weight is assigned. The weights are assigned according to the dispersion ratio and then normalized.
[0043] Specifically, if a company's science and technology innovation index is ≥95, the company is rated A++; if it is 90... If a company's science and technology innovation index is less than 95 points, its rating is A+; if it is 90 points... If a company's science and technology innovation index is less than 95 points, its rating is A+; if it is 85 points... If a company's science and technology innovation index is less than 90 points, its rating is A; if it is 80 points... If a company's science and technology innovation index is less than 85 points, its rating is B++; if it is 75 points... If a company's science and technology innovation index is less than 80 points, its rating is B+; if it is 70 points... If a company's science and technology innovation index is less than 75 points, its rating is B; if it is 65 points... If a company's innovation index is less than 70 points, its rating is C++; if it is 60... If a company's science and technology innovation index is less than 65 points, its rating is C+; if it is 55 points... If a company's science and technology innovation index is less than 60 points, its grade is C; if it is 50... If a company's science and technology innovation index is less than 55 points, its rating is D++; if it is 45... If an enterprise's science and technology innovation index is less than 50 points, the enterprise is rated as D+; if the enterprise's science and technology innovation index is less than 45 points, the enterprise is rated as D.
[0044] Furthermore, the specific content of inputting the enterprise's scientific and technological innovation index into the interpretive AI to obtain the enterprise's future improvement direction is as follows: the enterprise's scientific and technological innovation index and its constituent sub-indicators are submitted as input feature vectors to the interpretive AI analysis module. The interpretive AI analysis module uses the feature attribution method of attention mechanism to perform interpretive analysis on the input indicators, and outputs suggestions for future improvement directions based on the interpretive analysis results.
[0045] Specifically, the enterprise science and technology innovation index and its constituent sub-indicators, including the technology innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index, are used to form a feature vector. The data is then submitted to the interpretive AI analysis module for analysis.
[0046] The explanatory AI analysis module adopts a feature attribution method based on the attention mechanism or SHAP (Shapley Additive Explanations) algorithm to quantitatively explain the influence of each dimension of the input feature vector on the overall enterprise innovation index. The specific steps include: (1) Internal model weight attribution analysis: using the attention weight distribution or SHAP value to calculate the marginal contribution of the four types of input indicators, identifying which factor in the four dimensions contributes more to the result index and has less room for improvement, and which factor scores lower, contributes less but has greater potential for impact, thereby locking in potential improvement directions. (2) Constructing an explanatory mapping map: generating a visual feature explanation map (such as radar chart, importance heat map, etc.) based on the attribution results, and comparing the current performance of the enterprise with the industry average or best practice enterprise to assess the degree of gap and the urgency of improvement. (3) Generating improvement direction suggestions: combining the characteristics of the industry in which the enterprise is located, the current indicator performance, the historical improvement path and the domain knowledge graph, outputting targeted improvement suggestions through the rule engine or recommendation model. (4) Generate a corporate development recommendation report: Integrate the above analysis results and generate a corporate development recommendation report that includes current status diagnosis, key weakness identification, interpretable charts and future improvement paths, for reference by corporate management or investment decision-makers.
[0047] The recommendations may include: increasing R&D resource allocation and strengthening the layout of high-value patents; optimizing product market structure and improving brand communication efficiency; simplifying decision-making processes and improving the automation level of internal processes; expanding the scope of ESG governance and strengthening energy conservation and social responsibility fulfillment capabilities.
[0048] It should be noted that inputting the enterprise innovation index into explanatory AI can enhance the transparency of enterprise management: visualized attribution results help enterprises clearly identify internal weaknesses and provide strategic improvement guidance; the model output can be used to discover future potential directions and guide the optimal allocation of resources; it can improve the trustworthiness of the evaluation system: compared with traditional scoring systems, the output of explanatory AI has clear logic and traceable evidence; and it can support personalized development path planning: by combining industry and historical data, it can provide differentiated improvement strategies for different types of enterprises.
[0049] Furthermore, the proposed directions for future improvement include increasing R&D investment, optimizing patent portfolio, improving market strategies, enhancing organizational collaboration efficiency, or strengthening social responsibility practices.
[0050] Specifically, this invention can be applied to (1) jointly modeling and launching data financial products with financial institutions; (2) assisting financial institutions in acquiring customers and serving financial institutions in conducting preliminary due diligence; (3) applying to risk compensation policies to provide data support for enterprise access and risk monitoring; and (4) providing data reference for departmental (park) policy formulation and industry analysis.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for predicting scientific and technological innovation based on multi-dimensional data analysis, characterized in that, include: Construct a dynamic enterprise resource database, and obtain enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information based on the dynamic enterprise resource database; The enterprise's technological innovation information, market performance information, operation and management information, and corporate social responsibility information are transmitted to the data processing module for preprocessing to obtain preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed corporate social responsibility information. The preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information are input into a deep learning model based on the Transformer architecture to obtain the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index. The technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index are transmitted to the data fusion module to obtain the enterprise's scientific and technological innovation index; By inputting the enterprise's technological innovation index into an interpretive AI, we can obtain directions for future improvement for the enterprise.
2. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 1, characterized in that, The specific content of constructing the dynamic enterprise resource database is as follows: using Python scripts and API calls to obtain enterprise information in real time from the State Intellectual Property Office, industrial and commercial and financial platforms, open paper databases, news and social media, and form a dynamic enterprise resource database.
3. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 2, characterized in that, The enterprise's technological innovation information includes its growth potential and strength. The company's market performance information includes market share, market growth rate, and revenue share of core products; The enterprise operation and management information includes the capabilities of the enterprise team; The corporate social responsibility information includes ESG ratings, energy conservation and emission reduction achievements, corporate credit, and corporate risks.
4. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 3, characterized in that, The specific content of transmitting enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information to the data processing module for preprocessing is as follows: Data cleaning algorithms are used to identify and remove duplicates, errors, and incomplete records from corporate technology innovation information, corporate market performance information, corporate operation and management information, and corporate social responsibility information. Then, NLP technology is applied to extract keywords, themes, and sentiment. Subsequently, outlier data is identified and processed using statistical rules and machine learning algorithms. Finally, normalization is performed to obtain preprocessed corporate technology innovation information, preprocessed corporate market performance information, preprocessed corporate operation and management information, and preprocessed corporate social responsibility information.
5. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 4, characterized in that, The specific content of using data cleaning algorithms to identify and remove duplicates, erroneous values, and incomplete records from enterprise technological innovation information, enterprise market performance information, enterprise operation and management information, and enterprise social responsibility information is as follows: The system compares enterprise technology innovation information, market performance information, operation and management information, and corporate social responsibility information from multiple sources using hash matching, unique key merging, and field similarity calculation algorithms. Records with identical field content or similarity exceeding a preset threshold are identified as duplicates, automatically marked, and only one master record with the highest confidence is retained, while other duplicate information is deleted. Preliminary screening is performed using interval rules and statistical distribution methods, combined with logical rules for error judgment. For identified erroneous items, interpolation estimation or replacement correction is performed based on historical trends and industry averages. If estimation is not possible, the record is discarded. Records with more than 30% of the total number of missing fields are directly discarded; records with less than 30% of the total number of missing information are completed using multiple imputation methods.
6. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 5, characterized in that, The specific content of the outlier data identification and processing method based on statistical rules and machine learning algorithms is as follows: Outlier identification is performed using box plots and standard deviation rules. If a data deviation is detected but can be corrected by fitting a historical trend line, time interpolation or industry mean is used for correction. If a misrecorded or structural error is found, the record is directly removed. If an unexpected classification is found in the label field, it is remapped to another class or a missing class.
7. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 6, characterized in that, The preprocessed enterprise technological innovation information, preprocessed enterprise market performance information, preprocessed enterprise operation and management information, and preprocessed enterprise social responsibility information are input into a deep learning model based on the Transformer architecture to obtain a technological innovation capability evaluation index, a market competitiveness measurement index, a management efficiency index, and a social responsibility index, expressed as: ; in, , Expressed as an index for evaluating technological innovation capabilities, (*) represents the first processing function. This is expressed as pre-processed enterprise technological innovation information; ; in, Expressed as a market competitiveness index, (*) is expressed as the second processing function. This is expressed as pre-processed information on the company's market performance. ; in, Expressed as a management efficiency index, (*) is expressed as the third processing function. This is expressed as pre-processed enterprise operation and management information; ; in, Expressed as a social responsibility index, (*) is expressed as the fourth processing function. This is expressed as pre-processed corporate social responsibility information.
8. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 7, characterized in that, The process involves transmitting the technological innovation capability evaluation index, market competitiveness measurement index, management efficiency index, and social responsibility index to the data fusion module to obtain the enterprise's scientific and technological innovation index, expressed as: ; in, As an index of enterprise scientific and technological innovation, As the first weight, As the second weight, As the third weight, It is the fourth weight.
9. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 8, characterized in that, The specific content of inputting the enterprise's scientific and technological innovation index into the interpretive AI to obtain the enterprise's future improvement direction is as follows: The enterprise's scientific and technological innovation index and its constituent sub-indicators are submitted as input feature vectors to the interpretive AI analysis module. The interpretive AI analysis module uses the feature attribution method of attention mechanism to perform interpretive analysis on the input indicators, and outputs suggestions for future improvement directions based on the interpretive analysis results.
10. The method for predicting scientific and technological innovation based on multi-dimensional data analysis according to claim 9, characterized in that, The proposed directions for future improvement include increasing R&D investment, optimizing patent portfolio, improving market strategies, enhancing organizational collaboration efficiency, or strengthening social responsibility practices.