Machine learning-based enterprise technology risk diagnosis method and system

By using machine learning methods to dynamically correlate and analyze internal R&D data and external technical intelligence, risk patterns are identified and verified, and visualized risk reports are generated. This solves the problem of inaccurate risk pattern identification results in existing technologies and improves the stability and accuracy of risk pattern identification.

CN122491938APending Publication Date: 2026-07-31SUZHOU JIANG KEY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU JIANG KEY TECH CO LTD
Filing Date
2026-06-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to dynamically correlate and analyze continuously changing corporate R&D data and external technical intelligence in enterprise technology risk analysis, resulting in insufficient accuracy and stability of risk pattern identification results.

Method used

By employing machine learning methods, internal R&D data and external technical intelligence are collected, cleaned, and structured to generate a basic enterprise technical dataset. This dataset is then analyzed for trends, information contribution, and redundancy. Effective features are selected and integrated, and candidate risk patterns are identified using random forests. Consistency verification and reliability screening are performed to generate risk pattern recognition results. Finally, a visualized risk report is generated, and risk mitigation measures are developed, forming a closed loop for risk management.

Benefits of technology

It improves the dynamic adaptability of the risk pattern recognition process to data changes, enhances the stability and accuracy of risk pattern recognition results, and improves the reliability and inference efficiency of risk scoring and rating results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491938A_ABST
    Figure CN122491938A_ABST
Patent Text Reader

Abstract

This invention discloses a machine learning-based method and system for diagnosing enterprise technology risks, relating to the field of enterprise risk management technology. The method includes: analyzing the changing trends, information contribution, and redundancy of an enterprise technology dataset; screening and fusing effective features; completing standardized mapping according to preset matrix rules to generate a comprehensive feature matrix; performing machine learning analysis on the comprehensive feature matrix; identifying candidate risk patterns using random forests; and performing consistency verification, conflict detection, and reliability screening on these candidate risk patterns to generate risk pattern recognition results. Finally, risk inference is performed based on the risk pattern recognition results to generate a risk score set and risk level results. This invention improves the accuracy of risk pattern recognition and the reliability of risk scores and risk level results by utilizing random forests to identify risk patterns and combining consistency verification, conflict detection, and pattern reliability analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise risk management technology, and in particular to an enterprise technology risk diagnosis method and system based on machine learning. Background Technology

[0002] As enterprises increasingly digitize their R&D activities, the scale of internal R&D data, industry technical intelligence, and publicly available patent information continues to grow. Data-driven technology risk identification methods are gradually being applied to enterprise R&D management, technology decision-making, and risk assessment. Currently, some enterprises are utilizing machine learning methods to perform correlation analysis on R&D progress, test results, and market feedback to improve the efficiency of technology risk identification and assist enterprises in optimizing technology pathways and providing early warnings of risks.

[0003] Existing technologies in enterprise technology risk analysis often employ static feature analysis or single-model identification methods, which make it difficult to conduct dynamic correlation analysis of continuously changing enterprise R&D data and external technical intelligence. This results in risk pattern identification results being easily affected by data fluctuations and feature redundancy, thereby reducing the accuracy and stability of risk inference results and making it difficult to form a continuously updated dynamic diagnostic capability for enterprise technology risks. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a machine learning-based enterprise technology risk diagnosis method to address the problem of insufficient dynamic identification capability of risk patterns in the process of enterprise technology risk diagnosis.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: Firstly, this invention provides a machine learning-based method for diagnosing enterprise technology risks. The method includes: collecting internal R&D data and external technical intelligence, cleaning and structuring them to generate an enterprise technology foundation dataset; analyzing the trend, information contribution, and redundancy of the enterprise technology foundation dataset, screening and fusing effective features, and completing standardized mapping according to preset matrix rules to generate a comprehensive feature matrix; performing machine learning analysis on the comprehensive feature matrix, identifying candidate risk patterns through random forests, and verifying consistency, detecting conflicts, and screening reliability of the candidate risk patterns to generate risk pattern recognition results; then performing risk inference based on the risk pattern recognition results to generate a risk score set and risk level results; performing correlation analysis on the risk score set and risk level results, and using interpretability methods to analyze the risk impact and risk transmission relationships between the comprehensive feature matrix information to generate a visualized risk report; formulating risk mitigation measures based on the visualized risk report and the comprehensive feature matrix information, optimizing the adjustment of the comprehensive feature matrix information based on the risk mitigation measures, and updating the enterprise technology foundation dataset using the adjusted comprehensive feature matrix information to form a risk management closed loop.

[0007] As a preferred embodiment of the enterprise technology risk diagnosis method based on machine learning described in this invention, the internal R&D data of the enterprise includes enterprise R&D plans, project progress, test results, and R&D documents. The external technical intelligence includes publicly available patent information, industry reports, and market feedback.

[0008] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the generation of the enterprise technology basic dataset specifically includes, The system standardizes the format of internal R&D data and external technical intelligence, fills in missing values ​​and removes outliers, and processes them uniformly to generate structured datasets. Features related to R&D progress, project indicators, and external intelligence are extracted from structured datasets for derivative calculations, feature enhancement, and redundancy filtering to generate an enhanced feature set. Multidimensional consistency verification and anomaly removal are performed on the enhanced feature set to generate an enterprise technology foundation dataset.

[0009] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the generation of comprehensive feature matrix information specifically includes: Perform feature change analysis, feature contribution analysis, and feature correlation analysis on the enterprise's technical foundation dataset, and filter out effective features to generate a preliminary feature set; The effective features in the initial feature set are subjected to association fusion and derivation enhancement processing to generate a fused feature set; The fused feature set is summarized and formatted, and feature mapping is completed according to preset matrix rules to generate comprehensive feature matrix information.

[0010] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the generation of risk pattern recognition results specifically includes: Important data is filtered and abnormal data is removed from the comprehensive feature matrix information. The feature evaluation weights are dynamically adjusted according to the feature change rate, fluctuation amplitude and change stability within different time windows. The data change trend is analyzed and a preliminary risk feature set is generated. The random forest method is used to identify and fuse risk patterns in the initial risk feature set, and simultaneously match the comprehensive feature matrix information to generate a candidate risk pattern set. Consistency verification, conflict detection, and pattern reliability analysis are performed on the candidate risk pattern set to eliminate unstable and conflicting patterns and generate risk pattern recognition results.

[0011] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the generation of the risk score set and risk level results specifically includes: Based on the correlation between the risk pattern recognition results and the comprehensive feature matrix information, the risk probability is calculated using the comprehensive feature matrix information to generate a preliminary risk score set; The preliminary risk score set is normalized and analyzed, and the degree of data correlation and risk deviation are calculated by combining the comprehensive feature matrix information to generate a risk score set. The risk scores are divided into corresponding risk levels based on the risk score set, and the abnormal changes in the risk score set are analyzed to dynamically correct the risk score set and generate risk level results.

[0012] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the generation of the visualized risk report specifically includes: Based on the risk score set and risk level results, a correlation mapping is performed to analyze the risk impact relationship between the comprehensive feature matrix information and generate a risk impact factor mapping set. The correlation between the risk impact factor mapping set and the comprehensive feature matrix information is analyzed to determine the risk transmission relationship between the comprehensive feature matrix information and generate correlation analysis results. The correlation analysis results are visualized, and the risk correlation and trend changes are mapped using the risk score set and risk level results to generate a visualized risk report, which includes a risk hotspot map, a trend analysis map, and a correlation path diagram.

[0013] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the formulation of risk mitigation measures specifically includes, Based on the visualized risk report, correlation analysis results, and comprehensive feature matrix information, correlation analysis is performed and historical risk handling records are matched to generate a set of candidate risk mitigation measures. Based on the candidate risk mitigation measure set, predictive analysis is performed on the degree of change in the risk score set and the degree of adjustment of the comprehensive feature matrix information. Risk handling methods that meet the change conditions of the risk score set and the adjustment of the comprehensive feature matrix information meet the stability conditions are selected, and risk mitigation measures are generated.

[0014] As a preferred embodiment of the machine learning-based enterprise technology risk diagnosis method of the present invention, the formation of a risk management closed loop specifically includes: Based on the risk mitigation measures, the comprehensive feature matrix information is adjusted, and the adjusted comprehensive feature matrix information is written into the corresponding risk processing record; The enterprise's technical infrastructure dataset is updated based on business data following the implementation of risk mitigation measures, and the comprehensive feature matrix information is reconstructed. The updated enterprise technology infrastructure dataset is continuously monitored and dynamically analyzed, and the comprehensive feature matrix information is adjusted based on the risk score set and risk level results to form a risk management closed loop.

[0015] Secondly, this invention provides a machine learning-based enterprise technology risk diagnosis system, comprising: a data fusion module for collecting internal R&D data and external technical intelligence, cleaning and structuring them to generate an enterprise technology foundation dataset; a feature analysis module for analyzing the enterprise technology foundation dataset for trend analysis, information contribution, and correlation redundancy, filtering and fusing effective features, and completing standardized mapping according to preset matrix rules to generate comprehensive feature matrix information; a risk reasoning module for performing machine learning analysis on the comprehensive feature matrix information, identifying candidate risk patterns through random forests, and performing consistency verification, conflict detection, and reliability screening on the candidate risk patterns to generate risk pattern recognition results, and then performing risk reasoning based on the risk pattern recognition results to generate a risk score set and risk level results; an intelligent analysis module for performing correlation analysis on the risk score set and risk level results, and analyzing the risk impact relationship and risk transmission relationship between the comprehensive feature matrix information through interpretability methods to generate a visualized risk report; and a risk management module for formulating risk mitigation measures based on the visualized risk report and comprehensive feature matrix information, optimizing the adjustment of the comprehensive feature matrix information based on the risk mitigation measures, and updating the enterprise technology foundation dataset using the adjusted comprehensive feature matrix information to form a risk management closed loop.

[0016] The beneficial effects of this invention are as follows: By using the random forest method to identify and fuse risk patterns in the initial risk feature set, and simultaneously matching comprehensive feature matrix information, it is possible to continuously perform correlation analysis on multidimensional data changes in enterprise R&D activities, thereby improving the dynamic adaptability of the risk pattern identification process to data changes; by performing consistency verification, conflict detection, and pattern reliability analysis on the candidate risk pattern set, the stability and continuity of the risk pattern identification results can be enhanced, making the risk pattern identification results more accurately reflect the risk change status in enterprise technology activities, and improving the reliability of risk scoring and risk level results and the efficiency of risk reasoning. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a machine learning-based enterprise technology risk diagnosis method.

[0019] Figure 2 This is a schematic diagram of a machine learning-based enterprise technology risk diagnosis system.

[0020] Figure 3 A flowchart generated for the enterprise's technical infrastructure dataset.

[0021] Figure 4 This is a flowchart for the closed loop of risk analysis and risk management. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a machine learning-based enterprise technology risk diagnosis method, including the following steps: S1. Collect internal R&D data and external technical intelligence from enterprises, and clean and structure them to generate enterprise technical basic datasets.

[0026] S1.1 Internal R&D data of an enterprise includes enterprise R&D plans, project progress, test results and R&D documents.

[0027] Specifically, enterprise R&D plans are collected from R&D plan tables, task breakdown records, and resource allocation records. R&D objectives, task nodes, and resource allocation parameters are obtained through data reading and record extraction to generate R&D plan characteristic information. Project progress is collected from project execution records, milestone records, and task status records. Task completion rate, stage delays, and milestone achievement status are obtained through periodic synchronization to generate project progress characteristic information. Test results are collected from test reports, defect records, and verification records. The number of defects, defect repair cycle, and test pass rate are obtained through result parsing to generate test result characteristic information. R&D documents are collected from technical solution documents, R&D process records, and archived results. Technical solution change records, technical difficulty descriptions, and R&D result descriptions are obtained through text parsing to generate R&D document characteristic information. Based on project identifiers and time identifiers, the R&D plan characteristic information, project progress characteristic information, test result characteristic information, and R&D document characteristic information are correlated and matched to generate internal enterprise R&D data.

[0028] S1.2 External technical intelligence includes publicly available patent information, industry reports, and market feedback.

[0029] Specifically, publicly available patent information is collected from patent publications, patent application records, and patent legal status records. Field extraction methods are used to obtain technical field classifications, patent layouts, and technological development directions, generating publicly available patent characteristic information. Industry reports are collected from industry research materials, technology development analysis materials, and industry statistics. Word segmentation, topic identification, and semantic analysis are performed on the industry report text to extract information on industry development trends, technological maturity, and competitive landscape, generating industry report characteristic information. Market feedback is collected from product evaluation records, customer feedback records, and market research records. Word segmentation, entity identification, and semantic analysis are performed on the market feedback text to extract information on changes in technological needs, product usage issues, and market focus areas, generating market feedback characteristic information. Based on technical field identifiers and time identifiers, publicly available patent characteristic information, industry report characteristic information, and market feedback characteristic information are correlated, matched, and uniformly organized to generate external technical intelligence.

[0030] S1.3 Standardize the format of internal R&D data and external technical intelligence, fill in missing values ​​and remove outliers, and process them uniformly to generate a structured dataset.

[0031] Specifically, the process involves converting and unifying the field formats and encoding of internal R&D data and external technical intelligence. Time information is converted into a unified time format, and text information is converted into a unified field format to generate standardized data. Based on the field relationships in the standardized data, missing content is filled in. Continuous numerical fields are filled with the average of nearby time data, and discrete category fields are filled with high-frequency categories to generate complete data. The degree of abnormal deviation is calculated based on the statistical distribution and field relationships of the complete data, and abnormal records exceeding the preset threshold are removed to generate clean data. The clean data is then uniformly collected and sorted according to project identifiers, technical field identifiers, and time identifiers to generate a structured dataset.

[0032] S1.4 Extract features related to R&D progress, project indicators, and external intelligence from the structured dataset, perform derivative calculations, feature enhancement, and redundancy filtering, and generate an enhanced feature set.

[0033] Specifically, features related to R&D progress, project indicators, and external intelligence are extracted from structured datasets. Based on time, project, and technology dimensions, the rate of change in R&D progress, the deviation of project indicators, and the matching degree of external intelligence are calculated to generate derived features. These derived features are then cross-combined and trend-quantified to construct enhanced features reflecting the relationship between R&D activities and the external environment, generating a candidate feature set. Correlation analysis, information gain analysis, and variance contribution analysis are performed on the candidate feature set to obtain correlation coefficients, information gain values, and variance contribution degrees, respectively. The degree of feature redundancy is calculated based on the correlation coefficient, the information contribution degree of the features is judged based on the information gain value, and the explanatory power of the features for data changes is evaluated based on the variance contribution degree. Finally, features are screened based on correlation coefficient thresholds, information gain thresholds, and variance contribution degree thresholds to generate an enhanced feature set.

[0034] It should also be noted that the correlation coefficient threshold is determined based on the degree of information redundancy among candidate feature sets, the information gain threshold is determined based on the degree of information contribution of candidate feature sets, and the variance contribution threshold is determined based on the explanatory power of candidate feature sets for data changes; for example, the correlation coefficient threshold can be set to 0.80 to 0.95, the information gain threshold can be set to 0.005 to 0.05, and the variance contribution threshold can be set to 0.5% to 5%.

[0035] S1.5 Perform multidimensional consistency verification and abnormal association removal on the enhanced feature set to generate the enterprise technology foundation dataset.

[0036] Specifically, the R&D progress-related features, project indicator-related features, and external intelligence-related features in the enhanced feature set are matched according to time, project, and technology dimensions. After normalizing the features in each dimension, the degree of difference in feature values ​​is calculated, and the Pearson correlation coefficient is used to calculate the degree of consistency between features. Based on the distribution results of the degree of difference and the degree of consistency, feature records deviating from the normal distribution range are identified and removed, retaining features that meet the consistency requirements to generate a consistent feature set. Feature association relationships are constructed based on the consistent feature set, and the mutual information method is used to calculate the association strength between features. The sliding time window statistical method is used to calculate the direction and magnitude of change. Based on the distribution results of association strength and the distribution results of change trends, abnormal association features are identified and removed to generate a cleaned feature set. The feature data in the cleaned feature set are uniformly sorted, collected, and encoded according to project identifier, technology identifier, and time identifier. The cleaned feature set is converted into a unified matrix structure to generate the enterprise technology foundation dataset.

[0037] It should also be noted that the consistency requirement indicates the criteria for determining whether features maintain a stable correlation and a reasonable range of variation across time, project, and technical dimensions. The normal distribution interval is determined based on the mean and standard deviation of the distribution results of the degree of difference and the degree of correlation consistency. For example, the mean ± 2 times the standard deviation can be used as the normal distribution interval. Feature records that exceed the normal distribution interval are judged as abnormal feature records.

[0038] S2. Analyze the changing trends, information contribution, and correlation redundancy of the enterprise's technical foundation dataset, screen and integrate effective features, complete the standardized mapping according to the preset matrix rules, and generate comprehensive feature matrix information.

[0039] S2.1 Perform feature change analysis, feature contribution analysis, and feature correlation analysis on the enterprise technology foundation dataset, and select effective features to generate a preliminary feature set.

[0040] Specifically, time-series analysis is performed on the R&D progress-related features, project indicator-related features, and external intelligence-related features in the enterprise technology foundation dataset. A sliding time window is used to calculate the feature change rate and stability, generating trend features. Based on the correlation between trend features and the enterprise technology foundation dataset, the information gain method is used to calculate the information contribution of each feature, generating contribution features. Based on the correlation between contribution features and features in the enterprise technology foundation dataset, the Pearson correlation coefficient is used to calculate the redundancy between features, generating redundancy features. The weights of trend features are calculated based on the feature change rate within the sliding time window, the weights of contribution features are calculated based on the information gain value, and the weights of redundancy features are calculated based on the degree of redundancy between features. A comprehensive evaluation index is constructed by combining the weights of trend features, contribution features, and redundancy features, where the redundancy feature weight is negatively correlated with the comprehensive evaluation index. Effective features are then selected based on the comprehensive evaluation index, generating a preliminary feature set.

[0041] S2.2. Perform association fusion and derivation enhancement processing on the effective features in the preliminary feature set to generate a fused feature set.

[0042] Specifically, cross-combining the R&D progress features, project indicator features, and external intelligence features in the preliminary feature set to construct cross-dimensional correlation features; calculating the time change rate, deviation degree, and matching degree based on the correlation features, and statistically analyzing the impact of feature changes on historical risk level changes, combining the time change rate, deviation degree, matching degree, and risk impact degree to form derived features; calculating the comprehensive feature weight based on the derived features and correlation features, and using the comprehensive feature weight for fusion processing to generate a fused feature set.

[0043] S2.3. Summarize and format the fused feature set as a whole, and complete the feature mapping according to the preset matrix rules to generate comprehensive feature matrix information.

[0044] Specifically, the features in the fusion feature set are uniformly sorted and aggregated according to project identifier, technical field identifier, and time identifier to generate feature summary results; based on the feature summary results, features of different dimensions are normalized and uniformly encoded to generate standard feature results; matrix row and column indexes are constructed according to project identifier, technical field identifier, time window, and feature field, and feature values ​​are mapped to corresponding matrix positions to generate comprehensive feature matrix information.

[0045] It should also be noted that the matrix row and column indexes are jointly determined by the project identifier, technical field identifier, time window, and feature fields. The row index is used to identify the feature records of different projects under the corresponding time window, and the column index is used to identify different feature fields.

[0046] S3. Perform machine learning analysis on the comprehensive feature matrix information, identify candidate risk patterns through random forest, and perform consistency verification, conflict detection and reliability screening on the candidate risk patterns to generate risk pattern recognition results. Then, perform risk inference based on the risk pattern recognition results to generate risk score sets and risk level results.

[0047] S3.1. Perform important data screening and abnormal data removal on the comprehensive feature matrix information, and dynamically adjust the feature evaluation weights according to the feature change rate, fluctuation amplitude and change stability within different time windows, analyze the data change trend, and generate a preliminary risk feature set.

[0048] Specifically, statistical analysis is performed on each feature in the comprehensive feature matrix information to calculate the dispersion and distribution of each feature in different samples, obtaining feature statistical results. Based on the feature statistical results, the information gain value between each feature and the risk label is calculated, and the variance contribution of each feature to the overall change of the comprehensive feature matrix information is calculated. The information gain value and variance contribution are standardized to eliminate dimensional differences, and a comprehensive evaluation value is calculated based on the standardized information gain value and variance contribution. The features are ranked according to the comprehensive evaluation value to form feature evaluation results. The feature importance of each feature is calculated based on the feature evaluation results, and features with an importance value higher than the feature importance threshold are retained to generate important feature results. Based on the important feature results, the mean, standard deviation, and range of change of each feature in the historical time series are statistically analyzed to calculate the degree of abnormal deviation, and feature data with an abnormal deviation degree exceeding the abnormal deviation threshold are removed to generate cleaned feature results. The evaluation weights of each feature in the cleaned feature results are dynamically adjusted according to the feature change rate, fluctuation amplitude, and change stability within different time windows, and time series analysis is performed on the cleaned feature results to extract continuous growth features, continuous decline features, and abnormal fluctuation features, generating a preliminary risk feature set.

[0049] It should also be noted that the feature importance threshold is determined based on the degree of contribution of each feature to the information of risk labeling and its ability to explain the overall changes in the comprehensive feature matrix information. For example, the feature importance threshold can be a normalized evaluation value of 0.60 to 0.85. Risk labels are determined based on the occurrence of risks in historical risk event records and combined with historical risk assessment results to form corresponding risk category labels, where the historical risk assessment results are derived from historical sample data.

[0050] S3.2. Using the random forest method, risk patterns in the initial risk feature set are identified and fused, and the comprehensive feature matrix information is matched simultaneously to generate a candidate risk pattern set.

[0051] Specifically, the preliminary risk feature set is processed using a random forest method. The preliminary risk feature set is input into multiple decision trees within the random forest, which classify and judge the combinations of risk features in the preliminary risk feature set, obtaining the risk category and risk probability results output by each decision tree. The risk category results are voted on and statistically analyzed, and the risk probability results are averaged to determine the risk pattern category and risk occurrence probability corresponding to the preliminary risk feature set, generating multiple risk identification results. These multiple risk identification results are then fused through voting and confidence calculations to select risk feature combinations with high consistency, generating preliminary risk patterns. Based on the comprehensive feature matrix information, the feature states and correlations corresponding to the preliminary risk patterns are extracted, and the preliminary risk patterns are matched with the comprehensive feature matrix information to obtain risk pattern matching results. Based on the risk pattern matching results, the corresponding risk pattern category, risk occurrence probability, feature state, and correlations are extracted, aggregated, and organized to generate a candidate risk pattern set. Here, the multiple risk identification results represent the set of risk identification results output by the multiple decision trees in the random forest after classifying and predicting the preliminary risk feature set, used for subsequent voting fusion and risk pattern determination.

[0052] It should also be noted that high consistency means that the risk category results output by multiple decision trees have a high degree of consistency. For example, the number of decision trees that output the same risk category results accounts for 70% to 95% of the total number of decision trees.

[0053] S3.3 Perform consistency verification, conflict detection, and pattern reliability analysis on the candidate risk pattern set, screen out unstable and conflicting patterns, and generate risk pattern recognition results.

[0054] Specifically, cross-validation is performed on various risk patterns in the candidate risk pattern set to calculate the consistency of matching results of different risk patterns in multiple samples, generating consistency analysis results. Based on the consistency analysis results, the risk level, risk change direction, and risk occurrence conditions corresponding to each risk pattern are compared to identify risk patterns with contradictory judgment relationships, generating conflict detection results. Combining the conflict detection results, the frequency of occurrence, identification accuracy, and continuous stability of each risk pattern in historical data are statistically analyzed to calculate the pattern reliability evaluation value, generating pattern reliability analysis results. Based on the pattern reliability analysis results, unstable patterns with pattern reliability evaluation values ​​below the reliability threshold are removed, and conflict patterns with contradictory relationships in the conflict detection results are also removed. Risk patterns that meet the consistency requirements and whose pattern reliability evaluation values ​​reach the reliability threshold are retained, generating risk pattern identification results.

[0055] Among them, the consistency of matching results is used to characterize the consistency of the identification results of the same risk pattern in different samples; the pattern reliability evaluation value is used to characterize the credibility of the risk pattern, which is calculated by weighting the frequency of occurrence, historical matching degree and continuous stability, where the historical matching degree represents the degree of conformity between the risk pattern identification result and the historical risk event record; the reliability threshold is determined based on the historical verification results, for example, it can be set to 0.70 to 0.90.

[0056] S3.4. Based on the correlation between the risk pattern recognition results and the comprehensive feature matrix information, calculate the risk probability of the comprehensive feature matrix information to generate a preliminary risk score set.

[0057] Specifically, based on the risk pattern recognition results, the risk occurrence conditions, risk impact weights, and risk change trends corresponding to each risk pattern are extracted and matched with the corresponding features in the comprehensive feature matrix information to generate risk association results. Based on the risk association results, the degree of matching of each feature in the comprehensive feature matrix information to meet the risk occurrence conditions is calculated to obtain the feature risk probability. The degree of matching indicates the degree of conformity between each feature in the comprehensive feature matrix information and the corresponding feature state and association relationship of the risk pattern. When the corresponding feature state and association relationship are consistent with the feature combination corresponding to the risk pattern recognition results, the risk occurrence conditions are determined to be met. The feature risk probability is weighted and calculated based on the risk impact weights to obtain the risk probability value corresponding to each risk pattern. The risk probability values ​​corresponding to each risk pattern are summarized to form a preliminary risk score set corresponding to the comprehensive feature matrix information.

[0058] S3.5. Perform normalization analysis on the preliminary risk score set, and calculate the degree of data correlation and risk deviation by combining the comprehensive feature matrix information to generate a risk score set.

[0059] Specifically, the probability values ​​of each risk in the preliminary risk score set are normalized to their minimum values, transforming them into a unified scoring range to generate standard risk scores. Based on the standard risk scores and the comprehensive feature matrix information, the Pearson correlation coefficient between each feature is calculated. The formula for calculating the correlation coefficient is as follows: ; in, Representation of features With features The correlation coefficient between them Represents the first element in the comprehensive feature matrix information. Item features, Represents the first element in the comprehensive feature matrix information. Item features, Representation of features In the The values ​​in each sample Representation of features In the The values ​​in each sample Representation of features The average value across all samples. Representation of features The average value across all samples. Indicates the sample number. This indicates the total number of samples involved in the calculation.

[0060] The correlation strength is calculated based on the absolute value of the correlation coefficient. The formula for calculating the correlation strength is as follows: ; in, Representation of features With features The strength of the association between them is used to determine the degree of data association.

[0061] The analysis examines the differences between the standard risk score results and the historical variation patterns of the comprehensive feature matrix information based on the degree of data correlation. It also examines the differences between the standard risk score results and the historical risk score distribution based on the degree of data correlation. Using the historical risk score mean as a reference benchmark, the analysis calculates the degree of deviation of each risk score from the historical distribution center by combining the historical risk score standard deviation. When the degree of deviation exceeds a preset deviation threshold, it is judged as an abnormal deviation, and a risk deviation result is generated. The risk score set is then generated by combining the standard risk score results, the degree of data correlation, and the risk deviation result with weighted correction.

[0062] It should also be noted that the deviation threshold is determined based on the distribution of historical risk scores, and the deviation range corresponding to the mean of historical risk scores and 2 to 3 times the standard deviation can be used as the basis for judgment.

[0063] S3.6. Divide the risk level according to the risk score set, analyze the abnormal change data in the risk score set, dynamically correct the risk score set, and generate the risk level result.

[0064] Specifically, the risk score values ​​in the risk score set are matched with risk level intervals to divide the risk score values ​​into different risk levels, generating preliminary risk level results. Based on the preliminary risk level results, the risk score value change sequence within the corresponding time window is extracted, and the magnitude and rate of change between adjacent time nodes are calculated. Abnormal change data with magnitudes exceeding the abnormal change threshold are identified, generating abnormal change analysis results. Based on the abnormal change analysis results, the magnitude, rate, and duration of the abnormal change data are extracted, and the influence weight of the abnormal change data on the risk score value is calculated by combining the feature changes in the comprehensive feature matrix information. Based on the influence weight, the corresponding risk score values ​​in the risk score set are dynamically weighted and corrected, and the risk score values ​​are recalculated to generate corrected risk score results. Based on the corrected risk score results, risk level interval matching is performed again to generate the final risk level result.

[0065] It should also be noted that different risk levels represent risk status categories based on risk score values, used to characterize the severity of a company's technological risks, such as low risk, medium risk, high risk, and extremely high risk levels; The threshold for abnormal changes is determined based on the statistical distribution of the historical risk score changes. For example, the average historical change plus 2 to 3 times the standard deviation can be used as the judgment boundary, and the corresponding range of abnormal change thresholds can be set to 1.5 to 3 times the historical average change.

[0066] S4. Conduct correlation analysis on the risk score set and risk level results, and use interpretability methods to analyze the risk impact relationship and risk transmission relationship between the comprehensive feature matrix information to generate a visualized risk report.

[0067] S4.1. Based on the risk score set and risk level results, perform correlation mapping, analyze the risk impact relationship between the comprehensive feature matrix information, and generate a risk impact factor mapping set.

[0068] Specifically, the risk score values ​​and risk level results in the risk score set are correlated to determine the high-risk and low-risk characteristics corresponding to different risk levels, generating risk level correlation results. Based on the risk level correlation results, the correlation relationships between corresponding features in the comprehensive feature matrix information are extracted, and the influence of each feature on the risk score value is calculated in conjunction with changes in the risk score value, generating risk impact analysis results. Based on the risk impact analysis results, the correlation relationships between influence paths and influence weights between features are established, and the risk propagation path is constructed based on the chronological relationship of changes in the correlated features in the historical time series, determining the direction of risk propagation and the intensity of risk effect, generating risk relationship mapping results. The influencing features, affected features, influence weights, and risk propagation directions in the risk relationship mapping results are aggregated to generate a risk impact factor mapping set.

[0069] It should also be noted that the risk intensity represents the degree to which the influencing feature transmits risk to the affected feature along the influence path. It can be quantitatively characterized by influence weight. The larger the influence weight, the higher the risk intensity, and the smaller the influence weight, the lower the risk intensity. For example, the risk intensity can be a value between 0 and 1.

[0070] S4.2 Analyze the correlation between the risk impact factor mapping set and the comprehensive feature matrix information, analyze the risk transmission relationship between the comprehensive feature matrix information, and generate correlation analysis results.

[0071] Specifically, the process involves matching the influencing features, affected features, influence weights, and risk propagation directions in the risk impact factor mapping set with the corresponding features in the comprehensive feature matrix information to generate feature correlation results. Based on the correlation path, correlation strength, and correlation level information in the feature correlation results, the number of correlation nodes traversed between the influencing and affected features is counted to obtain the influence path length. The influence strength is calculated based on the correlation strength between the influencing and affected features. The transmission range is calculated based on the number of correlation features that the influencing feature can directly or indirectly influence. Risk transmission analysis results are generated based on the influence path length, influence strength, and transmission range. Based on the risk transmission analysis results, transmission nodes and risk propagation paths in the risk propagation chain are identified, and the risk transmission direction and risk diffusion relationship between the comprehensive feature matrix information are determined to generate risk relationship analysis results. Finally, the risk transmission direction, risk diffusion relationship, and transmission nodes in the risk relationship analysis results are aggregated to generate correlation analysis results.

[0072] It should also be noted that the influence path length represents the number of transmission levels through which the influencing feature acts on the affected feature through risk transmission relationships. It can be characterized by the number of associated nodes between the influencing feature and the affected feature. The more associated nodes, the longer the influence path length.

[0073] S4.3 Visualize the correlation analysis results and use the risk score set and risk level results to map the risk correlation and trend changes to generate a visualized risk report. The visualized risk report includes a risk hotspot map, a trend analysis map, and a correlation path diagram.

[0074] Specifically, based on the correlation analysis results, the risk transmission direction, risk diffusion relationship, and transmission nodes are extracted and visualized according to the impact intensity to generate risk correlation display results; based on the risk score values ​​in the risk score set and the corresponding features in the comprehensive feature matrix information, position matching and color-coded display are performed to form risk hotspot distribution information, and a risk hotspot map is generated based on the risk hotspot distribution information; based on the historical risk score value change sequence in the risk score set and the level change in the risk level results, a time axis mapping is performed to form risk change trend information, and a trend analysis chart is generated based on the risk change trend information; combined with the risk transmission direction, risk diffusion relationship, and transmission nodes in the risk correlation display results, a risk propagation path is constructed to form risk path information, and a correlation path diagram is generated based on the risk path information; the risk hotspot map, trend analysis chart, and correlation path diagram are collected and displayed to generate a visualized risk report.

[0075] S5. Based on the visualized risk report and comprehensive feature matrix information, formulate risk mitigation measures, optimize the adjustment of the comprehensive feature matrix information according to the risk mitigation measures, and update the enterprise's technical foundation dataset using the adjusted comprehensive feature matrix information to form a risk management closed loop.

[0076] S5.1. Based on the visualized risk report, correlation analysis results and comprehensive feature matrix information, perform correlation analysis and match historical risk handling records to generate a set of candidate risk mitigation measures.

[0077] Specifically, based on the visualized risk report, high-risk areas are extracted from the risk hotspot map, risk change trends from the trend analysis map, and risk propagation paths from the correlation path diagram. Correlation analysis and aggregation of the characteristics corresponding to high-risk areas, risk change trends, and risk propagation paths are then performed to generate risk feature analysis results. These results are then matched with the risk transmission direction, risk diffusion relationship, and transmission nodes from the correlation analysis results, and further correlated with the feature correlation relationships in the comprehensive feature matrix information to generate risk management requirement results. Based on these requirements, corresponding risk types, risk impact ranges, and risk formation causes are extracted and matched with the risk scenarios, management measures, and management effects in historical risk management records to generate risk management matching results. Based on the risk management matching results, risk management measures that meet the required management effects are selected and aggregated according to risk type to generate a candidate risk mitigation measure set. The candidate risk mitigation measure set includes risk type, risk impact range, risk formation cause, risk management measures, management effect, adjustment parameters, and adjustment objectives.

[0078] It should also be noted that historical risk management records represent a collection of data on risk events, causes of risks, risk management measures taken, and corresponding effects that an enterprise has recorded in its past risk management processes. Risk management measures that meet the requirements are those that, after implementation, improve the corresponding risk score to a preset quantile threshold of the statistical distribution of historical management effects for similar risks, and maintain a stable or continuously improving risk score over multiple consecutive monitoring periods. The quantile threshold is determined based on the statistical distribution of historical management effects for similar risks. For example, the 60th, 70th, or 80th percentile of historical improvement effects can be used as the evaluation benchmark.

[0079] S5.2 Based on the candidate risk mitigation measure set, predict and analyze the degree of change in the risk score set and the degree of adjustment of the comprehensive feature matrix information, screen the risk handling methods that meet the change conditions of the risk score set and the adjustment of the comprehensive feature matrix information meets the stability conditions, and generate risk mitigation measures.

[0080] Specifically, the process involves extracting adjustment parameters and historical effects for each risk management measure from the candidate risk mitigation measure set, constructing adjustment schemes based on risk scores in the risk score set, and generating candidate adjustment results. Based on these candidate adjustment results, target features and corresponding adjustment parameters are determined, and corresponding adjustment schemes are extracted from the candidate risk mitigation measure set. The trend of feature changes after the implementation of risk mitigation measures is predicted based on historical effects and feature correlations, and an adjusted comprehensive feature matrix is ​​constructed based on the predicted trend. Risk pattern recognition and risk probability calculation are re-executed based on the adjusted comprehensive feature matrix to obtain adjusted risk scores. The magnitude of change in risk scores before and after adjustment is compared to generate risk score prediction results. The degree of change in the risk score set is calculated based on the risk score prediction results, and the magnitude of feature fluctuations and the degree of change in correlations are calculated based on the adjusted comprehensive feature matrix to generate adjustment stability analysis results. Finally, risk management methods that meet the change conditions for the risk score set and whose adjustments to the comprehensive feature matrix satisfy the stability conditions are selected based on the risk score prediction results and adjustment stability analysis results, thus generating risk mitigation measures.

[0081] It should also be noted that the change condition indicates the criteria for determining whether the risk score set achieves the expected improvement effect after the implementation of risk mitigation measures. For example, it can be set as the risk score value decreases by 20% to 80% and the risk level result decreases by at least one level. The stability condition indicates the criteria for determining whether the comprehensive feature matrix information maintains normal operation after adjustment. For example, it can be set as the feature fluctuation amplitude does not exceed 1.2 times the historical average fluctuation amplitude and the degree of change in correlation does not exceed 15%.

[0082] S5.3. Adjust the comprehensive feature matrix information according to the risk mitigation measures, and write the adjusted comprehensive feature matrix information into the corresponding risk processing record.

[0083] Specifically, based on the risk mitigation measures, the corresponding adjustment schemes and target objects are extracted, and the target features corresponding to the risk mitigation measures in the comprehensive feature matrix information are located to generate feature adjustment schemes. Based on the feature change patterns of similar risk scenarios in historical risk handling records and the correlation between target features, the target feature status after the implementation of risk mitigation measures is predicted to generate feature prediction results. Based on the feature prediction results, the feature status, feature weights, and feature correlations corresponding to the target features are updated, and the feature matrix structure is reconstructed to generate adjusted comprehensive feature matrix information. The adjusted comprehensive feature matrix information is correlated with the corresponding risk mitigation measures, risk score change results, and risk level change results and written into the risk handling records to generate risk handling correlation records.

[0084] S5.4 Update the enterprise technology foundation dataset based on business data after the implementation of risk mitigation measures, and reconstruct the comprehensive feature matrix information.

[0085] Specifically, based on the business data generated after the implementation of risk mitigation measures, corresponding R&D progress data, project indicator data, and external technical intelligence data are collected and matched with historical data in the enterprise's technical foundation dataset to generate data update results; based on the data update results, the corresponding records in the enterprise's technical foundation dataset are supplemented and updated to generate an updated enterprise technical foundation dataset; based on the updated enterprise technical foundation dataset, feature information is re-extracted, and a corresponding feature matrix is ​​constructed according to project identifier, technical field identifier, and time identifier to generate updated comprehensive feature matrix information.

[0086] S5.5 Continuously monitor and dynamically analyze the updated enterprise technology foundation dataset, and adjust the comprehensive feature matrix information based on the risk score set and risk level results to form a risk management closed loop.

[0087] Specifically, the updated enterprise technology infrastructure dataset is continuously updated according to the monitoring cycle to collect new business data and extract feature information to generate dynamic monitoring results; based on the dynamic monitoring results, the comprehensive feature matrix information is reconstructed and risk analysis is performed to generate risk feedback results; based on the risk feedback results, risk mitigation measures are optimized, and the enterprise technology infrastructure dataset is updated based on the new business data generated by the optimized risk mitigation measures. The process of dynamic monitoring, risk analysis, and risk mitigation measure optimization is repeated to form a risk management closed loop.

[0088] This embodiment also provides a machine learning-based enterprise technology risk diagnosis system, comprising: a data fusion module for collecting internal R&D data and external technical intelligence, cleaning and structuring them to generate an enterprise technology foundation dataset; a feature analysis module for analyzing the enterprise technology foundation dataset for trend analysis, information contribution, and correlation redundancy, filtering and fusing effective features, completing standardized mapping according to preset matrix rules, and generating comprehensive feature matrix information; a risk reasoning module for performing machine learning analysis on the comprehensive feature matrix information, identifying candidate risk patterns through random forests, and performing consistency verification, conflict detection, and reliability screening on the candidate risk patterns to generate risk pattern recognition results, and then performing risk reasoning based on the risk pattern recognition results to generate a risk score set and risk level results; an intelligent analysis module for performing correlation analysis on the risk score set and risk level results, and using interpretability methods to analyze the risk impact relationship and risk transmission relationship between the comprehensive feature matrix information to generate a visualized risk report; and a risk management module for formulating risk mitigation measures based on the visualized risk report and comprehensive feature matrix information, optimizing the adjustment of the comprehensive feature matrix information based on the risk mitigation measures, and updating the enterprise technology foundation dataset using the adjusted comprehensive feature matrix information to form a risk management closed loop.

[0089] In summary, this invention utilizes the random forest method to identify and fuse risk patterns in the initial risk feature set, and simultaneously matches comprehensive feature matrix information. This enables continuous correlation analysis of multidimensional data changes in enterprise R&D activities, improving the dynamic adaptability of the risk pattern identification process to data changes. Furthermore, by performing consistency verification, conflict detection, and pattern reliability analysis on the candidate risk pattern set, the stability and continuity of the risk pattern identification results are enhanced, allowing the results to more accurately reflect the risk changes in enterprise technology activities, and improving the reliability of risk scoring and risk level results and the efficiency of risk reasoning.

[0090] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A machine learning-based method for diagnosing enterprise technology risks, characterized in that: include, Collect internal R&D data and external technical intelligence from enterprises, and clean and structure them to generate enterprise technical foundation datasets; The data set of enterprise technology infrastructure is analyzed for its changing trends, information contribution, and correlation redundancy. Effective features are selected and integrated, and standardized mapping is completed according to preset matrix rules to generate comprehensive feature matrix information. Machine learning analysis is performed on the comprehensive feature matrix information. Candidate risk patterns are identified through random forest. Consistency verification, conflict detection, and reliability screening are performed on the candidate risk patterns to generate risk pattern recognition results. Risk inference is then performed based on the risk pattern recognition results to generate risk score sets and risk level results. A correlation analysis is performed on the risk score set and risk level results, and the risk impact relationship and risk transmission relationship between the comprehensive feature matrix information are analyzed through interpretability methods to generate a visualized risk report; Based on the visualized risk report and comprehensive feature matrix information, risk mitigation measures are formulated, and the comprehensive feature matrix information is adjusted according to the risk mitigation measures. The adjusted comprehensive feature matrix information is then used to update the enterprise's technical infrastructure dataset, forming a risk management closed loop.

2. The enterprise technology risk diagnosis method based on machine learning as described in claim 1, characterized in that: The internal R&D data of the enterprise includes the enterprise's R&D plan, project progress, test results, and R&D documents; The external technical intelligence includes publicly available patent information, industry reports, and market feedback.

3. The enterprise technology risk diagnosis method based on machine learning as described in claim 2, characterized in that: The generated enterprise technology foundation dataset specifically includes, The system standardizes the format of internal R&D data and external technical intelligence, fills in missing values ​​and removes outliers, and processes them uniformly to generate structured datasets. Features related to R&D progress, project indicators, and external intelligence are extracted from structured datasets for derivative calculations, feature enhancement, and redundancy filtering to generate an enhanced feature set. Multidimensional consistency verification and anomaly removal are performed on the enhanced feature set to generate an enterprise technology foundation dataset.

4. The enterprise technology risk diagnosis method based on machine learning as described in claim 3, characterized in that: The generated comprehensive feature matrix information specifically includes, Perform feature change analysis, feature contribution analysis, and feature correlation analysis on the enterprise's technical foundation dataset, and filter out effective features to generate a preliminary feature set; The effective features in the preliminary feature set are subjected to association fusion and derivation enhancement processing to generate a fused feature set; The fused feature set is summarized and formatted, and feature mapping is completed according to preset matrix rules to generate comprehensive feature matrix information.

5. The enterprise technology risk diagnosis method based on machine learning as described in claim 1, characterized in that: The generated risk pattern recognition results specifically include, Important data is filtered and abnormal data is removed from the comprehensive feature matrix information. The feature evaluation weights are dynamically adjusted according to the feature change rate, fluctuation amplitude and change stability within different time windows. The data change trend is analyzed and a preliminary risk feature set is generated. The random forest method is used to identify and fuse risk patterns in the initial risk feature set, and simultaneously match the comprehensive feature matrix information to generate a candidate risk pattern set. Consistency verification, conflict detection, and pattern reliability analysis are performed on the candidate risk pattern set to eliminate unstable and conflicting patterns and generate risk pattern recognition results.

6. The enterprise technology risk diagnosis method based on machine learning as described in claim 5, characterized in that: The generation of the risk score set and risk level results specifically includes, Based on the correlation between the risk pattern recognition results and the comprehensive feature matrix information, the risk probability is calculated using the comprehensive feature matrix information to generate a preliminary risk score set; The preliminary risk score set is normalized and analyzed, and the degree of data correlation and risk deviation are calculated by combining the comprehensive feature matrix information to generate a risk score set. The risk scores are divided into corresponding risk levels based on the risk score set, and the abnormal changes in the risk score set are analyzed to dynamically correct the risk score set and generate risk level results.

7. The enterprise technology risk diagnosis method based on machine learning as described in claim 6, characterized in that: The generation of the visualized risk report specifically includes, Based on the risk score set and risk level results, a correlation mapping is performed to analyze the risk impact relationship between the comprehensive feature matrix information and generate a risk impact factor mapping set. The correlation between the risk impact factor mapping set and the comprehensive feature matrix information is analyzed to determine the risk transmission relationship between the comprehensive feature matrix information and generate correlation analysis results. The correlation analysis results are visualized, and the risk correlation and trend changes are mapped using the risk score set and risk level results to generate a visualized risk report, which includes a risk hotspot map, a trend analysis map, and a correlation path diagram.

8. The enterprise technology risk diagnosis method based on machine learning as described in claim 1, characterized in that: The aforementioned risk mitigation measures specifically include, Based on the visualized risk report, correlation analysis results, and comprehensive feature matrix information, correlation analysis is performed and historical risk handling records are matched to generate a set of candidate risk mitigation measures. Based on the candidate risk mitigation measure set, predictive analysis is performed on the degree of change in the risk score set and the degree of adjustment of the comprehensive feature matrix information. Risk handling methods that meet the change conditions of the risk score set and the adjustment of the comprehensive feature matrix information meet the stability conditions are selected, and risk mitigation measures are generated.

9. The enterprise technology risk diagnosis method based on machine learning as described in claim 8, characterized in that: The formation of a closed-loop risk management system specifically includes, Based on the risk mitigation measures, the comprehensive feature matrix information is adjusted, and the adjusted comprehensive feature matrix information is written into the corresponding risk processing record; The enterprise's technical infrastructure dataset is updated based on business data following the implementation of risk mitigation measures, and the comprehensive feature matrix information is reconstructed. The updated enterprise technology infrastructure dataset is continuously monitored and dynamically analyzed, and the comprehensive feature matrix information is adjusted based on the risk score set and risk level results to form a risk management closed loop.

10. A machine learning-based enterprise technology risk diagnosis system, based on the machine learning-based enterprise technology risk diagnosis method according to any one of claims 1 to 9, characterized in that: include, The data fusion module is used to collect internal R&D data and external technical intelligence, and to clean and structure them to generate a basic dataset of enterprise technology. The feature analysis module is used to analyze the changing trends, information contribution, and correlation redundancy of the enterprise's technical foundation dataset, filter and integrate effective features, complete the standardized mapping according to the preset matrix rules, and generate comprehensive feature matrix information. The risk reasoning module is used to perform machine learning analysis on the comprehensive feature matrix information, identify candidate risk patterns through random forest, and perform consistency verification, conflict detection and reliability screening on the candidate risk patterns to generate risk pattern recognition results. Then, risk reasoning is performed based on the risk pattern recognition results to generate risk score sets and risk level results. The intelligent analysis module is used to perform correlation analysis on the risk score set and risk level results, and to analyze the risk impact relationship and risk transmission relationship between the comprehensive feature matrix information through interpretable methods, generating a visual risk report; The risk management module is used to formulate risk mitigation measures based on the visualized risk report and comprehensive feature matrix information, optimize the adjustment of the comprehensive feature matrix information based on the risk mitigation measures, and update the enterprise's technical infrastructure dataset using the adjusted comprehensive feature matrix information to form a risk management closed loop.