Risk assessment system based on machine learning and enterprise multi-source data fusion

By constructing a three-dimensional judgment system, dynamic repair, and a two-factor fusion algorithm, the problems of data acquisition adaptability and anomaly handling in enterprise multi-source data fusion were solved, achieving efficient risk assessment and trend adaptation, and improving data quality and the accuracy of risk assessment.

CN121961233APending Publication Date: 2026-05-01HANGZHOU ZHIKE FEICHUANG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU ZHIKE FEICHUANG INFORMATION TECH CO LTD
Filing Date
2026-01-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing machine learning and enterprise multi-source data fusion technologies suffer from problems such as insufficient data acquisition adaptability, rigid outlier handling, poor data fusion reliability, incomplete feature engineering, weak model generalization ability, and poor real-time deployment and maintenance in general enterprise scenarios, which cannot meet the trust requirements of enterprise decision-making.

Method used

By constructing a three-dimensional universal judgment system, designing a dynamic adjustment mechanism for three-dimensional repair thresholds, a two-factor fusion algorithm, and a trend-driven interpretable ensemble model, dynamic data acquisition, differentiated repair, cross-data source cross-validation, and closed-loop optimization are achieved, thereby improving the accuracy and reliability of data quality and risk assessment.

Benefits of technology

It significantly improves the targeting and effectiveness of data collection, fully preserves the core characteristics of risk-type change data, enhances the reliability and risk orientation of fused data, and realizes the system's dynamic adaptation to changes in risk trends and visual traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961233A_ABST
    Figure CN121961233A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of risk assessment, and discloses a risk assessment system based on machine learning and enterprise multi-source data fusion, and the system comprises the steps: building an association pair through a basic frequency and trigger enhancement dynamic adjustment mechanism, determining four types of risk data core rules, and completing the classification judgment; a three-dimensional repair threshold dynamic adjustment mechanism is adopted, differential repair is carried out on abnormal data, and jump risk association degree is calculated to carry out jump risk grading; performing cross-data-source cross check, retaining core features of risk type jump data, and calculating and screening the core features through feature-jump risk association density; a feature system of four types of risk data is perfected through feature extension and sample enhancement, a trend-driven interpretable integration model is constructed, a three-layer integration model is constructed, a two-dimensional evaluation result is output, a trend risk coefficient is fused, model deviation is corrected through closed-loop iteration optimization, and the dynamic risk evaluation precision of the system on various trend scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Risk assessment system based on machine learning and enterprise multi-source data fusion Technical Field

[0001] This invention relates to the field of risk assessment technology, specifically to a risk assessment system based on machine learning and the fusion of multi-source enterprise data. Background Technology

[0002] Addressing the core shortcomings of existing machine learning and enterprise multi-source data fusion technologies in general enterprise scenarios, existing technologies suffer from insufficient data acquisition adaptability, failing to match the temporal characteristics of enterprise data jumps, leading to the omission of key jump values ​​and contextual data; rigid handling of outliers and jump values, employing fixed rule-based filtering mechanisms that fail to distinguish between risk signal-type and noise-type anomalies, resulting in the loss of core risk signals and data distortion; a lack of coordination between enterprise multi-source data fusion and anomaly handling in existing technologies, inconsistent semantics of heterogeneous data, and the absence of anomaly cross-validation before fusion, resulting in poor reliability of fused data; incomplete feature engineering representation, neglecting jump temporal correlation and cross-source correlation features, and lacking a risk-oriented feature selection mechanism, leading to the curse of dimensionality and the loss of key risk features; weak model generalization ability and interpretability, insufficient adaptation to small-sample risk jumps, failure to address concept drift caused by scenario changes, and inability to meet the trust requirements of enterprise decision-making; poor real-time deployment and maintenance, lack of a closed-loop mechanism for evaluation, feedback, and optimization, and inability of the system to adapt to new risk jumps; therefore, a risk assessment system based on machine learning and enterprise multi-source data fusion is needed. Summary of the Invention

[0003] The purpose of this invention is to provide a risk assessment system based on machine learning and enterprise multi-source data fusion. To solve the aforementioned problems in the prior art, this invention achieves this through the following technical solution: Firstly, the risk assessment system based on machine learning and enterprise multi-source data fusion provided in this embodiment of the invention specifically includes the following modules: A data judgment module: Through a dynamic adjustment mechanism of basic frequency and trigger enhancement, and by constructing correlation pairs, a three-dimensional universal judgment system is established to clarify the core patterns of four types of risk data and complete the classification judgment; A repair quantification module: Based on the core patterns of the four types of risk data, a three-dimensional repair threshold dynamic adjustment mechanism is adopted to implement differentiated repair of abnormal data. The system is divided into several modules: First, it integrates and calculates the correlation degree of risk abrupt changes to quantify and classify these risks. Second, it combines the correlation degree of risk abrupt changes with cross-data source cross-validation, using a two-factor fusion algorithm to retain the core features of risk-type risk abrupt change data, extracting multi-dimensional risk correlation time-series features and cross-data source collaborative risk features, and selecting core features through feature-risk abrupt change correlation density calculation. Third, it uses an evaluation and calibration module to improve the feature system of the four types of risk data through feature expansion and sample enhancement, constructs a trend-driven interpretable ensemble model, builds a three-layer ensemble model, outputs two-dimensional evaluation results, incorporates trend risk coefficients, and corrects model bias through closed-loop iterative optimization.

[0004] Secondly, the risk assessment method based on machine learning and enterprise multi-source data fusion provided in this embodiment of the invention specifically includes the following steps: Step 1: Through a dynamic adjustment mechanism of basic frequency and trigger enhancement, and by constructing correlation pairs, a three-dimensional general judgment system is established to clarify the core patterns of four types of risk data and complete the classification judgment; Step 2: Based on the core patterns of the four types of risk data, a three-dimensional repair threshold dynamic adjustment mechanism is adopted to implement differentiated repair of abnormal data, and the correlation degree of jump risk is calculated to quantify and classify jump risk; Step 3: Cross-data source cross-validation is performed in combination with the correlation degree of jump risk. Using a two-factor fusion algorithm, the core features of risk-type jump data are retained, multi-dimensional risk correlation time-series features and cross-data source collaborative risk features are extracted, and core features are screened through feature-jump risk correlation density calculation; Step 4: The feature system of the four types of risk data is improved through feature expansion and sample enhancement, a trend-driven interpretable ensemble model is constructed, a three-layer ensemble model is constructed, a two-dimensional assessment result is output, a trend risk coefficient is incorporated, and the model bias is corrected through closed-loop iterative optimization.

[0005] The beneficial effects of this invention are as follows: 1. It breaks through the limitations of traditional fixed-frequency sampling and proposes a dynamic adjustment mechanism combining base frequency and trigger enhancement. Normal transition periods of enterprises are tagged and managed, and the sampling frequency is adjusted synchronously. This ensures complete capture of the entire abnormal transition process while effectively avoiding misjudging normal transitions as risk signals, significantly improving the targeting and effectiveness of data collection. Through precise timestamp association between core data and full-dimensional context data, a complete data support chain is provided for risk attribution. 2. It proposes a three-dimensional universal judgment system based on temporal characteristics, multi-source corroboration, and enterprise rules, systematically sorting out the core patterns of four types of risk data, achieving accurate classification of risk data, and adapting to various enterprise scenarios. 3. It designs a dynamic adjustment mechanism for amplitude-temporal-multi-source three-dimensional repair thresholds for different types of risk data, implementing differentiated repair strategies. Risk-type data retains its original characteristics and is labeled with risk tags, solving the problem that traditional fixed-rule filtering mechanisms easily delete valid risk data or retain invalid noise data, thus improving data quality and the accuracy of risk judgment. 1. Accuracy; 2. Design a two-level semantic mapping dictionary with a general semantic layer and an enterprise-customized layer to effectively solve the semantic differences of multi-source heterogeneous data of enterprises, realize the unified semantic transformation of data from different systems and of different types, introduce a general jump risk correlation degree to carry out cross-data source cross-validation, and combine a risk weight and data reliability dual-factor fusion algorithm to strengthen the proportion of high-risk and high-reliability data in the fusion, fully retain the core features of risk jump data, and improve the reliability and risk orientation of fused data; 3. Construct a trend-driven interpretable integration model to incorporate the trend features of risk data into the evaluation system, and capture key nodes in trend changes through a time-series attention mechanism; 4. Design a dual-dimensional fusion algorithm of jump risk and trend risk to calculate the final risk assessment value, simultaneously output the risk level and key driving factors, and generate a trend-risk correlation heat map to realize the visual traceability of risk causes, establish a closed-loop iterative optimization mechanism based on trend misjudgment rate, continuously optimize feature extraction rules and model weights, and realize the dynamic adaptation of the system to risk trend changes. Attached Figure Description

[0006] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0007] Figure 1 is a flowchart of the steps of the risk assessment system based on machine learning and enterprise multi-source data fusion provided in Embodiment 1 of the present invention; Figure 2 is a schematic diagram of the structure of the risk assessment system based on machine learning and enterprise multi-source data fusion provided in Embodiment 2 of the present invention. Detailed Implementation

[0008] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0009] Example 1: As shown in Figure 1, the risk assessment system based on machine learning and enterprise multi-source data fusion provided in this embodiment of the invention specifically includes the following modules: Data Judgment Module: Through a dynamic adjustment mechanism of base frequency and trigger enhancement, and by constructing correlation pairs, a three-dimensional general judgment system is established to clarify the core patterns of four types of risk data and complete the classification judgment; In a specific embodiment, combined with the jump time sequence characteristics of enterprise multi-source data, a dynamic adjustment mechanism of base frequency and trigger enhancement is constructed. By analyzing the enterprise's historical multi-source data, the duration and frequency characteristics of jumps in different types of multi-source data are statistically analyzed, and a base sampling frequency is set; If data fluctuations are detected to exceed the enterprise's normal baseline, the baseline is combined with the enterprise's historical normal operation data statistics, and the fluctuation threshold is set. Adapting to the enterprise's business characteristics, an enhanced sampling mode is automatically triggered. The enhanced sampling frequency is [5, 10] times the base frequency, and the enhanced collection duration is set according to the continuous characteristics of the enterprise's data jumps to ensure complete capture of the entire jump process. For known normal jump phases of the enterprise, such as periodic business peaks and regular operation and maintenance periods, they are marked as normal jump period labels, and the sampling frequency is adjusted synchronously to adapt to the jump characteristics to avoid misjudging normal jumps as risk signals. While collecting core data, the enterprise's full-dimensional related context data is collected simultaneously to construct core data-context data pairs. Core data includes, but is not limited to, enterprise business data, operation and maintenance data, operational data, and external industry data. Context data includes, but is not limited to, enterprise business operation status, The most recent maintenance records, enterprise business cycle stages, industry policy changes, and regional market environment data; enterprise business operation status including but not limited to: routine operation, peak operation, and anomaly handling; all contextual data are precisely linked to core data through high-precision timestamps to avoid viewing jump data in isolation and provide complete data support for risk attribution; an enterprise-level hierarchical desensitization mechanism is used to fully preserve jump characteristics while ensuring data compliance; sensitive enterprise information, including but not limited to: core business data identifiers, personnel information, and core financial data, is desensitized through character replacement and masking without changing the numerical characteristics and temporal sequence of the jump data; external public... Open data is marked for traceability, and internal data is protected by access control; all de-identification operations are logged synchronously to ensure that the entire data processing process is traceable and meets the enterprise's data security and compliance requirements; a three-dimensional general judgment system based on time-series characteristics, multi-source corroboration, and enterprise rules is constructed. First, the core patterns of four types of risk data are identified: jump values, outliers, continuous periodic outliers, and regular periodic outliers. Then, classification and judgment are completed to adapt to various enterprise scenarios. The core patterns of the four types of risk data are analyzed as follows: Jump value pattern: It shows a strong correlation between amplitude and duration. If the duration exceeds 50% of the enterprise scenario adaptation threshold, the risk probability increases, and it is accompanied by risk-type jumps with synchronous fluctuations of multi-source data or noise-type jumps with instantaneous fluctuations of a single data source.Outlier patterns: These are categorized into isolated single-source anomalies and multi-source collaborative anomalies, with the frequency of isolated anomalies negatively correlated with data source reliability. It should be noted that isolated single-source anomalies without multi-source corroboration are considered noise. Multi-source collaborative anomalies indicate synchronous anomalies across multiple data sources, representing a high-risk level. Continuous periodic outlier patterns: These exhibit periodic repetition and continuous accumulation, meaning outliers appear continuously within a fixed period, with the magnitude of the anomalies gradually increasing with each period. This is related to system overload and business process congestion. Regular periodic outlier patterns: These are strongly tied to the enterprise's business cycle, with the anomaly occurrence period overlapping with the business cycle by 80% or more. The anomaly magnitude conforms to the normal fluctuation range within the business cycle but exceeds the threshold for non-periodic periods, leading to misjudgment as a risk jump. Based on the core patterns of four types of risk data, the three-dimensional judgment system is optimized as follows: It combines the cyclical characteristic dimension to perform time-series characteristic analysis, calculating the duration, trend, and cyclical overlap of jumps / anomalies; where the cyclical overlap is the ratio of the overlap duration of the anomaly occurrence period plus the business cycle period to the business cycle duration; it strengthens cyclical synergy through multi-source corroboration verification, querying whether there are synchronous anomalies in related data sources within the same cycle; it uses enterprise rules to filter and supplement cyclical rules, combining the enterprise business cycle table to mark regular cyclical anomalies and continuous cyclical anomalies with specific tags, excluding regular anomalies within normal business cycles; if the duration of a jump exceeds the enterprise scenario adaptation threshold and shows a continuous fluctuation / upward trend, or the continuous cyclical overlap is greater than or equal to 0.6, or the regular... When the period overlap is greater than or equal to 0.8 and there is no supporting business rules, it is initially identified as a potential risk. The repair quantification module: based on the core patterns of the four types of risk data, a three-dimensional repair threshold dynamic adjustment mechanism is adopted to implement differentiated repair for abnormal data, integrate and calculate the correlation of jump risks, and quantify and classify jump risks. Based on the core patterns of the four types of risk data, a classification repair and period calibration strategy is designed to avoid the filter mechanism of fixed rules from mistakenly deleting valid risk data. Specifically, a three-dimensional repair threshold dynamic adjustment mechanism of amplitude, time series, and multi-source is adopted. For any enterprise data, if the target enterprise data is single-source, the amplitude is less than or equal to 30% of the enterprise scenario adaptation threshold, and the duration is less than one collection cycle, it is marked as noisy jump data. For noisy jump data, adjacent period smoothing and multi-source benchmark calibration are used for repair. The average of the target enterprise data mean of the first 3 periods and the target enterprise data mean of the last 3 periods is calculated and multiplied with the multi-source benchmark coefficient to obtain the noisy jump repair data. The multi-source benchmark coefficient is the ratio of the mean of the associated data source during the same period to the mean of the historical target enterprise data during the same period. If the target enterprise data is from multiple sources, the amplitude is greater than the preset enterprise scenario adaptation threshold of 50%, and the duration is greater than or equal to 2 collection periods, it is marked as risky jump data. For risky jump data, the original target enterprise data is retained and a jump risk label is marked. A single-source-multi-source differentiated repair model is constructed. Single-source isolated anomalies are repaired using time-series interpolation and reliability weighting.Multi-source collaborative anomalies are not repaired; instead, a collaborative anomaly risk label is used. For continuous periodic anomalies, a two-factor approach of period overlap and anomaly superposition is introduced to determine repair. The difference between the current period's anomaly amplitude and the previous period's anomaly amplitude is calculated and then compared with the previous period's anomaly amplitude to obtain the anomaly overlap. If the period overlap is greater than or equal to 0.6 and the anomaly superposition is greater than 0, it is considered risky data, the data is retained, and an operational warning is triggered, such as continuous periodic anomalies caused by system performance degradation. If the anomaly superposition is 0 and meets the preset operational cycle rules, it is considered normal and smoothed using the mean of the current cycle, such as continuous anomalies caused by weekly system maintenance. For regular periodic anomalies, a matching verification mechanism between the business cycle and the anomaly cycle is constructed. The model verifies the matching degree between abnormal cycles and business cycles using an enterprise business cycle dictionary. If the matching degree is greater than or equal to 0.8, it is judged as normal business fluctuation and marked as a normal cycle fluctuation label; if the matching degree is less than 0.8, it is judged as a risk anomaly, and the original data is retained. The enterprise business cycle dictionary includes monthly business cycles, quarterly business cycles, and annual business cycles. It integrates time-series trend correction to calculate the correlation degree of jump risks, achieving collaborative anomaly classification and risk attribution in general enterprise scenarios. The jump risk correlation degree R is obtained through the formula: R=T×M×S×(1+λ); where the jump risk correlation degree takes the value [0,1], and the larger the jump risk correlation degree, the higher the risk; T is the time-series correlation coefficient, taking the value [0,1]. The minimum value of the ratio of the jump duration to the enterprise scenario adaptation threshold, with a maximum value of 1, is used. The enterprise scenario adaptation threshold is set based on historical data statistics of the enterprise. M is the multi-source corroboration coefficient, with a value of [0, 1]. It is dynamically assigned based on the number of associated corroborations. For example, 0 is assigned for no associated corroboration, 0.5 for one associated corroboration, and 1 for two or more associated corroborations. S is the enterprise rule coefficient, with a value of [0, 1]. 1 is assigned for compliance with enterprise risk rules, 0 for compliance with normal enterprise jump rules, and 0.3 for others. λ is the time-series trend correction coefficient, with a value of [0, 0.2]. For example, 0.2 is assigned for an upward trend, 0.1 for a fluctuating trend, and 0 for a downward trend. It is used to strengthen the risk weight of trend-based jumps. Through jump risk... The risk correlation degree is used to determine the level of abrupt risk in general enterprise scenarios. For example, if the abrupt risk correlation degree is greater than or equal to 0.7, it is considered high abrupt risk; if it is greater than or equal to 0.4 and less than 0.7, it is considered medium abrupt risk; and if it is less than 0.4, it is considered low abrupt risk. At the same time, it realizes the preliminary correlation verification of multi-source data, solving the triple problems of inaccurate anomaly classification, missing risk attribution, and data distortion. The fusion correlation module combines the abrupt risk correlation degree to perform cross-data source cross-verification. Using a two-factor fusion algorithm, it retains the core features of risky abrupt data, extracts multi-dimensional risk correlation time-series features and cross-data source collaborative risk features, and filters core features through feature-abrupt risk correlation density calculation.A two-tier semantic mapping dictionary, consisting of a general semantic layer and an enterprise-customized layer, is constructed to address the semantic differences in multi-source heterogeneous data within enterprises. The general semantic layer unifies the definition and format of basic data, ensuring semantic consistency. Basic data includes, but is not limited to, amounts, dates, data identifiers, and timestamps. The enterprise-customized layer supplements semantic mapping rules for personalized business data, unifying statistical definitions and naming conventions across different internal systems. Through semantic alignment, heterogeneous data, including enterprise business data, operational data, and external industry data, is transformed into fusionable data with a unified semantic standard. Before multi-source data fusion, the correlation of jump risks is considered. Cross-data source cross-validation is performed; for each jump value detected by a single data source, it is matched one by one with the contemporaneous data of related data sources. If there is a synchronization anomaly in the related data source, the jump risk correlation is increased by 0.2; if there is no anomaly in the related data source and it conforms to the normal business rules of the enterprise, the jump risk correlation is decreased by 0.1; if the related data source is missing data, the original R value is retained and a data incomplete label is marked; cross-validation is used to realize multi-source verification of outliers, avoid misjudgment from a single data source, and improve the accuracy of anomaly judgment; based on the verified jump risk correlation, a two-factor fusion algorithm of risk weight and data reliability is designed, using the formula: Calculate the fused data value ,in, The correlation degree of the jump risk after the i-th data source is verified; The original data value of the i-th data source; Let be the reliability coefficient of the i-th data source, with a value in the range [0.8, 1]. This coefficient is calculated based on the historical accuracy statistics of the data source. Total number of data types This system indexes data types; leverages the synergistic effect of risk correlation and reliability coefficients to increase the proportion of data sources with high risk correlation and strong reliability in the fusion process, fully preserving the core characteristics of risky jump data and reducing interference from noisy and low-reliability data; through the collaborative fusion of multi-source data, it verifies the propagation characteristics of outliers in multi-source data, improving the reliability and universal adaptability of the fused data; combining the general temporal characteristics of enterprise jump data, it extracts multi-dimensional risk correlation temporal features to adapt to various enterprise data types; it calculates the difference between the jump peak and the normal baseline to obtain the jump amplitude, obtains the jump frequency by obtaining the number of jumps per unit time, and calculates the rate of data change during the jump process to obtain the jump trend. The slope is used to obtain the time from the end of the jump to the data returning to the normal baseline, thus obtaining the recovery time after the jump. All time-series features are associated and labeled with the enterprise's historical risk event records to clarify the risk orientation of the features and ensure that the features accurately match the enterprise's risk assessment needs. Combining multi-source fusion data, a general cross-data source risk association feature is constructed, supplemented with the periodic association features of four types of risk data to characterize collaborative risk characteristics. Collaborative risk characteristics include: basic association features including: business data-operation and maintenance data association coefficient, core business-auxiliary business association coefficient, and internal data-external data association coefficient; periodic features including jump cycle collaboration coefficient, continuous cycle abnormal association degree, regular cycle matching deviation coefficient, and periodic anomaly. The propagation coefficients are as follows: Specifically, the jump cycle synergy coefficient is the ratio of the overlap between the jump occurrence period and the multi-source data jump period to the jump duration, characterizing the multi-source cycle synchronicity of jumps; the continuous cycle anomaly correlation coefficient is the ratio of the frequency of anomalies in related data sources occurring concurrently to the number of consecutive cycles when a continuous cycle anomaly occurs, used to strengthen the risk verification of continuous cycle anomalies; the regular cycle matching deviation coefficient is the ratio of the absolute value of the difference between the abnormal cycle amplitude and the preset normal business cycle amplitude threshold to the preset normal business cycle amplitude threshold, used to distinguish between normal cycle fluctuations and risky cycle anomalies; the cycle anomaly propagation coefficient is the ratio of the time delay of anomalies in other related data sources after a data source cycle anomaly occurs to the cycle duration, used to characterize... This paper describes the multi-source propagation risk of cyclical anomalies. By combining basic correlation features with cyclical innovation features, it achieves a comprehensive risk correlation representation of four types of risk data in multi-source enterprise data, making up for the limitations of single data sources and non-cyclical features. It strengthens the screening of cyclical features through general risk feature screening, and designs a feature-jump risk correlation density calculation method that integrates timeliness and cyclical characteristics to achieve the screening of core risk features of four types of risk data. The general feature-jump risk correlation density D is calculated by the formula: D=(C / T)×K×(1+τ)×(1+ω), with a value range of [0, 1], where ω is the cyclical feature correction factor with a value range of [0, 0.2], which is used to strengthen the weight of cyclical features.For example, the rules for ω are as follows: if the feature is the correlation coefficient between the jump cycle and the continuous cycle anomaly, and the co-occurrence frequency with the risk-type jump cycle is greater than or equal to 0.7, then ω = 0.2; if the co-occurrence frequency with the risk-type jump cycle is within [0.3, 0.7), then ω = 0.1; if the co-occurrence frequency with the risk-type jump cycle is less than 0.3, then ω = 0; for non-periodic features, ω = 0, C is the co-occurrence frequency of the feature with the risk-type jump, T is the total occurrence frequency of the feature, K is the fit between the feature change magnitude and the risk level, and τ is the feature timeliness correction factor; a general enterprise scenario screening threshold is set to retain the synergistic risk characteristics of general features—jump risk correlation density higher than the threshold, and training is performed to avoid high The system addresses the disaster of dimensionality in data, preserving collaborative risk characteristics, dynamic timeliness features, and cyclical correlation features to enhance the machine learning model's ability to identify continuous and regular cyclical anomalies. The assessment and calibration module refines the feature system of the four types of risk data through feature expansion and sample enhancement, constructs a trend-driven interpretable ensemble model, builds a three-layer ensemble model, outputs dual-dimensional assessment results, incorporates trend risk coefficients, and corrects model bias through closed-loop iterative optimization. It also performs full-dimensional feature expansion and sample enhancement on the four types of risk data, expanding the full-dimensional feature system of the four types of risk data to provide data support for machine learning model upgrades. Furthermore, it expands the basic features of the four types of risk data, supplementing them with data type identifiers and variation levels; for example, mild variation. The amplitude is less than or equal to the enterprise scenario adaptation threshold of 30%; moderate variation: amplitude greater than the enterprise scenario adaptation threshold of 30% and less than or equal to the enterprise scenario adaptation threshold of 60%; severe variation: amplitude greater than the enterprise scenario adaptation threshold of 60%; Analyze the trend characteristics of the four types of risk data: calculate the change rate of the frequency of occurrence of each category of the four types of risk data within a unit time, such as: by calculating the difference between the frequency of the last 3 periods and the average frequency of the last 10 periods and then comparing it with the average frequency of the last 10 periods, the change rate of risk occurrence frequency is obtained; use a 3-period sliding window to calculate the sliding growth rate of the total number of the four types of risk data within a unit time, and obtain the risk quantity growth rate; calculate the change rate of the proportion of severe variation data in the four types of risk data within a unit time and the sum of the variation amplitude. The sum of the growth rates yields the characteristic value of the degree of variation. Sample augmentation and optimization are performed based on the trend characteristics of four types of risk data: Four types of risk data trend samples are generated by combining these trend characteristics. Virtual samples of different trend types are generated using the historical distribution of these four types of risk data trend characteristics of the enterprise. Simultaneously, desensitized trend samples of the same industry's four types of risk data are introduced for transfer learning to ensure the adaptability of the machine learning model to various trend scenarios. The augmented sample set can meet the training requirements of the machine learning model. A trend-driven interpretable ensemble model is constructed by combining machine learning. A three-layer ensemble model is built by combining basic risk assessment, four types of risk data trend analysis, and fusion decision-making. A time-series attention mechanism and trend prediction component are integrated to achieve a deep upgrade of the machine learning model.Specifically, it uses machine learning models such as logistic regression, gradient boosting trees, long short-term memory networks, and random forests to output the risk level of abrupt changes; trend analysis of four types of risk data: adopting a combined architecture of temporal attention, LSTM, and trend classifier: combining temporal attention and LSTM, the frequency change rate, risk quantity growth rate, and variability feature values ​​of the four types of risk data are input, and the attention mechanism captures key time nodes in the trend changes, outputting the temporal trend coefficient Tr of the four types of risk data, with a value of [0, 1]. The larger the temporal trend coefficient, the more dangerous the trend; trend classifier: using gradient boosting trees, it determines the trend type based on trend features and outputs the trend risk level. This includes three risk levels: high-risk, medium-risk, and low-risk. A dual-dimensional fusion algorithm combining abrupt change risk and trend risk is designed to calculate the final risk assessment value, achieving in-depth risk assessment. The final risk assessment value Z is calculated using the formula: Z = R × (1 + Tr) × θ, where θ is the trend type correction coefficient. For example, 1.3 is used for rapid rise, 1.1 for rise, 1.0 for fluctuation, 0.9 for stability, and 0.7 for decline. The algorithm also outputs a combined assessment result of the four risk data types: abrupt change risk level, trend risk level, and key drivers. A trend-risk correlation explanation mechanism is added to generate a heatmap of the trend characteristics and risk contribution of the four risk data types, showcasing the core correlations and meeting the enterprise's needs for risk assessment. The need for in-depth tracing of risk causes; combining a dynamic calibration mechanism driven by the trends of four types of risk data, incorporating the trend characteristics of the four types of risk data to achieve a deep upgrade of the calibration mechanism; calculating the general dynamic calibration coefficient α using the formula: α=A×B×(1+γ)×(1+Tr), with a value range of [0.8,1.5], expanding the value range to adapt to trend-driven risk fluctuations, where Tr is the time series trend coefficient of the four types of risk data, with a value range of [0,1], output by the trend analysis module, used to integrate trend risk quantification into the calibration process; the definitions of other parameters are the same as before: A is the jump type adaptation coefficient; B is the scenario change coefficient; γ is the sample confidence correction term; the risk assessment value is compared with the general... The dynamic calibration coefficients are multiplied to obtain the risk assessment correction value, ensuring that the three-layer integrated model adapts to change types and scenarios, responds to the trend characteristics of four types of risk data, and improves the accuracy of dynamic risk assessment. The trend analysis results of the four types of risk data are incorporated into the optimization driving factors. The ratio of the number of deviations between the trend assessment level and the actual trend risk level to the total number of assessments is calculated to obtain the trend misjudgment rate. The actual trend risk level is determined by retrospective analysis of subsequent risk events of the enterprise. The trend misjudgment rate calculated by combining the trend characteristics of the four types of risk data with the preset deviation correction coefficient is used for dual-drive closed-loop iterative optimization. The deviation correction coefficient is weighted to simultaneously take into account the optimization needs of both change risk assessment and trend assessment.For four types of risk data with a false positive rate exceeding a preset false positive rate threshold, the feature extraction rules and input weights of the three-layer integrated model corresponding to the four types of risk data are optimized to ensure that the system continuously adapts to the trend changes of the four types of risk data. For example, the periodic window for calculating the degree of anomaly superposition is adjusted for continuous periodic anomalies, thereby improving the depth and accuracy of long-term risk assessment. Example 2: As shown in Figure 2, the risk assessment method based on machine learning and enterprise multi-source data fusion provided in this embodiment of the invention specifically includes the following steps: Step 1: Through the dynamic adjustment mechanism of basic frequency and trigger enhancement, and by constructing correlation pairs, a three-dimensional general judgment system is established to clarify the core laws of the four types of risk data and complete the classification judgment; Step 2: Based on the core laws of the four types of risk data The process involves four steps: Step 1: A three-dimensional repair threshold dynamic adjustment mechanism is adopted to perform differentiated repair on abnormal data. The correlation degree of abrupt change risks is calculated and quantified to classify the risks. Step 2: Cross-data source cross-validation is performed based on the correlation degree of abrupt change risks. A two-factor fusion algorithm is used to retain the core features of risk-type abrupt change data, extracting multi-dimensional risk correlation time-series features and cross-data source collaborative risk features. Core features are selected through feature-abrupt change risk correlation density calculation. Step 3: The feature system of the four types of risk data is improved through feature expansion and sample enhancement. A trend-driven interpretable ensemble model is constructed, a three-layer ensemble model is built, and a two-dimensional evaluation result is output, incorporating trend risk coefficients. The model bias is corrected through closed-loop iterative optimization.

[0010] The above provides a detailed description of one embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and should not be considered as limiting the scope of the present invention. The above formulas are all dimensionless numerical calculations, and the formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world situation. The preset parameters in the formulas are set by those skilled in the art based on actual conditions and historical experience, and can be adjusted according to actual conditions. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. All equivalent changes and improvements made in accordance with the scope of the present invention should still fall within the patent coverage of the present invention.

Claims

1. A risk assessment system based on machine learning and enterprise multi-source data fusion, characterized in that, It includes the following modules: Data Judgment Module: Through a dynamic adjustment mechanism of basic frequency and trigger enhancement, and by constructing correlation pairs, a three-dimensional general judgment system is established to clarify the core laws of four types of risk data and complete the classification judgment; Repair Quantification Module: Based on the core laws of four types of risk data, a three-dimensional repair threshold dynamic adjustment mechanism is adopted to implement differentiated repair of abnormal data, integrate and calculate the correlation degree of jump risk, and quantify and classify jump risk. Fusion and correlation module: Combines cross-data source cross-validation with the correlation degree of jump risk, uses a two-factor fusion algorithm to retain the core features of risk-type jump data, extracts multi-dimensional risk correlation time series features and cross-data source collaborative risk features, and filters core features by calculating the feature-jump risk correlation density. Assessment and calibration module: Improve the feature system of four types of risk data through feature expansion and sample enhancement, construct a trend-driven interpretable ensemble model, build a three-layer ensemble model, output two-dimensional assessment results, incorporate trend risk coefficients, and correct model bias through closed-loop iterative optimization.

2. The risk assessment system based on machine learning and enterprise multi-source data fusion as described in claim 1, characterized in that, The method for establishing a three-dimensional universal judgment system is as follows: A dynamic adjustment mechanism for basic frequency and trigger enhancement is constructed. By analyzing historical multi-source data of the enterprise, the duration and frequency characteristics of jumps in different types of multi-source data are statistically analyzed to set a basic sampling frequency. While collecting core data, the enterprise's full-dimensional related context data is collected simultaneously to construct core data-context data pairs. A three-dimensional universal judgment system based on time-series characteristics, multi-source corroboration, and enterprise rules is constructed to clarify the core patterns of four types of risk data: jump values, outliers, continuous periodic outliers, and regular periodic outliers, thus completing the classification and judgment.

3. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for identifying the core patterns of the four types of risk data is as follows: Time-series characteristic analysis is performed by combining the cyclical characteristic dimension to calculate the duration, trend, and cyclical overlap of jumps / anomalies; where the cyclical overlap is the ratio of the overlap between the anomaly occurrence period and the business cycle period to the duration of the business cycle; Cyclical synergy is strengthened through multi-source corroboration verification by querying whether there are synchronous anomalies in related data sources within the same cycle; Cyclical rules are supplemented using enterprise rules, and specific tags are marked for regular cyclical anomalies and continuous cyclical anomalies based on the enterprise business cycle table, excluding regular anomalies within normal business cycles.

4. The risk assessment system based on machine learning and enterprise multi-source data fusion as described in claim 1, characterized in that, The method for implementing differentiated repair is as follows: design a classification repair and periodic calibration strategy, and adopt a dynamic adjustment mechanism for the three-dimensional repair threshold of amplitude-time-multi-source; construct a single-source-multi-source differentiated repair model, and use time-series interpolation and reliability-weighted repair for single-source isolated anomalies; Multi-source collaborative anomalies are not repaired, but are labeled with a collaborative anomaly risk tag; for continuous periodic anomalies, a two-factor judgment and repair method is introduced based on periodic overlap and anomaly superposition.

5. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for quantifying and classifying the risk of abrupt changes is as follows: For regular periodic outliers, a verification model for matching business cycles and outlier cycles is constructed, and the matching degree between outlier cycles and business cycles is verified through the enterprise business cycle dictionary. By integrating time-series trend correction to calculate the correlation of jump risks, we can achieve collaborative anomaly classification and risk attribution in general enterprise scenarios; The analysis yielded the correlation degree of the risk of abrupt changes.

6. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for cross-data source cross-validation is as follows: construct a two-level semantic mapping dictionary with a general semantic layer and an enterprise-customized layer, and unify the definition and format of basic data in the general semantic layer; For abrupt values ​​detected by a single data source, they are matched one by one with contemporaneous data from related data sources; multi-source verification of outliers is achieved through cross-validation; based on the correlation of abrupt risk after verification, a two-factor fusion algorithm of risk weight and data reliability is designed to calculate the fused data value. By leveraging the synergistic effect of risk correlation and reliability coefficient, the proportion of data sources with high risk correlation and strong reliability in the fusion process is increased, fully preserving the core characteristics of risky abrupt data. Through the collaborative fusion of multi-source data, the propagation characteristics of outliers in multi-source data are verified.

7. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for retaining the core features of risky jump data is as follows: calculate the difference between the jump peak and the normal baseline to obtain the jump amplitude; obtain the number of jumps occurring per unit time to obtain the jump frequency; calculate the rate of data change during the jump to obtain the jump trend slope; and obtain the time from the end of the jump to the data returning to the normal baseline to obtain the recovery time after the jump. All time-series features are linked and labeled with the company's historical risk event records to clearly define the risk orientation of the features.

8. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for selecting core features is as follows: combining multi-source fusion data to construct a general cross-data source risk association feature, supplementing the periodic association features of four types of risk data, and characterizing the collaborative risk characteristics; the jump period synergy coefficient is the ratio of the overlap between the jump occurrence time and the jump time of multi-source data to the jump duration, characterizing the multi-source periodic synchronization of jumps. By combining basic correlation features with cyclical innovation features, a comprehensive risk correlation representation of four types of risk data in multi-source enterprise data is achieved; by strengthening the screening of cyclical features through general risk feature screening, a feature-jump risk correlation density calculation method that integrates timeliness and cyclical characteristics is designed to screen the core risk features of the four types of risk data; The general feature-jump risk association density is calculated; Set a general enterprise scenario screening threshold, retain the general feature - the collaborative risk characteristic with the jump risk correlation density higher than the threshold, and train to retain the collaborative risk characteristic, dynamic timeliness characteristic and periodic correlation characteristic.

9. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for constructing the three-layer integrated model is as follows: basic features are extended for the four types of risk data, and data type identifiers and degree of variation of the four types of risk data are added; trend samples of the four types of risk data are generated by combining the trend features of the four types of risk data; virtual samples of different trend types are generated by the historical trend feature distribution of the four types of risk data of the enterprise; and the desensitized trend samples of the four types of risk data of the same industry are introduced for transfer learning. A trend-driven, interpretable ensemble model is constructed by combining machine learning. A three-layer ensemble model is built by combining basic risk assessment, trend analysis of four types of risk data, and fusion decision-making, and incorporates a time-series attention mechanism and a trend prediction component. The architecture combines time-series attention, LSTM, and a trend classifier: the frequency change rate, risk quantity growth rate, and variability feature values ​​of the four types of risk data are input by combining time-series attention and LSTM. The attention mechanism captures key time nodes in the trend changes and outputs the time-series trend coefficients of the four types of risk data. trend Classifier: Gradient boosting tree is used to determine the trend type based on trend features and output the trend risk level; The algorithm is designed to combine two dimensions of risk abruptness and risk trend to calculate the final risk assessment value, thereby achieving in-depth risk assessment and obtaining the final risk assessment value. Simultaneously, it outputs a combined assessment result of four types of risk data: jump risk level, trend risk level, and key drivers; An added trend-risk correlation explanation mechanism was developed to generate heatmaps showing the trend characteristics and risk contribution of four types of risk data, demonstrating the core correlations.

10. The risk assessment system based on machine learning and enterprise multi-source data fusion according to claim 1, characterized in that, The method for correcting model bias is as follows: combining a dynamic calibration mechanism driven by the trends of four types of risk data, incorporating the trend characteristics of the four types of risk data, and achieving a deep upgrade of the calibration mechanism. Calculate the general dynamic calibration coefficient, and multiply the risk assessment value by the general dynamic calibration coefficient to obtain the risk assessment correction value; The trend analysis results of the four types of risk data are incorporated into the optimization driving factors. The ratio of the number of deviations between the trend assessment level and the actual trend risk level to the total number of assessments is calculated to obtain the trend misjudgment rate. The actual trend risk level is determined by retrospective analysis of subsequent risk events of the enterprise. The trend misjudgment rate calculated by combining the trend characteristics of four types of risk data is used for dual-drive closed-loop iterative optimization with the preset deviation correction coefficient.