Enterprise full life cycle credit risk decision method and system based on multi-source data

The enterprise full lifecycle credit risk decision-making system, which uses multi-source data, solves the problems of fragmented data processes and lax access control in existing technologies. It achieves standardized operation processes, refined access control, intensive resource management, and intelligent anomaly handling, thereby improving the efficiency and security of enterprise credit risk assessment.

CN122155830APending Publication Date: 2026-06-05SHENZHEN CREDIT INFORMATION SERVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN CREDIT INFORMATION SERVICE CO LTD
Filing Date
2026-03-04
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies in corporate credit risk assessment suffer from problems such as fragmented data processes, lax access control, poor tool compatibility, weak anomaly handling capabilities, cumbersome variable management, and insufficient security and audit isolation, making it difficult to meet the needs of financial institutions for systematic extraction and modeling of risk characteristics throughout the entire life cycle.

Method used

The enterprise full lifecycle credit risk decision-making system based on multi-source data includes a data aggregation module, a visualization navigation module, a variable overview module, a variable filtering module, a sample matching module, a sandbox modeling module, a table lifecycle management module, and a system management module. Combined with the government cloud modeling platform and the state-owned assets cloud modeling platform, it realizes a unified management and isolation mechanism for multi-source data, and supports custom algorithm programming and multi-level access control.

Benefits of technology

It has achieved standardized operation processes, refined access control, intensive resource management, personalized function adaptation, and intelligent anomaly handling, thereby improving data security, compliance, and the stability of the modeling process, and meeting the diverse business needs of financial institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155830A_ABST
    Figure CN122155830A_ABST
Patent Text Reader

Abstract

The application provides an enterprise full life cycle credit risk decision method and system based on multi-source data, the system gathers multi-source heterogeneous data such as industry and finance, and breaks the information island; the system extracts the characteristics of nine dimensions such as enterprise basic information and operation behavior, constructs super web variable, traces back the history growth context of the enterprise, and identifies the dynamic risk characteristics of the enterprise in different life cycle stages; the system supports self-service training of quantitative models, and can be applied to product innovation, risk decision and post-loan management and other credit whole processes. The application realizes the dynamic risk identification and self-service decision of the credit whole process by deeply fusing multi-source heterogeneous data to construct the enterprise full life cycle credit portrait, and effectively helps financial institutions to break the data island and realize digital transformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, specifically to a method and system for enterprise full lifecycle credit risk decision-making based on multi-source data. Background Technology

[0002] In the course of financial lending, accurate assessment of a company's credit status and risk level is a fundamental step in ensuring the safety of credit assets and improving business efficiency. With the increase in the number of companies, the diversification of business models, and the increasing dynamic changes in their operating conditions, relying on traditional methods to conduct corporate risk analysis often makes it difficult to simultaneously meet the requirements of accuracy, timeliness, and comprehensiveness in the assessment.

[0003] In the past, financial institutions relied heavily on static financial statements, limited qualification documents, and manual surveys submitted by enterprises for risk analysis. This approach was prone to problems such as limited information sources, outdated information, and difficulty in reflecting changes in business operations and dynamic risk signals. Consequently, risk assessment biases were identified, which further constrained product policy formulation and the development of differentiated strategies.

[0004] Meanwhile, enterprise-related information is objectively scattered across multiple data sources, including government, commerce, and finance. For example, government data includes business registration, administrative penalties, water and electricity usage, and social security / housing provident fund information; commerce data includes intellectual property, investment and financing, customs, judicial data, and audited financial statements; and finance data includes bank credit history and credit inquiry records. However, under current technological conditions, it is often difficult to achieve efficient aggregation, unified governance, and traceable feature backtracking of multi-source heterogeneous information within a unified framework. This hinders financial institutions' systematic extraction and modeling of enterprise "full lifecycle" risk characteristics (such as risk decision-making and post-loan early warning).

[0005] The shortcomings of existing technology: (a) Shortcomings of traditional "multi-tool" solutions: The most common approach in existing technologies is to use Excel / SPSS for data filtering and basic exploration, train models in a local Python environment, manage tables using native database tools, and control accounts using a third-party permission management system.

[0006] The main drawback of this type of solution is: 1. Fragmented processes require manual data migration between multiple tools, which can easily lead to data loss, formatting errors, and overall low efficiency.

[0007] 2. The lack of a unified access control system and the independent access rules for different tools make it difficult to achieve secure isolation of data throughout the entire process, posing a risk of data leakage.

[0008] 3. Poor tool compatibility: Local algorithm package configuration, model results and database table management are difficult to integrate smoothly, often requiring manual secondary processing, resulting in poor adaptability.

[0009] 4. Lack of intelligent anomaly handling; problems such as memory overflow and duplicate data often rely on manual troubleshooting and repair, making it difficult to guarantee stability.

[0010] (II) Shortcomings of general-purpose data modeling platform solutions While some general-purpose modeling platforms have basic exploration, training, and storage capabilities, they are usually not customized to meet the compliance, security, and localization needs of the financial industry.

[0011] The main drawback of this type of solution is: 1. The access control is too rudimentary, making it difficult to achieve full-chain isolation of "account-database-operating environment" and meet the compliance requirements of financial institutions.

[0012] 2. Insufficient business adaptability, with inadequate support for localized matching range and personalized algorithm extensions (such as self-installation of special algorithm packages), making it difficult to cover diverse business scenarios.

[0013] 3. Resource management is fragmented, with a lack of centralized collection and refined management of resources throughout the modeling process, such as original samples, intermediate data, and parameter configurations, making model iteration and compliance audit traceability difficult.

[0014] 4. The ability to handle exceptions is weak. There is a lack of automated repair mechanisms for common problems such as memory overflow and data matching failure, which rely on manual intervention.

[0015] (III) Common shortcomings in scenarios addressing the "enterprise lifecycle characteristics" 1. On the variable management side, traditional processes often suffer from cumbersome management procedures and unclear permission boundaries, making it difficult to support batch / real-time updates and controlled downloads of variable dictionaries.

[0016] 2. On the sample backtracking and table creation side, large-scale variable / sample processing is prone to triggering table creation memory overflow and process interruption, and lacks automatic repair capabilities, affecting continuity and availability.

[0017] 3. On the security and audit side, the "loose permissions and insufficient isolation" common in traditional platforms can lead to cross-account data access risks, making it difficult to meet the requirements of financial institutions for compliance audit and security traceability.

[0018] Therefore, existing technologies have shortcomings and need further improvement. Summary of the Invention

[0019] To address the problems existing in the prior art, this invention provides a method and system for enterprise full lifecycle credit risk decision-making based on multi-source data.

[0020] To achieve the above objectives, the specific solution of the present invention is as follows: This invention provides a corporate full lifecycle credit risk decision-making system based on multi-source data, comprising: The data aggregation module is used to collect and aggregate heterogeneous information from multiple sources, such as government data, industry data, commercial data and financial data, based on a distributed data processing architecture. It also creates profiles of enterprises according to preset feature dimensions, generates enterprise profile variables, and supports the retrospective analysis of enterprise credit risk feature data. The visual navigation module is used to respond to user operation commands after user login, call system function interfaces, and generate a step-by-step operation guidance interface. The variable overview module is used to provide a global overview of the data coverage of the enterprise profile variables according to the month dimension; The variable exploration module is used to perform distribution exploration and quantile statistics display on the target variable to obtain the coverage and value distribution characteristics of the target variable in different time periods. The variable filtering module is used to filter from the enterprise profile variables to form a set of target variables based on filtering conditions such as dimension, coverage and data type. The sample matching module is used to receive a sample table containing the unified social credit code and the retrospective date, and perform table creation and database storage and feature retrospective processing on the target variable set to generate a sample feature table. The sandbox modeling module is used to perform model training, parameter tuning, and model validation on the sample feature table in an isolated modeling environment based on a visual interactive interface, to obtain a credit risk scoring model and output a risk score for risk decision-making and post-loan management. The table lifecycle management module is used to manage the lifecycle of the sample feature table and perform dependency verification and state locking when a deletion instruction is received. The model management module is used to centrally collect and manage the sample data, intermediate processing data, variable feature data, parameter configuration data, iteration process data, and model output results generated during the modeling process. The system management module is used to manage accounts, roles, and permissions throughout their entire lifecycle, and to manage the isolated configuration of sandbox accounts.

[0021] Furthermore, it also includes the government cloud modeling platform and the state-owned assets cloud modeling platform. The two are respectively configured with data exchange, model training, model deployment and account partitioning functions, and interact with the government cloud database and the state-owned assets cloud database respectively, so as to realize cross-data domain data exchange and domain-based management of modeling resources. The sandbox modeling module is a Python-based Jupyter visual interface with built-in supervised and unsupervised algorithm libraries and supports custom algorithm programming. It also supports joint modeling of financial institutions' internal labels and the enterprise profile variables. An isolation mechanism is adopted that binds accounts to database permissions one-to-one, so that each user account can only access the specified database and table resources bound to it; The system management module assigns an independent Python runtime partition to each sandbox account and allows the installation of algorithm packages within the runtime partition. It also records account operations, role configurations, and permission change logs for auditing and traceability. The model management module performs regular backups of the collected modeling data and model results.

[0022] Furthermore, the variable overview module filters the variable coverage by month based on nine feature dimensions and uses a spectral heatmap to display the distribution of data coverage percentage for each month. The variable exploration module supports both hierarchical drill-down filtering based on first-level, second-level, and third-level dimensions, as well as precise filtering based on variable names. After filtering, it displays the latest monthly non-empty sample size distribution, data coverage, and quantile statistics of the target variable arranged in monthly order.

[0023] Furthermore, the sample matching module includes variable upload management, sample upload matching, and backtracking status monitoring sub-modules, and satisfies: Sample uploads only support sample tables containing "Unified Social Credit Code" and "Backtrack Date", and complete duplicate samples with the same Unified Social Credit Code and Backtrack Date will be checked and blocked. Automatic splitting is performed during table creation, database storage, and feature backtracking to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3. When a memory overflow error is detected during the table creation process, a one-click rerun mechanism is provided, which automatically breaks down the original table creation statement into multiple sub-table creation statements and re-executes them.

[0024] Furthermore, it also includes a remote sensing feature extraction module, an industrial IoT feature extraction module, and a trusted acquisition and remote verification module, wherein: (1) The remote sensing feature extraction module is used to: determine the remote sensing region of interest (ROI) centered on the geographical coordinates obtained by parsing the business address or factory address of the enterprise entity, and rasterize the ROI with a preset spatial grid; perform monthly aggregation feature generation on the remote sensing image covering the ROI, and the generated features include at least the monthly average value of nighttime light intensity, the month-on-month change rate of nighttime light intensity, and the change rate of building outline; wherein, the grid side length of the spatial grid is 100m to 250m, and when the effective pixel ratio of the remote sensing image in the ROI is less than 0.6, the remote sensing features of the corresponding month are set to missing and a missing indicator variable is generated; (2) The industrial IoT feature extraction module is used to: collect industrial operation signals, including at least the power load time series and the key equipment vibration time series, at the enterprise-side edge gateway, and perform windowed feature processing on the industrial operation signals at the edge gateway to obtain daily industrial operation features; wherein, the sampling period of the power load time series is 1min to 5min, the sampling frequency of the key equipment vibration time series is 1kHz to 10kHz, the window length of the windowed feature processing is 10s to 60s and at least outputs frequency domain energy, spectral entropy, kurtosis and outlier density; and upload the daily industrial operation features to the data aggregation module on a daily basis, and aggregate them into monthly industrial operation features on a monthly basis; (3) The trusted acquisition and remote verification module is used to: generate trusted verification information for remote sensing features and industrial operation day features. The trusted verification information includes at least the device identity identifier, firmware metric hash, feature data digest hash, and a remote verification token formed by a timestamped signature. Before writing the remote sensing features and the industrial operation day features into the data storage, the remote verification token is verified. It is allowed to be stored only when the verification is successful. Otherwise, it is rejected and an audit log is recorded. The audit log is stored in a chain hash structure, and each audit record includes at least the previous hash, the current record digest hash, and a timestamp. (4) The remote sensing features, the monthly industrial operation features, and the enterprise profile variables are aligned using a unified retrospective approach: the retrospective dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned remote sensing features and monthly industrial operation features are included as new enterprise profile variables in the enterprise profile variables, so that the variable screening module can screen them and the sample matching module can retrospectively generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

[0025] Furthermore, it also includes a network security attack surface characterization module, a supply chain logistics flow characterization module, and a decision audit and evidence retention module, and meets the following constraints and processing rules: (1) The network security attack surface feature module is used to: periodically acquire the public network assets of the enterprise entity and generate attack surface features. The public network assets include at least domain names, subdomains, IP addresses and externally open ports. The mapping between the public network assets and the enterprise entity adopts at least two types of evidence consistency verification. The two types of evidence include at least two of the following: ICP filing information, digital certificate organization information, WHOIS query information and subdomain resolution relationship. The attribution confidence score is calculated for each mapping result. The corresponding network asset is included in the attack surface calculation only when the attribution confidence score is not lower than a preset threshold. The attack surface features include at least the number of externally open ports, the distribution of external service types, the number of high-risk vulnerabilities, the statistics of vulnerability repair time, and the number of suspected counterfeit / phishing domain names and the number of certificate anomalies. The periodic acquisition frequency is once every 7 to 30 days, and the results are aggregated into monthly network security features. (2) The supply chain logistics flow feature module is used to: acquire logistics and transportation data associated with the enterprise entity and generate supply chain flow features. The logistics and transportation data includes at least one of the following: maritime automatic identification system (AIS) trajectory data, waybill trajectory data, and vehicle diagnostic (OBD) data. The association mapping between the enterprise entity and the logistics and transportation data is based on at least two of the following for consistency verification: unified social credit code, enterprise name, transportation vehicle registration information, carrier contract number, or waybill number. The supply chain flow features include at least the monthly number of shipments, monthly number of arrivals, monthly abnormal delay rate, number of berthings, port stay duration, and regional dispersion of shipments and receipts. The regional dispersion of shipments and receipts is characterized by regional entropy or equivalent indicators, and the results are aggregated monthly into monthly supply chain flow features. (3) The decision audit and evidence retention module is used to: perform traceable audit and retention of the generation process, version configuration and storage process of the network security monthly features and the supply chain flow monthly features, as well as the process of participating in modeling and risk decision-making based on these two types of features. The audit retention includes at least the data source identifier, feature generation configuration version number, variable dictionary version number, sample backtracking caliber identifier, model version number, parameter configuration summary and risk decision output result; the audit retention adopts an immutable storage strategy, including at least one of chained hash log or WORM write-once-read-many storage, each audit record includes at least the previous hash, current record summary hash and timestamp, and forms alarm records for abnormal events; (4) The network security monthly features, the supply chain circulation monthly features, and the enterprise profile variables are aligned using a unified backtracking caliber: the backtracking dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned network security monthly features and supply chain circulation monthly features are included as new enterprise profile variables in the enterprise profile variables, so that the variable filtering module can filter them and the sample matching module can backtrack to generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

[0026] This invention also provides a method for enterprise full lifecycle credit risk decision-making based on multi-source data, which, based on the above system, includes: S1. Collect and aggregate heterogeneous information from multiple sources, including government data, industry data, commercial data, and financial data. Based on a distributed data processing architecture, generate enterprise profile variables according to nine feature dimensions to form enterprise credit risk feature data that supports the longest historical growth trajectory of up to nine years. S2. After the user logs in, a step-by-step operation guide and core function entry points are provided through a visual navigation module; S3. The variable overview module displays the data coverage of enterprise profile variables by month, and the variable exploration module performs distribution exploration and quantile statistics on the target variable to obtain variable coverage and distribution characteristics. S4. The target variable set is formed by filtering based on conditions such as dimension, coverage and data type through the variable filtering module; S5. Receive a sample table containing the unified social credit code and the retrospective date, perform table creation and database storage and feature retrospective for the target variable set, generate a sample feature table, and verify and intercept completely duplicate samples. S6. In the sandbox modeling module, perform model training, parameter tuning and model validation on the sample feature table to obtain the credit risk scoring model and output the enterprise risk score; S7. Collect the modeling process data, parameter configurations and model output results into the model management module for unified management, and perform table lifecycle management and model retention management through the table lifecycle management module and the model management module according to the retention strategy. S8. Make risk decisions based on the enterprise risk score, and implement risk warnings based on changes in enterprise characteristic data in post-loan management.

[0027] Furthermore, in step S3: the variable overview module uses a spectral heatmap to display the distribution of data coverage percentage for each month; the variable exploration module uses a combined graph to display the non-empty sample size and data coverage, and outputs the quantile statistics of the target variable in monthly order.

[0028] Furthermore, in step S5: when the number of variables to be processed exceeds a preset threshold, multiple target tables are automatically generated to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3; when a memory overflow error occurs during the table creation process, the original table creation statement is automatically disassembled and multiple sub-table creation statements are generated and re-executed, thereby completing feature backtracking.

[0029] Furthermore, in steps S6 and S7: a one-to-one binding mechanism between account and database permissions is adopted, so that users can only access the database and table resources bound to their accounts; an independent Python running partition is allocated to each sandbox account to achieve user isolation, and account operation, role configuration and permission change logs are recorded; whether to retain the sample feature table and credit risk scoring model is determined according to the preset retention strategy. When not retained, the corresponding table resources are deleted through the table lifecycle management module. When retained, the model resources are collected and backed up regularly through the model management module.

[0030] The technical solution of this invention has the following beneficial effects: 1. Standardize operating procedures to improve ease of use and work efficiency. This invention sets up a visual navigation module as the core functional navigation hub, which, together with standardized operation guidance, enables new users to quickly complete key processes such as variable viewing, variable exploration, screening and sample matching step by step; and through visualization mechanisms such as spectral graphs, combination graphs, and quantile statistics, the coverage and distribution characteristics of variables are presented intuitively, reducing the analysis threshold and shortening the data exploration time.

[0031] 2. Refined access control ensures data security, compliance, and auditability. This invention constructs a multi-layered isolation system of "account—database—operating environment—core resources": it avoids cross-account data access from the source by binding accounts to database permissions one-to-one; it restricts administrator permissions for core modules such as model management and system management; it adopts Python multi-partition isolation for the sandbox modeling environment and introduces a mechanism for binding sandbox accounts with feature accounts; and it combines operation log traceability, secondary confirmation for deletion, and duplicate data verification and interception mechanisms to form a closed-loop security chain, meeting the stringent requirements of financial institutions for data compliance and audit traceability.

[0032] 3. Resource management is centralized, enabling controllable retention, reusability, and traceability of resources throughout the modeling process; This invention uses a model management module to automatically collect and manage all resources in the modeling process, covering original samples, intermediate processing data, variable feature data, parameter configurations, iteration process data, and model output results. It also supports multi-dimensional retrieval, categorized archiving, and regular backups to avoid loss and management chaos caused by resource dispersion. At the same time, it uses the lifecycle management and visual cleanup mechanism of the table lifecycle management module to reduce redundant table usage and optimize database storage resources and operating efficiency.

[0033] 4. Personalized functionality to meet the diverse business needs and localized scenario requirements of financial institutions. This invention provides a Jupyter-based sandbox modeling environment on the modeling side, with built-in common algorithms and support for custom algorithm programming and installation of special algorithm packages to adapt to the model building needs of different business scenarios; on the variable side, it sets differentiated permissions for administrators and financial institution users, taking into account both variable dictionary management and business-side variable filtering and download; on the sample side, it supports multiple date formats and limits the matching range (e.g., the range of business entities in a specific region), while the system management module supports multiple role presets and multiple role bindings to adapt to permission configurations under complex organizational structures.

[0034] 5. Enhanced intelligence and stability in anomaly handling ensure reliable operation of large-scale backtracking and modeling processes; This invention addresses the high-risk aspects of large-scale variable table creation and feature backtracking by incorporating an automatic splitting mechanism to control the number of variables in a single table and the number of associated sub-tables, thereby effectively avoiding table creation memory overflow. It also provides an automatic "one-click rerun" repair capability for common errors such as memory overflow (automatically disassembling table creation statements and retrying execution), and is combined with a mechanism for real-time updating of backtracking status and viewing error causes, significantly reducing the failure rate, reducing manual troubleshooting costs, and improving the continuity and reliability of modeling and data processing workflows. Attached Figure Description

[0035] Figure 1 This is a diagram of the overall system architecture. Figure 2 A diagram illustrating multi-cloud deployment and sandbox isolation mechanisms; Figure 3 Here is a flowchart of the sample matching and feature backtracking process; Figure 4 A flowchart for the reliable acquisition and verification of physical world data; Figure 5 A flowchart for data consistency verification and auditing in the digital world; Figure 6 A flowchart for corporate credit risk decision-making methods throughout the entire life cycle; Detailed Implementation The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, and not all of them.

[0036] Combination Figures 1-6 As shown, this invention provides an enterprise full lifecycle credit risk decision-making system based on multi-source data, comprising: The data aggregation module is used to collect and aggregate heterogeneous information from multiple sources, such as government data, industry data, commercial data and financial data, based on a distributed data processing architecture. It also creates profiles of enterprises according to preset feature dimensions, generates enterprise profile variables, and supports the retrospective analysis of enterprise credit risk feature data. The visual navigation module is used to respond to user operation commands after user login, call system function interfaces, and generate a step-by-step operation guidance interface. The variable overview module is used to provide a global overview of the data coverage of the enterprise profile variables according to the month dimension; The variable exploration module is used to perform distribution exploration and quantile statistics display on the target variable to obtain the coverage and value distribution characteristics of the target variable in different time periods. The variable filtering module is used to filter from the enterprise profile variables to form a set of target variables based on filtering conditions such as dimension, coverage and data type. The sample matching module is used to receive a sample table containing the unified social credit code and the retrospective date, and perform table creation and database storage and feature retrospective processing on the target variable set to generate a sample feature table. The sandbox modeling module is used to perform model training, parameter tuning, and model validation on the sample feature table in an isolated modeling environment based on a visual interactive interface, to obtain a credit risk scoring model and output a risk score for risk decision-making and post-loan management. The table lifecycle management module is used to manage the lifecycle of the sample feature table and perform dependency verification and state locking when a deletion instruction is received. The model management module is used to centrally collect and manage the sample data, intermediate processing data, variable feature data, parameter configuration data, iteration process data, and model output results generated during the modeling process. The system management module is used to manage accounts, roles, and permissions throughout their entire lifecycle, and to manage the isolated configuration of sandbox accounts.

[0037] It also includes the government cloud modeling platform and the state-owned assets cloud modeling platform. The two are respectively configured with data exchange, model training, model deployment and account partitioning functions, and interact with the government cloud database and the state-owned assets cloud database respectively to realize cross-data domain data exchange and domain-based management of modeling resources. The sandbox modeling module is a Python-based Jupyter visual interface with built-in supervised and unsupervised algorithm libraries and supports custom algorithm programming. It also supports joint modeling of financial institutions' internal labels and the enterprise profile variables. An isolation mechanism is adopted that binds accounts to database permissions one-to-one, so that each user account can only access the specified database and table resources bound to it; The system management module assigns an independent Python runtime partition to each sandbox account and allows the installation of algorithm packages within the runtime partition. It also records account operations, role configurations, and permission change logs for auditing and traceability. The model management module performs regular backups of the collected modeling data and model results.

[0038] The variable overview module filters variable coverage by month based on nine feature dimensions and uses a spectral heatmap to display the percentage distribution of data coverage in each month. The variable exploration module supports both hierarchical drill-down filtering based on first-level, second-level, and third-level dimensions, as well as precise filtering based on variable names. After filtering, it displays the latest monthly non-empty sample size distribution, data coverage, and quantile statistics of the target variable arranged in monthly order.

[0039] The sample matching module includes sub-modules for variable upload management, sample upload matching, and backtracking status monitoring, and satisfies the following: Sample uploads only support sample tables containing "Unified Social Credit Code" and "Backtrack Date", and complete duplicate samples with the same Unified Social Credit Code and Backtrack Date will be checked and blocked. Automatic splitting is performed during table creation, database storage, and feature backtracking to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3. When a memory overflow error is detected during the table creation process, a one-click rerun mechanism is provided, which automatically breaks down the original table creation statement into multiple sub-table creation statements and re-executes them.

[0040] It also includes a remote sensing feature extraction module, an industrial IoT feature extraction module, and a trusted acquisition and remote verification module, and meets the requirements of... Constraints and processing rules: (1) The remote sensing feature extraction module is used to: determine the remote sensing region of interest (ROI) centered on the geographical coordinates obtained by parsing the business address or factory address of the enterprise entity, and rasterize the ROI with a preset spatial grid; perform monthly aggregation feature generation on the remote sensing image covering the ROI, and the generated features include at least the monthly average value of nighttime light intensity, the month-on-month change rate of nighttime light intensity, and the change rate of building outline; wherein, the grid side length of the spatial grid is 100m to 250m, and when the effective pixel ratio of the remote sensing image in the ROI is less than 0.6, the remote sensing features of the corresponding month are set to missing and a missing indicator variable is generated; (2) The industrial IoT feature extraction module is used to: collect industrial operation signals, including at least the power load time series and the key equipment vibration time series, at the enterprise-side edge gateway, and perform windowed feature processing on the industrial operation signals at the edge gateway to obtain daily industrial operation features; wherein, the sampling period of the power load time series is 1min to 5min, the sampling frequency of the key equipment vibration time series is 1kHz to 10kHz, the window length of the windowed feature processing is 10s to 60s and at least outputs frequency domain energy, spectral entropy, kurtosis and outlier density; and upload the daily industrial operation features to the data aggregation module on a daily basis, and aggregate them into monthly industrial operation features on a monthly basis; (3) The trusted acquisition and remote verification module is used to: generate trusted verification information for remote sensing features and industrial operation day features. The trusted verification information includes at least the device identity identifier, firmware metric hash, feature data digest hash, and a remote verification token formed by a timestamped signature. Before writing the remote sensing features and the industrial operation day features into the data storage, the remote verification token is verified. It is allowed to be stored only when the verification is successful. Otherwise, it is rejected and an audit log is recorded. The audit log is stored in a chain hash structure, and each audit record includes at least the previous hash, the current record digest hash, and a timestamp. (4) The remote sensing features, the monthly industrial operation features, and the enterprise profile variables are aligned using a unified retrospective approach: the retrospective dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned remote sensing features and monthly industrial operation features are included as new enterprise profile variables in the enterprise profile variables, so that the variable screening module can screen them and the sample matching module can retrospectively generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

[0041] It also includes a network security attack surface characterization module, a supply chain logistics flow characterization module, and a decision audit and evidence retention module, and meets the following constraints and processing rules: (1) The network security attack surface feature module is used to: periodically acquire the public network assets of the enterprise entity and generate attack surface features. The public network assets include at least domain names, subdomains, IP addresses and externally open ports. The mapping between the public network assets and the enterprise entity adopts at least two types of evidence consistency verification. The two types of evidence include at least two of the following: ICP filing information, digital certificate organization information, WHOIS query information and subdomain resolution relationship. The attribution confidence score is calculated for each mapping result. The corresponding network asset is included in the attack surface calculation only when the attribution confidence score is not lower than a preset threshold. The attack surface features include at least the number of externally open ports, the distribution of external service types, the number of high-risk vulnerabilities, the statistics of vulnerability repair time, and the number of suspected counterfeit / phishing domain names and the number of certificate anomalies. The periodic acquisition frequency is once every 7 to 30 days, and the results are aggregated into monthly network security features. (2) The supply chain logistics flow feature module is used to: acquire logistics and transportation data associated with the enterprise entity and generate supply chain flow features. The logistics and transportation data includes at least one of the following: maritime automatic identification system (AIS) trajectory data, waybill trajectory data, and vehicle diagnostic (OBD) data. The association mapping between the enterprise entity and the logistics and transportation data is based on at least two of the following for consistency verification: unified social credit code, enterprise name, transportation vehicle registration information, carrier contract number, or waybill number. The supply chain flow features include at least the monthly number of shipments, monthly number of arrivals, monthly abnormal delay rate, number of berthings, port stay duration, and regional dispersion of shipments and receipts. The regional dispersion of shipments and receipts is characterized by regional entropy or equivalent indicators, and the results are aggregated monthly into monthly supply chain flow features. (3) The decision audit and evidence retention module is used to: perform traceable audit and retention of the generation process, version configuration and storage process of the network security monthly features and the supply chain flow monthly features, as well as the process of participating in modeling and risk decision-making based on these two types of features. The audit retention includes at least the data source identifier, feature generation configuration version number, variable dictionary version number, sample backtracking caliber identifier, model version number, parameter configuration summary and risk decision output result; the audit retention adopts an immutable storage strategy, including at least one of chained hash log or WORM write-once-read-many storage, each audit record includes at least the previous hash, current record summary hash and timestamp, and forms alarm records for abnormal events; (4) The network security monthly features, the supply chain circulation monthly features, and the enterprise profile variables are aligned using a unified backtracking caliber: the backtracking dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned network security monthly features and supply chain circulation monthly features are included as new enterprise profile variables in the enterprise profile variables, so that the variable filtering module can filter them and the sample matching module can backtrack to generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

[0042] This invention also provides a method for enterprise full lifecycle credit risk decision-making based on multi-source data, which, based on the above system, includes: S1. Collect and aggregate heterogeneous information from multiple sources, including government data, industry data, commercial data, and financial data. Based on a distributed data processing architecture, generate enterprise profile variables according to nine feature dimensions to form enterprise credit risk feature data that supports the longest historical growth trajectory of up to nine years. S2. After the user logs in, a step-by-step operation guide and core function entry points are provided through a visual navigation module; S3. The variable overview module displays the data coverage of enterprise profile variables by month, and the variable exploration module performs distribution exploration and quantile statistics on the target variable to obtain variable coverage and distribution characteristics. S4. The target variable set is formed by filtering based on conditions such as dimension, coverage and data type through the variable filtering module; S5. Receive a sample table containing the unified social credit code and the retrospective date, perform table creation and database storage and feature retrospective for the target variable set, generate a sample feature table, and verify and intercept completely duplicate samples. S6. In the sandbox modeling module, perform model training, parameter tuning and model validation on the sample feature table to obtain the credit risk scoring model and output the enterprise risk score; S7. Collect the modeling process data, parameter configurations and model output results into the model management module for unified management, and perform table lifecycle management and model retention management through the table lifecycle management module and the model management module according to the retention strategy. S8. Make risk decisions based on the enterprise risk score, and implement risk warnings based on changes in enterprise characteristic data in post-loan management.

[0043] In step S3: the variable overview module uses a spectral heatmap to display the distribution of data coverage percentage for each month; the variable exploration module uses a combined graph to display the non-empty sample size and data coverage, and outputs the quantile statistics of the target variable in monthly order.

[0044] In step S5: when the number of variables to be processed exceeds the preset threshold, multiple target tables are automatically generated to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3; when a memory overflow error occurs during the table creation process, the original table creation statement is automatically disassembled and multiple sub-table creation statements are generated and re-executed, thereby completing feature backtracking.

[0045] In steps S6 and S7: a one-to-one binding mechanism between account and database permissions is adopted, so that users can only access the database and table resources bound to their accounts; an independent Python running partition is allocated to each sandbox account to achieve user isolation, and account operation, role configuration and permission change logs are recorded; whether to retain the sample feature table and credit risk scoring model is determined according to the preset retention strategy. When not retained, the corresponding table resources are deleted through the table lifecycle management module. When retained, the model resources are collected and backed up regularly through the model management module.

[0046] Example 1: Enterprise Credit Risk Feature Extraction, Modeling, and Risk Decision Making across the Entire Enterprise Lifecycle using an Enterprise Credit Decision Feature Platform I. System Deployment and Overall Structure In this embodiment, an "Enterprise Credit Decision-Making Feature Platform" is constructed, which provides financial institutions and credit reporting agencies with capabilities for enterprise credit risk feature management, sample backtesting, sandbox modeling, and risk decision-making. The platform relies on a distributed Hive / data warehouse to aggregate multi-source heterogeneous data and derives over 10,000 variables according to nine feature dimensions, covering approximately 4.5 million business entities and supporting backtesting of up to nine years of historical growth.

[0047] 1. User layer: including bank users and credit reporting agency users; 2. Business Layer: Provides two types of business entry points: joint modeling and automatic modeling (joint modeling is used to integrate "in-house labels + platform features", while automatic modeling is used for self-training and parameter tuning of the platform's built-in algorithms). 3. Application Layer: It should include at least a visual navigation module, a variable overview module, a variable exploration module, a variable filtering module, a sample matching module, a sandbox modeling module, and a table lifecycle management module. It can also be expanded to include a model management module and a system management module for resource aggregation and access control.

[0048] 4. Data Layer: This includes the government cloud data warehouse and the state-owned assets cloud data warehouse. Each data domain has basic support capabilities such as data exchange, variable dictionary / stored procedure / task scheduling / data storage; and provides data reading and backtracking calculation capabilities to the upper layer of the platform through a controlled data exchange channel.

[0049] To ensure secure isolation, this embodiment adopts a "one-to-one binding of account and database permissions" mechanism: each user account is only authorized to access the database and table resources bound to it; the sandbox modeling environment (Jupyter) only allows access to the bound database, preventing cross-account data access from the root.

[0050] II. Data Model and Key Parameter Conventions In this embodiment, to standardize the system processing logic, the following unified conventions are made for core data objects: 1. Enterprise entity identifier: Unified Social Credit Code (USCC), string length 18, serving as the unique key for the enterprise entity; 2. Backtrack date: backtrack_date, date type; when the platform organizes feature data by "monthly snapshot", backtrack_date is mapped to the statistical caliber at the end of the corresponding month (e.g., 2026-01-15 is mapped to 2026-01-31). The mapping rules are fixed in the platform configuration items. 3. Variables (Features): Defined by a variable dictionary, with fields including at least: var_code, var_name, dim_l1 / dim_l2 / dim_l3, data_type, agg_rule, coverage_rule, update_cycle, null_fill_rule, source_table, and source_field. Administrators can upload / update the variable dictionary in batches, while users can only filter and download.

[0051] 4. Nine primary dimensions: basic information, operating costs, operating results, operating risks, operating behavior, human resource profile, financing behavior, asset status, and qualification innovation; in this embodiment, all variables belong to one of the above.

[0052] 5. Sample table: Contains only two core fields, uscc and backtrack_date; completely duplicate records are not allowed; the matching scope is limited to commercial entities in Shenzhen.

[0053] III. Implementation Process: Feature Exploration → Variable Selection → Sample Backtracking → Modeling → Risk Decision Making Step M1: Multi-source data aggregation and dimensional variable generation The platform integrates at least three types of data in a distributed Hive / data warehouse: Government data: such as business registration, administrative penalties, water and electricity consumption, social security, housing provident fund, etc.; Business data: such as intellectual property, investment and financing, customs, judiciary, audited financial statements, etc.; Financial data: such as credit history, credit inquiries, and offline visits.

[0054] In this embodiment, the data access scale is as follows: the government side comes from 37 commissions and bureaus, the commercial side comes from 27 cooperative data providers, and it is integrated with credit reporting and financial data to form a basic full profile of enterprises.

[0055] The platform runs ETL / feature derivation tasks monthly (triggered by the task scheduling module), generating a "nine-dimensional-over-10,000-dimensional" feature set for each enterprise entity in each statistical month, and storing the results in a "long table + dictionary constraint" format.

[0056] Step M2: Variable Overview and Variable Exploration After logging in, users can enter the cockpit and view the data coverage by nine primary dimensions and by month through the variable overview module. The coverage distribution is displayed using spectral visualization.

[0057] Then, in the variable exploration module, drill down by dimension or search by variable name to view the non-empty sample size and coverage of the variable in the latest month, and view the quantile statistics and coverage statistics for the past year to confirm whether the variable meets the modeling caliber and stability requirements.

[0058] Users explored variables such as "number of administrative penalties in the past 12 months", "number of abnormal fluctuations in electricity consumption in the past 6 months", and "number of patents granted in the past 24 months" to confirm coverage and distribution characteristics before deciding to include them in the candidate variable set.

[0059] Step M3: Variable Filtering and Variable Upload Management Users can filter target variables by dimension, coverage, data type, and other criteria in the variable filtering module and download them to their local machine for review. Administrators can upload and update variable dictionaries in batches or edit single variable information in the variable dictionary management interface to ensure consistency and traceability.

[0060] In this embodiment, the user forms a candidate variable set V, for example, |V|=1200 (which exceeds the capacity of a single table and will trigger automatic splitting later).

[0061] Step M4: Sample Upload and Sample Matching (Generate a wide feature table on the specified backtracking date) 1. Sample Upload: Users download the sample template and fill in sample table S, where each row contains only uscc and backtrack_date; Platform Verification: whether it is a Shenzhen business entity, whether the date format is compliant, and whether there are completely duplicate rows.

[0062] 2. Start matching and database creation: The user enters a custom table name (e.g., tb_case_202601_bankA) and clicks "Start Matching". The platform performs the following actions: Create a sample temporary storage table tb_case_202601_bankA_sample(uscc,backtrack_date) in the database bound to the user account and write the sample to it; Based on the variable dictionary, the candidate variable set V is mapped to the data retrieval path of the feature long table, and backtracking SQL / stored procedure call parameters are generated; Convert “Sample × Variable × Backtracking Month” into a wide feature table of “Sample × Variable” required for modeling (use var_code for wide table field names to avoid name conflicts).

[0063] 3. Automatic Splitting Mechanism (Key Stability Parameter): To avoid memory overflow during table creation, the platform automatically splits the target table when there are too many variables to be processed: ensuring that the number of variables in a single table does not exceed 500 and limiting the number of associated sub-tables to no more than 3.

[0064] In this embodiment, |V|=1200, therefore the platform partitions the variable set into V1 (≤500), V2 (≤500), and V3 (≤500), generating three feature sub-tables respectively: tb_case_202601_bankA_fea_1(uscc,backtrack_date,V1...) tb_case_202601_bankA_fea_2(uscc,backtrack_date,V2...) tb_case_202601_bankA_fea_3(uscc,backtrack_date,V3...) During the training phase, the three sub-tables are joined by uscc and backtrack_date to form a logical wide table view vw_case_202601_bankA_feature, thereby making over 1,000-dimensional features available without exceeding resource thresholds.

[0065] 4. Status Monitoring and Anomaly Handling: Users can refresh the status during the matching process; if an error occurs, the cause of the error can be viewed. For common "memory overflow" errors, the platform provides "one-click rerun," which automatically breaks down the original erroneous table creation statement into multiple sub-table creation statements and re-executes them without manual intervention.

[0066] Step M5: Sandbox Modeling (Jupyter) and Co-modeling Examples The platform provides a Python-based Jupyter visual sandbox modeling environment, with dozens of built-in supervised / unsupervised algorithms and support for custom algorithm programming.

[0067] Meanwhile, sandbox access permissions are constrained by a "one-to-one account-database binding": Jupyter can only access tables such as vw_case_202601_bankA_feature in the database bound to that account.

[0068] This embodiment provides a workflow example for "joint modeling" (fusion of bank internal tags and platform features): 1. Prepare labels (internal data): The bank prepares a label table tb_label_bankA(uscc,label_date,y) in its bound database; where y∈{0,1} indicates whether a default / non-performing loan has occurred within T days after label_date (T is configured by the bank, such as 90 days). The label generation criteria are fixed and archived by the bank's internal rules. 2. Constructing the training set: In Jupyter, execute SQL to align vw_case_202601_bankA_feature and tb_label_bankA according to USCC and date mapping rules to obtain the training set D: If label_date is a daily level, first map it to the corresponding month-end snapshot and then connect it (the mapping rule is the same as step M1). 3. Missing value handling: For numeric variables, the null_fill_rule from the variable dictionary is applied (e.g., 0-filling / median-filling / missing indicator variables); for categorical variables, one-hot encoding is performed. 4. Model training and parameter tuning: For example, use logistic regression, XGBoost, or random forest (choose one or a combination), employing five-fold cross-validation and grid search; output metrics such as AUC, KS, and recall. 5. Risk Decision Output: The decision result is obtained by comparing the risk score (ranging from 0 to 1) output by the model with the policy threshold θ. When the score is greater than or equal to the threshold θ, at least one of the following actions is taken: rejection, increased guarantee, or manual review. score<θ: Credit limit calculation via / entry; The threshold θ is configured and archived by the strategy formulation interface to ensure consistency between pre-loan decisions and post-loan monitoring.

[0069] Step M6: Model / Table Resource Management and Traceable Retention (Administrator Side) To facilitate compliance audits and model iterations, administrators can use the model management module to collect data and results from the entire modeling process (original samples, intermediate data, variable feature data, parameter configurations, model files, evaluation metrics, etc.), and support retrieval and archiving by time / model type / business scenario. At the same time, resources can be backed up regularly.

[0070] For historical tables, redundant tables, or test tables that are no longer needed, users can initiate deletion in the table lifecycle management module and require secondary confirmation to keep the database clean and reduce storage usage.

[0071] Example 2: Trusted Feature Enhancement and Joint Modeling Based on Remote Sensing + Industrial Internet of Things + Trusted Hardware I. System Structure and New Module Composition In this second embodiment, in addition to the modules in the first embodiment, the enterprise credit decision feature platform adds three new modules and completes the data entry and alignment with the data storage subsystem through the data aggregation module.

[0072] 1. Remote sensing feature extraction module Input: Business address or factory address (from business registration / government data), and remote sensing image data (nighttime light image, visible light / multispectral image).

[0073] Output: A table of remote sensing features aggregated by month, tb_rs_feature_monthly.

[0074] The rules for generating the Region of Interest (ROI) are as follows: Address resolution yields the coordinates (lat, lon) of the company's factory center point; A rectangular ROI of 2km×2km is constructed around this center point (configurable from 1km×1km to 5km×5km; in this embodiment, it is fixed at 2km×2km). The ROI is rasterized using a spatial grid with a grid side length of 200m, so each ROI is divided into 10×10=100 grid cells.

[0075] 2. Industrial IoT Feature Extraction Module Input: Industrial operation signals collected by the enterprise-side edge gateway, including: power load time series and vibration time series of key equipment; Output: Daily industrial operation feature table tb_iot_feature_daily, which is then aggregated by the platform into a monthly industrial operation feature table tb_iot_feature_monthly.

[0076] 3. Trusted data collection and remote verification module This module acquires remote sensing feature data, daily industrial operation feature data, and reliable measurement information of edge devices; This module generates the monthly feature table tb_sc_feature_monthly for supply chain circulation.

[0077] This module verifies the trustworthiness of the data source before writing to the data storage subsystem. It only allows data to be written into the database if the verification passes; otherwise, it rejects the data and records the audit log.

[0078] II. Key Parameters and Data Standards The following parameters are fixed in this embodiment: 1. Month-end snapshot mapping rules: Map any `backtrack_date` to the last day of the same month, `month_end_date`. For example: 2026-01-15→2026-01-31; 2026-02-02→2026-02-28 (non-leap year); Map the backtrack_date to the end date of the month it belongs to, month_end_date.

[0079] 2. Threshold for effective pixel ratio in remote sensing: valid_pixel_ratio_threshold=0.6.

[0080] If the proportion of valid pixels within a ROI is less than 0.6 in a given month, the remote sensing features for that month are set to missing, and a missing indicator variable rs_missing_flag=1 is generated; otherwise, rs_missing_flag=0.

[0081] 3. Sampling and Windowing Parameters for Industrial IoT: Electricity load sampling period: 1 minute (P_load=1min); Vibration sampling frequency: 5kHz (meets the range of 1kHz~10kHz; F_vib=5000Hz); Windowing: Window length 30 seconds, sliding step 30 seconds (no overlap; W=30s, step=30s); Daily feature upload frequency: once a day (upload_daily=1 / day), and the uploaded content is the feature vector of the day.

[0082] 4. Trusted Proof and Signature Parameters: The edge gateway has a TPM or equivalent trusted module, and the device identity is device_id; Firmware metric hash algorithm: SHA-256; Feature data digest hash algorithm: SHA-256; Signature algorithm: ECDSA (curve P-256); The remote proof token field is fixed as follows: device_id,firmware_hash,data_hash,ts,signature.

[0083] Trusted verification strategy: Verify that the signature is valid, the firmware_hash is in the whitelist, the time deviation is within the allowed time deviation range (±10 minutes), and the data_hash is consistent with the feature data to be added to the database.

[0084] 5. Audit chained hash log fields: Each audit record always contains: prev_hash,record_hash,ts,event_type,payload_hash, where record_hash=SHA256(prev_hash||payload_hash||ts||event_type); The first record, prev_hash, is always set to all zeros (a 64-bit hexadecimal string).

[0085] III. Implementation Steps (Complete Closed Loop: Generation → Validation → Storage → Alignment → Modeling) The following steps are illustrated using "single company, backtracking across three months" as an example. The enterprise entity is uscc=9144XXXXXXXXXXXXXX, and the sample backtracking date set is {2025-11-20, 2025-12-15, 2026-01-10}. The mapped month-end snapshot is {2025-11-30, 2025-12-31, 2026-01-31}.

[0086] Step S21: Remote Sensing Feature Extraction (by Month) 1. ROI Generation and Rasterization The platform obtains (lat,lon) from the company's registered address, constructs a 2km×2km ROI with a grid side length of 200m, and obtains 100 grid cells.

[0087] 2. Generation of Nighttime Light Features For the nighttime light images corresponding to the month captured at the end of each month, calculate within the ROI: ntl_mean_m: Monthly average of nighttime light intensity within the ROI; ntl_mom_m: Rate of change of nighttime light intensity, in accordance with... ntl_mom_m=(ntl_mean_m-ntl_mean_{m-1}) / max(ntl_mean_{m-1},ε), The value of ε = 0.01 is used to prevent instability caused by a denominator of 0.

[0088] 3. Generation of Building Profile Change Rate Building segmentation is performed on visible / multispectral imagery within the ROI to obtain the monthly building cell set B_m; calculation: bld_change_m=|B_m△B_{m-1}| / max(|B_{m-1}|,1), Where △ represents the difference in symmetry, and |·| represents the number of pixels.

[0089] 4. Determination of effective pixel ratio and missing pixel marker Calculate valid_pixel_ratio_m=valid_pixels_m / total_pixels_m.

[0090] If valid_pixel_ratio_m < 0.6: set ntl_mean_m, ntl_mom_m, and bld_change_m as missing, and set rs_missing_flag = 1; Otherwise, rs_missing_flag=0 and the feature value is retained.

[0091] Generate monthly feature records and write them to the candidate table (not yet stored in the database): rs_feature_m={uscc,month_end_date,ntl_mean_m,ntl_mom_m,bld_change_m,rs_missing_flag}.

[0092] Step S22: Industrial IoT Feature Extraction (Daily Features from the Edge Side → Monthly Aggregation from the Platform) 1. Signal Acquisition The edge gateway collects the power load load(t) every minute; the vibration of critical equipment is sampled at 5 kHz to collect vib(t).

[0093] 2. Windowed Feature Generation (Edge Side) The vibration signal is processed in 30-second windows, and the following values ​​are calculated for each window: E_f: Frequency domain energy (sum of squared amplitudes after FFT); SpecEntropy: Spectral entropy; Kurtosis: kurtosis; AnomDensity: Outlier density (MAD threshold determination, threshold = 3 × MAD).

[0094] Summarize all window statistics into the characteristics of the day (e.g., take the mean / maximum / 95th percentile): vib_energy_p95_d,vib_entropy_mean_d,vib_kurtosis_max_d,vib_anom_rate_d.

[0095] Daily characteristics of electricity load calculation: load_peak_val_d (daily peak value), load_valley_val_d (daily valley value), load_peak_valley_ratio_d = load_peak_val_d / max(load_valley_val_d, 0.1); load_drop_count_d: When load(t) drops by more than 30% within 5 minutes, it is counted as a sudden drop, and the number of drops within the day is counted.

[0096] Daily characteristic record formation: iot_feature_d={uscc,day_date,load_peak_valley_ratio_d,load_drop_count_d,vib_energy_p95_d,vib_entropy_mean_d,vib_kurtosis_max_d,vib_anom_rate_d}.

[0097] 3. Daily Feature Upload and Platform Monthly Aggregation The edge gateway uploads its iot_feature_d data to the platform daily. The platform aggregates this data monthly. For logarithmic values, retrieve the monthly average or maximum value (with a fixed aggregation rule, agg_rule, in the variable dictionary). For count-type data, calculate the monthly cumulative total; For example: load_drop_count_m = Σload_drop_count_d (monthly cumulative) vib_anom_rate_m=mean(vib_anom_rate_d) (monthly average).

[0098] Generate monthly feature records: iot_feature_m={uscc,month_end_date,load_peak_valley_ratio_m,load_drop_count_m,vib_energy_p95_m,vib_entropy_mean_m,vib_kurtosis_max_m,vib_anom_rate_m}.

[0099] Step S23: Trusted data collection and remote verification (verification before data entry) 1. Generate remote proof token The edge gateway generates the following for each daily feature packet and each monthly remote sensing feature packet: firmware_hash=SHA256(firmware_binary); data_hash=SHA256(serialize(feature_payload)); ts is the generated timestamp; signature=ECDSA_sign(private_key,firmware_hash||data_hash||ts); form attestation_token={device_id,firmware_hash,data_hash,ts,signature}.

[0100] 2. Platform-side verification Before writing data to the government cloud database or state-owned assets cloud database, the platform performs the following steps on the attestation_token: Signature verification passed; firmware_hash is in the whitelist; The time difference between ts and the current platform time shall not exceed ±10 minutes; The data_hash is consistent with the feature data to be entered into the database; Writes to the data storage subsystem are allowed only if all four conditions are met; otherwise, writes are rejected.

[0101] 3. Chained hash audit log Regardless of whether the data entry is approved or rejected, the platform records an audit log: payload_hash=SHA256(serialize({uscc,month_end_date,feature_payload,token})); Calculate the record_hash and write it to the tb_audit_hashlog; The event_type value can be one of ACCEPT_RS, REJECT_RS, ACCEPT_IOT, or REJECT_IOT.

[0102] Step S24: Align with existing enterprise profile variables and incorporate them into sample backtesting. The platform will verify the data entered into the database. tb_rs_feature_monthly and tb_iot_feature_monthly The button (uscc, month_end_date) is aligned with the platform's existing enterprise profile variable table to form an expanded set of enterprise profile variables: FeatureSet_ext = FeatureSet_base ∪ {rs features, iot features}.

[0103] Subsequently, the sample table is backtracked in the sample matching module to output the sample feature table (still following the splitting rule of single table ≤ 500 variables and sub-table ≤ 3; if the number of variables exceeds the threshold, the table will be automatically split without changing the variable definition).

[0104] Step S25: Sandbox Joint Modeling and Risk Decision Output Load the sample feature table into the Jupyter sandbox and add inline labels y for joint modeling: The training set is divided by time: 2025-11 and 2025-12 are for training, and 2026-01 is for validation; The model can be either logistic regression or gradient boosting tree; Evaluation metrics: AUC, KS, recall; Output the risk score ∈ [0,1] and the threshold decision result (the threshold θ is configured by the strategy and written to the audit log for traceability).

[0105] Example 3: Enhanced External Risk Characteristics and Traceable Risk Decision-Making Based on Cybersecurity + Supply Chain Logistics + Decision Audit Retention I. System Structure and New Module Composition In this embodiment 3, in addition to the modules in embodiment 1, the enterprise credit decision feature platform adds three new modules, all of which enter the data storage subsystem through the data aggregation module, and the decision audit and evidence retention module leaves traces of key processes.

[0106] 1. Network security attack surface signature module Input: Public network asset information of the enterprise entity (domain name, subdomain name, IP address, externally accessible ports, service fingerprint, vulnerability information, certificate information, clues of suspected counterfeit domain names).

[0107] Output: Monthly cybersecurity feature table tb_cyber_feature_monthly.

[0108] 2. Supply Chain Logistics Flow Characteristics Module Input: Logistics and transportation data associated with the enterprise entity (at least one of the following: ocean shipping AIS trajectory, waybill trajectory, and vehicle networking OBD data).

[0109] Output: Monthly feature table of supply chain flow tb_sc_feature_monthly.

[0110] 3. Decision Audit and Evidence Retention Module Inputs: Feature generation process, configuration version, variable dictionary version, sample backtracking criteria, model version and parameter summary, risk decision output.

[0111] Output: Unalterable audit record tb_decision_audit_immut, and abnormal alarm record tb_audit_alert.

[0112] II. Key Parameters and Data Definitions The following parameters are fixed settings: 1. Month-end snapshot mapping rules: Map the backtrack_date to the end date of the month in which it is located, for example, map 2026-01-10 to 2026-01-31.

[0113] 2. Confidence threshold for network asset ownership: conf_threshold=0.70.

[0114] Network assets are only included in the attack surface calculation when the confidence score of the enterprise entity and the network asset attribution is ≥0.70.

[0115] 3. Types and weights of evidence for ownership of online assets (used to calculate confidence levels): At least two types of attribution evidence must be consistent for confidence to be calculated. This embodiment uses four types of evidence and assigns weights accordingly: ICP filing entities are consistent: weight 0.35 Consistent digital certificate organization information (certificate O field): Weight 0.25 WHOIS registered organization consistency: weight 0.20 Subdomain DNS resolution / reverse DNS lookup is consistent with the corporate domain: weight 0.20 Confidence score calculation: conf_score=Σ(weight_imatch_i), where match_i∈{0,1} indicates whether the evidence is consistent.

[0116] 4. Network mapping and vulnerability scanning frequency: once every 14 days (within the range of 7 to 30 days), with results aggregated monthly.

[0117] 5. Definition of vulnerability severity and remediation time: Vulnerability severity is defined as severity∈{low,medium,high,critical}; high-risk vulnerabilities are defined as severity∈{high,critical}.

[0118] The fix duration `fix_days` is defined as the number of days between the first detection of the same vulnerability in the same asset and the period after which the vulnerability is not detected in two consecutive scans.

[0119] 6. Supply chain data consistency rules: The association between the enterprise entity and logistics / transportation data is based on at least two consistency checks of the following fields: USCC, Company Name (standard name for business registration), Waybill Number / Contract Number, Vehicle Registration Information (IMO / MMSI for ships or VIN / License Plate for vehicles).

[0120] When using "company name" for matching, the name standardization rules are fixed as follows: remove company type suffixes (such as "Limited Company"), unify full-width and half-width characters, unify uppercase and lowercase letters, and unify spaces and punctuation.

[0121] 7. Monthly Statistical Criteria for Supply Chain Characteristics: monthly shipment count (ship_cnt_m): the number of shipment node records in the current month's waybills; Monthly delivery count recv_cnt_m: The number of signed delivery records in the current month's waybills; Monthly abnormal delay rate delay_rate_m: Number of delayed shipments in the current month / Total number of shipments in the current month. The delay judgment threshold is fixed as "scheduled arrival time + 24 hours without being signed for". Number of berthing calls (port_call_cnt_m) and length of stay in port (port_stay_hours_m): extracted from AIS berthing events and accumulated monthly; The geographical dispersion of shipping and receiving locations, geo_entropy_m, is characterized by geographical entropy and is calculated using the formula: geo_entropy_m=-Σ(p_rln(p_r)), where p_r is the percentage of the receiving / shipping region r in the current month (provincial or municipal level can be configured, this embodiment is fixed as municipal level).

[0122] 8. Immutable audit retention strategy: adopts chained hash log (which can be stored as WORM storage or equivalent immutable storage).

[0123] Each audit record has the following fixed fields: prev_hash,record_hash,ts,audit_type,config_version,dict_version,mapping_rule_id,model_version,param_digest,decision_digest; in record_hash=SHA256(prev_hash||ts||audit_type||config_version||dict_version||mapping_rule_id||model_version||param_digest||decision_digest); The first record has all zeros in its prev_hash.

[0124] III. Implementation Steps (Complete Closed Loop: Attribution → Feature → Aggregation → Alignment → Modeling → Audit Retention) The following steps take a single financial institution user conducting a post-loan early warning for a certain enterprise during the period from December 2025 to January 2026 as an example. The enterprise entity is uscc=9144YYYYYYYYYYYYYY, the backtracking date set is {2025-12-18, 2026-01-12}, and the month-end snapshot after mapping is {2025-12-31, 2026-01-31}.

[0125] Step S31: Public Network Asset Acquisition and Ownership Mapping 1. Asset Acquisition A network mapping and vulnerability scan is performed every 14 days to collect relevant domain names, subdomains, IPs, ports, service fingerprints, certificate information, and vulnerability lists for the enterprise, forming the raw asset record asset_raw_k.

[0126] 2. Evidence extraction and consistency verification of attribution Four types of evidence are extracted from each asset record: ICP filing entity, certificate O field, WHOIS registered organization, and parsing relationship; Extract the registration standard name and uscc from the enterprise entity, and obtain org_norm according to the name standardization rules; Calculate match_icp, match_cert, match_whois, and match_dns ∈ {0,1} respectively, and then calculate... conf_score=0.35match_icp+0.25match_cert+0.20match_whois+0.20match_dns.

[0127] Only if: If at least two types of evidence match_i=1 and conf_score≥0.70, the asset record is attributed to the enterprise entity and denoted as asset_owned_k; otherwise, it is discarded or placed in the "awaiting manual verification" queue.

[0128] 3. Asset deduplication and fingerprint merging The assets attributable to the user are merged and deduplicated according to {domain / ip,port,service_fingerprint} to form the set of assets attributable to the user in this scan, A_k.

[0129] Step S32: Cybersecurity Month Feature Generation and Aggregation Aggregate all scan results {A_k} within the same month using the month-end snapshot to generate a monthly cybersecurity feature record: open_port_cnt_m = the number of externally open ports in the assets attributable to the current month; service_type_dist_m = Service type distribution (stored as a count vector or expanded into several service count variables); high_vuln_cnt_m = Number of high-risk vulnerabilities in the current month (high + critical); fix_days_mean_m = the average time it takes to fix vulnerabilities that have been fixed in the current month; phish_domain_cnt_m = Number of suspected phishing / spoofing domains in the current month (counted by similar domain rules or blacklist hits); cert_anom_cnt_m = Number of certificate anomalies in the current month (count of rule hits such as short-term certificates / frequent replacements).

[0130] Write tb_cyber_feature_monthly(uscc,month_end_date,...).

[0131] Step S33: Supply Chain Logistics Data Collection and Enterprise Association (Supply Chain Logistics Flow Characteristics Module) 1. Data Collection Obtain at least one type of logistics and transportation data: Maritime AIS: Includes MMSI / IMO, timestamp, location, speed, and berthing events; Waybill tracking: Includes waybill number, shipping node, receipt node, planned arrival time, actual arrival time, and node location; Vehicle-to-Everything (OBD) system includes vehicle VIN / license plate information, mileage, fault codes, and active duration.

[0132] 2. Enterprise Association Mapping The association is completed using the "two-way consistency check" rule: If the waybill data contains a contract number and the contract number matches the enterprise's purchase / sales contract database, and the enterprise name is standardized and consistent, then the association is successful; If the AIS data can be linked to a company from the ship registration information (e.g., long-term lease registration of the vessel), and the company names are consistent, then the linking is successful; If the OBD data can be linked to the company's fleet list from the vehicle registration VIN / license plate, and the USCC matches, then the linking is successful.

[0133] Records that fail to meet the two consistency checks are not included in the feature statistics.

[0134] 3. Data cleaning and deduplication Deduplicat waybills by waybill number + node type + timestamp; deduplicat AIS berthing events by port + arrival time; aggregate OBD records by vehicle + date into daily records.

[0135] Step S34: Generation and aggregation of monthly characteristics of supply chain flow Monthly characteristic records of supply chain flow are obtained by aggregating the end-of-month snapshots: ship_cnt_m: Number of shipments per month; recv_cnt_m: Number of deliveries per month; delay_rate_m: Monthly abnormal delay rate, with a delay threshold of "scheduled arrival + 24 hours without receipt"; port_call_cnt_m: Number of port calls; port_stay_hours_m: Total stay in port (hours); geo_entropy_m: Geographical dispersion of receiving and shipping (geographical entropy).

[0136] Write tb_sc_feature_monthly(uscc,month_end_date,...).

[0137] Step S35: Feature Alignment with Sample Backtracking and Joint Modeling 1. Alignment Rules Map backtrack_date in the sample table to month_end_date, and use (uscc, month_end_date) as the key. The platform includes enterprise profile variables; Characteristics of Cybersecurity Month; Monthly characteristics of supply chain circulation; Merge them into an extended feature set FeatureSet_ext2.

[0138] 2. Sample backtracking In the sample matching module, FeatureSet_ext2 is used to create tables, store data, and backtrack features to generate sample feature tables. If the number of variables exceeds the threshold of a single table, the tables are split according to the existing splitting strategy without changing the variable definition.

[0139] 3. Modeling and Risk Decision Output Load the sample feature table in the sandbox modeling module and use a fixed training strategy: Time segmentation: the previous 12 months were for training, and the most recent month was for validation; Model: Gradient boosting tree or logistic regression, with numerical features standardized; Output a risk score ∈ [0,1], and output the warning result with a threshold θ=0.65 (platform policy configuration): A post-loan warning is triggered when the score is greater than or equal to 0.65. score < 0.65: No warning is triggered.

[0140] Step S36: Decision Audit and Evidence Preservation (Immutable and Traceable Records) Generate audit logs and write them to tb_decision_audit_immut for the following key processes: Cybersecurity Month Feature Generation: Records config_version (scanning frequency, confidence threshold, evidence weight, and other configuration versions), dict_version (variable dictionary version), and mapping_rule_id (attribution mapping rule number); Monthly characteristics of supply chain flow: Records configurations such as enterprise-related rule versions and delay thresholds; Sample backtracking and modeling: Record model version (model_version), parameter digest (param_digest, such as hash digest of normalization switch, regularization coefficient, etc.), and decision digest (decision_digest, sample count, threshold θ, hash of summary of list of companies triggering early warning). The record_hash and prev_hash are calculated using chain hashing to ensure that the audit chain is continuous and verifiable.

[0141] Write to tb_audit_alert and associate it with the audit record hash when one of the following exceptions occurs: Insufficient confidence in the ownership of online assets led to the removal of more than 50% of assets. Missing supply chain data resulted in both ship_cnt_m and recv_cnt_m being 0 for the current month, and not being 0 for the previous month; The model output distribution is abnormal (e.g., the average score for the current month deviates by more than 0.2 compared to the previous month).

[0142] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of the present invention.

Claims

1. A corporate full lifecycle credit risk decision-making system based on multi-source data, characterized in that, include: The data aggregation module is used to collect and aggregate heterogeneous information from multiple sources, such as government data, industry data, commercial data and financial data, based on a distributed data processing architecture. It also creates profiles of enterprises according to preset feature dimensions, generates enterprise profile variables, and supports the retrospective analysis of enterprise credit risk feature data. The visual navigation module is used to respond to user operation commands after user login, call system function interfaces, and generate a step-by-step operation guidance interface. The variable overview module is used to provide a global overview of the data coverage of the enterprise profile variables according to the month dimension; The variable exploration module is used to perform distribution exploration and quantile statistics display on the target variable to obtain the coverage and value distribution characteristics of the target variable in different time periods. The variable filtering module is used to filter from the enterprise profile variables to form a set of target variables based on filtering conditions such as dimension, coverage and data type. The sample matching module is used to receive a sample table containing the unified social credit code and the retrospective date, and perform table creation and database storage and feature retrospective processing on the target variable set to generate a sample feature table. The sandbox modeling module is used to perform model training, parameter tuning, and model validation on the sample feature table in an isolated modeling environment based on a visual interactive interface, to obtain a credit risk scoring model and output a risk score for risk decision-making and post-loan management. The table lifecycle management module is used to manage the lifecycle of the sample feature table and perform dependency verification and state locking when a deletion instruction is received. The model management module is used to centrally collect and manage the sample data, intermediate processing data, variable feature data, parameter configuration data, iteration process data, and model output results generated during the modeling process. The system management module is used to manage accounts, roles, and permissions throughout their entire lifecycle, and to manage the isolated configuration of sandbox accounts.

2. The system according to claim 1, characterized in that, It also includes the government cloud modeling platform and the state-owned assets cloud modeling platform. The two are respectively configured with data exchange, model training, model deployment and account partitioning functions, and interact with the government cloud database and the state-owned assets cloud database respectively to realize cross-data domain data exchange and domain-based management of modeling resources. The sandbox modeling module is a Python-based Jupyter visual interface with built-in supervised and unsupervised algorithm libraries and supports custom algorithm programming. It also supports joint modeling of financial institutions' internal labels and the enterprise profile variables. An isolation mechanism is adopted that binds accounts to database permissions one-to-one, so that each user account can only access the specified database and table resources bound to it; The system management module assigns an independent Python runtime partition to each sandbox account and allows the installation of algorithm packages within the runtime partition. It also records account operations, role configurations, and permission change logs for auditing and traceability. The model management module performs regular backups of the collected modeling data and model results.

3. The system according to claim 1, characterized in that, The variable overview module filters variable coverage by month based on nine feature dimensions and uses a spectral heatmap to display the percentage distribution of data coverage in each month. The variable exploration module supports both hierarchical drill-down filtering based on first-level, second-level, and third-level dimensions, as well as precise filtering based on variable names. After filtering, it displays the latest monthly non-empty sample size distribution, data coverage, and quantile statistics of the target variable arranged in monthly order.

4. The system according to claim 1, characterized in that, The sample matching module includes sub-modules for variable upload management, sample upload matching, and backtracking status monitoring, and satisfies the following: Sample uploads only support sample tables containing "Unified Social Credit Code" and "Backtrack Date", and complete duplicate samples with the same Unified Social Credit Code and Backtrack Date will be checked and blocked. Automatic splitting is performed during table creation, database storage, and feature backtracking to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3. When a memory overflow error is detected during the table creation process, a one-click rerun mechanism is provided, which automatically breaks down the original table creation statement into multiple sub-table creation statements and re-executes them.

5. The system according to claim 1, characterized in that, It also includes a remote sensing feature extraction module, an industrial IoT feature extraction module, and a trusted acquisition and remote verification module, among which: (1) The remote sensing feature extraction module is used to: determine the remote sensing region of interest (ROI) centered on the geographical coordinates obtained by parsing the business address or factory address of the enterprise entity, and rasterize the ROI with a preset spatial grid; perform monthly aggregation feature generation on the remote sensing image covering the ROI, and the generated features include at least the monthly average value of nighttime light intensity, the month-on-month change rate of nighttime light intensity, and the change rate of building outline; wherein, the grid side length of the spatial grid is 100m to 250m, and when the effective pixel ratio of the remote sensing image in the ROI is less than 0.6, the remote sensing features of the corresponding month are set to missing and a missing indicator variable is generated; (2) The industrial IoT feature extraction module is used to: collect industrial operation signals, including at least the power load time series and the key equipment vibration time series, at the enterprise-side edge gateway, and perform windowed feature processing on the industrial operation signals at the edge gateway to obtain daily industrial operation features; wherein, the sampling period of the power load time series is 1min to 5min, the sampling frequency of the key equipment vibration time series is 1kHz to 10kHz, the window length of the windowed feature processing is 10s to 60s and at least outputs frequency domain energy, spectral entropy, kurtosis and outlier density; and upload the daily industrial operation features to the data aggregation module on a daily basis, and aggregate them into monthly industrial operation features on a monthly basis; (3) The trusted acquisition and remote verification module is used to: generate trusted verification information for remote sensing features and industrial operation day features. The trusted verification information includes at least the device identity identifier, firmware metric hash, feature data digest hash, and a remote verification token formed by a timestamped signature. Before writing the remote sensing features and the industrial operation day features into the data storage, the remote verification token is verified. It is allowed to be stored only when the verification is successful. Otherwise, it is rejected and an audit log is recorded. The audit log is stored in a chain hash structure, and each audit record includes at least the previous hash, the current record digest hash, and a timestamp. (4) The remote sensing features, the monthly industrial operation features, and the enterprise profile variables are aligned using a unified retrospective approach: the retrospective dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned remote sensing features and monthly industrial operation features are included as new enterprise profile variables in the enterprise profile variables, so that the variable screening module can screen them and the sample matching module can retrospectively generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

6. The system according to claim 1, characterized in that, It also includes a network security attack surface characterization module, a supply chain logistics flow characterization module, and a decision audit and evidence retention module, and meets the following constraints and processing rules: (1) The network security attack surface feature module is used to: periodically acquire the public network assets of the enterprise entity and generate attack surface features. The public network assets include at least domain names, subdomains, IP addresses and externally open ports. The mapping between the public network assets and the enterprise entity adopts at least two types of evidence consistency verification. The two types of evidence include at least two of the following: ICP filing information, digital certificate organization information, WHOIS query information and subdomain resolution relationship. The attribution confidence score is calculated for each mapping result. The corresponding network asset is included in the attack surface calculation only when the attribution confidence score is not lower than a preset threshold. The attack surface features include at least the number of externally open ports, the distribution of external service types, the number of high-risk vulnerabilities, the statistics of vulnerability repair time, and the number of suspected counterfeit / phishing domain names and the number of certificate anomalies. The periodic acquisition frequency is once every 7 to 30 days, and the results are aggregated into monthly network security features. (2) The supply chain logistics flow feature module is used to: acquire logistics and transportation data associated with the enterprise entity and generate supply chain flow features. The logistics and transportation data includes at least one of the following: maritime automatic identification system (AIS) trajectory data, waybill trajectory data, and vehicle diagnostic (OBD) data. The association mapping between the enterprise entity and the logistics and transportation data is based on at least two of the following for consistency verification: unified social credit code, enterprise name, transportation vehicle registration information, carrier contract number, or waybill number. The supply chain flow features include at least the monthly number of shipments, monthly number of arrivals, monthly abnormal delay rate, number of berthings, port stay duration, and regional dispersion of shipments and receipts. The regional dispersion of shipments and receipts is characterized by regional entropy or equivalent indicators, and the results are aggregated monthly into monthly supply chain flow features. (3) The decision audit and evidence retention module is used to: perform traceable audit and retention of the generation process, version configuration and storage process of the network security monthly features and the supply chain flow monthly features, as well as the process of participating in modeling and risk decision-making based on these two types of features. The audit retention includes at least the data source identifier, feature generation configuration version number, variable dictionary version number, sample backtracking caliber identifier, model version number, parameter configuration summary and risk decision output result; the audit retention adopts an immutable storage strategy, including at least one of chained hash log or WORM write-once-read-many storage, each audit record includes at least the previous hash, current record summary hash and timestamp, and forms alarm records for abnormal events; (4) The network security monthly features, the supply chain circulation monthly features, and the enterprise profile variables are aligned using a unified backtracking caliber: the backtracking dates in the sample table are mapped to the corresponding month-end snapshot dates, and the aligned network security monthly features and supply chain circulation monthly features are included as new enterprise profile variables in the enterprise profile variables, so that the variable filtering module can filter them and the sample matching module can backtrack to generate a sample feature table, and the sandbox modeling module can perform joint modeling or automatic modeling.

7. A method for enterprise full lifecycle credit risk decision-making based on multi-source data, based on the system described in claim 1, characterized in that, include: S1. Collect and aggregate heterogeneous information from multiple sources, including government data, industry data, commercial data, and financial data. Based on a distributed data processing architecture, generate enterprise profile variables according to nine feature dimensions to form enterprise credit risk feature data that supports the longest historical growth trajectory of up to nine years. S2. After the user logs in, a step-by-step operation guide and core function entry points are provided through a visual navigation module; S3. The variable overview module displays the data coverage of enterprise profile variables by month, and the variable exploration module performs distribution exploration and quantile statistics on the target variable to obtain variable coverage and distribution characteristics. S4. The target variable set is formed by filtering based on conditions such as dimension, coverage and data type through the variable filtering module; S5. Receive a sample table containing the unified social credit code and the retrospective date, perform table creation and database storage and feature retrospective for the target variable set, generate a sample feature table, and verify and intercept completely duplicate samples. S6. In the sandbox modeling module, perform model training, parameter tuning and model validation on the sample feature table to obtain the credit risk scoring model and output the enterprise risk score; S7. Collect the modeling process data, parameter configurations and model output results into the model management module for unified management, and perform table lifecycle management and model retention management through the table lifecycle management module and the model management module according to the retention strategy. S8. Make risk decisions based on the enterprise risk score, and implement risk warnings based on changes in enterprise characteristic data in post-loan management.

8. The method according to claim 7, characterized in that, In step S3: the variable overview module uses a spectral heatmap to display the distribution of data coverage percentage for each month; the variable exploration module uses a combined graph to display the non-empty sample size and data coverage, and outputs the quantile statistics of the target variable in monthly order.

9. The method according to claim 7, characterized in that, In step S5: when the number of variables to be processed exceeds the preset threshold, multiple target tables are automatically generated to ensure that the number of variables in a single table does not exceed 500 and the number of associated sub-tables does not exceed 3; when a memory overflow error occurs during the table creation process, the original table creation statement is automatically disassembled and multiple sub-table creation statements are generated and re-executed, thereby completing feature backtracking.

10. The method according to claim 7, characterized in that, In steps S6 and S7: a one-to-one binding mechanism between account and database permissions is adopted, so that users can only access the database and table resources bound to their accounts; an independent Python running partition is allocated to each sandbox account to achieve isolation between users, and account operation, role configuration and permission change logs are recorded; The system determines whether to retain the sample feature table and credit risk scoring model based on the preset retention strategy. If not retained, the corresponding table resources are deleted through the table lifecycle management module. If retained, the model resources are collected and backed up periodically through the model management module.