Multi-source data entity identification and context grading method

By generating multi-source heterogeneous feature sets and entity feature primitive association sets, and combining them with PCIe 4.0 bus transmission, the problems of single entity identification dimension and low transmission efficiency in credit risk assessment are solved, and the efficiency of credit approval in dynamic risk assessment and high-concurrency scenarios is improved.

CN121981815APending Publication Date: 2026-05-05BEIJING HUARONG XINNING TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HUARONG XINNING TECH CO LTD
Filing Date
2026-01-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing credit risk assessment schemes fail to delve into the intrinsic relationships between features from multiple data sources, resulting in a single dimension of entity identification, reliance on static thresholds that cannot be dynamically adjusted, low transmission efficiency, and impact on the real-time performance and accuracy of credit approval.

Method used

A multi-source heterogeneous feature set is generated, an initial context risk cascade table is constructed through the entity feature primitive association set, and the PCIe 4.0 bus is used for transmission. Combined with the interrupt controller, dynamic calibration and rapid risk assessment are achieved.

Benefits of technology

It achieves deep correlation and integration of multi-source data features, dynamically adjusts risk assessment results, improves the real-time performance and accuracy of credit approval, and meets the transmission requirements in high-concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981815A_ABST
    Figure CN121981815A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source data entity identification and context grading method, which comprises the following steps of: generating a multi-source heterogeneous feature set according to credit investigation data, asset and liability data and consumer transaction flow data in a credit and loan scene; performing entity feature primitive extraction on the multi-source heterogeneous feature set to generate an entity feature primitive association set; generating an initial context risk cascade table based on the entity feature primitive association set; performing dynamic calibration on the initial contextual risk cascade table based on the real-time credit transaction data to generate a calibrated contextual risk cascade table; and outputting the calibrated context risk cascade table to a credit approval system in the form of a risk assessment comprehensive data packet through a PCIe4.0 bus based on a credit risk early warning mechanism of an interrupt controller. According to the method, the dynamic adaptation of risk assessment is realized, the change of the customer risk state can be reflected in time, the risk misjudgment probability caused by insufficient data timeliness is reduced, and the timeliness and rationality of the risk assessment result are better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method for multi-source data entity recognition and context classification. Background Technology

[0002] In the field of credit risk assessment, with the development of financial technology, the credit approval process needs to integrate multi-source heterogeneous information such as credit data, asset and liability data, and consumer transaction data to comprehensively depict a customer's credit status and repayment ability. Currently, financial institutions are continuously increasing their requirements for the precision and efficiency of credit risk management. How to effectively process effective information from multi-source data, accurately identify the relationships between data entities, and scientifically classify risk levels based on dynamic data changes has become a key requirement for supporting credit approval decisions and reducing credit default risk, directly impacting the compliance and operational efficiency of credit operations.

[0003] Existing multi-source data processing solutions for credit risk assessment typically involve first performing independent feature filtering on credit data, asset and liability data, and consumer transaction data, removing fields deemed irrelevant to form a single data source feature subset; then, simply concatenating the feature subsets from each single data source according to the customer identification field to form a unified multi-source data set; finally, based on this data set, using pre-set fixed risk assessment indicators and thresholds, a static risk assessment result is generated.

[0004] However, the existing solutions mentioned above have several drawbacks. First, their processing of multi-source data remains at the level of independent screening and simple splicing, failing to delve into the inherent relationships between the characteristics of different data sources. This results in a single dimension for identifying data entities, making it difficult to form a comprehensive characterization of customer risk. Second, risk assessment relies on static thresholds, making it impossible to dynamically adjust risk assessment results based on real-time credit transaction data. This can easily lead to misjudgments of risk due to insufficient data timeliness. Furthermore, data transmission uses conventional interfaces, which cannot meet the need for rapid transmission of risk information to the approval system in high-concurrency credit business scenarios, affecting the real-time performance and accuracy of credit approval. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a method for multi-source data entity recognition and context-based classification, which can at least alleviate the aforementioned technical problems.

[0006] The technical solutions provided in this application are as follows:

[0007] This application provides a method for multi-source data entity recognition and context classification, which includes:

[0008] Step 1: Generate a multi-source heterogeneous feature set based on credit data, asset and liability data, and consumer transaction data in the credit scenario;

[0009] Step 2: Extract entity feature primitives from the multi-source heterogeneous feature set to generate an entity feature primitive association set;

[0010] Step 3: Generate an initial context risk cascade table based on the entity feature primitive association set;

[0011] Step 4: Based on real-time credit transaction data, perform dynamic calibration on the initial context risk cascade table to generate a calibrated context risk cascade table;

[0012] Step 5: The credit risk early warning mechanism based on the interrupt controller outputs the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus.

[0013] The technical solution in this application has the following technical advantages:

[0014] I. Technical effectiveness of independently filtering and simply concatenating multi-source data without exploring feature relationships, resulting in a single dimension of entity recognition.

[0015] In the background, traditional solutions simply splice together credit data, asset and liability data, and consumer transaction data after independently filtering features, without establishing relationships between features from different data sources. This results in a single dimension for data entity identification, making it difficult to comprehensively portray customer risk. This application generates a multi-source heterogeneous feature set in step 1, rather than processing each data source in isolation. Instead, it integrates the three types of core credit data into a unified feature set, laying the foundation for feature association mining. Further, step 2 extracts entity feature primitives and generates an association set, enabling the mining of inherent business relationships between different features from the multi-source heterogeneous feature set (such as the relationship between overdue records in credit data and abnormal transaction frequencies in consumer transaction data), forming an association set of entity feature primitives containing related relationships. Compared to the traditional simple splicing method, this solution achieves deep association and integration of multi-source data features, enriching the dimensions of data entity identification, providing a more comprehensive portrayal of customer risk, and effectively alleviating the problem of a single identification dimension in traditional solutions.

[0016] II. Technical effectiveness in addressing risks misjudgment caused by reliance on static threshold assessments and the inability to dynamically adjust based on real-time data.

[0017] In the background, traditional solutions generate static results based on fixed risk assessment indicators and thresholds, which cannot respond to changes in real-time credit transaction data and are prone to misjudgment due to insufficient data timeliness. This application generates an initial contextual risk cascade table based on the association set in step 3, instead of relying on preset static thresholds. Instead, it constructs an initial risk assessment framework based on the association relationships of entity feature primitives. Then, in step 4, the initial cascade table is calibrated using real-time credit transaction data. This allows for dynamic adjustment of the assessment parameters in the risk cascade table based on real-time data (such as the customer's latest consumption transactions and asset changes), ensuring that the risk assessment results are updated synchronously with the real-time data. Compared to traditional static assessment methods, this solution achieves dynamic adaptation of risk assessment, can promptly reflect changes in the customer's risk status, reduces the probability of misjudgment due to insufficient data timeliness, and makes the timeliness and rationality of the risk assessment results superior.

[0018] III. Technical Effects on Poor Real-Time Performance in High-Concurrency Scenarios Regarding Conventional Interface Transmission, Which Impacts Approval Efficiency

[0019] In the background technology, traditional solutions use conventional interfaces to transmit risk assessment results, which is difficult to meet the real-time transmission requirements in high-concurrency credit business scenarios, affecting the real-time performance and accuracy of credit approval. This application addresses this issue by using interrupt controller warnings and PCIe 4.0 bus transmission in step 5. On one hand, the interrupt controller can quickly trigger risk warnings and data transmission commands, reducing command response latency; on the other hand, the PCIe 4.0 bus has higher transmission bandwidth and lower transmission latency compared to conventional transmission interfaces (such as USB and ordinary Ethernet interfaces), enabling rapid transmission of the calibrated risk cascade table to the credit approval system in high-concurrency scenarios. Compared to traditional conventional interface transmission methods, this solution significantly improves the transmission efficiency of risk information, ensuring the real-time acquisition requirements of the approval system for risk information in high-concurrency credit business scenarios, and effectively alleviating the problem of poor real-time performance in traditional solutions. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a multi-source data entity recognition and context-based classification method according to an embodiment of this application.

[0021] Figure 2 This application provides a schematic diagram of the structure of a multi-source data entity recognition and context-based hierarchical device. Detailed Implementation

[0022] Figure 1 This is a flowchart illustrating a multi-source data entity recognition and context-based classification method according to an embodiment of this application. Figure 1 As shown, it includes:

[0023] Step 1: Generate a multi-source heterogeneous feature set based on credit data, asset and liability data, and consumer transaction data in the credit scenario;

[0024] Step 2: Extract entity feature primitives from the multi-source heterogeneous feature set to generate an entity feature primitive association set;

[0025] Step 3: Generate an initial context risk cascade table based on the entity feature primitive association set;

[0026] Step 4: Based on real-time credit transaction data, perform dynamic calibration on the initial context risk cascade table to generate a calibrated context risk cascade table;

[0027] Step 5: The credit risk early warning mechanism based on the interrupt controller outputs the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus.

[0028] Optionally, step 1 involves generating a multi-source heterogeneous feature set based on credit data, asset and liability data, and consumer transaction data in the credit scenario. Specifically, this involves calling the multi-source data heterogeneous feature deconstructor integrated on the FPGA acceleration unit to perform heterogeneous feature deconstruction on the credit data, asset and liability data, and consumer transaction data in the credit scenario, thereby generating a multi-source heterogeneous feature set.

[0029] Optionally, step 1 involves calling the multi-source heterogeneous feature deconstructor integrated on the FPGA acceleration unit to perform heterogeneous feature deconstruction on credit data, asset and liability data, and consumer transaction data in a credit scenario, generating a multi-source heterogeneous feature set, specifically including:

[0030] Step 11: Preload dynamic adaptation firmware that supports multiple protocols onto the FPGA acceleration unit. This dynamic adaptation firmware has built-in multi-protocol adaptive matching logic to identify the format type of credit data, asset and liability data, and consumer transaction flow data, and to separate the credit metadata, asset and liability metadata, and consumer transaction flow metadata from them.

[0031] Step 12: Based on the configured risk-related feature items, the multi-source data heterogeneous feature destructor calls the pre-loaded list of enterprise credit risk-related feature items to perform heterogeneous feature structure on credit metadata, asset and liability metadata, and consumer transaction flow metadata in order to filter out metadata related to credit risk as risk-related feature items;

[0032] Step 13: Calculate the co-occurrence dependency weights of risk-related feature terms to establish a feature temporal dependency chain, mark the user unique identifier index on the feature temporal dependency chain, and generate a multi-source heterogeneous feature set.

[0033] Preferably, the specific implementation process of step 11 is as follows:

[0034] After the FPGA acceleration unit powers on and initializes, it automatically loads pre-compiled, dynamically adaptable firmware that supports multiple protocols. This firmware contains a built-in enterprise data format feature fingerprint library, which stores format features categorized by enterprise data type—the enterprise credit data fingerprint subset includes tag features from XML format enterprise credit reports (such as `...`). <unifiedsocialcreditcode>(Unified Social Credit Code) <enterprisecreditrating>The fingerprint subset of enterprise asset and liability data includes the header features of corporate financial statements in Excel / CSV format (such as the header features of total assets, total liabilities, return on net assets, and financial reporting period), and the field name features of financial data in database table format (such as the corporate_id, balance_sheet_date, operating_cash_flow field). The fingerprint subset of enterprise consumption transaction data includes the header features of corporate account transactions in TCP / IP message format (such as the identifiers transactionType=corporateTransfer and counterpartyType=enterprise), and the message body field features (such as the transactionAmount, transactionTimeStamp, and counterpartyUnifiedCode fields).

[0035] The dynamically adaptable firmware initiates a built-in multi-protocol adaptive matching logic. It first receives input multi-source raw enterprise data (enterprise credit data, enterprise asset and liability data, and enterprise consumer transaction data). For each piece of raw enterprise data, it performs format feature extraction—extracting top-level tags / key names as format features from enterprise credit data, table headers / field names from enterprise asset and liability data, and message header identifiers / message body field names from enterprise consumer transaction data. The extracted format features are compared with preset features in the enterprise data format feature fingerprint database. When the matching degree exceeds a preset threshold (e.g., 85%), the format type of the enterprise's raw data is determined, forming an enterprise raw data-format type correspondence table.

[0036] Based on the determined format type, the firmware dynamically adapts to call the corresponding enterprise metadata parsing module: for enterprise credit data in XML / JSON format, it extracts data according to the enterprise credit metadata parsing rules. <unifiedsocialcreditcode>(Unified Social Credit Code) <reportissuinginstitution>(Reporting Institution) <reportissuedate>(Date of report issuance) <creditratingagency>The field value of "(Credit rating agency)" is used as enterprise credit metadata; for enterprise asset and liability data in Excel / CSV / database table format, the field values ​​of the audit opinion type (such as standard unqualified opinion, qualified opinion) and data reporting agency (such as accounting firm, enterprise financial system) of the unified social credit code financial reporting period are extracted as enterprise asset and liability metadata according to the enterprise financial metadata parsing rules; for enterprise consumption transaction flow data in TCP / IP message format, the field values ​​of the unified social credit code transaction flow collection system identifier (such as bank Corporate Banking System) and flow data generation time data verification code are extracted as enterprise consumption transaction flow metadata according to the enterprise flow metadata parsing rules.

[0037] The extracted enterprise credit information metadata, enterprise asset and liability metadata, and enterprise consumption transaction flow metadata are separated from the payloads in the corresponding original enterprise data (such as credit rating details of enterprise credit information data, specific financial indicator values ​​of asset and liability data, and specific transaction records of transaction flow) to form a set of enterprise data type-enterprise metadata-payload triples. This set is used as the processing object in the subsequent step 12.

[0038] Preferably, in the specific technical implementation of step 12:

[0039] The multi-source data heterogeneous feature deconstructor pre-loads a list of corporate credit risk-related features. This list is configured by credit business experts based on the needs of corporate credit risk assessment and is categorized by data type, including features strongly correlated with corporate risk. Among them, the corporate credit risk feature-related subset includes features such as the number of credit defaults in the past 3 years, the proportion of external guarantees, and whether there are administrative penalty records; the corporate asset and liability risk feature-related subset includes features such as asset-liability ratio, current ratio, quick ratio, net operating cash flow, and return on net assets; and the corporate consumption transaction flow risk feature-related subset includes features such as the frequency of large transfers to corporate accounts in the past 30 days, the proportion of upstream and downstream enterprise transactions in the past 90 days, the timeliness of tax payments in the past month, and the proportion of counterparties involved in litigation. Each feature is marked with the corresponding data source type (e.g., corporate credit rating corresponding to corporate credit data source) and extraction rule identifier (e.g., asset-liability ratio = total liabilities / total assets rule identifier).

[0040] The multi-source data heterogeneous feature destructor calls the enterprise metadata-risk feature association mapping engine. Taking the enterprise data type-enterprise metadata-payload triplet set generated in step 11 as input, it matches the data source type in the enterprise credit risk association feature item list based on the enterprise data type to determine the risk association features to be extracted. Then, based on the extraction rule identifier corresponding to the feature item, it calls the specific rules in the feature item extraction rule library to extract the feature item value from the payload of the triplet set. This value is then associated with the unified social credit code data generation time field in the corresponding enterprise metadata, forming a unified social credit code-risk feature association item name-feature item value-data generation time quadruple. For example, for the asset-liability ratio feature item, after matching the enterprise asset and liability data source, the asset-liability ratio = total liabilities / total assets rule is called. The total liabilities and total assets values ​​are extracted from the payload of the enterprise asset and liability data to calculate the asset-liability ratio. This is then associated with the unified social credit code financial reporting period (data generation time) to form the corresponding quadruple.

[0041] For all generated quadruple pairs (Unified Social Credit Code - Risk Feature Name - Feature Value - Data Generation Time), perform a validity check on the enterprise risk feature association. This involves calling the enterprise feature validation rule base, which includes rules for the numerical range of enterprise features (e.g., the debt-to-equity ratio must be between 0 and 3, and the current ratio must be greater than 0), logical constraints (e.g., when net operating cash flow is negative, a field indicating the reason for cash flow shortage must be associated), and timeliness rules (e.g., the enterprise credit rating must be issued within the last 6 months). For each quadruple pair, validate it one by one according to the rule base, removing quadruple pairs with abnormal values, logical contradictions, or those exceeding the time limit, and retaining valid quadruple pairs. Integrate the Unified Social Credit Code - Risk Feature Name - Feature Value - Data Generation Time from the valid quadruple pairs to form an enterprise risk-related feature set, which will be the processing object in subsequent step 13.

[0042] Preferably, in a scenario, step 13 is specifically implemented as follows:

[0043] Using the set of enterprise risk-related features generated in step 12 as the processing object, the co-occurrence dependency weight calculation model for enterprise features is first invoked. This model is designed with feature correlation calculation logic for enterprise credit scenarios. Based on the unified social credit code, all risk feature-related feature names, feature values, and data generation time records of the same enterprise are classified into enterprise feature groups. For each enterprise feature group, the frequency of any two risk feature-related features appearing simultaneously in the same time period (such as the same financial quarter or the same month) is counted (co-occurrence frequency). Combined with the risk event occurrence rate when the two features co-occur in historical enterprise credit risk cases (such as the default rate when high debt-to-asset ratio and negative net operating cash flow co-occur), the co-occurrence dependency weight of the two is calculated (co-occurrence dependency weight = (co-occurrence frequency / total number of records) × (risk event occurrence rate / industry average risk rate)). The weight value ranges from 0 to 1. The higher the value, the higher the risk correlation between the two features. For example, if a manufacturing company's debt-to-equity ratio and net operating cash flow appear three times in the past four financial quarters, with a co-occurrence frequency of 0.75, and the default rate of this combination in historical cases is 30%, while the industry average default rate is 10%, then the co-occurrence dependency weight of the two is 0.75 × (30% / 10%) = 2.25 (if it exceeds 1, it is taken as 1), forming a unified social credit code - risk characteristic association item pair - co-occurrence dependency weight association table.

[0044] Based on the Unified Social Credit Code-Risk Feature Association Pair-Co-occurrence Dependency Weight Association Table, a corporate feature-time-series dependency chain is constructed: with the data generation time as the time axis (horizontal axis), risk feature association items of the same enterprise are arranged in chronological order; for risk feature association pairs with a co-occurrence dependency weight ≥ 0.5, a directed line segment with a weight label is used to connect them, with the starting point of the line segment being the feature item with an earlier time and the ending point being the feature item with a later time, and the co-occurrence dependency weight is labeled next to the line segment; the corresponding Unified Social Credit Code (unique identifier for enterprise users) is labeled at the beginning of each feature-time-series dependency chain to ensure that each dependency chain is bound to a specific enterprise, forming a Unified Social Credit Code-Enterprise Feature-Time-Series Dependency Chain Correspondence Table.

[0045] The Unified Social Credit Code-Enterprise Characteristics-Time Dependency Chain Correspondence Table is integrated with the risk feature association item numerical data generation time field from the enterprise risk association feature item set retained in Step 12. All association information is aggregated according to the Unified Social Credit Code to generate a multi-source heterogeneous feature set. Each record in this feature set contains the Unified Social Credit Code Enterprise Characteristics-Time Dependency Chain Risk Feature Association Item Name, Risk Feature Association Item Numerical Data Generation Time, and Data Source Type (Enterprise Credit Investigation / Assets and Liabilities / Transaction Flow) fields. These can be directly used as input data for extracting enterprise entity feature primitives in Step 2, ensuring that the risk association information of the enterprise's multi-source data is completely preserved and transferred to subsequent processing stages.

[0046] Optionally, step 2 involves extracting entity feature primitives from the multi-source heterogeneous feature set to generate an entity feature primitive association set. Specifically, this involves extracting entity feature primitives from the multi-source heterogeneous feature set to generate a collapsed feature subset, and generating an entity feature primitive association set based on the credit business scenario tag library and the collapsed feature subset.

[0047] Optionally, step 2 involves extracting entity feature primitives from the multi-source heterogeneous feature set to generate a collapsed feature subset, and generating an entity feature primitive association set based on the credit business scenario label library and the collapsed feature subset. Specifically, this includes:

[0048] Step 21: The GPU parallel computing node assigns weight coefficients according to the real-time impact priority of credit risk assessment, and performs salted hash encoding on the risk feature association items in the multi-source heterogeneous feature set based on the weight coefficients to obtain a set of risk feature association items with weight encoding.

[0049] Step 22: Remove duplicates from the set of risk feature associations with weighted codes by coarse screening based on the coding prefix, and remove duplicates from the set of risk feature associations with weighted codes based on the uniqueness of the complete coding to obtain a collapsed feature subset;

[0050] Step 23: Based on the credit business scenario tag library, match the collapsed feature subset with the scenario tag library to generate entity feature primitives. Then, based on the support coefficient of the entity feature primitives for the risk assessment conclusion, analyze the business logic dependency and co-occurrence rules of the entity feature primitives to generate entity feature primitive association set.

[0051] Preferably, the specific implementation process of step 21 is as follows:

[0052] GPU parallel computing nodes load the "Enterprise Credit Risk Feature Priority Configuration Library". This library is categorized according to enterprise credit risk assessment scenarios, including "Priority Subset for Manufacturing Enterprises", "Priority Subset for Trading Enterprises", and "Priority Subset for Technology Enterprises". Each subset defines the real-time impact priority of risk feature related items. Among them, "High-priority feature items" are features that have a significant impact on the enterprise's short-term repayment ability (such as "Net Operating Cash Flow in the Past 30 Days", "Debt Repayment Ratio", and "Tax Credit Rating"); "Medium-priority feature items" are features that reflect the enterprise's medium-term operational stability (such as "Revenue Growth Rate in the Past 6 Months" and "Asset-Liability Ratio Trend"); and "Low-priority feature items" are features that reflect the enterprise's long-term qualifications but have weaker real-time characteristics (such as "Years of Establishment of the Enterprise" and "Industry Qualification Certification").

[0053] Based on the priority configuration library, the GPU parallel computing nodes assign weight coefficients to the risk feature association items in the "multi-source heterogeneous feature set" generated in step 13: high priority feature items are assigned a weight coefficient of 0.6-0.8, medium priority feature items are assigned a weight coefficient of 0.3-0.5, and low priority feature items are assigned a weight coefficient of 0.1-0.2. Furthermore, feature items within the same priority are assigned differently according to industry characteristics (e.g., the weight of "accounts receivable turnover rate" is higher for trading companies than for manufacturing companies, and the weight of "fixed asset utilization rate" is higher for manufacturing companies than for technology companies).

[0054] For each risk feature association item, "salted hash encoding" is performed: First, a unique "enterprise salt value" is generated for each enterprise user (generated using an irreversible hash algorithm based on the unified social credit code). The risk feature association item name, feature value, weight coefficient, and enterprise salt value are concatenated to form a string. Then, this string is encrypted using the SHA-256 hash algorithm to generate a fixed-length hash value. Finally, a binary identifier of the weight coefficient is embedded in the prefix position of the hash value (e.g., "11" for high priority, "10" for medium priority, and "01" for low priority), forming a "weighted risk feature association item". All weighted risk feature association items are grouped according to the "unified social credit code" to form a "weighted risk feature association item set". This set retains the risk weight information of the original feature items and achieves a standardized representation of the feature items through hash processing.

[0055] Preferably, in the specific technical implementation of step 22:

[0056] Using the "set of risk feature associations with weighted codes" generated in step 21 as the processing object, the "coarse screening and deduplication of coding prefixes" is first performed: extract the prefix identifier (including the weight coefficient identifier and the first 16 bits of the hash value) of each risk feature association with weighted codes, group them by "unified social credit code", and count the frequency of occurrence of the same prefix identifier; when the frequency of occurrence of a certain prefix identifier exceeds a preset threshold (such as 3 times), it is determined to be a cluster of highly similar feature items. The risk feature association with weighted codes with the smallest deviation between the feature item value and the average value of the cluster is retained from the cluster, and the rest are removed to form the "set of feature items after coarse screening of prefixes".

[0057] Perform "Complete Code Uniqueness Refinement Deduplication" on the "Prefix Coarse Screening Feature Item Set": Compare the complete hash values ​​of all risk feature association items with weighted codes in the set. If the complete hash values ​​of the two codes are completely identical, they are determined to be duplicate features. For duplicate features, retain the latest record based on the "data generation time" of the corresponding original risk feature association item (such as retaining features generated within the last 30 days, or retaining records with higher data source credibility under the same timestamp - enterprise credit data has higher credibility than self-reported data), and remove the remaining duplicates to form the "Refinement Screening Feature Item Set".

[0058] The weighted risk feature associations in the "refined feature set" are associated with their corresponding original risk feature association information (feature name, feature value, weight coefficient, unified social credit code, and data generation time). These associations are then sorted using the combination key "unified social credit code + feature name" to generate a "collapsed feature subset." Compared to the original multi-source heterogeneous feature set, this subset reduces the data volume by 30%-60% (due to duplicate removal) while preserving the integrity of high-weight risk feature associations, thus laying an efficient data foundation for subsequent entity feature primitive extraction.

[0059] Preferably, in a scenario, step 23 is specifically implemented as follows:

[0060] Load the "Enterprise Credit Business Scenario Tag Library", which includes "Enterprise Credit Product Tag Set" (such as "Working Capital Loan", "Fixed Asset Loan", "Accounts Receivable Pledged Loan"), "Enterprise Industry Classification Tag Set" (such as "Manufacturing", "Wholesale and Retail", "Information Technology Services"), and "Risk Assessment Dimension Tag Set" (such as "Debt Repayment Ability", "Profitability", "Operational Capability"). Each tag is associated with "Feature Item Matching Rules" (such as matching the "Accounts Receivable Pledged Loan" tag with features such as "Accounts Receivable Turnover Ratio" and "Accounts Receivable Aging Distribution").

[0061] The risk feature association items in the "collapsed feature subset" generated in step 22 are matched with the "enterprise credit business scenario tag library": First, the initial matching tags are determined based on the credit application information of enterprise users (such as the type of loan product applied for and the industry to which the enterprise belongs); then, the correlation degree between each risk feature association item and the tag is calculated through the "feature item-tag correlation degree algorithm" (correlation degree = frequency of occurrence of feature item in the scenario corresponding to the tag × contribution of feature item to the risk assessment of the scenario); when the correlation degree exceeds the preset threshold (such as 0.6), the risk feature association item is bound to the tag to generate a "feature item-scenario tag" association pair.

[0062] Based on the "feature item - scenario label" association pair, risk feature association items under the same scenario label are aggregated to generate "entity feature primitives" - each entity feature primitive includes "primary name" (such as "manufacturing - working capital loan - debt repayment ability primitive"), "list of included risk feature association items", "primary weight coefficient" (calculated by weighting the included risk feature association item weight coefficients) and "associated unified social credit code".

[0063] Load the "Enterprise Risk Element Support Assessment Model". This model calculates the support assessment coefficient of each entity characteristic element to the final risk assessment conclusion based on historical enterprise credit cases (the value ranges from 0 to 1, and the higher the value, the greater the impact of the element on the risk conclusion). For example, the support assessment coefficient of "Manufacturing - Debt Solvency Element" is usually higher than that of "Service Industry - Operational Capability Element".

[0064] Based on the support evaluation coefficient, "business logic dependency analysis" is performed on entity feature primitives: identifying causal relationships between different primitives (e.g., changes in "profitability primitives" lead to changes in "solvency primitives"), with directed arrows representing the direction of dependency; simultaneously, "co-occurrence pattern analysis" is performed: statistically analyzing the co-occurrence probability of different primitives in historical risk events, with association strength values ​​(0-1) representing the co-occurrence probability. The entity feature primitives, the business logic dependencies between primitives, the co-occurrence pattern association strength, and the support evaluation coefficient are integrated to generate an "entity feature primitive association set." This set clearly presents the inherent association structure of enterprise risk characteristics, providing structured input for the subsequent generation of the initial context risk cascade table.

[0065] This technology differs from traditional feature extraction methods in two ways: First, it ensures the retention of high-value enterprise risk features through "weighted encoding + two-level deduplication" while significantly reducing the amount of data. Second, it transforms scattered enterprise risk features into a feature primitive network with business logic through scenario label matching and primitive association parsing, which is more in line with the business decision-making logic of enterprise credit risk assessment.

[0066] Optionally, step 3, generating an initial context risk cascade table based on the entity feature primitive association set, specifically includes: based on the entity feature primitive association set, constructing a credit scenario-specific hierarchical dimension through a context risk cascade mapper built into the DDR4 cache, and generating a context risk cascade table.

[0067] Optionally, step 3, based on the entity feature primitive association set, constructs a credit scenario-specific hierarchical dimension through a context risk cascade mapper built into the DDR4 cache, generating a context risk cascade table, specifically including:

[0068] Step 31: The context risk cascade mapper on the DDR4 cache analyzes the current credit business risk assessment indicator system and scenario requirements, identifies and determines the core hierarchical dimensions including user credit stage, business risk type, and data timeliness characteristics, and generates a dimension attribute description table.

[0069] Step 32: Using historical credit risk case datasets as training samples, generate a cascaded factor mapping rule base by mining the mapping relationship between core grading dimensions and risk outcomes;

[0070] Step 33: Match each primitive in the entity feature primitive association set with the cascade factor mapping rule base according to its business attributes to determine the associated risk cascade factor;

[0071] Step 34: Based on the associated risk cascading factors, generate a contextual risk cascading table and store it in the DDR4 cache.

[0072] Preferably, the specific implementation process of step 31 is as follows:

[0073] Taking the current corporate credit business risk assessment indicator system document and corporate scenario-based requirements specification as the processing objects, the risk assessment indicator system document is first structuredly parsed through the corporate indicator system decomposition module. This extracts the primary risk indicators for enterprises (such as corporate solvency, corporate operational stability, and corporate industry adaptability), secondary sub-indicators (such as the debt-to-asset ratio, current ratio, and interest coverage ratio under corporate solvency), and indicator calculation rules (such as interest coverage ratio = (total profit + financial expenses) / financial expenses), forming a corporate indicator decomposition table of corporate indicator hierarchy, indicator definition, and calculation rules.

[0074] Based on the business scenario descriptions in the enterprise scenario-based requirements specification (such as manufacturing enterprises needing to focus on fixed asset turnover and raw material inventory for credit, and service enterprises needing to focus on cash flow and customer concentration), the core grading dimensions strongly correlated with the indicators are identified through enterprise-level correlation analysis: the original user credit stage is adjusted to the enterprise credit rating dimension (aligning with the characteristics of third-party enterprise rating), the business risk type is refined to the enterprise industry risk type dimension (reflecting the differences between enterprise industries), and the data timeliness characteristics are optimized to the enterprise data timeliness characteristics dimension (adapting to the periodicity of enterprise data), ultimately determining three core grading dimensions.

[0075] For each core grading dimension, define attribute information: For the enterprise credit rating dimension, define attributes including rating classification standards (referencing third-party rating agencies, such as AAA: no default records and extremely strong debt repayment ability; AA: occasional minor overdue payments but relatively strong debt repayment ability; A: no major defaults but average debt repayment ability; B / C: overdue payments or weak debt repayment ability), and rating determination fields (third-party credit rating reports, enterprise historical credit default records, guarantee and compensation status); For the enterprise industry risk type dimension, define attributes including risk type classification standards (referencing industry risk coefficients, such as high-risk industries: commodity trading, real estate development). The system is divided into three risk levels: medium-risk industries (equipment manufacturing and wholesale / retail) and low-risk industries (public utilities and medical services). For medium-risk industries, the system includes fields for determining the industry type (industry code and regulatory policy documents). For data timeliness, attributes are defined including timeliness level classification standards (real-time data: corporate account statements, real-time order data (collected < 24 hours ago); recent data: monthly financial statements, monthly tax data (24 hours ≤ collected < 30 days ago); historical data: annual audit reports, quarterly industry analysis data (collected ≥ 30 days ago)). Timeliness determination fields include data collection timestamps and data cycle identifiers. These three core risk levels and their attribute information are integrated to generate a table describing the enterprise dimension attributes: dimension name - attribute definition - enterprise judgment basis. This provides a specific standard for the subsequent association between dimensions and risk factors.

[0076] Preferably, in the specific technical implementation of step 32:

[0077] The dataset of historical corporate credit risk cases is used as the processing object. This dataset contains corporate credit case records from the past 3-5 years. Each record needs to cover the core hierarchical dimension labels determined in step 31 (corporate credit rating label, corporate industry risk type label, corporate data timeliness characteristic label), risk outcome labels (normal repayment / overdue / extension / bad debt), and also include quantitative indicators of risk impact (such as overdue days, bad debt amount as a percentage of credit limit, number of extensions).

[0078] First, the historical corporate credit risk case dataset is grouped according to corporate dimensions. Cases are grouped by a three-dimensional combination of corporate credit rating label, corporate industry risk type label, and corporate data timeliness characteristic label. For example, AA-level - equipment manufacturing (medium risk) - monthly financial data (recent) is one group, and A-level - bulk commodity trading (high risk) - real-time transaction data (real-time) is another group, forming several corporate dimension combination case groups to ensure that each group contains a sufficient number of samples (e.g., ≥50 samples per group to avoid small sample bias).

[0079] For each enterprise dimension combination case group, the average value of risk impact indicators within the group is calculated through quantitative statistical analysis of enterprise risk impact (such as the average number of overdue days in the overdue group and the average bad debt ratio in the bad debt group). This average value is then compared with the enterprise risk impact benchmark value (the average value of risk impact indicators for all historical cases, such as an average overdue day of 15 days and an average bad debt ratio of 5% for all cases) to calculate the relative risk coefficient (relative risk coefficient = group average risk impact value / enterprise risk impact benchmark value). This relative risk coefficient is the enterprise risk cascading factor corresponding to that enterprise dimension combination. For example, the average bad debt ratio of the B-level - Real Estate Development (High Risk) - Quarterly Financial Report Data (Recent) group is 15%, and the enterprise risk impact benchmark value is 5%, so the enterprise risk cascading factor for this combination = 15% / 5% = 3.0; while the average bad debt ratio of the AAA-level - Public Utilities (Low Risk) - Monthly Tax Data (Recent) group is 1%, so the risk cascading factor = 1% / 5% = 0.2.

[0080] By integrating all enterprise dimension combinations, corresponding enterprise risk cascade factors, and calculation bases (group sample size, group average risk impact value, and benchmark value), a cascade factor mapping rule base of enterprise dimension combination - risk cascade factor - calculation base is generated. This rule base differs from the traditional fixed threshold model and can be updated regularly (e.g., quarterly) as new enterprise credit cases accumulate, ensuring that the risk cascade factors maintain a high degree of matching with the actual risk level of the enterprise credit market.

[0081] Preferably, in a scenario, step 33 is specifically implemented as follows:

[0082] Using the entity feature primitive association set generated in step 23 (limited to enterprise entity feature primitives) and the cascaded factor mapping rule base generated in step 32 as processing objects, the enterprise business attribute label of each enterprise entity feature primitive is first extracted from the enterprise entity feature primitive association set. This label is composed of the enterprise industry type corresponding to the primitive (such as equipment manufacturing primitive and public utility primitive), the associated enterprise risk indicators (such as the debt-to-asset ratio primitive and the debt-paying ability indicators associated with the enterprise), and the data source period (such as monthly revenue primitive and real-time flow primitive). For example, the enterprise business attribute label of the equipment manufacturing-debt-to-asset ratio-monthly primitive is equipment manufacturing industry-debt-paying ability indicators-monthly data.

[0083] Based on enterprise business attribute tags, the core hierarchical dimension value combination corresponding to each enterprise entity feature element is determined through enterprise element-dimension matching logic: For the equipment manufacturing-asset-liability ratio-monthly element, the medium risk (equipment manufacturing) value of the enterprise industry risk type dimension is matched according to the equipment manufacturing industry in the tag, the recent data (monthly) value of the enterprise data timeliness feature dimension is matched according to the monthly data, and then the third-party rating report of the enterprise to which the element belongs (such as the rating of AA) is matched with the AA level value of the enterprise credit rating dimension, and finally the enterprise dimension value combination of the element is formed (AA level - medium risk (equipment manufacturing) - recent data (monthly)).

[0084] The enterprise dimension value combination is precisely compared with the enterprise dimension combination field in the cascade factor mapping rule base: if the two are completely consistent (e.g., there is an AA-level-medium risk (equipment manufacturing)-recent data (monthly) combination in the rule base), the corresponding enterprise risk cascade factor in the rule base is directly extracted as the associated risk cascade factor of the enterprise entity feature primitive; if there are some dimension value differences (e.g., the dimension combination of the primitive is AA-level-medium risk (equipment manufacturing)-real-time data, there is no completely matching item in the rule base, but there are AA-level-medium risk (equipment manufacturing)-recent data and AA-level-medium risk (equipment manufacturing)-historical data), then factor interpolation is used for calculation (weighted calculation based on the proximity of data timeliness, such as the timeliness difference between real-time data and recent data is small, the weight is 0.7, and the weight of historical data is 0.3, to obtain the interpolated risk cascade factor).

[0085] The unified social credit code, element name, enterprise dimension value combination, and associated risk cascade factors of each enterprise entity feature element are integrated to form an enterprise element-risk cascade factor correspondence table. This table directly provides core data support for the subsequent generation of the initial context risk cascade table, ensuring that the risk factors of each element are consistent with the actual industry attributes and credit status of the enterprise.

[0086] Preferably, the specific implementation process of step 34 is as follows:

[0087] Taking the enterprise primitive-risk cascade factor correspondence table generated in step 33 and the enterprise business attribute label calibration trajectory reserved field of the enterprise entity feature primitive association set in step 23 as the processing objects, the enterprise identifier aggregation module first aggregates the records in the enterprise primitive-risk cascade factor correspondence table according to the unified social credit code (unique enterprise identifier), and gathers all enterprise entity feature primitive names, enterprise dimension value combinations and associated risk cascade factors of the same enterprise into the same enterprise group, forming an enterprise-level data group of unified social credit code - enterprise primitive set - dimension combination set - risk factor set.

[0088] Based on enterprise-level data grouping, construct the header structure of the initial context risk cascade table. The header fields need to cover all dimensions of enterprise credit risk assessment requirements, specifically including: core identifier area (unified social credit code, enterprise name, credit application number), basic element information area (enterprise entity characteristic basic element name, enterprise business attribute label), dimension association area (enterprise credit rating dimension value, enterprise industry risk type dimension value, enterprise data timeliness characteristic dimension value), risk factor area (associated risk cascade factors, factor calculation basis), and calibration reservation area (calibration status identifier, reserved calibration value field, data generation timestamp).

[0089] Fill in the information from the enterprise-level data groups one by one into the corresponding fields in the table header: For example, for an enterprise with the unified social credit code 91110000XXXXXXXXX, fill in the Enterprise Entity Characteristic Element Name field for Equipment Manufacturing - Asset-Liability Ratio - Monthly Element, fill in the Enterprise Industry Risk Type Dimension Value field for Medium Risk (Equipment Manufacturing), fill in the Related Risk Cascade Factor field for 2.1 (Associated Risk Cascade Factor), and fill in the Data Generation Timestamp field for 2024-05-2014:30:00 (Data Generation Time) to complete the filling of a single record.

[0090] Perform enterprise data integrity checks on the completed initial context risk cascade table: check whether there are null values ​​in the core identifier fields (such as unified social credit code and loan application number), whether the associated risk cascade factors in the risk factor area are valid values ​​(range 0-5, values ​​outside the range are considered abnormal), and whether the values ​​in the dimension association area are consistent with the actual enterprise information (e.g., verify credit rating values ​​by calling the enterprise credit reporting interface). For null fields, use enterprise data completion rules (e.g., temporarily fill in "unrated" for null credit ratings, and temporarily use the average factor of the same dimension combination in the same industry for null risk factor values); for abnormal value fields, mark them as pending review and trigger manual verification reminders.

[0091] After verification, the initial context risk cascade table is indexed according to the unified social credit code and stored in the preset corporate credit data block of the DDR4 high-speed cache (this block is pre-allocated with fixed memory space, supporting fast query and batch reading by index), to ensure that the data in the table can be efficiently called when performing dynamic calibration in the subsequent step 4, and to avoid the impact of storage delay on the real-time performance of corporate credit risk assessment.

[0092] Optionally, step 4 involves performing dynamic calibration on the initial context risk cascade table based on real-time credit transaction data to generate a calibrated context risk cascade table. Specifically, this involves calling the dynamic confidence calibration model deployed on the ASIC chip and combining it with the real-time credit transaction data accessed via the RDMA protocol to perform dynamic calibration on the initial context risk cascade table, thereby generating a calibrated context risk cascade table.

[0093] Optionally, step 4 involves calling the dynamic confidence calibration model deployed on the ASIC chip and using the real-time credit transaction data accessed via the RDMA protocol to perform dynamic calibration on the initial context risk cascade table, generating a calibrated context risk cascade table, specifically including:

[0094] Step 41: Extract real-time feature parameters corresponding to primitives in the initial context risk cascade table from the real-time credit transaction data accessed via the RDMA protocol, in order to establish a real-time feature-primitive mapping relationship;

[0095] Step 42: Based on the real-time feature-primitive mapping relationship, compare the dynamic deviation between the risk cascade factors and the real-time feature parameters in the initial context risk cascade table, and combine the credit business risk sensitivity feature threshold to determine the target primitive whose dynamic deviation exceeds the feature threshold.

[0096] Step 43: Evaluate the dynamic deviation of the target primitive, match it with the hierarchical calibration rule base, update the risk cascade factors in the initial context risk cascade table, retain the calibration trajectory, and generate a post-calibration context risk cascade table.

[0097] Preferably, the specific implementation process of step 41 is as follows:

[0098] The data processed focuses on real-time corporate credit transaction data accessed via the RDMA protocol. This data specifically refers to dynamically generated real-time operational data within corporate credit scenarios, encompassing four core categories: 1) Real-time corporate bank account transaction data (including transaction amount, counterparty's unified social credit code, transaction type (e.g., payment for goods, tax payment), and transaction timestamp); 2) Real-time corporate order data (including order amount, order fulfillment progress, downstream customer name, and delivery cycle); 3) Real-time corporate collateral valuation data (including collateral type (e.g., factory buildings, equipment, accounts receivable), current valuation, valuation change range, and valuation institution identifier); and 4) Real-time corporate liability adjustment data (including new bank credit lines, maturing debt amounts, and changes in guarantee liabilities). All data carries the corporate unified social credit code data source system identifier (e.g., bank corporate banking system, third-party valuation system).

[0099] First, a dual validation of enterprise real-time credit transaction data is performed: The first validation is a format compliance check, which calls the enterprise's real-time data format rule library to verify the completeness of data fields (e.g., corporate transaction records must include the counterparty's unified social credit code, and order data must include fulfillment progress) and the correctness of data types (e.g., transaction amounts must be positive, and valuation changes must be percentages). The second validation is a business logic check, which calls the enterprise's business logic rule library to verify the logical consistency between data (e.g., the amount of maturing debt cannot exceed the enterprise's current total credit line, and the valuation change of collateral cannot exceed the average fluctuation range of the past three months) and the legality of counterparties (by connecting to the enterprise credit reporting system to verify whether the counterparty is a defaulting enterprise). Data with format violations, logical contradictions, or abnormal counterparties is removed, retaining only the valid set of real-time enterprise data.

[0100] From the initial context risk cascade table generated in step 34, extract the enterprise entity feature primitive name, enterprise unified social credit code primitive, and associated data type (such as enterprise operating cash flow primitive associated with corporate bank account data type, enterprise collateral value primitive associated with collateral valuation data type), forming an enterprise primitive-identifier-data type correspondence table.

[0101] The system invokes a real-time enterprise feature parameter extraction engine, taking a valid set of real-time enterprise data and an enterprise primitive-identifier-data type mapping table as input. It matches the data source system identifier of the valid data with the associated data type of the primitive in the mapping table, and then matches the entity feature primitive of the same enterprise through the enterprise's unified social credit code. The system extracts values ​​directly associated with the primitive from the valid data as real-time enterprise feature parameters—for example, after matching the enterprise operating cash flow primitive with corporate bank transaction data, it extracts the net inflow amount of the corporate bank account in the past 24 hours as a parameter; after matching the enterprise collateral value primitive with collateral valuation data, it extracts the current valuation of the collateral as a parameter; and after matching the enterprise order fulfillment primitive with order data, it extracts the real-time order fulfillment rate (fulfilled order amount / total order amount) as a parameter.

[0102] By integrating the enterprise's unified social credit code, enterprise entity feature primitive name, enterprise real-time feature parameters, and data collection timestamp, a real-time feature-primary mapping relationship table is established. This table directly links the enterprise's real-time operating data with the primitives in the initial cascade table, providing a data association foundation specific to the enterprise for subsequent deviation calculations.

[0103] Preferably, in the specific technical implementation of step 42:

[0104] Using the enterprise real-time feature-primitive mapping relationship table generated in step 41 and the initial context risk cascade table generated in step 34 as processing objects, the associated risk cascade factor of each enterprise entity feature primitive is first extracted from the initial cascade table. The industry type of the enterprise unified social credit code primitive (such as the asset-liability ratio primitive of manufacturing enterprises and the cash flow primitive of service enterprises) is formed to form an enterprise primitive-factor-industry correspondence table.

[0105] By linking the enterprise real-time feature-primitive mapping table with the enterprise primitive-factor-industry correspondence table using the enterprise's unified social credit code and enterprise entity feature primitive names, a five-dimensional correlation table is obtained: Enterprise-Primitive-Initial Risk Cascade Factor-Real-Time Feature Parameter-Industry. Differentiated dynamic deviation calculation logic is adopted for primitive types in different industries.

[0106] For numerical primitives (such as the debt-to-equity ratio primitive for manufacturing enterprises and the net cash flow primitive for service enterprises), the relative deviation formula is used: Dynamic deviation = |Real-time characteristic parameters of the enterprise - Initial factor benchmark value| / Initial factor benchmark value (the initial factor benchmark value is the actual business benchmark corresponding to the initial risk cascade factor, such as the initial factor of 0.6 for the debt-to-equity ratio primitive corresponding to a benchmark value of 60%).

[0107] For frequency / ratio-type primitives (such as the order default frequency primitive for trading companies and the R&D investment ratio primitive for technology companies), the absolute deviation formula is adopted: Dynamic deviation = |Real-time characteristic parameters of the enterprise - Initial factor benchmark frequency / ratio| (e.g., the initial factor of 0.5 for the order default frequency primitive corresponds to a benchmark frequency of 2 times / quarter).

[0108] After the calculation is completed, a dynamic deviation field is added to the five-dimensional association table to form the enterprise primitive deviation association table.

[0109] Load the enterprise credit business risk sensitivity characteristic threshold library. This library sets thresholds according to two dimensions: enterprise industry type and element name, reflecting industry risk differences.

[0110] High-risk industries (such as real estate development and commodity trading): Set the threshold for the debt-to-equity ratio and cash flow to a lower value (such as 15% or 20%), because these industries are more sensitive to changes in debt and cash flow.

[0111] For medium-risk industries (such as equipment manufacturing and wholesale and retail): set the threshold for accounts receivable recovery of order fulfillment units to a medium value (such as 25% or 30%).

[0112] For low-risk industries (such as utilities and healthcare services): the threshold for the tax base of revenue stability should be set to a higher value (e.g., 35% or 40%). The threshold pool can be adjusted according to industry regulatory policies; for example, when the real estate industry is tightened, the threshold for its financing scale base can be lowered.

[0113] The enterprise element deviation association table is linked to the enterprise credit business risk sensitivity feature threshold library through the enterprise entity feature element name of the industry type to which the enterprise belongs. A corresponding industry-specific sensitivity feature threshold is matched for each deviation record. The dynamic deviation is compared with the industry threshold: if the dynamic deviation > the industry threshold, the element is determined to be an enterprise target element requiring calibration; if the dynamic deviation ≤ the industry threshold, it is determined to be an enterprise element that does not require calibration. All target elements are integrated using the enterprise unified social credit code, element name, dynamic deviation, and industry threshold to form an enterprise target element calibration list. This clarifies the scope of elements that need subsequent adjustment, and the industry to which the element belongs must be indicated in the list, providing a basis for differentiated calibration.

[0114] Preferably, in a scenario, step 43 is specifically implemented as follows:

[0115] Using the enterprise target primitive calibration list generated in step 42 and the initial context risk cascade table generated in step 34 as processing objects, the enterprise hierarchical calibration rule library is loaded first. This library divides enterprise-specific calibration strategies according to the range of dynamic deviation, and refines the rules in combination with industry characteristics, including three core strategies:

[0116] Slight deviation calibration strategy (dynamic deviation > industry threshold and ≤ 1.5 times industry threshold): The industry ratio adjustment method is adopted. The adjustment coefficient is set according to the industry to which the enterprise belongs (0.8 for high-risk industries, 0.9 for medium-risk industries, and 1.0 for low-risk industries). The adjusted factor = initial risk cascade factor × (1 + dynamic deviation × adjustment coefficient) -- For example, in the manufacturing industry (medium risk), the dynamic deviation of a certain element is 20%, the industry threshold is 15% (1.33 times the threshold), and the adjustment coefficient is 0.9. Then the adjusted factor = initial factor × (1 + 20% × 0.9) = initial factor × 1.18;

[0117] Moderate deviation calibration strategy (dynamic deviation > 1.5 times the industry threshold and ≤ 2 times the industry threshold): adopt the benchmark reset method to re-obtain the industry benchmark value corresponding to the primitive (such as obtaining the primitive benchmark factor of enterprises of the same size through the industry association database). The adjusted factor = industry benchmark value × (1 + dynamic deviation × 0.5).

[0118] Severe deviation calibration strategy (dynamic deviation > 2 times industry threshold): The enterprise case matching method is adopted. From the historical enterprise high deviation case library, high deviation cases of the same industry, same size and same element type as the current enterprise are matched. The average adjustment range of the elements in the case is referenced (e.g., the average adjustment range of the case is +0.3). Combined with the current dynamic deviation, the adjustment direction and value are determined to ensure that the adjusted factor is consistent with the actual risk level of the industry.

[0119] For each target primitive in the enterprise's target primitive calibration list, the deviation range is determined based on its dynamic deviation. This is then combined with the corresponding calibration strategy in the matching rule library for the enterprise's industry type. For example, if a real estate enterprise (high risk) has a collateral value primitive with a dynamic deviation of 35% and an industry threshold of 15% (2.33 times the threshold, which is considered severe deviation), then the enterprise case matching method is used. Cases with a deviation of 30%-40% in real estate-collateral value primitives are found in the case library. The average adjustment range is +0.25, and the adjustment range for the current primitive is determined to be +0.28 (because the current deviation is slightly higher than the case average, the adjustment range is slightly higher).

[0120] Calculate the enterprise-adjusted risk cascade factor for the target primitive based on the matched calibration strategy, and replace the associated risk cascade factor of the corresponding primitive in the initial context risk cascade table with the adjusted factor. Simultaneously, add enterprise calibration trajectory information to each calibration record. This information must include the calibration timestamp, initial factor value, adjusted factor value, dynamic deviation, calibration strategy adopted, and calibration basis (such as real-time data record ID, industry benchmark source, and matching case number) to ensure the traceability of the enterprise calibration process. For example, the calibration trajectory of the R&D investment ratio primitive for a technology company needs to be recorded as follows: 2024-06-10 09:30:00, initial factor 0.4, adjusted factor 0.52, deviation 25%, strategy: slight deviation - industry ratio adjustment (adjustment coefficient 1.0), basis: real-time R&D investment data ID=RD20240610001.

[0121] After updating the factors of all enterprise target primitives in the initial cascade table, the original factors of primitives that do not need to be calibrated are retained to form a cascade table of enterprise calibrated context risk. At the same time, all calibration trajectory information is linked to the calibration trajectory field of the calibrated table according to the enterprise's unified social credit code and primitive name, or a separate enterprise calibration trajectory detail table is generated and linked to the calibrated table through the enterprise identifier. This ensures that the reasons for the adjustment of enterprise risk factors can be traced during subsequent credit approval, and improves the enterprise scenario adaptability of risk assessment results.

[0122] Optionally, step 5, the credit risk early warning mechanism based on the interrupt controller, outputs the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus, specifically as follows:

[0123] The interrupt controller performs weighted aggregation on the risk cascade factors corresponding to the same user ID in the calibrated context risk cascade table to generate the user risk assessment total score.

[0124] The total risk assessment score is mapped to a specific risk level and encapsulated with the context risk cascade table to form a comprehensive risk assessment data package, which is then transmitted to the credit approval system via the PCIe 4.0 bus. The comprehensive risk assessment data package includes user identification, risk level, details of risk cascade factors, and calibration trajectory information.

[0125] Preferably, the specific implementation process of step 5 is as follows:

[0126] The enterprise-calibrated contextual risk cascade table generated in step 43 is used as the processing object. This table contains core fields such as the enterprise's unified social credit code, enterprise entity characteristic element name, adjusted risk cascade factor calibration trajectory information, and element support assessment coefficient (derived from step 23, reflecting the degree of influence of elements on risk conclusions). First, the records in the calibrated table are grouped according to the enterprise's unified social credit code through the enterprise identification index module, ensuring that all risk cascade factors of the same enterprise belong to the same group, forming an enterprise-element-factor-support group set. Each group corresponds to a unique enterprise and contains all risk assessment element information of that enterprise.

[0127] Load the enterprise primitive dynamic weight configuration table. This table combines the primitive support evaluation coefficient from step 23 with the characteristics of the enterprise's industry, assigning an aggregation weight coefficient to each primitive: primitives with high support evaluation coefficients (e.g., 0.7-1.0) are assigned higher aggregation weights (0.6-0.8), primitives with medium support (0.3-0.6) are assigned medium aggregation weights (0.3-0.5), and primitives with low support (0-0.2) are assigned lower aggregation weights (0.1-0.2). Simultaneously, adjustments are made according to industry differences; for example, the fixed asset utilization rate primitive weight for manufacturing enterprises is increased by 10%, and the accounts receivable turnover rate primitive weight for trading enterprises is increased by 10%. For example, a manufacturing enterprise with a debt-to-asset ratio primitive support of 0.8 and a default aggregation weight of 0.7 will have an actual weight of 0.77 after the industry adjustment.

[0128] For each enterprise group in the enterprise-element-factor-support grouping set, perform risk cascade factor weighted aggregation: multiply each adjusted risk cascade factor of the enterprise by its corresponding aggregation weight coefficient to obtain a weighted factor value; sum the weighted factor values ​​of all elements to generate the enterprise's total risk assessment score. For example, an enterprise contains 3 elements: Factor 1 = 2.0 (weight 0.7), Factor 2 = 1.5 (weight 0.5), and Factor 3 = 0.8 (weight 0.2), then the total score = 2.0 × 0.7 + 1.5 × 0.5 + 0.8 × 0.2 = 1.4 + 0.75 + 0.16 = 2.31. During the aggregation process, the calculation details of the enterprise's total score must be recorded, including the factor value, weight, weighted value, and accumulation process for each element, to ensure the traceability of the total score.

[0129] Preferably, in the specific technical implementation of step 5:

[0130] Using the enterprise's total risk assessment score and industry type as the processing objects, an industry-differentiated risk level mapping rule library is loaded. This library divides risk level ranges according to industry type, reflecting the differences in risk tolerance among different industries:

[0131] High-risk industries (such as real estate development and bulk commodity trading): The risk level is divided into 5 levels with relatively low thresholds (e.g., total score <1.0 is low risk, 1.0-2.0 is low-to-medium risk, 2.0-3.0 is medium risk, 3.0-4.0 is medium-high risk, and >4.0 is high risk).

[0132] Medium-risk industries (such as equipment manufacturing and wholesale and retail): The risk level is also level 5, but the range threshold is moderate (e.g., <1.5 is low risk, 1.5-2.5 is medium-low risk, and so on).

[0133] Low-risk industries (such as public utilities and medical services): The range threshold is relatively high (e.g., <2.0 is low risk, 2.0-3.0 is medium-low risk, and so on).

[0134] The rule base can be dynamically adjusted according to macroeconomic policies. For example, when support policies for micro and small enterprises are introduced, the corresponding industry level thresholds can be appropriately relaxed.

[0135] The company's total risk assessment score is matched with the risk level range of its industry to determine its specific risk level and description (e.g., medium risk - requires additional collateral; high risk - credit is recommended to be rejected). For example, a real estate company (high-risk industry) with a total score of 2.8 matches a medium risk level and is described as medium risk - requires an additional guarantor; a public utility company (low-risk industry) with a total score of 2.8 matches a medium-low risk level and is described as medium-low risk - credit can be granted normally.

[0136] Preferably, in a given scenario, step 5 is specifically implemented as follows:

[0137] Using the enterprise's specific risk level, the contextual risk cascade table after enterprise calibration, the enterprise's total score calculation details, and the enterprise's unified social credit code as the processing objects, a structural framework for a comprehensive enterprise risk assessment data package is constructed, containing four core data segments:

[0138] Identifier segment: Stores the enterprise's unified social credit code, enterprise name, loan application number, and data packet generation timestamp (accurate to milliseconds) to ensure that the data packet is uniquely associated with the enterprise's loan application;

[0139] Level Segment: The specific risk level of the storage enterprise, the total risk assessment score, the level description, and the industry type are presented in a clear manner to show the assessment conclusion;

[0140] Detailed section: Stores key information of the enterprise's calibrated contextual risk cascade table, including the enterprise entity feature primitive name, adjusted risk cascade factor, primitive support assessment coefficient, and retains the core basis for risk assessment;

[0141] Track segment: Stores calibration track information (derived from step 43) and total score calculation details (derived from the weighted aggregation process), including the initial factor, adjustment reason, and weighted calculation process for each primitive, ensuring that the evaluation process is auditable.

[0142] Perform enterprise data compression and verification on data packets: Use industry-specific compression algorithms (for high-risk industry data packets that need to be frequently accessed, the compression rate is set to 30%-40% to balance transmission speed and decompression efficiency; for low-risk industry data packets, the compression rate is set to 50%-60% to save bandwidth). After compression, add a CRC32 checksum (calculated based on the data packet content) for the receiving end to verify data integrity.

[0143] Configure an interrupt controller triggering mechanism: Set interrupt priorities based on the enterprise's specific risk level. High-risk levels (e.g., high-medium-high risk) trigger the highest priority interrupt (IRQ0-15) to ensure priority processing by the credit approval system; medium-risk levels trigger medium-priority interrupts (IRQ16-31); and low-risk levels trigger low-priority interrupts (IRQ32-47). The interrupt signal carries the enterprise's unified social credit code hash value and data packet storage address, facilitating rapid data location by the approval system.

[0144] Transmitting comprehensive enterprise risk assessment data packets via PCIe 4.0 bus: Utilizing the high-speed serial bus characteristics of PCIe 4.0 (single-channel rate 16GT / s), data packets are queued for transmission according to interrupt priority, with high-priority data packets inserted at the front of the transmission queue; during transmission, a data packet fragmentation and reassembly mechanism is enabled to fragment data packets exceeding the maximum transmission unit (MTU) (such as large enterprise data packets containing a large number of calibration tracks), and the receiving end reassembles them according to the fragment sequence number to ensure data integrity.

[0145] After receiving the data packet, the credit approval system first verifies the CRC32 checksum. If the verification passes, it reads the complete data packet based on the storage address in the interrupt signal, parses and displays the enterprise's risk level, factor details, and calibration trajectory, providing a structured and traceable basis for enterprise risk assessment for credit approval decisions. This technology differs from traditional fixed-format transmission schemes. Through industry-differentiated level mapping priority interrupt transmission with full trajectory traceability, it improves the accuracy of enterprise credit risk assessment and approval efficiency, while also meeting the financial regulatory requirements for auditable risk assessment.

[0146] Figure 2 This is a schematic diagram of the structure of a multi-source data entity recognition and context-based hierarchical device according to an embodiment of this application. Figure 2 As shown, it includes:

[0147] Feature generation unit: used to generate multi-source heterogeneous feature sets based on credit data, asset and liability data, and consumer transaction data in credit scenarios;

[0148] Primitive Extraction Unit: Used to extract entity feature primitives from multi-source heterogeneous feature sets to generate entity feature primitive association sets;

[0149] Cascade table generation unit: used to generate an initial context risk cascade table based on the entity feature primitive association set;

[0150] Dynamic calibration unit: used to perform dynamic calibration on the initial context risk cascade table based on real-time credit transaction data to generate a calibrated context risk cascade table;

[0151] Data output unit: used for the credit risk early warning mechanism based on the interrupt controller, outputting the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus.

[0152] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.< / creditratingagency> < / reportissuedate> < / reportissuinginstitution> < / unifiedsocialcreditcode> < / enterprisecreditrating> < / unifiedsocialcreditcode>

Claims

1. A method for multi-source data entity recognition and context-based hierarchical classification, characterized in that, include: Step 1: Generate a multi-source heterogeneous feature set based on credit data, asset and liability data, and consumer transaction data in the credit scenario; Step 2: Extract entity feature primitives from the multi-source heterogeneous feature set to generate an entity feature primitive association set; Step 3: Generate an initial context risk cascade table based on the entity feature primitive association set; Step 4: Based on real-time credit transaction data, perform dynamic calibration on the initial context risk cascade table to generate a calibrated context risk cascade table; Step 5: The credit risk early warning mechanism based on the interrupt controller outputs the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus.

2. The method according to claim 1, characterized in that, Step 1: Generate a multi-source heterogeneous feature set based on credit data, asset and liability data, and consumer transaction data in the credit scenario. Specifically, call the multi-source data heterogeneous feature deconstructor integrated on the FPGA acceleration unit to perform heterogeneous feature deconstruction on the credit data, asset and liability data, and consumer transaction data in the credit scenario to generate a multi-source heterogeneous feature set.

3. The method according to claim 2, characterized in that, Step 1: Call the multi-source heterogeneous feature destructor integrated on the FPGA acceleration unit to perform heterogeneous feature destructor on credit data, asset and liability data, and consumer transaction data in the credit scenario, generating a multi-source heterogeneous feature set, specifically including: Step 11: Preload dynamic adaptation firmware that supports multiple protocols onto the FPGA acceleration unit. This dynamic adaptation firmware has built-in multi-protocol adaptive matching logic to identify the format type of credit data, asset and liability data, and consumer transaction flow data, and to separate the credit metadata, asset and liability metadata, and consumer transaction flow metadata from them. Step 12: Based on the configured risk-related feature items, the multi-source data heterogeneous feature destructor calls the pre-loaded list of enterprise credit risk-related feature items to perform heterogeneous feature structure on credit metadata, asset and liability metadata, and consumer transaction flow metadata in order to filter out metadata related to credit risk as risk-related feature items; Step 13: Calculate the co-occurrence dependency weights of risk-related feature terms to establish a feature temporal dependency chain, mark the user unique identifier index on the feature temporal dependency chain, and generate a multi-source heterogeneous feature set.

4. The method according to claim 1, characterized in that, Step 2: Extract entity feature primitives from the multi-source heterogeneous feature set to generate an entity feature primitive association set, specifically: Entity feature primitives are extracted from multi-source heterogeneous feature sets to generate collapsed feature subsets, and entity feature primitive association sets are generated based on the credit business scenario label library and the collapsed feature subsets.

5. The method according to claim 4, characterized in that, Step 2: Extract entity feature primitives from the multi-source heterogeneous feature set to generate a collapsed feature subset, and based on the credit business scenario label library, generate an entity feature primitive association set according to the collapsed feature subset, specifically including: Step 21: The GPU parallel computing node assigns weight coefficients according to the real-time impact priority of credit risk assessment, and performs salted hash encoding on the risk feature association items in the multi-source heterogeneous feature set based on the weight coefficients to obtain a set of risk feature association items with weight encoding. Step 22: Remove duplicates from the set of risk feature associations with weighted codes by coarse screening based on the coding prefix, and remove duplicates from the set of risk feature associations with weighted codes based on the uniqueness of the complete coding to obtain a collapsed feature subset; Step 23: Based on the credit business scenario tag library, match the collapsed feature subset with the scenario tag library to generate entity feature primitives. Then, based on the support coefficient of the entity feature primitives for the risk assessment conclusion, analyze the business logic dependency and co-occurrence rules of the entity feature primitives to generate entity feature primitive association set.

6. The method according to claim 1, characterized in that, Step 3: Generate an initial context risk cascade table based on the entity feature primitive association set. Specifically, this includes: constructing a credit scenario-specific hierarchical dimension based on the entity feature primitive association set through a context risk cascade mapper built into the DDR4 cache, and generating a context risk cascade table.

7. The method according to claim 6, characterized in that, Step 3: Based on the entity feature primitive association set, construct a credit scenario-specific hierarchical dimension through the context risk cascade mapper built into the DDR4 cache, and generate a context risk cascade table, specifically including: Step 31: The context risk cascade mapper on the DDR4 cache analyzes the current credit business risk assessment indicator system and scenario requirements, identifies and determines the core hierarchical dimensions including user credit stage, business risk type, and data timeliness characteristics, and generates a dimension attribute description table. Step 32: Using historical credit risk case datasets as training samples, generate a cascaded factor mapping rule base by mining the mapping relationship between core grading dimensions and risk outcomes; Step 33: Match each primitive in the entity feature primitive association set with the cascade factor mapping rule base according to its business attributes to determine the associated risk cascade factor; Step 34: Based on the associated risk cascading factors, generate a contextual risk cascading table and store it in the DDR4 cache.

8. The method according to claim 1, characterized in that, Step 4: Based on real-time credit transaction data, perform dynamic calibration on the initial context risk cascade table to generate a calibrated context risk cascade table. Specifically, call the dynamic confidence calibration model deployed on the ASIC chip, and combine it with the real-time credit transaction data accessed via the RDMA protocol to perform dynamic calibration on the initial context risk cascade table to generate a calibrated context risk cascade table.

9. The method according to claim 8, characterized in that, Step 4: Invoke the dynamic confidence calibration model deployed on the ASIC chip, and perform dynamic calibration on the initial context risk cascade table using real-time credit transaction data accessed via the RDMA protocol, generating a calibrated context risk cascade table, specifically including: Step 41: Extract real-time feature parameters corresponding to primitives in the initial context risk cascade table from the real-time credit transaction data accessed via the RDMA protocol, in order to establish a real-time feature-primitive mapping relationship; Step 42: Based on the real-time feature-primitive mapping relationship, compare the dynamic deviation between the risk cascade factors and the real-time feature parameters in the initial context risk cascade table, and combine the credit business risk sensitivity feature threshold to determine the target primitive whose dynamic deviation exceeds the feature threshold. Step 43: Evaluate the dynamic deviation of the target primitive, match it with the hierarchical calibration rule base, update the risk cascade factors in the initial context risk cascade table, retain the calibration trajectory, and generate a post-calibration context risk cascade table.

10. The method according to claim 1, characterized in that, Step 5: The credit risk early warning mechanism based on the interrupt controller outputs the calibrated context risk cascade table to the credit approval system in the form of a comprehensive risk assessment data packet via the PCIe 4.0 bus. Specifically: The interrupt controller performs weighted aggregation on the risk cascade factors corresponding to the same user ID in the calibrated context risk cascade table to generate the user risk assessment total score. The total risk assessment score is mapped to a specific risk level and encapsulated with the context risk cascade table to form a comprehensive risk assessment data package, which is then transmitted to the credit approval system via the PCIe 4.0 bus. The comprehensive risk assessment data package includes user identification, risk level, details of risk cascade factors, and calibration trajectory information.

Citation Information

Patent Citations

  • Financial risk intelligent early warning and risk control system based on FPGA hardware acceleration

    CN110610099A

  • Investment decision-making system and method based on federated learning and multi-modal risk control

    CN120013680A

  • Real-time credit risk early warning system based on multi-source heterogeneous data fusion

    CN120471706A

  • Network security and data security comprehensive analysis method and system based on large model

    CN120825344A

  • Credit risk dynamic assessment and early warning method based on multi-dimensional data analysis

    CN120975904A