A method, system, equipment, and medium for optimizing bank factoring business based on data tagging governance.

CN122089448APending Publication Date: 2026-05-26EVERGROWING BANK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EVERGROWING BANK CO LTD
Filing Date
2025-12-11
Publication Date
2026-05-26

Smart Images

  • Figure CN122089448A_ABST
    Figure CN122089448A_ABST
Patent Text Reader

Abstract

The application provides a bank factoring business optimization method, system, device and medium based on data tag management, belongs to the technical field of bank factoring business, and specifically comprises the following steps: preprocessing internal and external data, constructing an enterprise relationship graph, generating initial tags in combination with historical data and rules, dynamically updating upstream and downstream tags by using node similarity and a decay factor through a tag propagation algorithm, reflecting real-time risk states of enterprises, setting abnormal detection rules, triggering tag re-inspection to maintain accuracy, and adjusting credit limits and freezing abnormal financing prompts according to tags. The application improves the risk identification accuracy, response speed and control flexibility of the factoring business, ensures the timeliness of risk assessment through abnormal detection, accurately matches the risk control measures with the state of enterprises, and improves the service efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bank factoring business technology, specifically relating to a method, system, equipment and medium for optimizing bank refactoring business based on data tag governance. Background Technology

[0002] Refactoring is an important component of supply chain finance, referring to a factoring company acquiring accounts receivable already acquired by other factoring companies and further providing services such as financing, accounts receivable management, and bad debt guarantees. This business model helps to revitalize factoring assets, diversify risks, and improve capital utilization efficiency, and is widely used in various industries such as manufacturing, wholesale and retail, healthcare, and electronics.

[0003] In related technologies, when banks assess the risks of reinsurance business, they often analyze data such as the company's financial statements and credit records, neglecting the inter-company relationships such as transactions, equity, and guarantees. For example, if a core company experiences credit risk, its upstream suppliers may fall into crisis due to uncollectible accounts. However, if the assessment only focuses on the supplier's own data, it cannot provide early warnings, resulting in a one-sided risk assessment and easily overlooking systemic risks that can spread across enterprises.

[0004] While related technologies define risk labels for businesses, they struggle to capture sudden, unexpected risks. For example, if a supplier's monthly transaction volume suddenly surges by 500%, but the label isn't updated in time, banks might still provide financing based on the old label, potentially leading to fraud. Furthermore, a 300% increase in transaction volume during peak season might be normal for a supplier in a seasonal industry, but a fixed threshold could be misjudged as abnormal, increasing risk. Summary of the Invention

[0005] This invention provides a method for optimizing bank refactoring business based on data tagging governance, which improves the accuracy of risk identification in bank refactoring business and enhances the service efficiency for upstream and downstream enterprises in the supply chain.

[0006] The methods include: S101: Acquire internal and external data and perform preprocessing; S102: Construct and store an enterprise relationship graph based on the preprocessed data; S103: Based on the preprocessed data, establish a labeling system and mapping rules, and generate initial labels; S104: Based on the tag propagation algorithm, combined with the node similarity matrix and decay factor, establish a tag propagation method and update the tags of upstream and downstream enterprises in the supply chain; Node similarity is configured based on transaction size, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises; through an asynchronous update strategy, the labels of unlabeled nodes are made to approach the weighted average of the labels of their neighboring nodes until the label distribution is stable; S105: Establish outlier detection rules. When an outlier exceeds a set threshold, trigger tag re-detection. The tag re-detection includes updating the initial tag and recalculating the tag propagation. S106: Based on the labels of each enterprise, take risk prevention and credit control measures in the factoring business; the measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating, and issuing a financing freeze notice when the enterprise's accounts receivable period is abnormal.

[0007] It should be further explained that in step S104, the node similarity is calculated based on the transaction scale, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises; through an asynchronous update strategy, the labels of unlabeled nodes are made to approach the weighted average of the labels of their neighboring nodes until the label distribution is stable. For firm i and firm j, the similarity is... The formula is expressed as follows:

[0008] in, : Total historical transaction amount between company i and company j; Maximum one-sided transaction amount; Number of transactions; Industry-adaptive decay rate; : Total historical guarantee amount of Enterprise i and Enterprise j; Maximum guarantee amount; Number of guarantees; : Current time - last transaction time; Equity linkage strength coefficient; Distance between equity levels; Equity decay coefficient; Equity weighting; Compliance coefficient.

[0009] It should be further noted that the industry adaptive decay rate The calculation method is as follows: ,in Industry risk volatility; in, The half-life of the risk impact; industry default rate standard deviation over the past year

[0010] Daily rate of change of the industry risk index; : Average daily rate of change of the industry risk index; 𝑇: Observation window; k: Adjustment coefficient; Equity decay coefficient.

[0011] Shareholding ratio The calculation method is as follows:

[0012] Compliance coefficient The calculation method is as follows:

[0013] Basic compliance items , indicating the satisfaction status of the k-th regulatory rule; [This refers to the penalty coefficient for high-risk transactions;] ; : represents the historical violation decay rate; Violation_count is the sum of the number of violations committed by enterprise i or j over the past 12 months.

[0014] It should be further explained that in S104, the label propagation algorithm propagates labels between nodes, so that the labels of unlabeled nodes gradually approach the weighted average of the labels of their neighboring nodes until the label distribution is stable. For each unlabeled enterprise j, calculate the label-weighted average of its neighbors:

[0015]

[0016] For companies that have already been labeled, their labels will remain unchanged, and the initial value Y will be forcibly applied. ; Repeat until convergence or the maximum number of iterations is reached: ; For each unlabeled node, select the label with the highest probability in its final label distribution vector as its predicted label; Unlabeled node j final label for: .

[0017] It should be further explained that step S103 specifically includes: Based on a company's credit data, financial data, transaction data, and compliance information, multiple judgment conditions are set, each corresponding to a specific indicator threshold; when a company meets any combination of judgment conditions, it is assigned a corresponding hard label. Predict the risk probability of enterprises and output a probability value that reflects the risk level of enterprises as a soft label; The final risk index is generated by merging hard and soft labels, and weights are calculated based on historical accuracy. The hard and soft labels are then weighted and summed to obtain the final risk index. An outlier detection method is set up so that when an enterprise's transaction data or other key indicators show abnormal changes exceeding the threshold, the initial labels are recalculated and the label fusion process is re-examined to update the final risk index.

[0018] It should be further noted that step S103 also includes: Extract enterprise credit data, supply chain transaction behavior data, and industry characteristic data, perform field validation and content extraction respectively, and form a standardized dataset for tag generation; The labeling system is divided into a basic layer, a risk layer, and a compliance layer, with each layer further subdivided into 27 specific subcategories of labels. For companies that have already been labeled, their label vectors are determined by verifying them using hard indicators such as historical bad debt rates and records of legal disputes; the label trigger thresholds are adjusted according to industry risk fluctuations.

[0019] It should be further explained that step S105 specifically includes: Based on the risk scenarios of reinfactoring business, specific detection indicators are defined to cover transaction anomalies, financial anomalies, and compliance anomalies. Adjust the trigger thresholds for each abnormal indicator based on industry cycles, company type, and characteristics of historical abnormal events; When an anomaly is detected in a company, the tags of its upstream and downstream related companies are obtained through the company relationship graph, and the anomaly is analyzed to see if it has spread to related companies. Based on the severity of the anomaly indicators, different tag re-inspection processes are triggered in stages.

[0020] It should be further explained that step S106 specifically includes: Based on the company's latest tag combination, retrieve the corresponding upper and lower credit limits from the credit limit configuration table and generate a credit limit range; When the abnormal payment period indicator in the tag reaches the preset warning value, a freeze command is generated and synchronized to the monitoring cloud. Based on the company's list of pledgeable assets, credit overflow records, and factoring balance, a credit enhancement plan is pushed to the financing party, along with the expected unlocking amount and time information; When a company's label falls back to a safe range and meets the credit enhancement conditions during the continuous observation period, the unfreezing process is executed, and an electronic report is generated.

[0021] This application also provides a bank refactoring business optimization system based on data tagging governance, the system comprising: The data preprocessing module is used to acquire internal and external data and perform preprocessing. The relationship graph construction module builds and stores enterprise relationship graphs based on preprocessed data. The tag initialization module establishes a tag system and mapping rules based on the preprocessed data, and generates initial tags; The tag propagation and iteration module, based on the tag propagation algorithm, combines node similarity matrix and decay factor to establish tag propagation methods and update the tags of upstream and downstream enterprises in the supply chain; The anomaly detection and re-detection module is used to establish anomaly detection rules. When an anomaly exceeds a set threshold, a tag re-detection is triggered. The tag re-detection includes updating the initial tag and recalculating the tag propagation. The risk control module is used to take risk prevention and credit control measures in the refactoring business based on the labels of each enterprise. The measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating and issuing a financing freeze warning when the enterprise's accounts receivable period is abnormal.

[0022] According to another embodiment of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the data tag governance-based bank factoring business optimization method.

[0023] According to another embodiment of this application, a storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the bank refactoring business optimization method based on data tag governance.

[0024] As can be seen from the above technical solutions, the present invention has the following advantages: The bank refactoring business optimization method provided by this invention abstracts the relationships between enterprises, such as equity, transactions, and guarantees, into a graph structure. Initial tags cover financial, credit, and compliance aspects, and the tags are corrected using a tag propagation mechanism based on relationships, solving the problem that tags cannot reflect the real-time risk status of enterprises. Anomaly detection and threshold calibration address the false alarm / missed alarm issues caused by fixed thresholds. The entire process is controlled through dynamic credit limit adjustment, immediate freezing of funds in case of anomalies, flexible credit enhancement support, and risk resolution traceability.

[0025] By establishing an anomaly detection method and setting thresholds, tags are recalculated and propagated when an anomaly is triggered. If the authenticity of a transaction is questionable, a financing freeze alert is simultaneously triggered. This method constructs a complete process from data processing and relationship modeling to tag generation and risk control, improving the accuracy of risk identification, the transparency of risk transmission, and the refinement of risk control in bank factoring business, thereby enhancing service efficiency for upstream and downstream enterprises in the supply chain. Attached Figure Description

[0026] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 Flowchart of the optimization method for bank refactoring business based on data tagging governance; Figure 2 This is a diagram illustrating interbank reinsurance. Figure 3 A flowchart illustrating an implementation example of a data-label-based governance method for optimizing bank refactoring business; Figure 4 This is a schematic diagram of an electronic device. Detailed Implementation

[0028] The bank refactoring business optimization method involved in this application constructs an enterprise relationship graph based on internal and external data, establishes a labeling system and mapping rules, and generates initial labels, such as credit rating, transaction stability, and financial risk. Based on label propagation algorithms, node similarity matrices, and decay factors, a label propagation mechanism and model are established. The labels of upstream and downstream enterprises in the supply chain are updated. Based on the aforementioned enterprise labels, corresponding risk prevention and credit control measures are taken in the refactoring business with enterprises. For example, credit lines are dynamically adjusted according to changes in enterprise risk ratings, and financing is frozen in case of abnormal accounts receivable periods.

[0029] It should be noted that, as Figure 2 As shown, interbank refactoring refers to a commercial bank purchasing accounts receivable held by another bank, with the payer being a legal person or non-legal person organization. The subject of the transaction is the accounts receivable under the underlying transaction. The bank selling the accounts receivable is called the factoring bank, and the bank purchasing the accounts receivable is called the refactoring bank. Accounts receivable include both domestic and foreign currency accounts receivable purchased by the holding bank from the initial creditor, commercial factoring company, or other banks.

[0030] The Label Propagation Algorithm (LPA) is a semi-supervised learning algorithm that, based on the assumption that similar nodes tend to share the same label, propagates known labels (such as high-risk / low-risk) to unlabeled nodes (supply chain enterprises) through a graph structure. Introducing the LPA into the application of risk labeling for upstream and downstream enterprises in the supply chain of bank refactoring can effectively utilize the relationships within the supply chain network to achieve dynamic propagation and updating of risk labels.

[0031] The following describes in detail the bank refactoring business optimization method based on data tagging governance involved in this application. Specific details such as particular system architectures and technologies are presented for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details.

[0032] It should be understood that, when used in this specification, terms include indicating the presence of a described feature, integral, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. The terms include, encompass, have, and variations thereof mean including but not limited to, unless otherwise specifically emphasized.

[0033] The statements such as "one embodiment" or "some embodiments" described in this application mean that one or more embodiments of this application include the specific features, structures, or characteristics described in that embodiment. Therefore, the statements such as "in one embodiment," "in some embodiments," "in other embodiments," and "in still other embodiments" in this application do not necessarily refer to the same embodiment, but rather mean one or more, but not all, embodiments, unless otherwise specifically emphasized.

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Please see Figure 1 The diagram shows a flowchart of a bank refactoring business optimization method based on data tagging governance in a specific embodiment. The method includes: S101: Acquire internal and external data and perform preprocessing.

[0036] In some embodiments, the collected internal data includes historical transaction records of enterprises in the banking credit system, internal credit scores, and balance sheets and profit statements submitted by enterprises in the financial system; external data includes credit reports, enterprise registration change information from the industrial and commercial departments, and transaction statistics data from industry databases.

[0037] Preprocessing can involve data cleaning and standardization. Data cleaning removes duplicate transaction records, corrects typos in company names, and adds missing financial indicators. Standardization unifies different formats and standardizes the units of financial indicators. This improves the efficiency of subsequent steps and ensures the reliability of tag generation, risk control, and other processes.

[0038] S102: Construct and store an enterprise relationship graph based on the preprocessed data.

[0039] Specifically, in the enterprise relationship graph, enterprise entities are nodes, and behavioral data between enterprises are directed edges; the attributes of the nodes include basic enterprise information and business registration information; the behavioral data includes equity relationships, transaction transactions, and guarantee relationships; the weight of the directed edges is quantified according to the relationship strength, which includes transaction amount, guarantee ratio, equity amount and percentage.

[0040] It should be noted that, in addition to basic information and business registration information, node attributes also include the company's establishment time, registered capital, legal representative, and actual controller; the direction of directed edges clearly reflects the flow of relationships, and the weight of equity relationships is based on the shareholding ratio; a graph database is used for storage, supporting fast addition, deletion, modification, and querying of nodes and edges.

[0041] S103: Based on the preprocessed data, establish a labeling system and mapping rules, and generate initial labels.

[0042] Optionally, the labeling system includes five major categories and 27 subcategories: supply chain health index, credit index, financial indicators, industry sensitivity, and compliance; the initial labels include credit rating, transaction stability, and financial risk; for labeled enterprises, their label vectors are directly determined; for unlabeled enterprises, initial labels are assigned, which are random values ​​or global average labels.

[0043] This embodiment is based on standardized data and a preset labeling system. Through mapping rules, it transforms the multi-dimensional characteristics of enterprises into quantifiable initial labels, assigns clear label vectors to labeled enterprises, provides reasonable initial values ​​for unlabeled enterprises, forms the starting point for label dissemination, and ensures that the labels can initially reflect the basic risks and health status of enterprises.

[0044] S104: Based on the tag propagation algorithm, combined with the node similarity matrix and decay factor, a tag propagation method is established to update the tags of upstream and downstream enterprises in the supply chain.

[0045] In this embodiment, node similarity is calculated based on transaction size, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises. Through an asynchronous update strategy, the labels of unlabeled nodes are made to approach the weighted average of the labels of their neighboring nodes until the label distribution is stable.

[0046] Optionally, when calculating node similarity, the transaction size is calculated as the average of the two companies' mutual transaction volume over the past year / the average of their respective annual total transaction volume. The transaction frequency is calculated based on the number of transactions in the past 12 months. The most recent transaction time is calculated as the inverse of (current time - last transaction time) (the more recent the time, the larger the value). During asynchronous updates, the labels of core enterprises and other labeled nodes are fixed first, and then the new labels for unlabeled nodes are calculated as the sum of (neighboring node labels × similarity weights) / the sum of similarity weights. The criterion for judging iterative convergence is that the change in all node labels is less than the preset change threshold. If convergence is not achieved after multiple iterations, the process is forcibly stopped and the current label is adopted. Label propagation can reflect the transmission path of risks; iterative convergence ensures the stability and reliability of labels, improving the consistency and accuracy of labels across the entire supply chain.

[0047] S105: Establish outlier detection rules. When an outlier exceeds a set threshold, trigger tag re-detection. The tag re-detection includes updating the initial tag and recalculating the tag propagation. Optionally, the outlier includes a supplier's monthly transaction volume increasing by 500%.

[0048] The outlier detection rules in this embodiment include: a monthly transaction amount increasing by more than 300% compared to the average of the past 6 months; a liquidity ratio decreasing by more than 50% compared to the previous month; newly added execution information; and sudden termination of cooperation between enterprises and suppliers with whom they have cooperated for more than 5 years. The thresholds are set differently according to industry and enterprise type.

[0049] In this way, by monitoring a company's transactions, finances, credit, etc., when the indicators exceed the preset threshold, it is judged as abnormal and the label re-examination is initiated. The initial label and the propagation process are recalculated so that the label can reflect the sudden changes of the company in a timely manner.

[0050] In some specific embodiments, step S105 specifically includes the following steps: S1051: Based on the risk scenarios of reinfactoring business, define specific detection indicators covering transaction anomalies, financial anomalies, and compliance anomalies; including indicators such as monthly transaction volume growth rate and proportion of large transactions in the transaction dimension, sudden changes in accounts payable turnover days and net cash flow volatility in the financial dimension, and the number of failed trade background verifications and new regulatory penalty records in the compliance dimension.

[0051] Optionally, the monthly transaction volume growth rate = (this month's transaction volume - last month's transaction volume) / last month's transaction volume × 100%; the proportion of large transactions = (the sum of the amounts of a single transaction exceeding the set threshold) / this month's total transaction volume × 100%.

[0052] S1052: Adjust the trigger thresholds of each abnormal indicator based on industry cycle, enterprise type, and characteristics of historical abnormal events.

[0053] Optionally, threshold= Benchmark threshold × industry adjustment factor × Enterprise type coefficient; The industry adjustment coefficient is determined based on the industry prosperity index. ; The enterprise type coefficient represents an enterprise's risk resistance capability. A higher risk resistance capability corresponds to a more relaxed threshold, while a lower risk resistance capability corresponds to a tighter threshold.

[0054] S1053: When an anomaly is detected in a company, the tags of its upstream and downstream related companies are obtained through the company relationship graph to analyze whether the anomaly may be transmitted to related companies. For example, when a supplier's monthly transaction volume surges by 500%, the core company is checked for excessive payment instructions, and the accounts receivable collection risk tags of related companies are assessed to determine whether they need to be updated accordingly.

[0055] S1054: Based on the severity of the abnormal indicators, different label re-inspection processes are triggered in different levels.

[0056] If the monthly transaction volume increases by 200%, only the system will automatically review the label; if it increases by 400%, manual intervention is required to verify the transaction contract; if it increases by 500%, the relevant financing will be frozen immediately and a special audit will be initiated.

[0057] It can be seen that by monitoring indicators, adapting thresholds, linking upstream and downstream labels, and triggering different label re-examination processes according to the degree of abnormality, the company can accurately identify abnormal behavior of enterprises in the factoring business, and improve the comprehensiveness of risk identification through correlation analysis.

[0058] S106: Based on the enterprise's label, take risk prevention and credit control measures in the factoring business; the measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating, and automatically freezing financing when the enterprise's accounts receivable period is abnormal.

[0059] This embodiment directly links enterprise labels with specific risk control measures. Changes in labels trigger dynamic adjustments to the measures, ensuring that the level of control matches the actual risk level of the enterprise. By adjusting credit limits in advance, freezing credit, and enhancing creditworthiness, the probability of default in factoring business is reduced.

[0060] In some specific embodiments, step S106 specifically includes the following steps: Step S1061: Based on the latest combination of enterprise tags, retrieve the corresponding upper and lower limits of credit in the credit limit configuration table, and generate the credit limit range.

[0061] Step S1062: When the abnormal payment period indicator in the tag reaches the preset warning value, a freeze instruction is generated and synchronized to the monitoring cloud to ensure that the fund stoppage takes effect within the same accounting cycle.

[0062] Step S1063: Based on the company's list of pledgeable assets, the company's credit overflow records, and the re-factoring balance, push the credit enhancement plan to the financing party, along with the expected unlocking amount and time information.

[0063] Step S1064: When the enterprise's label falls back to the safe range and meets the credit enhancement conditions during the continuous observation period, execute the unfreezing process and generate an electronic report containing the reason for the freeze, the credit enhancement process, and the basis for the unfreezing.

[0064] This embodiment uses the latest combination of enterprise labels as the core basis, and determines a reasonable credit limit range by establishing a mapping relationship between labels and credit limits, thereby achieving precise allocation of credit limits. For risk signals such as abnormal payment periods, it promptly triggers freezing orders and ensures the timeliness of fund stoppage, quickly preventing the escalation of risk. It also pushes customized credit enhancement solutions based on the enterprise's pledgeable assets, credit history, and business balance. After the enterprise's risk is mitigated and conditions are met, the freeze is lifted in a standardized manner, and a complete process record is formed, creating a full-process mechanism from risk identification, control, and mitigation to closed-loop management. This ensures that risk control measures dynamically match the enterprise's actual risk status, improving the risk controllability and operational efficiency of the reinsurance business.

[0065] In one embodiment of the present invention, based on step S103, the following will provide a possible embodiment and its specific implementation will be described in a non-limiting manner. Step S103 specifically includes: S1031: Based on the enterprise's credit data, financial data, transaction data and compliance information, multiple judgment conditions are set, and each judgment condition corresponds to a specific indicator threshold; when an enterprise meets any combination of judgment conditions, it is assigned a corresponding hard label, which is used to directly identify the enterprise's high-risk or low-risk status.

[0066] S1032: Predict the risk probability of an enterprise and output a probability value that reflects the risk level of the enterprise as a soft label; S1033: Integrate hard and soft labels to generate the final risk index, calculate weights based on historical accuracy, and sum the hard and soft labels using the weights to obtain the final risk index; when the hard label identifies the risk as high and the company is on the blacklist, the final risk index is determined to be high risk.

[0067] It should be noted that the hard label generation adopts a Boolean rule algorithm, which judges the enterprise risk status through the logical combination of multiple judgment conditions. When any set of rule combinations is met, the hard label is high risk; otherwise, it is low risk. The judgment conditions are set based on specific indicator thresholds.

[0068] The soft tag generation uses a machine learning model to obtain a base value by accumulating the prediction results of multiple decision trees. Then, the base value is converted into a probability value between 0 and 1 by the sigmoid function. The higher the probability value, the higher the risk to the enterprise.

[0069] S1034: Set an outlier detection method. When the company's transaction data or other key indicators show abnormal changes exceeding the threshold, the initial label will be recalculated and the label fusion process will be re-examined to update the final risk index.

[0070] Optionally, the final risk index is calculated using a weighted fusion formula, which can be based on the sum of the weight multiplied by the soft label probability value and the hard label probability value. The weight is calculated based on the accuracy of the model in the previous period and the accuracy of the hard indicator rule in the previous period. The higher the accuracy of both, the greater the corresponding weight.

[0071] This embodiment directly identifies clear risk signals through hard indicator rules, uses model prediction to uncover potential risk trends, and then combines the advantages of both through dynamic weighting to form a corporate risk label that balances certainty and predictability. At the same time, it sets up an anomaly re-examination method to ensure that the label can respond to sudden changes in the corporate status, providing an accurate basis for risk assessment of refactoring business.

[0072] Step S103 in this embodiment also includes the following implementation method and steps: S2031: Extract enterprise credit data, supply chain transaction behavior data, and industry characteristic data, perform field validation and content extraction respectively, and form a standardized dataset that can be used for tag generation.

[0073] S2032: The labeling system is divided into a basic layer, a risk layer, and a compliance layer, with each layer further subdivided into 27 specific subcategories of labels.

[0074] Optionally, the basic layer involves fundamental attributes such as enterprise registration information and equity structure. The risk layer involves risk assessment indicators such as credit rating, financial risk, and transaction stability. The compliance layer involves regulatory red-line indicators, such as single-entity risk exposure limits and verification results of the authenticity of trade background.

[0075] S2033: For labeled companies, their label vectors are determined by verifying them through hard indicators such as historical bad debt rate and judicial litigation records; for unlabeled companies, initial labels that conform to the general characteristics of the industry are assigned based on the indicator distribution and cluster analysis results of companies in the same industry.

[0076] Optionally, the indicator distribution includes the industry average debt-to-equity ratio and current ratio. Cluster analysis results include the credit rating distribution of similar companies.

[0077] S2034: Adjust the tag trigger threshold according to industry risk fluctuations.

[0078] It should be noted that the statistical Z-score method is used to calculate the degree of deviation of the indicator.

[0079] Optionally, when the industry average accounts receivable delinquency rate increases by 5% compared to the benchmark, the delinquency rate threshold for the "poor transaction stability" label will be raised from 15% to 20%; at the same time, the effectiveness of the labels will be assessed by matching historical labels with actual risks, and the rules for labels with deviations exceeding 10% will be recalibrated.

[0080] The hard tag rule configuration method in this embodiment can use logical operations to combine judgment conditions. For example, the trigger conditions for a high-risk tag are an asset-liability ratio > 70% or the number of defaults in the past two years ≥ 2. The corresponding logical expression is: A hard label = 1 (high risk) if and only if (debt-to-equity ratio > 70%) ∨ (number of defaults in the past two years ≥ 2), otherwise it is 0 (low risk).

[0081] Combining the soft label probability P and the hard label result H, the weight β is adjusted based on the model's recent accuracy and rule accuracy. The final label is: Final label = βP + (1-β ) H.

[0082] This embodiment constructs a tagging system covering a company's basic attributes, risk characteristics, and compliance requirements by using credit records, supply chain transaction behavior, industry macro indicators, and historical risk cases. For already tagged companies, quantitative indicators are extracted from their historical data and mapped to the tagging system; for untagged companies, initial tags are generated using industry cluster analysis to ensure that the tags not only meet the requirements but also reflect the company's true risk level.

[0083] Furthermore, as a refinement and extension of the specific implementation method of the above-mentioned bank factoring business optimization method, in order to fully explain the specific implementation process in this embodiment, such as... Figure 3 As shown, the optimization method for bank reinsurance business also includes the following steps: Acquire and preprocess internal and external data. Establish a labeling system comprising 5 major categories and 27 subcategories, including supply chain health index, credit index, financial indicators, industry sensitivity, and compliance. Let the total number of enterprises be N, and the number of label categories be C, then the label matrix Y∈R N×C .

[0084] For a labeled company i, its label vector is Y. i For unlabeled company j, its label vector Y j An initial label can be assigned, such as a random value or a global average label.

[0085] The label is quantified by combining the calculation methods of hard and soft indicators, as well as model predictions and expert rules. For example, if a company's cash flow cycle is greater than the industry average of 2σ, then the company is labeled as having a long cash flow cycle.

[0086] Define a set of Boolean rules:

[0087] glx(X) is the judgment condition, such as the debt-to-asset ratio > 70%.

[0088] ###################################### rules=[ lambdax:x.credit_score<600,#Rule1:Lowcreditscore lambdax:x.last_year_default_times>=2,#Rule2:Frequentdefaults lambdax:x.legal_risk==True,#Rule3:Haslegalrisks lambdax:x.cash_flow_ratio<0.2,#Rule4:Poorcashflow lambdax:x.debt_to_equity>2.0,#Rule5:Highleverage lambdax:x.operating_margin<0.05,#Rule6:Lowprofitability lambdax:x.industry_risk_level=="high",#Rule7:High-riskindustry lambdax:x.audit_issues>=3,#Rule8:Multipleauditissues lambdax:x.management_turnover_rate>0.3,#Rule9:Highmanagementturnover lambdax:x.asset_quality_rating=="poor"#Rule10:Poorassetquality ...... ] defhard_label(x): if(x.debt_ratio>0.7)or(x.overdue_times>=2): return1#High Risk elif(x.cash_flow_ratio<0.2)and(x.industry_risk_level=='high'): return1 ...... else: return0#Low Risk ####################################### Output probability values ​​using the XGBoost model:

[0089]

[0090] ####################################### import xgboostasxgb params={ 'objective':'binary:logistic', 'eta':0.05, 'max_depth':6, 'subsample':0.8, 'lambda': 1.5# L2 regularization enhances anti-interference capabilities } dtrain=xgb.DMatrix(X_train,label=y_train) model=xgb.train(params,dtrain,num_boost_round=200) #Soft Tags proba=model.predict(xgb.DMatrix(x)) ####################################### Ultimately, the ability to provide refactoring services to enterprises and purchase accounts receivable depends on the risk index calculated by the final model. This risk index is a final risk score that integrates all the aforementioned quantitative labels and incorporates regulatory compliance rules.

[0091]

[0092] 𝛽∈[0,1] represents the credibility weight (suggested initial value 0.6), with dynamic adjustment rules:

[0093] ####################################### defget_final_label(x): #Soft Tags proba_x=proba #Hard Labels hard_x = hard_label(x) #Dynamic weights (Recent model AUC = 0.88, rule accuracy = 0.82) beta = 0.88 / (0.88 + 0.82) #Conflict Resolution ifhard==1andx['in_blacklist']: return1# Regulatory red line forcibly triggered else: returnbeta*proba+(1-beta)*hard ####################################### This embodiment performs an accuracy backtracking test, and the results are shown in Table 1.

[0094] Table 1: Accuracy Retrospective Indicator Information

[0095] Establish outlier detection rules. When an outlier exceeds a set threshold, such as a supplier suddenly experiencing a 500% increase in monthly transaction volume, a tag re-examination is triggered. This re-examination includes updating the initial tag and calculating tag propagation.

[0096] This embodiment also involves a tag propagation algorithm. In refactoring, key information such as transaction size, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises can be used to calculate the similarity between enterprise nodes.

[0097] A graph network construction algorithm is used to calculate the similarity between firms and construct a similarity matrix. For example, for firm i and firm j, their similarity is... The similarity formula is expressed as follows: Similarity = Transaction intensity * Transaction frequency * Time decay * Guarantee intensity * Guarantee frequency * Equity relationship * Compliance adjustment.

[0098] Specifically,

[0099] In the formula, : Total historical transaction amount between company i and company j; The largest single-sided transaction amount across the entire network; : Number of transactions (+1 to prevent zero frequency); Industry-adaptive decay rate; : Total historical guarantee amount of Enterprise i and Enterprise j; The largest guarantee amount across the entire network; : Number of guarantees (+1 to prevent zero frequency); : Current time - last transaction time (days); : Equity linkage strength (shareholding ratio / control coefficient); : Distance between equity levels (e.g., direct control = 1, indirect control >= 2); Equity decay coefficient; Equity weighting; Compliance coefficient (between 0 and 1).

[0100] Industry adaptive decay rate The calculation method is as follows ,in This refers to industry risk volatility.

[0101] in, The half-life of risk impact is set at 60, based on a comprehensive consideration of multiple factors such as international benchmarking, regulatory requirements, industry optimization needs, and supply chain finance reform.

[0102] industry default rate standard deviation over the past year

[0103] In the formula, Daily rate of change of the industry risk index; : Average daily rate of change of the industry risk index; 𝑇: Time observation window.

[0104] Table 2 shows the ways in which the coefficient k can be selected.

[0105] Table 2: Methods for determining the value of coefficient k

[0106] Equity decay coefficient In supply chain finance risk models, determining the equity decay coefficient μ requires combining empirical data and statistical methods on risk transmission across equity levels. Based on maximum likelihood estimation using historical corporate default data, μ is calculated to be 0.2, indicating moderate control and dependence on the supply chain.

[0107] Shareholding ratio for:

[0108] Compliance coefficient for:

[0109] The basic hard rules are The satisfaction status of the k-th regulatory rule (1 = satisfied) For example, R1: The authenticity of the trade background is verified (invoices and contracts match); R2: It does not involve enterprises on the negative list of the State Financial Supervision and Administration Bureau; R3: The risk exposure of a single enterprise is ≤15% of its net capital;

[0110] [High-risk transaction penalty coefficient, set to 0.3 here. High-risk transaction penalty] ; Specifically, it refers to the historical violation decay rate; Violation_count: the sum of the number of violations committed by enterprise i or j in the past 12 months.

[0111] The similarity matrix T is achieved through label propagation between nodes, causing the labels of unlabeled nodes to gradually approach the weighted average of the labels of their neighboring nodes until the label distribution stabilizes. This process is based on the principle of information diffusion in graph theory and satisfies the clustering assumption.

[0112] This embodiment uses an asynchronous update strategy to calculate the label-weighted average of the neighbors for each unlabeled enterprise j:

[0113]

[0114] For companies that have already been labeled, their labels will remain unchanged, which means that Y will be forced to remain unchanged. j Initial value.

[0115] Repeat until convergence or the maximum number of iterations is reached:

[0116] For each unlabeled node, the label with the highest probability in its final label distribution vector is selected as its predicted label. For unlabeled node j, its final label is:

[0117] This embodiment updates enterprise risk labels in real time based on the above method, providing early warnings of high-risk suppliers. Based on the propagated risk labels, it adjusts factoring financing amounts, interest rates, and fees. It also identifies risk transmission paths, such as how downstream enterprise risks affect upstream enterprises.

[0118] In the field of supply chain finance, this method can accurately assess transaction risks and provide a reliable basis for the rational allocation of funds, whether in accounts receivable financing, inventory pledge financing, or prepayment financing. In trade finance, it can efficiently process complex trade documents and transaction data, enabling rapid verification of the authenticity of trade backgrounds and reducing risks.

[0119] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0120] The following are embodiments of the bank refactoring business optimization system based on data tag governance provided in this disclosure. This system and the bank refactoring business optimization method based on data tag governance in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the bank refactoring business optimization system based on data tag governance, please refer to the embodiments of the bank refactoring business optimization method based on data tag governance described above.

[0121] The system includes: The data preprocessing module is used to acquire internal and external data and perform preprocessing. The relationship graph construction module builds and stores enterprise relationship graphs based on preprocessed data. The tag initialization module establishes a tag system and mapping rules based on the preprocessed data, and generates initial tags; The tag propagation and iteration module, based on the tag propagation algorithm, combines node similarity matrix and decay factor to establish tag propagation methods and update the tags of upstream and downstream enterprises in the supply chain; The anomaly detection and re-detection module is used to establish anomaly detection rules. When an anomaly exceeds a set threshold, a tag re-detection is triggered. The tag re-detection includes updating the initial tag and recalculating the tag propagation. The risk control module is used to take risk prevention and credit control measures in the refactoring business based on the labels of each enterprise. The measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating and issuing a financing freeze warning when the enterprise's accounts receivable period is abnormal.

[0122] like Figure 4As shown, this application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored in the memory and executable on the processor 101. When the processor 101 executes the program, it implements the steps of a bank refactoring business optimization method based on data tag governance.

[0123] In embodiments of the present invention, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments described and / or claimed herein.

[0124] In this embodiment, processor 101 may be implemented using at least one of an application-specific integrated circuit, a programmable logic device, a field-programmable gate array, a processor, a controller, a microcontroller, a microprocessor, or an electronic unit designed to perform the functions described herein. In some cases, such an implementation may be implemented within a controller. For software implementation, implementations such as processes or functions may be implemented with separate software modules that allow the performance of at least one function or operation. Software code may be implemented by a software application (or program) written in any suitable programming language, and the software code may be stored in memory and executed by the controller.

[0125] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like.

[0126] The memory 102 can be used to store software programs and various data. The memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0127] This application also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the data tag governance-based bank factoring business optimization method.

[0128] The storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0129] In a storage medium, a readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0130] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for optimizing bank refactoring business based on data tagging governance, characterized in that, The methods include: S101: Acquire internal and external data and perform preprocessing; S102: Construct and store an enterprise relationship graph based on the preprocessed data; S103: Based on the preprocessed data, establish a labeling system and mapping rules, and generate initial labels; S104: Based on the tag propagation algorithm, combined with the node similarity matrix and decay factor, establish a tag propagation method and update the tags of upstream and downstream enterprises in the supply chain; Node similarity is configured based on transaction size, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises; through an asynchronous update strategy, the labels of unlabeled nodes are made to approach the weighted average of the labels of their neighboring nodes until the label distribution is stable; S105: Establish outlier detection rules. When an outlier exceeds a set threshold, trigger tag re-detection. The tag re-detection includes updating the initial tag and recalculating the tag propagation. S106: Based on the labels of each enterprise, take risk prevention and credit control measures in the factoring business; the measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating, and issuing a financing freeze notice when the enterprise's accounts receivable period is abnormal.

2. The method for optimizing bank factoring business based on data tagging governance according to claim 1, characterized in that, In step S104, the node similarity is calculated based on the transaction scale, transaction frequency, most recent transaction time, compliance, shareholding amount, and shareholding ratio between enterprises; through an asynchronous update strategy, the labels of unlabeled nodes are made to approach the weighted average of the labels of their neighboring nodes until the label distribution is stable. For firm i and firm j, the similarity is... The formula is expressed as follows: in, : Total historical transaction amount between company i and company j; Maximum one-sided transaction amount; Number of transactions; Industry-adaptive decay rate; : Total historical guarantee amount of enterprise i and enterprise j; Maximum guarantee amount; Number of guarantees; : Current time - last transaction time; Equity linkage strength coefficient; Distance between equity levels; Equity decay coefficient; Shareholding weight; Compliance coefficient.

3. The method for optimizing bank factoring business based on data tagging governance according to claim 2, characterized in that, Industry adaptive decay rate The calculation method is as follows: ,in Industry risk volatility; in, The half-life of the risk impact; industry default rate standard deviation over the past year Daily rate of change of the industry risk index; : Average daily rate of change of the industry risk index; 𝑇: Observation window; k: Adjustment coefficient; Equity decay coefficient; Shareholding ratio The calculation method is as follows: Compliance coefficient The calculation method is as follows: Basic compliance items , indicating the satisfaction status of the k-th regulatory rule; [This refers to the penalty coefficient for high-risk transactions;] ; : represents the historical violation decay rate; Violation_count is the sum of the number of violations committed by enterprise i or j over the past 12 months; The label propagation algorithm propagates labels between nodes, causing the labels of unlabeled nodes to gradually approach the weighted average of the labels of their neighboring nodes until the label distribution stabilizes. For each unlabeled enterprise j, calculate the label-weighted average of its neighbors: For companies that have already been labeled, their labels will remain unchanged, and the initial value Y will be forcibly applied. ; Repeat until convergence or the maximum number of iterations is reached: ; For each unlabeled node, select the label with the highest probability in its final label distribution vector as its predicted label; Unlabeled node j final label for: 。 4. The method for optimizing bank factoring business based on data tagging governance according to claim 1 or 2, characterized in that, Step S103 specifically includes: Based on a company's credit data, financial data, transaction data, and compliance information, multiple judgment conditions are set, each corresponding to a specific indicator threshold; when a company meets any combination of judgment conditions, it is assigned a corresponding hard label. Predict the risk probability of enterprises and output a probability value that reflects the risk level of enterprises as a soft label; The final risk index is generated by merging hard and soft labels, and weights are calculated based on historical accuracy. The hard and soft labels are then weighted and summed to obtain the final risk index. An outlier detection method is set up so that when an enterprise's transaction data or other key indicators show abnormal changes exceeding the threshold, the initial labels are recalculated and the label fusion process is re-examined to update the final risk index.

5. The method for optimizing bank factoring business based on data tagging governance according to claim 1 or 2, characterized in that, Step S103 also includes: Extract enterprise credit data, supply chain transaction behavior data, and industry characteristic data, perform field validation and content extraction respectively, and form a standardized dataset for tag generation; The labeling system is divided into a basic layer, a risk layer, and a compliance layer, with each layer further subdivided into 27 specific subcategories of labels. For companies that have already been labeled, their label vectors are determined by verifying them using hard indicators such as historical bad debt rates and records of legal disputes; the label trigger thresholds are adjusted according to industry risk fluctuations.

6. The method for optimizing bank factoring business based on data tagging governance according to claim 1 or 2, characterized in that, Step S105 specifically includes: Based on the risk scenarios of reinfactoring business, specific detection indicators are defined to cover transaction anomalies, financial anomalies, and compliance anomalies. Adjust the trigger thresholds for each abnormal indicator based on industry cycles, company type, and characteristics of historical abnormal events; When an anomaly is detected in a company, the tags of its upstream and downstream related companies are obtained through the company relationship graph, and the anomaly is analyzed to see if it has spread to related companies. Based on the severity of the anomaly indicators, different tag re-inspection processes are triggered in stages.

7. The method for optimizing bank factoring business based on data tagging governance according to claim 1 or 2, characterized in that, Step S106 specifically includes: Based on the company's latest tag combination, retrieve the corresponding upper and lower credit limits from the credit limit configuration table and generate a credit limit range; When the abnormal payment period indicator in the tag reaches the preset warning value, a freeze command is generated and synchronized to the monitoring cloud. Based on the company's list of pledgeable assets, credit overflow records, and factoring balance, a credit enhancement plan is pushed to the financing party, along with the expected unlocking amount and time information; When a company's label falls back to a safe range and meets the credit enhancement conditions during the continuous observation period, the unfreezing process is executed, and an electronic report is generated.

8. A bank refactoring business optimization system based on data tagging governance, characterized in that, The system is used to implement the bank refactoring business optimization method based on data tag governance as described in any one of claims 1 to 7; The system includes: The data preprocessing module is used to acquire internal and external data and perform preprocessing. The relationship graph construction module builds and stores enterprise relationship graphs based on preprocessed data. The tag initialization module establishes a tag system and mapping rules based on the preprocessed data, and generates initial tags; The tag propagation and iteration module, based on the tag propagation algorithm, combines node similarity matrix and decay factor to establish tag propagation methods and update the tags of upstream and downstream enterprises in the supply chain; The anomaly detection and re-detection module is used to establish anomaly detection rules. When an anomaly exceeds a set threshold, a tag re-detection is triggered. The tag re-detection includes updating the initial tag and recalculating the tag propagation. The risk control module is used to take risk prevention and credit control measures in the refactoring business based on the labels of each enterprise. The measures include dynamically adjusting the credit limit according to changes in the enterprise's risk rating and issuing a financing freeze warning when the enterprise's accounts receivable period is abnormal.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the bank refactoring business optimization method based on data tag governance as described in any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the bank refactoring business optimization method based on data tag governance as described in any one of claims 1 to 7.