Automobile repayment method and system, electronic equipment and storage medium

By combining information entropy processing, multi-dimensional decision trees, and Monte Carlo simulation, the static nature of traditional auto finance repayment strategies and insufficient credit assessment are solved. This enables dynamic risk-return assessment and smart contract execution, thereby improving decision-making efficiency and risk control in auto finance business.

CN120823067APending Publication Date: 2025-10-21BEIJING JINHUI TECH CO LTD

Patent Information

Application Number
CN202510975463.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Traditional auto finance repayment strategies rely on human experience and static rules, which cannot dynamically adapt to differences in sales models or vehicle models. The credit assessment system lacks time-series analysis, resulting in delayed risk warnings and unoptimized strategy recommendations.

Method used

The system uses a natural language processing algorithm based on information entropy to analyze the correlation between policy texts and financial payment data, constructs a multi-dimensional decision tree rule base, evaluates the income distribution through a Monte Carlo simulation engine, calculates credit scores using a machine learning model, and generates a smart contract to execute the payment process.

Benefits of technology

It has achieved standardization of policy terms and data integration, dynamically assessed the risk and return of repayment strategies, accurately identified credit risk levels, ensured that contract terms match credit ratings, and improved decision-making efficiency and risk control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823067A_ABST
    Figure CN120823067A_ABST
Patent Text Reader

Abstract

The invention discloses an automobile money returning method and system, electronic equipment and a storage medium, and relates to the field of automobile finance. The method comprises the following steps: firstly, acquiring a policy text and financial refunding flow data, analyzing the policy text through a natural language processing algorithm, and associating the policy text with the financial refunding flow data to generate a structured data set; and then, constructing a decision tree rule base, and randomly sampling a rebate period and a rebate proportion by using a Monte Carlo simulation engine to generate a profit distribution map. And then dealer qualification data and money return behavior data are obtained, a credit score value is calculated through a machine learning model, and a target scheme is screened from the income distribution diagram according to the credit score value. And finally, based on the target scheme, an electronic contract matched with the credit rating of the dealer is automatically generated through an intelligent contract, and a money return process is executed. By implementing the technical scheme provided by the invention, a solution integrating policy analysis, risk income prediction and credit assessment can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of automobile finance, and in particular to a method, system, electronic device and storage medium for automobile payment collection. Background Art

[0002] Amid the digital transformation of auto finance, optimizing collection strategies and managing dealer credit have become core challenges for OEMs in improving capital efficiency and controlling risk. As market competition intensifies, companies must simultaneously address complex variables such as regional economic disparities, accelerated vehicle model iterations, and diversified sales models (e.g., the coexistence of direct sales, agency, and distribution).

[0003] In the auto finance sector, traditional collection strategy development and dealer credit management rely heavily on manual experience and static rules, resulting in significant technical flaws. Policy text parsing typically relies on manual entry, storing terms like payment terms and rebate ratios in isolated tables. Decisions are then made based on only a single parameter (such as region). This disconnects policy terms from actual collection data and prevents dynamic adaptation to sales models or vehicle models. Credit assessment systems are limited to static qualification data (such as registered capital) or simple overdue payment statistics, lacking dynamic behavioral characteristics such as collection cycle volatility. Inadequate time series analysis capabilities lead to delayed risk warnings. Strategy recommendations are disconnected from risk-return assessments, lacking a unified framework for linking credit ratings and return distribution. This can lead to suboptimal policy combinations (such as high rebates and long payment terms), triggering a risk-return imbalance. This fragmented and static technical approach makes it difficult for companies to achieve a dynamic risk-return balance in a complex market environment.

[0004] Therefore, there is an urgent need for a solution that integrates policy analysis, risk-return forecasting, and credit assessment. Summary of the Invention

[0005] The present application provides a method, system, electronic device and storage medium for automobile repayment, providing a solution that integrates policy analysis, risk-return prediction and credit assessment.

[0006] In a first aspect of the present application, a method for collecting automobile payments is provided, which adopts the following technical solution: Obtaining policy text and financial collection flow data, parsing the policy text using a natural language processing algorithm based on information entropy, and associating the parsed results with the financial collection flow data to generate a structured dataset. The structured dataset includes the number of days in the payment period, vehicle type, rebate ratio, and penalty rate for breach of contract. Construct a decision tree rule base based on preset multi-dimensional rules, including sales model, capital attributes and regional characteristics; Based on the structured data set and the decision tree rule base, randomly sampling the collection period and the rebate ratio in the financial collection flow data through a Monte Carlo simulation engine to generate a profit distribution graph; Obtaining dealer qualification data and payment collection behavior data, calculating a credit score using a pre-trained machine learning model, and screening a target solution from the revenue distribution map based on the credit score; Based on the target solution, an electronic contract matching the dealer's credit rating is automatically generated through a smart contract and a payment collection process is executed.

[0007] By adopting the above technical solutions, a natural language processing algorithm based on information entropy accurately links unstructured policy texts with financial collection flow data, generating a structured data set containing core elements such as payment period days and vehicle type, significantly improving the efficiency and standardization of data integration. Secondly, through the combination of a decision tree rule base constructed through multi-dimensional rules and a Monte Carlo simulation engine, it is possible to quantitatively evaluate the profit distribution of different collection strategies, dynamically simulate risk scenarios, and provide a scientific basis for optimizing collection cycles and rebate ratios. At the same time, a credit scoring mechanism based on a machine learning model can accurately identify the credit risk level of dealers and, combined with the profit distribution map, screen out target solutions with controllable risks and optimal returns, thereby reducing the probability of default. Finally, the automated execution of smart contracts seamlessly connects electronic contract generation with the collection process, eliminating manual operation delays and errors, ensuring that contract terms strictly match credit ratings, and realizing transparent and traceable intelligent collection management throughout the entire process, comprehensively improving the decision-making efficiency, risk control capabilities, and execution reliability of auto finance business.

[0008] Optionally, parsing the policy text and associating the parsing results with the financial payment flow data to generate a structured data set specifically includes: Performing a preset keyword matching operation on the policy text to obtain a structured table, wherein the preset keywords include policy ID, number of days in the account period, rebate ratio, penalty rate for breach of contract, and applicable vehicle model; The policy ID is matched with the collection records in the financial collection flow data, and the actual collection rate and bad debt rate of each policy are calculated to obtain the structured data set.

[0009] By adopting the above technical solution, first, the policy text is accurately matched and extracted through preset keywords, and a structured table containing key fields is quickly generated, which greatly reduces the cost of manual analysis and avoids subjective errors, ensuring the standardization and consistency of policy information; secondly, through the automatic association of policy ID and financial collection flow data, the actual collection rate and bad debt rate of each policy are directly quantified, and abstract policy terms are converted into analyzable digital indicators, providing highly reliable data support for subsequent strategy optimization.

[0010] Optionally, the decision tree rule base is constructed based on preset multi-dimensional rules, specifically including: Using the regional features as the root node of the decision tree, generating an initial decision tree structure; Bind the sales model and vehicle type classification constraints under the root node to form a secondary leaf node; Configuring policy parameters matching the multi-dimensional rules under the second-level leaf node to generate a third-level leaf node, wherein the policy parameters include the number of days in the account period, the rebate ratio, the prepayment ratio, and the penalty rate for breach of contract; The inventory turnover rate and the regional economic index are obtained in real time through an external data interface, the weights and constraints of all nodes of the decision tree are updated, and the decision tree rule base is generated. All nodes include the root node, the second-level leaf node and the third-level leaf node.

[0011] By adopting the above technical solution, a decision tree is constructed with regional characteristics as the root node, and regional economic differences are used as the core basis for strategy formulation to ensure that the collection strategy is deeply adapted to regional characteristics; secondly, by binding sales models, vehicle classifications and specific policy parameters step by step, a hierarchical decision rule library is formed, and complex business logic is broken down into quantifiable and traceable node relationships, greatly improving the transparency and explainability of the strategy.

[0012] Optionally, based on the structured data set and the decision tree rule base, randomly sampling the collection period and the rebate ratio in the financial collection flow data through a Monte Carlo simulation engine to generate a profit distribution map, specifically including: Extracting the value ranges of the number of days in the account period, the rebate ratio, and the penalty rate for breach of contract from the structured data set to construct the variable constraint boundaries of the Monte Carlo simulation engine; Generate a sampling parameter set based on the regional characteristics and vehicle type classification rules of the decision tree rule base; Based on the sampling parameter set, performing a first number of random samplings on the collection period and the rebate ratio by the Monte Carlo simulation engine to calculate a net income value and an income standard deviation; The net income value and the income standard deviation are mapped into an income distribution graph.

[0013] By adopting the above technical solution, firstly, based on the variable constraint boundaries of the structured data set, a scientific and reasonable parameter boundary is provided for the Monte Carlo simulation engine, avoiding the distortion of simulation results due to sampling range deviation, and significantly improving the accuracy of risk-return assessment; secondly, the sampling parameter set is generated through the regional characteristics and vehicle classification rules of the decision tree rule library, so that the simulation process is closely aligned with the actual business scenario, ensuring that the sampling results are business interpretable; at the same time, through random sampling, the net income fluctuation range and standard deviation of different collection period and rebate ratio combinations are quantified, intuitively presenting the distribution pattern of high returns accompanied by high risks, providing automobile companies with a decision-making basis for the "return-risk" balance.

[0014] Optionally, performing a first number of random samplings on the collection period and the rebate ratio by the Monte Carlo simulation engine to calculate the net income value and the income standard deviation specifically includes: Obtaining a single vehicle profit from a vehicle model cost database, matching it with the vehicle type in the structured dataset, and obtaining a single vehicle sales volume corresponding to the single vehicle profit; Calculating the cost of capital according to the preset weighted average cost of capital of the server, and calculating the bad debt loss based on the bad debt rate and the default penalty rate; Calculating a net income value based on the bicycle profit, the bicycle sales volume, the capital cost, and the bad debt loss; The net income average of the first number of samplings is calculated, and the income standard deviation is calculated based on the net income average and the net income value of each sampling.

[0015] By adopting the above technical solution, firstly, through the precise matching of the vehicle model cost database and the structured data set, the profit of each vehicle is dynamically bound to the sales volume, and the actual profit contribution of different vehicle models is quantified, thus avoiding the profit deviation caused by the vague classification of vehicle models in traditional estimation methods; secondly, the capital cost is calculated based on the weighted average cost of capital, and the bad debt loss is accurately calculated by combining the bad debt rate and the default penalty rate, covering the hidden costs of capital utilization and default risk, so that the net income calculation is more in line with the actual financial scenario.

[0016] Optionally, obtaining the dealer's qualification data and payment collection behavior data, calculating a credit score using a pre-trained machine learning model, and screening a target solution from the revenue distribution map based on the credit score may include: Extracting the dealer's qualification data, including registered capital, years of cooperation, and regional economic index, and converting it into a numerical feature vector; Extracting the dealer's payment collection behavior data, including the on-time payment collection rate, the number of overdue payments, and the fluctuation rate of the payment collection cycle, to generate a time series behavior feature vector; Inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value; Classifying the credit score value into high, medium and low levels according to the credit score value and a preset credit level threshold; Based on the divided credit score values, a set of matching solutions is matched from the benefit distribution graph to obtain the target solution.

[0017] By adopting the above technical solution, by converting the dealer's qualification data and dynamic collection behavior data into numerical feature vectors and time series feature vectors respectively, we can achieve structured integration of multi-dimensional data, provide a comprehensive and fine-grained input basis for credit scoring, and avoid the one-sidedness of single-dimensional evaluation; secondly, the pre-trained machine learning model accurately quantifies the dealer's credit risk by integrating static qualifications and dynamic behavior characteristics, generating an interpretable credit score value, replacing traditional manual experience judgment, and significantly improving evaluation efficiency and objectivity.

[0018] Optionally, inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value specifically includes: Extracting time series features from the time series behavior feature vector, and calculating the collection cycle volatility and overdue frequency through a sliding window to generate a behavior statistics vector; Concatenating the numerical feature vector with the behavioral statistical vector to form a multi-dimensional fusion feature vector; Inputting the fused feature vector into a pre-trained machine learning model, and outputting a predicted value of the probability of default through a tree structure hierarchical decision; The default probability prediction value is converted into the credit score value based on a nonlinear mapping function.

[0019] By adopting the above technical solution, through sliding window analysis of time series behavior feature vectors, key time series indicators such as collection cycle volatility and overdue frequency are dynamically extracted, the dealer's performance behavior pattern is accurately captured, and the timeliness and granularity of behavioral risk prediction are effectively improved; secondly, by fusing static qualification data and dynamic behavior statistical vectors into multi-dimensional features, the limitation of the separation of static data and dynamic behavior in traditional credit assessment is broken through, and a fusion feature system that comprehensively reflects the dealer's credit risk is constructed; at the same time, the machine learning model based on tree-structured hierarchical decision-making makes interpretable hierarchical judgments on the fusion features, outputs a high-confidence default probability prediction value, and avoids the decision-making blind spots of the "black box model".

[0020] In a second aspect of the present application, a system for collecting automobile payments is provided, specifically comprising: A policy data parsing module is used to obtain policy text and financial collection flow data, parse the policy text using a natural language processing algorithm based on information entropy, and associate the parsing results with the financial collection flow data to generate a structured data set. The structured data set includes the number of days in the account period, vehicle type, rebate ratio, and penalty rate for breach of contract. A multi-dimensional rule configuration module is used to build a decision tree rule library based on preset multi-dimensional rules, wherein the multi-dimensional rules include sales model, capital attributes and regional characteristics; A revenue prediction module, configured to randomly sample the collection period and the rebate ratio in the financial collection flow data based on the structured data set and the decision tree rule base through a Monte Carlo simulation engine to generate a revenue distribution graph; A credit scoring module is used to obtain dealer qualification data and payment collection behavior data, calculate a credit score value through a pre-trained machine learning model, and screen target solutions from the profit distribution map based on the credit score value; The electronic contract generation and execution module is used to automatically generate an electronic contract that matches the dealer's credit rating through a smart contract based on the target solution and execute the collection process.

[0021] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.

[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, any one of the methods described above is executed.

[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: First, a natural language processing algorithm based on information entropy accurately links unstructured policy text with financial collection flow data into a structured dataset, overcoming the limitations of traditional manual analysis, which is inefficient and prone to large errors, and achieving standardization of policy terms. Second, by combining a decision tree rule base constructed from multi-dimensional rules with a Monte Carlo simulation engine, it quantitatively evaluates the benefit-risk distribution of different collection strategies and dynamically simulates the impact of market fluctuations on the capital chain, providing automakers with a scientific basis for decision-making and avoiding the rigidity of experience-driven strategies. Furthermore, a credit scoring mechanism based on a pre-trained machine learning model integrates dealer qualification data with dynamic collection behavior data. Through time-series feature extraction and tree-structured hierarchical decision-making, it generates highly reliable credit scores, accurately identifying high-risk dealers. Combined with the profit distribution map, it selects target solutions with manageable risks and optimal returns, achieving differentiated strategy matching. Finally, smart contracts automatically generate electronic contracts strictly tied to credit ratings and execute the collection process, ensuring transparency and auditability throughout the entire process. This results in a solution that integrates policy analysis, risk-return prediction, and credit assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a schematic diagram of the system architecture of an embodiment of a method or system for collecting automobile payments using the present application; Figure 2 This is a flow chart of a method for collecting automobile payments disclosed in an embodiment of the present application; Figure 3 This is a module diagram of a car payment collection system disclosed in an embodiment of the present application; Figure 4 This is a structural diagram of an electronic device disclosed in an embodiment of the present application.

[0025] Explanation of the accompanying drawings: 301, policy data analysis module; 302, multi-dimensional rule configuration module; 303, profit prediction module; 304, credit scoring module; 305, electronic contract generation and execution module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0027] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0028] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0030] Terminal devices 101, 102, and 103 can be hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Group Audio Layer 3) players, MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Group Audio Layer 4) players, laptop computers, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or they can be implemented as a single software or software module. No specific limitation is made here.

[0031] When terminals 101, 102, and 103 are hardware, they may also be equipped with a video capture device. The video capture device may be any device capable of capturing video, such as a camera, a sensor, and the like. Users can use the video capture device on terminals 101, 102, and 103 to capture video.

[0032] The server 105 may be a server that provides various services, such as a background server that processes data displayed on the terminal devices 101, 102, and 103. The background server may analyze and process the received data, and may feed back the processing results (e.g., recognition results) to the terminal device.

[0033] It should be noted that a server can be either hardware or software. When a server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When a server is software, it can be implemented as multiple software programs or software modules (for example, multiple software programs or software modules used to provide distributed services), or as a single software program or software module. This is not specifically limited here.

[0034] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is merely illustrative. Any number of terminal devices, networks, and servers may be used as needed. In particular, if target data does not need to be acquired remotely, the above system architecture may not include a network, but may instead include only terminal devices or servers.

[0035] This embodiment discloses a method for collecting automobile payments. Figure 2 This is a flow chart of a method for collecting automobile payments disclosed in an embodiment of the present application. Figure 2 As shown, the method includes the following steps: S201. Obtain a policy text and financial collection flow data, parse the policy text using a natural language processing algorithm based on information entropy, and associate the parsed results with the financial collection flow data to generate a structured data set, wherein the structured data set includes the number of days in the account period, vehicle type, rebate ratio, and penalty rate for breach of contract; First, obtain policy texts and financial collection flow data in batches through the API (application programming interface) interface or database, unify the format of the policy text (such as converting it into a UTF-8 encoded TXT file) and clean up missing values ​​and abnormal records in the financial data; then, use the TF-IDF (term frequency-inverse document frequency) algorithm combined with information entropy calculation (based on the Shannon entropy formula to screen high-information vocabulary) to extract keywords in the policy text (such as payment period days and rebate ratio), build a domain dictionary, and parse the policy terms based on the rule engine (regular expression) or sequence labeling model (such as BiLSTM-CRF (bidirectional long short-term memory network-conditional random field)) to extract payment period days, vehicle type , rebate ratio, and penalty rate, etc., to generate a preliminary structured table; then, the parsed results are matched with the financial collection flow data through SQL (Structured Query Language) or Pandas (Python data analysis library), and the actual collection rate (total collection amount / total collection amount) and bad debt rate (overdue collection amount / total collection amount) of each policy are calculated. The data consistency is ensured through information entropy thresholds (such as entropy value > 2.0) and cross-validation (such as comparison with historical data); finally, the associated data is stored in JSON (lightweight data exchange format) or database table format to generate a structured data set containing field types, value ranges, and statistical indicators for downstream modules to call.

[0036] Optionally, parsing the policy text and associating the parsing result with the financial payment collection flow data to generate a structured data set includes: Performing a preset keyword matching operation on the policy text to obtain a structured table, wherein the preset keywords include policy ID, number of days in the account period, rebate ratio, penalty rate for breach of contract, and applicable vehicle model; The policy ID is matched with the collection records in the financial collection flow data, and the actual collection rate and bad debt rate of each policy are calculated to obtain the structured data set.

[0037] Specifically, the policy text must first be obtained from the enterprise ERP system or document management system. For scanned PDF files, text recognition must be performed using OCR tools (such as Tesseract) to convert them into parseable text content. After data preprocessing, regular expression rules are used to match preset keywords. For example, the policy ID is extracted using the fixed format "POL-\d{4}-\d{4}", such as "POL-2023-1001". "POL" is a fixed prefix, indicating the policy category identifier. The first "\d{4}" is the year the policy was issued, and the second "\d{4}" is the specific policy serial number, which is spliced ​​with "-"; the number of days in the account period is obtained by matching the value associated with "account period". The rule is "account period\D*(\d+)\D* days". "Account period" is a literal character, and the keyword "\D*" is located. Non-numeric characters are skipped, and "(\d+)" captures the number of days in the payment period. "Day" is a fixed unit, constraining the matching range. The rebate ratio is captured by capturing the percentage value in the phrase "rebate ratio ≥ 5%" using the rule "rebate ratio\D*([\d.]+)%." "[\d.]+" matches integers or decimals, and "%" matches the percent sign to ensure the correct data type. The penalty rate is extracted using "penalty rate\D*([\d.]+)%." Applicable vehicle models require Named Entity Recognition (NER) technology to accurately locate the vehicle model description (e.g., "new energy vehicle model") from the clause. For semantically ambiguous clauses (e.g., negotiable extension of the payment period), the BERT model is used for contextual analysis to extract implicit constraints (e.g., the maximum extension period). The extracted data must be correlated and verified with the actual collection rate and bad debt rate in the financial system, and manually reviewed using labeling tools (such as Label Studio) to correct matching errors. Finally, the structured data (including fields such as policy ID and payment period days) must be output in Excel or database table format to ensure data traceability and support subsequent analysis.

[0038] Furthermore, the structured policy data, which includes fields such as the policy ID and the number of days in the payment period, is left-linked to the financial system's payment flow sheet (including fields such as the policy ID, receivable amount, actual amount collected, receivable date, actual collection date, and overdue status) via the policy ID to ensure that each payment collection record is accurately linked to the corresponding policy. For this linked data, invalid records (such as entries with missing policy IDs or incorrect formats) must be cleaned and the field legitimacy verified. Next, the overdue status is dynamically calculated based on the number of days in the payment period and the payment collection date: if the actual collection date is earlier than or equal to the receivable date plus the number of days in the payment period, the payment is marked as "on time payment"; if the actual collection date is later than the number of days in the payment period but within the grace period (such as 30 days), the payment is marked as "overdue"; if the payment is still not collected after the grace period, the payment is marked as "bad debt." Key metrics are then aggregated and calculated by grouping by policy ID: the actual collection rate is the ratio of the sum of the actual collection amounts for all "on-time collection" and "overdue" records under that policy to the total receivables, and the bad debt rate is the ratio of the total receivables for "bad debt" records. To ensure calculation accuracy, scripts (such as Python) are used to verify logical consistency (for example, the sum of the collection rate and bad debt rate does not exceed 100%). Alerts are triggered for abnormal data (such as negative collection values ​​or zero payment period days), and annotation tools are used to manually review and correct them. The resulting structured dataset includes fields such as policy ID, payment period days, actual collection rate, bad debt rate, and number of associated orders. It can be exported as an Excel file or written to a database and synchronized to the enterprise data warehouse using ETL (Extract-Transform-Load) tools. BI (Business Intelligence) tools can also be used to build visual dashboards to analyze collection health across multiple dimensions, such as policy and time period.

[0039] S202: Construct a decision tree rule base based on preset multi-dimensional rules, wherein the multi-dimensional rules include sales model, capital attributes, and regional characteristics; Specifically, first, define the input logic for multi-dimensional rules, clarifying the value ranges for each dimension, including sales model (e.g., direct sales, distribution), funding attributes (e.g., full payment, installment), and regional characteristics (e.g., East China, South China, or economic tiers based on GDP). Priorities and combination conditions between rules are set (e.g., "IF Sales Model = Distribution AND Region = East China, THEN Account Period ≤ 60 Days"). Next, construct rule templates and decision tree structures, using IF-THEN logic to describe rules. For example, "IF Sales Model = Distribution AND Funding Attribute = Installment AND Region = First-tier City THEN Rebate Percentage ≥ 5%." Rules are then converted into decision tree nodes: each non-leaf node represents a dimensional condition (e.g., "Is the region East China"), and a leaf node represents the final decision outcome (e.g., number of days in the account period, rebate percentage). Next, implement the storage and dynamic loading of the rule base, storing the rules in a structured format (e.g., JSON or XML) in a database (e.g., MySQL) or a rule engine (e.g., Drools). A rule configuration interface is also developed to allow business personnel to dynamically edit rules using a visual tool, ensuring maintainability. Rule conflicts and overlays require preset resolution strategies, such as sorting by priority or overwriting historical rules by the most recent update time, and verifying the logical consistency of rules through automated test scripts.

[0040] Optionally, the building of a decision tree rule base based on preset multi-dimensional rules includes: Using the regional features as the root node of the decision tree, generating an initial decision tree structure; Bind the sales model and vehicle type classification constraints under the root node to form a secondary leaf node; Configuring policy parameters matching the multi-dimensional rules under the second-level leaf node to generate a third-level leaf node, wherein the policy parameters include the number of days in the account period, the rebate ratio, the prepayment ratio, and the penalty rate for breach of contract; The inventory turnover rate and the regional economic index are obtained in real time through an external data interface, the weights and constraints of all nodes of the decision tree are updated, and the decision tree rule base is generated. All nodes include the root node, the second-level leaf node and the third-level leaf node.

[0041] Specifically, first, determine the regional division dimensions according to business needs (such as the administrative regions of "East China, South China" or the economic indicators of "high GDP (gross domestic product), medium GDP, low GDP"), define a unique code for each region and record the regional attributes; then, use JSON (JavaScript Object Notation, a lightweight data exchange format) or Python dictionary to design a tree data structure, mark the root node type as "root", and set the branch condition as the regional code; then, store the generated decision tree structure in a relational database (such as MySQL), and record the node ID, parent node ID, node type and condition description through the database table to ensure that the hierarchical relationship between parent and child nodes is traceable; finally, use the visualization tool Graphviz (a graph visualization software) to convert the tree structure into a schematic diagram, verify the integrity and logical correctness of the regional branches, and ensure that all preset areas are covered and there are no omissions.

[0042] Furthermore, first, in the decision tree node table of the relational database (such as MySQL, a relational database management system), the fields are expanded to add two columns, sales_mode (sales mode) and vehicle_type (vehicle type classification), to store the constraint conditions (such as direct sales / distribution and fuel vehicle / new energy vehicle) respectively; then, through SQL (Structured QueryLanguage (Structured Query Language) creates second-level child nodes for each regional root node. For example, under the East China region, add a sales model branch point {"node_id":"sales_EC01_D","parent_id":"root_EC01","node_type":"sales","condition":"sales_mode=Distribution","children":[]}, and further bind a vehicle model classification child node {"node_id":"vehicle_EC01_D_EV","parent_id":"sales_EC01_D","node_type":"vehicle","condition":"vehicle_type=New Energy Vehicle"} to it. Next, use a script to traverse all regional root nodes and verify that each region contains a complete sales model and vehicle model classification combination (such as direct sales + fuel vehicles) to prevent missing or redundant branches. Finally, use Graphviz to regenerate the decision tree diagram to verify the accuracy of the logical hierarchy and binding relationships of the second-level nodes to ensure full coverage of business rules.

[0043] Furthermore, the policy parameter table is expanded in the relational database, with new fields such as term_days (account period), rebate_rate (rebate ratio), prepay_rate (prepayment ratio), and penalty_rate (default penalty rate). These fields are linked to the decision tree node table through node_id (node ​​unique identifier). Subsequently, for each second-level node that combines sales model and vehicle classification (such as the new energy vehicle node under the East China regional distribution model), a third-level leaf node is created and bound to specific parameters. The sample data insertion statement is INSERT INTO policy_params (node_id, term_days,rebate_rate) VALUES ('node_103', 60, 5.0), indicating that the account period under this node is 60 days and the rebate ratio is 5%. Next, a Python script is used to implement dynamic parameter matching logic, querying the region, sales model, and vehicle classification path layer by layer from the root node, ultimately hitting the third-level leaf node and returning the parameters. The sample code queries the database through the ORM (Object-Relational Mapping) framework, with the input conditions being region=East China, When sales_mode = distribution and vehicle_type = new energy vehicle, the corresponding payment period and rebate value are output. Then, an integrity verification script is written to ensure that there must be at least one third-level leaf node under each second-level node and the parameter field is not empty. Otherwise, an alarm is triggered and the missing item is recorded. Finally, Graphviz is used to generate a decision tree diagram with parameter annotations to verify the matching of the third-level node parameters with business rules. For example, whether the new energy vehicle node is bound to a longer payment period or a higher rebate strategy, to ensure that the policy configuration meets business expectations.

[0044] Furthermore, an API call module is designed to connect to the inventory turnover rate interface and third-party regional economic index interface of the warehouse system in real time. The latest data is regularly obtained and parsed into structured numerical values ​​through the Python requests library. Subsequently, the fields are expanded in the decision tree node table of the relational database, adding economy_index (economic index), inventory_turnover (inventory turnover rate), and weight (weight value). The weights of all nodes are updated according to the preset formula (such as weight = economic index × 0.6 + inventory turnover rate × 0.4). The sample SQL statement is UPDATE decision_nodes SET weight = 0.6 * %s + 0.4 * %s WHERE node_id = %s; At the same time, constraints are dynamically adjusted (for example, when the economic index falls below a threshold, the payment period of the third-level leaf node is automatically modified). The policy parameter table is updated through SQL. Then, the message queue Kafka (a distributed stream processing platform) is used to publish rule change events, triggering the business system to reload the decision tree rules to ensure real-time effectiveness. Finally, the database transaction log records the change history of each weight and parameter, supports rolling back abnormal operations by time point, and generates a weighted decision tree diagram using Graphviz to verify the impact of the economic index and inventory turnover rate on the node logic (for example, whether regions with a weight greater than 70% are preferentially matched with a relaxed payment period policy), ensuring that the rule base dynamically responds to external data changes.

[0045] S203: Based on the structured data set and the decision tree rule base, randomly sampling the collection period and the rebate ratio in the financial collection flow data through a Monte Carlo simulation engine to generate a profit distribution graph; Specifically, based on a structured data set (including fields such as historical collection period, rebate ratio, profit per vehicle, sales volume and bad debt rate) and a decision tree rule library (XML / JSON format rule file), the Monte Carlo simulation engine (a numerical simulation method based on random sampling) first initializes the simulation parameters, sets the collection period to be randomly sampled according to a uniform distribution within 30 to 90 days, and the rebate ratio to be concentratedly sampled according to a normal distribution between 1% and 10% (the probability is higher near the mean), and presets an annualized cost of capital rate of 5% and the number of simulations to 10,000 times; then, a random number matrix is ​​generated through Python's NumPy library, and combinations of collection period and rebate ratio (such as 60 days + 3% or 45 days + 8%) are extracted one by one, combined with the profit per vehicle in the data set. The system dynamically calculates the net income value of each sampling based on the sales data, specifically by deducting rebate costs from total revenue, converting interest on funds occupied based on the collection period, and subtracting bad debt losses. Finally, it summarizes all simulation results and calculates the mean, standard deviation, and quantile of the net income to identify high-return, low-risk strategy intervals (such as the top 10% with high returns and risks below the mean). Finally, it uses Matplotlib to draw a return-risk scatter plot (with return volatility on the horizontal axis and expected net income on the vertical axis), overlays Seaborn to generate a heat map to display the distribution density, and uses red contour lines to mark the Pareto optimal area (high return and low risk). Finally, it generates an interactive HTML5 chart embedded in the front-end decision-making interface, allowing users to click to filter high-quality strategy combinations (such as the top 10% return-risk ratio plan) and adjust parameters in real time to verify the effect.

[0046] Optionally, the generating of the income distribution graph by randomly sampling the collection period and the rebate ratio in the financial collection flow data based on the structured data set and the decision tree rule base through a Monte Carlo simulation engine further includes: Extracting the value ranges of the number of days in the account period, the rebate ratio, and the penalty rate for breach of contract from the structured data set to construct the variable constraint boundaries of the Monte Carlo simulation engine; Generate a sampling parameter set based on the regional characteristics and vehicle type classification rules of the decision tree rule base; Based on the sampling parameter set, performing a first number of random samplings on the collection period and the rebate ratio by the Monte Carlo simulation engine to calculate a net income value and an income standard deviation; The net income value and the income standard deviation are mapped into an income distribution graph.

[0047] Specifically, by statistically analyzing the account period days segment, we can determine its actual distribution range (e.g., minimum 30 days, maximum 90 days), and use it as a Monte Carlo simulation engine. The uniform sampling boundaries of the account period variable in the simulation engine (a random numerical simulation method based on probability statistics) are set. Secondly, based on the distribution characteristics of the "rebate ratio" field (e.g., mean 5%, standard deviation 2%) and business rule constraints (e.g., rebate lower limit 1%, upper limit 10%), a normal distribution sampling range for the rebate ratio is set (mean ± 3 times the standard deviation, truncated to a reasonable range of 1%-10%). At the same time, based on the actual value of the "default penalty rate" field (e.g., 0.5% to 5%) and contract terms, the upper and lower limits of its uniform distribution are defined (e.g., 0.5%-5%). Finally, the constraint boundaries of the above variables (account period ∈ [30, 90] days, rebate ratio ∈ [1%, 10%], default penalty rate ∈ [0.5%, 5%]) are written into the simulation engine configuration file (XML / JSON format rule file). This ensures that subsequent random sampling strictly adheres to the actual data range and business logic, avoids generating out-of-bounds or invalid strategy combinations, and provides a reliable variable constraint foundation for subsequent simulation and deduction.

[0048] Furthermore, based on the regional characteristics (such as geographical divisions such as East China and North China) and vehicle classification rules (such as new energy vehicles and fuel vehicles) in the decision tree rule base, the conditional constraints in the rule base are first parsed to extract parameter restrictions under different region-vehicle combinations (such as the lower limit of the rebate ratio for new energy vehicles in East China is 5%, and the upper limit of the payment period is 60 days). Then, a sampling parameter set is dynamically generated according to the rule logic. For example, the rebate ratio range for the "East China + New Energy" combination is set to 5%-10%, and the payment period range is set to 30-60 days, while the rebate ratio for the "North China + Fuel Vehicle" combination is set to 1%-8%, and the payment period is 45-90 days. At the same time, ancillary parameters such as the penalty rate are bound. Finally, the parameter constraint sets under all region-vehicle classifications (such as the upper and lower limits of the rebate ratio, payment period, and penalty rate) are integrated into the configuration file of the Monte Carlo simulation engine to ensure that the subsequent sampling process strictly matches the business rules, avoids policy conflicts or out-of-bounds risks, and provides accurate parameter input for multi-dimensional strategy deduction.

[0049] Furthermore, first, based on the preset number of sampling times (e.g., the first number of times is 10,000 times), a random parameter combination is independently generated for each region-vehicle classification rule (e.g., "East China + new energy vehicle models"), specific days are uniformly drawn from the collection period (e.g., 30-60 days), the rebate value is drawn from the normal distribution range of the rebate ratio (e.g., mean 5%, standard deviation 2%, truncated to 1%-10%), and the penalty ratio is uniformly drawn according to the penalty rate; then, for each sampling parameter combination (e.g., collection 45 days + rebate 6 % + penalty rate 2%), combined with the per-vehicle profit, sales volume, and historical bad debt rate in the business dataset, dynamically calculate the net profit value, which specifically includes: total revenue (sales volume × per-vehicle profit × (1-rebate ratio)), capital cost (per-vehicle profit × sales volume × collection days × annualized capital cost rate / 365), and bad debt loss (sales volume × per-vehicle profit × bad debt rate × penalty rate). The net profit of a single simulation is obtained by adding these three together. Finally, the results of all 10,000 simulations are summarized, the mean and standard deviation of the net profit are calculated, and all intermediate data is stored as a structured result set.

[0050] Furthermore, to map the net returns generated by the Monte Carlo simulation and their corresponding standard deviations (a risk indicator measuring return volatility) into a return distribution chart, a visualization framework was first constructed using Python's Matplotlib library: the horizontal axis represents the standard deviation, the vertical axis represents the expected net return, and each scattered point represents the result of a single sampling simulation (e.g., 10,000 points). To enhance the representation of distribution density, the kernel density estimation function of the Seaborn library was used to generate a heat map, using color depth to distinguish high-probability areas (e.g., dark red indicates dense areas). Simultaneously, an algorithm was used to screen for Pareto optimal solutions (i.e., strategy combinations with returns above the mean and risks below the mean), demarcating this area with red contour lines, and annotating the coordinates of typical high-quality strategies (e.g., the highest return point and the lowest risk point) in the chart. Finally, the chart was converted to an interactive HTML5 format, allowing users to hover their mouse to view specific parameter combinations (e.g., payment period, rebate ratio, penalty rate), click to filter the top 10% high-return, low-risk solutions, and link the front-end interface to update the strategy comparison panel in real time, ensuring that decision makers have an intuitive understanding of the return and risk distribution characteristics and key strategy boundaries.

[0051] Optionally, performing a first number of random samplings on the payment collection period and the rebate ratio through the Monte Carlo simulation engine to calculate the net income value and the income standard deviation further includes: Obtaining a single vehicle profit from a vehicle model cost database, matching it with the vehicle type in the structured dataset, and obtaining a single vehicle sales volume corresponding to the single vehicle profit; Calculating the cost of capital according to the preset weighted average cost of capital of the server, and calculating the bad debt loss based on the bad debt rate and the default penalty rate; Calculating a net income value based on the bicycle profit, the bicycle sales volume, the capital cost, and the bad debt loss; The net income average of the first number of samplings is calculated, and the income standard deviation is calculated based on the net income average and the net income value of each sampling.

[0052] Specifically, extract the per-vehicle profit data of all vehicle models from the vehicle cost database (such as MySQL). Key fields include the vehicle model unique identification code (such as the VIN code, i.e., the vehicle identification number), vehicle type classification (such as new energy sedan is type A, fuel SUV is type B), and per-vehicle profit; at the same time, obtain historical sales data matching the vehicle type from the structured data set (such as the Hive table or CSV file in the data warehouse), with fields covering vehicle type, sales, region, and time period; use ETL tools to clean and associate the two types of data, perform JOIN operations with vehicle type as the unique key (if the vehicle model identification is inconsistent, correct it through fuzzy matching or rule mapping), and generate an integrated data set; finally output the structured matching result table (field examples: vehicle type, per-vehicle profit, sales, region, time period).

[0053] Furthermore, when calculating the cost of capital based on the server's preset weighted average cost of capital (WACC), the preset annualized WACC value (e.g., 5%) is first read from the configuration file. For each Monte Carlo simulation sampling period (T), combined with the profit per vehicle (P) and sales volume (Q) of the corresponding vehicle model, the formula is used: , calculate the interest cost of funds occupied under this strategy (convert the annualized interest rate into the actual day cost); at the same time, based on the historical bad debt rate in the structured data set (such as 2%) and the default penalty rate generated by sampling (such as 3%), according to the formula: , calculate the potential loss (for example, when the profit of a single vehicle is 10,000 yuan and the sales volume is 100 units, the loss is 10,000 × 100 × 2% × 3% = 600 yuan); finally, summarize the capital cost and bad debt loss into the net income calculation logic.

[0054] Furthermore, it is necessary to first extract the profit per vehicle (the difference between the unit price and the total cost), the sales volume per vehicle (historical or predicted data) and the annualized cost of capital provided by the finance department from the enterprise ERP system and sales management platform, and perform data cleaning and verification in combination with the bad debt rate predicted by the credit model, eliminate outliers and fill in missing data. Subsequently, set the Monte Carlo simulation parameters, including the collection period (uniformly distributed between 30 and 90 days) and the rebate ratio (normal distribution with a mean of 5% and a standard deviation of 2%, limited to the range of 1% to 10%), and cover multiple scenarios through 10,000 random samplings. In each simulation, the adjusted sales (taking rebates into account), the cost of capital (annualized rate converted based on the number of days to collect) and the bad debt loss are dynamically calculated and substituted into the formula: , generating a single result. After batch simulation, the net return distribution is summarized and the Pareto front is plotted to screen for high-return, low-risk strategy combinations (e.g., 6% rebate + 60-day payment collection). Sensitivity analysis is also used to assess the elasticity of rebates and cycle times. The validation phase uses historical backtesting (calculating mean squared error) and extreme scenario testing (e.g., doubling the bad debt ratio) to ensure model accuracy and robustness. Ultimately, through visual heat maps and dynamic threshold mechanisms (e.g., triggering strategy adjustments when bad debt exceeds 5%), businesses are provided with real-time, data-driven decision support, achieving risk-controlled return maximization.

[0055] Furthermore, after completing the Monte Carlo simulation, first store all the net profit sampling results for the preset number of times into an array or list, and then calculate the mean and standard deviation of the returns in two steps: the mean is obtained by summing all the net profit values ​​and dividing them by the number of samples, using the formula: Where N is the total number of samples, μ is the mean, and X is the net income. Based on the mean, the square of the difference between each net income value and the mean is calculated item by item and accumulated, and finally the standard deviation formula is applied: , σ is the standard deviation; in actual implementation, the data can be traversed twice through the programming language (the first time is to accumulate and calculate the mean, and the second time is to calculate the square difference), or the mathematical library function can be directly called.

[0056] S204: Obtain the dealer's qualification data and payment collection behavior data, calculate a credit score using a pre-trained machine learning model, and select a target solution from the profit distribution map based on the credit score; Specifically, first extract the dealer's qualification data (such as registered capital, years of operation, compliance document status, industry certification level) and collection behavior data (on-time collection rate, overdue days, historical bad debt count) from the enterprise database or CRM system, fill missing values ​​with the median of the same type or fill with the "not provided" label, standardize continuous data (such as registered capital), and one-hot encode categorical data (such as industry certification level); then load the pre-trained machine learning model (such as XGBoost gradient boosting tree or random forest model), input the preprocessed data into the model, output the credit score value (0-100 points), and screen the eligible dealer group based on the preset threshold (such as a score ≥80 points is low risk); finally, combine the profit distribution map generated by Monte Carlo simulation to extract the high-yield-low-risk policy combination corresponding to this group (such as rebate ratio ≤6%, collection period ≤60 days, and net profit mean higher than the overall level, and standard deviation lower than the set threshold). Through automated scripts, the screening results are integrated with the business system (such as ERP or BI platform) to achieve dynamic strategy recommendation.

[0057] Optionally, obtaining the dealer's qualification data and payment collection behavior data, calculating a credit score using a pre-trained machine learning model, and selecting a target solution from the revenue distribution map based on the credit score further includes: Extracting the dealer's qualification data, including registered capital, years of cooperation, and regional economic index, and converting it into a numerical feature vector; Extracting the dealer's payment collection behavior data, including the on-time payment collection rate, the number of overdue payments, and the fluctuation rate of the payment collection cycle, to generate a time series behavior feature vector; Inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value; Classifying the credit score value into high, medium and low levels according to the credit score value and a preset credit level threshold; Based on the divided levels of the credit score values, a set of matching solutions is matched from the benefit distribution graph to obtain the target solution.

[0058] Specifically, first obtain the original data from the enterprise database or data platform, fill the missing values ​​of registered capital and years of cooperation with the median of similar data, and fill in the missing values ​​of regional economic index with the average value of the region; then perform Z-score standardization on continuous fields (registered capital, years of cooperation, regional economic index) to eliminate dimensional differences. If the regional economic index is categorized (such as A / B / C grades), it is converted into multi-column binary features through one-hot encoding; finally, the processed fields are spliced ​​into a numerical feature vector of unified dimension, and an automated process (such as Python's Scikit-learn library Pipeline) is used to ensure consistent processing logic. During verification, the vector dimension and the statistical characteristics after standardization (mean approaching 0, standard deviation approaching 1) are checked. At the same time, it supports dynamic expansion of new fields (such as adding credit ratings) to adapt to business needs and provide standardized input for subsequent machine learning models.

[0059] Furthermore, when extracting dealers' payment collection behavior data (including on-time payment collection rate, number of overdue payments, and payment cycle volatility) and generating time series behavior feature vectors, it is necessary to first extract historical payment collection records from the enterprise's financial system or database, align the data according to fixed time windows (such as monthly / quarterly), and fill in missing periods through forward filling or linear interpolation; then, based on the aligned time series, calculate the core indicators in each window: on-time payment collection rate (percentage of on-time payment collection amount), number of overdue payments (number of overdue payment collection orders), and payment cycle volatility (standard deviation of actual and agreed payment collection days), and further calculate the core indicators in each window through a sliding window (such as a 12-month ) extract statistical features (mean, variance, maximum value, trend slope) and cyclical features; finally, the statistical features or model encoding features are spliced ​​with the basic indicators, and the fixed-dimensional time series behavior feature vector (for example, a standardized 45-dimensional vector) is generated through Z-score standardization or normalization. The feature library is regularly updated through an automated pipeline (Python script or Spark job), combined with dimensional consistency checks (ensuring that the output dimensions of each dealer are the same) and business logic verification (for example, the volatility characteristics of high-risk dealers must be significantly higher) to ensure that the feature quality adapts to the input requirements of the subsequent credit scoring model.

[0060] Furthermore, the preprocessed numerical feature vectors and time series behavioral feature vectors are first loaded from the feature storage system. A feature concatenation layer then merges the two into a high-dimensional joint feature vector in a fixed order (for example, if the numerical features are 3-dimensional and the time series statistical features are 12-dimensional periodic indicators plus 8-dimensional trend features, the combined vector is a 23-dimensional vector). Subsequently, a pretrained machine learning model (such as a deep neural network) is called. This model must have undergone feature importance screening and hyperparameter tuning in advance and be persistently stored in ONNX (Open Neural Network eXchange) or Pickle (Python object serialization format). During the inference phase, the input vector must be consistent with the feature distribution during model training. Dynamic dimensionality verification (for example, checking that the input vector length is 23) is implemented to prevent feature drift, and batch mode is enabled to increase throughput (for example, processing 1024 samples at a time). After the model outputs the raw credit score, it is converted to a standardized score value (e.g., a 0-100 scale) using a sigmoid function (S-shaped growth curve function) or quantile mapping. The sigmoid function is: , where x is the model credit score and S is the standardized score value.

[0061] Furthermore, the system first sets credit rating thresholds based on business needs or historical data analysis (for example, high rating: credit score ≥ 80, medium rating: 60 ≤ credit score < 80, and low rating: credit score < 60) and stores these thresholds in a rule database. Subsequently, conditional logic is used to compare dealers' credit scores (numerical evaluation results output by a pre-trained machine learning model) against the preset thresholds. If the score meets the lower limit of the high-level threshold, it is classified as high; if it falls within the medium-level range, it is marked as medium; and if it falls below the lowest threshold, it is classified as low. During this process, the system dynamically loads threshold parameters through the rules engine and adaptively calibrates the thresholds based on the distribution characteristics of credit scores (such as standard deviation and quantiles) to ensure the rationality of the grading. For example, if the overall credit rating of dealers in a certain region is high, the system can dynamically adjust the threshold range by obtaining the Regional Economic Index (REI) through an external data interface to avoid imbalances in the grading distribution. Ultimately, the resulting credit rating will serve as the core basis for subsequent solution screening, triggering differentiated collection strategies for different grading levels through smart contracts.

[0062] Furthermore, the system first maps credit ratings to the risk labels and return parameters of each solution in the return distribution chart, for example, high ratings correspond to low-risk solutions. Subsequently, the rules engine loads multi-dimensional screening rules corresponding to the credit rating, including risk tolerance matching (e.g., high ratings only allow risk coefficients (RC) ≤ 0.3), return range constraints (e.g., high ratings require ≥ 500,000 yuan), and stability requirements (e.g., high ratings require standard deviations (SD) ≤ 15%). Furthermore, by integrating regional economic indices and vehicle type classifications obtained from external data interfaces, solutions that are not suitable for the region or vehicle type are dynamically eliminated. Next, the Analytic Hierarchy Process (AHP), a multi-criteria decision analysis method, is used to comprehensively score candidate solutions, assigning weights based on expected net return (40%), risk coefficient (30%), standard deviation (20%), and regional suitability (10%). Consistency checks (e.g., CR ≤ 0.1) are performed to ensure the rationality of the weights. Finally, a ranking list is generated in descending order of score. The system selects the solution with the highest overall score as the target solution. In case of a tie, the solution with the shortest payment period or the highest rebate ratio is prioritized.

[0063] Optionally, inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value further includes: Extracting time series features from the time series behavior feature vector, and calculating the collection cycle volatility and overdue frequency through a sliding window to generate a behavior statistics vector; Concatenating the numerical feature vector with the behavioral statistical vector to form a multi-dimensional fusion feature vector; Inputting the fused feature vector into a pre-trained machine learning model, and outputting a predicted value of the probability of default through a tree structure hierarchical decision; The default probability prediction value is converted into the credit score value based on a nonlinear mapping function.

[0064] Specifically, the system first sets the sliding window length (e.g., 30 days) and step size (e.g., 1 day), and then sequentially traverses the dealer's collection behavior time series data. Within each window, the actual cycle number of all collection records is extracted, and the standard deviation (SD, a statistical indicator that measures the degree of data dispersion) is calculated as the collection cycle volatility. The specific formula is: Among them, X i is the single payment collection period, μ is the mean of the payment collection period within the window, and N is the total number of payments. At the same time, the number of overdue payments within the window is counted, and the overdue frequency is obtained by calculating the ratio of the number of overdue payments to the total number of payments. For example, if there are 2 overdue payments in 10 payments within the window, then OF is 20%. In addition, the derived supplementary features include the maximum number of overdue days and the average payment collection period within the window. After completing all window traversals, the missing data is filled in by interpolation, and the volatility, overdue frequency and other indicators are Z-Score standardized to ensure uniform feature scales. The volatility, overdue frequency, maximum overdue days and average payment period of each window are spliced ​​in chronological order to form a multi-dimensional behavioral statistical vector (for example, SD = 5.2 days, OF = 8%, maximum overdue days = 15 days, average period = 25 days).

[0065] Furthermore, the numerical feature vector (including static data such as registered capital, years of cooperation, and the regional economic index REI) and the behavioral statistical vector (including dynamic time-series features such as SD and OF) must be fused through feature alignment. First, ensure that the standards of the two feature types are consistent, and then directly concatenate them in dimensional order. For example, if the numerical feature vector is registered capital = 5 million yuan, REI = 1.2, and the behavioral statistical vector is SD = 5.2 days, OF = 8%, maximum overdue days = 15 days, and average period = 25 days, then the multi-dimensional fused feature vector is 500, 1.2, 5.2, 8, 15, 25, with the first two dimensions being static numerical features and the last four being dynamic behavioral features.

[0066] Next, a trained tree model (such as a gradient boosted tree (GBDT) or random forest) is loaded. This model consists of multiple decision trees, each of which constructs hierarchical decision logic using feature splitting rules. After the fused feature vector is input into the model, each tree processes the data in turn. Within a single tree, the model starts at the root node and, based on the feature splitting criteria (such as REI ≤ 1.0 or SD ≥ 5 days), determines whether the input data should enter the left or right subtree layer by layer until it reaches a leaf node. For example, if the splitting rule for a node is REI ≤ 1.0 and the current input REI value is 1.2, the data will enter the right subtree and continue to evaluate the next layer of rules. Each leaf node of the tree stores the percentage of default samples collected during the training phase as an initial probability (for example, if 20 out of 100 samples in a leaf node during training resulted in defaults, the initial default probability is 0.2). For ensemble models (such as random forests), the probability prediction results of all trees are integrated by taking the average; for gradient boosting trees, the output of each tree is accumulated through an additive model and then converted to a default probability value between 0 and 1 through the Sigmoid function.

[0067] Furthermore, the system designed scoring rules based on piecewise nonlinear functions. In the low-risk range (default probability ≤ 5%), a gentle logarithmic function is used to slowly reduce the score from 100 to 80 (e.g., probability 1% → 100, 5% → 80). In the medium-risk range (5% < probability ≤ 20%), an exponential function is used to accelerate the deduction of points (e.g., probability 10% → 65, 20% → 30). In the high-risk range (probability > 20%), a sigmoid function is used to rapidly compress the score to 0-30 points (e.g., probability 30% → 9), highlighting the trend of worsening risk. The system automatically constrains extreme values ​​(e.g., over-limit scores are forcibly truncated to 100 or 0) and uses polynomial interpolation to smooth sudden changes at interval boundaries (e.g., 5% and 20%) to ensure continuous and stable scoring. Finally, the KS value is used to verify the score's ability to discriminate against defaulting customers (e.g., customers in the high-score segment have significantly lower default rates than those in the low-score segment). The PSI index is used to monitor the monthly stability of the score distribution and dynamically adjust the mapping parameters. For example, a default probability of 2% is converted to 95 points after nonlinear transformation, triggering a long-term account strategy; a probability of 35% is 5 points, forcing prepayment to achieve a precise match between risk and return.

[0068] S205. Based on the target solution, an electronic contract matching the dealer's credit rating is automatically generated through a smart contract and a payment collection process is executed.

[0069] Specifically, the credit rating matching clauses (such as a high credit rating corresponding to "60-day account period, 5% rebate, and 1% daily penalty") are encoded as executable logic, and the on-chain oracle (Oracle) is called to obtain real-time regional economic index (REI) and vehicle sales data to automatically calibrate contract details (such as an additional 2% increase in regional rebates for new energy vehicles); after the contract is deployed to the blockchain network, it is connected to the payment system through an API to automatically trigger payment verification when the account expires - if the dealer repays on time, the contract will immediately issue a rebate to the designated account and update the credit score; if overdue, a fine is calculated on a daily basis (fine = uncollected amount × daily penalty rate × number of overdue days), automatically deducted from the pre-deposited deposit, triggering a risk control warning, and freezing the dealer's subsequent vehicle pick-up rights; all contract execution records (such as repayment time, fine details) are stored on the chain in real time, supporting the generation of visual audit reports by region and vehicle model. For example, in the target scenario, a high-credit dealer matches "60-day payment period + 5% rebate". The contract will automatically verify the bank statement at 23:59 on the 60th day. If the full amount is recovered, the rebate will be released at 00:01 the next day and the dealer's credit rating will be upgraded.

[0070] This embodiment also discloses a system for collecting automobile payments. Figure 3 This is a module diagram of a car payment collection system disclosed in an embodiment of the present application. Figure 3 As shown, the system includes: The policy data parsing module 301 is used to obtain policy text and financial collection flow data, parse the policy text using a natural language processing algorithm based on information entropy, and associate the parsed results with the financial collection flow data to generate a structured data set. The structured data set includes the number of days in the payment period, vehicle type, rebate ratio, and penalty rate for breach of contract. A multi-dimensional rule configuration module 302 is used to construct a decision tree rule base based on preset multi-dimensional rules, wherein the multi-dimensional rules include sales model, capital attributes, and regional characteristics; The revenue prediction module 303 is configured to randomly sample the collection period and the rebate ratio in the financial collection flow data based on the structured data set and the decision tree rule base through a Monte Carlo simulation engine to generate a revenue distribution graph; A credit scoring module 304 is configured to obtain dealer qualification data and payment collection behavior data, calculate a credit score using a pre-trained machine learning model, and select a target solution from the revenue distribution map based on the credit score; The electronic contract generation and execution module 305 is used to automatically generate an electronic contract that matches the credit rating of the dealer through a smart contract based on the target solution and execute the collection process.

[0071] Optionally, the policy data parsing module 301 is specifically configured to: Performing a preset keyword matching operation on the policy text to obtain a structured table, wherein the preset keywords include policy ID, number of days in the account period, rebate ratio, penalty rate for breach of contract, and applicable vehicle model; The policy ID is matched with the collection records in the financial collection flow data, and the actual collection rate and bad debt rate of each policy are calculated to obtain the structured data set.

[0072] Optionally, the multi-dimensional rule configuration module 302 is specifically configured to: Using the regional features as the root node of the decision tree, generating an initial decision tree structure; Bind the sales model and vehicle type classification constraints under the root node to form a secondary leaf node; Configuring policy parameters matching the multi-dimensional rules under the second-level leaf node to generate a third-level leaf node, wherein the policy parameters include the number of days in the account period, the rebate ratio, the prepayment ratio, and the penalty rate for breach of contract; The inventory turnover rate and the regional economic index are obtained in real time through an external data interface, the weights and constraints of all nodes of the decision tree are updated, and the decision tree rule base is generated. All nodes include the root node, the second-level leaf node and the third-level leaf node.

[0073] Optionally, the revenue prediction module 303 is specifically configured to: Extracting the value ranges of the number of days in the account period, the rebate ratio, and the penalty rate for breach of contract from the structured data set to construct the variable constraint boundaries of the Monte Carlo simulation engine; Generate a sampling parameter set based on the regional characteristics and vehicle type classification rules of the decision tree rule base; Based on the sampling parameter set, performing a first number of random samplings on the collection period and the rebate ratio by the Monte Carlo simulation engine to calculate a net income value and an income standard deviation; The net income value and the income standard deviation are mapped into an income distribution graph.

[0074] Optionally, the revenue prediction module 303 is specifically configured to: Obtaining a single vehicle profit from a vehicle model cost database, matching it with the vehicle type in the structured dataset, and obtaining a single vehicle sales volume corresponding to the single vehicle profit; Calculating the cost of capital according to the preset weighted average cost of capital of the server, and calculating the bad debt loss based on the bad debt rate and the default penalty rate; Calculating a net income value based on the bicycle profit, the bicycle sales volume, the capital cost, and the bad debt loss; The net income average of the first number of samplings is calculated, and the income standard deviation is calculated based on the net income average and the net income value of each sampling.

[0075] Optionally, the credit scoring module 304 is specifically configured to: Extracting the dealer's qualification data, including registered capital, years of cooperation, and regional economic index, and converting it into a numerical feature vector; Extracting the dealer's payment collection behavior data, including the on-time payment collection rate, the number of overdue payments, and the fluctuation rate of the payment collection cycle, to generate a time series behavior feature vector; Inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value; Classifying the credit score value into high, medium and low levels according to the credit score value and a preset credit level threshold; Based on the divided levels of the credit score values, a set of matching solutions is matched from the benefit distribution graph to obtain the target solution.

[0076] Optionally, the credit scoring module 304 is specifically configured to: Extracting time series features from the time series behavior feature vector, and calculating the collection cycle volatility and overdue frequency through a sliding window to generate a behavior statistics vector; Concatenating the numerical feature vector with the behavioral statistical vector to form a multi-dimensional fusion feature vector; Inputting the fused feature vector into a pre-trained machine learning model, and outputting a predicted value of the probability of default through a tree structure hierarchical decision; The default probability prediction value is converted into the credit score value based on a nonlinear mapping function.

[0077] This embodiment also discloses an electronic device, referring to Figure 4 The electronic device may include: at least one processor 401 , at least one communication bus 402 , a user interface 403 , a network interface 404 , and at least one memory 405 .

[0078] The communication bus 402 is used to implement the connection and communication between these components.

[0079] The user interface 403 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0080] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0081] Processor 401 may include one or more processing cores. Using various interfaces and circuits, processor 401 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in memory 405, as well as accesses data stored in memory 405, to perform various server functions and process data. Optionally, processor 401 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). Processor 401 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 401 and implemented as a separate chip.

[0082] Memory 405 may include random access memory (RAM) or read-only memory (ROM). Optionally, memory 405 may include non-transitory computer-readable storage medium. Memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. Memory 405 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch control, sound playback, image playback, etc.), and instructions for implementing the aforementioned method embodiments. The data storage area may store data related to the aforementioned method embodiments. Memory 405 may also optionally be at least one storage device located remotely from the processor 401. As shown in the figure, memory 405, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application for auto payment repayment.

[0083] exist Figure 4In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 401 can be used to call an application for automobile repayment stored in the memory 405. When executed by one or more processors 401, the electronic device executes one or more methods in the above embodiments.

[0084] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the technical field that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for collecting automobile payments, characterized in that: Applied to a server, the method includes: Obtaining policy text and financial collection flow data, parsing the policy text using a natural language processing algorithm based on information entropy, and associating the parsed results with the financial collection flow data to generate a structured dataset. The structured dataset includes the number of days in the payment period, vehicle type, rebate ratio, and penalty rate for breach of contract. Construct a decision tree rule base based on preset multi-dimensional rules, including sales model, capital attributes and regional characteristics; Based on the structured data set and the decision tree rule base, randomly sampling the collection period and the rebate ratio in the financial collection flow data through a Monte Carlo simulation engine to generate a profit distribution graph; Obtaining dealer qualification data and payment collection behavior data, calculating a credit score using a pre-trained machine learning model, and screening a target solution from the revenue distribution map based on the credit score; Based on the target solution, an electronic contract matching the dealer's credit rating is automatically generated through a smart contract and a payment collection process is executed.

2. The method according to claim 1, characterized in that Parsing the policy text and associating the parsing results with the financial payment flow data to generate a structured data set includes: Performing a preset keyword matching operation on the policy text to obtain a structured table, wherein the preset keywords include policy ID, number of days in the account period, rebate ratio, penalty rate for breach of contract, and applicable vehicle model; The policy ID is matched with the collection records in the financial collection flow data, and the actual collection rate and bad debt rate of each policy are calculated to obtain the structured data set.

3. The method according to claim 1, characterized in that The construction of a decision tree rule base based on preset multi-dimensional rules further includes: Using the regional features as the root node of the decision tree, generating an initial decision tree structure; Bind the sales model and vehicle type classification constraints under the root node to form a secondary leaf node; Configuring policy parameters matching the multi-dimensional rules under the second-level leaf node to generate a third-level leaf node, wherein the policy parameters include the number of days in the account period, the rebate ratio, the prepayment ratio, and the penalty rate for breach of contract; The inventory turnover rate and the regional economic index are obtained in real time through an external data interface, the weights and constraints of all nodes of the decision tree are updated, and the decision tree rule base is generated. All nodes include the root node, the second-level leaf node and the third-level leaf node.

4. The method according to claim 1, wherein The generating of the income distribution graph by randomly sampling the collection period and the rebate ratio in the financial collection flow data based on the structured data set and the decision tree rule base through a Monte Carlo simulation engine further includes: Extracting the value ranges of the number of days in the account period, the rebate ratio, and the penalty rate for breach of contract from the structured data set to construct the variable constraint boundaries of the Monte Carlo simulation engine; Generate a sampling parameter set based on the regional characteristics and vehicle type classification rules of the decision tree rule base; Based on the sampling parameter set, performing a first number of random samplings on the collection period and the rebate ratio by the Monte Carlo simulation engine to calculate a net income value and an income standard deviation; The net income value and the income standard deviation are mapped into an income distribution graph.

5. The method according to claim 4, characterized in that The performing a first number of random samplings on the payment collection period and the rebate ratio by the Monte Carlo simulation engine to calculate the net income value and the income standard deviation further includes: Obtaining a single vehicle profit from a vehicle model cost database, matching it with the vehicle type in the structured dataset, and obtaining a single vehicle sales volume corresponding to the single vehicle profit; Calculating the cost of capital according to the preset weighted average cost of capital of the server, and calculating the bad debt loss based on the bad debt rate and the default penalty rate; Calculating a net income value based on the bicycle profit, the bicycle sales volume, the capital cost, and the bad debt loss; The net income average of the first number of samplings is calculated, and the income standard deviation is calculated based on the net income average and the net income value of each sampling.

6. The method according to claim 1, characterized in that The obtaining of dealer qualification data and payment collection behavior data, calculating a credit score using a pre-trained machine learning model, and screening a target solution from the revenue distribution diagram based on the credit score also includes: Extracting the dealer's qualification data, including registered capital, years of cooperation, and regional economic index, and converting it into a numerical feature vector; Extracting the dealer's payment collection behavior data, including the on-time payment collection rate, the number of overdue payments, and the fluctuation rate of the payment collection cycle, to generate a time series behavior feature vector; Inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value; Classifying the credit score value into high, medium and low levels according to the credit score value and a preset credit level threshold; Based on the divided credit score values, a set of matching solutions is matched from the benefit distribution graph to obtain the target solution.

7. The method according to claim 6, characterized in that Inputting the numerical feature vector and the time series behavior feature vector into a pre-trained machine learning model to calculate the credit score value further includes: Extracting time series features from the time series behavior feature vector, and calculating the collection cycle volatility and overdue frequency through a sliding window to generate a behavior statistics vector; Concatenating the numerical feature vector with the behavioral statistical vector to form a multi-dimensional fusion feature vector; Inputting the fused feature vector into a pre-trained machine learning model, and outputting a predicted value of the probability of default through a tree structure hierarchical decision; The default probability prediction value is converted into the credit score value based on a nonlinear mapping function.

8. A system for collecting automobile payments, characterized in that: Specifically include: A policy data parsing module is used to obtain policy text and financial collection flow data, parse the policy text using a natural language processing algorithm based on information entropy, and associate the parsing results with the financial collection flow data to generate a structured data set. The structured data set includes the number of days in the account period, vehicle type, rebate ratio, and penalty rate for breach of contract. A multi-dimensional rule configuration module is used to build a decision tree rule library based on preset multi-dimensional rules, wherein the multi-dimensional rules include sales model, capital attributes and regional characteristics; A revenue prediction module, configured to randomly sample the collection period and the rebate ratio in the financial collection flow data based on the structured data set and the decision tree rule base through a Monte Carlo simulation engine to generate a revenue distribution graph; A credit scoring module is used to obtain dealer qualification data and payment collection behavior data, calculate a credit score value through a pre-trained machine learning model, and screen target solutions from the profit distribution map based on the credit score value; The electronic contract generation and execution module is used to automatically generate an electronic contract that matches the dealer's credit rating through a smart contract based on the target solution and execute the collection process.

9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, a user interface and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Rating system and method for matching credit rating with default loss rate based on SMAA-DS

    CN113610638A

  • Tax information processing method and related equipment

    CN114971833A

  • Financial risk assessment management method and system

    CN117593142A

  • Contract multi-return prediction method, device and equipment and medium

    CN119313487A

  • Financial industry chain risk identification and analysis method

    CN119379424A

Cited By

  • Rail transit station power distribution and transformation electric energy quality comprehensive evaluation method and system

    CN121189654A

  • Rail transit station power distribution quality comprehensive evaluation method and system

    CN121189654B