Intelligent litigation method and system for small financial creditor's right dispute based on block chain
By constructing a multi-dimensional borrower user portrait and using the LightGBM model to generate personalized litigation plans, and combining blockchain technology to manage loan data, the problem of the lack of personalized litigation model for existing microfinance debt dispute litigation models is solved, and the litigation effect and resource allocation efficiency are improved.
Patent Information
- Application Number
- CN202510132541.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing microfinance debt dispute litigation model lacks personalized plans and cannot effectively deal with complex and changeable dispute cases, resulting in poor litigation results.
By constructing a multi-dimensional borrower user portrait and using the LightGBM binary classification model to generate personalized litigation plans, combining blockchain technology to manage and track loan data, automatically triggering the overdue logic.
It realizes the accurate portrayal of the borrower's in-depth information, improves the objectivity and quantitative nature of the litigation strategy, dynamically adjusts the litigation plan, and improves the efficiency of litigation resource allocation.
Smart Images

Figure CN120070033A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of blockchain, and particularly to a smart litigation method and system for small - amount financial creditor's rights disputes based on blockchain. Background Art
[0002] In the field of small - amount financial lending, there is information asymmetry between institutions and individuals, and the default risk of borrowers is relatively high. Once overdue occurs, financial institutions need to invest a large amount of manpower and material resources in collecting debts from borrowers, prosecuting, and enforcing, which has a long cycle, great difficulty, and low recovery rate. At the same time, due to the small amount of the subject matter of small - amount lending and the scattered locations of borrowers, professional institutions such as law firms are not very enthusiastic about taking on such cases, resulting in insufficient supply of legal services for post - loan debt collection, further exacerbating the difficulty of handling financial disputes.
[0003] The existing small - claim litigation models usually adopt a "pipeline" operation, applying a standardized litigation process to each case, ignoring the individual differences of borrowers and the particularity of cases. This "one - size - fits - all" approach is difficult to effectively handle complex and changeable dispute cases, resulting in less than satisfactory litigation results. With the rise of fintech, new technologies such as smart contracts, big data analysis, and machine learning have brought hope for the intelligent resolution of small - amount financial disputes, but the relevant applied research is still in its infancy and has not yet formed a mature and complete technical solution.
[0004] Specifically for personalized litigation, there is currently a lack of technical means to accurately depict borrowers from a large amount of heterogeneous data. Most credit reporting systems only focus on borrowers' credit scores and default records, ignoring deep - level information such as borrowers' behavioral characteristics and risk preferences, resulting in a single - dimensional user portrait and inaccurate judgment of default reasons and litigation risks. At the same time, most existing litigation strategies are based on experience summary and subjective judgment, lacking an objective and quantitative risk assessment and strategy optimization mechanism, and it is difficult to dynamically adjust litigation plans according to the characteristics of cases, leading to improper allocation of litigation resources. Summary of the Invention
[0005] Aiming at the lack of personalized litigation solutions in small - amount financial creditor's rights disputes in the prior art, this application provides a smart litigation method and system for small - amount financial creditor's rights disputes based on blockchain, by constructing a multi - dimensional borrower user portrait and using the LightGBM binary classification model to generate personalized litigation plans.
[0006] The objectives of this application are achieved through the following technical solutions.
[0007] An aspect of the present application provides a smart litigation method for small - scale financial creditor's rights disputes based on blockchain, including: collecting loan business data; setting up a smart contract according to the loan business data, deploying the smart contract to the consortium blockchain, with all parties participating in the loan as nodes; obtaining the repayment data of the borrower, synchronizing the repayment data to the blockchain, and updating the loan status field in the smart contract according to the comparison result between the repayment data and the repayment plan; wherein the loan status field includes normal and overdue; obtaining the loan status field in the smart contract, determining whether it includes an overdue field, and if so, calculating the number of overdue days and the overdue penalty receivable; for borrowers with the loan status field being overdue, constructing a borrower user portrait based on the borrower information in the loan application materials and the corresponding credit report; using the borrower user portrait as the input feature X, the litigation result of the historical closed - case loan dispute training samples as the target variable Y, and the litigation time consumption of the samples as the training weight W, training a LightGBM binary classification model, and generating a litigation difficulty score using the trained LightGBM binary classification model; in step S7, according to the litigation difficulty score, matching a preset litigation strategy template to generate a personalized litigation plan.
[0008] Among them, a smart contract is a self - executing contract based on blockchain technology, and the contract terms are written into the blockchain in the form of a computer program. In the present application, the smart contract is used to manage and track various data and repayment status of the loan business. Once the borrower defaults, the corresponding processing logic will be automatically triggered. Through the smart contract, the rights and responsibilities of all parties participating in the loan are clearer and more transparent, and the contract execution is more efficient and rigorous. A consortium blockchain is a form of blockchain between a public blockchain and a private blockchain. In this solution, all parties participating in the loan business (such as loan institutions, guarantors, regulatory departments, etc.) form a consortium and jointly act as nodes of the blockchain network to participate in the recording and verification of blocks. Compared with a public blockchain, the consortium blockchain has better scalability and privacy protection; compared with a private blockchain, its consensus mechanism is more decentralized, with higher security and credibility.
[0009] The loan status field is a key status variable in the smart contract, used to record the real-time repayment status of each loan. The two states designed in this scheme are "normal" and "overdue". The borrower's repayment data for each period will be synchronized to the blockchain. By comparing the repayment records with the preset repayment plan, the status field will be updated accordingly. This provides a basis for subsequent calculation of overdue fines, litigation risk assessment, etc. LightGBM is a gradient boosting machine learning algorithm based on decision trees, with unique advantages in performance and training speed. It is used in this scheme to train a binary classification model for evaluating litigation risks. Through the borrower's user portrait features (input X), the litigation results of historical closed disputes (target Y), and the litigation time (sample weight W), the model can learn the associations between different features and litigation results, litigation costs, so as to make a score prediction for the litigation difficulty of new overdue cases.
[0010] Furthermore, construct the borrower's user portrait, including: extracting the borrower's loan application materials from the deposit contract on the blockchain, and parsing the loan application materials into a structured borrowing feature vector X a ; where the loan application materials include the borrower's basic information, borrowing purpose, income certificate, and asset certificate; obtaining the borrower's credit report from a third-party credit reporting platform, and parsing the credit report into a structured credit feature vector X b ; where the credit report includes the borrower's credit score, historical overdue records, and debt-to-income ratio; according to the borrower's borrowing feature vector X a and credit feature vector X b , construct the user portrait as the input feature X.
[0011] Furthermore, construct the user portrait as the input feature X, including: vertically concatenating the borrowing feature vector X a and the credit feature vector X b to obtain the initial user portrait feature set X 0 ; perform min-max normalization processing on the continuous variables in X 0 ; perform One-hot encoding processing on the discrete variables in X 0 to obtain the user portrait feature set X 1 ; use the recursive feature elimination algorithm RFE to perform feature selection on X 1 and select the top N features most relevant to the target variable Y to obtain the dimensionality-reduced user portrait feature set X 2 ; for X 2Perform non-linear mapping to project the N-dimensional features onto a two-dimensional plane to obtain the user portrait distribution map; use the K-Means algorithm to cluster the user portrait distribution map to obtain K user portrait clusters; conduct statistical analysis on the K typical user portrait clusters to obtain the feature distribution of borrowers within each cluster, and generate K portrait labels based on the features of the clustering centers of each cluster; perform One-hot encoding on the K portrait labels to generate a K-dimensional binary feature vector X 3 ; Combine the N-dimensional user portrait feature set X 2 and the K-dimensional binary feature vector X 3 horizontally to obtain the extended user portrait feature set X 4 , which is used as the input feature X.
[0012] Furthermore, generate a litigation difficulty score, including: preprocess the input feature X, the target variable Y, and the training weights {w i} to obtain the training set D train and the validation set D val ; Set the hyperparameters of the LightGBM binary classification model, and use the K-fold cross-validation method to divide the training set D train into K mutually exclusive subsets; among them, the hyperparameters include the learning rate, the number of leaf nodes, the maximum tree depth, and the regularization coefficient; on each fold of the training subset, set the weight of each training sample according to the training weights {w i}, train a LightGBM binary classification sub-model, repeat K times to obtain K LightGBM binary classification sub-models; use a logistic regression model to perform weighted averaging on the K LightGBM binary classification sub-models to generate a LightGBM binary classification model; use the LightGBM binary classification model to predict the validation set D val to obtain the winning probability value of each sample; select multiple different decision thresholds, calculate the precision and recall rate under each decision threshold, and draw the P-R curve; calculate the Youden index according to the P-R curve, and select the winning probability value corresponding to the maximum Youden index as the optimal decision threshold t best ; According to the borrowers with the loan status field being overdue, extract the corresponding user portraits as the input feature X, and use the trained LightGBM binary classification model to predict to obtain the winning probability value p i of each overdue borrower; Compare the predicted winning probability value p i with the optimal decision threshold t best to generate the winning risk level; Generate the litigation difficulty score according to the winning risk level, the winning probability value, and the overdue days.
[0013] Among them, the decision threshold is a critical point for converting the probability prediction value into a binary classification prediction label (win or lose in this solution). For each possible threshold, the probability higher than the threshold is predicted as a win, and the probability lower than the threshold is predicted as a loss, and a set of prediction labels can be obtained. By comparing with the true labels, the precision and recall of the model under this threshold can be calculated. By trying multiple thresholds and plotting the P-R curve, the optimal decision point of the model can be found. The Youden index is a comprehensive index for evaluating the performance of a binary classification model, defined as the sum of sensitivity and specificity minus 1, that is: Youden's Index = Sensitivity + Specificity - 1. Among them, sensitivity is equivalent to recall, and specificity is equal to 1 - false positive rate. The larger the Youden index, the better the balance of the model in terms of sensitivity and specificity. Calculate the Youden index through the precision and recall of each point on the P-R curve, and find the decision threshold corresponding to the maximum value as the optimal threshold of the model.
[0014] The probability of winning is a value between 0 and 1 that the model predicts the possibility of winning based on the borrower's user profile characteristics. The closer it is to 1, the more similar the characteristics the model believes the borrower has to the historical winning cases, and the greater the possibility of winning; otherwise, the possibility of winning is lower. This probability value essentially reflects the correlation between factors such as the borrower's credit status and default situation and the litigation result.
[0015] The winning risk level is different levels divided for the winning risk of each overdue borrower according to the predicted probability of winning value of the model, such as high risk, medium risk, low risk, etc. By comparing the probability of winning value with the optimal decision threshold, the risk interval where the probability is located can be judged. For example, two thresholds p1 and p2 can be set. When the probability of winning is less than p1, it is high risk; when it is greater than p2, it is low risk; and when it is between the two, it is medium risk. The division of the risk level helps the subsequent matching and decision-making of litigation strategies.
[0016] Furthermore, preprocess the input feature X, the target variable Y, and the training weights {w i}, including: normalizing the input feature X; binarizing the target variable Y, marking the winning result as 1 and the losing result as 0; performing a logarithmic transformation on the training weights {w i}.
[0017] Furthermore, performing a logarithmic transformation on the training weights {w i} includes: for each training sample i, extracting the corresponding litigation time t i; For the litigation time t i perform normalization to obtain t i '; For the normalized litigation time t i ' perform logarithmic transformation to obtain the training weights {w i}, and the calculation formula of w i is: where α is the parameter of the logarithmic transformation.
[0018] Furthermore, obtain the training set D train and the validation set D val , including: According to the litigation results of the training samples of historical closed loan disputes, for each winning sample s i , randomly select one s i from the M 1 nearest neighbor winning samples of s j ; where, Y[s i =1; Randomly select an interpolation point x i between the user portrait features x j corresponding to the winning sample s i and x j = x new = x i + rand(0,1)×(x j - x i ), and add x new as a new winning sample to the training set; Repeat until the number of winning samples is greater than the number of losing samples; For each losing sample f i in the litigation results of the training samples of historical closed loan disputes, calculate the Euclidean distance from the user portrait x i corresponding to f i to all samples in the sample set; Select the M 2 samples with the closest distance, and determine the majority class corresponding to the M 2 nearest neighbor samples according to the majority voting rule; where, Y[f i =0; If the class label 0 of f i is inconsistent with the majority class among the M 2 nearest neighbors, then remove the user portrait x i corresponding to the f i set from the training set, and repeat until all losing samples are traversed to obtain the training set D train,balanced ; Randomly divide the training set D train,balanced into the training set D train and the validation set D val .
[0019] Furthermore, draw the P-R curve, including: According to the LightGBM binary classification model for the validation set D valThe predicted default probability value, select multiple different decision thresholds, and convert the predicted default probability value into a binary prediction label according to each decision threshold; for each decision threshold, calculate the precision Precision and recall Recall according to the corresponding binary prediction label and the corresponding true label: Precision = TP / (TP + FP), Recall = TP / (TP + FN), F1 Score = 2 * Precision * Recall / (Precision + Recall), where TP represents the number of samples correctly predicted as winning the lawsuit, FP represents the number of samples wrongly predicted as winning the lawsuit, and FN represents the number of samples wrongly predicted as losing the lawsuit; plot the calculation results of precision Precision and recall Recall under each decision threshold as a P-R curve.
[0020] Further, calculate the Youden index according to the P-R curve, and select the winning probability value corresponding to the maximum Youden index as the optimal decision threshold t best , including: for each point (Precision i , Recall i ) on the P-R curve, calculate its corresponding comprehensive evaluation index F1 Score i : where i is the number of the point on the P-R curve; for each point (Precision i , Recall i ) on the P-R curve, calculate its corresponding Youden index J i : J i = Precision i + Recall i - 1, respectively obtain the maximum values F1 max and J max of the F1 Score and the Youden index J, and extract the corresponding decision thresholds t F1 and t J ; calculate the optimal decision threshold t best according to the following formula: t best = λ × t F1 + (1 - λ) × t J , where λ is the weight coefficient.
[0021] Another aspect of the present application also provides a blockchain-based intelligent litigation system for small financial claim disputes, which is used to execute a blockchain-based intelligent litigation method for small financial claim disputes of the present application.
[0022] Compared with the prior art, the advantages of the present application are as follows:
[0023] When constructing the borrower user portrait, on the one hand, structured borrowing features are extracted from the loan application materials stored on the blockchain, and on the other hand, the credit report of the borrower is obtained from a third-party credit investigation platform to form a comprehensive and three-dimensional user portrait. The fusion of dual-source heterogeneous data makes up for the deficiencies of a single data dimension and improves the information integrity and description accuracy of the user portrait. At the same time, algorithms such as recursive feature elimination, T-SNE dimensionality reduction, and K-Means clustering are used to optimize the user portrait, so that the extracted user features can also have strong discriminative ability under small sample conditions.
[0024] When training the litigation result prediction model, the LightGBM algorithm is used to construct a binary classification model. On the one hand, the ensemble learning framework of decision trees is used to capture the non-linear relationship between the user portrait features and the litigation results, and on the other hand, the Gradient Boosting mechanism is introduced for iterative optimization and parameter adjustment, which performs well in complex and changeable small-claim litigation scenarios. At the same time, the sample litigation time consumption is innovatively introduced as the training weight, and the weight is logarithmically transformed to highlight the importance of high-risk cases; strategies such as K-fold cross-validation and logistic regression weighting are used to control overfitting and improve the generalization ability of the model; data augmentation and balancing are performed on the training samples to alleviate the problem of class imbalance.
[0025] When generating the litigation difficulty score, the trained LightGBM model is used to predict the litigation results of new overdue cases to obtain the winning probability distribution. The P-R curve and Youden index are innovatively introduced, and by weighted averaging the optimal threshold of the F1 score and the optimal threshold of the Youden index, the litigation win-loss judgment boundary is adaptively selected, improving the recall rate of the judgment while controlling the risk of misjudgment, making the difficulty score more in line with the actual business needs. At the same time, multiple factors such as the winning probability and overdue days are comprehensively considered to construct a multi-dimensional scoring matrix to refine the litigation difficulty level and make the subsequent strategy matching more accurate. Brief Description of the Drawings
[0026] This application will be further described in the form of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0027] Figure 1 is an exemplary flowchart of a smart litigation method for small financial creditor's rights disputes based on blockchain according to some embodiments of the present application;
[0028] Figure 2 is an exemplary flowchart of generating an extended user portrait feature set X4 according to some embodiments of the present application;
[0029] Figure 3It is an exemplary flowchart for generating the winning risk level shown in some embodiments of the present application. Detailed implementation manners
[0030] The methods and systems provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0031] As Figure 1 shown, collect loan business data; according to the loan business data, set up a smart contract, deploy the smart contract to the consortium blockchain, and use the loan participating parties as nodes; obtain the borrower's repayment data, synchronize the repayment data to the blockchain, and update the loan status field in the smart contract according to the comparison result between the repayment data and the repayment plan; wherein, the loan status field includes normal and overdue; obtain the loan status field in the smart contract, determine whether it includes an overdue field, if it does, calculate the overdue days and the overdue fine receivable; for the borrower with the loan status field being overdue, construct a borrower user profile according to the borrower information and the corresponding credit report in the loan application materials; use the borrower user profile as the input feature X, use the litigation result of the training samples of the historical closed loan disputes as the target variable Y, and use the litigation time-consuming of the samples as the training weight W, train the LightGBM binary classification model, and use the trained LightGBM binary classification model to generate a litigation difficulty score; according to the litigation difficulty score, match the preset litigation strategy template to generate a personalized litigation plan.
[0032] Specifically, collect loan business data; extract structured data such as the borrower's basic information table, loan application form, loan contract table, and repayment plan table from the core business system database. Obtain the image scans of the borrower's identity certificate, income certificate, asset certificate, etc. submitted by the borrower from the document storage system. Obtain the borrower's credit report, including credit score, debt situation, overdue record, etc., from the external credit reporting system API interface. Clean and standardize the collected structured data, extract key information from the unstructured data, and integrate them into a unified loan business data set.
[0033] According to the loan business data, set up a smart contract, deploy the smart contract to the consortium blockchain, and use the loan participating parties as nodes; set the state variables of the smart contract according to the elements in the loan contract table, such as loan amount, interest rate, term, repayment method, etc. Set the regular execution function of the smart contract according to the repayment plan, such as automatically deducting from the borrower's account on the monthly repayment date. Set the event trigger function of the smart contract according to the collection rules, such as sending collection text messages and calls when overdue. Compile and package the set smart contract code into a unified deployment file. Through the contract management tool of the consortium chain, distribute the smart contract deployment file to each participating node. Each participating node runs the deployment file, installs the smart contract on the local blockchain node, and starts the corresponding blockchain node service.
[0034] Obtain the repayment data of the borrower and update the smart contract: From the core business system, extract the actual repayment amount of the borrower on each repayment date with the borrower ID as the primary key. Compare the actual repayment amount with the repayment plan to calculate the completion status (normal / overdue / partial repayment) of each period's repayment. For each borrower, generate a repayment status change transaction and synchronize the repayment status to the blockchain. After receiving the repayment status change transaction, the blockchain node automatically calls the status update function of the corresponding smart contract. The smart contract update function modifies the loan status field inside the contract according to the repayment status and records the update timestamp. The loan status field is set to "normal" or "overdue" according to the repayment completion situation and triggers the corresponding contract event.
[0035] Obtain the loan status field of the smart contract and calculate the overdue information: The blockchain node calls the query function of the smart contract through the RPC interface to obtain the loan status field of each borrower. For borrowers whose status field contains "overdue", extract the timestamp To when the "overdue" status first appears from the transaction records of the blockchain. According to the current timestamp Tc and the earliest overdue timestamp To, calculate the overdue days Td = (Tc - To) / 86400. Query the daily interest rate Rd of the borrower from the smart contract and calculate the overdue penalty Fp = principal * Rd * Td. Save the calculated overdue days and overdue penalty as additional fields of the borrower's overdue status to the blockchain or business database.
[0036] As Figure 2 shown, for borrowers with an overdue loan status field, construct a borrower user profile based on the borrower information in the loan application materials and the corresponding credit report, including: Extract the loan application materials of the borrower from the blockchain deposit contract and parse them into a borrower feature vector X a: Through the event monitoring mechanism of the blockchain, monitor the MaterialEvidence event in the deposit contract to obtain the transaction hash of the materials submitted by each borrower. According to the transaction hash, call the query method Get Material of the deposit contract to obtain the IPFS hash of the loan application materials uploaded by the borrower. Through the IPFS protocol, access the distributed storage network and download the loan application material file corresponding to the hash to obtain a PDF or picture file containing the borrower's basic information, borrowing purpose, income certificate, and asset certificate. Use OCR (Optical Character Recognition) technology to perform text extraction and structured parsing on the borrower's basic information and borrowing purpose parts. The basic information fields include name, age, gender, education level, occupation, marital status, etc., and the borrowing purpose fields include labels such as consumption, business, house purchase, and medical treatment. Use object detection and semantic segmentation algorithms to extract key information from pictures such as payslips, bank statements, and tax returns in the income certificate part. The extracted fields include company name, start and end dates, pre-tax income, after-tax income, etc. Use OCR and NLP technologies to extract and identify the text and amounts of materials such as property ownership certificates, vehicle certificates, and stocks and funds in the asset certificate part. The extracted fields include asset type, asset value, and net asset value. Construct the extracted basic information, borrowing purpose, income certificate, and asset certificate fields into structured key-value pairs to form a unified borrowing feature vector X a 。
[0037] Obtain the borrower's credit report from a third-party credit reporting platform and parse it into a credit feature vector X b : Sign a data cooperation agreement with qualified credit reporting platforms such as the People's Bank of China Credit Information Center and Baihang Credit. Through the API interface provided by the credit reporting platform, pass in unique identifiers such as the borrower's ID number and name. After receiving the identity information, the credit reporting platform queries the credit report data of the borrower in its credit database, including fields such as credit score, historical overdue records, and debt-to-income ratio. The credit reporting platform performs desensitization processing on the queried credit report data, masking sensitive information of the borrower such as ID card and mobile phone number, and returns a structured JSON-format credit report. Parse the returned JSON-format credit report and extract the key credit information fields. The credit score field includes the score value (such as 650) and the score grade (such as A-, B+, etc.), the historical overdue record field includes the number of overdue accounts, overdue amount, and overdue duration, and the debt-to-income ratio field includes loan balance, credit card balance, and average monthly income. According to business rules and expert experience, set reasonable thresholds and risk levels for each credit information field. For example, a credit score below 600 is considered high risk, and a debt-to-income ratio exceeding 50% is considered high risk, etc. Construct the parsed credit information fields and the corresponding risk levels into structured key-value pairs to form a unified credit feature vector X b 。
[0038] According to the borrower's borrowing feature vector X a and credit feature vector X b , construct a user portrait as the input feature X; vertically splice the borrowing feature vector X a and credit feature vector X b to form the initial user portrait feature set X 0 . Among them, X a includes features such as the borrower's basic information, borrowing purpose, income certificate, and asset certificate, and X b includes features such as the borrower's credit score, historical overdue records, and debt-to-income ratio. The spliced X 0 integrates the multi-dimensional information of the borrower in a vector space, laying a foundation for the next step of processing.
[0039] Perform min-max normalization on the continuous features (such as age, income, debt ratio, etc.) in X 0 , scale the value range to the interval [0, 1], and eliminate the influence of dimension. Perform One-hot encoding on the discrete features (such as education level, occupation, marital status, etc.) in X 0 , and convert the categorical variable into a 0-1 binary feature. After preprocessing, the normalized user portrait feature set X 1 is obtained, and the feature value distribution is more balanced, which is conducive to the convergence and generalization of subsequent machine learning algorithms.
[0040] Use the Recursive Feature Elimination (RFE) algorithm for X 1Feature screening is performed. After each training, several of the least relevant features are removed, and this is recursively repeated for multiple rounds until N of the most relevant features are retained. Among them, feature relevance is measured by the prediction contribution of a machine learning model (such as logistic regression, decision tree, etc.) to the target variable Y (such as whether to default, whether to be high-risk, etc.). Through RFE, a feature subset with the largest amount of information and the lowest redundancy can be automatically selected, reducing the subsequent computational complexity and improving the representational ability of the user profile. In this embodiment, there is a loan user dataset containing 1000 samples and 20 features. The feature list is as follows: 1. Age, 2. Gender, 3. Education level, 4. Marital status, 5. Occupation, 6. Monthly income, 7. Housing situation, 8. Vehicle situation, 9. Credit score, 10. Credit card limit, 11. Number of historical overdue times, 12. Historical overdue amount, 13. Loan amount, 14. Loan term, 15. Loan interest rate, 16. Loan purpose, 17. Guarantee method, 18. Debt-to-income ratio, 19. Number of credit report inquiries. The target variable Y is whether the loan defaults, with values of 1 (default) or 0 (non-default). Our goal is to select a subset from these 20 features that has the greatest contribution to default prediction and construct a simple and effective user profile. Denote the original feature set as X1, with a total of 20-dimensional features. Select a basic machine learning model, such as logistic regression (LR). Initialize an empty list S to store the selected features. Based on the current feature set X 1 , train an LR model to obtain the weight coefficients of each feature. Sort the features in descending order according to the absolute value of the weight coefficients. Remove the m features with the smallest weights (m can be selected according to the actual situation, such as m = 1, 2, 5, etc.) to obtain a new feature set X 1 '. Delete the m removed features from S. Based on the new feature set X 1 ', repeat until X 1Only N features are left. The value of N can be determined through cross-validation, such as N = 5, 10, 15, etc. The finally obtained N features are used as the selected feature subset S to construct a user profile. For example: Income level: According to the monthly income, users are divided into three categories: high income (>100,000), medium income (50,000 - 100,000), and low income (<50,000). Credit rating: According to the credit score, users are divided into four categories: excellent (800 - 950 points), good (700 - 800 points), average (650 - 700 points), and poor (<650 points). Historical overdue: According to the number of overdue times, users are divided into three categories: never overdue (0 times), occasionally overdue (1 - 2 times), and frequently overdue (>=3 times). Loan amount: According to the loan amount, users are divided into three categories: high loan amount (>500,000), medium loan amount (100,000 - 500,000), and low loan amount (<100,000). Debt ratio: According to the debt-to-income ratio, users are divided into three categories: low debt (<30%), medium debt (30% - 50%), and high debt (>50%). These labels can intuitively reflect the credit risk levels of different users, help business personnel quickly identify and manage risks, and also provide a reference for subsequent credit strategy formulation and product recommendation. At the same time, since only 5 core features are used, the data dimension and calculation overhead are greatly reduced, and the interpretability and practicality of the profile are improved.
[0041] Through methods such as RFE, 5 most relevant features are selected to form the user profile feature matrix X 2 , which contains 1000 user samples. Next, calculate the similarity. For any two user samples x 2 and x i in X j , calculate their Euclidean distance in the 5-dimensional feature space: where x ik represents the value of user i on the k-th feature. Repeat step 1 for all sample pairs (i, j) to obtain a 1000x1000 distance matrix D. D is a symmetric matrix, and the diagonal elements are 0, indicating that the distance of each user to itself is 0. The smaller the element d ij in D, the closer the features of users i and j are, and the higher the similarity.
[0042] According to the distance matrix D, use the t-SNE algorithm to calculate the similarity matrix P between samples:
[0043] where α is the degree-of-freedom parameter of the t-distribution, usually taking values of 1 or 2. p ij can be understood as the conditional probability of taking x i as the neighbor of x j given x i . When x i and xj When the distance is very close, p ij will be relatively large; when x i and x j are at a relatively far distance, p ij will be relatively small.
[0044] Coordinate mapping. Randomly initialize the coordinates (y 1 , y 2 ) of 1000 sample points on a two-dimensional plane to obtain the initial mapping matrix Y. Similarly, calculate the similarity q ij between sample i and j under the low-dimensional coordinates of Y: Define the optimization objective as minimizing the KL divergence between P and Q: Intuitively, we hope that the similarity q ij under the two-dimensional coordinates is as close as possible to the similarity p ij in the high-dimensional feature space. Use the gradient descent method to update the coordinates in Y to continuously reduce the KL divergence. The formula is: where Y (t) represents the coordinates at the t-th iteration, η is the learning rate, is the gradient of the KL divergence with respect to Y. Repeat until the KL divergence converges or reaches the maximum number of iterations. Finally, obtain the optimized mapping matrix Y, and visualize Y to get the two-dimensional scatter plot of the user distribution. In the figure, similar users will be more clustered, while different users will be more dispersed.
[0045] Portrait clustering. Perform K-Means clustering on the optimized Y. First, randomly select K samples as the initial clustering centers, assuming K = 4. Calculate the Euclidean distance from each sample y i to the 4 clustering centers, and divide y i into the cluster with the closest distance. Calculate the coordinate average of the samples within each cluster and update the clustering centers. Repeat until the clustering centers no longer change. Finally, obtain 4 user portrait clusters.
[0046] Describe the portraits. For users within the first portrait cluster, calculate the mean and median of five features including their monthly income and credit score respectively. Assume the results are as follows: mean monthly income: 25,000 yuan, median: 18,000 yuan; mean credit score: 782 points, median: 796 points; mean number of overdue times: 1.5 times, median: 1 time; mean loan amount: 80,000 yuan, median: 50,000 yuan; mean debt ratio: 35%, median: 30%. Based on the above feature distributions, the typical features of this portrait cluster can be summarized as: below-average income, good credit, occasional overdue, moderate loan amount, and moderate debt ratio. Generate a portrait description such as "low to medium income - good credit - occasional overdue - moderate debt". Repeat steps 1-3 for the other three portrait clusters to obtain a set of typical user portrait labels. These labels characterize the characteristics of different user groups from multiple dimensions, facilitating the formulation of targeted marketing and risk control strategies. In this application, the dimensionality reduction mapping from high-dimensional user features to a low-dimensional space intuitively shows the distribution structure of users; the K-Means algorithm is used to achieve automatic clustering and partitioning of users to discover potential user groups; feature analysis and natural language description are used to summarize the feature portraits and labels of each user group.
[0047] One-hot encode the labels of K typical portrait clusters to generate a K-dimensional binary feature vector X 3 , representing the portrait category to which each user belongs. Horizontally concatenate the original N-dimensional feature X 2 and the newly generated K-dimensional feature X 3 to obtain the extended user portrait feature X 4 , where each user is represented by an (N + K)-dimensional vector. X4 combines the original features of the user and the portrait category, and can more comprehensively characterize the credit risk status of the user, providing more accurate input for subsequent default prediction and credit granting decisions.
[0048] As Figure 3 shown, use the borrower user portrait as the input feature X, use the litigation result of the training samples of historical closed loan disputes as the target variable Y, and use the litigation time-consuming of the samples as the training weight {w i}, train a LightGBM binary classification model. Using the trained LightGBM binary classification model, generate a litigation difficulty score, including: preprocess the input feature X, target variable Y, and training weight {w i} to obtain the training set D train and the validation set D val ; the input data includes: the user portrait feature matrix X, which contains N samples and M features; the target variable Y, representing the litigation result corresponding to each sample, 1 for winning the lawsuit and 0 for losing the lawsuit; the litigation time-consuming vector T, representing the litigation time-consuming corresponding to each sample, in days;
[0049] Normalize the input feature X, scaling the value of each feature to the interval [0, 1]: where x max and x min are the maximum and minimum values of this feature respectively. Normalization can eliminate the influence of the dimensions of different features, making the model more stable. Binarize the target variable Y, marking the winning result as 1 and the losing result as 0. This step is to transform the litigation result into a classification problem. Perform a logarithmic transformation on the litigation time T to obtain the training weight W: First, normalize T, scaling the litigation time of each sample to the interval [0, 1]: Then perform a logarithmic transformation on the normalized litigation time t i ' to obtain the training weight w: where α is the parameter of the logarithmic transformation, usually taking the value of 10. The logarithmic transformation can weaken the influence of outliers and at the same time amplify the weight of samples with short time consumption. After preprocessing, obtain the normalized feature matrix X', the binarized litigation result Y', and the logarithmically transformed training weight W.
[0050] Generate samples. According to the litigation result Y', divide the samples into a winning sample set S and a losing sample set F. For each winning sample s i , perform the following operations: Calculate the Euclidean distance between s i and all other winning samples, and select the nearest M 1 samples as the nearest neighbor set N(s i ). Randomly select a sample s i from N(s j ), and randomly select an interpolation point x i between the corresponding feature vectors x j and x i and x j : x new = x new + rand(0, 1) × (x i - x j - x i ), where rand(0, 1) represents a randomly generated number in the interval [0, 1]. The interpolation method can generate new samples based on the original samples to expand the dataset. Add x new as a new winning sample to S, with the corresponding weight being Repeat the above steps until the number of samples in S is greater than the number of samples in F to achieve sample balance.
[0051] For each losing sample f i , perform the following operations: Calculate the Euclidean distance between f i and all samples, and select the nearest M 2A sample is used as the nearest neighbor set N(f i ), and the majority class label y i in N(f maj ) is counted. If the class label (0) of f i is inconsistent with y maj , then f i is removed from F. This step can remove noisy samples and improve data quality. Repeat the above steps until all losing samples are traversed. Combine the processed winning sample set S and losing sample set F to obtain the balanced training set D balanced . Randomly divide D balanced into a training set D train and a validation set D val . The ratio can be set according to specific circumstances, such as 7:3 or 8:2. The validation set is used to evaluate the model performance during training to avoid overfitting.
[0052] Set the hyperparameters of the LightGBM binary classification model, mainly including: Learning rate (learning_rate): The step size of gradient descent in each iteration. A smaller learning rate can improve the training stability but the convergence speed is slower. Such as 0.01, 0.05, 0.1, etc. Number of leaves (num_leaves): The maximum number of leaf nodes in each tree. A larger value can improve the model's expressive ability but may lead to overfitting. Such as 31, 63, 127, etc. Maximum tree depth (max_depth): The maximum depth of each tree. A larger value can increase the model complexity but the training time will also increase. Such as 5, 7, 9, etc. Regularization coefficient (λ L1 , λ L2 ): The weights of the L1 and L2 regularization terms. A larger value can control the model complexity and prevent overfitting. Such as 0.1, 0.5, 1, etc. The optimal combination of these hyperparameters needs to be determined through repeated experiments and tuning. Generally, a benchmark value can be set first, and then grid search or random search can be performed on this basis.
[0053] Adopt the K-fold cross-validation method to divide the training set D train into K mutually exclusive subsets. Commonly used K values are 5 or 10. First, randomly shuffle D train , and then evenly divide it into K subsets in order; the number of samples in each subset is approximately 1 / K of the total number of samples; there is no overlap between any two subsets, that is, they do not contain the same samples; Cross-validation can make full use of the limited training data, avoid the influence of contingency at the same time, and improve the reliability of the model performance.
[0054] Sub-model training, using K-fold cross-validation, each time select K - 1 of these subsets as the training set E k , and the remaining 1 subset as the validation set Vk Repeat K times to obtain K sets of training-validation pairs. For the k-th set of training-validation pairs, according to the training weights W, set the weights of each training sample: For the training set E k for each sample x i in it, its weight is w i ; for each sample in the validation set V k , its weight is 1; the sample weights reflect the importance of different samples for model training. Samples with higher weights contribute more to the objective function and have a greater impact on the model. Use the weighted training set E k to train a LightGBM binary classification sub-model M k : Pass the hyperparameter dictionary to the training function of LightGBM, such as lgb.train(); Pass the features, labels, and weights of the training set Ek to the parameters of the training function respectively, such as train_set = lgb.Dataset(X, label = y, weight = w); Set the validation set V k , evaluate the performance of the model on the validation set during training, such as valid_sets = [valid_set]; Specify the binary classification task, such as objective = 'binary'; Train to obtain the sub-model M k ; Repeat until K sub-models M 1 , M 2 ,....., M K are obtained. These K sub-models are trained on different training-validation pairs respectively and have certain differences and complementarities.
[0055] Model integration. To combine the prediction results of the K sub-models and improve the generalization performance, a logistic regression model is adopted: Train K different types of sub-models, such as decision trees, random forests, neural networks, etc., to obtain the predicted probability values p 1 , p 2 ,....., p K . Calibrate the predicted probability values of each sub-model, such as using the PlattScaling method: For the k-th sub-model, use its predicted probability value p k and the true label y to train a logistic regression model: y = sigmoid(a k ×p k +b k ), where k represents the number of the k-th sub-model, and the value range is from 1 to K, where K is the total number of sub-models. p k represents the predicted probability value of the k-th sub-model for a certain sample, that is, the probability that the sample belongs to the positive example (y = 1). p kThe value range of k is [0, 1]. y represents the true label of the sample, which is a binary variable and takes values of 0 (negative example) or 1 (positive example). a represents the weight parameter of the calibration logistic regression model, which is used to adjust the predicted probability value p k 's scale. a is a real number, which can be positive or negative, and its specific value is learned by the model. b represents the bias parameter of the calibration logistic regression model, which is used to translate the predicted probability value p k 's center. b is a real number, which can be positive or negative, and its specific value is learned by the model. sigmoid function: represents the logistic function, which is used to convert a real number into a probability value within the interval (0, 1). Its definition is: where x is a real number, and exp represents the natural exponential function. After obtaining the calibration parameters a and b, p k is converted into the calibrated probability value p ': For each sample x val in the validation set D i : Use K sub-models to make predictions respectively, and obtain the calibrated predicted probability values Extract the meta-features of each sub-model, such as the AUC value, hyperparameters, the number of samples, etc., denoted as Combine the predicted probability values and meta-features into a new feature vector z i : z i = {p 1 ', p 2 ',......, p K ', m i1 , m i2 ,......, m iM}; Use z i and the true label y i as the training samples of logistic regression (z i , y i ); Use the new training set {(z i , y i )} to train a logistic regression model: Set the hyperparameters of logistic regression, such as the regularization strength, optimization algorithm, etc.
[0056] The weight vector θ = [θ 1 , θ 2 ,......, θ K , θ 1 ', θ 2 ',......, θ M '] is obtained by training; For the new test sample x: Use K sub-models to make predictions respectively, and obtain the calibrated predicted probability values p 1 ', p 2 ',......, p K'; Extract the meta-features m of each sub-model i1 , m i2 ,......, m iM ; Combine the feature vector z i = {p 1 ', p 2 ',......, p K ', m i1 , m i2 ,......, m iM}; Use the weight vector θ of the logistic regression model to calculate the weighted average probability:
[0057] p = sigmoid(θ 1 p 1 '+ θ 2 p 2 '+......+ θ K p K '+ θ 1 'm i1 + θ 2 'm i2 +......+ θ M 'm iM ); Where θ represents the weight vector of the integrated logistic regression model, which contains the weighted coefficients for the predicted probability values of the sub-models and the meta-features. The dimension of θ is (K + M), where K is the number of sub-models and M is the number of meta-features. θ 1 , θ 2 ,......, θ K represents the weight coefficients corresponding to the predicted probability values of the K sub-models in the integrated logistic regression model. θ k represents the importance of the predicted probability value p k ' of the k-th sub-model in the integrated model. θ 1 ', θ 2 ',......, θ M ' represents the weight coefficients corresponding to the M meta-features in the integrated logistic regression model. θ m ' represents the importance of the m-th meta-feature m im in the integrated model. x represents a new test sample, that is, the target sample to be predicted. x contains the feature information of the sample and is used for input into the sub-models and the integrated model for prediction.
[0058] p 1 ', p 2 ',......, p K ' represent the new probability values obtained after calibrating the predicted probability values of the K sub-models for the test sample x. p k ' represents the calibrated probability prediction that the k-th sub-model predicts the sample x to belong to the positive example. mi1 , m i2 ,......, m iM denotes M meta - feature values extracted from the i - th sub - model. m im denotes the value of the m - th meta - feature of the i - th sub - model, which reflects a certain characteristic or property of the sub - model. z represents the combined feature vector, which contains the calibrated prediction probability values of K sub - models and M meta - feature values. The dimension of z is (K + M), representing the comprehensive representation of the test sample x on different sub - models and meta - features. p represents the final prediction probability value of the integrated logistic regression model for the test sample x, that is, the comprehensive probability that the sample x belongs to the positive example. p is obtained by weighted averaging the calibrated prediction probability values p k ' of K sub - models and M meta - feature values m im . p is used as the final prediction probability for output; error analysis is performed on the final prediction results to find the misclassified samples and their feature distributions. Based on the results of error analysis, the training process of the sub - model is adjusted, such as increasing the weights of misclassified samples, modifying feature engineering, etc. Continuously repeat to optimize the sub - model and the integration effect until satisfactory performance is achieved.
[0059] This integrated learning method based on LightGBM and logistic regression can effectively combine the prediction results of multiple sub - models, improving the generalization performance and robustness of the model. Among them, the LightGBM sub - model is responsible for learning the internal patterns and rules of the samples, and the logistic regression model is responsible for learning the weight assignment and combination strategy of the sub - models. Through cross - validation and integrated learning, the information in the data can be fully mined while reducing the risk of overfitting.
[0060] Use the LightGBM binary classification model to predict the validation set D val , obtaining the winning probability value of each sample; select multiple different decision thresholds, calculate the precision and recall rate under each decision threshold, and draw the P - R curve; including: use the trained LightGBM binary classification model to predict each sample x val in the validation set D i , obtaining its winning probability value p i . Select multiple different decision thresholds t k , such as 0.1, 0.2,..., 0.9. For each threshold t k : Compare the winning probability value p i with t k . If p i ≥t k , then set the predicted label y i ' to 1 (winning), otherwise set it to 0 (losing); according to the predicted label y i ' and the true label y i, calculate the following evaluation metrics: TP (True Positive): the number of samples with a true label of 1 and a predicted label of 1; FP (False Positive): the number of samples with a true label of 0 and a predicted label of 1; TN (True Negative): the number of samples with a true label of 0 and a predicted label of 0; FN (False Negative): the number of samples with a true label of 1 and a predicted label of 0.
[0061] Calculate the precision, recall, and F1-score: Precision = TP / (TP + FP); Recall = TP / (TP + FN); F1-score:; Take the precision and recall at each decision threshold t k as the horizontal and vertical coordinates to plot the P-R curve. The horizontal axis is recall, and the vertical axis is precision; each point on the curve corresponds to the precision and recall at a decision threshold; the closer the curve is to the upper right corner, the better the comprehensive performance of the model; the P-R curve intuitively shows the precision-recall trade-off of the model at different decision thresholds and is an important tool for evaluating binary classification models.
[0062] Calculate the Youden index based on the P-R curve and select the winning probability value corresponding to the maximum Youden index as the optimal decision threshold t_best; including: for each point (Precision i , Recall i ) on the P-R curve, calculate its F1-score and Youden index: F1-score: Youden index:
[0063] J i = Precision i + Recall i - 1; respectively find the maximum values of the F1-score and Youden index, denoted as F1_max and J_max, and the corresponding decision thresholds t F1 and t J . According to the preset weight coefficient λ, calculate the optimal decision threshold: t best = λ × t F1 +(1 - λ) × t J ; where λ ranges from [0, 1] and represents the relative importance attached to the F1-score and Youden index. When λ = 1, it is completely based on the F1-score; when λ = 0, it is completely based on the Youden index; when λ = 0.5, both are considered with equal weights; the optimal decision threshold t bestIt combines the requirements of precision and recall, and balances the accuracy and recall rate of the model. In practical applications, the value of λ can be adjusted according to business requirements and cost-benefit.
[0064] According to the loan status field, all overdue borrower samples are screened out, and their user portrait features are extracted to form the test set Dtest. Using the trained LightGBM binary classification model, each sample xi in the test set Dtest is predicted to obtain its winning probability value p i . The winning probability value p of each overdue borrower i is compared with the optimal decision threshold t best : If p i ≥ t best , the risk level is "high risk", and it is predicted that the borrower will win; if p i < t best , the risk level is "low risk", and it is predicted that the borrower will lose; generate the winning risk level label for each overdue borrower.
[0065] For each overdue borrower, according to its winning risk level, winning probability value and overdue days, design a litigation difficulty scoring rule, such as: the winning risk level is "high risk", and the winning probability value >= 0.8, and the overdue days >= 90 days, then the litigation difficulty score is 5 (the most difficult); the winning risk level is "high risk", and the winning probability value >= 0.7, and the overdue days >= 60 days, then the litigation difficulty score is 4 (more difficult); the winning risk level is "high risk", and the winning probability value >= 0.6, and the overdue days >= 30 days, then the litigation difficulty score is 3 (medium); the winning risk level is "low risk", and the winning probability value < 0.6, and the overdue days < 30 days, then the litigation difficulty score is 2 (easier); in other cases, the litigation difficulty score is 1 (the easiest); generate the litigation difficulty score for each overdue borrower. For borrowers with high winning risk and great litigation difficulty, legal means can be taken preferentially to recover; while for borrowers with low winning risk and small litigation difficulty, risks can be resolved through flexible methods such as negotiation and communication. At the same time, these prediction results can also be used as a reference for post-loan management, such as adjusting the loan amount, optimizing the approval rules, etc., so as to improve the overall credit risk control level.
[0066] So far, a litigation difficulty score has been generated for each overdue borrower, with a value range of 1-5. We can preset a set of litigation strategy templates, each of which corresponds to a different litigation difficulty level, including specific litigation processes, measures and precautions. The following five templates are used as examples for explanation: Template 1 (score = 1, easiest): Pre-litigation stage: send a collection notice to guide the two parties to negotiate and reconcile; litigation stage: file a simplified procedure lawsuit separately and adopt a trial in absentia; execution stage: apply for compulsory execution and give priority to deducting bank deposits. Template 2 (score = 2, easier): Pre-litigation stage: send a lawyer's letter, inform the litigation risks, and promote debt recognition; litigation stage: file a common procedure lawsuit separately, and provide appropriate evidence and cross-examination; execution stage: apply for compulsory execution, deduct bank deposits and seal movable property. Template 3 (score = 3, medium): Pre-litigation stage: notarize the delivery of lawyer's letters, and take preservation measures when necessary; litigation stage: choose to sue separately or jointly according to the situation, and provide comprehensive evidence and cross-examination; execution stage: apply for compulsory execution, deduct deposits, and seal movable and immovable property. Template 4 (score = 4, more difficult): Pre-litigation stage: notarization and delivery of lawyer's letter, and application for pre-litigation preservation measures at the same time; litigation stage: joint litigation, comprehensive and in-depth evidence and cross-examination, and application for appraisal when necessary; execution stage: application for compulsory execution, deduction of deposits, seizure and auction of movable and immovable property. Template 5 (score = 5, most difficult): Pre-litigation stage: notarization and delivery of lawyer's letter, pre-litigation preservation, and prosecution of guarantors when necessary; litigation stage: joint litigation, comprehensive and in-depth evidence and cross-examination, and application for appraisal as much as possible; execution stage: application for compulsory execution, deduction of deposits, seizure and auction of movable and immovable property, and additional persons to be executed. According to the litigation difficulty score of overdue borrowers, we can match the corresponding litigation strategy template to generate a personalized litigation plan: extract the litigation difficulty score of overdue borrowers. According to the score value, match the preset litigation strategy template: if score = 1, match template 1; if score = 2, match template 2; if score = 3, match template 3; if score = 4, match template 4; if score = 5, match template 5. Extract key elements from the matched template, such as litigation process, measures and precautions, and fill them into the corresponding parts of the personalized litigation plan.
[0067] In the personalized litigation plan, specific content and suggestions can also be added according to other characteristics of overdue borrowers, such as overdue amount, region, occupation, etc. For example, if the overdue amount is large (such as exceeding 500,000), joint litigation is recommended during the litigation stage; if the borrower is located in a remote area, sufficient time and costs are recommended to be reserved during the execution stage; if the borrower is a corporate legal person, the company and the individual are recommended to be sued simultaneously during the pre-litigation stage. Generate a complete personalized litigation plan, the content including but not limited to: basic information of the borrower: name, ID number, address, contact information, etc.; basic information of the loan: loan amount, interest rate, term, overdue time, overdue amount, etc.; collection records: collection time, method, result, etc.; litigation difficulty assessment: probability of winning, litigation difficulty score, main basis, etc.; litigation strategy: specific measures and processes in each stage of pre-litigation, litigation, and execution; special suggestions: additional suggestions and tips based on the characteristics of the borrower. Risk warnings: possible legal risks, time costs, execution difficulties, etc. This application automatically generates a personalized litigation plan that fits the risk characteristics of overdue borrowers according to their litigation difficulty scores. These plans can provide reference and guidance for lawyers and collection staff, improve litigation efficiency, reduce legal risks, and enhance the odds of collection success.
Claims
1. A blockchain-based smart litigation method for small financial debt disputes, characterized by: include: Collect loan business data; According to the loan business data, set up smart contracts and deploy them to the alliance blockchain, with all parties involved in the loan as nodes; Obtain the borrower's repayment data, synchronize the repayment data to the blockchain, and update the loan status field in the smart contract based on the comparison results of the repayment data and the repayment plan; the loan status field includes normal and overdue; Get the loan status field in the smart contract and determine whether it contains the overdue field. If it does, calculate the overdue days and penalty interest receivable. For borrowers whose loan status field is overdue, a borrower user profile is constructed based on the borrower information in the loan application materials and the corresponding credit report; The borrower user profile is used as the input feature X, the litigation results of the historical loan dispute training samples that have been closed are used as the target variable Y, and the litigation time of the samples is used as the training weight W. The LightGBM binary classification model is trained and the litigation difficulty score is generated using the trained LightGBM binary classification model. Based on the litigation difficulty score, the preset litigation strategy template is matched to generate a personalized litigation plan.
2. The blockchain-based smart litigation method for small financial debt disputes according to claim 1 is characterized by: Build a borrower user profile, including: Extract the borrower’s loan application materials from the blockchain’s evidence contract and parse the loan application materials into a structured loan feature vector X a ; Among them, the loan application materials include the borrower's basic information, loan purpose, income proof and asset proof; Obtain the borrower’s credit report from a third-party credit reporting platform and parse the credit report into a structured credit feature vector X b ; Among them, the credit report includes the borrower's credit score, historical delinquency record and debt-to-income ratio; According to the borrower's borrowing feature vector X a and the credit feature vector X b , construct user portraits as input features X.
3. The blockchain-based smart litigation method for small financial debt disputes according to claim 2 is characterized by: Construct a user profile as input feature X, including: The borrowing feature vector X a and the credit feature vector X b Vertical splicing to obtain the initial user portrait feature set X0; Perform minimum-maximum normalization on the continuous variables in X0; perform One-hot encoding on the discrete variables in X0 to obtain the user portrait feature set X1; The recursive feature elimination algorithm RFE is used to select features from X1, and the top N features most relevant to the target variable Y are selected to obtain the user portrait feature set X2 after dimensionality reduction; Perform nonlinear mapping on X2, project the N-dimensional features onto a two-dimensional plane, and obtain a user portrait distribution map; Use K-Means algorithm to cluster the user portrait distribution map and obtain K user portrait clusters; Perform statistical analysis on K typical user portrait clusters to obtain the characteristic distribution of borrowers in each cluster, and generate K portrait labels based on the characteristics of the cluster center of each cluster; One-hot encode the K image labels to generate a K-dimensional binary feature vector X3; The N-dimensional user portrait feature set X2 and the K-dimensional binary feature vector X3 are horizontally spliced to obtain the expanded user portrait feature set X4 as the input feature X.
4. The blockchain-based smart litigation method for small financial debt disputes according to claim 2 is characterized by: Generates a litigation difficulty score including: For input feature X, target variable Y and training weight {w i }Preprocessing to obtain the training set D train and validation set D val ; Set the hyperparameters of the LightGBM binary classification model and use K-fold cross validation to train the training set D train Divided into K mutually exclusive subsets; among them, the hyperparameters include learning rate, number of leaf nodes, maximum tree depth and regularization coefficient; On each fold training subset, according to the training weight {w i } Set the weight of each training sample, train a LightGBM binary classification sub-model, repeat K times, and obtain K LightGBM binary classification sub-models; Use the logistic regression model to perform weighted averaging on K LightGBM binary classification sub-models to generate a LightGBM binary classification model; Use the LightGBM binary classification model to classify the validation set D val Make predictions to obtain the winning probability value of each sample; select multiple different decision thresholds, calculate the precision and recall under each decision threshold, and draw a PR curve; Calculate the Youden index based on the PR curve and select the winning probability value corresponding to the maximum Youden index as the optimal decision threshold t best ; According to the borrowers whose loan status field is overdue, the corresponding user portrait is extracted as the input feature X, and the trained LightGBM binary classification model is used for prediction to obtain the winning probability value p of each overdue borrower. i ; The predicted probability of winning p i and the optimal decision threshold t best Make comparisons and generate a risk level of winning the case; Generate a litigation difficulty score based on the risk level of winning, probability of winning and number of overdue days.
5. The blockchain-based smart litigation method for small financial debt disputes according to claim 4 is characterized by: For input feature X, target variable Y and training weight {w i }Preprocessing, including: Normalize the input feature X; Binarize the target variable Y, mark the winning result as 1 and the losing result as 0; For training weights {w i } Perform logarithmic transformation.
6. The blockchain-based smart litigation method for small financial debt disputes according to claim 5 is characterized by: For training weights {w i } Perform logarithmic transformation, including: For each training sample i, extract the corresponding litigation time t i ; Time consuming litigation i After normalization, we get t i '; The normalized litigation time t i 'Perform logarithmic transformation to obtain training weights {w i },w i The calculation formula is: Among them, α is the parameter of logarithmic transformation.
7. The blockchain-based smart litigation method for small financial debt disputes according to claim 5 is characterized by: Get the training set D train and validation set D val ,include: According to the litigation results of the historical loan dispute training samples that have been closed, for each winning sample s i , from s i A random sample of s is selected from the winning samples of M1 nearest neighbors j ; Among them, Y[s i ]=1; In the winning sample i and winning samples j Corresponding user profile feature x i and x j Randomly select an interpolation point x between new =x i +rand(0,1)×(x j -x i ), and x new Add the new winning samples to the training set; repeat until the number of winning samples is greater than the number of losing samples; For each losing case f in the litigation results of the historical loan dispute training samples that have been closed i , calculate f i Corresponding user portrait x i The Euclidean distance to all samples in the sample set; select the M2 samples with the closest distance, and determine the majority class corresponding to the M2 nearest neighbor samples according to the majority voting rule; where Y[f i ]=0; If f i The category label 0 is inconsistent with the majority category in M2 neighbors, then f i The user portrait x corresponding to the set i Remove from the training set, repeat until all the losing samples are traversed, and get the training set D train,balanced ; The training set D train,balanced , randomly divided into training set D train and validation set D val .
8. The blockchain-based smart litigation method for small financial debt disputes according to claim 4 is characterized by: Draw a PR curve, including: According to the LightGBM binary classification model, the validation set D val , select multiple different decision thresholds, and convert the default probability prediction value into a binary prediction label according to each decision threshold; For each decision threshold, the precision and recall are calculated based on the corresponding binary predicted label and the corresponding true label: Precision=TP / (TP+FP) Recall=TP / (TP+FN) F1Score=2*Precision*Recall / (Precision+Recall) Among them, TP represents the number of samples correctly predicted as winning, FP represents the number of samples incorrectly predicted as winning, and FN represents the number of samples incorrectly predicted as losing; The calculation results of the precision rate Precision and recall rate Recall under each decision threshold are plotted as a PR curve.
9. The blockchain-based smart litigation method for small financial debt disputes according to claim 4 is characterized by: Calculate the Youden index based on the PR curve and select the winning probability value corresponding to the maximum Youden index as the optimal decision threshold t best ,include: According to each point on the PR curve (Precision i ,Recall i ), calculate the corresponding comprehensive evaluation index F1Score i : Where i is the number of the point on the PR curve; For each point on the PR curve (Precision i ,Recall i ), calculate the corresponding Youden index J i : J i =Precision i +Recall i -1 Get the maximum value of F1 Score and Youden index J respectively max and J max , extract the corresponding decision threshold t F1 and t J ; The optimal decision threshold t is calculated according to the following formula best : t best =λ×t F1 +(1-λ)×t J Among them, λ is the weight coefficient.
10. A blockchain-based smart litigation system for small financial debt disputes, characterized by: include: At least one processing unit; used to execute instructions to implement the blockchain-based smart litigation method for small financial debt disputes as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Debt collection method, system and device based on block chain, medium and product
CN112767141A
Block chain-based micro financial creditor's right dispute smart litigation system
CN113793208A
Portrayal generation method and device, server and storage medium
CN114902212A
Overdue risk assessment method and device based on user portrait, and electronic equipment
CN118521399A
Loan risk control method, electronic device and readable storage medium
WO2019061989A1