Asset data grouping method and device, electronic equipment and storage medium

By applying a classified decision tree model in asset data management, the risk categories of asset data are automatically identified, and the problems of inefficiency and high error rates caused by traditional manual dependence are solved, and more efficient and accurate asset data grouping is achieved.

CN120146997APending Publication Date: 2025-06-13SHENZHEN LEXIN SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317908.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional asset data management processes rely on manual experience, resulting in low processing efficiency and prone to human errors when facing large amounts of data and complex financial markets.

Method used

The classified decision tree model is used to automatically identify the risk categories of asset data, and through data cleaning and feature extraction, manual intervention is reduced and processing efficiency is improved.

Benefits of technology

It realizes the automation of the asset data grouping process, improves the accuracy and efficiency of grouping, and reduces the occurrence of human errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146997A_ABST
    Figure CN120146997A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data request processing, and discloses an asset data grouping method, which comprises the following steps: acquiring a plurality of risk indexes corresponding to asset data according to identifiers of the asset data, and a preset mapping relationship between the risk indexes and a classification decision tree model; extracting a first key feature and a second key feature of the asset data, and obtaining a classification decision tree model corresponding to each risk index of the asset data; inputting the first key feature and the second key feature into a classification decision tree model, performing path retrieval along a tree structure of the classification decision tree model to obtain a risk category of the asset data, and obtaining a data grouping rule of the asset data from a preset data grouping rule base according to the risk category; and matching the first key feature and the second key feature by using a data grouping rule to obtain a data group to which the asset data belongs, and storing the asset data in the data group. According to the invention, automatic asset data grouping is realized, and the grouping efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data request processing, and particularly to an asset data grouping method, device, electronic device, and storage medium. Background Art

[0002] In the financial industry, the management of asset data and risk control are crucial for ensuring the stability and profitability of financial institutions.

[0003] In the traditional asset data management process, asset data grouping usually relies on the experience of analysts and is completed by manually reviewing a large number of financial statements, market information, and other relevant materials. However, with the increasing complexity of the financial market and the explosive growth of data volume, this manual-based approach gradually reveals its limitations.

[0004] For example, financial institutions receive thousands of new loan applications (asset data) every day. To group each asset data into a risk category, it is necessary to conduct a detailed analysis of various information provided by the applicants. These information include but are not limited to customer attributes such as income level, employment status, credit history, debt ratio, and value attributes related to the loan itself such as loan amount and term.

[0005] Analysts will conduct a preliminary screening based on their own experience and standard scorecards, and then senior approval personnel will make the final decision. Due to the different experiences and preferences of each analyst, it may lead to different grouping results for asset data under the same conditions. At the same time, when facing a large number of new loan applications, the processing time is long and human errors are prone to occur.

[0006] Therefore, how to reduce the need for manual grouping of asset data and improve the automation level of the asset data grouping process has become an urgent technical problem to be solved. Summary of the Invention

[0007] In view of the above, it is necessary to provide an asset data grouping method, which aims to automatically identify the risk category of asset data through a classification decision tree model, reduce manual intervention, improve processing efficiency, comprehensively group asset data from multiple dimensions, and provide a more accurate grouping accuracy rate.

[0008] The asset data grouping method provided by the present invention includes:

[0009] Receiving a grouping request for grouping asset data to be grouped, obtaining the identifier of the asset data from the grouping request, and obtaining multiple risk indicators corresponding to the asset data and the mapping relationship between the preset risk indicators and the classification decision tree model from a preset database according to the identifier;

[0010] Perform data cleaning on the asset data, extract the first key features associated with customer attributes and the second key features associated with asset value attributes from the asset data after data cleaning, and obtain the classification decision tree models corresponding to the respective risk indicators of the asset data from a preset model library according to the multiple risk indicators corresponding to the asset data and the mapping relationship;

[0011] Input the first key features and the second key features into the classification decision tree models corresponding to the respective risk indicators, perform path retrieval along the tree structure of the classification decision tree models to obtain the risk categories of the asset data, and obtain the data grouping rules of the asset data from a preset data grouping rule library according to the risk categories;

[0012] Match the first key features and the second key features by using the data grouping rules to obtain the data group to which the asset data belongs, store the asset data in the data group to obtain the grouping result of the asset data, and feedback the grouping result to the terminal corresponding to the grouping request.

[0013] Optionally, the extraction of the first key features associated with customer attributes includes:

[0014] Perform word segmentation on the asset data after data cleaning to obtain multiple word segments;

[0015] Select the first keywords highly associated with customer attributes from the multiple word segments;

[0016] Convert the first keywords into feature vectors to obtain the first key features.

[0017] Optionally, the extraction of the second key features associated with asset value attributes includes:

[0018] Identify the accounting statements of a preset type in the asset data after data cleaning, and extract the derived indicators representing the asset characteristics from the accounting statements;

[0019] Convert the numerical data corresponding to the derived indicators into feature vectors to obtain the second key features.

[0020] Optionally, the classification decision tree model is obtained in the following manner:

[0021] Obtain historical asset data and the multiple risk indicators corresponding to the historical asset data from a preset historical asset database;

[0022] Construct a training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators;

[0023] Divide the training sample set corresponding to each risk indicator into a training set and a test set according to a preset ratio;

[0024] Use a preset decision tree algorithm to construct an initial decision tree model for each risk indicator respectively;

[0025] Use the training set corresponding to each risk indicator to train the corresponding initial decision tree model, and use the test set corresponding to each risk indicator to test the trained initial decision tree model. Stop training until the evaluation indicators of the initial decision tree model meet the preset criteria, and obtain the classification decision tree model corresponding to each risk indicator.

[0026] Optionally, the constructing of the training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators includes:

[0027] Perform data cleaning on the historical asset data, perform word segmentation on the cleaned historical asset data to obtain multiple word segments, select the first keywords highly correlated with customer attributes from the multiple word segments, convert the first keywords into feature vectors, and obtain the first initial features;

[0028] Identify the preset type of financial statements in the cleaned historical asset data, extract the derivative indicators representing asset characteristics from the financial statements, convert the numerical data corresponding to the derivative indicators into feature vectors, and obtain the second initial features;

[0029] Extract historical samples from the cleaned historical asset data according to each risk indicator to obtain historical samples corresponding to each risk indicator category;

[0030] Construct the labels of the historical samples corresponding to each risk indicator according to the first initial features, the second initial features and the corresponding risk indicators, and obtain the training sample set corresponding to each risk indicator.

[0031] Optionally, the using of the preset decision tree algorithm to construct an initial decision tree model for each risk indicator respectively includes:

[0032] According to the first key features and the second key features corresponding to each risk indicator, select the preset Gini coefficient as the standard for the splitting nodes of the preset decision tree algorithm;

[0033] Starting from the root node of the preset decision tree algorithm, for each risk indicator, use the preset Gini coefficient to calculate each splitting point, and select the best splitting point for splitting;

[0034] Repeat the step of using the preset Gini coefficient to calculate each splitting point and selecting the best splitting point for splitting until the stopping condition is met, and obtain the initial decision tree model.

[0035] Optionally, inputting the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, and performing path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data includes:

[0036] Concatenate the first key feature and the second key feature, and input the concatenated feature vector obtained by concatenation into the classification decision tree models corresponding to the respective risk indicators;

[0037] According to the concatenated feature vector, perform path selection along the tree structure of the classification decision tree model until reaching the final leaf node, and obtain the risk category of the asset data for the corresponding risk indicator.

[0038] To solve the above problems, the present invention further provides an asset data grouping device, and the device includes:

[0039] A receiving module, configured to receive a grouping request for grouping the asset data to be grouped, obtain the identifier of the asset data from the grouping request, and obtain the multiple risk indicators corresponding to the asset data and the mapping relationship between the preset risk indicators and the classification decision tree models from a preset database;

[0040] An extraction module, configured to perform data cleaning on the asset data, extract the first key feature associated with the customer attribute and the second key feature associated with the asset value attribute from the asset data after data cleaning, and obtain the classification decision tree models corresponding to the respective risk indicators of the asset data from a preset model library according to the multiple risk indicators corresponding to the asset data and the mapping relationship;

[0041] A retrieval module, configured to input the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, perform path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data, and obtain the data grouping rule of the asset data from a preset data grouping rule library according to the risk category;

[0042] A grouping module, configured to use the data grouping rule to match the first key feature and the second key feature, obtain the data group to which the asset data belongs, store the asset data in the data group, obtain the grouping result of the asset data, and feed back the grouping result to the terminal corresponding to the grouping request.

[0043] To solve the above problems, the present invention further provides an electronic device, and the electronic device includes:

[0044] At least one processor; and,

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores an asset data grouping program executable by the at least one processor, and the asset data grouping program is executed by the at least one processor to enable the at least one processor to execute the above-mentioned asset data grouping method.

[0047] To solve the above problems, the present invention also provides a computer-readable storage medium, on which an asset data grouping program is stored, and the asset data grouping program can be executed by one or more processors to implement the above-mentioned asset data grouping method.

[0048] Compared with the prior art, the present invention performs data cleaning and feature extraction on asset data, analyzes the key features of the asset data using a classification decision tree to obtain the risk categories of the asset data, obtains the data grouping rules for the asset data according to the risk categories and performs a grouping operation on the asset data to obtain the data groups to which the asset data belongs, stores the asset data in the data groups to obtain the grouping result of the asset data, and feeds back the grouping result to the terminal corresponding to the grouping request, realizing the full-process automation from data cleaning, feature extraction to classification, reducing manual intervention, improving the processing speed and accuracy of asset data grouping, comprehensively considering multiple risk indicators, and constructing a dedicated classification decision tree model for each risk indicator to ensure comprehensive and accurate risk identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic flowchart of the asset data grouping method provided by an embodiment of the present invention;

[0050] Figure 2 It is a schematic module diagram of the asset data grouping device provided by an embodiment of the present invention;

[0051] Figure 3 It is a schematic structural diagram of an electronic device for implementing the asset data grouping method provided by an embodiment of the present invention;

[0052] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] It should be noted that the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes, and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0055] Referring to Figure 1 As shown, it is a schematic flowchart of an asset data grouping method provided by an embodiment of the present invention. This method is executed by an electronic device.

[0056] In this embodiment, the asset data grouping method includes:

[0057] S1. Receive a grouping request for grouping the asset data to be grouped, obtain the identifier of the asset data from the grouping request, and obtain multiple risk indicators corresponding to the asset data and the mapping relationship between the preset risk indicators and the classification decision tree model from a preset database according to the identifier.

[0058] In this embodiment, after the server receives a grouping request for grouping the asset data to be grouped, it obtains the identifier of the asset data from the grouping request. The grouping request may come from the business requirements within the enterprise or the requirements of external customers. The identifier of the asset data is a unique identification code used to distinguish the identity of the asset package. The identifier includes loan number, product code, ID, batch number, etc.

[0059] Use the identifier as a query condition to search in the pre-set database for multiple risk indicators corresponding to the asset data and the mapping relationship between the preset risk indicators and the classification decision tree model.

[0060] The multiple risk indicators corresponding to the asset data refer to various quantitative or measurable factors that can reflect the risk levels of the asset in aspects such as operation, market fluctuations, credit, etc. The risk indicators include credit risk indicators (credit score, default probability), market risk indicators (market volatility, interest rate sensitivity), liquidity risk indicators (liquidity ratio, trading activity), and operation risk indicators, etc.

[0061] The mapping relationship between the preset risk indicators and the classification decision tree model is a corresponding rule. It determines which classification decision tree model should be used for data processing, risk classification, and grouping operations for the preset risk indicators or combinations of risk indicators. For example, for asset data within a high credit score range (such as 700 - 850 points) in credit risk indicators, it is mapped to a relatively simple classification decision tree model.

[0062] The mapping relationship between the preset risk indicators and the classification decision tree model can be obtained in the following way: Use decision tree algorithms (such as C4.5, CART, etc.) to train a large amount of asset data. During the training process, the algorithm will automatically construct a decision tree model based on the risk indicators in the data and determine which risk indicators are used as judgment conditions at the nodes of the tree. For example, when training credit card consumption credit data, the algorithm may find that risk indicators such as the cardholder's consumption amount in the last 6 months and whether there is an overdue repayment record have an important impact on credit risk, thus constructing a corresponding decision tree model and establishing the mapping relationship between these risk indicators and the model.

[0063] S2. Perform data cleaning on the asset data. From the asset data after data cleaning, extract the first key features associated with customer attributes and the second key features associated with asset value attributes. According to the multiple risk indicators corresponding to the asset data and the mapping relationship, obtain the classification decision tree models corresponding to each risk indicator of the asset data from the preset model library.

[0064] In this embodiment, data cleaning is performed on the asset data. Data cleaning includes removing duplicate records, filling in missing values, correcting incorrect data, standardizing formats, and eliminating invalid features. From the asset data after data cleaning, extract the first key features associated with customer attributes and the second key features associated with asset value attributes. According to the multiple risk indicators corresponding to the asset data and the mapping relationship, obtain the classification decision tree models corresponding to each risk indicator of the asset data from the preset model library.

[0065] In the preset model library, there are multiple trained classification decision tree models. Each classification decision tree model is designed for a preset type of risk indicator, such as default probability, expected loss rate, etc. Use the multiple risk indicators corresponding to the asset data and the mapping relationship to select a classification decision tree model suitable for current asset data for risk assessment from the model library.

[0066] In one embodiment, the extraction of the first key features associated with customer attributes includes:

[0067] Perform word segmentation on the asset data after data cleaning to obtain multiple word segments;

[0068] Select the first keyword that is highly correlated with customer attributes from multiple word segments;

[0069] Convert the first keyword into a feature vector to obtain the first key feature.

[0070] Perform word segmentation on the text content in the asset data after data cleaning to obtain multiple word segments, and select the first keyword that is highly correlated with customer attributes from the multiple word segments according to word frequency statistics or TF-IDF.

[0071] Word frequency statistics refer to selecting word segments with higher occurrence frequencies. TF-IDF refers to selecting words that appear frequently in a document in the asset data but are relatively rare in the entire document set. Convert the selected first keyword into a feature vector in numerical form to obtain the first key feature. Convert the first keyword into a feature vector according to the following method: create a vocabulary, and then generate a vector for each document, where each dimension corresponds to a word in the vocabulary, and the value represents the number of times the word appears in the document. Query the vocabulary with the first keyword to obtain the feature vector of the first keyword.

[0072] In one embodiment, the extraction of the second key feature associated with the asset value attribute includes:

[0073] Identify the preset type of financial statements in the asset data after data cleaning, and extract the derived indicators representing the asset characteristics from the financial statements;

[0074] Convert the numerical data corresponding to the derived indicators into a feature vector to obtain the second key feature.

[0075] Identify the preset type of financial statements in the asset data after data cleaning. The preset type of financial statements includes the balance sheet, personal property registration form, cash flow statement, etc. For each type of statement, select the derived indicator that best reflects the asset characteristics. The following are the derived indicators for different types of statements: The derived indicators of the balance sheet include total assets, shareholders' equity, current ratio, and quick ratio. The derived indicators of the personal property registration form include the proportion of major asset categories and asset dispersion. The derived indicators of the cash flow statement include net cash flow from operating activities, free cash flow, and ending balance of cash and cash equivalents.

[0076] Convert the numerical data corresponding to the derived indicators into a feature vector to obtain the second key feature.

[0077] In one embodiment, the classification decision tree model is obtained according to the following method:

[0078] Obtain historical asset data and multiple risk indicators corresponding to the historical asset data from a preset historical asset database;

[0079] Construct a training sample set corresponding to each risk indicator based on the historical asset data and the multiple risk indicators;

[0080] Divide the training sample set corresponding to each risk indicator into a training set and a test set according to a preset ratio;

[0081] Use a preset decision tree algorithm to construct an initial decision tree model for each risk indicator respectively;

[0082] Use the training set corresponding to each risk indicator to train the corresponding initial decision tree model, and use the test set corresponding to each risk indicator to test the trained initial decision tree model. Stop training until the evaluation indicators of the initial decision tree model meet the preset criteria, and obtain the classification decision tree model corresponding to each risk indicator.

[0083] Extract historical asset data from a pre-established database containing a large number of past asset records. The historical asset data includes not only customer attributes and asset value attributes, but also a variety of risk indicators related to the historical asset data. The risk indicators include but are not limited to default probability, expected loss rate, credit score change, etc. Default probability: the possibility that a borrower fails to repay the loan on time. Expected loss rate: the proportion of loss expected to occur if a default occurs. Credit score change: the change in the borrower's credit score during the loan period.

[0084] Create a set of labeled training sample sets for each risk indicator. For each risk type (such as default probability), each sample data in the training sample set contains two sets of features (i.e., customer attributes and asset value attributes) and the actual risk result (label) corresponding to the asset. These training sample sets will be used to train the initial decision tree model to learn how to predict risks based on the input features.

[0085] Further divide the training sample set of each risk indicator into two parts, a training set and a test set. Use a pre-selected decision tree algorithm (such as CART, ID3, C4.5, etc.) to create an initial decision tree model for each risk indicator. A decision tree is a supervised learning method that determines the final class or numerical output through a series of rules (node splitting). Use the data in the training set to adjust the structure and parameters of the initial decision tree model so that the initial decision tree model can better fit the training data. After each training, use the test set to evaluate the performance of the model and calculate evaluation indicators such as accuracy, recall rate, F1 score, etc. If the performance of the initial decision tree model reaches the preset standard (such as the error rate is lower than a certain threshold), it is considered that the training is completed; otherwise, continue to adjust and optimize until the requirements are met.

[0086] In other embodiments, during the process of constructing a classification decision tree model, pre-pruning is performed on the classification decision tree model to stop the growth of the tree in advance to avoid creating an overly complex tree. Specifically, it includes: during the process of constructing a classification decision tree model, restricting the maximum depth of the tree to prevent the tree from growing too deep. It is possible to stipulate the minimum number of samples that a node needs to contain to continue splitting, stipulate the minimum number of samples that must be contained in a leaf node, and restrict the total number of leaf nodes of the entire tree. Applying the above methods directly during the training process simplifies the tree structure, removes those branches that contribute little or no contribution to the prediction accuracy, thereby reducing the complexity of the classification decision tree model and improving its performance on unseen data.

[0087] In one embodiment, the constructing a training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators includes:

[0088] Performing data cleaning on the historical asset data, performing word segmentation processing on the historical asset data after data cleaning to obtain a plurality of word segments, selecting first keywords highly correlated with customer attributes from the plurality of word segments, converting the first keywords into feature vectors to obtain first initial features;

[0089] Identifying accounting statements of a preset type in the historical asset data after data cleaning, extracting derivative indicators representing asset characteristics from the accounting statements, and converting the numerical data corresponding to the derivative indicators into feature vectors to obtain second initial features;

[0090] Extracting historical samples from the historical asset data after data cleaning according to each risk indicator to obtain historical samples corresponding to each category of risk indicators;

[0091] Constructing labels for the historical samples corresponding to each risk indicator according to the first initial features, the second initial features, and the corresponding risk indicators to obtain a training sample set corresponding to each risk indicator.

[0092] Performing word segmentation processing on the text content in the historical asset data after data cleaning to obtain a plurality of word segments, and selecting first keywords highly correlated with customer attributes from the plurality of word segments according to word frequency statistics or TF-IDF. Converting the selected first keywords into feature vectors in numerical form to obtain first key features. The first keywords are converted into feature vectors according to the following method: creating a vocabulary, and then generating a vector for each document, where each dimension corresponds to a word in the vocabulary, and the value represents the number of times the word appears in the document. Querying the vocabulary with the first keywords to obtain the first initial features.

[0093] Identify the accounting statements of a preset type in the historical asset data after data cleaning, extract the derived indicators representing the asset characteristics from the accounting statements, convert the numerical data corresponding to the derived indicators into feature vectors to obtain the second initial features. Select appropriate samples from the historical asset data after data cleaning according to the first initial features, the second initial features, and each risk indicator (such as default probability, expected loss rate, etc.), ensuring that each risk category has sufficient representative data for training. The specific operations include: dividing the data into multiple categories according to different risk indicators, or randomly or systematically extracting a certain number of data points from each category as training samples. Ensure that the sample size of each category is large enough to ensure that the model can learn the characteristics of that category.

[0094] Assign a label to each extracted historical sample, and the label reflects the true situation or result of the historical sample (such as whether it defaults). Through this step, a labeled training sample set is created for training the supervised learning model. Providing the corresponding training sample set for each risk indicator can ensure that the subsequent classification decision tree model training is based on high-quality data, improve the accuracy and reliability of the classification decision tree model, and thus better support the risk assessment and grouping work.

[0095] In one embodiment, constructing an initial decision tree model for each risk indicator by using a preset decision tree algorithm includes:

[0096] According to the first key feature and the second key feature corresponding to each risk indicator, select the preset Gini coefficient as the criterion for the splitting node of the preset decision tree algorithm;

[0097] Starting from the root node of the preset decision tree algorithm, for each risk indicator, calculate each splitting point using the preset Gini coefficient, and select the best splitting point for splitting;

[0098] Repeat the step of calculating each splitting point using the preset Gini coefficient and selecting the best splitting point for splitting until the stopping condition is met to obtain the initial decision tree model.

[0099] Analyze the first key feature of each risk indicator. For example, the credit risk indicator may be related to features such as the customer's credit score and repayment record. Analyze the second key feature market of each risk indicator. For example, the risk indicator may be related to features such as the volatility of the asset and interest rate sensitivity. According to these first key features and second key features, determine which feature variables have an important impact on the classification of the risk indicator.

[0100] Based on these first and second key features (feature variables) that have an important impact on the classification of risk indicators, selecting the preset Gini coefficient as the criterion for the splitting nodes of the preset decision tree algorithm is crucial for the effectiveness and performance of the initial decision tree model. Selecting the Gini coefficient as the splitting criterion means that at each split, the decision tree algorithm will attempt to find the feature and threshold that can minimize the impurity of the subset.

[0101] The construction of the decision tree starts from the root node. At this stage, all possible key features and their values are considered, and the Gini coefficient of each possible split point is calculated. Specifically: 1. Calculate the Gini coefficient: For each key feature, take its values, divide the data of the current node into two subsets, and calculate the weighted average Gini coefficient of these two subsets in terms of the values. 2. Select the best split point: Select the split point that makes the weighted average Gini coefficient the smallest as the best split point, which means this split point can best separate samples of different classes, thereby improving the purity of the child nodes.

[0102] Once the best split point is determined and the split is made, the new child nodes will become the starting point for the next round of splitting. This process will be repeated continuously until any of the following stopping conditions is met, and finally a preliminarily constructed decision tree model is obtained, ensuring that each step of splitting can maximize the purity of the nodes, thus providing an effective tool for subsequent risk assessment and asset classification.

[0103] In step S2, obtaining the key features associated with the asset grouping can remove unnecessary information, reduce the impact of noise on the classification decision tree model. The selection of key features makes the classification decision tree model more concise, reduces unnecessary computational volume, improves the processing speed, and also helps to prevent overfitting.

[0104] S3. Input the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, perform path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data, and obtain the data grouping rule of the asset data from the preset data grouping rule library according to the risk category.

[0105] In this embodiment, the first key features (such as customer age, occupation) and the second key features (such as credit score, loan amount, term, etc.) of the asset data are input into the classification decision tree models corresponding to the respective risk indicators. Each risk indicator has a specially trained decision tree model for evaluating the risk of this category.

[0106] After the first key feature and the second key feature are input, the classification decision tree model starts from the root node of the decision tree and gradually moves down according to the conditional tests on each internal node, selecting the appropriate branch to continue. When the path retrieval reaches a leaf node, the risk category represented by this leaf node is the risk category of the asset data. Each leaf node usually corresponds to a predefined risk level (such as low risk, medium risk, high risk), which reflects the assessment of the risk level of this asset based on the input key features. Search for the corresponding asset grouping rule pre-stored in the asset grouping rule library according to the determined risk category. The asset grouping rule clearly stipulates how the assets belonging to a certain risk level should be grouped. For example, all asset data determined to be of low risk may be grouped into the "high-quality asset package", while the asset data of high risk may require more stringent management and monitoring measures.

[0107] In one embodiment, the inputting the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, and performing path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data includes:

[0108] Concatenate the first key feature and the second key feature, and input the concatenated feature vector obtained into the classification decision tree models corresponding to the respective risk indicators;

[0109] According to the concatenated feature vector, perform path selection along the tree structure of the classification decision tree model until reaching the final leaf node, and obtain the risk category of the asset data for the corresponding risk indicator.

[0110] Arranging the first key feature and the second key feature in a preset order and combining them into a concatenated feature vector ensures that all relevant information is integrated together to form a complete input representation. Input the concatenated feature vector into the classification decision tree models corresponding to the respective risk indicators. The decision tree algorithm starts from the root node and gradually moves down according to the conditional tests on each internal node, selecting the appropriate branch to continue. This process is similar to answering a series of "if... then..." questions until finally reaching a leaf node. When the path retrieval reaches a leaf node, the risk category represented by this leaf node is the risk category of the asset data. Each leaf node usually corresponds to a predefined risk level (such as low risk, medium risk, high risk), which reflects the assessment of the risk level of this asset based on the input key features.

[0111] In one embodiment, the obtaining the data grouping rule of the asset data from the preset data grouping rule library according to the risk category includes:

[0112] Obtain a mapping list between preset risk indicators and data grouping rules;

[0113] According to the mapping list, obtain the data grouping rules corresponding to the category results of each risk indicator from a preset data grouping rule library, where the data grouping rules include criteria for dividing asset data into different categories or groups.

[0114] The asset grouping rule library is a pre-established database containing various different grouping rules. The mapping list is a predefined rule set or table that establishes the correspondence between the risk indicator category results and specific asset grouping rules. The mapping list is formulated by domain experts based on experience and historical data analysis to ensure that each risk category has clear grouping guiding principles.

[0115] The asset grouping rules include criteria for dividing asset data into different categories or groups. The asset grouping rules are obtained in the following way: collect all relevant information related to the asset data (financial statements, market performance, credit ratings, historical transaction records, etc.) and conduct in-depth analysis on it, and define specific grouping criteria based on the above analysis results. These criteria should be able to clearly describe the common characteristics that the assets within each group should possess.

[0116] Once the risk category of each asset data is determined, find the asset grouping rule that matches this risk category from the preset data grouping rule library to decide which specific group the asset data should be assigned to. For example, if an asset is evaluated as "high risk", then the grouping rules specifically for high-risk assets will be applied. According to the specific asset grouping rules obtained from the rule library, analyze the key characteristics of the asset data (such as credit score, debt-to-income ratio, years of work experience, etc.) and divide it into the corresponding group accordingly. This step ensures that all assets can be correctly organized according to the established criteria, thus supporting subsequent risk management, investment decisions, or other business activities.

[0117] In step S3, through the analysis of the key characteristics by the classification decision tree model, the risk category of the asset data is obtained, and according to the risk category, the asset grouping rules of the asset data are obtained, which not only improves the efficiency and accuracy of asset management, but also reduces the uncertainty brought by human intervention.

[0118] S4. Use the data grouping rules to match the first key characteristic and the second key characteristic to obtain the data group to which the asset data belongs, store the asset data in the data group to obtain the grouping result of the asset data, and feedback the grouping result to the terminal corresponding to the grouping request.

[0119] In this embodiment, the first key feature and the second key feature are matched one by one with the conditions in the selected asset grouping rule. For example, if the asset grouping rule stipulates that "for loan applicants with a high risk level, their credit score must be lower than a certain threshold", then it is necessary to check whether the credit score of each loan applicant meets this condition. Alternatively, a grouping rule may involve multiple feature conditions. Therefore, it is necessary to comprehensively evaluate all relevant feature vectors to determine whether they jointly meet all the requirements of the rule. For example, in addition to the credit score, factors such as income level and employment stability may also need to be considered.

[0120] In one embodiment, the step of using the data grouping rule to match the first key feature and the second key feature to obtain the data group to which the asset data belongs includes:

[0121] Substitute the first key feature and the second key feature into the data grouping rule for matching in sequence according to a preset order;

[0122] When the first key feature and the second key feature meet the preset criteria of the data grouping rule, obtain the data group to which the asset data belongs.

[0123] Substitute the first and second key features into the grouping rule for matching in sequence according to a certain order (which can be sorted according to the priority of the rule or the risk level). Taking a housing loan application as an example, assume that the first key feature of an applicant is a monthly income of 8,000 yuan and a good credit history (no overdue), and the second key feature is a loan amount of 800,000 yuan and a term of 20 years.

[0124] First, look at the first grouping rule. If the rule is that a monthly income higher than 7,000 yuan and a loan amount lower than 700,000 yuan belong to the low-risk group, since the loan amount of this applicant does not meet the requirement, it does not match this rule.

[0125] Only when the first and second key features meet all the grouping rules can it be determined that the asset data belongs to this group. Continuing with the above example, then look at another rule: a monthly income higher than 6,000 yuan, a good credit history, and a loan amount between 700,000 and 1,000,000 yuan belong to the medium-risk group. This applicant meets all the conditions, so it is determined that the asset data belongs to the medium-risk group.

[0126] Obtain the data group to which the asset data belongs, store the asset data in the data group, obtain the grouping result of the asset data, and feedback the grouping result to the terminal corresponding to the grouping request. Users can view the grouping result in real time, perform further operations or make decisions, so as to better meet the actual needs of financial institutions in asset management and risk control. For example, when business personnel of a financial institution initiate an asset data grouping request through an online loan application platform, after the grouping is completed, the grouping result is displayed on the operation interface of the business personnel. At the same time, the system also provides a data export function, and business personnel can export the grouping result as an Excel file for further risk assessment or report generation. In addition, the system will record the grouping result in the historical database for subsequent query and analysis.

[0127] In another embodiment, setting a change rule for the grouped asset data can identify and report changes in the grouped asset data that do not meet expectations or the normal pattern under preset conditions, specifically including:

[0128] 1. Set the change rule: Clearly define when the asset data is considered "abnormal". The change rule can be based on multiple conditions, such as: changes in risk level, abnormal financial performance, changes in market environment, etc.

[0129] 2. Configure the trigger conditions: Configure specific trigger conditions for each change rule of the asset data, including thresholds (such as a credit score decrease exceeding 20%) and time windows (such as multiple overdue occurrences within 3 consecutive months).

[0130] 3. Real-time monitoring and evaluation: Continuously track the performance of the asset data and regularly evaluate whether the asset data meets the preset change rules. For newly input data, the system will check for any abnormal situations while performing regular classification. When a change is detected, an instant notification can be sent to relevant personnel via email, text message, or the internal messaging system to ensure timely response.

[0131] Finally, the asset data grouping in the financial scenario is used as an example to illustrate the present invention, and no limitation is made on the application scenario of the present invention:

[0132] An online loan application platform receives a batch of new loan applications. Each application has a unique identifier (such as an application number). According to the unique identifier of the loan application, multiple risk indicators of the corresponding loan application are retrieved from the database, as well as the mapping relationship between the preset risk indicators and the loan application. The obtained asset data is cleaned to ensure data quality.

[0133] After data cleaning, tokenize the credit report text in the loan application to obtain multiple tokens. Select the first keywords highly correlated with the credit status (such as "good credit", "repay on time") from these tokens, and convert these keywords into the first key features. Identify and extract the derivative indicators in the balance sheet (such as current ratio, return on net assets, etc.). Convert the numerical data corresponding to these derivative indicators into the second key features.

[0134] Obtain the classification decision tree models corresponding to each risk indicator related to this application from the preset model library. These models are pre-trained based on historical loan data and their corresponding risk categories. Concatenate the above-extracted first key features and second key features into a complete feature vector, and input it into the classification decision tree models corresponding to each risk indicator. Select the path according to the feature vector along the tree structure of the decision tree model until reaching the final leaf node, so as to determine the risk category of each loan (such as high risk, medium risk, low risk).

[0135] According to the determined risk category, select the corresponding asset grouping rules from the preset data grouping rule library. For example, for "high-risk" loan applicants, there may be more stringent approval processes and higher interest rate requirements. Use the selected asset grouping rules to match the first key features and second key features. For example, check whether the applicant's credit score is lower than a certain threshold, whether the annual income is within a certain range, etc.

[0136] Once all conditions are met or not met, the specific group to which the asset data should be assigned can be obtained according to the asset grouping rules. For example, if a loan meets all the conditions of the "medium risk" category, then it will be classified into a subgroup such as "stable occupation, medium income group".

[0137] In steps S1 - S4, the present invention performs data cleaning and feature extraction on the asset data, analyzes the key features of the asset data using a classification decision tree, obtains the risk category of the asset data, obtains the asset grouping rules of the asset data according to the risk category, and performs a grouping operation on the asset data to obtain the data group to which the asset data belongs, realizing the full process automation from data cleaning, feature extraction to classification, reducing manual intervention, improving the processing speed and accuracy of asset data grouping, comprehensively considering multiple risk indicators, constructing a dedicated classification decision tree model for each risk indicator, and ensuring comprehensive and accurate risk identification.

[0138] As Figure 2 shown, it is a schematic diagram of the modules of an asset data grouping device provided by an embodiment of the present invention.

[0139] The asset data grouping device 100 according to the present invention can be installed in an electronic device. According to the implemented functions, the asset data grouping device 100 may include a receiving module 110, an extraction module 120, a retrieval module 130, and a grouping module 140. The modules in the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0140] In this embodiment, the functions of each module / unit are as follows:

[0141] The receiving module 110 is configured to receive a grouping request for grouping asset data to be grouped, obtain an identifier of the asset data from the grouping request, obtain a plurality of risk indicators corresponding to the asset data from a preset database according to the identifier, and a mapping relationship between the preset risk indicators and a classification decision tree model;

[0142] The extraction module 120 is configured to perform data cleaning on the asset data, extract a first key feature associated with customer attributes and a second key feature associated with asset value attributes from the asset data after data cleaning, and obtain a classification decision tree model corresponding to each risk indicator of the asset data from a preset model library according to the plurality of risk indicators corresponding to the asset data and the mapping relationship;

[0143] The retrieval module 130 is configured to input the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, perform path retrieval along the tree structure of the classification decision tree models to obtain a risk category of the asset data, and obtain a data grouping rule of the asset data from a preset data grouping rule library according to the risk category;

[0144] The grouping module 140 is configured to match the first key feature and the second key feature by using the data grouping rule to obtain a data group to which the asset data belongs, store the asset data in the data group to obtain a grouping result of the asset data, and feed back the grouping result to a terminal corresponding to the grouping request.

[0145] In one embodiment, the extraction of the first key feature associated with customer attributes includes:

[0146] Performing word segmentation processing on the asset data after data cleaning to obtain a plurality of word segments;

[0147] Selecting a first keyword highly associated with customer attributes from the plurality of word segments;

[0148] Converting the first keyword into a feature vector to obtain the first key feature.

[0149] In one embodiment, the extraction of the second key features associated with the asset value attributes includes:

[0150] Identifying the financial statements of a preset type in the asset data after data cleaning, and extracting the derivative indicators representing the asset characteristics from the financial statements;

[0151] Converting the numerical data corresponding to the derivative indicators into feature vectors to obtain the second key features.

[0152] In one embodiment, the classification decision tree model is obtained according to the following method:

[0153] Obtaining historical asset data and multiple risk indicators corresponding to the historical asset data from a preset historical asset database;

[0154] Constructing a training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators;

[0155] Dividing the training sample set corresponding to each risk indicator into a training set and a test set according to a preset ratio;

[0156] Using a preset decision tree algorithm to construct an initial decision tree model for each risk indicator respectively;

[0157] Training the corresponding initial decision tree model using the training set corresponding to each risk indicator, and testing the trained initial decision tree model using the test set corresponding to each risk indicator. Stop training until the evaluation index of the initial decision tree model meets the preset standard, and obtain the classification decision tree model corresponding to each risk indicator.

[0158] In one embodiment, the construction of the training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators includes:

[0159] Performing data cleaning on the historical asset data, performing word segmentation processing on the historical asset data after data cleaning to obtain multiple word segments, selecting the first keywords highly associated with the customer attributes from the multiple word segments, and converting the first keywords into feature vectors to obtain the first initial features;

[0160] Identifying the financial statements of a preset type in the historical asset data after data cleaning, extracting the derivative indicators representing the asset characteristics from the financial statements, and converting the numerical data corresponding to the derivative indicators into feature vectors to obtain the second initial features;

[0161] Extracting historical samples from the historical asset data after data cleaning according to each risk indicator to obtain historical samples corresponding to each risk indicator category;

[0162] Construct the labels of the historical samples corresponding to each risk indicator based on the first initial feature, the second initial feature, and the corresponding risk indicators, and obtain the training sample set corresponding to each risk indicator.

[0163] In one embodiment, constructing an initial decision tree model for each risk indicator by using a preset decision tree algorithm includes:

[0164] According to the first key feature and the second key feature corresponding to each risk indicator, select the preset Gini coefficient as the criterion for the splitting node of the preset decision tree algorithm;

[0165] Starting from the root node of the preset decision tree algorithm, for each risk indicator, calculate each splitting point using the preset Gini coefficient, and select the best splitting point for splitting;

[0166] Repeat the step of calculating each splitting point using the preset Gini coefficient and selecting the best splitting point for splitting until the stop condition is met to obtain the initial decision tree model.

[0167] In one embodiment, inputting the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, and performing path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data, includes:

[0168] Concatenate the first key feature and the second key feature, and input the concatenated feature vector obtained by concatenation into the classification decision tree models corresponding to the respective risk indicators;

[0169] According to the concatenated feature vector, perform path selection along the tree structure of the classification decision tree model until reaching the final leaf node, and obtain the risk category of the asset data for the corresponding risk indicator.

[0170] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the asset data grouping method provided by an embodiment of the present invention.

[0171] In this embodiment, the electronic device 1 includes, but is not limited to, a memory 11, a processor 12, and a network interface 13 that can communicate with each other through a system bus. The memory 11 stores an asset data grouping program 10, and the asset data grouping program 10 can be executed by the processor 12. Figure 3 Only the electronic device 1 with components 11 - 13 and the asset data grouping program 10 is shown. Those skilled in the art can understand that Figure 3The structure shown does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown in the figure, or combine certain components, or have a different component arrangement.

[0172] Among them, the memory 11 includes a memory and at least one type of readable storage medium. The memory provides a cache for the operation of the electronic device 1; the readable storage medium can be a non-volatile storage medium such as a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the readable storage medium may be an internal storage unit of the electronic device 1; in other embodiments, the non-volatile storage medium may also be an external storage device of the electronic device 1, such as a plug-in hard disk equipped on the electronic device 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. In this embodiment, the readable storage medium of the memory 11 is generally used to store the operating system and various application software installed on the electronic device 1, such as storing the code of the asset data grouping program 10 in an embodiment of the present invention, etc. In addition, the memory 11 can also be used to temporarily store various data that have been output or will be output.

[0173] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with other devices, etc. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the asset data grouping program 10, etc.

[0174] The network interface 13 may include a wireless network interface or a wired network interface, and this network interface 13 is used to establish a communication connection between the electronic device 1 and a terminal (not shown in the figure).

[0175] Optionally, the electronic device 1 may further include a user interface, which may include a display, an input unit such as a keyboard, and optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0176] It should be understood that the above embodiments are only for illustrative purposes and are not limited by this structure in the scope of the patent application.

[0177] The asset data grouping program 10 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 12, it can implement:

[0178] Receiving a grouping request for grouping the asset data to be grouped, obtaining the identifier of the asset data from the grouping request, obtaining a plurality of risk indicators corresponding to the asset data from a preset database according to the identifier, and a mapping relationship between the preset risk indicators and the classification decision tree model;

[0179] Performing data cleaning on the asset data, extracting a first key feature associated with customer attributes and a second key feature associated with asset value attributes from the asset data after data cleaning, and obtaining a classification decision tree model corresponding to each risk indicator of the asset data from a preset model library according to the plurality of risk indicators corresponding to the asset data and the mapping relationship;

[0180] Inputting the first key feature and the second key feature into the classification decision tree models corresponding to the respective risk indicators, performing path retrieval along the tree structure of the classification decision tree models to obtain the risk category of the asset data, and obtaining the data grouping rule of the asset data from a preset data grouping rule library according to the risk category;

[0181] Using the data grouping rule to match the first key feature and the second key feature to obtain the data group to which the asset data belongs, storing the asset data in the data group to obtain the grouping result of the asset data, and feeding back the grouping result to the terminal corresponding to the grouping request.

[0182] Specifically, for the specific implementation method of the above asset data grouping program 10 by the processor 12, reference can be made to Figure 1Descriptions of relevant steps in corresponding embodiments are not elaborated herein.

[0183] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or non-volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory).

[0184] An asset data grouping program 10 is stored on the computer-readable storage medium. The asset data grouping program 10 can be executed by one or more processors. The specific implementation manner of the computer-readable storage medium of the present invention is basically the same as that of the above-described embodiments of the asset data grouping method, and will not be elaborated herein.

[0185] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0186] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] In addition, in each embodiment of the present invention, the functional modules can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0188] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0189] Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0190] In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices recited in the system claims can also be implemented by one element or device through software or hardware. The terms such as "second" are used to denote names and do not denote any particular order.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for grouping asset data, characterized in that: The method comprises: receiving a grouping request for grouping asset data to be grouped, obtaining an identifier of the asset data from the grouping request, and obtaining a plurality of risk indicators corresponding to the asset data from a preset database according to the identifier, as well as a mapping relationship between the preset risk indicators and a classification decision tree model; Performing data cleaning on the asset data, extracting a first key feature associated with the customer attribute and a second key feature associated with the asset value attribute from the cleaned asset data, and acquiring a classification decision tree model corresponding to each risk indicator of the asset data from a preset model library according to a plurality of risk indicators corresponding to the asset data and the mapping relationship; Inputting the first key feature and the second key feature into the classification decision tree model corresponding to each risk indicator, performing path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data, and obtaining the data grouping rule of the asset data from a preset data grouping rule library according to the risk category; The first key feature and the second key feature are matched using the data grouping rule to obtain the data group to which the asset data belongs, the asset data is stored in the data group, the grouping result of the asset data is obtained, and the grouping result is fed back to the terminal corresponding to the grouping request.

2. The asset data grouping method according to claim 1, characterized in that: The extracting the first key feature associated with the customer attribute includes: Perform word segmentation on the asset data after data cleaning to obtain multiple word segments; Select the first keyword that is highly related to the customer attributes from multiple segmented words; The first keyword is converted into a feature vector to obtain the first key feature.

3. The asset data grouping method according to claim 1, characterized in that: The extracting of the second key feature associated with the asset value attribute includes: Identify account statements of a preset type in the asset data after data cleaning, and extract derived indicators representing asset characteristics from the account statements; The numerical data corresponding to the derived indicator is converted into a feature vector to obtain the second key feature.

4. The asset data grouping method according to claim 1, characterized in that: The classification decision tree model is obtained according to the following method: Acquire historical asset data and a plurality of risk indicators corresponding to the historical asset data from a preset historical asset database; Constructing a training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators; The training sample set corresponding to each risk indicator is divided into a training set and a test set according to a preset ratio; Use the preset decision tree algorithm to construct an initial decision tree model for each risk indicator; The training set corresponding to each risk indicator is used to train the corresponding initial decision tree model, and the test set corresponding to each risk indicator is used to test the trained initial decision tree model until the evaluation index of the initial decision tree model meets the preset standard. The training is stopped to obtain the classification decision tree model corresponding to each risk indicator.

5. The asset data grouping method according to claim 4, characterized in that: The step of constructing a training sample set corresponding to each risk indicator according to the historical asset data and the multiple risk indicators includes: Cleaning the historical asset data, performing word segmentation processing on the cleaned historical asset data to obtain multiple word segments, selecting a first keyword that is highly associated with the customer attribute from the multiple word segments, converting the first keyword into a feature vector, and obtaining a first initial feature; Identify account statements of a preset type in the historical asset data after data cleaning, extract derived indicators representing asset characteristics from the account statements, convert numerical data corresponding to the derived indicators into feature vectors, and obtain second initial features; According to each risk indicator, historical samples are extracted from the historical asset data after data cleaning to obtain historical samples of the corresponding categories of each risk indicator; A label of a historical sample corresponding to each risk indicator is constructed according to the first initial feature, the second initial feature and the corresponding risk indicator, so as to obtain a training sample set corresponding to each risk indicator.

6. The asset data grouping method according to claim 4, characterized in that: The method of using a preset decision tree algorithm to construct an initial decision tree model for each risk indicator includes: According to the first key feature and the second key feature corresponding to each risk indicator, a preset Gini coefficient is selected as a criterion for splitting nodes of the preset decision tree algorithm; Starting from the root node of the preset decision tree algorithm, for each risk indicator, each split point is calculated using the preset Gini coefficient, and the best split point is selected for splitting; Repeat the steps of using the preset Gini coefficient to calculate each split point and selecting the best split point for splitting until the stopping condition is met to obtain an initial decision tree model.

7. The asset data grouping method according to claim 1, characterized in that: The step of inputting the first key feature and the second key feature into the classification decision tree model corresponding to each risk indicator, performing path retrieval along the tree structure of the classification decision tree model, and obtaining the risk category of the asset data includes: splicing the first key feature and the second key feature, and inputting the spliced ​​feature vector obtained by splicing into the classification decision tree model corresponding to each risk indicator; According to the concatenated feature vector, a path is selected along the tree structure of the classification decision tree model until the final leaf node is reached, thereby obtaining the risk category of the asset data in the corresponding risk indicator.

8. An asset data grouping device, characterized in that: The device comprises: A receiving module, configured to receive a grouping request for grouping asset data to be grouped, obtain an identifier of the asset data from the grouping request, and obtain a plurality of risk indicators corresponding to the asset data from a preset database according to the identifier, as well as a mapping relationship between the preset risk indicators and the classification decision tree model; An extraction module is used to perform data cleaning on the asset data, extract a first key feature associated with the customer attribute and a second key feature associated with the asset value attribute from the cleaned asset data, and obtain a classification decision tree model corresponding to each risk indicator of the asset data from a preset model library according to a plurality of risk indicators corresponding to the asset data and the mapping relationship; A retrieval module, used to input the first key feature and the second key feature into the classification decision tree model corresponding to each risk indicator, perform path retrieval along the tree structure of the classification decision tree model to obtain the risk category of the asset data, and obtain the data grouping rule of the asset data from a preset data grouping rule library according to the risk category; A grouping module is used to match the first key feature and the second key feature using the data grouping rule to obtain the data group to which the asset data belongs, store the asset data in the data group, obtain the grouping result of the asset data, and feed back the grouping result to the terminal corresponding to the grouping request.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores an asset data grouping program that can be executed by the at least one processor, and the asset data grouping program is executed by the at least one processor to enable the at least one processor to perform the asset data grouping method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: An asset data grouping program is stored on the computer-readable storage medium, and the asset data grouping program can be executed by one or more processors to implement the asset data grouping method as described in any one of claims 1 to 7.