Risk classification method and device

By constructing the TF-IDF matrix of the abnormal resource sample set and generating the number of risk topics, the problem of inaccurate risk classification of transaction resource information in the existing technology is solved, more efficient risk identification and management is achieved, and the stability of the financial market is ensured.

CN120147007APending Publication Date: 2025-06-13CHINA ELECTRONICS JINXIN DIGITAL TECH GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217099.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The difficulty in accurately identifying and classifying risk categories of transaction resource information in the prior art has led to insufficient financial institutions in risk management and market stability.

Method used

By obtaining the sample risk characteristics of the abnormal resource sample set, the first TF-IDF matrix is ​​constructed, and a second risk classification system is generated based on the number of risk topics. Finally, the target risk classification system is generated based on the two jointly, and accurate risk classification of transaction resource information is achieved.

Benefits of technology

It improves the accuracy and comprehensiveness of risk classification, enhances the ability of financial institutions to identify and manage risks, and ensures the stable operation of the financial market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147007A_ABST
    Figure CN120147007A_ABST
Patent Text Reader

Abstract

The invention provides a risk classification method and device. The method comprises the steps of obtaining a first risk classification system corresponding to an abnormal resource sample set; according to each segmented word in each abnormal resource sample, constructing a first word frequency-inverse document frequency TF-IDF matrix corresponding to the abnormal resource sample set; generating a second risk classification system according to the number interval corresponding to the risk theme and the first TF-IDF matrix; generating a target risk classification system according to the first risk classification system and the second risk classification system; wherein the target risk classification system is used for carrying out risk classification on the transaction resource information, so that the accuracy and comprehensiveness of the target risk classification system are improved, risk classification is carried out on the transaction resource information based on the target risk classification system, the risk classification accuracy is improved, and the risk classification efficiency is improved. Therefore, a financial institution can better identify and manage risks, and stable operation of a financial market is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of risk control technologies, and in particular, to a risk classification method and apparatus. Background Art

[0002] With the continuous expansion of the scale of financial transactions and the increasing diversification of transaction services and transaction methods, financial institutions need to deeply mine and analyze a large amount of financial transaction resource information to accurately identify the risk categories to which the transaction resource information belongs, and take corresponding prevention and control measures for the transaction resource information with higher risk categories, which not only helps to improve the risk management level of financial institutions, but also provides a strong guarantee for the stable and healthy development of the financial market; therefore, how to classify the risk of transaction resource information is very important. Summary of the Invention

[0003] The present disclosure provides a risk classification method and apparatus to at least partly solve one of the technical problems in the related art. The technical solution of the present disclosure is as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a risk classification method is provided, including: obtaining a first risk classification system corresponding to an abnormal resource sample set; wherein, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set; constructing a first term frequency-inverse document frequency TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample; wherein, the element in the i-th row and the j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample; generating a second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix; generating a target risk classification system according to the first risk classification system and the second risk classification system; wherein, the target risk classification system is used to classify the risk of transaction resource information.

[0005] According to a second aspect of the embodiments of the present disclosure, a risk classification device is provided, including: an acquisition module, configured to acquire a first risk classification system corresponding to a set of abnormal resource samples; wherein, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the set of abnormal resource samples; a construction module, configured to construct a first term frequency-inverse document frequency (TF-IDF) matrix corresponding to the set of abnormal resource samples according to each word segment in each of the abnormal resource samples; wherein, the element in the i-th row and j-th column of the TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample; a first generation module, configured to generate a second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix; a second generation module, configured to generate a target risk classification system according to the first risk classification system and the second risk classification system; wherein, the target risk classification system is used to classify the risk of transaction resource information.

[0006] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the risk classification method as described in the embodiments of the first aspect of the present disclosure.

[0007] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the risk classification method as described in the embodiments of the first aspect of the present disclosure.

[0008] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including: a computer program, which when executed by a processor, implements the risk classification method as described in the embodiments of the first aspect of the present disclosure.

[0009] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0010] In this technical solution, by obtaining the first risk classification system generated based on the sample risk characteristics in each abnormal resource sample, the core characteristics of risks in actual business are accurately captured, ensuring that the classification results are close to the actual business requirements. Furthermore, the first TF-IDF matrix of the abnormal resource sample set is constructed, and the second risk classification system is generated by combining the quantity interval corresponding to the risk theme and the first TF-IDF matrix, realizing the numerical processing of abnormal resource samples, quantifying the key characteristics in abnormal resource samples and highlighting the important factors related to risks, providing a scientific basis for generating the second risk classification system, and enhancing the objectivity and comprehensiveness of risk classification. Finally, the target risk classification system is jointly generated based on the first risk classification system and the second risk classification system, improving the accuracy and comprehensiveness of the target risk classification system. Thus, the risk classification of transaction resource information is carried out based on the target risk classification system, improving the accuracy of risk classification, which helps financial institutions better identify and manage risks and ensure the stable operation of the financial market. Among them, when generating the second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix, the number of risk themes is determined based on the quantity interval and the second TF-IDF matrix (the TF-IDF matrix after dimensionality reduction), reducing the computational complexity, ensuring the rationality and applicability of the number of risk themes, and avoiding the deviation caused by subjective setting. Furthermore, a probability generation model is used to combine the target quantity and the first TF-IDF matrix to predict the topic probability distribution and the topic word probability distribution of risk themes. By analyzing the topic word probability distribution, multiple risk topic words under each risk theme are determined, and thus the second risk classification system is generated based on the risk topic words, enhancing the interpretability of risk classification and at the same time improving the accuracy and comprehensiveness of the second risk classification system.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0012] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0013] Figure 1 is a schematic flowchart of the risk classification method shown in the first embodiment of the present disclosure;

[0014] Figure 2 is a schematic flowchart of the risk classification method shown in the second embodiment of the present disclosure;

[0015] Figure 3 is a schematic flowchart of the risk classification method shown in the third embodiment of the present disclosure;

[0016] Figure 4 It is a schematic flowchart of the risk classification method shown in the fourth embodiment of the present disclosure;

[0017] Figure 5 It is a schematic diagram of the principle of the risk classification method shown in the embodiments of the present disclosure;

[0018] Figure 6 It is a schematic structural diagram of the risk classification device shown in the fifth embodiment of the present disclosure;

[0019] Figure 7 It is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0020] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0022] It should be noted that in the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processing are all carried out on the premise of obtaining the consent of the user, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0023] With the continuous expansion of the scale of financial transactions and the increasing diversification of transaction operations and transaction methods, financial institutions need to deeply mine and analyze a large amount of financial transaction resource information to accurately identify the risk categories to which the transaction resource information belongs. However, due to the lack of a unified risk classification standard, it is impossible to accurately identify the risk classification of these transaction resource information.

[0024] Currently, the risk classification system for classifying the risk of transaction resource information is mainly constructed based on the following two aspects:

[0025] 1. Manually construct a risk classification system only based on expert experience;

[0026] 2. Construct a risk classification system only based on big data analysis;

[0027] Among them, when manually constructing a risk classification system based on expert experience, due to the multi-dimensionality and complexity of transaction resource information, and the cognitive limitations or restricted experience scope of relevant experts, the constructed risk classification system may be difficult to comprehensively cover all relevant risk factors. Moreover, when processing large-scale data manually, the efficiency is low and it relies too much on subjective judgment, which may affect the consistency and objectivity of the results. Constructing a risk classification system based on big data analysis depends on a large amount of data input and complex algorithm models, which may lead to the uncertainty and unverifiability of the results.

[0028] In view of the above problems, the present disclosure proposes a risk classification method and device.

[0029] The following describes the risk classification method and device of the embodiments of the present disclosure with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flowchart of the risk classification method shown in the first embodiment of the present disclosure. Among them, it should be noted that the execution subject of the embodiments of the present disclosure can be a risk classification device, and this risk classification device can be applied to any electronic device with computing capabilities, so that the electronic device can perform the risk classification function.

[0031] As Figure 1 shown, the risk classification method includes the following steps:

[0032] Step 101, obtain a first risk classification system corresponding to the abnormal resource sample set.

[0033] Among them, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set.

[0034] In the embodiments of the present disclosure, obtain a risk classification framework generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set, that is, obtain the first risk classification system. Among them, the abnormal resource sample set is a set including multiple abnormal resource samples, and the abnormal resource sample is transaction information (such as transaction text) generated by a transaction behavior that does not conform to public order and good customs. For example, the abnormal resource sample is an illegal financing sample, and the sample risk characteristics are such as: transaction amount, transaction frequency, transaction counterparty information, geographical location, etc.

[0035] Among them, the first risk classification system can be generated by the following steps:

[0036] 1. Obtain reference risk information associated with each abnormal resource sample;

[0037] In the embodiments of the present disclosure, risk information related to each abnormal resource sample is collected, and the risk information related to each abnormal resource sample is referred to as reference risk information. Among them, the reference risk information may include, but is not limited to: transaction subject information related to the corresponding abnormal resource sample, historical transaction amount, user behavior pattern, historical transaction frequency, etc.

[0038] 2. Based on the reference risk information, feature extraction is performed on each abnormal resource sample to obtain the sample risk features of each abnormal resource sample;

[0039] Further, guided by the reference risk information, features highly correlated with the risk attributes are extracted from each abnormal resource sample, and this feature is used as the sample risk feature of the corresponding abnormal resource sample.

[0040] 3. According to the sample risk features of each abnormal resource sample, an initial risk classification system is generated;

[0041] Furthermore, based on the extracted sample risk features, a preliminary risk classification framework is constructed. For example, samples with similar risk features are grouped into one category, and according to the distribution of the sample risk features, different risk categories in the initial risk classification system are defined.

[0042] 4. According to the initial risk classification system, a first risk classification system is generated.

[0043] Finally, since the initial risk classification system may be redundant or insufficient, in order to improve the accuracy of the first risk classification system, based on the initial risk classification system, combined with business requirements, expert experience or further optimized algorithms, the final first risk classification system is generated.

[0044] Step 102, according to each word segment in each abnormal resource sample, construct a first term frequency-inverse document frequency (TF-IDF) matrix corresponding to the abnormal resource sample set.

[0045] Among them, the element in the i-th row and j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample.

[0046] As a possible implementation manner, in order to further improve the accuracy and applicability of risk classification, a TF-IDF matrix (i.e., the first TF-IDF matrix) is constructed based on each word segment in the unstructured abnormal resource sample to quantify the sample risk features in the abnormal resource sample while highlighting the key risk factors, thereby providing a more scientific data basis for subsequent risk analysis and classification; among them, the element in the i-th row and j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample.

[0047] Step 103: Generate a second risk classification system according to the quantity range corresponding to the risk theme and the first TF-IDF matrix.

[0048] To ensure that the risk classification result is close to the actual business requirements, as a possible implementation, a second risk classification system is generated based on the quantity range of the risk theme set according to the actual requirements and the first TF-IDF matrix.

[0049] As an example, obtain the quantity range corresponding to the risk theme, where the quantity range is used to indicate the expected number range of the risk themes of the to-be-generated second risk classification system. For example, the quantity range corresponding to the risk theme is [15, 20]. Then, based on the quantity range and the TF-IDF matrix, jointly generate the second risk classification system.

[0050] Step 104: Generate a target risk classification system according to the first risk classification system and the second risk classification system.

[0051] Among them, the target risk classification system is used to classify the risk of transaction resource information.

[0052] To improve the accuracy of the risk classification system, in the embodiments of the present disclosure, the target risk classification system is jointly generated by combining the first risk classification system and the second risk classification system.

[0053] The risk classification method of the embodiments of the present disclosure realizes the accurate capture of the core features of risks in actual business by obtaining the first risk classification system generated based on the sample risk features in each abnormal resource sample, ensuring that the classification result is close to the actual business requirements. Furthermore, by constructing the first TF-IDF matrix of the abnormal resource sample set and combining the quantity range corresponding to the risk theme and the first TF-IDF matrix to generate the second risk classification system, the numerical processing of the abnormal resource sample is realized, the key features in the abnormal resource sample are quantified, and the important factors related to risks are highlighted, providing a scientific basis for generating the second risk classification system and enhancing the objectivity and comprehensiveness of risk classification. Finally, the target risk classification system is jointly generated based on the first risk classification system and the second risk classification system, improving the accuracy and comprehensiveness of the target risk classification system. Therefore, the risk of transaction resource information is classified based on the target risk classification system, improving the accuracy of the risk classification of transaction resource information, which helps financial institutions better identify and manage risks and ensure the stable operation of the financial market.

[0054] To clearly illustrate how the second risk classification system is generated according to the quantity range corresponding to the risk theme and the TF-IDF matrix in the above embodiments, the present disclosure proposes another risk classification method.

[0055] Figure 2It is a schematic flowchart of the risk classification method shown in the second embodiment of the present disclosure.

[0056] As Figure 2 shown, the risk classification method includes the following steps:

[0057] Step 201, obtain the first risk classification system corresponding to the abnormal resource sample set.

[0058] Among them, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set.

[0059] Step 202, construct the first TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample.

[0060] Among them, the element in the i-th row and j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample.

[0061] In order to accurately construct the TF-IDF matrix corresponding to the abnormal resource sample set, as a possible implementation, perform word segmentation processing on each abnormal resource sample to obtain multiple word segments of each abnormal resource sample; among them, it should be noted that when the multiple word segments of each abnormal resource sample include stop words (such as "de", "le", "ne", "you", etc.), the stop words in the multiple word segments of each abnormal resource sample need to be removed.

[0062] Furthermore, for any abnormal resource sample, count the number of occurrences of each word segment in any abnormal resource sample in any abnormal resource sample to obtain the word frequency statistical result of each word segment in any abnormal resource sample; according to the word frequency statistical result of each word segment in any abnormal resource sample, calculate the TF value of each word segment in any abnormal resource sample; according to the document frequency of each word segment in the abnormal resource sample set in any abnormal resource sample, calculate the IDF value of each word segment in any abnormal resource sample; generate the first TF-IDF matrix corresponding to the abnormal resource sample set according to the TF value and IDF value of any word segment in each abnormal resource sample.

[0063] For example, for any abnormal resource sample, count the occurrence times of each word segment in it. Taking an abnormal resource sample "Large amount transfer to unknown account" as an example, the word segment results are: ["Large amount", "Transfer", "To", "Unknown", "Account"], a total of 5 word segments. The word frequency statistics results are: {"Large amount": 1, "Transfer": 1, "To": 1, "Unknown": 1, "Account": 1}. According to the word frequency statistics results, calculate the TF value of each word segment in this sample. The TF value of "Large amount" = 1 / 5 = 0.2, the TF value of "Transfer" = 1 / 5 = 0.2, the TF value of "To" = 1 / 5 = 0.2, the TF value of "Unknown" = 1 / 5 = 0.2, the TF value of "Account" = 1 / 5 = 0.2. Furthermore, according to the distribution of each word segment in the entire abnormal resource sample set, calculate the IDF value of each word segment. Assume that there are 10 samples in the abnormal resource sample set, and "Large amount" appears in 3 abnormal resource samples, then the Further, multiply the TF value of each word segment by the IDF value to obtain the TF-IDF value of this word segment. For example, the TF-IDF value of "Large amount" = 0.2 × log(2.5); finally, integrate the TF-IDF values of each word segment in all abnormal resource samples into a TF-IDF matrix.

[0064] Step 203, obtain the second TF-IDF matrix.

[0065] Among them, the second TF-IDF matrix is obtained by performing dimensionality reduction processing on the first TF-IDF matrix.

[0066] In order to reduce the computational complexity, in the embodiments of the present disclosure, a specified dimensionality reduction algorithm is used to perform dimensionality reduction processing on the first TF-IDF matrix to obtain the dimensionality-reduced TF-IDF matrix, that is, the second TF-IDF matrix is obtained.

[0067] For example, use t-SNE to perform dimensionality reduction on the first TF-IDF matrix. Assume that the first TF-IDF matrix is 1000×10000 (1000 samples, 10000 word segments). After t-SNE dimensionality reduction, a second TF-IDF matrix of 1000×2 may be obtained.

[0068] Step 204, determine the target quantity according to the quantity interval and the second TF-IDF matrix.

[0069] Among them, the target quantity is used to indicate the number of risk themes of the second risk classification system to be generated.

[0070] In order to improve the flexibility and accuracy of determining the number of risk topics while meeting the actual business requirements, in the embodiments of the present disclosure, within the preset number range corresponding to the risk topics, the dimension-reduced TF-IDF matrix is used as input data, and a cluster set including different numbers of risk topics is generated through a clustering algorithm. Furthermore, by combining evaluation metrics (such as the silhouette coefficient, Calinski-Harabasz index, etc.) or through quantitative evaluation of each cluster set by relevant experts, the number of clusters in the optimal cluster set is finally selected as the target number.

[0071] Step 205: Generate a second risk classification system according to the target number and the first TF-IDF matrix.

[0072] In order to improve the rationality and comprehensiveness of the risk classification system, in the embodiments of the present disclosure, a probability generation model (such as the Latent Dirichlet Allocation (LDA) topic model) is used to perform probability prediction on the first TF-IDF matrix in combination with the target number, obtaining the topic probability distribution of the risk topics including the target number and the topic word probability distribution under each topic. Thus, based on the topic probability distribution of the risk topics and the topic word probability distribution under each topic, a second risk classification system is generated.

[0073] Step 206: Generate a target risk classification system according to the first risk classification system and the second risk classification system.

[0074] Among them, the target risk classification system is used to classify the risk of transaction resource information.

[0075] It should be noted that the execution processes of step 201 and step 206 can be implemented in any way in the respective embodiments of the present disclosure. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.

[0076] In summary, by constructing the first TF-IDF matrix, the text features in the abnormal resource samples are quantified, and the keywords related to risks are highlighted, thus providing a scientific basis for subsequent analysis. Furthermore, based on the number range and the dimension-reduced TF-IDF matrix, the target number is determined, reducing the computational complexity, ensuring the rationality and applicability of the number of risk topics, and avoiding the deviation caused by subjective setting. Finally, a second risk classification system is generated according to the target number and the TF-IDF matrix, improving the comprehensiveness and accuracy of the second risk classification system.

[0077] In order to clearly illustrate how to generate the second risk classification system according to the target number and the TF-IDF matrix in the above embodiments, the present disclosure proposes another risk classification method.

[0078] Figure 3It is a schematic flowchart of the risk classification method shown in the third embodiment of the present disclosure.

[0079] As Figure 3 shown, the risk classification method includes the following steps:

[0080] Step 301, obtain the first risk classification system corresponding to the abnormal resource sample set.

[0081] Among them, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set.

[0082] Step 302, construct the first TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample.

[0083] Among them, the element in the i-th row and j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample.

[0084] Step 303, obtain the second TF-IDF matrix.

[0085] Among them, the second TF-IDF matrix is obtained by performing dimensionality reduction processing on the first TF-IDF matrix.

[0086] Step 304, determine the target quantity according to the quantity interval and the second TF-IDF matrix.

[0087] Among them, the target quantity is used to indicate the number of risk themes of the second risk classification system to be generated.

[0088] In order to improve the flexibility and rationality of the number of risk themes, as a possible implementation, within the set quantity interval of risk themes, generate multiple different numbers of risk themes, and determine the number of risk themes of the second risk classification system to be generated from the multiple different numbers of risk themes.

[0089] As an example, based on the quantity interval, cluster each element in the second TF-IDF matrix to obtain multiple candidate cluster sets; determine the target cluster set from the multiple candidate cluster sets; count the number of clusters in the target cluster set to obtain the target quantity.

[0090] That is to say, cluster the dimensionality-reduced TF-IDF matrix based on the quantity interval to generate multiple candidate cluster sets, evaluate the multiple candidate cluster sets to select the optimal target cluster set, and count the number of clusters in the target cluster set to obtain the target quantity; among them, the target quantity is used to indicate the number of risk themes of the second risk classification system to be generated.

[0091] Step 305: Using a probability generation model, based on the target quantity and the first TF-IDF matrix, predict the topic probability distribution of the risk topics corresponding to the abnormal resource sample set and the topic word probability distribution under each risk topic.

[0092] For the accuracy and reliability of risk classification, as a possible implementation, by predicting the topic probability distribution of the risk topics corresponding to the abnormal resource sample set and the topic word probability distribution under each risk topic, provide a quantitative basis for risk classification to avoid the biases that may be brought by traditional rules or manual settings.

[0093] As an example, input the target quantity as a parameter into the probability generation model, where the target quantity is used to specify the number of risk topics to be generated, and use the first TF-IDF matrix as the input data of the probability generation model. The probability generation model can automatically infer the topic probability distribution of the risk topics corresponding to the abnormal resource sample set, as well as the topic word probability distribution of multiple topic words under each risk topic.

[0094] Step 306: According to the topic word probability distribution under each risk topic, determine multiple risk topic words under each risk topic.

[0095] It should be noted that according to the topic word probability distribution under each risk topic, the probability value of each topic word under each risk topic can be determined. Among them, the higher the probability value of the topic word, the stronger the correlation between the topic word and the corresponding risk topic.

[0096] As a possible implementation, set the screening rules according to actual needs, and select the most relevant topic words from the topic word probability distribution under each risk topic as risk topic words.

[0097] As an example, for any risk topic, based on the topic word probability distribution under the any risk topic, select the topic words with a probability value greater than a certain threshold as the risk topic words under the any risk topic.

[0098] As another example, for any risk topic, based on the topic word probability distribution under the any risk topic, select the top N topic words with the highest probability values as the risk topic words under the any risk topic.

[0099] As yet another example, for any risk topic, combine the experience of experts in the relevant field and the topic word probability distribution under the any risk topic to screen out the risk topic words with practical significance.

[0100] Step 307: Generate a second risk classification system according to the multiple risk topic words under each risk topic.

[0101] To improve the integrity and accuracy of the second risk classification system, as a possible implementation, based on multiple risk topic words under each risk topic, determine the first-level classification labels and the second-level classification labels under the first-level classification labels in the to-be-generated second risk classification system.

[0102] As an example, for any risk topic, analyze and refine the multiple risk topic words under any risk topic to obtain the topic name of any risk topic; use each topic name as multiple first-level classification labels in the to-be-generated second risk classification system; determine the second-level classification labels under the first-level classification labels corresponding to each risk topic according to the multiple risk topic words under each risk topic; generate the second risk classification system according to the multiple first-level classification labels and the second-level classification labels under the multiple first-level classification labels.

[0103] That is to say, for each risk topic, analyze and refine the multiple risk topic words under the risk topic, summarize the topic names that can summarize the core features of the risk topic, and use these topic names as the first-level classification labels in the second risk classification system; then, further subdivide based on the risk topic words under each risk topic to determine the second-level classification labels corresponding to each first-level classification label to achieve more refined risk classification; finally, integrate all the first-level classification labels and their subordinate second-level classification labels to form a complete second risk classification system; it should be noted that when determining the second-level classification labels corresponding to each first-level classification label based on the risk topic words under each risk topic, if there are cases where the risk topic words under the risk topic have similar meanings or ambiguous meanings, the risk topic words with similar meanings or ambiguous meanings can be processed (such as merging, replacing, etc.), and then the processed risk topic words are used as the corresponding second-level classification labels.

[0104] Step 308, generate a target risk classification system according to the first risk classification system and the second risk classification system.

[0105] Among them, the target risk classification system is used to classify the risk of transaction resource information.

[0106] It should be noted that the execution processes of Step 301 to Step 302 and Step 308 can be implemented in any way in the respective embodiments of the present disclosure. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.

[0107] In summary, determining the target quantity based on the quantity interval and the second TF-IDF matrix ensures the scientificity and rationality of the number of risk themes, and avoids the deviation that may be caused by subjectively setting the number of risk themes. Furthermore, a probability generation model is used to combine the target quantity and the first TF-IDF matrix to predict the theme probability distribution and the theme word probability distribution of the risk themes, and multiple risk theme words under each risk theme are determined by analyzing the theme word probability distribution, so as to generate the second risk classification system based on the risk theme words, enhancing the interpretability of the risk classification and improving the accuracy and comprehensiveness of the second risk classification system at the same time.

[0108] To clearly illustrate how the target risk classification system is generated according to the first risk classification system and the second risk classification system in the above embodiments, the present disclosure proposes another risk classification method.

[0109] Figure 4 It is a schematic flowchart of the risk classification method shown in the fourth embodiment of the present disclosure.

[0110] As Figure 4 shown, the risk classification method includes the following steps:

[0111] Step 401, obtain the first risk classification system corresponding to the abnormal resource sample set.

[0112] Among them, the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set.

[0113] Step 402, construct the first TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample.

[0114] Among them, the element in the i-th row and j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample.

[0115] Step 403, generate the second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix.

[0116] Step 404, compare the first risk classification system and the second risk classification system to determine the same content and different content between the first risk classification system and the second risk classification system.

[0117] To improve the accuracy of the target risk classification system, as a possible implementation manner, the target risk classification system is jointly generated based on the first risk classification system and the second risk classification system.

[0118] In an embodiment of the present disclosure, a target risk classification system is generated based on the comparison result between the first risk classification system and the second risk classification system; that is, the first risk classification system and the second risk classification system are compared to obtain the same content and different content between the first risk classification system and the second risk classification system. Thus, based on the same content and different content, a target risk classification system is generated.

[0119] Among them, it should be noted that the different content may include missing risk categories, extra risk categories, differences in classification criteria, etc.

[0120] Step 405, generate a target risk classification system according to the same content and different content.

[0121] In order to improve the accuracy and comprehensiveness of the target risk classification system, as a possible implementation manner, the same content and different content are integrated to generate a target risk classification system.

[0122] As an example, a first sub-system is generated according to the same content; a second sub-system is generated according to the different content; the first sub-system and the second sub-system are combined to obtain a target risk classification system.

[0123] That is to say, the same content is the content that exists in both the first risk classification system and the second risk classification system. The credibility of this same content is relatively high, so the same content is directly used as a part of the target risk classification system, that is, the first sub-system. At the same time, the different content is the content that the first risk classification system is different from the second risk classification system. The different content is adjusted or optimized to obtain the second sub-system. For example, missing risk types are added, irrelevant risk classifications are deleted, similar risk classifications are merged, etc. Furthermore, the first sub-system and the second sub-system are combined to obtain a target risk classification system.

[0124] It should be noted that the execution processes of steps 401 to 403 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.

[0125] In summary, the first risk classification system and the second risk classification system are compared to determine the same content and different content between the first risk classification system and the second risk classification system; a target risk classification system is generated according to the same content and different content. Thus, by integrating the same content and the adjusted different content, a target risk classification system is generated, avoiding the omissions and subjective judgment biases that may exist in a single classification system, and enhancing the accuracy and comprehensiveness of the target risk classification system.

[0126] Based on any embodiment of the present disclosure, the risk classification method of the embodiment of the present disclosure can also be implemented based on the following steps:

[0127] Step 1: Construct a multi-source integrated data integration library

[0128] That is, by integrating and fusing transaction information generated by transaction behaviors that do not conform to the good customs of the process from various channels, an abnormal resource sample set is constructed;

[0129] Step 2: Generate an initial risk classification system

[0130] Step 2.1: Carry out the work of defining the key connotation boundaries, determine the connotation definition and extension characteristics of risk classification, and form the basic definitions of each risk classification.

[0131] That is, collect risk information related to each abnormal resource sample, and call the risk information related to each abnormal resource sample reference risk information, where the reference risk information may include, but is not limited to: transaction subject information related to the corresponding abnormal resource sample, historical transaction amount, user behavior pattern, historical transaction frequency, etc.

[0132] Step 2.2: Data risk feature analysis. Split each abnormal resource sample in the abnormal resource sample set into key risk information one by one, and extract keywords / characters from the risk information; use expert business experience to perform data merging, risk item type integration and splitting, etc. on the case risk items to form an initial risk classification system.

[0133] As an example, based on the reference risk information, extract features from each abnormal resource sample to obtain the sample risk features of each abnormal resource sample. Furthermore, based on the extracted sample risk features, construct a preliminary risk classification framework.

[0134] Step 2.3: Internal expert demonstration of risk classification. After completing the initial risk classification system, conduct multiple rounds of expert discussion and demonstration, and optimize the risk classification according to the conclusion of the demonstration to form the first risk classification system.

[0135] Step 3: Use machine learning technical methods for risk classification

[0136] To verify the effectiveness and accuracy of the classification, as Figure 5 shown, the machine learning method will be adopted to perform risk classification on the abnormal resource sample set. Finally, through verifying and comparing the classification results (the second risk classification system) after machine learning with the first risk classification system, demonstration and optimization are carried out to form the final risk classification.

[0137] Step 3.1: Preprocess the abnormal resource samples, including word segmentation, stop word removal, etc., and construct a TF-IDF matrix.

[0138] Step 3.2: Use t-SNE to reduce the dimension of the TF-IDF matrix. Considering the possible clustering overlap problem between the subjects, the clustering results are supplemented with manual adjustment to obtain the optimal number of topics.

[0139] Step 3.3: Use the number of clusters of t-SNE as the number of topics, define the parameters of the LDA model, and input the TF-IDF matrix into the LDA model to obtain the topic distribution of the abnormal resource sample set output by the LDA model and the topic word distribution of each topic. Output the 15 keywords with the highest frequency under each topic, and manually define the subject names according to the keywords under each topic as the first-level classification labels.

[0140] Step 3.4: Considering the possible cases of similar word meanings and ambiguous word meanings between the keywords, manually merge and combine the high-frequency words, and use the result as the second-level classification label under the first-level classification label.

[0141] Step 4: Compare the risk classification results (the second risk classification system) generated by the t-SNE and LDA machine learning technologies with the risk classification results (the first risk classification system) formed by expert experience judgment, conduct expert demonstration and optimization, verify the rationality of the risk classification system of expert experience judgment, generate the final risk classification system (the target risk classification system). Furthermore, the final risk classification system can be used to classify the risk of transaction resource information.

[0142] Corresponding to the risk classification method provided in the above embodiment, the present disclosure also provides a risk classification device. Since the risk classification device provided in the embodiment of the present disclosure corresponds to the risk classification method provided in the above embodiment, the implementation manner of the risk classification method is also applicable to the risk classification device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.

[0143] Figure 6 It is a schematic structural diagram of the risk classification device shown in the fifth embodiment of the present disclosure.

[0144] As Figure 6 shown, the risk classification device 600 includes: an acquisition module 610, a construction module 620, a first generation module 630, and a second generation module 640.

[0145] Among them, an acquisition module 610 is configured to acquire a first risk classification system corresponding to a set of abnormal resource samples, where the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the set of abnormal resource samples; a construction module 620 is configured to construct a first term frequency-inverse document frequency (TF-IDF) matrix corresponding to the set of abnormal resource samples according to each word segment in each abnormal resource sample, where the element in the i-th row and the j-th column of the first TF-IDF matrix is used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample; a first generation module 630 is configured to generate a second risk classification system according to a quantity interval corresponding to a risk theme and the first TF-IDF matrix; a second generation module 640 is configured to generate a target risk classification system according to the first risk classification system and the second risk classification system, where the target risk classification system is used to classify the risk of transaction resource information.

[0146] As a possible implementation manner of an embodiment of the present disclosure, the first generation module 630 is configured to acquire a second TF-IDF matrix, where the second TF-IDF matrix is obtained by performing dimensionality reduction processing on the first TF-IDF matrix; determine a target quantity according to the quantity interval and the second TF-IDF matrix, where the target quantity is used to indicate the number of risk themes of the second risk classification system to be generated; and generate the second risk classification system according to the target quantity and the first TF-IDF matrix.

[0147] As a possible implementation manner of an embodiment of the present disclosure, the first generation module 630 is configured to adopt a probabilistic generation model to predict a theme probability distribution of the target number of risk themes corresponding to the set of abnormal resource samples and a theme word probability distribution under each risk theme based on the target quantity and the first TF-IDF matrix; determine a plurality of risk theme words under each risk theme according to the theme word probability distribution under each risk theme; and generate the second risk classification system according to the plurality of risk theme words under each risk theme.

[0148] As a possible implementation manner of an embodiment of the present disclosure, the first generation module 630 is configured to, for any risk theme, analyze and refine a plurality of risk theme words under the any risk theme to obtain a theme name of the any risk theme; use each theme name as a plurality of first-level classification labels in the second risk classification system to be generated; determine second-level classification labels under the first-level classification labels corresponding to each risk theme according to the plurality of risk theme words under each risk theme; and generate the second risk classification system according to the plurality of first-level classification labels and the second-level classification labels under the plurality of first-level classification labels.

[0149] As a possible implementation manner of the embodiment of the present disclosure, the first generation module 630 is configured to cluster each element in the second TF-IDF matrix based on a quantity interval to obtain a plurality of candidate cluster sets; determine a target cluster set from the plurality of candidate cluster sets; and count the number of clusters in the target cluster set to obtain a target quantity.

[0150] As a possible implementation manner of the embodiment of the present disclosure, the second generation module 640 is configured to compare the first risk classification system and the second risk classification system to determine the same content and different content between the first risk classification system and the second risk classification system; and generate a target risk classification system according to the same content and different content.

[0151] As a possible implementation manner of the embodiment of the present disclosure, the second generation module 640 is configured to generate a first sub-system according to the same content; generate a second sub-system according to the different content; and combine the first sub-system and the second sub-system to obtain a target risk classification system.

[0152] As a possible implementation manner of the embodiment of the present disclosure, the first risk classification system is generated by using the following module: a third generation module.

[0153] Wherein, the third generation module is configured to obtain reference risk information associated with each abnormal resource sample; extract features from each abnormal resource sample based on the reference risk information to obtain sample risk features of each abnormal resource sample; generate an initial risk classification system according to the sample risk features of each abnormal resource sample; and generate a first risk classification system according to the initial risk classification system.

[0154] As a possible implementation manner of the embodiment of the present disclosure, the construction module 620 is configured to, for any abnormal resource sample, count the number of occurrences of each word segment in any abnormal resource sample in any abnormal resource sample to obtain a word frequency statistical result of each word segment in any abnormal resource sample; calculate the TF value of each word segment in any abnormal resource sample according to the word frequency statistical result of each word segment in any abnormal resource sample; calculate the IDF value of each word segment in any abnormal resource sample according to the document frequency of each word segment in the abnormal resource sample set; and generate a TF-IDF matrix corresponding to the abnormal resource sample set according to the TF value and IDF value of any word segment in each abnormal resource sample.

[0155] The risk classification device according to the embodiments of the present disclosure captures the core features of risks in actual business accurately by obtaining the first risk classification system generated based on the sample risk features in each abnormal resource sample, ensuring that the classification results are close to the actual business requirements. Furthermore, by constructing the first TF-IDF matrix of the abnormal resource sample set and generating the second risk classification system in combination with the quantity interval corresponding to the risk theme and the first TF-IDF matrix, the numerical processing of the abnormal resource sample is realized, the key features in the abnormal resource sample are quantified, and the important factors related to risks are highlighted, providing a scientific basis for generating the second risk classification system and enhancing the objectivity and comprehensiveness of risk classification. Finally, the target risk classification system is jointly generated based on the first risk classification system and the second risk classification system, improving the accuracy and comprehensiveness of the target risk classification system. Thus, the risk classification of transaction resource information is performed based on the target risk classification system, improving the accuracy of risk classification, which helps financial institutions better identify and manage risks and ensure the stable operation of the financial market.

[0156] In an exemplary embodiment, an electronic device is further proposed.

[0157] Wherein, the electronic device includes:

[0158] A processor;

[0159] A memory for storing executable instructions of the processor;

[0160] Wherein, the processor is configured to execute instructions to implement the risk classification method proposed in any of the foregoing embodiments.

[0161] As an example, Figure 7 is a schematic structural diagram of the electronic device 700 shown in an exemplary embodiment of the present disclosure. As Figure 7 shown, the above-mentioned electronic device 700 may further include:

[0162] A memory 710 and a processor 720, a bus 730 connecting different components (including the memory 710 and the processor 720). The memory 710 stores a computer program, and when the processor 720 executes the program, the risk classification method described in the embodiments of the present disclosure is implemented.

[0163] The bus 730 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0164] The electronic device 700 typically includes a variety of electronically readable media. These media can be any available media accessible to the electronic device 700, including volatile and non-volatile media, removable and non-removable media.

[0165] The memory 710 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 740 and / or cache memory 750. The server 700 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 760 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 7 not shown, commonly referred to as a "hard disk drive"). Although Figure 7 not shown in the figure, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 730 through one or more data media interfaces. The memory 710 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0166] A program / utility 780 having a set (at least one) of program modules 770 can be stored, for example, in the memory 710. Such program modules 770 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 770 generally perform the functions and / or methods in the embodiments described in the present disclosure.

[0167] The electronic device 700 can also communicate with one or more external devices 790 (such as a keyboard, a pointing device, a display 791, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 792. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 793. As shown in the figure, the network adapter 793 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0168] The processor 720 executes various functional applications and data processing by running the programs stored in the memory 710.

[0169] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the risk classification method of the embodiments of the present disclosure, and details will not be repeated here.

[0170] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, and the above instructions can be executed by the processor of the electronic device to complete the risk classification method proposed in any of the above embodiments. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0171] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the risk classification method proposed in any of the above embodiments is implemented.

[0172] Those skilled in the art will readily think of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0173] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A risk classification method, characterized in that: include: Obtaining a first risk classification system corresponding to the abnormal resource sample set; wherein the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set; According to each word segment in each abnormal resource sample, a first term frequency-inverse document frequency TF-IDF matrix corresponding to the abnormal resource sample set is constructed; wherein the j-th element in the i-th row of the first TF-IDF matrix is ​​used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample; Generate a second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix; A target risk classification system is generated according to the first risk classification system and the second risk classification system; wherein the target risk classification system is used to classify the risks of transaction resource information.

2. The method according to claim 1, characterized in that The generating a second risk classification system according to the quantity interval corresponding to the risk theme and the TF-IDF matrix includes: Obtain a second TF-IDF matrix; wherein the second TF-IDF matrix is ​​obtained by performing dimensionality reduction processing on the first TF-IDF matrix; Determine a target quantity according to the quantity interval and the second TF-IDF matrix; wherein the target quantity is used to indicate the number of risk topics of the second risk classification system to be generated; A second risk classification system is generated according to the target quantity and the first TF-IDF matrix.

3. The method according to claim 2, characterized in that Generating a second risk classification system according to the target quantity and the first TF-IDF matrix includes: Using a probability generation model, based on the target number and the first TF-IDF matrix, predict the topic probability distribution of the target number of risk topics corresponding to the abnormal resource sample set and the topic word probability distribution under each of the risk topics; Determining a plurality of risk keywords under each of the risk topics according to the probability distribution of the keywords under each of the risk topics; A second risk classification system is generated based on a plurality of risk keywords under each of the risk themes.

4. The method according to claim 3, characterized in that Generating a second risk classification system according to the plurality of risk keywords under each of the risk themes includes: For any risk theme, multiple risk subject words under the risk theme are parsed and refined to obtain a subject name of the risk theme; Using each of the subject names as a plurality of primary classification labels in the second risk classification system to be generated; According to the multiple risk keywords under each of the risk themes, determine the secondary classification labels under the primary classification labels corresponding to each of the risk themes; The second risk classification system is generated according to the multiple first-level classification labels and the secondary classification labels under the multiple first-level classification labels.

5. The method according to claim 2, characterized in that: The determining the target quantity according to the quantity interval and the second TF-IDF matrix includes: Based on the quantity interval, clustering each element in the second TF-IDF matrix to obtain a plurality of candidate cluster sets; Determine a target cluster set from the multiple candidate cluster sets; The number of clusters in the target cluster set is counted to obtain a target number.

6. The method according to claim 1, characterized in that Generating a target risk classification system according to the first risk classification system and the second risk classification system includes: Comparing the first risk classification system with the second risk classification system to determine the same contents and the different contents between the first risk classification system and the second risk classification system; The target risk classification system is generated based on the identical content and the different content.

7. The method according to claim 6, characterized in that The generating the target risk classification system according to the same content and the different content includes: According to the same content, a first subsystem is generated; generating a second subsystem according to the difference content; The first subsystem and the second subsystem are combined to obtain the target risk classification system.

8. The method according to claim 1, characterized in that The first risk classification system is generated by the following steps: Acquiring reference risk information associated with each of the abnormal resource samples; Based on the reference risk information, feature extraction is performed on each of the abnormal resource samples to obtain a sample risk feature of each of the abnormal resource samples; Generating an initial risk classification system according to the sample risk characteristics of each of the abnormal resource samples; The first risk classification system is generated according to the initial risk classification system.

9. The method according to any one of claims 1 to 8, characterized in that The step of constructing a term frequency-inverse document frequency TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample includes: For any abnormal resource sample, count the number of occurrences of each word in the abnormal resource sample to obtain a word frequency statistical result of each word in the abnormal resource sample; Calculate the TF value of each of the segmented words in any of the abnormal resource samples according to the word frequency statistics of each of the segmented words in any of the abnormal resource samples; Calculate the IDF value of each of the segmented words in any abnormal resource sample according to the document frequency of each of the segmented words in any abnormal resource sample in the abnormal resource sample set; According to the TF value and IDF value of any word segment in each of the abnormal resource samples, a TF-IDF matrix corresponding to the abnormal resource sample set is generated.

10. A risk classification device, characterized in that: include: An acquisition module, used to acquire a first risk classification system corresponding to the abnormal resource sample set; wherein the first risk classification system is generated according to the sample risk characteristics of each abnormal resource sample in the abnormal resource sample set; A construction module, used to construct a first term frequency-inverse document frequency TF-IDF matrix corresponding to the abnormal resource sample set according to each word segment in each abnormal resource sample; wherein the j-th element in the i-th row of the first TF-IDF matrix is ​​used to indicate the TF-IDF value of the j-th word segment in the i-th abnormal resource sample; A first generating module, configured to generate a second risk classification system according to the quantity interval corresponding to the risk theme and the first TF-IDF matrix; The second generating module is used to generate a target risk classification system according to the first risk classification system and the second risk classification system; wherein the target risk classification system is used to classify the risks of transaction resource information.