Transaction opponent identification method based on data mining, electronic equipment and storage medium

By processing bank internal account transaction data using data mining techniques, association rules are generated to identify counterparties, solving the problem of low efficiency in counterparty identification within internal accounts and achieving automated and accurate counterparty identification.

CN121504604APending Publication Date: 2026-02-10CHINA CONSTRUCTION BANK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511582503.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, the efficiency of identifying the real counterparty in bank internal account transaction records is low, which cannot meet the requirements of automatic deduction and repayment mechanisms.

Method used

By using data mining techniques, inflow and outflow data of internal users are obtained and preprocessed to determine multiple data item sets. Association rules are generated based on support and confidence to identify counterparties. If the preset rules cannot determine the counterparties, the association rules are used to further identify counterparties.

Benefits of technology

It enables automated and accurate identification of real counterparties in internal account transaction records, improves counterparty identification efficiency, and meets the needs of automatic deduction and repayment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504604A_ABST
    Figure CN121504604A_ABST
Patent Text Reader

Abstract

The invention provides a transaction opponent identification method based on data mining, electronic equipment and a storage medium, and relates to the technical field of data mining. The method comprises the following steps: acquiring inflow pipeline data of an internal household, and performing preprocessing according to the inflow pipeline data to obtain preprocessed data; determining a plurality of data item sets according to the preprocessed data and a plurality of target attributes of the target business; determining a target frequent item set according to the support degrees of the plurality of data item sets and a support degree threshold value; determining an association rule according to the confidence of the target frequent item set; determining a transaction opponent of the target business according to the inflow flow data and a preset identification rule; and if the transaction opponent cannot be determined according to the preset identification rule, determining the transaction opponent of the target service according to the inflow pipeline data and the association rule. The transaction opponent of the target business is automatically identified from the inflow flow data of the internal user, and the identification efficiency of the transaction opponent is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data mining technology, and in particular to a data mining-based counterparty identification method, electronic device, and storage medium. Background Technology

[0002] Internal accounts are dedicated accounts used by banks to account for internal fund transactions, primarily serving the institution's own fund clearing, asset management, and transitional business processing. However, corporate loan repayments sometimes involve using internal accounts for transitional repayments, thus requiring the identification of the actual counterparty in the internal account's transaction records.

[0003] Currently, manual identification is used to sift through a large amount of transaction data, resulting in low efficiency. Improving the efficiency of identifying genuine counterparties within internal accounts is a pressing issue. Summary of the Invention

[0004] This application provides a data mining-based counterparty identification method, electronic device, and storage medium to solve the problem of low efficiency in identifying real counterparties in internal account transaction data in the prior art.

[0005] In a first aspect, embodiments of this application provide a counterparty identification method based on data mining, including:

[0006] Obtain the inflow data of internal users, preprocess the inflow data to obtain preprocessed data, and determine multiple data item sets based on the preprocessed data and multiple target attributes of the target business.

[0007] The target frequent itemset is determined based on the support and support threshold of the multiple data itemsets; the association rule is determined based on the confidence of the target frequent itemset.

[0008] The counterparty for the target business is determined based on the inflow data and preset identification rules;

[0009] If the counterparty cannot be determined according to the preset identification rules, the counterparty of the target business is determined according to the inflow data and the association rules.

[0010] Secondly, embodiments of this application also provide a counterparty identification device based on data mining, comprising:

[0011] The data acquisition module is used to acquire inflow and outflow data of internal users;

[0012] The preprocessing module is used to preprocess the inflow water data to obtain preprocessed data;

[0013] The data itemset determination module is used to determine multiple data itemsets based on the preprocessed data and multiple target attributes of the target service;

[0014] The frequent itemset determination module is used to determine a target frequent itemset based on the support of the multiple data itemsets and a support threshold.

[0015] The association rule generation module is used to determine association rules based on the confidence level of the target frequent itemset;

[0016] The primary classification module is used to determine the counterparty of the target business based on the inflow data and preset identification rules;

[0017] The secondary classification module is used to determine the counterparty of the target business based on the inflow data and the association rules if the counterparty cannot be determined according to the preset identification rules.

[0018] Thirdly, embodiments of this application also provide an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the data mining-based counterparty identification method as shown in embodiments of this application.

[0019] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the data mining-based counterparty identification method as shown in the embodiments of this application.

[0020] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the data mining-based counterparty identification method shown in embodiments of this application.

[0021] The counterparty identification method based on data mining provided in this application preprocesses the inflow transaction data to obtain preprocessed data; determines multiple data itemsets based on the preprocessed data and multiple target attributes of the target business; determines target frequent itemsets based on the support and support threshold of the multiple data itemsets; determines association rules based on the confidence of the target frequent itemsets; and determines the counterparty of the target business based on the inflow transaction data and preset identification rules. If the counterparty cannot be determined based on the preset identification rules, the counterparty of the target business is determined based on the inflow transaction data and the association rules. Compared to the current method of manually screening real transaction accounts, this application can first determine multiple data itemsets based on the inflow transaction data of internal accounts and multiple target attributes of the target business; determine target frequent itemsets based on the support and support threshold of the multiple data itemsets; and determine association rules between transaction data and counterparties based on the target frequent itemsets with high confidence. Then, it identifies counterparties from the inflow transaction data according to the preset identification rules; if the preset identification rules cannot determine counterparties, the counterparty of the target business is determined based on the inflow transaction data and the association rules. It accurately identifies the real counterparties in the inflow data, automates the identification of counterparties for target businesses from the inflow data of internal accounts, and improves the efficiency of counterparty identification. Attached Figure Description

[0022] Figure 1 This is the flowchart of the data mining-based counterparty identification method provided in the embodiments of this application. Figure 1 ;

[0023] Figure 2 This is the flowchart of the data mining-based counterparty identification method provided in the embodiments of this application. Figure 2 ;

[0024] Figure 3 This is the flowchart of the data mining-based counterparty identification method provided in the embodiments of this application. Figure 3 ;

[0025] Figure 4 This is a schematic diagram of the structure of the data mining-based counterparty identification device provided in the embodiments of this application. Figure 1 ;

[0026] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0028] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0029] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0030] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0031] The information collected in this invention embodiment is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, and necessary measures have been taken to ensure compliance with public order and good morals. A corresponding operation entry point is provided for users to choose to authorize or refuse. Users are provided with a corresponding operation entry point to choose to agree to or refuse the automated decision-making result; if the user chooses to refuse, the process enters the expert decision-making process to avoid relevant legal and public opinion risks.

[0032] The technical terms used in this application are explained below.

[0033] Internal account: A dedicated account used by a bank to account for internal fund transactions, primarily serving the institution's own fund clearing, asset management, and transitional business processing. Specifically, internal accounts are special accounts set up by the bank for the collection of accounts receivable.

[0034] Counterparties: Counterparties are one of the main participants in a transaction. In any transaction, whether buying or selling commodities, securities, or financial derivatives, there will be a buyer and a seller. These two parties are the counterparties. They transact through trading platforms or intermediaries, jointly determining the price and quantity of the transaction.

[0035] Related companies: These are companies that form close relationships with other companies through control or influence, typically including equity control, board seats, family ties, etc.

[0036] Corporate loans are a type of financial service offered to businesses or other organizations, also known as corporate loans or institutional loans. Their main characteristic is that the application is made in the name of the company or organization, requiring the provision of relevant qualifications, financial statements, collateral, and other documentation. After review, the bank provides a loan of a certain amount. Corporate loans are typically used for daily operating funds, fixed asset investment, and technological upgrades.

[0037] Accounts receivable refers to the amounts owed by a company to its customers for the sale of goods, products, or services in the ordinary course of business. This includes taxes payable by the purchasing or service-receiving unit, and various freight and miscellaneous charges advanced on behalf of the buyer. Accounts receivable is a claim that arises with a company's sales activities. Accounts receivable includes both existing and future claims. The former refers to claims that have already occurred and are clearly established, while the latter refers to claims that have not yet occurred but will certainly occur in the future.

[0038] Debt: This refers to a contract or agreement signed between a borrower and a creditor that stipulates the borrower's future repayment to the creditor. Debt typically includes the principal, interest, and other fees payable by the borrower. Debt can be a personal loan or corporate bond, and its issuance is usually for the purpose of raising funds for investment or business expansion.

[0039] Support: refers to the frequency with which an itemset or rule appears in all transactions.

[0040] Confidence level: refers to the probability that another itemset will appear given that one itemset is included.

[0041] The inventors discovered that internal bank accounts cannot identify the true source of the inflow account (counterparty) when making repayments. This is primarily because the name of the inflow account has no direct correlation with the debt being repaid, and there is a time lag between the transferred funds and the debt's maturity date. Furthermore, there is information confusion within the funds pool; for example, funds transferred uniformly from accounts of finance companies or affiliated companies cannot be identified as being used to repay any future loan debt.

[0042] Currently, the system typically uses precise identification based on the account holder's name and manual counter verification. However, this method suffers from low system coverage and low processing efficiency, failing to meet the needs of current automatic deduction and repayment mechanisms. Improving the efficiency of identifying genuine counterparties within internal accounts has become an urgent problem to be solved.

[0043] The technical solution of this application and how it solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The processing of user identifiers, user card numbers, and other information involved in this application all comply with national laws and regulations. The embodiments of this application will be described below with reference to the accompanying drawings.

[0044] Figure 1 The flowchart of the data mining-based counterparty identification method provided in the embodiments of this application Figure 1 This embodiment can be applied to scenarios where real trading counterparties are identified based on the inflow and outflow data of internal accounts, and where voice navigation lists are optimized to provide better service to customers. This method can be executed by electronic devices such as servers. Figure 1 As shown, the data mining-based counterparty identification method provided in this application includes:

[0045] Step S101: Obtain the inflow data of internal users.

[0046] Internal account inflow data includes transfer summary information, counterparty details, inflow time, and inflow amount. In addition to inflow data, the following data can be obtained to execute subsequent steps: Information on outstanding debts of the borrower, including maturity date, debt amount, and handling institution; Centralized relationship tree information of the core enterprise, including the core enterprise and the actual buyer's related party relationships; Accounts receivable information, including counterparty details, accounts receivable amount, and maturity date; and Business registration information, including the equity relationship of the counterparty.

[0047] Step S102: Preprocess the inflow water data to obtain preprocessed data.

[0048] Inflow data can be real-time or historical.

[0049] Optionally, preprocessing can be performed on the inflow data to obtain preprocessed data, which can be implemented as follows:

[0050] The inflow data is cleaned of missing and outlier values ​​to obtain preprocessed data.

[0051] In addition to cleaning the inflow data, we can also clean the information on outstanding debts of borrowers, the centralized relationship tree information of core enterprises, accounts receivable information, and business registration information to remove missing and outlier values.

[0052] The above implementation method can clean up outliers and differences in the incoming data, thereby improving the reliability of the preprocessed data and the robustness of the association rules.

[0053] Step S103: Determine multiple data item sets based on the preprocessed data and multiple target attributes of the target service.

[0054] Optionally, multiple data item sets can be determined based on the preprocessed data and multiple target attributes of the target service, which can be implemented as follows:

[0055] The first item set is determined based on the fact that the account name in the preprocessed data matches the name of the accounts receivable counterparty or the name of the core enterprise in the debt in the industrial and commercial equity structure; the second item set is determined based on the fact that the extracted name in the preprocessed data matches the accounts receivable number or the name of the accounts receivable counterparty or the name of the core enterprise in the debt; the third item set is determined based on the fact that the amount in the preprocessed data matches the accounts receivable balance; and the fourth item set is determined based on the fact that the date in the preprocessed data matches the accounts receivable due date range.

[0056] The above implementation method can generate itemsets according to corporate loan scenarios, thereby improving the reliability of itemsets in corporate loan scenarios.

[0057] Step S104: Determine the target frequent itemset based on the support and support threshold of the multiple data itemsets.

[0058] Optionally, the target frequent itemset is determined based on the support and support threshold of the plurality of data itemsets, including:

[0059] Step 1: Determine frequent 1-itemsets based on the support and support threshold of individual itemsets in the multiple data item sets.

[0060] Specifically, the dataset is scanned once to count the support of each individual item (1-itemset). Frequent 1-itemsets are then selected based on a set support threshold, with 1-itemsets having a support greater than or equal to the threshold.

[0061] Step 2: Perform a self-join based on the frequent 1-itemset and the k-1-itemset to obtain candidate itemsets, where the k-1-itemset is the itemset in the multiple data itemsets other than the frequent 1-itemset.

[0062] Specifically, frequent (k-1) itemsets are self-joined to generate candidate k itemsets. The candidate k itemsets are then pruned to remove itemsets containing infrequent subsets.

[0063] Step 3: Determine the frequent k-itemsets based on the support of the candidate itemsets and the support threshold. Repeat steps 1 and 2 until no new frequent itemsets can be generated.

[0064] Scan the dataset: Scan the dataset a second time to calculate the support of each candidate k-itemset. Filter frequent k-itemsets: Based on the set support threshold, filter k-itemsets with support greater than or equal to the threshold as frequent k-itemsets. Repeat steps 1 and 2 until no new frequent itemsets can be generated. Each iteration generates candidate k-itemsets and filters out frequent k-itemsets.

[0065] The above implementation method can obtain frequent k-itemsets through multiple rounds of iterative updates based on the support of the itemsets and the support threshold, thereby improving the accuracy of frequent k-itemsets.

[0066] Step S105: Determine the association rule based on the confidence level of the target frequent itemset.

[0067] Optionally, the association rule can be determined based on the confidence level of the target frequent itemset, which can be implemented as follows:

[0068] Calculate the confidence score of the non-empty subsets of each frequent itemset; based on the confidence score and confidence threshold of the non-empty subsets of each frequent itemset, determine the non-empty subsets with a confidence score greater than the confidence threshold; determine the association rule based on the non-empty subsets and the frequent itemsets to which the non-empty subsets belong.

[0069] For each frequent itemset, calculate the confidence score for all its non-empty subsets. The confidence score represents the probability that two itemsets occur simultaneously. Based on a set confidence score threshold, filter out association rules with a confidence score greater than or equal to the threshold.

[0070] The above implementation can generate association rules based on the confidence of non-empty subsets of frequent itemsets, thereby improving the accuracy of association rules.

[0071] Step S106: Determine the counterparty of the target business based on the inflow data and preset identification rules.

[0072] Figure 2 This is the flowchart of the data mining-based counterparty identification method provided in the embodiments of the present invention. Figure 2 .like Figure 2 As shown, after collecting the aforementioned multiple data sources, primary classification is performed first. If the primary classification cannot identify the counterparty and therefore cannot label it, secondary classification is then executed. In secondary classification, counterparty identification is performed using an association rule model.

[0073] Optionally, determining the counterparty of the target business based on the inflow data and preset identification rules can be implemented as follows:

[0074] The system matches payment accounts with preset matching information from the inflow data. The preset matching information includes: account information for outstanding debts, centralized relationship tree information of core enterprises, or accounts receivable information. If the match is successful, the counterparty is determined based on the matched account information.

[0075] The counterparty information in the inflow of funds can be matched one by one with the counterparty information of outstanding debts, the centralized relationship tree information of core enterprises, and accounts receivable information. If a match is found, the counterparty can be identified, and then tagging can be performed to mark the identified counterparty. If all three matches fail, tagging cannot be performed, and step S107 is executed.

[0076] The above implementation method can match the outstanding debt account information, the core enterprise centralized relationship tree information, or the accounts receivable information with the payment account in the inflow data in sequence, thereby realizing the identification of the counterparty, achieving rapid identification of the counterparty, and improving the identification efficiency.

[0077] Step S107: If the counterparty cannot be determined according to the preset identification rules, the counterparty of the target business is determined according to the inflow data and the association rules.

[0078] Association rules can be obtained in advance through training, or they can be applied in real time based on the association rules.

[0079] Optionally, determining the counterparty of the target business based on the inflow data and the association rules can be implemented as follows:

[0080] The counterparty of the target business is determined based on the target attributes of the target business and the confidence level of the association rule.

[0081] Specifically, by comparing the confidence levels of the association rules between different debts and account inflows, the counterparty with the strongest correlation to the debt information is identified.

[0082] The above implementation method can accurately determine the counterparty from the transaction data based on the target attributes of the target business and the confidence level of the association rules, thereby improving the accuracy of counterparty identification.

[0083] Furthermore, before determining the counterparty of the target business based on the inflow data and the association rules, the process also includes:

[0084] The data flow is divided into multiple training and test sets; association rules are determined based on the multiple training sets; and the association rules are verified based on the test sets.

[0085] The original dataset is divided into training and test sets. K-Fold Cross-Validation is employed, dividing the dataset into K subsets. In each cross-validation, K-1 subsets are used as the training set, and the remaining subset is used as the test set. In each cross-validation, the association planning algorithm is rerun using the training set to generate new frequent itemsets and association rules. Then, the accuracy and stability of the generated association rules are validated using the test set. The accuracy of each cross-validation is calculated as the number of correctly classified samples divided by the total number of samples. The results of all cross-validations are summarized, and the average performance index is calculated to obtain the overall performance of the association rule algorithm on the entire dataset, thus yielding the final association rules. The association rule algorithm used in this application can be the Apriori algorithm.

[0086] The above implementation method can optimize association rules multiple times using known training and test sets during the association rule generation process, thereby improving the reliability of association rules.

[0087] The above implementation method can determine the data itemset based on the inflow data, generate the target frequent itemset based on the support, and generate association rules based on the target frequent itemset and the confidence level. This enables data mining of the inflow data based on the association rule algorithm to obtain the association rules of the target business and improve the accuracy of counterparty identification.

[0088] This application provides a data mining-based counterparty identification method that obtains inflow transaction data of internal accounts, preprocesses the inflow transaction data to obtain preprocessed data, determines multiple data item sets based on the preprocessed data and multiple target attributes of the target business, determines target frequent itemsets based on the support and support threshold of the multiple data item sets, determines association rules based on the confidence of the target frequent itemsets, determines the counterparty of the target business based on the inflow transaction data and preset identification rules, and if the counterparty cannot be determined based on the preset identification rules, the counterparty of the target business is determined based on the inflow transaction data and the association rules. Compared to the current method of manually screening real transaction accounts, this application can first determine multiple data item sets based on the inflow transaction data of internal accounts and multiple target attributes of the target business, determine target frequent itemsets based on the support and support threshold of the multiple data item sets, and determine association rules between transaction data and counterparties based on the target frequent itemsets with higher confidence. Then, the counterparty is identified from the inflow data according to preset identification rules. If the preset identification rules cannot determine the counterparty, the counterparty of the target business is determined based on the inflow data and the association rules. This accurately identifies the real counterparty in the inflow data, automates the identification of the counterparty of the target business from the inflow data of internal users, and improves the efficiency of counterparty identification.

[0089] Figure 3 The flowchart of the data mining-based counterparty identification method provided in the embodiments of this application Figure 3 .like Figure 4 As shown, the data mining-based counterparty identification method includes the following steps:

[0090] Step S301: Obtain historical inflow data of internal users, and divide the historical inflow data into multiple training sets and test sets; determine association rules based on the multiple training sets.

[0091] Step S302: For the historical inflow data in the training set, clean the historical inflow data of missing values ​​and outliers to obtain preprocessed data.

[0092] Step S303: Determine the first item set based on the fact that the account name in the preprocessed data matches the name of the accounts receivable counterparty or the name of the core enterprise in the debt in the industrial and commercial equity structure; determine the second item set based on the fact that the extracted summary name in the preprocessed data matches the accounts receivable number or the name of the accounts receivable counterparty or the name of the core enterprise in the debt; determine the third item set based on the fact that the amount in the preprocessed data matches the accounts receivable balance; determine the fourth item set based on the fact that the date in the preprocessed data matches the accounts receivable due date range.

[0093] Step S304: Determine frequent 1-itemsets based on the support and support threshold of individual itemsets in the multiple data item sets; perform self-joins between the frequent 1-itemsets and k-1 itemsets to obtain candidate itemsets; determine frequent k-itemsets based on the support and support threshold of the candidate itemsets, and repeat the above steps until no new frequent itemsets can be generated.

[0094] Wherein, the k-1 itemset is the itemset in the plurality of data itemsets other than the frequent 1-itemset.

[0095] Step S305: Calculate the confidence level of the non-empty subsets of each frequent itemset; determine the non-empty subsets with a confidence level greater than the confidence threshold based on the confidence level and the confidence threshold of the non-empty subsets of each frequent itemset; determine the association rules based on the non-empty subsets and the frequent itemsets to which the non-empty subsets belong.

[0096] Step S306: Validate the association rules based on the test set.

[0097] Step S307: Obtain real-time internal user inflow data.

[0098] Step S308: Determine the counterparty of the target business based on the confidence level of the target attributes of the target business and the association rule.

[0099] Step S309: If the counterparty cannot be determined according to the preset identification rules, the counterparty of the target business shall be determined according to the inflow data and the verified association rules.

[0100] S310. Generate a receipt notification message based on the counterparty and output the notification message.

[0101] The data mining-based counterparty identification method provided in this application can accurately identify the real counterparties in the real-time inflow data of internal accounts, automatically identifying the counterparties of target businesses from the inflow data of internal accounts and improving the efficiency of counterparty identification. Based on this, it can achieve real-time payment arrival notifications, improving the accuracy and timeliness of notifications.

[0102] Figure 4 This is a schematic diagram of the structure of a data mining-based counterparty identification device provided in an embodiment of this application. Figure 4 As shown, the data mining-based counterparty identification device includes: a data acquisition module 41, a preprocessing module 42, a data itemset determination module 43, a frequent itemset determination module 44, an association rule generation module 45, a primary classification module 46, and a secondary classification module 47.

[0103] Data acquisition module 41 is used to acquire inflow data of internal users;

[0104] Preprocessing module 42 is used to preprocess the inflow water data to obtain preprocessed data;

[0105] The data item set determination module 43 is used to determine multiple data item sets based on the preprocessed data and multiple target attributes of the target service;

[0106] The frequent itemset determination module 44 is used to determine a target frequent itemset based on the support of the multiple data itemsets and the support threshold.

[0107] The association rule generation module 45 is used to determine association rules based on the confidence level of the target frequent itemset;

[0108] The primary classification module 46 is used to determine the counterparty of the target business based on the inflow data and preset identification rules;

[0109] The secondary classification module 47 is used to determine the counterparty of the target business based on the inflow data and the association rules if the counterparty cannot be determined according to the preset identification rules.

[0110] In some embodiments, the frequent itemset determination module 44 performs the following steps:

[0111] Frequent one-itemsets are determined based on the support of individual itemsets and the support threshold in the multiple data item sets;

[0112] A self-join is performed based on the frequent 1-itemset and the k-1-itemset to obtain candidate item sets, wherein the k-1-itemset is the item set in the multiple data item sets other than the frequent 1-itemset;

[0113] Based on the support of the candidate itemsets and the support threshold, determine the frequent k-itemsets, and repeat the above steps until no new frequent itemsets can be generated.

[0114] In some embodiments, the association rule generation module 45 is used for:

[0115] Calculate the confidence score of the non-empty subset of each frequent itemset;

[0116] Based on the confidence level and confidence threshold of the non-empty subsets of each frequent itemset, determine the non-empty subsets with a confidence level greater than the confidence threshold;

[0117] Association rules are determined based on the non-empty subset and the frequent itemsets to which the non-empty subset belongs.

[0118] In some embodiments, the secondary classification module 47 is used for:

[0119] The counterparty of the target business is determined based on the target attributes of the target business and the confidence level of the association rule.

[0120] In some embodiments, the data item set determination module 43 is used for:

[0121] The first item set is determined based on the situation where the account name in the preprocessed data is the same as the name of the counterparty in the accounts receivable transaction and the name of the core enterprise in the debt item in the industrial and commercial equity structure.

[0122] The second item set is determined based on the situation where the extracted names in the preprocessed data match the accounts receivable number or the name of the accounts receivable counterparty or the name of the core enterprise in the debt item.

[0123] The third itemset is determined based on the matching of amounts with accounts receivable balances in the preprocessed data;

[0124] The fourth itemset is determined based on the matching of dates with the due dates of accounts receivable in the preprocessed data.

[0125] In some embodiments, the preprocessing module 42 is used for:

[0126] The inflow data is cleaned of missing and outlier values ​​to obtain preprocessed data.

[0127] In some embodiments, the primary classification module 46 is used for:

[0128] The data is matched based on the payment account in the inflow data and preset matching information; the preset matching information includes: account information for outstanding debts, centralized relationship tree information of core enterprises, or accounts receivable information;

[0129] If a match is found, the counterparty will be determined based on the matched account information.

[0130] In some embodiments, the system further includes a dataset partitioning module and a verification module.

[0131] The dataset partitioning module is used to divide the inflow data into multiple training and test sets.

[0132] The verification module is used to determine association rules based on the multiple training sets and to verify the association rules based on the test set.

[0133] The data mining-based counterparty identification device provided in this application includes: a preprocessing module 42 for preprocessing the inflow data to obtain preprocessed data; a data itemset determination module 43 for determining multiple data itemsets based on the preprocessed data and multiple target attributes of the target business; a frequent itemset determination module 44 for determining a target frequent itemset based on the support and support threshold of the multiple data itemsets; an association rule generation module 45 for determining association rules based on the confidence level of the target frequent itemset; a data acquisition module 41 for acquiring inflow data of internal users; a primary classification module 46 for determining the counterparty of the target business based on the inflow data and preset identification rules; and a secondary classification module 47 for determining the counterparty of the target business based on the inflow data and the association rules if the counterparty cannot be determined based on the preset identification rules. Compared to the current method of manually screening real transaction accounts, this application first determines multiple data item sets based on the inflow transaction data of internal accounts and multiple target attributes of the target business. Then, it determines target frequent item sets based on the support and support threshold of these data item sets. Finally, it determines association rules between transaction data and counterparties based on the frequent item sets with high confidence. Next, it identifies counterparties from the inflow transaction data according to preset identification rules. If the preset identification rules cannot determine the counterparty, it determines the counterparty of the target business based on the inflow transaction data and the association rules. This accurately identifies real counterparties in the inflow transaction data, automating the identification of counterparties for the target business from the inflow transaction data of internal accounts and improving the efficiency of counterparty identification.

[0134] The counterparty identification device based on data mining provided in this application can be used to execute the technical solution of the counterparty identification method based on data mining in the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0135] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the secondary classification module 47 can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its function can be called and executed by a processing element of the above device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.

[0136] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device may include: transceiver 51, processor 52, and memory 53.

[0137] Processor 52 executes computer execution instructions stored in memory, causing processor 52 to perform the scheme in the above embodiments. Processor 52 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0138] The memory 53 is connected to the processor 52 via the system bus and completes communication between them. The memory 53 is used to store computer program instructions.

[0139] The transceiver 51 can be used to interact with clients.

[0140] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.

[0141] This application also provides a chip for executing instructions, which is used to execute the technical solution of the counterparty identification method based on data mining in the above embodiments.

[0142] This application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer performs the technical solution of the data mining-based counterparty identification method described in the above embodiments.

[0143] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solution of the counterparty identification method based on data mining in the above embodiments.

[0144] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0145] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for identifying trading counterparties based on data mining, characterized in that, include: Obtain the inflow data of internal users, and preprocess the inflow data to obtain preprocessed data; Multiple data item sets are determined based on the preprocessed data and multiple target attributes of the target business; The target frequent itemset is determined based on the support and support threshold of the multiple data itemsets; Association rules are determined based on the confidence level of the target frequent itemsets; The counterparty for the target business is determined based on the inflow data and preset identification rules; If the counterparty cannot be determined according to the preset identification rules, the counterparty of the target business is determined according to the inflow data and the association rules.

2. The method according to claim 1, characterized in that, The target frequent itemset is determined based on the support and support threshold of the multiple data itemsets, including: Frequent one-itemsets are determined based on the support and support threshold of individual itemsets in the multiple data item sets; a self-join is performed on the frequent one-itemsets and k-1 itemsets to obtain candidate itemsets, where the k-1 itemsets are itemsets in the multiple data item sets other than the frequent one-itemsets; Based on the support of the candidate itemsets and the support threshold, determine the frequent k-itemsets, and repeat the above steps until no new frequent itemsets can be generated.

3. The method according to claim 1, characterized in that, Determine association rules based on the confidence level of the target frequent itemsets, including: Calculate the confidence score of the non-empty subset of each frequent itemset; Based on the confidence level and confidence threshold of the non-empty subsets of each frequent itemset, determine the non-empty subsets with a confidence level greater than the confidence threshold; Association rules are determined based on the non-empty subset and the frequent itemsets to which the non-empty subset belongs.

4. The method according to claim 1, characterized in that, The counterparty for the target business is determined based on the inflow data and the association rules, including: The counterparty of the target business is determined based on the target attributes of the target business and the confidence level of the association rule.

5. The method according to claim 1, characterized in that, Multiple data item sets are determined based on the preprocessed data and multiple target attributes of the target service, including: The first item set is determined based on the situation where the account name in the preprocessed data is the same as the name of the counterparty in the accounts receivable transaction and the name of the core enterprise in the debt item in the industrial and commercial equity structure. The second item set is determined based on the situation where the extracted names in the preprocessed data match the accounts receivable number or the name of the accounts receivable counterparty or the name of the core enterprise in the debt item. The third itemset is determined based on the matching of amounts with accounts receivable balances in the preprocessed data; The fourth itemset is determined based on the matching of dates with the due dates of accounts receivable in the preprocessed data.

6. The method according to claim 1, characterized in that, Preprocessing is performed on the inflow water data to obtain preprocessed data, including: The inflow data is cleaned of missing and outlier values ​​to obtain preprocessed data.

7. The method according to claim 1, characterized in that, The counterparty for the target business is determined based on the inflow data and preset identification rules, including: The data is matched based on the payment account in the inflow data and preset matching information; the preset matching information includes: account information for outstanding debts, centralized relationship tree information of core enterprises, or accounts receivable information; If a match is found, the counterparty will be determined based on the matched account information.

8. The method according to claim 1, characterized in that, Before determining the counterparty for the target business based on the inflow data and the association rules, the process also includes: The data is divided into multiple training and testing sets based on the inflow water data; Association rules are determined based on the multiple training sets; the association rules are validated based on the test set.

9. A counterparty identification device based on data mining, characterized in that, include: The data acquisition module is used to acquire the inflow and outflow data of internal users; The preprocessing module is used to preprocess the inflow water data to obtain preprocessed data; The data itemset determination module is used to determine multiple data itemsets based on the preprocessed data and multiple target attributes of the target service; The frequent itemset determination module is used to determine a target frequent itemset based on the support of the multiple data itemsets and a support threshold. The association rule generation module is used to determine association rules based on the confidence level of the target frequent itemset; The primary classification module is used to determine the counterparty of the target business based on the inflow data and preset identification rules; The secondary classification module is used to determine the counterparty of the target business based on the inflow data and the association rules if the counterparty cannot be determined according to the preset identification rules.

10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.

12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-8.