Target information mining method based on visualization
By adopting visual-based target information mining methods in financial institutions, the abnormal patterns in transaction data are automatically identified and analyzed, and the limitations of financial institutions relying on manual and rules in abnormal transaction monitoring are solved, and fast and accurate abnormal transaction detection and early warning are achieved.
Patent Information
- Application Number
- CN202510661366.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
AI Technical Summary
In monitoring of abnormal transaction data, it is difficult for financial institutions to quickly and effectively discover data clues in massive complex data. The traditional manual analysis model relies on preset rules and is difficult to cover all abnormal patterns, resulting in legal transactions being wrongly marked as abnormal and the rules are slowly adjusted.
A visual-based target information mining method is adopted, by collecting historical transaction records, extracting feature data, performing standardized processing and clustering analysis, automatically identifying transaction patterns, dynamically adjusting thresholds, identifying abnormal transactions, and visually displaying abnormal transactions through visual charts.
Real-time analysis of each transaction is realized, quickly identify potential abnormal transactions, reduce response delays, reduce manual intervention needs, provide more comprehensive abnormal detection, and have high adaptability and reduce false positives and missed reports.
Smart Images

Figure CN120179882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to a method for mining target information based on visualization. Background Art
[0002] In the aspect of monitoring abnormal transaction data by financial institutions, it is relatively difficult to trace a relatively complete transaction data path. Abnormal transactions have gradually formed a complete upstream, middle, and downstream link, and the channels and forms are also more variable. After data extraction, feature analysis, and abnormal data determination, the transaction data and transaction link of abnormal transactions become more obvious.
[0003] For financial institutions, relying solely on the traditional manual analysis mode, it is difficult to quickly and effectively discover data clues in the face of massive and complex data. As time goes by, abnormal transactions have greatly improved both in scale and intelligence value. Judging whether a transaction is abnormal based on preset rules, such as the transaction amount exceeding a certain threshold, frequent small transactions, etc., although simple and effective, it is difficult to cover all abnormal patterns. Over-reliance on rules will cause many legal transactions to be wrongly marked as abnormal, resulting in customer dissatisfaction and business interruption. When new fraud behaviors appear, the adjustment of rules is relatively slow and requires manual intervention and continuous rule updates. In the abnormal data output report, it is impossible to intuitively display the hierarchical situation of the transaction group and the complete transaction link.
[0004] Therefore, people need to use a method for mining target information based on visualization to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for mining target information based on visualization to solve the problems raised in the prior art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for mining target information based on visualization, the method comprising the following steps: S100. Collect the historical transaction records of a number of accounts within a certain time range, and use the transaction data collected from the historical transaction records as the original data; Further, the step of collecting the original data is; S101. Extract the historical transaction records within the time period t with a period of T. The historical transaction records include transaction numerical data, transaction time data, and transaction frequency data, and use the extracted data as the original data.
[0007] S200. According to the extracted original data, establish a big data model to extract data features from the original data to obtain feature data; Further, the specific steps of extracting features from the original data are: S201. Extract the specific times of several transactions of a single account H, the value of each single transaction, from the original data within the time period t. Divide the time period t evenly into p trading time periods, and the length of each trading time period is L. p , extract the number of transactions y that occurred in the z-th trading time period, and calculate the frequency P of transactions occurring in the z-th time period. z , , where z ∈ [1, p] and z is a positive integer; Aggregate the frequencies of each trading time period to obtain the frequency change set P = {P1, P2, P3..., P p} of account H within the time period t, where P1, P2, P3...P p are the frequencies of the 1st, 2nd, 3rd,..., and p-th trading time periods respectively; S202. By extracting all the single transaction values of each trading time period within the time period t, calculate the average value S of the single transaction values within the z-th trading time period; Aggregate the S values of each trading time period to obtain the set S = {S1, S2, S3..., S p} of the average single transaction values of account H within the time period t, where S1, S2, S3...S p are the average single transaction values of the 1st, 2nd, 3rd,..., and p-th trading time periods respectively; S203. For each trading time period, calculate the change rate B between the n-th and (n - 1)-th single transaction values ’ , , calculate the average value of the change rates of the single transaction values within the z-th trading time period, and take the average value of the change rates of the single transaction values within the z-th trading time period as the transaction value change rate B of the z-th trading time period, to obtain the transaction value change rates of account H within p trading time periods within the time period t; Aggregate the transaction value change rates B of each trading time period to obtain the set B = {B1, B2, B3..., B p} of the average single transaction values of account H within the time period t, where B1, B2, B3...B p are the average single transaction values of the 1st, 2nd, 3rd,..., and p-th trading time periods respectively; S204. Integrate the sets P, B, and S to obtain a feature dataset A = {S1:P1;B1, S2:P2:B2,..., S P :P P :B p} established based on account H, S1:P1:B1, S2:P2:B2,..., S P :P P:B p represent the 1st, 2nd, …, and pth elements in the feature dataset A, where S P , P P , B p As the three types of feature data of account H, S P :P P :B p Extract the feature data of m accounts to obtain a set M = {A1, A2, ..., Am}, where A1, A2, A3, …, and Am represent the feature datasets of the 1st, 2nd, 3rd, …, and mth accounts respectively.
[0008] Through the big data model, it can be automatically cleaned and standardized without relying on manually set rules. Therefore, it can process larger-scale and more complex datasets to extract meaningful features.
[0009] S300. By standardizing the feature data, eliminate the difference in data dimensions between different feature data to obtain standardized data; Furthermore, the specific steps to quantify the feature data to obtain standardized data are as follows: S301. Standardize each type of feature data within each feature dataset in set M. The formula is , where μ is the average value of each type of feature data, σ is the standard deviation of each type of feature data. Through the standardization calculation formula, Z represents the result after standardizing each type of feature data. The standardized results are pooled to obtain a standardized dataset M ’ .
[0010] S400. Perform clustering analysis on the standardized data and classify the standardized data according to the analysis results; Furthermore, the specific steps to classify the standardized data according to the analysis results and set the anomaly threshold are as follows: S401. Set each element in the standardized feature dataset A as a data point x. Each data point has three types of feature data. Each data point x is a three-dimensional vector, denoted as x = (x i1 , x i2 , x i3 ), to obtain a data point set X = {x1, x2, ..., x n}. Randomly select k data points as a new set C = {c1, c2, ..., c k}. Calculate the distances from n data points to each data point in set C, denoted as d = {x n, c1, x n, c2, ... x n, c k}, x n, ck Represents the distance between the nth data point in set X and the kth data point in set C, using the Euclidean distance calculation formula: ;
[0011] where c k1, c k2 , c k3 represents the 3D vector of each type of data point in set C. The calculated distances are represented in a two-dimensional array established with the data points in set X arranged vertically and the data points in set C arranged horizontally. The distances are the data in the array. Compare the sizes of the data in each row horizontally, take the data point corresponding to the minimum value, and perform vertical classification to obtain a set X ’ , and calculate the mean of the 3D vectors of the data points in k sets X ’ . Take the mean as the data points in the new set C ’ . Calculate the distances between the data points in set X and the data points in set C ’ again until the means of the 3D vectors of the data points in the k sets no longer change, then stop the calculation to obtain the set C of mean data points ” , and k sets X ” established with the data points in set C as representatives ” ; S402. Calculate the distances between the data points in k sets X ” and the corresponding data points in set C ” . Sort them in ascending order to obtain k distance sets D = {d1, d2,..., d m}, where d m represents the distance corresponding to the mth data in set X ” . Set the distance threshold using quartiles. Denote the distance value of the first quartile in D as Q1 and the distance value of the third quartile as Q3. Calculate the interquartile range IQR: IQR = Q3 - Q1. Set the upper threshold as Ub = Q3 + 1.5×IQR and the lower threshold as Lb = Q3 - 1.5×IQR. For a data point x m , when its corresponding , determine that the data point x m corresponding to d ” in set X m is an outlier. When d m ∈[Lb, Ub], the data point x m is a normal data point; Many traditional methods rely on manually set thresholds and rules, such as setting fixed upper or lower limits according to the amount size, transaction frequency, etc. These rules are not flexible enough for data changes and need to be adjusted manually regularly. Through automated clustering algorithms, transaction patterns can be automatically identified based on the data distribution without manually setting rules. The clustering results are based on the natural structure of the data and are not affected by artificially set rules, making them more adaptable; By calculating the distance between data points and cluster centers, clustering analysis can accurately identify which data points deviate from the normal pattern, thus improving the accuracy of anomaly detection. By dynamically adjusting thresholds and distance calculations, more complex abnormal behaviors can be identified.
[0012] S500. Based on the categories divided by clustering analysis, determine whether the real-time transaction data is an abnormal transaction; Further, the specific steps for determining the real-time transaction data are as follows: S501. For the real-time transaction data, perform feature extraction and standardization processing on the real-time transaction data to obtain real-time standardized feature data. Based on step S401, divide the real-time data points into the kth set X'', calculate the distance between the real-time data points and the data points in the corresponding set C'' of the set X'', and determine whether it is an abnormal transaction through a threshold.
[0013] S600. Visualize all the real-time transaction data to generate a visualization chart, and mark and display the abnormal transactions in the visualization chart; Further, the specific steps for visualization are as follows: S601. Use Matplotlib technology to visualize all the real-time transaction data to generate a visualization chart; S602. Mark and display the abnormal transactions in the visualization chart; Through the visualization of real-time transaction data, abnormal transactions can be discovered in a timely manner and marked, which is particularly important for discovering abnormal data and can help detect abnormal behaviors.
[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. It can perform real-time analysis on each transaction and make abnormal judgments at the moment of transaction occurrence. This feature enables the rapid identification of potential abnormal transactions and reduces the response delay in traditional risk control technologies; 2. Once a transaction is determined to be abnormal, an alarm is automatically triggered and the abnormal transaction is marked. Through automated processing, the need for manual intervention is greatly reduced, and the manual review burden is reduced.
[0015] 3. By comprehensively analyzing the features of multiple dimensions such as transaction amount, transaction frequency, and amount volatility, more comprehensive anomaly detection can be provided; 4. By adopting a machine learning-based model, it can be automatically adjusted and optimized through training on historical data, has high adaptability when facing new risks, and reduces false alarms and missed alarms. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic flowchart of a method for visual-based target information mining according to the present invention; Figure 2 It is a schematic diagram of an embodiment of a method for visual-based target information mining according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Embodiment: As Figure 1 - Figure 2 shown, the present invention provides a technical method, a method for visual-based target information mining, and the method includes the following steps: A method for visual-based target information mining, the method includes the following steps: S100. Collect historical transaction records of a number of accounts within a certain time range, and use the transaction data collected from the historical transaction records as the original data; Furthermore, the step of collecting the original data is; S101. Extract historical transaction records within a time period t with a period of T. The historical transaction records include transaction numerical data, transaction time data, and transaction frequency data, and use the extracted data as the original data.
[0018] S200. According to the extracted original data, establish a big data model to extract data features from the original data to obtain feature data; Furthermore, the specific steps for extracting features from the original data are: S201. Extract the specific time when several transactions of a single account H occur, the value of a single transaction from the original data within the time period t, divide the time period t into p transaction time periods on average, and the time length of each transaction time period is L p , extract the number of transactions y that occur in the z-th transaction time period, and calculate the frequency P of transactions that occur in the z-th time period z , , where z ∈ [1, p], and z is a positive integer; Aggregate the frequencies of each trading period to obtain the frequency change set P = {P1, P2, P3..., P p} for account H in period t, where P1, P2, P3... P p are the frequencies of the 1st, 2nd, 3rd,..., and pth trading periods respectively; S202. Calculate the average value S of the single - transaction values in the zth trading period by extracting all single - transaction values of each trading period within period t; Aggregate the S values of each trading period to obtain the set S = {S1, S2, S3..., S p} of the average single - transaction values for account H in period t, where S1, S2, S3... S p are the average single - transaction values of the 1st, 2nd, 3rd,..., and pth trading periods respectively; S203. For each trading period, calculate the change rate B between the nth and (n - 1)th single - transaction values ’ , , calculate the average value of the single - transaction value change rates in the zth trading period, and use the average value of the single - transaction value change rates in the zth trading period as the transaction value change rate B for the zth trading period, to obtain the transaction value change rates for p trading periods of account H in period t; Aggregate the transaction value change rates B of each trading period to obtain the set B = {B1, B2, B3..., B p} of the average single - transaction values for account H in period t, where B1, B2, B3... B p are the average single - transaction values of the 1st, 2nd, 3rd,..., and pth trading periods respectively; S204. Integrate the sets P, B, and S to obtain a feature data set A = {S1:P1;B1, S2:P2:B2,..., S P :P P :B p} based on account H, where S1:P1:B1, S2:P2;B2,..., S P :P P :B p represent the 1st, 2nd,..., and pth elements in the feature data set A, where S P , P P , B p serve as three types of feature data of account H, and S P :P P :B pExtract the feature data of m accounts to obtain a set M = {A1, A2,..., Am}, where A1, A2, A3,..., and Am represent the feature data sets of the 1st, 2nd, 3rd,..., and mth accounts respectively; Through the big data model, it can be automatically cleaned and standardized without relying on manual rule setting. Therefore, it can process larger-scale and more complex data sets to extract meaningful features.
[0019] S300. By standardizing the feature data, eliminate the difference in data dimensions between different feature data to obtain standardized data; Furthermore, the specific steps to quantify the feature data to obtain standardized data are as follows: S301. Standardize each type of feature data within each feature data set in set M. The formula is , where μ is the average value of each type of feature data, σ is the standard deviation of each type of feature data. Through the standardization calculation formula, Z represents the result after standardizing each type of feature data. The standardized results are pooled to obtain a standardized data set M ’ ; For the feature data set A = {S1: P1; B1, S2: P2: B2,..., S P : P P : B p} established for account H, Extract all S1 - S P to obtain the mean value W1 and standard deviation W2 of S1 - S p . Calculate the standardized result of the pth one, Sp’ = (Sp - W1) / W2, to obtain a standardized data sequence S 1’ - S p’ . Similarly, obtain P 1’ - P p’ , B 1’ - B p’ . Pool them to obtain a set A’ = {S 1’ : P 1’ ; B 1’ , S 2’ : P 2’ : B 2’ ,..., S P’ : P P’ : B p’}. Extract the standardized feature data of m accounts to obtain a set M’ = {A1’, A2’,..., Am’} S400. Conduct a clustering analysis on the standardized data and classify the standardized data according to the analysis results; Further, the specific steps for classifying the standardized data according to the analysis results and setting the anomaly threshold are as follows: S401. Set each element in the standardized feature data set A as a data point x. Each data point has three types of feature data. Each data point x is a three-dimensional vector, denoted as x = (x i1 , x i2 , x i3 ), to obtain a data point set X = {x1, x2,..., x n}. Randomly select k data points to form a new set C = {c1, c2,..., c k}. Calculate the distances from the n data points to each data point in set C, denoted as d = {x n, c1, x n, c2,... x n, c k}, x n, c k represents the distance between the nth data point in set X and the kth data point in set C. Use the Euclidean distance calculation formula: ;
[0020] where c k1, c k2 , c k3 represents the three-dimensional vectors of each type of data point in set C. Represent the calculated distances in a two-dimensional array established with the data points in set X arranged vertically and the data points in set C arranged horizontally. The distances are the data in the array. Compare the magnitudes of the data in each row horizontally, take the data point corresponding to the minimum value, and perform vertical classification to obtain a set X ’ . Calculate the means of the three-dimensional vectors of the data points in the k sets X ’ . Use the means as the data points in the new set C ’ . Calculate the distances from the data points in set X to the data points in set C ’ again until the means of the three-dimensional vectors of the data points in the k sets no longer change, then stop the calculation to obtain the mean data point set C ” , and the k sets X ” established with the data points in set C ” as representatives; S402. Calculate the distances between the data points in the k sets X ” and the corresponding data points in set C ” . Sort them in ascending order to obtain k distance sets D = {d1, d2,..., d m}, d m represents the distance between set X ”The distance corresponding to the m-th data in [description], use quartiles to set the distance threshold. Denote the distance value of the first quartile in D as Q1, and the distance value of the third quartile as Q3. Calculate the interquartile range IQR: IQR = Q3 - Q1. Set the upper threshold as Ub = Q3 + 1.5×IQR, and the lower threshold as Lb = Q3 + 1.5×IQR. For a data point x m , when its corresponding is [condition], determine d m corresponding set X ” The data point x m in is an outlier. When d m ∈ [Lb, Ub], the data point x m is a normal data point; Many traditional methods rely on manually set thresholds and rules. For example, set fixed upper or lower limits according to the amount size, transaction frequency, etc. These rules are not flexible enough for data changes and need to be adjusted manually regularly. Through an automated clustering algorithm, transaction patterns can be automatically identified based on the data distribution without manually setting rules. The clustering results are based on the natural structure of the data and are not affected by artificially set rules, making them more adaptable; By calculating the distance between data points and the cluster center, clustering analysis can accurately identify which data points deviate from the normal pattern, thereby improving the accuracy of anomaly detection. By dynamically adjusting the threshold and distance calculation, more complex abnormal behaviors can be identified.
[0021] S500. Based on the categories divided by clustering analysis, determine whether real-time transaction data is an abnormal transaction; Further, the specific steps for determining real-time transaction data are as follows: S501. For real-time transaction data, perform feature extraction and standardization processing on the real-time transaction data to obtain real-time standardized feature data. Based on step S401, divide the real-time data points into the k-th set X”. Calculate the distance between the real-time data points and the data points in the corresponding set C” of the set X”. Determine whether it is an abnormal transaction through the threshold.
[0022] S600. Visualize all real-time transaction data to generate a visualization chart, and mark and display the abnormal transactions in the visualization chart; Further, the specific steps for visualization are as follows: S601. Use Matplotlib technology to visualize all real-time transaction data to generate a visualization chart; S602. Mark and display the abnormal transactions in the visualization chart; Visualizing real-time transaction data enables the timely detection of abnormal transactions and their marking, which is particularly important for discovering abnormal data and can help detect abnormal behaviors.
[0023] In an embodiment, an object information mining method based on visualization is also proposed, and an object information mining system based on visualization includes a data extraction module, a feature extraction module, a feature processing module, an algorithm module, a judgment and warning module, and a display module. The data processing module is used to collect transaction data of historical transaction records of a number of accounts within a certain time range. The feature extraction module is used to extract features from the extracted data using a big data model to obtain feature data. The feature processing module is used to perform standardization processing on the feature data to eliminate the difference in data dimensions between different feature data and obtain standardized data. The algorithm module is used to use a clustering algorithm for the standardized data and classify the standardized data based on the calculated results. The judgment and warning module is used to compare the real-time transaction data with a threshold value to judge the transaction data. The display module displays abnormal transactions in a visualization chart. The algorithm module includes a distance calculation unit, a threshold dynamic calculation unit, and an abnormality determination unit. The distance calculation unit is used to calculate the distance between standardized data using an algorithm. The threshold dynamic calculation unit is used to set an abnormal threshold based on historical transaction data using a clustering algorithm. The abnormality determination unit is used to compare the real-time data with the abnormal threshold to judge the real-time transaction data. The judgment and warning module includes an alarm generation unit and a blocking linkage unit. The alarm generation unit is used to generate a standardized alarm event combination by combining the determined abnormal transactions with the original data and abnormal features. The blocking linkage unit is used to call the business system interface to suspend the transaction when an abnormal transaction is determined.
[0024] Embodiment: Now analyze the transaction data of users in a certain financial institution. First, for accounts A1 and A2, n transaction data within the transaction time period are respectively extracted and standardized to obtain feature data: A11, A12, A13, A21, A22, A23. Each piece of data takes 2 features, denoted as 1A1, 2A1, 1A2, 2A2. The corresponding relationship is A1-(A11-(1A1, 2A1), A12-(1A1, 2A1), A13-(1A1, 2A1)). A2-(A21-(1A2,2A2), A22-(1A2,2A2), A23-(1A2,2A2)), with 6 pieces of data as 6 data points. Randomly select two data points A11 and A21 as the center points, and calculate the distances from A11, A12, A13, A21, A22, and A23 to the center points respectively. Arrange the data points vertically and the center points horizontally to establish a two-dimensional array, with the distances as the data in the array. Compare the sizes of the data in each row horizontally, take the data point corresponding to the minimum value, and classify them vertically. Suppose the obtained classifications are L1-(A11, A13, A21) and L2-(A12, A22, A23). For example, if the data corresponding to A12 is (d1, d2) and the data corresponding to A23 is (d3, d4), and d1 > d2, so A12 is classified into L2 class; d3 > d4, so A23 is classified into L2 class; Calculate the characteristic means L1' and L2' of the data points in L1 and L2 respectively. Take L1' and L2' as the new center points and calculate the distances from the data points to the new center points. For example, obtain the new classifications L1-(A11, A13, A21, A12) and L2-(A22, A23). Calculate the characteristic means L1'' and L2'' of the data points again, calculate the distances from the data points to L1'' and L2'', and obtain the new classifications L11-(A11, A13, A21, A12) and L22-(A22, A23). Since the characteristic means of the data points in L11 and L22 classes no longer change, stop the classification. The result will obtain two final center points L1'' and L2'', and two types of data points divided based on L1'' and L2''; Calculate the distances between L1'' and L2'' and the data points in their corresponding classes. Set the upper and lower limits of the threshold according to the quartiles. When a new data A31 appears, classify it, calculate the distance to the class center point. When the distance is not within the range between the upper and lower limits of the threshold, it is recorded as an anomaly.
[0025] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for mining target information based on visualization, characterized in that, The method includes: S100. Collect the historical transaction records of a number of accounts within a certain time range, and use the transaction data collected from the historical transaction records as the original data; S200. According to the extracted original data, establish a big data model to extract data features from the original data and obtain feature data; S300. By performing standardization processing on the feature data, eliminate the data dimension differences between different feature data to obtain standardized data; S400. Perform clustering analysis on the standardized data and classify the standardized data according to the analysis results; S500. Based on the categories divided by the clustering analysis, judge whether the real-time transaction data is an abnormal transaction; S600. Visualize all the real-time transaction data to generate a visualization chart, and mark and display the abnormal transactions in the visualization chart.
2. The method for mining target information based on visualization according to claim 1, characterized in that: The specific steps for collecting the original data in S100 are: S101. Extract the historical transaction records within the time period t with a period of T. The historical transaction records include transaction value data, transaction time data, and transaction frequency data, and use the extracted data as the original data.
3. The method for mining target information based on visualization according to claim 2, characterized in that The specific steps for extracting features in S200 are: S201. Extract the specific times when several transactions of a single account H occurred, the value of each single transaction, from the original data within the time period t, and evenly divide the time period t into p trading time periods, with the time length of each trading time period being L p , extract the number of transactions y that occurred in the z-th trading time period, and calculate the transaction occurrence frequency P in the z-th time period z , , where z ∈ [1, p], and z is a positive integer; Aggregate the frequencies of each trading period to obtain the frequency change set P = {P1, P2, P3..., P p} for account H in period t, where P1, P2, P3... P p are the frequencies of the 1st, 2nd, 3rd,..., and pth trading periods respectively; S202. By extracting all single-transaction values in each transaction time period within the time period t, calculate the average value S of the single-transaction values in the z-th transaction time period; Aggregate the S values for each trading period to obtain the set of single - transaction value means S = {S1, S2, S3..., S p} during time period t for account H, where S1, S2, S3... S p are the single - transaction value means for the 1st, 2nd, 3rd,..., and pth trading periods respectively; S203. For each trading time period, calculate the change rate B of the nth single transaction value and the (n - 1)th single transaction value ’ , , calculate the average value of the change rates of single transaction values within the zth trading time period, use the average value of the change rates of single transaction values within the zth trading time period as the change rate B of the trading value in the zth trading time period, and obtain the change rates of trading values of p trading time periods of account H within the time period t Aggregate the transaction value change rate B for each trading period to obtain the set B of the single - transaction value means of account H in period t, where B = {B1, B2, B3..., B p}, where B1, B2, B3... B p are the single - transaction value means of the 1st, 2nd, 3rd,..., and pth trading periods respectively; S204. Integrate sets P, B, and S to obtain a feature dataset A = {S1:P1;B1, S2:P2:B2,..., S P :P P :B p}, where S1:P1:B1, S2:P2:B2,..., S P :P P :B p represent the 1st, 2nd,..., and pth elements in the feature dataset A, where S P , P P , and B p are three types of feature data of the account H. Extract the feature data of m accounts to obtain a set M = {A1, A2,..., Am}, where A1, A2, A3,..., and Am represent the feature datasets of the 1st, 2nd, 3rd,..., and mth accounts respectively.
4. The method for mining target information based on visualization according to claim 3, characterized in that: The specific steps for performing standardization processing on the feature data in S300 to eliminate the data dimension differences between different feature data and obtain standardized data are: S301. Standardize each type of feature data in each feature data set in set M. The formula is , where μ is the mean of each type of feature data, σ is the standard deviation of each type of feature data. Through the standardization calculation formula, Z represents the result after standardizing each type of feature data. The standardized results are pooled to obtain a standardized data set M ’ .
5. The method for mining target information based on visualization according to claim 4, characterized in that: The specific steps for performing clustering analysis on the standardized data and classifying in S400 are: S401. Set each element in the standardized feature data set A as a data point x. Each data point has three types of feature data. Each data point x is a three-dimensional vector, denoted as x = (x i1 , x i2 , x i3 ). Obtain a data point set X = {x1, x2,..., x n}. Randomly select k data points to form a new set C = {c1, c2,..., c k}. Calculate the distances from n data points to each data point in set C, denoted as d = {x n, c1, x n, c2,... x n, c k}, x n, c k represents the distance between the nth data point in set X and the kth data point in set C. Use the Euclidean distance calculation formula: ; where c k1, c k2 , c k3 represents the 3D vectors of each data point in set C. The calculated distances are represented in a two-dimensional array with the data points in set X arranged vertically and the data points in set C arranged horizontally. The distances are the data in the array. Compare the magnitudes of the data in each row horizontally, take the data point corresponding to the minimum value, and classify them vertically to obtain a set X ’ , calculate k sets of X ’ Calculate the mean of the 3D vectors of the data points in X, and use the mean as the new set C ’ as the data points in C ’ Calculate the distances between the data points in set X and the data points in set C again until the means of the 3D vectors of the data points in the k sets no longer change, then stop the calculation to obtain the set C of mean data points ” , and k sets of X ” established with the data points in set C as representatives ” ; S402. Calculate the distances between the data points in k sets X ” and the corresponding data points in set C ” . Obtain k distance sets D = {d1, d2,..., d m} through forward sorting. Here, d m represents the distance corresponding to the m-th data in set X ” . Set the distance threshold using quartiles. Denote the distance value of the first quartile in D as Q1 and the distance value of the third quartile as Q3. Calculate the interquartile range IQR: IQR = Q3 - Q1. Set the upper threshold as Ub = Q3 + 1.5×IQR and the lower threshold as Lb = Q3 - 1.5×IQR. For a data point x m , when its corresponding , determine that the data point x m corresponding to d ” in set X m is an outlier. When d m ∈[Lb, Ub], the data point x m is a normal data point.
6. The method for mining target information based on visualization according to claim 5, characterized in that: In step S500, the steps for judging abnormal transactions for the real-time transaction data are: S501. For the real-time transaction data, perform feature extraction and standardization processing on the real-time transaction data to obtain real-time standardized feature data. Based on step S401, divide the real-time data points into the k-th set X'', calculate the distance between the real-time data points and the data points in the corresponding set C'' of the set X'', and determine whether it is an abnormal transaction through a threshold.
7. The method for mining target information based on visualization according to claim 6, characterized in that: The specific steps for marking and displaying the abnormal transactions in the visualization chart in step S600 include: S601. Use Matplotlib technology to visualize all the real-time transaction data to generate a visualization chart; S602. Mark and display the abnormal transactions in the visualization chart.
Citation Information
Patent Citations
Business abnormal data analysis method based on financial management system
CN119312177A
Method and system for real-time, false positive resistant, load independent and self-learning anomaly detection of measured transaction execution parameters like response times
US20150032752A1